Description of feature
Hi!
Thanks so much for this amazing pipeline, love it! It works seamlessly so far and everything is really clean and tidy.
I have been processing different data types with Immcantation suite for a few years now and I'm used to filtering out sequences with very low umi counts / consensus counts. I guess it's sometimes hard to distinguish sequencing errors vs real SHM in rare clonotypes, but I prefer to err on the side of conservatism and remove sequences that don't have enough reads to back them up.
I believe the pipeline already removes sequences that have < 2 representative sequences (the documentation says it's in "presto/10-splitseq" but I think for bulk it's in the trust4 folder? ). However, I'm referring to a later step, once the final clonotypes are assembled, before they are piped into the clonal analysis. You already apply quite a few filters based on sequence quality, productive or not, etc, but I was wondering whether it would be easy enough for you to add an optional "consensus_count" and/or "umi_count" filter (you could set the default to 0).
I've implemented this in a local copy of
/researchers/laura.twomey/Tools/omics_tools/nf-core/nf-core-airrflow/5_1_0/bin/reveal_mod_3_junction.R
(attached), but it would be probably cleaner if this were a separate module.
Just a suggestion! If this is not something you recommend I'd also would like to know your opinion on it - do you trust sequences with only 1 count // 1 umi ?
Thanks so much!
Laura
reveal_mod_3_junction.txt
Description of feature
Hi!
Thanks so much for this amazing pipeline, love it! It works seamlessly so far and everything is really clean and tidy.
I have been processing different data types with Immcantation suite for a few years now and I'm used to filtering out sequences with very low umi counts / consensus counts. I guess it's sometimes hard to distinguish sequencing errors vs real SHM in rare clonotypes, but I prefer to err on the side of conservatism and remove sequences that don't have enough reads to back them up.
I believe the pipeline already removes sequences that have < 2 representative sequences (the documentation says it's in "presto/10-splitseq" but I think for bulk it's in the trust4 folder? ). However, I'm referring to a later step, once the final clonotypes are assembled, before they are piped into the clonal analysis. You already apply quite a few filters based on sequence quality, productive or not, etc, but I was wondering whether it would be easy enough for you to add an optional "consensus_count" and/or "umi_count" filter (you could set the default to 0).
I've implemented this in a local copy of
/researchers/laura.twomey/Tools/omics_tools/nf-core/nf-core-airrflow/5_1_0/bin/reveal_mod_3_junction.R
(attached), but it would be probably cleaner if this were a separate module.
Just a suggestion! If this is not something you recommend I'd also would like to know your opinion on it - do you trust sequences with only 1 count // 1 umi ?
Thanks so much!
Laura
reveal_mod_3_junction.txt