I installed the pipeline following the Installation wiki page, but with the addition of the taxonomy-switching step provided in issue #47 to avoid the "Invalid taxonomic rank: domain" error. There were no errors when I ran the pipeline on the test flavivirus dataset with the default RVDB database, and the output .tsv file exactly matches the provided example except for the presence of NA in the top_viral_lineage and top_viral_coverage columns. Looking at the source code, I'm not sure what's causing this issue.
In an attempt to solve the issue, I built an environment with the latest version of diamond (and re-installed the databases), thinking it might be an incompatibility issue between the packages and databases, but the flavivirus example returned the same output, complete with NAs. Any assistance in fixing this issue would be very appreciated! I know I could pull the protein ID from top_viral_desc and the find the virus taxonomy, but the taxonomy provided in top_viral_lineage and top_viral_coverage doesn't seem to exactly match up to top_viral_desc in the flavivirus-specific database example.
Here's what I got as an output for the flavivirus test example:
eve_id confidence eve_score suggests because locus eve_length percent_contig top_evalue top_pident top_desc top_viral_desc top_viral_lineage top_viral_coverage max_count_phylum
ATLV01_cut_EVE001 high 92 viral (18), maybe-viral (2) VDB (16), glycoprotein protein (2), virus (2) ATLV01019207.1_1147-1404:- 258 3.5 1.71e-17 48.6 putative glycoprotein [Gambie virus] acc=AOR51379.1 putative glycoprotein [Gambie virus] acc=AOR51379.1 NA NA Negarnaviricota
ATLV01_cut_EVE002 high 91 viral (18), maybe-viral (2) VDB (15), glycoprotein protein (2), virus (2), viral (1) ATLV01019207.1_2615-3409:- 795 10.77 5.509999999999999e-112 60.5 putative glycoprotein [Anopheles darlingi virus] acc=QBK47202.1 putative glycoprotein [Anopheles darlingi virus] acc=QBK47202.1 NA NA Negarnaviricota
ATLV01_cut_EVE003 high 85 viral (20), maybe-viral (4) VDB (20), polyprotein protein (3), uncharacterized protein (1) ATLV01019207.1_3580-4280:- 701 9.5 9.180000000000001e-73 54.7 polyprotein [Karumba virus] acc=YP_009388577.1 polyprotein [Karumba virus] acc=YP_009388577.1 NA NA Kitrinoviricota
I installed the pipeline following the Installation wiki page, but with the addition of the taxonomy-switching step provided in issue #47 to avoid the "Invalid taxonomic rank: domain" error. There were no errors when I ran the pipeline on the test flavivirus dataset with the default RVDB database, and the output .tsv file exactly matches the provided example except for the presence of NA in the
top_viral_lineageandtop_viral_coveragecolumns. Looking at the source code, I'm not sure what's causing this issue.In an attempt to solve the issue, I built an environment with the latest version of diamond (and re-installed the databases), thinking it might be an incompatibility issue between the packages and databases, but the flavivirus example returned the same output, complete with NAs. Any assistance in fixing this issue would be very appreciated! I know I could pull the protein ID from top_viral_desc and the find the virus taxonomy, but the taxonomy provided in top_viral_lineage and top_viral_coverage doesn't seem to exactly match up to top_viral_desc in the flavivirus-specific database example.
Here's what I got as an output for the flavivirus test example: