HostFinder: Large-scale statistical detection of virus-host interactions across the Sequence Read Archive
| dc.contributor.author | Mascarenhas, Shyan | |
| dc.date.accessioned | 2026-09-16T15:44:21Z | |
| dc.date.issued | 2026-09-16 | |
| dc.date.submitted | 2026-09-10 | |
| dc.description.abstract | The Sequence Read Archive (SRA) is the largest public sequencing repository and represents an unparalleled record of global biodiversity. Despite the success of prior studies in discovering novel viruses through the SRA, the development of reliable methods for inferring virus–host interactions remain an important unmet need. Here, we present HostFinder, a computational framework that infers virus-host associations at scale by analysing co-occurrence patterns across the entire SRA. HostFinder generates taxonomic profiles using STAT (SRA Taxonomic Analysis Tool), quantifies host and viral abundances per dataset, and calculates association strength through k-mer thresholds and log2-fold enrichment scores. Using a curated virus-eukaryote interaction dataset as well as VirusHost DB, we evaluated whether co-occurrence/enrichment scores could distinguish validated biological associations from false ones. Based on the enrichment score alone, the method achieved 0.93 ROC-AUC on the smaller 50 curated virus-host interaction dataset, whereas on the larger 8,508 Virus-HostDB dataset it achieved a 0.73 ROC-AUC with a 0.79 AUPRC. Scaling across the entire SRA, we analysed 3.78 billion virus-host pairs for potential interactions. Our analysis identified 7.8 M high-confidence interactions (log2-fold enrichment score > 3, FDR < 1%) involving 53,025 viral species and 15,143 host species. As a case study, we discovered novel Partitiviridae associations with multiple insect hosts that are unreported in existing literature. These predictions were independently validated through viral genome assembly and phylogenetic analysis. HostFinder demonstrates that large-scale co-occurrence analysis of public sequencing repositories can reveal the hidden structure of virus-host networks and accelerate discovery of novel viral associations. | |
| dc.identifier.uri | https://hdl.handle.net/10012/24301 | |
| dc.language.iso | en | |
| dc.pending | false | |
| dc.publisher | University of Waterloo | en |
| dc.relation.uri | https://github.com/s2mascar/HostFinder | |
| dc.subject | virus-host interactions | |
| dc.subject | taxonomic profiling | |
| dc.subject | co-occurrence analysis | |
| dc.subject | Sequence Read Archive | |
| dc.subject | bioinformatics | |
| dc.title | HostFinder: Large-scale statistical detection of virus-host interactions across the Sequence Read Archive | |
| dc.type | Master Thesis | |
| uws-etd.degree | Master of Science | |
| uws-etd.degree.department | Biology | |
| uws-etd.degree.discipline | Biology | |
| uws-etd.degree.grantor | University of Waterloo | en |
| uws-etd.embargo.terms | 1 year | |
| uws.contributor.advisor | Doxey, Andrew | |
| uws.contributor.affiliation1 | Faculty of Science | |
| uws.peerReviewStatus | Unreviewed | en |
| uws.published.city | Waterloo | en |
| uws.published.country | Canada | en |
| uws.published.province | Ontario | en |
| uws.scholarLevel | Graduate | en |
| uws.typeOfResource | Text | en |