HostFinder: Large-scale statistical detection of virus-host interactions across the Sequence Read Archive

dc.contributor.authorMascarenhas, Shyan
dc.date.accessioned2026-09-16T15:44:21Z
dc.date.issued2026-09-16
dc.date.submitted2026-09-10
dc.description.abstractThe Sequence Read Archive (SRA) is the largest public sequencing repository and represents an unparalleled record of global biodiversity. Despite the success of prior studies in discovering novel viruses through the SRA, the development of reliable methods for inferring virus–host interactions remain an important unmet need. Here, we present HostFinder, a computational framework that infers virus-host associations at scale by analysing co-occurrence patterns across the entire SRA. HostFinder generates taxonomic profiles using STAT (SRA Taxonomic Analysis Tool), quantifies host and viral abundances per dataset, and calculates association strength through k-mer thresholds and log2-fold enrichment scores. Using a curated virus-eukaryote interaction dataset as well as VirusHost DB, we evaluated whether co-occurrence/enrichment scores could distinguish validated biological associations from false ones. Based on the enrichment score alone, the method achieved 0.93 ROC-AUC on the smaller 50 curated virus-host interaction dataset, whereas on the larger 8,508 Virus-HostDB dataset it achieved a 0.73 ROC-AUC with a 0.79 AUPRC. Scaling across the entire SRA, we analysed 3.78 billion virus-host pairs for potential interactions. Our analysis identified 7.8 M high-confidence interactions (log2-fold enrichment score > 3, FDR < 1%) involving 53,025 viral species and 15,143 host species. As a case study, we discovered novel Partitiviridae associations with multiple insect hosts that are unreported in existing literature. These predictions were independently validated through viral genome assembly and phylogenetic analysis. HostFinder demonstrates that large-scale co-occurrence analysis of public sequencing repositories can reveal the hidden structure of virus-host networks and accelerate discovery of novel viral associations.
dc.identifier.urihttps://hdl.handle.net/10012/24301
dc.language.isoen
dc.pendingfalse
dc.publisherUniversity of Waterlooen
dc.relation.urihttps://github.com/s2mascar/HostFinder
dc.subjectvirus-host interactions
dc.subjecttaxonomic profiling
dc.subjectco-occurrence analysis
dc.subjectSequence Read Archive
dc.subjectbioinformatics
dc.titleHostFinder: Large-scale statistical detection of virus-host interactions across the Sequence Read Archive
dc.typeMaster Thesis
uws-etd.degreeMaster of Science
uws-etd.degree.departmentBiology
uws-etd.degree.disciplineBiology
uws-etd.degree.grantorUniversity of Waterlooen
uws-etd.embargo.terms1 year
uws.contributor.advisorDoxey, Andrew
uws.contributor.affiliation1Faculty of Science
uws.peerReviewStatusUnrevieweden
uws.published.cityWaterlooen
uws.published.countryCanadaen
uws.published.provinceOntarioen
uws.scholarLevelGraduateen
uws.typeOfResourceTexten

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Mascarenhas_Shyan.pdf
Size:
5.04 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
6.4 KB
Format:
Item-specific license agreed upon to submission
Description:

Collections