Skip to main content
ARS Home » Midwest Area » Ames, Iowa » National Animal Disease Center » Virus and Prion Research » Research » Research Project #449616

Research Project: Scalable algorithms for epidemiology, reassortment risk modeling, and proactive influenza A virus vaccine design

Location: Virus and Prion Research

Project Number: 5030-32000-231-128-A
Project Type: Cooperative Agreement

Start Date: Aug 1, 2026
End Date: Jul 31, 2027

Objective:
To develop computational algorithms and improve data analysis and reporting on the genetic diversity and reassortment patterns of influenza A virus, including highly pathogenic avian influenza A viruses. Algorithms will be developed that can: 1) objectively select the most representative viruses for inclusion in vaccines, and evaluate the evolutionary breadth of viruses remaining uncovered by current formulations; 2) develop a pathogen-agnostic algorithm for genomic surveillance data that accurately identifies virus identity, estimates relative abundances of viruses in samples, and performs precise read assignments; and 3) develop algorithms that can rapidly assess reassortment of influenza A viruses, including HPAI and endemic influenza A viruses in swine, and establish a software that can ingest and analyze data in near-real time to identify genetically novel influenza A virus.

Approach:
1) Vaccines are essential for controlling influenza viruses in wildlife and agricultural animals, yet rapid viral evolution and the diversity of circulating strains make it challenging to ensure effective vaccine protection. To address this need, an existing algorithm (PARNAS) will be revised by formulating a novel optimization problem that incorporates phylogenetic diversity (PD) to evaluate the evolutionary breadth of lineages remaining uncovered by current vaccines. PD serves as a scalable proxy for antigenic novelty, identifying lineages that represent significant evolutionary "blind spots" in the current vaccine portfolio. Implementing this PD-aware framework for genome-scale phylogenies of tens of thousands of taxa presents substantial computational challenges. For this objective, efficient algorithms will be designed to handle genome-scale datasets and rank unprotected viruses based on the magnitude of their genetic novelty. This work will deliver a production-grade software tool that will objectively design vaccines and identify when current vaccines do not adequately cover virus diversity detected in genomic surveillance. 2) The circulation of influenza A viruses (IAVs) in wildlife and livestock presents a significant threat due to their impact on agricultural animals. Accurate classification of viral subtypes and characterization of within-host diversity are crucial for risk assessment and vaccine development. Although metagenomic sequencing facilitates early detection of some animal pathogens, current pipelines often discard critical information and do not consider sequencing quality in analysis. A novel probabilistic alignment-based framework for high-resolution viral genome identification will be developed. This approach will integrate advanced string data structures for efficient alignment with a quality-score-aware Expectation-Maximization algorithm. The approach will accurately identify source strains, estimate relative abundances of different viruses in samples, and perform precise read assignment and genome assembly. This objective will result in a pathogen-agnostic diagnostic tool that can accurately identify and automatically flag novel reassorted viruses, thereby improving the detection of emerging pathogens that may impact animal agriculture. 3) A significant volume of genomic surveillance data are generated by USDA animal surveillance efforts, and these data have revealed the fundamental importance of reassortment on the emergence of novel variants. Reassortment across the segmented influenza A genome generates transient, host-restricted genotypes that obscure the distinction between extinction, lineages with sustained transmission, and lineages that have interspecies transmission potential. This project will implement distinct domain-specific Large Language Models (LLM) to automate identification of reassorted genotypes that have epidemiological relevance. The models will provide a discrete risk score for viruses detected in genomic surveillance that captures the likelihood of persistence, dissemination, and risk for interspecies transmission.