Location: Virus and Prion Research
Project Number: 5030-32000-231-130-S
Project Type: Non-Assistance Cooperative Agreement
Start Date: Aug 1, 2026
End Date: Jul 31, 2027
Objective:
Objective: A central challenge in data science is the integration of complex data from public sources such as NCBI GenBank together with private genetic datasets, locally inferred secondary data (such as genetic clade classifications and phenotypic information), and other metadata that can be scraped from the internet. Additionally, access to comprehensive information for analyses requires merging and filtering data while unifying equivalent data, removing duplicates, and resolving inconsistencies.
Approach:
Phase 1: Parser combinators for data extraction with context-aware, ontology-driven type inference: We propose a solution that leverages three key techniques: parser combinators, ontology-driven grammatical linking, and statistical disambiguation. Parsers can be composed for complex tasks like parsing data from a phylogenetic tree. They can extract rich data as well: the octofludb IAV strain name parser extracts the strain name, year, optional US state, barcode, country and host from a strain. The ontologies can serve as rules for inferring the grammar of rows in tables (i.e., how columns are related to each other). For data with free forms (for example, terms in a FASTA header) a Markov model will be investigated for inferring the type of each term and then ontologies can be used for each path to infer grammatical relationships.
Phase 2: Enrich data analysis through ontology and rule-based inferencing: After knowledge is extracted from raw data and uploaded to a graph database, the next stage of knowledge engineering can begin. We employ two methods: ontology engineering, where relationships between entities are specified and used to infer new links; and logic programing where rules are specified that trigger the generation of new knowledge. For example, uploading the single statement that France is in Europe would allow any query for objects in Europe to also return objects in France. Similarly, an IAV strain may be flagged as "mixed" if it is linked to two different subtypes, e.g., H1N1 and H1N2. This declarative approach will allow much of the tedious, error prone data munging to be replaced with concise logical statements.
The outcome of these two linked phases will be techniques to allow inputs of raw data (e.g., tabular, GenBank, FASTA) and optional metadata into a graph database for genomic epidemiology. Using the domain ontology, relations will be inferred between terms in a "sentence." Parser combinators are used to infer data types and extracts nested information.The rule engine allows relationships between data to be expressed, anomalies to be flagged, and new data to be inferred. Overall, the goal of this agreement is to develop automated techniques for database management that facilitate visualization of the spatial and temporal trends in genetic diversity of influenza A viruses in swine. This work will provide critical information on when and where influenza A viruses are spreading and provide an early warning system for novel viruses. The early detection of novel pathogens, alongside other endemic viruses, will enhance US farmers ability to control infection, improve animal health, and enhance economic productivity.