Skip to main content
ARS Home » Northeast Area » Ithaca, New York » Robert W. Holley Center for Agriculture & Health » Plant, Soil and Nutrition Research » Research » Publications at this Location » Publication #434155

Research Project: Championing Improvement of Sorghum and Other Agriculturally Important Species through Data Stewardship and Functional Dissection of Complex Traits

Location: Plant, Soil and Nutrition Research

Title: An Incremental Multi-Agent AI Pipeline for Marker Knowledge Extraction and Ontology Harmonization to Support Sorghum Genotyping Array Design

Author
item CHOUGULE, KAPEEL - Cold Spring Harbor Laboratory
item OLSON, ANDREW - Cold Spring Harbor Laboratory
item Gladman, Nicholas
item KUMARI, SUNITA - Cold Spring Harbor Laboratory
item LU, ZHENYUAN - Cold Spring Harbor Laboratory
item WEI, SHARON - Cold Spring Harbor Laboratory
item Ware, Doreen

Submitted to: Cold Spring Harbor Meeting
Publication Type: Abstract Only
Publication Acceptance Date: 5/26/2026
Publication Date: 5/26/2026
Citation: Chougule, K., Olson, A., Gladman, N.P., Kumari, S., Lu, Z., Wei, S., Ware, D. 2026. An Incremental Multi-Agent AI Pipeline for Marker Knowledge Extraction and Ontology Harmonization to Support Sorghum Genotyping Array Design. Cold Spring Harbor Meeting. 90th Cold Spring Harbor Laboratory Symposium on Quantitative Biology: AI in Biology.

Interpretive Summary:

Technical Abstract: Sorghum is a climate-resilient crop of increasing importance for food, feed, and bioenergy production, but translating published marker evidence into reusable breeding resources remains a major bottleneck. Trait-associated marker information is often dispersed across PDFs, supplementary tables, and heterogeneous document formats, making manual curation slow, inconsistent, and difficult to maintain. We developed a modular multi-agent AI pipeline to automate extraction, normalization, scoring, deduplication, and ontology harmonization of sorghum marker knowledge for downstream genotyping array and marker panel design. The pipeline processes PDF, TXT, and XLSX inputs through coordinated agents for document ingestion, study classification, marker extraction, evidence scoring, record canonicalization, and ontology annotation. From narrative text and tabular content, it captures marker identifiers, marker type, associated genes, chromosome and genomic position, assay information, target traits, and source provenance. Trait and biological context are harmonized using the EMBL-EBI Ontology Lookup Service Model Context Protocol (OLS MCP) server, enabling annotation with Trait Ontology (TO), Plant Ontology (PO), and Gene Ontology (GO) terms for interoperability and downstream prioritization. Applied to an initial sorghum literature set of 14 papers, the system extracted 230 raw marker records that resolved to 73 canonical marker entities after normalization and deduplication, reducing redundant mentions by approximately 68%. The workflow produces traceable outputs including canonical and per-paper marker tables, paper-level summaries, and a complete JSON run report. To enable continuous curation, we implemented incremental orchestration using file-level content fingerprints; in a test update cycle, the system reused 10 cached documents and processed only 4 newly added or modified files. This pipeline is being deployed as a curation layer for a community-driven, pangenome-informed ~100K sorghum genotyping array integrating trait-associated markers from public and international breeding programs. By converting heterogeneous literature into a structured and auditable marker knowledge base, this framework supports scalable incorporation of GWAS loci, functional markers, cloned genes, QTL-linked variants, and pan-genome candidates into sorghum breeding resources. These results demonstrate how multi-agent AI systems can accelerate biological knowledge synthesis while preserving traceability, reproducibility, and compatibility with domain ontologies. More broadly, this work highlights the value of AI-assisted scientific curation as enabling infrastructure for crop genomics and molecular breeding.