Skip to main content
ARS Home » Midwest Area » Ames, Iowa » Corn Insects and Crop Genetics Research » Research » Publications at this Location » Publication #432077

Research Project: MaizeGDB - Database and Computational Resources for Maize Genetics, Genomics, and Breeding Research

Location: Corn Insects and Crop Genetics Research

Title: AI-ready genomics and multi-omic functional annotation at MaizeGDB

Author
item HALEY, OLIVIA - Oak Ridge Institute For Science And Education (ORISE)
item TIBBS-CORTES, LAURA - Oak Ridge Institute For Science And Education (ORISE)
item HARDING, STEPHEN - Oak Ridge Institute For Science And Education (ORISE)
item Cannon, Ethalinda
item Portwood Ii, John
item GARDINER, JACK - University Of Missouri
item Woodhouse, Margaret
item Andorf, Carson

Submitted to: Maize Genetics Conference Abstracts
Publication Type: Abstract Only
Publication Acceptance Date: 2/3/2026
Publication Date: 2/28/2026
Citation: Haley, O., Tibbs-Cortes, L., Harding, S., Cannon, E.K., Portwood II, J.L., Gardiner, J.M., Woodhouse, M.H., Andorf, C.M. 2026. AI-ready genomics and multi-omic functional annotation at MaizeGDB. Maize Genetics Conference Abstracts.

Interpretive Summary:

Technical Abstract: The integration of Artificial Intelligence (AI) into computational biology is improving agricultural research by enabling the extraction of biological insights from increasingly complex datasets. As a primary resource for the maize (Zea mays L.) community, the Maize Genetics and Genomics Database (MaizeGDB) is proactively establishing an AI-ready infrastructure designed to support machine learning applications. This strategic initiative involves standardizing multi-omic datasets, generating precomputed embeddings from state-of-the-art DNA and protein language models, and providing reproducible workflows via GitHub. New AI-driven functionalities include zero-shot variant-effect scoring and genome browser tracks for nucleotide conservation, which facilitate the interpretation of functional significance across the genome. These AI resources are deeply integrated with an expanded suite of functional annotation tools and visualization tracks within the MaizeGDB genome browser. New Syntenome tracks provide whole-genome alignments with other grasses and lifted gene models to support comparative analyses and gene model validation. Users can also access unmethylated regions, aligned protein fragments from the Maize Peptide Atlas, and high-confidence protein structure predictions from AlphaFold and ESMFold. To capture structural and functional diversity across lineages, the database incorporates pangenome and pan-gene annotations from the Nested Association Mapping (NAM) founder genomes. Furthermore, over 600 epigenetic and DNA-binding datasets provide critical regulatory context, including open chromatin regions and histone modifications. The integration of these data with tools such as SNPVersity and PanEffect allows researchers to navigate from sequence variation to functional impact. By combining AI-ready data with a robust framework for inter-species and cross-species comparisons, MaizeGDB empowers the maize genetics community to accelerate gene function discovery and the development of improved maize varieties. By delivering AI-ready maize genomics and multi-omic functional annotation with scalable, reproducible machine learning workflows, MaizeGDB advances USDA's priorities by accelerating research and breeding efforts for the discovery and deployment of higher-yielding, stress-tolerant, and quality-enhanced maize that boosts producer profitability, strengthens resilience to pests and diseases, and improves human health through better food quality and nutrition.