Skip to main content
ARS Home » Pacific West Area » Albany, California » Western Regional Research Center » Crop Improvement and Genetics Research » Research » Research Project #444625

Research Project: GrainGenes- A Global Data Repository for Small Grains

Location: Crop Improvement and Genetics Research

2025 Annual Report


Objectives
GrainGenes is an international, centralized crop database for peer-reviewed small grains data and information portal that serves the small grains research and breeding communities (wheat, barley, oat, and rye). The GrainGenes project ensures long-term data curation, accessibility, and sustainability so that small grains researchers can develop new, more nutritious, disease and pest resistant, high yielding cultivars. Objective 1: Accelerate small grains (wheat, oats, barley, and rye) trait, germplasm, genetics and genomics, and breeding data analysis and information by curating small grains genome sequences, germplasm diversity information, pangenomes, trait mapping information, and phenotype data into GrainGenes. Sub-objective 1.A: Integrate small grains genome assemblies, pangenomes, and annotations into GrainGenes. Sub-objective 1.B: Integrate genetic, diversity, functional, and phenotypic data into GrainGenes with a pangenome-centric focus. Objective 2: Develop computational and visualization tools to curate, integrate, and query the genetic, genomic, and phenotypic relationships in small grains germplasm, and deploy machine learning and artificial intelligence approaches to enhance functional annotations and discover biological interactions. Sub-objective 2.A: Develop methods and pipelines to link genetic, genomic, functional, and phenotypic information and to enhance pangenome-centric focus. Sub-objective 2.B: Implement web-based and computational tools to integrate and visualize genomic data linked with genetic, expression, functional, and diversity data. Objective 3: Collaborate with database developers and plant researchers to develop improved methods and mechanisms for open, standardized data and knowledge exchange to enhance database utility and interoperability. Sub-objective 3.A: Collaborate with data and germplasm repositories and organizations to facilitate the curation, sharing, and linking of data. Objective 4: Provide community support and training for small grains researchers through workshops, webinars, and other outreach activities. Sub-objective 4.A: Facilitate communication and information sharing among the small grain communities and GrainGenes to support research needs.


Approach
As a service project, the GrainGenes team does not perform hypothesis-driven research, but rather fulfills its long-term objectives by adding value to peer-reviewed data generated by others. It provides data curation, management and integration, long-term sustainability, and digital platforms as needed. Driven by stakeholder input, GrainGenes will maintain a central location for curated genomic, genetic, functional, and phenotypic data sets, downloadable in standardized formats, enhanced by intuitive query and visualization tools. Objective 1: Our approach will be to (a) curate genomic, pangenomic, and diversity data into GrainGenes database; (b) create new genome browsers, gene model pages to aggregate and link genomic and genetic data at GrainGenes; (c) curate high-impact, peer-reviewed genetic, trait, phenotypic data into GrainGenes; (d) visualize more accurate genetic maps at GrainGenes; and (e) curate functional and structural annotations (gene ontology, enzymatic functions, protein structure). For Objective 2: we will (a) create better search indexing and linking for data discovery at GrainGenes; (b) implement computational pipelines to link and align genomic and genetic features between different genome assemblies and GrainGenes pages; (c) implement computational pipelines to link and align genomic and genetic features between different genome assemblies and GrainGenes pages; (d) implement pipelines to facilitate data curation into the GrainGenes database; (e) implement and maintain genome browsers that allow comparative viewing using JBrowse2; (f) implement and maintain genome browsers to display tracks for multiple genome assemblies; and (g) create a BLAST plug-in that can be easily installed in JBrowse instances to allow users to align their sequences against small grains genome assemblies from JBrowse. For Objective 3: we will (a) enhance links and data sharing between GrainGenes and the Triticeae Toolbox for small grains data; (b) collaborate with other data and germplasm repositories, groups and organizations to facilitate the curation, sharing, and linking of data; (c) improve data interoperability and data sharing with WheatIS; (d) coordinate with ARS databases MaizeGDB and The Triticeae Toolbox to establish distributed infrastructure to serve users faster and more reliably; and (e) actively participate in the AgBioData Consortium. For Objective 4: we will (a) present GrainGenes tools and resources in conferences and site visits; (b) create training videos to teach users how they can use GrainGenes more efficiently; (c) organize annual meetings between GrainGenes and the GrainGenes Liaison Committee to receive community feedback; (d) maintain GrainGenes and OatMail e-mail lists to help the communication among the members of small grain communities; and (e) maintain and provide digital platforms to small grain researchers as needed.


Progress Report
This report documents the FY 2025 progress of project 2030-21000-056-000D, titled, “GrainGenes- A Global Data Repository for Small Grains”, which began in April 2023. In support of Sub-objective 1A, several genome and pangenome assemblies were acquired and displayed in GrainGenes. GrainGenes now has a total of more than 100 genome browsers for varieties of wheat, barley, oat, and rye (not including 35 genome browsers under embargo for the PanOat project). In FY25 alone, more than 25 new genome browsers were created for various varieties of wheat and oat, displaying their genomic features in a facilitated display enriched by linkages to other databases and GrainGenes pages. For Sub-objective 1B, the genetic information from over 30 journal articles were curated into GrainGenes and associated pages were created. The curated data includes 283 quantitative trait loci, 258 meta quantitative trait loci, and 39 genes for disease and agronomic traits related studies. In addition, over 30 Marker-Assisted Selection in Wheat protocols were brought in and linkages were created for small grains researchers. In support of Sub-objective 2A, GrainGenes personnel developed a framework called the GrainGenes Application Programming interface (GG-API) to facilitate browser feature searcing as a service to the collection of genome browsers in GrainGenes. Several improvements to indexing have been made to facilitate more efficient lookup as well as support the “Genome Browser Lookup by Gene/Transcript” feature through more efficient coding. For Sub-objective 2B, ARS GrainGenes personnel in Albany, California, have created a new pangenome page and revamped our genome browser landing page to facilitate connections between JBrowse1 (for single genome assemblies) and JBrowse2 (for comparative views between multiple genome assemblies). Through these new and revamped pages, GrainGenes users will be able to reach tracks in different genome browsers in a facilitated manner. In support of Sub-objective 3A, the data files that are accessible to Wheat Information System, which is operated under the Wheat initiative, were amplified with the addition of International Wheat Genome Sequencing Consortium’s Chinese Spring wheat variety, versions 1.0 and 2.1 genome assembly and annotations. The datasets include information about gene model ids, their chromosomal coordinates, as well as links back to the associated regions of GrainGenes genome browsers. For Sub-objective 4A, GrainGenes personnel created two new walk-through tutorials covering two web-based important resources for small grains researchers. The first tutorial introduces the Komugi database where Wheat Gene Catalogue, which is curated over decades, are easily accessible through a user-friendly website. The second tutorial provides an overview of the web resource titled Marker-Assisted Selection in Wheat protocols, in which protocols elucidating how to test for the presence of disease-resistance and other genes that are agronomically important traits in wheat varieties are stored.


Accomplishments
1. An improved web-based digital platform to help unlock hidden potential in ancient wheat varieties. Adapting wheat varieties to meet a variety of challenges, from shifting weather to new pathogens and pests, is essential to ensuring growers can keep America fed. However, progress in trait selection can be slowed by a lack of information, particularly with regard to genes linked to traits of interest. To expedite this critical research, ARS researchers in Albany, California, made significant changes to an online dashboard, providing DNA sequence information about Einkorn, one of the oldest ancestors of wheat, and Aegilops tauschii, a wild species with durable genes. This publicly available tool has been widely promoted and provides breeders a roadmap with the fastest route to better crops.


Review Publications
Poretsky, E., Blake, V.C., Andorf, C.M., Sen, T.Z. 2025. Assessing the performance of generative artificial intelligence in retrieving information against manually curated genetic and genomic data. Database: The Journal of Biological Databases and Curation. 2025. Article baaf011. https://doi.org/10.1093/database/baaf011.
Grewal, S., Yang, C., Krasheninnikova, K., Collins, J., Wood, J., Ashling, S., Scholefield, D., Kaithakottil, G.G., Swarbreck, D., Yao, E., Sen, T.Z., King, I., King, J. 2025. Chromosome-level haplotype-resolved genome assembly of bread wheat’s wild relative Aegilops mutica. Scientific Data. 12. Article 438. https://doi.org/10.1038/s41597-025-04737-y.