Location: Plant, Soil and Nutrition Research
2024 Annual Report
Objectives
Objective 1: Develop species-transferable models of gene expression and protein activity for crop plants (build using maize, sorghum, rice, wheat, Brachypodium, Setaria, tomato, Arabidopsis, soybean, and cassava data) derived solely from DNA sequence information.
Sub-objective 1.A: Develop bioinformatics to support pangenome comparisons, analysis, and imputation of haplotypes for molecular quantitative genetics and breeding.
Sub-objective 1.B: Develop and evaluate gene activity models for GWAS and GWP.
Sub-objective 1.C: Annotate pangenomes with gene activity estimates and make them accessible to molecular breeder analysis tools.
Objective 2: Apply species-transferable models of gene expression and protein activity to enhance allele mining and genomic selection for maize, minor crops, and specialty crops (including but not limited to sorghum, oat, cassava, table grapes, and others) for the identification of genetic variation controlling frost/heat tolerance and nitrogen/phosphorous recycling.
Sub-objective 2.A: Allele mining for temperature tolerance in maize and its wild relatives.
Sub-objective 2.B: Allele mining for nutrient recycling in maize and its wild relatives.
Sub-objective 2.C: Allele mining and genomic selection using gene activity for breeding across Breeding Insight species.
Approach
Plant breeding and genetics are poised to contribute to numerous goals in feeding the planet, adapting to climate change, and reducing the environmental impact of agriculture. To date, breeding models have been crop-specific, drastically limiting their efficacy and scope. Genomics and machine learning, however, have now developed to the point where the molecular activity of genetic variation, and thereby impacts on plant performance in the field, can be estimated solely from genome sequence information. The goal of this project now is to shift breeding models from being genetic variant-based to being molecular activity-based (e.g., protein structure and RNA/protein expression). New cross-species models will be generated and leveraged to predict whole-plant phenotypes and uncover environmental adaptation strategies, making prediction and discovery more accurate, efficient, and powerful. The project will develop bioinformatic and machine learning tools to infer the complete genomes of crop species (pangenomes) and then use that sequence to predict the protein structure and gene expression patterns in thousands of crops and related wild species. Molecular breeders working across species use these models through bioinformatic portals we develop.
This project will then allele mine for genes in maize and sorghum’s wild relatives that control frost and heat tolerance, as well as recycling nitrogen and phosphorous. These wild relatives, many of which are perennial, are more tolerant of extreme temperatures and are incredibly productive without needing chemical inputs. The candidate genes will be identified by genetic mapping within Tripsacum and Zea, environmental mapping using gene activity across the Andropogoneae and landraces, and expression/metabolite profiling in wild species. Successful identification of these genes could lead to a dramatic reduction in the need for fertilizer inputs and increases in maize yields by allowing earlier planting and avoiding and tolerating heat extremes.
Progress Report
Finding shared genes within and across species: In previous years, we demonstrated that AI Large Language Models (LLMs) for protein sequences were effective tools for identifying causal variants in protein coding regions, improving predictions for complex traits like heterosis and yield. However, these analyses were restricted by their inability to identify causal variants linked to gene regulation. Recently, inspired by advancements such as ChatGPT and Microsoft Copilot, we partnered with a leading team from Cornell's Computer Science department to create PlantCaduceus, a groundbreaking LLM with single base pair resolution. This foundational model offers resolution for all causal variants and is adaptable for other tasks, showing remarkable cross-species transferability. An added advantage of these LLM is their ability to perform robustly on new, unseen species after training on just a few species.
In parallel, we have been refining machine learning models designed to predict gene expression across species. Although these models predate PlantCaduceus and are less sensitive to single base variations, they have demonstrated notable accuracy. Over the next year, we anticipate a significant integration of all these models. Additionally, we are expanding our efforts to train models on eukaryotic proteins to pinpoint proteins that play roles in cold tolerance and protein storage.
Our Practical Haplotype Graph (PHG), a powerful way to represent the haplotype diversity of a crop, has been redesigned and streamlined from the ground up this year. This redesign was prompted by recent improvements in tools to store and compress genotypic information. Rebuilding the system also allowed our team to integrate standard software development practices like rigorous unit testing, code coverage, Continuous Integration and Continuous Delivery to create easy to install and robust software. Careful attention was also paid to make the User Interface simple to use. By doing so, users can start and finish a full build in roughly a day, compared to a few weeks before. A simple BrAPI compliant webserver is also included in the package which allows users to efficiently access data through the internet. A R package is also being developed to allow for easy data retrieval and visualization. In the coming year, the PHG will be merged with the AI causal variant identification.
Apply shared gene effects to increasing nitrogen efficiency and adapting to climate change: This project is leading an effort to increase the nitrogen efficiency of the entire agricultural system, which, if successful, will reduce farmer fertilizer costs and decrease water pollution and greenhouse gas emissions. While cropping, livestock, and manure systems are all key to make this happen, a crucial starting point is corn production (currently uses over half of U.S. nitrogen fertilizer). The CERCA (Circular Economy for Reimaging Corn Agriculture) project is focused on increasing system level nitrogen efficiency in corn through enhancing cold tolerance and recycling nitrogen at the end of season back to the soil. This project and our collaborators have made progress in these domains. We have completed assembly of 40 species related to maize that have both traits. Analysis of these genomes indicates that polyploidy is mostly a mechanism to avoid inbreeding depression, however, the Zea-Tripsacum lineage appears to be an exception where polyploidy opened new adaptive evolution avenues. Using these genomes, plus an expanded set of over 500 grass species at a lower quality, we are investigating the genetic basis of perennial-to-annual transitions and environmental adaptation to different climates. We have identified groups of genes with signal of differential selection for both, and we will explore them further in the next year.
For cold tolerance, we have identified 120 proteins that are more abundant in tissues that survive the winter. Many of these proteins have structural differences relevant to cold tolerance, including lipid transfer proteins, heat shock proteins, late embryogenesis abundant proteins, and other drought-related proteins. 30 of the leading candidates have entered transgenic pipelines for testing in maize. Additionally, 23 grass species were extensively evaluated in growth chamber experiments for their regulatory responses to frost. Preliminary results show a promising overlap with our previous list of candidates, plus additional new insights. We also evaluated 63 different highland maize landraces from the USDA-ARS GRIN (Germplasm Resources Information Network) germplasm repository, across 8 locations and planting dates, and performed RNA expression profiling, along with aerial UAV (unmanned aerial vehicle) and visual scoring. This year we will have a detailed analysis of this germplasm and additional germplasm from highland southwest U.S., and we will determine if useful variation can be found in maize or in wild species related to maize.
Biological nitrification inhibition (BNI) is a process that prevents nitrogen from being lost to the environment through its conversion to nitrate. We tested the hypothesis that substantial BNI might be found in perennial species, but contrary to our expectation the reverse was true. Annual species, both domesticated and wild, showed more BNI. This suggests it might be difficult to further enhance this process among our annual crops using perennial relatives. However, it may still be possible to learn how to recycle nitrogen back to soils at the end of the growing season (as perennial crops do) and identify the regulatory mechanisms required to remobilize nitrogen during senescence without hurting crop yield. Extensive field trials, where we measure nitrogen status and RNA profiles, point to common strategies that maize’s perennial wild relatives use in regulating their genes in roots and leaves throughout late summer and fall. Next year, we will continue to study the genetic basis of these strategies to identify the mechanisms that may be useful in maize.
Making allele mining accessible to breeding programs: We are working to make our newly developed tools and models accessible to breeders. The Breeder Genomics Hub is a JupyterHub based notebook system that enables creation of reproducible pipelines that can be easily shared with other collaborators. This Breeder Genomics Hub is designed to have a large suite of internal and community standard tools preinstalled so a breeder can work without needing to manage installation of different packages. All tools developed within our group are designed to be easily integrated into this Genomics Hub when possible.
Breeding Insight (BI) is the ARS initiative to increase the adoption of genomics, phenomics, and analytics tools (including data management software) in ARS specialty crop and animal breeding programs. BI is currently in year 6 (phase II), and its sister program, BI OnRamp, is in year 4. Together, BI and OnRamp provide breeding support services for 28 ARS species (blueberry, table grape, sweetpotato, alfalfa, rainbow trout, North American Atlantic salmon, honeybee, strawberry, cranberry, oat, pecan, lettuce, cucumber, sorghum, hemp, citrus, sugarcane, soybean, cotton, hydrangea, sugar beet, red clover, cover crops, hop, raspberry, blackberry, potato, and coffee), with BI providing support to ~60 breeder programs. The future goal is expansion to all ARS specialty crops, animal, and natural resource breeding programs.
In FY 2023-2024, BI’s most significant accomplishment was the widespread adoption of genomic resources by ARS scientists and their public counterparts for species that previously lacked these resources. This year, BI created custom genotyping resources for cranberry, pecan, lettuce, cucumber, honeybee, and sweetpotato potato weevil, which brings the total number of marker panels made by BI to 10 (alfalfa, blueberry, salmon, and sweetpotato panels were created in prior years of BI). This is a substantial and important leap forward from 2019, when there were few or no genetic markers available for these species. Since these resources have become publicly available, BI has genotyped more than 34,000 varieties in FY23-24 alone (~71,000 varieties over the life of BI). Genomic data for the species allows for the creation of genetic maps, GWAS, QTL analysis, and genomic predictions in ARS breeding programs while simultaneously enhancing stakeholder engagement, collaboration, and publications. These resources are also being utilized by private industry breeders in the US and around the globe. The adoption of the genotyping platform and pipeline benefits the entire global breeding effort such that all breeders have access to the same markers to improve data sharing with FAIR data principles. Given this success, Breeding Insight has already initiated new genotyping resources for several Phase 3 species, including blackberry, hemp, hop, red clover, and trout. Putting these powerful analyses and genomic tools into the hands of ARS’s excellent specialty crop and animal breeders improves breeding decisions and helps to meet public demand for more nutritious and flavorful foods.
BI’s most significant software development accomplishment is the release of version DeltaBreed 0.9 in January 2024, which includes germplasm loading functionalities, improved pedigree support for perennial and clonal crops, multi-location experimental trial support, additional breeding methods, sub-observation datasets, BrAPI-driven sample submission for genotypic experiments, Gigwa integration for haplotype genetic data storage, and a jobs module for monitoring the status of data uploads, downloads, and processes. The software team has made major improvements to the back-end communications between Field Book and DeltaBreed through the BrAPI connection, pushing for endpoint expansion when necessary. As with all BI software, its fully BrAPI compliant, open source and publicly available.
Accomplishments
1. Artificial intelligence in the form of large language models and applying DNA sequences. Artificial intelligence in the form of large language models have advanced dramatically in power in the last couple of years, but one of the greatest challenges were applying these to DNA sequences and large stretches of DNA have little relevant information and other regions single base changes matter. USDA scientists in Ithaca, New York, in collaboration with computer scientists from Cornell University developed a state-of-the-art large language model for analyzing the changes in DNA and have trained and tested the model on flowering plants. It exhibits state of the art transferability across flowering plants and a precision that rivals protein models. This model can be used to identify deleterious mutations, which are key to yield and hybrid vigor, and to precisely annotate the functional regions of plant genomes. This model and future iterations will allow researchers to apply advanced genome knowledge to all crops for improvement.
2. Breeding Insight expanded access to field and genomic tools for 28 specialty crops and animal species. One of the major challenges in breeding is the integration and processing of billions of genomic and field data points needed to make informed decisions. Breeding Insight (a USDA-ARS cooperative agreement with Cornell University) expanded again in 2024 to support 9 new the number of crops (sorghum, hemp, hop, sugar beet, hydrangea, coffee, potato, red clover, and cover crops) in the program with the establishment of data management systems for all species, new genomic tools are available for a quarter of the species, and field informatics tools for three-quarters of them. Many of the specialty crops (e.g., blueberry, alfalfa, strawberry, blackberry, potato, and sweetpotato) have genome duplications that make genomics tools challenging to apply, but this year the teams were able to apply these tools to these complex genomes. Putting these powerful analyses and genomic tools into the hands of ARS’s excellent specialty crop and animal breeders helps to improve breeding decisions and to meet public demand for more sustainable, nutritious, and flavorful foods.
Review Publications
Vorontsova, M.S., Peterson, K.B., Minx, P., Aubuchon-Elder, T.M., Romay, M., Buckler Iv, E.S., Kellogg, E.A. 2023. Reinstatement and expansion of the genus Anatherum (Andropogoneae, Panicoideae, Poaceae). Systematics and Biodiversity. 21(1). https://doi.org/10.1080/14772000.2023.2274386.
Morales, N., Anche, M.T., Kaczmar, N.S., Lepak, N.K., Ni, P., Romay, M., Santantonio, N., Buckler Iv, E.S., Gore, M.A., Mueller, L.A., Robbins, K.R. 2024. Spatio-temporal modeling of high-throughput multispectral aerial images improves agronomic trait genomic prediction in hybrid maize. Genetics. 227(1):iyae037. https://doi.org/10.1093/genetics/iyae037.