Skip to main content
ARS Home » Northeast Area » Ithaca, New York » Robert W. Holley Center for Agriculture & Health » Plant, Soil and Nutrition Research » Research » Research Project #445558

Research Project: The Triticeae Toolbox (T3) - Breeding Process and Data Facilitation to Accelerate Genetic Gain and Discovery

Location: Plant, Soil and Nutrition Research

2024 Annual Report


Objectives
Objective 1: Conduct research to simplify getting data into T3 and to expand T3 functionality for users. Sub-objective 1.A: Sample tracking with the small grains genotyping labs. Sub-objective 1.B: Data upload through Android Field Book, Coordinate, and Intercross apps. Sub-objective 1.C: Spatial analysis of individual field trials. Sub-objective 1.D: Target population of environments (TPE) determination. Sub-objective 1.E: Determine the value to selection gain from weighting field evaluations by their genetic correlation to the long-run TPE average. Sub-objective 1.F: Genomic mate selection (GMS). Objective 2: Develop and incorporate analyses into T3 that strengthen initial data curation. Sub-objective 2.A: Practical haplotype graph incorporation. Objective 3: Strengthen T3 interoperability with other biological databases. Sub-objective 3.A: Creation of private instances of T3. Sub-objective 3.B: Ensure T3 remains compliant with the BrAPI specification Objective 4: Assemble a T3 User Group to advise developers on advances in the breeding process and on data management feature needs. Sub-objective 4.A: Host a meeting of the T3 Advisory User Group (TAUG) twice each year.


Approach
Research and development work on Objective 1 to simplify getting data into T3 and to expand T3 functionality includes some small subobjectives and some larger ones. We will develop interfaces to track samples coming from the small grains genotyping labs and to better interface with digital data acquisition devices like tablet apps. We will add functionality to run spatial analyses on phenotyping trials that have field layouts. These features will culminate in better data to determine target populations of environments (TPE) for breeding programs. With the TPE defined, we will seek to use multi-year and multi-location trail data to better weight data from trials in for selection. Finally, we will implement genomic mate selection. Work on Objective 2 really focuses on harmonizing marker data across genotyping platforms or protocols. We will do this by integrating the practical haplotype graph into T3. Over time, we have come to acknowledge that breeders do not always want their data to be made public rapidly. In Objective 3 to avoid that putting data into T3 forces that, we will create private instance of T3. Nevertheless, to ensure ongoing T3 interoperability with open-source breeding data analysis packages, we will ensure that T3 remains compliant with the BrAPI specification. Finally, we know T3 has faults that we are not fully aware of. To better appreciate and be able to address these faults, in Objective 4, we will assemble a T3 User Group to advise developers on advances in the breeding process and on data management feature needs. We will convene this group twice each year.


Progress Report
Overall, the Project Plan focuses strongly on improving The Triticeae Toolbox (https://triticeaetoolbox.org/) as a service for United States public sector small grains breeders. Consistent with this focus, there are several innovations running the gamut from small scale useability features to larger scale features still in process and on to broader efforts at modernizing the overall public sector small grains breeding enterprise in which T3 is embedded, efforts will enable this sector to better leverage cutting edge computational and sequencing innovations. Small scale useability features • Barcodes are a transformational productivity and data validity technology, dramatically reducing tracking errors relative to human-readable labels. T3 has a barcode design interface that enables many of the data types stored by T3 to be turned into barcode labels. That interface has been redesigned to increase flexibility and ease of use. • Historically, T3 used accession names to label plots. Breeders, however, routinely use entry numbers and requested that we add entry number features to trials in T3. We have done so, facilitating, for example, data upload and download in entry number order. • Breeders investing the effort to upload data to T3 want to track that effort, as do grant projects that use T3 as their data management solution. We have created a simple dashboard that gives annual summaries of data uploads by breeding programs facilitating reporting for both program and grant managers. • Finally, the Android Field Book app facilitates a seamless process between data collection in the field and upload to the database. Setting up the proper integration needed better validity testing and documentation, and we have provided that. Larger data functionality features The largest time sink in wrangling multiple datasets to run a joint analysis is getting all accession names to match correctly across trials. Consequently, it is important when uploading trials into T3, be they phenotyping or genotyping trials, to ensure that the accessions in those trials are either identified as new accessions or properly connected to previously existing accessions. This task is challenging because of the many ways in which accession names can be unwittingly changed by people, for example by adding or removing spaces, punctuation, or prefixes. Over the past year in T3 we have developed algorithms to consistently synonymize accessions and perform the task with warnings systematically upon data upload. This effort reduced our raw accession count by about 20% but now ensures much greater data validity and hence data analysis power. Experimental agricultural fields in which phenotyping trials are performed are heterogeneous and many studies have shown statistically accounting for that heterogeneity can improve the accuracy of evaluations. Consequently, T3 must be able to store field layout information and be able to account for it in internal analyses. New functionality warns users when they have not uploaded the data, we have integrated linear mixed model analyses to use it in spatial analyses. The most important agronomic trait, yield, is polygenic: there are no single variants that strongly affect it. But there are many single loci that affect an accession's suitability to become a variety. These are major disease resistance, vernalization, photoperiod sensitivity, and height loci. Some abiotic stress and quality traits also are affected by major loci. Breeders need to be able to track these loci, both to understand phenotypes as they develop in the field and to decide which accessions to cross choosing to cross parents carrying complementary major alleles. Functionality for tracking major alleles has been lacking in T3, partly because the process is complicated. Identifying the alleles often starts not with typing them directly but by typing DNA markers with which they are in linkage disequilibrium (assuming the causal variant itself is not known). These markers are "Known Informative Markers" or KIMs. In T3, we have now created the proper data structure to manage KIM and major locus allele data. That is the first step. We are now in collaboration with the directors of all four Small Grains Genotyping Labs (SGGLs) to develop the process whereby the SGGLs will upload KIM data and T3 will infer from that data alleles present at the major loci. From there, we will further develop functionality to search accessions based on their major alleles and to display known alleles for any accession. Modernizing public sector small grains breeding to leverage data management To make T3 development more responsive to users, we have organized the T3 Advisory User Group (TAUG). The TAUG is composed of small grains breeders, small grains data analysts, other database development scientists, and researchers who contribute to the data such as the directors of the small grains genotyping labs. TAUG's mandate is to keep the T3 team accountable on its curation of data and deployment of new features, and to suggest other areas of development, technologies and resources. The purpose of the first TAUG meeting (Nov. 29th, 2023) was to introduce T3 and its vision, to make sure all on the TAUG were on the same page and in agreement on what the TAUG can do for T3. In the second meeting (May 29th, 2024) the TAUG discussed features. In general, there is a consensus that progressive, steady improvement to T3 is needed. The TAUG discussed efforts important to the general community and features important more directly to breeders using T3 for their data management. Important efforts include comprehensive availability of historical cooperative nursery data and search, and access features focused on alleles present at major loci in accessions on T3. Scientists in Ithaca, New York, have now made substantial progress in these areas. Progress has also been made on documenting the use of the tablet app Field Book for joint data collection and upload to T3 and on expanded functionality surrounding entry numbers in field trials. Finally, the TAUG discussed new features. An improved ability to identify historical trials relevant to current sets of accessions, basic improvements in data upload and download interfaces, and better integration with third parties, such as the Small Grains Genotyping Labs, who will submit data to T3 going forward. A second organizational body that we participate in are the wheat and oat synergy executive committees (SECs). In addition to data management, but the SECs will also be a platform for communication and collaboration around opportunities for data synergy. One of the activities ripest for improvement from centralized data management are the USDA-coordinated cooperative nurseries. These nurseries test experimental lines from many USDA and Land Grant university breeding programs; a resource for understanding the adaptation of advanced experimental lines and exploring genotype by environment interaction. That exploration can only take place if the data are easily aggregated. In the past year we have engaged in extensive effort to obtain and curate cooperative nursery data (see Table) Table: Cooperative Nursery additions to T3 in the past 12 months Crop Cooperative Nursery Number of Trials Trial Years New Accessions in T3 With Genotype Without Geno Wheat Uniform Regional Scab 147 1995 - 2023 297 163 S Southern Regional Performance 16 2023 Northern Regional Performance 5 2023 Barley North American Barley Scab Evaluation 7 2022 23 152 Western Regional Spring Barley 79 2016 - 2023 Winter Malting Barley Trial 20 2023 Oat Uniform Early Oat Performance 10 2023 46 155 Uniform Midseason Oat Performance 11 2023 Uniform Winter Oat Yield 95 2004 - 2023 To facilitate this process, we have also begun discussion with cooperative nursery coordinators to use a uniform data template. This template has been developed over a few years by a data analyst. Changes in the processes of these nurseries are as much social as they are technological, and that kind of change takes time. The Wheat Coordinated Agricultural Project (WheatCAP) uses T3 as its central data management repository. An important effort of WheatCAP is to expand the use of drone-based high throughput phenotyping. Flying the drones and collecting images are well within the skill set of breeders but often image analysis to extract plot characteristics is challenging. The WheatCAP has centralized an analysis service based at Texas A&M, UASHub. The major progress is data extracted from images at UASHub is now directly uploaded to T3, without going through the breeders uploading the data. This direct upload has dramatically simplified the work of breeders and consequently increased the stream of data coming to T3, we view this collaboration as a model that we will work to extend to other third-party data providing services, particularly those run by USDA-ARS scientists. For example, see the description of the collaboration with the SGGLs on managing data from KIMs and inferring the alleles present at major loci. Managing the large amount of genotypic data coming from genotyping services is more challenging for breeding programs and facilitating that management through T3 holds promise to free resources within those programs to focus on evaluation and selection strategies. Having had success integrating T3 with the UASHub and now being well on our way to a similar integration with data coming from the SGGL, we are hopeful that integration with data generators can become a general model. Other targets could be, USDA grain quality labs, or the cereal disease labs that provide measurements to public sector breeders that could be tied into a general management platform such as T3. Likewise, as described for the cooperative nurseries, such integration can simplify collaboration across programs in data generation, analysis, and use, thereby accelerating gain from selection broadly.


Accomplishments
1. Maximizing genetic gain in plant breeding despite genotype by environment interaction. Genotype by environment interaction (GEI) decreases the gain from breeding that reaches farmer fields because it creates uncertainty about whether the best genotype is being selected by the breeder and by the farmer for a specific farm environment. USDA-ARS scientists in Ithaca, New York, participated in research to mitigate this problem in two ways. First, they developed breeding simulations to help breeders determine if gain would be greater when selecting for varieties adapted broadly across a geographic area versus narrowly in multiple specific subareas. The difference hinged on how great differential variety response to the subareas was. Second, they analyzed extensive historical data of oat variety trials performed over many locations in the upper Midwest to determine, on a zip code by zip code basis, which varieties would perform best in each location. This analysis provided a tool for farmers to select the best varieties for their farms. Stakeholders will benefit from this research in the short term by being able to select the best varieties for their farms and in the long term because the research will empower programs to optimize their breeding schemes regarding the amount of GEI prevalent in the regions they serve.

2. Exploring better methods to breed for cassava brown streak disease resistance. The cassava brown streak virus (CBSV), causes major harm to smallholder cassava farmers in east and central Africa by causing necrotic lesions in cassava storage roots that renders them inedible. Resistance to CBSV has traditionally been selected by visually scoring root symptoms on a 1 to 5 scale, a possibly imprecise and subjective measurement. USDA-ARS scientists in Ithaca, New York, participated in research to increase the objectivity and precision of the measurement of CBSV in two ways, first by scoring lesions using objective image analysis and second by directly measuring virus titer in infected roots. Surprisingly, the image analysis did not improve on the traditional 1 to 5 scoring, with the latter generating greater heritabilities. Thus, greater gains may be obtained with the traditional than the more technologically sophisticated method. The genetic correlation between virus titer and symptoms was low suggesting that different genetic mechanisms contribute to resistance – the ability to prevent the virus from replicating – than tolerance – the lack of symptoms despite high virus load. Future research will be needed to choose the best way forward with respect to the differences between selecting for tolerance or resistance. Cassava breeder stakeholders will benefit from this research because it helps them in choosing the best technologies to use in attaining their breeding goals in selecting new resistant cassava varieties.

3. Connecting The Triticeae Toolbox to external data providers. Public sector small grains breeders can accelerate gain from selection by sharing data in The Triticeae Toolbox (T3), creating larger datasets that enable more accurate predictions and powerful QTL detection. Getting data into T3 to share it, however, is an extra task to which breeders must allocate time. That time allocation can be eliminated if external data providers push the data directly to T3. USDA-ARS scientists in Ithaca, New York, developed this functionality along with the wheat Coordinated Agricultural Project (WheatCAP) Unoccupied Aerial System Hub (UASHub) team from collaborators by leveraging the BrAPI application program interface available to T3. In the first year after direct uploads from UASHub were enabled, observations coming to T3 from drone-based high throughput phenotyping accounted for 407,225 new observations from 50 trials. Public sector small grains breeders will increasingly benefit from the functionality as the dataset sizes available in T3 grow making genomic predictions more accurate. Furthermore, these datasets will be needed for data hungry deep learning applications going forward.


Review Publications
Braatz De Andrade, L.R., Bandeira E Sousa, M., Wolfe, M., Jannink, J., Vilela De Resende, M.D., Azevedo, C.F., De Oliveira, E.J. 2022. Increasing cassava root yield: Additive-dominant genetic models for selection of parents and clones. Frontiers in Plant Science. 13:1071156. https://doi.org/10.3389/fpls.2022.1071156.
Nandudu, L., Kawuki, R., Ogbonna, A., Kanaabi, M., Jannink, J. 2023. Genetic dissection of cassava brown streak disease in a genomic selection population. Frontiers in Plant Science. 13:1099409. https://doi.org/10.3389/fpls.2022.1099409.