Location: Hard Winter Wheat Genetics Research
Project Number: 3020-21000-012-030-S
Project Type: Non-Assistance Cooperative Agreement
Start Date: Jun 1, 2026
End Date: May 31, 2028
Objective:
The objective of this project is to integrate legacy and newly generated genotypic data to create reproducible bioinformatic pipelines to predict wheat performance. These pipelines will facilitate rapid and consistent genetic data curation for regional small grains breeding programs across the hard winter wheat region. This work directly supports American farmers by providing wheat breeders with rapid access to genetic information required to accelerate development of highly productive U.S. cultivars. This research strengthens national agricultural security by providing genetic information required to increase the frequency of wheat resistance to critical diseases and insect pests, as well as drought and heat stress in commercial cultivars. This project will construct new analysis pipelines that transform historical and contemporary data into modern database systems, ensuring that critical genetic information is standardized, searchable, and actionable for quantitative genetic analysis. The specific objectives of this project are:
1. Create novel bioinformatics pipelines in order to utilize the AgriSeq marker panel in predicting legacy known informative markers for use in genomic assisted breeding efforts.
2. Establish a continuous integration workflow for storage, query, and analysis of genetic information to ensure utility of future data collection.
3. Transform of historical genetic datasets into regularized formats compatible with relational database schema.
4. Upload and curation of transformed historical and contemporary genetic data to a T3 Breedbase instance for easy internal and external access by USDA members and collaborators alike.
5. Create a pipeline, associated with the newly derived T3 genetic database, to both forward and backward predict the haplotype of key loci.
Approach:
AgriSeq Pipeline and Marker Reports: To address Objective 1, the cooperator will first assess the state of incoming genetic data derived from the AgriSeq marker panel and work with USDA members to understand the workflow generated by ARS, The AgriSeq panel is a recently developed marker platform utilized across USDA-ARS Small Grains Genotyping Labs to replace the majority of kompetitive allele-specific PCR (KASP) assays. This transition poses a challenge to the small grains breeding and genetics community with fundamental changes in genotyping platforms, file structures, and data manipulation.
After evaluating the current AgriSeq workflow, a novel bioinformatics pipeline will be constructed to process the variant calling format (VCF) files output from the system to predict haplotypes associated with known informative markers (KIMs). Using an algorithm designed by the cooperator, the VCF will be transformed into a HapMap format utilizing standard International Union of Pure and Applied Chemistry (IUPAC) nucleotide calls. The cooperator and ARS personnel will iterate through several methodologies to develop a reference file assigned haplotype calls associated with known loci. This process will result in a pipeline script that analyzes AgriSeq VCF data to produce marker panels in a consistent and predictable manner.
T3 Breedbase Integration: To address Objectives 3 and 4, the cooperator will conduct an internal review of existing genetic data repositories local to the USDA-ARS. This audit will identify changes in data formats, missing metadata, and bottlenecks within historical datasets. Subsequently, the cooperator will utilize Bash, Python, and R scripting to develop pipelines facilitating the transition of datasets into regularized tidy formats compatible with relational database schemata.
These transformed contemporary genetic datasets will then be curated and uploaded to a T3 Breedbase instance to facilitate easy internal and external access by USDA personnel and collaborators. This integration will allow regional nursery data to be ingested into the database and processed into easily queried marker data, facilitating the rapid generation of marker reports for breeding decisions. Throughout this process, the cooperator will inventory upcoming changes in genotyping platforms and consult with USDA-ARS members regarding reporting methods that best serve the interests of the USDA and its constituency.
Predictive Haplotyping.
To address Objectives 2 and 5, the cooperator will establish a continuous integration workflow for the storage, query, and analysis of genetic information. This architecture will include the creation of a predictive haplotyping pipeline, associated with the newly derived T3 genetic database, to both forward and backward predict the haplotype of key loci using unbalanced historical information.
The cooperator will deploy these pipelines both locally and on the SCINet high-performance computing cluster. This approach ensures that computationally intensive tasks, such as genotype file manipulation, imputation, and production, are reproducible and scalable to the larger service area.