Skip to main content
ARS Home » Northeast Area » Beltsville, Maryland (BARC) » Beltsville Agricultural Research Center » Soybean Genomics & Improvement Laboratory » Research » Publications at this Location » Publication #430684

Research Project: Characterization and Utilization of Genetic Diversity in Soybean and Common Bean and Management and Utilization of the National Rhizobium Genetic Resource Collection

Location: Soybean Genomics & Improvement Laboratory

Title: Classification-based genomic prediction for early identification of high-yielding and stable soybean genotypes

Author
item MARMO, GONCALVES - University Of Arkansas
item ACUNA, ANDREA - University Of Arkansas
item WU, CHENGJUN - University Of Arkansas
item FLOREZ-PALACIOS, LILIANA - University Of Arkansas
item HARRISON, DERRICK - University Of Arkansas
item ROGERS, DANIEL - University Of Arkansas
item SAGAE, VITOR - University Of Florida
item ROBERTS, TRENTON - University Of Arkansas
item ROSS, JEREMY - University Of Arkansas
item ZHANG, QINGYANG - University Of Arkansas
item Song, Qijian
item JARQUIN, DIEGO - University Of Florida
item VIEIRA, CAIO - University Of Arkansas

Submitted to: Frontiers in Plant Science
Publication Type: Peer Reviewed Journal
Publication Acceptance Date: 4/1/2026
Publication Date: 4/29/2026
Citation: Marmo, G.R., Acuna, A., Wu, C., Florez-Palacios, L., Harrison, D., Rogers, D., Sagae, V.S., Roberts, T.L., Ross, J., Zhang, Q., Song, Q., Jarquin, D., Vieira, C.C. 2026. Classification-based genomic prediction for early identification of high-yielding and stable soybean genotypes. Frontiers in Plant Science. 17. Article e1770360. https://doi.org/10.3389/fpls.2026.1770360.
DOI: https://doi.org/10.3389/fpls.2026.1770360

Interpretive Summary: Soybean breeders develop thousands of new soybean lines each year, hoping to find the highest-yielding and most stable lines. However, the limited number of seeds for each new line makes early testing difficult, limiting the number of testing sites and reducing the reliability of results. This study shows that modern DNA-based tools and machine learning methods can help breeders identify the most promising soybean lines earlier. Researchers from the University of Arkansas, the University of Florida, and the USDA-ARS, Beltsville, Maryland evaluated more than 1,800 soybean genotypes in ten field settings and created a simple scoring system that categorized each line into high-yielding, uncertain-yielding, or low-yielding categories. They then trained two computer models that used DNA markers rather than relying solely on field trials to predict these categories. Both models—a Generalized Linear Model Network and a Random Forest—were able to correctly classify most soybean lines with extremely low error rates. Importantly, these models were able to accurately identify most low-yielding lines before costly multi-site trials. This approach provides breeders with a practical and economical tool that allows them to focus their resources on soybean lines most likely to succeed in future trials, ultimately providing farmers with higher-yielding varieties.

Technical Abstract: mproving grain yield remains the central objective of soybean breeding programs. During early-stage yield trials, breeders often evaluate thousands of genotypes; however, limited seed availability constrains the number of tested environments and replications, reducing selection accuracy. Genomic prediction offers a promising approach to identify high-yielding and stable genotypes earlier in the breeding pipeline. In this study, 1,803 genotypes ranging from maturity groups III to V were evaluated for grain yield across ten environments (year × location combination) in 2023 and 2024. Best Linear Unbiased Predictors (BLUPs) were obtained for each genotype in each environment, and a selection index (MSI) was calculated as the average yield deviation from the checks’ mean across tested environments, centered at zero. Genotypes were then classified as high-yielding (MSI = –5), uncertain (–5 > MSI = –15), or low-yielding (MSI < –15). Two classification-based genomic prediction models, Generalized Linear Model via Elastic Net Regularization (GLMNet) and Random Forest (RF), were trained using the SoySNP3K BeadChip markers as predictors and the MSI-based yield classes as response categories. GLMNet and RF showed strong predictive ability, with balanced accuracies of 0.70 and 0.62, respectively. Both models demonstrated high specificity (0.78 and 0.75) and very low rates of extreme misclassifications, indicating reliable discrimination between low- and high-yielding genotypes. These results suggest that classification-based genomic prediction is an effective strategy for early-stage soybean breeding, enabling more efficient resource allocation and more targeted advancement of promising genotypes.