Location: Genetics and Sustainable Agriculture Research
Title: Maize and soybean yield prediction using machine learning methods: a systematic literature reviewAuthor
![]() |
SHARMA, RAMANDEEP - Mississippi State University |
![]() |
KAUR, JASLEEN - Chitkara University |
![]() |
Feng, Guanglong |
![]() |
Huang, Yanbo |
![]() |
KUMAR, CHANDAN - Mississippi State University |
![]() |
WANG, YI - University Of Wisconsin |
![]() |
SHARMA, SANDHIR - Chitkara University |
![]() |
Jenkins, Johnie |
![]() |
DHILLON, JAGMANDEEP - Mississippi State University |
|
Submitted to: Discover Agriculture
Publication Type: Peer Reviewed Journal Publication Acceptance Date: 4/10/2025 Publication Date: 4/30/2025 Citation: Sharma, R.K., Kaur, J., Feng, G.G., Huang, Y., Kumar, C., Wang, Y., Sharma, S., Jenkins, J.N., Dhillon, J. 2025. Maize and soybean yield prediction using machine learning methods: a systematic literature review. Discover Agriculture. 64(3):1-29. https://doi.org/10.1007/s44279-025-00215-6. DOI: https://doi.org/10.1007/s44279-025-00215-6 Interpretive Summary: Improvements in crop management practices through research rely on the in-season prediction of grain yield or other crop performance metrics. As crop yield depends on many factors including weather, genetics, and management factors, hence, accurate yield prediction is challenging task. However, machine learning (ML) techniques emerged out to be a powerful tool for efficiently predicting yield to aid farmers on-farm decision making. Since there are many ML models are available to apply, it is always difficult for the agronomic researchers to select any particular model suitable for the crop data at hand. The studies reviewing the different ML models for identifying the most suitable models are scarce. Also, crop-targeted approach suggesting most suitable models, input parameters, and accuracy measures to help researcher ability of model and parameter selection are needed. Therefore, this review study has thoroughly and systematically scrutinized the literature to find top used ML model, input parameters, accuracy measures, and the challenges encountered in maize and soybean yield prediction research to guide the future research. Technical Abstract: In-season crop yield prediction is a challenging task due to the intertwined nature of numerous variables. However, today’s agronomy is data-rich, and machine learning (ML) provides the ability to efficiently predict crop yields, utilizing high-volume data to optimize agricultural decision-making. Numerous ML models are employed in yield prediction research, yet systemized know-how on its crop-targeted utilizations is lacking, specifically for soybean and maize, world’s vital crops. Henceforth, this systematic literature review (SLR) is performed to retrieve and consolidate the ML techniques and key features utilized in maize and soybean yield prediction research. Study’s search criteria utilized four electronic databases including ProQuest, Wiley, Science Direct, and EBSCOhost, totally producing 1859 related articles, which were finally reduced to 82 articles following SLR’s inclusion and exclusion criteria. Aligning with the study objectives, all papers were thoroughly analysed for generating common consensus and future research recommendations. The SLR analysis noted that ML is gaining popularity in the studied domain with a significant increase in its adoption after 2019. This study revealed the temperature, precipitation, historical crop yield, normalized difference vegetation index (NDVI), and soil pH to be the most utilized variables in ML for yield prediction research. The Random Forest (RF), Artificial Neural Networks (ANN), Support Vector Machines (SVM), and Extreme Gradient Boosting (XG-Boost) were identified as the mostly used ML algorithms. Most often applied deep learning (DL) techniques include long short-term memory (LSTM) and convolutional neural networks (CNN). In the utilized models, the most used performance assessment measures were noted as the coefficient of determination (R2), root absolute error (RAE), root mean square error (RMSE), and mean absolute error (MAE). Most applied software for building ML models includes Python, MATLAB, Weka, R, and SPSS. Altogether, there is a rising trend among ML researchers towards leveraging ensemble techniques in for the betterment of model performance and reliability. |
