Location: Coastal Plain Soil, Water and Plant Conservation Research
Title: A framework for early-season crop yield prediction using machine learning and drought conditions in the Southeastern United StatesAuthor
![]() |
Sohoulande Djebou, Dagbegnon |
![]() |
KHEDUN, PRAKASH - Clemson University |
|
Submitted to: Farming System
Publication Type: Peer Reviewed Journal Publication Acceptance Date: 3/21/2026 Publication Date: 3/22/2026 Citation: Sohoulande Djebou, D.C., Khedun, P. 2026. A framework for early-season crop yield prediction using machine learning and drought conditions in the Southeastern United States. Farming System. Article 100224. https://doi.org/10.1016/j.farsys.2026.100224. DOI: https://doi.org/10.1016/j.farsys.2026.100224 Interpretive Summary: Drought is a major hazard with significant impacts on agriculture. Under climate change drought events are expected to increase in frequency, severity, duration and propagation. However, drought is very unpredictable and the link between drought conditions and crop yield is still poorly understood. Particularly, in a humid region such as the southeastern coastal plain of the United States (US) where rainfed agriculture is dominant, a better understanding of crop responses to drought could be a levee for crop yield predictions. Hence, this study aims to enlighten the implications of drought on crops by proposing a modeling framework for predicting corn, cotton, peanuts, and soybeans yields based on early-season drought conditions. The study considered a region of the US southeastern coastal plain including the States of Georgia, North Carolina, and South Carolina, then used county-level data of drought and crop yields over the period 1990 to 2018. Explicitly, a transformation was used to develop a classification scheme that distinguishes gradual levels of crop yields. Supervised Machine Learning models were applied distinctly to evaluate yield predictability based on drought. Results showed a disparity of performances depending on the model, the crop, and the prediction scales. For instance, the accuracy of predicting crop yields at regional and state scales was up to 77% with the models. However, the recommendation of a model may consider a trade-off between accuracy and complexity. Overall, the prediction accuracies were substantially high when the models were set for predicting three or two levels of yield anomalies. The model predictions could be useful for anticipating crop yield anomalies at the regional scale. Technical Abstract: Drought is a major hydroclimatic hazard with significant impacts on agriculture. Under climate change drought events are expected to increase in frequency, severity, duration and propagation. However, drought is very unpredictable and the tandem between drought conditions and crop yield is still poorly understood. Particularly, in a humid region such as the southeastern coastal plain of the United States (US) where rainfed agriculture is dominant, a better understanding of crop responses to drought could be a levee for crop yield predictions. Hence, this study aims to enlighten the implications of drought on crops by proposing a modeling framework for predicting corn, cotton, peanuts, and soybeans yields based on early-season drought conditions. The study considered a region of the US southeastern coastal plain including the States of Georgia, North Carolina, and South Carolina, then used county-level time series of monthly standardized precipitation and evapotranspiration index (SPEI) and detrended crop yields over the period 1990 to 2018. Explicitly, a z-score transformation was used to develop a classification scheme that distinguishes gradual levels of crop yields. Supervised Machine Learning models including the random forests (RF) and the k-nearest neighbors (kNN) were applied distinctly to evaluate yield predictability based on SPEI. Results showed a disparity of performances depending on the model, the crop, and the prediction scales. For instance, the accuracy of predicting crop yields at regional and state scales was up to 74% with RF and 77% with kNN models. This slight outperformance of the kNN models was consistent at all scales for each of the four crops indicating that the kNN modeling framework was more suitable. However, the structure of the structure of kNN is more complex compared to RF, and the recommendation of a model may consider a trade-off between accuracy and complexity. Overall, the prediction accuracies were substantially high when both RF and kNN models were set for predicting three or two levels of yield anomalies. The model predictions could be useful for anticipating crop yield anomalies at the regional scale. |
