Location: Soil Dynamics Research
Title: Content-free computation for agricultural AI: Distributed multimodal representation learning with Qwen2.5-VL and AgroGPT on Stampede3Author
![]() |
MANJUNATHA, H - The University Of Texas At Dallas |
![]() |
SUNDARAVADIVEL, P - University Of Texas At Tyler |
![]() |
Torbert Iii, Henry |
![]() |
TAMIL, L - The University Of Texas At Dallas |
|
Submitted to: Meeting Abstract
Publication Type: Proceedings Publication Acceptance Date: 4/15/2026 Publication Date: N/A Citation: N/A Interpretive Summary: The bottleneck in precision agriculture has quietly shifted. Low-altitude UAV platforms now generate field imagery at densities that outpace interpretive capacity. This research introduces a content-free distributed agricultural intelligence framework built and evaluated on Stampede3. Annotation-free visual structure discovery from UAV imagery is fused with structured agronomic metadata and routed through locally executed vision-language models. The reasoning architecture centers on two open-source deployments: Qwen2.5-VL (3B–7B), a natively multimodal model that co-encodes image tile representations alongside agronomic context without serial transcription overhead, and AgroGPT (~7B). The broader implications are direct: eliminating annotation dependency and maintaining fully on-premise inference substantially lowers entry barriers for smallholder farms, making precision agricultural AI a practical reality rather than a resource-gated aspiration. Technical Abstract: The bottleneck in precision agriculture has quietly shifted. Low-altitude UAV platforms now generate field imagery at densities that outpace interpretive capacity. This research introduces a content-free distributed agricultural intelligence framework built and evaluated on Stampede3. Annotation-free visual structure discovery from UAV imagery is fused with structured agronomic metadata and routed through locally executed vision-language models. The reasoning architecture centers on two open-source deployments: Qwen2.5-VL (3B–7B), a natively multimodal model that co-encodes image tile representations alongside agronomic context without serial transcription overhead, and AgroGPT (~7B). Under SLURM-managed mixed data, model, and pipeline parallelism on Stampede3, parallel efficiency declines from near-linear at low worker counts to approximately 53% at 32 GPUs, with communication overhead constituting roughly 39% of total runtime. The broader implications are direct: eliminating annotation dependency and maintaining fully on-premise inference substantially lowers entry barriers for smallholder farms, cooperative extension networks, and under-resourced research institutions that lack both labeling infrastructure and cloud API access making precision agricultural AI a practical reality rather than a resource-gated aspiration. |
