Skip to main content
ARS Home » Northeast Area » Ithaca, New York » Robert W. Holley Center for Agriculture & Health » Plant, Soil and Nutrition Research » Research » Publications at this Location » Publication #433506

Research Project: Championing Improvement of Sorghum and Other Agriculturally Important Species through Data Stewardship and Functional Dissection of Complex Traits

Location: Plant, Soil and Nutrition Research

Title: PangeneIndexer: A Scalable Framework for Consistent Genome Annotation Across Crop Pangenomes

Author
item CHOUGULE, KAPEEL - Cold Spring Harbor Laboratory
item WEI, SHARON - Cold Spring Harbor Laboratory
item LU, ZHENYUAN - Cold Spring Harbor Laboratory
item OLSON, ANDREW - Cold Spring Harbor Laboratory
item Ware, Doreen

Submitted to: CONFERENCE ON THE BIOLOGY OF GENOMES
Publication Type: Abstract Only
Publication Acceptance Date: 5/5/2026
Publication Date: 5/5/2026
Citation: Chougule, K., Wei, S., Lu, Z., Olson, A., Ware, D. 2026. PangeneIndexer: A Scalable Framework for Consistent Genome Annotation Across Crop Pangenomes. CONFERENCE ON THE BIOLOGY OF GENOMES. Conference on the Biology of Genomes.

Interpretive Summary:

Technical Abstract: As the number of high-quality crop assemblies continues to expand, the challenge of producing accurate, consistent, and scalable gene annotations across accessions has become increasingly critical. Single-reference annotations often fail to represent the full coding potential of a species, while traditional de novo pipelines, although precise, are computationally intensive and limited in sensitivity. To address this challenge, we developed PangeneIndexer, a pan-gene indexing workflow that leverages orthology-based gene family trees and representative models to efficiently propagate annotations across diverse assemblies. The workflow begins with the construction of gene family trees using the Ensembl Compara pipeline, from which representative models are selected based on curation priority and structural quality. During index construction, PangeneIndexer filters split or fused predictions and prioritizes curated models, ensuring both sensitivity and structural accuracy. These models form the foundation of a pan-gene index. Gene structures are then projected to new accessions with Liftoff and refined using PASA, which integrates transcriptomic evidence to resolve gene boundaries and splicing structures. We applied this framework to multiple crop pan-genomes, including 26 maize, 29 rice, and 18 sorghum accessions, generating pan-gene sets that partition into core, soft-core, and shell components. Compared to de novo annotations, the pan-gene index workflow demonstrated significant improvements in both coverage and uniformity, while reducing per-genome annotation time from 1–2 weeks to 2–3 days. In maize benchmarking, the pan-gene index encompassed over 116,000 pan-genes, partitioned into ~21,000 core, ~70,000 soft-core, and ~15,000 shell genes, reflecting the broad diversity across 26 accessions. Similarly, rice and sorghum pan-genomes showed tens of thousands of additional loci uncovered relative to reference-only annotations. Importantly, PangeneIndexer also recovered unique loci per accession with transcriptome or protein evidence support and also improved coverage for lineage-specific genes. To facilitate curation and evaluation, we employed the Gramene gene tree visualization platform, which enables rapid detection of inconsistent models within pan-gene families. This integration supports downstream evolutionary and functional studies while providing a scalable framework for incorporating additional accessions.