Location: Crop Bioprotection Research
Project Number: 5010-30400-001-001-S
Project Type: Non-Assistance Cooperative Agreement
Start Date: Jun 3, 2026
End Date: Sep 30, 2027
Objective:
The Beenome100 is an ARS driven project to generate genomes across the tree of life for bee species. The genomic data generated by the reference genomes will enable a better understanding of the biology, ecology, and evolution of known organisms, as well as the conservation, protection, and regeneration of biodiversity. Given these findings and considering the wealth of genomic data from diverse bee species that will soon be available through the Beenome100 project, this study aims to reconstruct the evolutionary history and do a genome-wide analysis of key detoxification gene families as well as the cellular stress-related genes. Additionally, this analysis will incorporate ecological and behavioral factors, such as resource-collection strategies and life history, to provide a more comprehensive understanding of how these genes have evolved.
The objective of this study is to investigate the detoxification gene families in bee species that have become available in the Beenome100 project. This will be achieved through genome-wide screening, and reconstruction of phylogenetic relationships to uncover evolutionary patterns and functional insights. The project will provide key insight into how bees adapt to stress and provide a foundation to improve bee health.
Approach:
A phylogenetic analysis will be performed using gene sequences and compared to a protein-based phylogenetic tree. Since protein-based trees are more effective for distantly related species, while genome-based trees provide higher resolution for closely related species (e.g., within the same family), this comparison will help determine the most appropriate approach. Initially, a preliminary test will be performed on a few species within the same bee family, followed by an extended analysis including all Beenome100 project species across different families to assess the consistency of the results. Also, a test with only one set of genes will be performed after the union of all genes in the same three.
A comparative analysis will be conducted to identify genes of interest, An additional similarity analysis will be performed using the protein sequences of the genes as query sequences. The most closely related sequences will then undergo conserved domain database searches to validate the presence of conserved domains. Ortholog identification will be done to further validate the selected gene sequences. Multiple software alignment tools will be tested for sequence alignment, and phylogenetic analysis will be conducted. The visualization and annotation of phylogenetic trees, including bootstrap values, will follow.
Finally, a species tree reconciliation will be performed to merge the gene or protein trees.
To access the conservation of the genome, a relative synteny analysis will be performed comparing the genes across different bee species. The intron-exon distribution will be analyzed and the conserved protein motifs will be determined.