[iMetaOmics] ExMODE (https://db.genomics.cn/exmode/) integrates 3518 samples to build a unified extremophile resource, which hosts 1.35 billion habitat-specific non-redundant genes, 5.25 million representative protein structures, 67,026 metagenome-assembled genomes (MAGs), and 164,132 biosynthetic gene clusters (BGCs).

By combining sequence-, structure-, and genome-based exploration across five extreme environments, it enables functional characterization of microbial dark matter through integrated comparative analyses and structure-guided searches. A cross-habitat microviridin case study demonstrates that ExMODE is an effective platform for discovering novel BGCs and biotechnologically valuable natural products.

To the Editor – Conditions such as extreme thermal gradients, hypersalinity, and high hydrostatic pressure render most habitats inhospitable to life, yet these very stressors define extreme environments. Despite these harsh physicochemical conditions, habitats including glaciers, hydrothermal vents, and acid mine drainage support specialized microbial communities in which extremophiles have evolved unique adaptive mechanisms [1].

The study of extremophiles has expanded our understanding of the limits of life and uncovered valuable biotechnological resources [2, 3]. Their unique enzymes and bioactive compounds have enabled advances in industrial biotechnology and drug discovery [3]. Notable examples include Taq polymerase [4], which transformed molecular biology, and recently identified deep-sea polyethylene terephthalate (PET) hydrolases that highlight the immense bioprospecting potential of extreme ecosystems [5].

Despite rapid advances in metagenomics and related microbiome sequencing approaches, several major challenges remain. Publicly available datasets are scattered across repositories with heterogeneous metadata and annotation standards, hindering large-scale integration and comparative analyses [6].

Existing databases are generally designed either for broad ecosystem surveys or for a single type of extreme environment, limiting systematic cross-environment investigations of microbial adaptation [7, 8]. Moreover, most resources remain sequence-based and lack complementary structural information, restricting the functional interpretation of microbial dark matter and the discovery of novel biomolecules.

To address these limitations, we developed Extremophile Multi-Omics DatabasE (ExMODE, https://db.genomics.cn/exmode/), a comprehensive resource that combines diverse microbial datasets with standardized functional annotations and predicted protein structures.

ExMODE harmonizes data across five major extreme environments and integrates genomes, genes, representative protein structures, and biosynthetic gene clusters (BGCs) into a unified platform, thereby enabling systematic cross-habitat comparisons. This multi-layered resource supports the exploration of extremophile adaptation and provides a foundation for functional discovery and bioprospecting from microbial dark matter.

Comprehensive profiling of genomic, functional, structural, and biosynthetic features of microbial genetic resources from five extreme biomes. (A) Standardized data processing and curation pipeline of the ExMODE platform. The workflow illustrates the systematic progression from environmental sequence acquisition to final multi-dimensional repository storage across four consecutive modules. (B) Functional annotation breakdown of the non-redundant gene sets. Values in each concentric donut chart represent the total unigenes (in millions, M) clustered at a 95% sequence similarity threshold from each extreme biome. Rings from inner to outer sequentially show the annotation proportions against CAZy, KEGG, eggNOG, and CARD databases, with “Unknown” indicating unigenes unannotated in any database. (C) Distribution of the representative protein structure repository across five biomes. Protein structures predicted via ESMFold are stratified into three distinct quality control layers based on their pLDDT (predicted Local Distance Difference Test) and pTM (predicted TM-score) metrics: High-confidence (pLDDT > 0.7 and pTM > 0.7), Good-confidence (pLDDT > 0.5 and pTM > 0.5), and Low-confidence (pLDDT ≤ 0.5 or pTM ≤ 0.5). (D) Taxonomic profiling at the species level for the 67,026 recovered MAGs using GTDB-Tk. MAGs assigned to named species (“s__xxx”) are shown as “Known,” whereas MAGs lacking species-level assignments (“s__”) in the current GTDB reference database are shown as “Novel.” Taxonomic breakdown of the recovered MAG repository, showing the distribution of the top 10 dominant bacterial (E) and archaeal (F) phyla across biomes, with all minority lineages consolidated under “Other.” (G) Total counts (×103) of biosynthetic gene clusters (BGCs) recovered from each of the five extreme biomes, with colors distinguishing different BGC classes. NRPS, nonribosomal peptide synthetase; PKS, polyketide synthase; RiPP, ribosomally synthesizied and post-translationally modified peptide. (H) Alluvial diagram visualizing the cross-linking of habitat types, recovered MAG phyla, and predicted BGC classes to show the biosynthetic and taxonomic context within the repository, with phylum-level taxa containing fewer than 1000 BGCs consolidated under “Other.” — iMetaOmics

Astrobiology, Genomics,

Explorers Club Fellow, ex-NASA Space Station Payload manager/space biologist, Away Teams, Journalist, Lapsed climber, Synaesthete, Na’Vi-Jedi-Freman-Buddhist-mix, ASL, Devon Island and Everest Base Camp...

Leave a comment

Your email address will not be published. Required fields are marked *