# Search database composition — AllOralsDB.v2026.215_A

*Auto-generated from build artifacts.*

## Methods summary

The AllOralsDB.v2026.215_A search database comprised 10,190,142 protein sequence(s) assembled with the maniFasta pipeline on 2026-08-03. Sequences were drawn from 12 curated sources spanning 8 categories (bacteria_archaea, contaminants, food, fungi, human, microeuks, obelisks, viruses). All publicly sourced sequences were retrieved on that build date. Across all sources the database represents 559 distinct NCBI taxa spanning 33 phyla, 111 families, 184 genera, 466 species. HOMD genomes were restricted to the oral body site(s) using the HOMD taxonomy table prior to any redundancy reduction, while non-HOMD sources were curated inclusion lists whose taxa were selected by the criteria noted per source below (e.g., cultivation or pathogenicity evidence, tissue tropism, or literature curation) rather than by automated taxonomic sweeps.

HOMD genomes were included without taxonomic dereplication (one proteome per assembly passing the site filter).

HOMD proteins carried PROKKA functional annotations as distributed by HOMD (release V11.03), with product descriptions and locus tags retained in the FASTA headers, while sequences contributed as curated FASTAs or accession lists retained their source-provided functional annotations. Taxonomic annotation was assigned by resolving each sequence's NCBI taxon ID to a full ranked lineage (NCBI Taxonomy) during the lineage-enrichment step. Common contaminant proteins were incorporated from the cRAP collection (set: ccp; 10.5281/ZENODO.15115102), contributing 125 sequence(s) merged into the final database alongside the biological sources.

## Origin of sequences

| Source | Category | Module / collection | Version | Sequences | Citation |
|---|---|---|---|---|---|
| FUNGI | fungi | mod_I | — | 106,156 | — |
| ENTAMOEBA | microeuks | mod_I | — | 14,338 | — |
| VIRUSES_HUMAN | viruses | mod_I | — | 4,300 | — |
| VIRUSES_FUNGAL_MICROEUK | viruses | mod_I | — | 114 | — |
| VIRUSES_DIETARY | viruses | mod_I | — | 66 | — |
| VIRUSES_HERVS | viruses | mod_II | — | 13 | — |
| TTENAX | microeuks | mod_III | Mpeyako_etal_2024 | 20,286 | 10.3389/fmicb.2024.1437572 |
| OBELISKS_ORAL | obelisks | mod_III | Zheludev_etal_2024 | 27 | 10.25740/WB363NT3637 |
| DAIRY | food | mod_III | Hendy_2019 | 151 | 10.15124/589742eb-287a-4576-a00a-30df33d9f52c |
| HUMAN_CI | human | human_uniprot | — | 42,547 | — |
| HOMD_ORAL_A | bacteria_archaea | HOMD | V11.03 | 10,002,019 | homd.org |
| CRAP | contaminants | cRAP | — | 125 | 10.5281/ZENODO.15115102 |

Source-specific provenance:

- **FUNGI** — Fungi detected in the oral samples.
- **ENTAMOEBA** — Proxy for Entamoeba gingivalis; E. gingivalis genome unavailable in NCBI therefore include genomes of other Entamoeba species.
- **VIRUSES_HUMAN** — Viruses that infect humans and have been detected in oral samples (sources include data from https://viralzone.expasy.org/ and the Human Virus Database http://computationalbiology.cn/humanVirusBase/).
- **VIRUSES_FUNGAL_MICROEUK** — Viruses that infect fungi and other microeukaryotes (sources Kinsella et al. and Keeler et al.)
- **VIRUSES_DIETARY** — Viruses infecting plants and tobacco products, detected in oral samples (sources include Aguado-García et al., Rivera-Gutierrez 2023 et al., and literature survey).
- **VIRUSES_HERVS** — Viruses that are endogenous in the human genome (HERVs).
- **TTENAX** — Proteins identified and provided by Mpeyako et al. as Suppl Table 2 and Suppl Data 3).
- **OBELISKS_ORAL** — Subset of proteins from protein calls for Obelisk genomes identified by Zheludev (their Supp Table 2), filtered and reviewed to include only Obelisk proteins from oral/oral-proximal datasets (27 of 11,581 pyrodigal ORFs).
- **DAIRY** — Proteins provided by Hendy et al. as curated dairy proteins, developed for their study of ancient dental calculus, Wilkin et al. 2020.
- **HOMD_ORAL_A** — HOMD PROKKA proteomes filtered to oral body site
- **CRAP** — Cambridge Centre for Proteomics cRAP set.
