BacCorpus100 BacCorpus100 is a large scale bacterial genome corpus for training and evaluating genome aware genomic language models. It contains quality controlled bacterial genomes represented at the genome level, with predicted protein coding sequences stored as protein translations and intergenic regions stored as DNA sequences. Genomes were deduplicated at 100% identity using genome sketching. BacCorpus100 spans approximately 7 million genomes, 20 billion protein coding features, 16 billion intergenic regions, more than 150,000 species, and over 10,000 environments. The dataset is intended for large scale pretraining and analysis of bacterial genomic language models. Dataset Description Repository: AllTheBacteria/BacCorpus100 Domain: bacterial genomics Data type: genome level bacterial sequence and annotation data Primary modalities: protein sequences and intergenic DNA Organisms: bacteria from isolate genomes and metagenome assembled genomes Scale: approximately 7 million quality controlled, deduplicated genomes Primary use case: pretraining and analysis of genomic language models in bacteria BacCorpus was built by combining genomes from several public bacterial genome resourc…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy