🧬 Carbon Pretraining Corpus Description 173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model. This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species. Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon's 6 mer tokenizer). A pre sampled 10B token eukaryote subset ( eukaryote generator 10B subset ) is also provided for smaller / faster runs. Subset Domain Source Size Rows Nucleotides : : : eukaryote generator Eukaryote genomes (≤ 100 kbp) GenerTeam / GENERATOR 191.9 GB 46,323,396 423.4 Gbp mrna evo2 Messenger RNA Arc Institute / OpenGenome2 54.8 GB 52,702,454 115.9 Gbp mrna splice evo2 mRNA + splice & promoter Arc Institute / OpenGenome2 92.9 GB 56,877,762 197.4 Gbp prokaryote evo2 Prokaryote genomes (GTDB + IMG/PR) Arc Institute / OpenGenome2 166.0 GB 17,408,059 357.5 Gbp eukaryote generator 10B subset Eukaryote subsample (10B tokens, natural species distribution) derived from eukaryote generator 27.6 GB 6,562,876 60.0 Gbp Nucleotides are count…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy