DNABERT 2 Pre trained model on multi species genome using a masked language modeling (MLM) objective. Disclaimer This is an UNOFFICIAL implementation of the DNABERT 2: Efficient Foundation Model and Benchmark for Multi Species Genomes by Zhihan Zhou, et al. The OFFICIAL repository of DNABERT 2 is at MAGICS LAB/DNABERT 2. [!TIP] The MultiMolecule team has confirmed that the provided model and checkpoints are producing the same intermediate representations as the original implementation. The team releasing DNABERT 2 did not write this model card for this model so this model card has been written by the MultiMolecule team. Model Details DNABERT 2 is a bert style model pre trained on a large corpus of multi species genome sequences in a self supervised fashion. This means that the model was trained on the raw nucleotides of DNA sequences only, with an automatic process to generate inputs and labels from those texts. Please refer to the Training Details section for more information on the training process. Model Specification Num Layers Hidden Size Num Heads Intermediate Size Num Parameters (M) FLOPs (G) MACs (G) Max Num Tokens 12 768 12 3072 117.07 125.83 62.92 512 Links Code : multimo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy