Comma v0.1 2T Comma v0.1 2T is a 7 billion parameter language model trained on 2 trillion tokens from the Comma v0.1 dataset, comprising of openly licensed text from the Common Pile. Comma v0.1 2T is a "base model" that can be used a the starting point for finetuning and post training. It performs comparably to budget matched models (7 billion parameters, 2 trillion tokens) trained on unlicensed data. Model ARC C ARC E MMLU BoolQ HSwag OBQA CSQA PIQA SIQA HEval MBPP Avg. OLMo Twin 45.2 67.5 28.2 71.7 73.4 48.0 61.8 77.9 48.5 18.2 27.5 51.6 Llama 2 48.5 69.5 45.8 80.2 76.2 48.4 62.8 76.7 50.8 26.1 28.5 55.8 Comma v0.1 2T 45.8 71.8 49.8 78.6 64.4 46.2 64.0 72.5 52.3 44.2 41.5 57.4 DeepSeekLLM 49.5 67.7 48.5 71.7 74.1 52.0 66.6 77.8 51.6 43.1 43.8 58.8 Training details Comma v0.1 2T is a decoder only transformer that uses the same architecture as Llama 3. Training was done in two stages: first on 1.93 trillion tokens with a cosine learning rate schedule, and second a "cool down" training phase on 75.5 billion tokens from high quality sources. The final model is the average of 10 checkpoints during this cool down phase. Both training phases use a batch size of 8.3 million tokens per st…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy