Gravity 16B A3B Preview Gravity 16B A3B Preview is a post trained language model built on Gravity 16B A3B Base by Trillion Labs. Starting from the base model, it underwent context length extension (32K → 128K), supervised fine tuning (SFT), and reinforcement learning (GRPO) focused on science and code. This is a preview release offering a strong balance of capability, efficiency, and long context support for its size. We are actively working on agentic capabilities for the full release. Model Summary Property Value Base Model Gravity 16B A3B Base Total Parameters 16.24B Active Parameters 3.16B Architecture GravityMoE Context Length 131,072 tokens (128K) Precision bf16 License Apache 2.0 For full architectural details (MLA, MoE routing, tokenizer, etc.), see the base model card. Post Training Pipeline Starting from Gravity 16B A3B Base (pretrained on ~5.5T tokens): 1. Context Length Extension — Extended from 32K to 128K tokens. 2. Supervised Fine Tuning (SFT) — Instruction tuning for general chat and task following capabilities. 3. Reinforcement Learning (GRPO) — Single step Group Relative Policy Optimization focused on science and code domains. Agentic RL and multi turn RL stages a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy