Technical Report 👁️ 1. Introduction Nanbeige4.2 3B is a compact agentic model built on Nanbeige4.2 3B Base, designed to combine strong agentic behavior with broad reasoning and alignment capabilities. Its Looped Transformer architecture reuses the transformer layers to increase model capacity without adding parameters. With only 3B non embedding parameters, the model delivers solid performance on general agent and code agent tasks. During supervised fine tuning (SFT), we expand the diversity of training environments through real world environment integrations and large scale environment synthesis. We further diversify task types, task assets, and the agentic scaffolds used for each task. To ensure training data quality, we apply filtering at both the trajectory and turn levels, combining test case based validation with rubric based assessment. During reinforcement learning (RL), we combine outcome and process rewards to improve training stability for the compact model. Key strengths include: Solid Agentic Behavior at the 3B Scale : Across complex tool use, office agent, and code agent benchmarks, Nanbeige4.2 3B outperforms larger models such as Qwen3.5 9B and Gemma4 12B. Strong Re…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy