Quasar 10B: Fully Linear Foundation Model Quasar 10B is a high performance foundation model developed by SILX AI . It is built upon the Qwen3.5 9B Base architecture, fundamentally re engineered to support extreme long context reasoning (2 Million+ tokens) while maintaining high computational efficiency. This model marks a major shift in the Quasar training stack, moving from traditional Softmax based attention to a Hybrid Gated Linear Attention (GLA) architecture. Model Overview Model Name: Quasar 10B Organization: SILX AI Base Model: Qwen3.5 9B Base Architecture Evolution The original Qwen3.5 architecture uses a combination of Gated Delta Attention and Softmax Gated Attention. To support the Quasar design requirements for infinite scaling and efficient state management, we performed a deep architectural swap: GLA Integration : Replaced the target attention layers with Gated Linear Attention (GLA) . NOPE (No Positional Embeddings) : Removed traditional RoPE (Rotary Positional Embeddings) to eliminate positional bias and enable native extrapolation to millions of tokens. [!NOTE] GLA was chosen as the core linear mechanism to maintain exact architectural parity with the Quasar 22B Mo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy