Check out the NVFP4+GPTQ weights by FuriosaAI ! ➡️ link K EXAONE 236B A23B Introduction We introduce K EXAONE , a large scale multilingual language model developed by LG AI Research. Built using a Mixture of Experts architecture, K EXAONE features 236 billion total parameters, with 23 billion active during inference. Performance evaluations across various benchmarks demonstrate that K EXAONE excels in reasoning, agentic capabilities, general knowledge, multilingual understanding, and long context processing. Key Features Architecture & Efficiency: Features a 236B fine grained MoE design (23B active) optimized with Multi Token Prediction (MTP) , enabling self speculative decoding that boosts inference throughput by approximately 1.5x. Long Context Capabilities: Natively supports a 256K context window , utilizing a 3:1 hybrid attention scheme with a 128 token sliding window to significantly minimize memory usage during long document processing. Multilingual Support: Covers 6 languages: Korean, English, Spanish, German, Japanese, and Vietnamese. Features a redesigned 150k vocabulary with SuperBPE , improving token efficiency by ~30%. Agentic Capabilities: Demonstrates superior tool us…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy