GigaChat Audio 10B (A1.8B) GigaChat Audio 10B is an audio native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture of Experts decoder, so the model keeps the text quality of its base while adding speech understanding. Capabilities: audio question answering and classification, temporal grounding (localization in long audio, timestamped event descriptions, audio summarization with timestamps), tool use, and text only tasks. The temporal grounding skills are trained on TimeGround 1M — a purpose built dataset of long form audio paired with time aligned annotations. Evaluation 1. Core audio tasks vs open models Task Set Metric GigaChat Audio (10B A1.8B) Voxtral (3B) Phi 4 (4B) Qwen3 Omni (30B A3B) : : : : Audio QA MMAU acc ↑ 62.2 59.8 68.3 74.7 Audio QA MMLU speech acc ↑ 50.3 38.8 35.1 72.2 Audio math MQA acc ↑ 72.5 35.3 42.0 86.7 Audio QA (ru) RuBQ acc ↑ 60.0 23.4 2.3 43.7 Temporal Localization ≤10m mIoU ↑ 40.3 3.4 0.2 12.9 Temporal Localization 20–60m mIoU ↑ 48.3 0.1 0.2 0.1 Emotion Dusha crowd acc ↑ 90.0 43.9 11.4 77.2 Emotion Dusha podcast acc ↑ 92.4 79.6 7.2 80.7 ASR (ru) Golos…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy