Kimi Audio π€ Kimi Audio 7B π€ Kimi Audio 7B Instruct π Paper Introduction We present Kimi Audio, an open source audio foundation model excelling in audio understanding, generation, and conversation . This repository hosts the model checkpoints for Kimi Audio 7B Instruct. Kimi Audio is designed as a universal audio foundation model capable of handling a wide variety of audio processing tasks within a single unified framework. Key features include: Universal Capabilities: Handles diverse tasks like speech recognition (ASR), audio question answering (AQA), audio captioning (AAC), speech emotion recognition (SER), sound event/scene classification (SEC/ASC) and end to end speech conversation. State of the Art Performance: Achieves SOTA results on numerous audio benchmarks (see our Technical Report). Large Scale Pre training: Pre trained on over 13 million hours of diverse audio data (speech, music, sounds) and text data. Novel Architecture: Employs a hybrid audio input (continuous acoustic + discrete semantic tokens) and an LLM core with parallel heads for text and audio token generation. Efficient Inference: Features a chunk wise streaming detokenizer based on flow matcβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy