MOSS Audio MOSS Audio is an open source audio understanding model from MOSI.AI, the OpenMOSS team, and Shanghai Innovation Institute. It performs unified modeling over complex real world audio, supporting speech understanding, environmental sound understanding, music understanding, audio captioning, time aware QA, and complex reasoning . In this release, we provide four models : MOSS Audio 4B Instruct , MOSS Audio 4B Thinking , MOSS Audio 8B Instruct , and MOSS Audio 8B Thinking . The Instruct variants are optimized for direct instruction following, while the Thinking variants provide stronger chain of thought reasoning capabilities. News 2026.4.13: 🎉🎉🎉 We have released MOSS Audio. Blog and paper coming soon! Contents Introduction Model Architecture DeepStack Cross Layer Feature Injection Time Aware Representation Released Models Evaluation Quickstart Environment Setup Basic Usage Gradio App SGLang Serving More Information Citation Introduction Understanding audio requires more than simply transcribing words — it demands the ability to perceive acoustic cues, recognize speakers and emotions, interpret environmental sounds, reason over temporal context, and handle complex multi s…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy