MiDashengLM 7B 0804 A general audio captioner MiDashengLM is an efficient audio language model that achieves holistic audio understanding through caption based alignment. It achieves state of the art performance on multiple audio understanding benchmarks while maintaining high inference efficiency—delivering 3.2× throughput speedup and supporting batch sizes up to 512. 📖 For more detailed introduction and technical report, please visit our GitHub repository. Note that for most applications, we strongly recommend using the BF16 version (mispeech/midashenglm 7b 0804 bf16) for optimal performance and efficiency. Usage Load Model Construct Prompt Generate Output Results The following evaluation results are based on the model version: mispeech/midashenglm 7b 0804 fp32 . Audio Captioning Results Domain Dataset MiDashengLM Qwen2.5 Omni 7B Kimi Audio Instruct : : : : : : : : : : Music MusicCaps 59.71 43.71 35.43 Music Songdescriber 45.39 45.31 44.63 Sound AudioCaps 62.18 60.79 49.00 Sound ClothoV2 49.20 47.55 48.01 Sound AutoACD 66.52 55.93 44.76 Metrics: FENSE (higher is better). Audio and Paralinguistic Classification Dataset Metric MiDashengLM Qwen2.5 Omni 7B Kimi Audio Instruct : : :…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy