VCB Bench: An Evaluation Benchmark for Audio Grounded Large Language Model Conversational Agents Introduction Voice Chat Bot Bench (VCB Bench) is a high quality Chinese benchmark built entirely on real human speech. It evaluates large audio language models (LALMs) along three complementary dimensions: (1) Instruction following : Text Instruction Following (TIF), Speech Instruction Following (SIF), English Text Instruction Following (TIF En), English Speech Instruction Following (SIF En) and Multi turn Dialog (MTD); (2) Knowledge : General Knowledge (GK), Mathematical Logic (ML), Discourse Comprehension (DC) and Story Continuation (SC). (3) Robustness : Speaker Variations (SV), Environmental Variations (EV), and Content Variations (CV). Getting Started Installation: Note: To evaluate Qwen3 omni, please replace it with the environment it requires. Download Dataset: Download the dataset from Hugging Face and place the 'vcb bench' into 'data/downloaded datasets'. Evaluation: This code is adapted from Kimi Audio Evalkit, where you can find more details about the evaluation commands. (1) Inference + Evaluation: For example: (2) Only Inference: For example: (3) Only Evaluation: For exampl…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy