MultiTalkBench MultiTalkBench is the first benchmark to jointly evaluate long, multi party, and bilingual full duplex dialogue. It tests speech to speech systems on: (a) Long interactions — conversations longer than ten minutes, with explicit probes for long range entity tracking and topic coherence. (b) One model many user multi party interaction — quantitative addressee selection and turn taking metrics. (c) Chinese–English bilingual ability. To our knowledge, no prior benchmark addresses these axes jointly in a fully interactive, end to end speech setting. Evaluation pipeline (judge LLM, scoring rubric, model adapters): . Splits English Chinese Total : : : samples 60 44 104 audio 32.95 h 23.57 h 56.52 h Loading To get every file on disk: metadata.jsonl fields Field Type Notes sample id str Speaker code (matches target speaker ) language "en" \ "zh" meeting id str e.g. EN2002a , R8002 M8002 target speaker str The participant the model plays audio str Relative path to FLAC alignment str Relative path to the alignment JSON for the target speaker persona str Relative path to the detailed persona prompt meeting profile str Path to meeting level profile JSON (shared by all speakers in…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy