Tech Report ACE Step Transcriber Description ACE Step Transcriber is the annotation model used by ACE Step v1.5 for training data labeling. It is a powerful multilingual audio transcription model capable of transcribing both speech and singing voice with high accuracy. Key Features 🌍 50+ Languages Support Covers major world languages and regional dialects 🎤 Speech Transcription Accurately transcribes spoken content 🎵 Singing Voice Transcription Specialized in lyrics transcription with musical structure annotations 🏷️ Structure Annotation Automatically identifies song sections (verse, chorus, bridge, etc.) Usage The usage is the same as Qwen2.5 Omni 7B. Prompt Format Use the following prompt to transcribe audio: Output Format The model outputs structured content in the following format: Example Output Supported Section Tags [Intro] , [Outro] [Verse 1] , [Verse 2] , etc. [Chorus] , [Pre Chorus] , [Post Chorus] [Bridge] [Guitar Interlude] , [Instrumental] [Spoken] Supported Languages (50+) The model supports transcription in over 50 languages, including but not limited to: Region Languages East Asia Chinese (zh), Japanese (ja), Korean (ko) Southeast Asia Vietnamese (vi), Thai (th)…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy