Step Audio EditX           ๐ฅ๐ฅ๐ฅ News!!๏ผ Jan 23, 2026: ๐ Training and inference for vLLM are now supported. Thanks to the vLLM team! Jan 23, 2026: ๐ป We release the GRPO training code. Jan 23, 2026: ๐งฉ New Model Release: Now supporting more paralinguistic tags. Nov 28, 2025: ๐ New Model Release: Now supporting Japanese and Korean languages. Nov 23, 2025: ๐ Step Audio Edit Benchmark Released! Nov 19, 2025: โ๏ธ We release a new version of our model, which supports polyphonic pronunciation control and improves the performance of emotion, speaking style, and paralinguistic editing. Nov 12, 2025: ๐ฆ We release the optimized inference code and model weights of Step Audio EditX (HuggingFace; ModelScope) and Step Audio Tokenizer (HuggingFace; ModelScope) Nov 07, 2025: โจ Demo Page ; ๐ฎ HF Space Playground Nov 06, 2025: ๐ We release the technical report of Step Audio EditX. Introduction We are open sourcing Step Audio EditX, a powerful 3B parameter LLM based Reinforcement Learning audio model specialized in expressive and iterative audio editing. It excels at editing emotion, speaking style, and paralinguistics, and also features robust zero shot text to speech (โฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy