MossAudioTokenizer This is the code for MOSS Audio Tokenizer presented in MOSS Audio Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models. MOSSAudioTokenizer is a unified discrete audio tokenizer based on the Cat ( C ausal A udio T okenizer with T ransformer) architecture. Scaling to 1.6 billion parameters, it functions as a unified discrete interface, delivering both lossless quality reconstruction and high level semantic alignment. Key Features: Extreme Compression & Variable Bitrate : It compresses 24kHz raw audio into a remarkably low frame rate of 12.5Hz. Utilizing a 32 layer Residual Vector Quantizer (RVQ), it supports high fidelity reconstruction across a wide range of bitrates, from 0.125kbps to 4kbps. Pure Transformer Architecture : The model features a "CNN free" homogeneous architecture built entirely from Causal Transformer blocks. With 1.6B combined parameters (Encoder + Decoder), it ensures exceptional scalability and supports low latency streaming inference. Large Scale General Audio Training : Trained on 3 million hours of diverse audio data, the model excels at encoding and reconstructing all audio domains, including speech, sound effects, and mus…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy