MOSS Audio Tokenizer Nano This repository contains the Hugging Face remote code implementation and weights for MOSS Audio Tokenizer Nano , the lightweight audio tokenizer used by MOSS TTS Nano . MOSS Audio Tokenizer Nano is a compact discrete audio tokenizer based on the Cat ( C ausal A udio T okenizer with T ransformer) architecture from MOSS Audio Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models. The checkpoint in this repository has 21,969,664 parameters (approximately 22M ), making it much smaller than the full size MOSS Audio Tokenizer while preserving the 48 kHz stereo tokenizer interface used by the MOSS TTS family. Key Features Small model size : approximately 22M parameters , including about 10.45M encoder parameters, 10.45M decoder parameters, and 1.07M quantizer parameters. Native high resolution audio : supports 48 kHz input and output with 2 channel stereo audio, helping reduce compression loss and improve listening quality. Low frame rate discrete codes : compresses 48 kHz stereo audio into a 12.5 Hz token stream with a downsample rate of 7,680 samples. Variable bitrate reconstruction : uses a residual quantizer stack with 16 codebooks and 1,024…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy