Kokoro 82M CoreML End to end CoreML export of hexgrad/Kokoro 82M at FP16, optimized for Apple Neural Engine. Requires iOS 18+ / macOS 15+. A single kokoro 5s.mlmodelc runs the full pipeline (BERT → duration prediction → fixed shape alignment → prosody → decoder) in one CoreML call. G2P (grapheme to phoneme) is a separate pair of CoreML models. Looking for a smaller variant? See aufklarer/Kokoro 82M CoreML INT8 — INT8 k means palettized, 83 MB vs 325 MB here, with log spec distance 0.42 vs this FP16 reference on a validation utterance. Model Parameter Value Parameters 82M Precision FP16 Max audio length 5 s (200 frames @ 40 fps) Sample rate 24 kHz Style dimension 256 Max phonemes per pass 128 Files File Size Description kokoro 5s.mlmodelc 325 MB Pre compiled E2E model (pre compiled, loads directly on device) G2PEncoder.mlmodelc 0.7 MB Grapheme to phoneme encoder G2PDecoder.mlmodelc 0.8 MB Grapheme to phoneme decoder voices/ 0.5 MB 54 preset voice embeddings (10 languages) vocab index.json 4 KB Phoneme vocabulary g2p vocab.json 4 KB G2P vocabulary us gold.json , us silver.json 6 MB English pronunciation dictionaries pipeline config.json 4 KB Swift pipeline config Voices 54 preset voi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy