Parakeet EOU 120M CoreML INT8 Streaming speech recognition with end of utterance detection, converted to CoreML for Apple Neural Engine inference. Part of speech swift — on device speech AI for Apple Silicon. Based on nvidia/parakeet realtime eou 120m v1 (FastConformer RNNT, cache aware streaming). Quick Start Or via CLI: Model Property Value Parameters 120M Architecture FastConformer RNNT (17 layer encoder, 1 layer LSTM decoder) Format CoreML (.mlmodelc) Quantization INT8 palettization (encoder) Vocabulary 1024 BPE + EOU + EOB + blank (1027 total) Sample rate 16 kHz Streaming chunk 320ms (configurable) Files File Size Description encoder.mlmodelc 102 MB Cache aware FastConformer encoder (INT8) decoder.mlmodelc 7.5 MB 1 layer LSTM prediction network joint.mlmodelc 2.7 MB RNNT joint network (1027 outputs) config.json token (ID 1024) emitted by the joint network. No external VAD required for utterance segmentation. Source Converted from nvidia/parakeet realtime eou 120m v1 using coremltools 8.3 with INT8 palettization. Links speech swift Apple SDK soniqo.audio Website blog.ivan.digital Blog Guide : soniqo.audio/guides/dictate Docs : soniqo.audio GitHub : soniqo/speech swift
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy