Note : For full duplex (real time) inference, use the 8 bit variant instead. 4 bit quantization degrades PersonaPlex response quality significantly — INT8 is both 30% faster (112ms vs 158ms/step) and produces coherent responses where INT4 generates garbled output. PersonaPlex 7B MLX 4 bit PersonaPlex 7B full duplex speech to speech model converted to MLX safetensors with 4 bit quantization for Apple Silicon. Converted from nvidia/personaplex 7b v1 (based on Kyutai Moshi architecture). Swift inference : soniqo/speech swift Library Docs : soniqo.audio Blog : PersonaPlex on Apple Silicon — Full Duplex Speech to Speech in Native Swift with MLX Model Details Component Architecture Size Temporal Transformer 32 layer, 4096d, 32 heads (7B params) ~3.5 GB (4 bit) Depformer 6 layer, 1024d, 16 heads, per codebook weights ~50 MB (fp16) Mimi Codec SEANet encoder/decoder + 8L transformer + 16 RVQ codebooks ~370 MB (fp16) Embeddings Text + 16 audio embeddings + output heads ~940 MB (fp16) Total ~4.9 GB Architecture Voices 18 voice presets available: Category Voices Natural Female NATF0, NATF1, NATF2, NATF3 Natural Male NATM0, NATM1, NATM2, NATM3 Variety Female VARF0, VARF1, VARF2, VARF3, VARF4 Va…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy