Built for vMLX — the MLX inferencer with KV cache quantization, prefix cache reuse, agentic tool calling, and hybrid sliding+full attention support. Free for macOS · vmlx.net Gemma 4 12B it — JANG 4M CRACK CRACK abliterated · JANG mixed precision (8 bit attention, 4 bit MLP) · Omni modal (text + image + audio + video) · 9.6 GB What Is This? This is Gemma 4 12B it by Google — a unified omni modal language model (text + image + audio + video, hybrid sliding/full attention, 48 layers, 128k context) that has been: 1. CRACK abliterated — safety refusal removed at the weight level. The model now complies across all task categories instead of refusing, while keeping its knowledge, reasoning, and multimodal capabilities intact. 2. JANG mixed precision (8 bit attention, 4 bit MLP) quantized for MLX on Apple Silicon — 9.6 GB. Results Evaluated through the Osaurus runtime on a Mac Studio M3 Ultra. Compliance graded via HarmBench text refusal classifier; MMLU via logit mode argmax over A/B/C/D token logits (matched on both base and CRACK with identical chat template rendering — no answer truncated). HarmBench compliance (70 prompts · 10 per category) Category CRACK ASR : Chemical / biological…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy