GPT OSS 120B Fable 5 Distilled — GGUF Trained & Converted by : cloudyu Compute Sponsor : AutoTrust AI Lab Source Model : cloudyu/gpt oss 120b Fable 5 Distilled Architecture : GPT OSS (OpenAI MoE, 36 layers, 128 experts, 4 active) Quantization : Q8 0 (115.7 GB) Q5 0 (67 GB) — GGUF V3 for llama.cpp GGUF is the universal format for LLM inference. These files run on llama.cpp, LM Studio, Ollama, GPT4All, text generation webui, llamafile, MLX LM, and any GGUF compatible runtime — across macOS, Windows, Linux, iOS, and Android, on CPU, CUDA, Metal, Vulkan, and ROCm backends. Model Overview This is a distilled variant of the OpenAI gpt oss 120b model, fine tuned using MLX with LoRA adapters (rank=16, targeting all attention projections, MoE router, and expert FFN layers). The training was performed in the MLX ecosystem using the MXFP4 quantized base model format. These GGUF files are the first community produced GGUF conversions of a LoRA fine tuned GPT OSS model from the MLX format. Observed Results Model Quant HumanEval Pass@1 Fable 5 Distilled Q5 0 99.39% (163/164) Fable 5 Distilled Q8 0 98.78% (162/164) gpt oss 120b (base) MXFP4 83.54% (137/164) Available Quantizations Format File Siz…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy