GPT OSS Nano Compact Reasoning Model with Mixture of Experts 9B parameters • 12 experts • 128K context • Chain of thought reasoning 🤗 Model 📖 Docs 🔮 Q GPT 📋 Model Description GPT OSS Nano is a fine tuned Mixture of Experts (MoE) language model optimized for step by step reasoning and problem solving. Built on the GPT OSS architecture with sparse expert activation, it achieves strong reasoning performance while using only ~3B active parameters per forward pass. ✨ Key Features Feature Description 🧠 Sparse MoE 12 experts, 4 active per token — efficient compute 📝 Chain of Thought Fine tuned on reasoning datasets with step by step solutions ⚡ 128K Context Long context with YaRN rope scaling 🔮 Q GPT Ready Compatible with quantum confidence estimation 📦 GGUF Available Run locally with llama.cpp or Ollama 🏗️ Architecture 💻 Usage Quick Start with Transformers ⚡ With Unsloth (2x Faster) 📦 With GGUF (llama.cpp) 🦙 With Ollama 🎓 Training Training Details Parameter Value Base Model openai/gpt oss 20b Method QLoRA (4 bit quantized LoRA) LoRA Rank 32 LoRA Alpha 32 Learning Rate 2e 4 Batch Size 2 (gradient accumulation: 8) Epochs 2 Framework Unsloth + TRL Hardware NVIDIA H200 Dataset:…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy