gpt oss puzzle 88B Model Overview Description gpt oss puzzle 88B is a deployment optimized large language model developed by NVIDIA, derived from OpenAI's gpt oss 120b. The model is produced using Puzzle, a post training neural architecture search (NAS) framework, with the goal of significantly improving inference efficiency for reasoning heavy workloads while maintaining or improving accuracy across reasoning budgets. The model is specifically optimized for long context and short context serving on NVIDIA H100 class hardware, where reasoning models are often bottlenecked by KV cache bandwidth and memory capacity rather than raw compute. Compared to its parent, gpt oss puzzle 88B: Reduces total parameters to ~88B (≈73% of the parent), Achieves 1.63× throughput improvement in long context (64K/64K) scenarios on an 8×H100 node, Achieves 1.22× throughput improvement in short context (4K/4K) scenarios, Delivers up to 2.82× throughput improvement on a single H100 GPU, Matches or slightly exceeds parent accuracy across reasoning efforts. Parameter count note. Hugging Face Hub may automatically show this model as ~91B parameters. We refer to it as 88B because the automatic count includes…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy