Behemoth X 123B v2.1 FP8 Dynamic Quantization This is an FP8 quantized version of TheDrummer/Behemoth X 123B v2.1 using llmcompressor with the FP8 DYNAMIC scheme. Model Details Base Model : TheDrummer/Behemoth X 123B v2.1 Quantization : FP8 DYNAMIC (W8A8) Format : compressed tensors (SafeTensors) Memory : ~50% of original BF16 size Quality : <1 2% degradation on benchmarks (typical) Quick Start vLLM (Recommended) Transformers Quantization Details This model was quantized using: Tool : llmcompressor Method : FP8 DYNAMIC (Round to Nearest) Targets : All Linear layers except lm head Scheme : W8A8 (8 bit weights and activations) Performance Memory Usage Original BF16 : ~2× size of FP8 FP8 Quantized : ~50% of original Savings : ~50% VRAM reduction Inference Speed Expect 1.3 1.8× faster inference vs BF16 2× higher throughput (more KV cache available) Use Cases Perfect for: ✅ Production inference on limited VRAM ✅ Running larger models on single GPU ✅ Cost effective API serving ✅ High throughput applications ✅ Extended context lengths (more KV cache) Hardware Requirements Minimum VRAM (approximate): 70B model: ~40 GB (RTX A6000, A100 40GB) 123B model: ~70 GB (A100 80GB, H100, H200) Recomm…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy