π OPAL Consumer GPU Quantization Harness This repo contains OPAL Quants of AlexWortega/SIQ 1 35B Tier Name Target Size Middle Layer Strategy π quality ~23 GB IQ4 XS βοΈ balanced ~25 GB Q5 K π¦ compact ~18 GB Q3 K π mini ~14 GB IQ2 S π¬ micro ~11 GB IQ1 M π€ nano ~12 GB IQ2 XXS π‘οΈ exp minimalist ~16 GB IQ2 XXS (Protected Edges) π Credits π APEX Quantization Method Ettore Di Giacinto & Richard Palethorpe (LocalAI Team). OPAL adapts the layer wise precision gradients and MoE aware tensor classification outlined in the APEX technical paper. π Bartowski and Lamim For the excellent semantic imatrix calibration dataset that powers OPAL's activation scaling. π llama.cpp Georgi Gerganov and contributors for the foundational inference and quantization engine. π HuggingFace Accelerate For the init empty weights() context manager that makes the 0 RAM "Ghost Model" possible. π Jackrong For the sexy markdown readme inspo. Support the Project A coffee in Ethereum would be cool! Although I don't drink coffeeβI think it tastes like burnt waterβbut a pink lemonade would be fire! π₯ 0xDEE7fa8C421BD038D32e4441ea1aDe72fE973706
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy