HRM Text 1B GGUF This repository contains a BF16 GGUF conversion of sapientinc/HRM Text 1B and validated Q8 0 , Q6 K , and Q5 K M quantizations derived from that BF16 GGUF. The GGUF files use: general.architecture = hrm text BF16 source tensor storage or standard llama.cpp quantized tensor storage the original tokenizer from tokenizer.json no injected chat template This is not a chat model and is not instruction tuned. "Useful output" for this repository means alignment with the original Transformers model on the same prompt, not chat assistant behavior. Compatibility Notice Standard upstream llama.cpp , Ollama, LM Studio, and llama cpp python are expected not to load this file until hrm text is supported upstream. Use the included patch: The patch was built against: Only the normal causal generation path is implemented in the patched runtime. Prefix LM bidirectional token type ids are not supported by the llama.cpp path in this release. Files File Description HRM Text 1B BF16.gguf BF16 GGUF conversion of sapientinc/HRM Text 1B HRM Text 1B Q8 0.gguf Validated Q8 0 quantization from BF16 HRM Text 1B Q6 K.gguf Validated Q6 K quantization from BF16 HRM Text 1B Q5 K M.gguf Validated Q5…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy