TriLM 1.5B — TQ2 0 GGUF for IBM Power TriLM 1.5B (Spectra Suite, ternary { 1, 0, +1} weights) quantized to TQ2 0 (2.06 bits per weight) for fast CPU inference with llama.cpp — optimized for IBM POWER9 and later with the LibrePower VSX ternary kernels. File : TriLM 1.5B TQ2 0.gguf (716 MiB, 1.52 B parameters) Quantization : TQ2 0 — exact ternary, no quality loss vs the original ternary weights No GPU required. Run it on IBM Power (ppc64le) Works with any recent llama.cpp on any architecture; the LibrePower build adds VSX acceleration on Power (4.3x prompt / 2.4x generation vs the generic path). AIX / big endian TriLM 1.5B TQ2 0 be.gguf is the big endian variant for IBM AIX ( dnf install llama aix from aix.librepower.org): Performance (IBM POWER9, Ubuntu 22.04, llama bench) Test Threads Tokens/s Prompt processing (pp64) 16 122.1 Generation (tg32) 16 33.7 Note: this is a base model (no instruction tuning) — use completion style prompts, not chat. Credits Model: Spectra Suite (Apache 2.0) TQ2 0 quantization & Power packaging: LibrePower
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy