This is a quantized variant of google/translategemma 12b it , created by The Kaitchup (newsletter: https://kaitchup.substack.com). More details (training recipe, benchmarks, and recommended settings) will be added later. In the meantime, here are the current notes and a working inference example. Status / limitations Quick smoke test only (not fully evaluated). RoPE parameters were removed for compatibility with vLLM . As a result, long context behavior may be degraded . I have not verified the impact yet. Chat template not supported (for now). To use the model in vLLM, call a completions endpoint and provide a fully formatted prompt . Serving with vLLM
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy