๐ฑ Run it on your phone or a GPU less PC โ POCKET ยท ๐ Try it live (CPU chat) VIDRAFT's on device family: a 35B model that runs on iPhone and on CPU with no GPU โ stock llama.cpp , no fork. Darwin 28B Coder โ GGUF (MTP enabled) GGUF builds of FINAL Bench/Darwin 28B Coder with the native Multi Token Prediction (MTP) head preserved, for self speculative decoding in llama.cpp . Requested in the base model discussion. Files File Quant Size Notes Darwin 28B Coder Q4 K M.gguf Q4 K M 16.8 GB recommended for most GPUs Darwin 28B Coder Q8 0.gguf Q8 0 29.0 GB near lossless Darwin 28B Coder F16.gguf F16 54.7 GB full precision All files include the MTP layer โ verified in metadata: general.architecture = qwen35 , qwen35.nextn predict layers = 1 , tensors blk.64.nextn. . Multi Token Prediction (MTP) This model ships with a trained MTP head (1 prediction layer). With a recent llama.cpp build that includes MTP support (merged in PR 22673), the nextn layer is used for self speculative decoding โ typically ~1.5โ2ร faster generation with identical output (the main model verifies every drafted token, so quality is unchanged). A standard (non MTP) GGUF does not contain the prediction head โ you need tโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy