๐ฑ Run it on your phone or a GPU less PC โ POCKET ยท ๐ Try it live (CPU chat) VIDRAFT's on device family: a 35B model that runs on iPhone and on CPU with no GPU โ stock llama.cpp , no fork. โ ๏ธ Use batch size=1. Some attention branches (NSA family) do not consume a padding mask, so batching padded sequences can silently corrupt results. Run inference one sequence at a time (batch size=1). Aether 7B 5Attn โ Base Aether family โ Which Aether model should I use? Model What it is Pick it if Aether 7B 5Attn 6.59B MoE base. Fully open weights + data recipe + training code + all 162k step logs + checkpoints You want to audit, verify or rebuild a foundation model end to end Aether 7B 5Attn it The same model, instruction tuned You want it to answer rather than continue text AETHER 7B 7Attn base Same 49 layer architecture, a different checkpoint. Open weights You want a second run of this architecture to compare against Aether 6B 11Attn base 121 layers, 11 sequence mixing mechanisms in one network attention, Mamba 2, Hyena, GDN, MLA on an 11x11 Latin square You research heterogeneous sequence mixing. It is a mid training research artifact All four load the same way: ๐ Part of the Aether Founโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy