ZAYA1 8B Checkpoint has been reshaped for transformers PR, special thank you to @JJJYmmm. The original checkpoint is at Zyphra/ZAYA1 8B legacy ZAYA1 8B is a small mixture of experts language model with 760M active parameters and 8.4B total parameters trained end to end by Zyphra. ZAYA1 8B sets a new standard of intelligence efficiency for its parameter count through a combination of novel architecture and innovations in pretraining and post training. ZAYA1 8B excels at detailed long form reasoning especially for mathematical and coding task. It punches heavily above its weight in these regimes and due to its inference efficiency and small size can be highly effective in test time compute harnesses. Due to its small total parameter count, ZAYA1 8B can also be deployed on device for local LLM applications. Learn more in our technical report and blog. This is the post trained reasoning version of ZAYA1 8B. The pretraining base can be found here. Performance ZAYA1 8B performs extremely strongly, especially in challenging mathematical, reasoning, and coding benchmarks. ZAYA1 8B is competitive with models several times its own size including frontier scale reasoning models at mathematica…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy