中文 | English 🖥️ Official Website 💬 GitHub Table of Contents Model Introduction Stronger Agent Capabilities More Reliable Product Experiences Benchmark Appendix News Model Links Quickstart Deployment vLLM SGLang Finetuning RL Post training Quantization License Contact Us Model Introduction Hy3 is a 295B parameter Mixture of Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. Following the Hy3 Preview launch in late April, we gathered feedback from 50+ products and scaled up post training with higher quality data. Today, we introduce Hy3, which outperforms similar size models and rivals flagship open source models with 2 5x parameters. It also shows significant gains in utility across various products and productivity tasks. Property Value : : Architecture Mixture of Experts (MoE) Total Parameters 295B Activated Parameters 21B MTP Layer Parameters 3.8B Number of Layers (excluding MTP layer) 80 Number of MTP Layers 1 Attention Heads 64 (GQA, 8 KV heads, head dim 128) Hidden Size 4096 Intermediate Size 13312 Context Length 25…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy