中文 | English 🖥️ Official Website 💬 GitHub Table of Contents Model Introduction Highlights Benchmark Results STEM & Reasoning Context Learning & Instruction Following Code & Agent News Model Links Quickstart Deployment vLLM SGLang Training Quantization License Contact Us Model Introduction Hy3 preview is a 295B parameter Mixture of Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. Hy3 preview is the first model trained on our rebuilt infrastructure, and the strongest we've shipped so far. It improves significantly on complex reasoning, instruction following, context learning, coding, and agent tasks. Property Value : : Architecture Mixture of Experts (MoE) Total Parameters 295B Activated Parameters 21B MTP Layer Parameters 3.8B Number of Layers (excluding MTP layer) 80 Number of MTP Layers 1 Attention Heads 64 (GQA, 8 KV heads, head dim 128) Hidden Size 4096 Intermediate Size 13312 Context Length 256K Vocabulary Size 120832 Number of Experts 192 experts, top 8 activated Supported Precisions BF16 Highlights STEM & Reasoni…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy