DeepSeek-V3 architecture with 4 layers + 8 experts per MoE + MTP module + FP8 weights from original model without further tuning
To be used in CI testing
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy
DeepSeek-V3 architecture with 4 layers + 8 experts per MoE + MTP module + FP8 weights from original model without further tuning
DeepSeek-V3 architecture with 4 layers + 8 experts per MoE + MTP module + FP8 weights from original model without further tuning
To be used in CI testing
Mirrored from an external registry.
Last synced 6/14/2026
No commits available