uyu 2 28B Introduction uyu 2 28B is a language model specialized for role playing. It is derived from google/gemma 4 31B it , structurally pruned with Global Iterative Structured Pruning (GISP), and then fine tuned on English and Korean question answering data. Model details Property Value Model name uyu 2 28b Parameters 28,181,549,312 Weight format BF16 Safetensors, 6 shards Weight size 56,363,098,744 bytes Layers 60 Hidden size 5,376 Configured maximum position length 262,144 tokens Validated serving length 2,048 tokens Modalities Text only Languages Korean and English Primary use Conversation and role play Base checkpoint google/gemma 4 31B it Pruning GISP structured pruning Recovery fine tuning LoRA on Korean and English question answering data, merged into the weights Pruning method: GISP This checkpoint was produced through a staged sequence of globally ranked structured pruning passes. Each pass used the same Korean role playing calibration set, BF16 inference, a fixed maximum sequence length, and teacher forced language model loss restricted to assistant response tokens. The model weights were frozen during calibration; only temporary scalar pruning gates received gradients…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy