T lite it 2.1 🚨 Users are advised to exercise caution and are responsible for any additional training and oversight required to ensure the model's responses meet acceptable ethical and safety standards. The responsibility for incorporating this model into industrial or commercial solutions lies entirely with those who choose to deploy it. Description T lite it 2.1 is an efficient Russian model built upon the Qwen 3 architecture, featuring significant improvements in instruction following and adds support for tool calling capabilities — a key advancement over T lite it 1.0, which lacks tool use support. Outperforms Qwen3 8B in tool calling scenarios, which is essential for agentic applications. Built for both general tasks and complex workflows, with higher Russian text generation throughput enabled by optimized tokenizer. More train details in our Habr: https://habr.com/ru/companies/tbank/articles/979650/ NOTE: This model supports only non thinking mode and does not generate in its output. Meanwhile, specifying enable thinking=False is no longer required. 📚 Dataset Instruction midtraining: 40B tokens of instruction data. Supervised Fine Tuning (SFT): ~670K high quality and divers…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy