PLLuM: A Family of Polish Large Language Models Overview PLLuM is a family of large language models (LLMs) specialized in Polish with additional English data incorporated for broader generalization. Developed through an extensive collaboration with various data providers, PLLuM models are built on high quality text corpora and refined through instruction tuning, preference learning, and advanced alignment techniques. These models are intended to generate contextually coherent text, offer assistance in various tasks (e.g., question answering, summarization), and serve as a foundation for specialized applications such as domain specific intelligent assistants. In 2024, PLLuM models were developed by the PLLuM consortium. Since 2025, their development has continued under HIVE AI, a broader alliance of research institutions and organizations delivering digital public services, focused on open language technologies for Polish public administration. Key Highlights Extensive Data Collection We gathered large scale, high quality text data in Polish and English, focusing on rigorous cleaning and deduplication to ensure the highest training standards. Organic Instruction Dataset We curated t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy