PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages Abstract: Truly multilingual safety moderation efforts for Large Language Models (LLMs) have been hindered by a narrow focus on a small set of languages (e.g., English, Chinese) as well as a limited scope of safety definition, resulting in significant gaps in moderation capabilities. To bridge these gaps, we release PolyGuard, a new state of the art multilingual safety model for safeguarding LLM generations, and the corresponding training and evaluation datasets. PolyGuard is trained on PolyGuardMix, the largest multilingual safety training corpus to date containing 1.91M samples across 17 languages (e.g., Chinese, Czech, English, Hindi). We also introduce PolyGuardPrompts, a high quality multilingual benchmark with 29K samples for the evaluation of safety guardrails. Created by combining naturally occurring multilingual human LLM interactions and human verified machine translations of an English only safety dataset (WildGuardMix; Han et al., 2024), our datasets contain prompt output pairs with labels of prompt harmfulness, response harmfulness, and response refusal. Through extensive evaluations across multiple safe…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy