Llama 3.1 Nemotron Safety Guard 8B v3 Model Overview Llama 3.1 Nemotron Safety Guard 8B v3 is a multilingual content safety model that moderates human LLM interaction content and classifies user prompts and LLM responses as safe or unsafe. If the content is unsafe, the model additionally returns a response with a list of categories that the content violates. It supports 9 languages: English, Spanish, Mandarin, German, French, Hindi, Japanese, Arabic, and Thai. The base large language model (LLM) is the multilingual Llama 3.1 8B Instruct model from Meta. NVIDIA’s optimized release is LoRa tuned on approved datasets and better conforms to NVIDIA’s content safety risk taxonomy and other safety risks in human LLM interactions. The model is trained using the Nemotron Safety Guard Dataset v3 dataset which is synthetically curated using the CultureGuard pipeline. The model shows strong zero shot generalization supporting over 20 languages (en, ar, de, es, fr, hi, ja, th, zh, it, ko, nl, cs, da, fi, iw, pt BR, pl, ru, sv). The model can be prompted using an instruction and a taxonomy of unsafe risks to be categorized. The instruction format for prompt moderation is shown below under input…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy