privacy filter multilingual v2 Fine tuned openai/privacy filter for fine grained PII extraction across 54 categories in 16 languages . This v2 checkpoint is the more performant successor to OpenMed/privacy filter multilingual , with stronger multilingual PII masking behavior while keeping the same 16 language, fine grained OpenMed label space and runtime interface. Base model : openai/privacy filter — 1.4B parameter MoE (50M active per token), BIOES token classification head Task : Token classification for PII detection (BIOES scheme) Languages (16) : Arabic, Bengali, Chinese, Dutch, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Spanish, Telugu, Turkish, Vietnamese Training data : The original language balanced multilingual OpenMed/AI4Privacy mix, followed by a v2 source balanced privacy masking adaptation mix from AI4Privacy OpenPII, Nemotron, Gretel, and Privy style PII data Recipe : opf train (OpenAI's official fine tuning CLI) — full fine tune, AdamW, balanced language and source sampling, bf16 Labels : 54 PII categories → 217 BIOES classes (1 O + 54 × B/I/E/S) The base model ships with 8 coarse PII categories and English only training. This model trade…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy