privacy filter multilingual Fine tuned openai/privacy filter for fine grained PII extraction across 54 categories in 16 languages . Base model : openai/privacy filter — 1.4B parameter MoE (50M active per token), BIOES token classification head Task : Token classification for PII detection (BIOES scheme) Languages (16) : Arabic, Bengali, Chinese, Dutch, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Spanish, Telugu, Turkish, Vietnamese Training data : Multilingual mix from AI4Privacy — pii masking 200k , pii masking 400k , and open pii masking 500k ai4privacy , language balanced Recipe : opf train (OpenAI's official fine tuning CLI) — full fine tune, AdamW, balanced language sampling, 5 epochs, bf16 Labels : 54 PII categories → 217 BIOES classes (1 O + 54 × B/I/E/S) The base model ships with 8 coarse PII categories and English only training. This model trades that for a 6.75× more granular vocabulary spanning identity, contact, address, financial, vehicle, digital, and crypto labels — all evaluated across 16 languages. Family at a glance. Same architecture, three runtimes: PyTorch (this repo) — CPU + CUDA, anywhere transformers runs. MLX BF16 — OpenMed/privac…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy