gravitee io/bert small pii detection 🚀 Token classification model for PII detection, fine tuned from prajjwal1/bert small on gravitee io/pii detection dataset . Label Set How to Use Quick start (pipeline) ONNX Intended use Detect personally identifiable information (PII) spans in english text. Suitable for privacy filtering, redaction pipelines, and data leak prevention particularly on structured data (JSON, HTML, XML, SQL, Document) Evaluation Metric Value F1 0.8686 Precision 0.8182 Recall 0.9256 Eval loss 0.0132 Limitations English focused; other languages will degrade Domain drift is real: audit on your own data Benchmarks External corpus evaluation (English only), seqeval. Last run: 2026 05 21. Benchmark Examples FP32 micro F1 FP32 macro F1 INT8 micro F1 INT8 macro F1 : : : : : gretelai/gretel pii masking en v1:test 5,000 0.9141 0.8971 0.9121 0.8860 gretelai/synthetic pii finance multilingual:test 2,962 0.7534 0.7354 0.7498 0.7351 DataikuNLP/kiji pii training data:test 1,033 0.9259 0.8685 0.9265 0.8725 beki/privy:test 28,843 0.8809 0.9694 0.8800 0.9680 beki/privy:test large 120,574 0.9833 0.9810 0.9825 0.9801 Per entity breakdown gretelai/gretel pii masking en v1:test Entity F…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy