High Accuracy Vietnamese Text Correction v1.3.1 Introduction ProtonX Text Correction (v1.3 NC) A specialized Vietnamese text correction model engineered for high accuracy normalization of legal and enterprise text. Optimized for OCR post processing (including PaddleOCR outputs), but also capable of cleaning broader Vietnamese text with diacritic restoration, segmentation repair, and correction of domain specific terminology. The model is optimized to clean up real world OCR mistakes such as: missing or incorrect diacritics broken word segmentation misrecognized legal terms punctuation artifacts formatting inconsistencies Built on a Seq2Seq Transformer architecture, the model is trained on 800,000 correction pairs, including 30,000 pairs manually annotated by expert Vietnamese annotators, covering: official legal documents OCR outputs from scanned PDFs colloquial → standardized legal text Strict constraints ensure: Correction ≠ rewriting meaning of legal text must never change no hallucination / no added legal terms confidence based correction no paraphrasing LICENSE This model is released under the ProtonX Text Correction Model License (v1.3 NC). See LICENSE.md for full terms, cond…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy