🖥️ API / Platform 📑 Blog 🗣️ Discord 🛠️ GitHub NuExtract3 is a unified 4B vision language reasoning model for document understanding. It combines strong structured information extraction with high quality image to Markdown conversion, making it suitable for extraction pipelines, OCR, and RAG preprocessing for all types of documents such as scans, receipts, forms, invoices, contracts or tables. Try it out in the 🤗 space! Overview Structured extraction : input (text/images) + JSON template + instructions JSON output Markdown conversion : input (text/images) Markdown Multimodal inputs : text, images, or text + images. Multilingual documents. Reasoning and non reasoning inference modes. Template generation for structured extraction from natural language or input document. Benchmark results Structured Extraction We benchmarked NuExtract on NuMind's internal structured benchmark, measuring model's performances on ~600 documents of diverse types including invoices, movie posters or floor plans. These documents and their ground truth cover diverse use cases testing model visual understanding, OCR, reasoning a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy