π Github π₯ Model Download π Paper Link π Arxiv Paper Link DeepSeek OCR 2: Visual Causal Flow Explore more human like visual encoding. Usage Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8οΌ vLLM Refer to πGitHub for guidance on model inference acceleration and PDF processing, etc. Support Modes Dynamic resolution Default: (0 6)Γ768Γ768 + 1Γ1024Γ1024 β (0 6)Γ144 + 256 visual tokens β Main Prompts Acknowledgement We would like to thank DeepSeek OCR, Vary, GOT OCR2.0, MinerU, PaddleOCR for their valuable models and ideas. We also appreciate the benchmark OmniDocBench. Citation bibtex @article{wei2025deepseek, title={DeepSeek OCR: Contexts Optical Compression}, author={Wei, Haoran and Sun, Yaofeng and Li, Yukun}, journal={arXiv preprint arXiv:2510.18234}, year={2025} } @article{wei2026deepseek, title={DeepSeek OCR 2: Visual Causal Flow}, author={Wei, Haoran and Sun, Yaofeng and Li, Yukun}, journal={arXiv preprint arXiv:2601.20552}, year={2026} }
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy