A unified 0.8B model for document parsing and OCR-related understanding.
English · 简体中文
Xiaomi-OCR-0 is a unified 0.8B vision-language model for document parsing and OCR-centric understanding. Starting from Qwen3.5-0.8B, it is trained on an approximately 170M-sample OCR-centric corpus with Q-Mask text anchoring, continued pretraining (CPT), and mixed-task reinforcement learning (Mix-RL).
git clone https://github.com/SeerRay-Lab/Xiaomi-OCR-0.git
cd Xiaomi-OCR-0- Document parsing: reads document pages and produces structured Markdown, with OTSL model outputs converted to HTML tables in Markdown and formulas in LaTeX.
- Key information extraction (KIE): extracts requested fields as JSON.
- OCR-centric visual question answering (VQA): answers questions about text and content in document images.
- Local agent tools: an MCP service handles prompts, PDF/image preparation, optional region detection, parallel crop inference, and output assembly.
- Browser demo: try page and region parsing through a local web interface.
For printed, regular documents, region parsing can detect and process regions concurrently. For scene text, handwriting, calligraphy, historical books, and irregular layouts, use whole-page parsing because incorrect region boundaries can lose context.
⚠️ Notice: With your authorization, this Skill may create a Python environment, install dependencies, download model weights, install and configure SGLang or vLLM, and register an MCP server. Model and layout downloads can be large. If you do not agree to this setup flow, do not execute the command below.
copy and send to your agent:
Read and execute https://raw.githubusercontent.com/SeerRay-Lab/Xiaomi-OCR-0/main/SKILL.md
The model will run on your machine.
You can run the browser demo without installing the Agent Skill or MCP server. The demo sends requests to a local SGLang or vLLM inference server at http://127.0.0.1:8000/v1.
Python 3.10 or newer is required.
If you already have a compatible local server, skip to step 2. Otherwise, choose one runtime and follow its installation guide for your operating system, GPU, and driver. These runtimes have hardware-specific requirements; the linked guides list supported platforms and installation options:
Use a release that supports the checkpoint architecture Qwen3_5ForConditionalGeneration. For macOS and non-NVIDIA hardware, check the runtime's platform-specific support before installing; the default commands below are for a compatible GPU installation.
Run one of the following commands in a terminal. On first launch, the runtime downloads the checkpoint from Hugging Face if it is not already cached. Keep this terminal open while using the demo.
SGLang:
python -m sglang.launch_server --model-path SeerRay-Lab/Xiaomi-OCR-0 \
--host 127.0.0.1 --port 8000 --context-length 16384vLLM:
vllm serve SeerRay-Lab/Xiaomi-OCR-0 \
--host 127.0.0.1 --port 8000 --max-model-len 16384Wait for the server to finish loading the model before continuing.
Open a second terminal in the repository directory and run:
python3 -m pip install -r requirements.txt
python3 demo/server.py --port 8787Open http://127.0.0.1:8787. The same browser page supports whole-page/region document parsing, KIE, VQA, and PDF parsing. Install the shared requirements above, including Pillow and pypdfium2 for images and PDFs. Region mode optionally needs PaddlePaddle/PaddleX and PP-DocLayoutV3 weights. See demo/README.md for dependencies and INSTALL.md for MCP setup.
| Path | Contents |
|---|---|
skills/xiaomi-ocr/ |
Agent Skill and MCP service for OCR, PDF parsing, KIE, and VQA |
demo/ |
Local browser demo and sample cases |
example_pics/ |
Example inputs, reference Markdown, and demo animations mirrored from the Hugging Face model repository |
pipeline/ |
Whole-page and region batch inference |
postprocess/ |
Shared table conversion, formula/text assembly and repetition handling |
| Hugging Face model | Model card, examples, and benchmark assets |
| Hugging Face Space | Space assets and project-page source |
examples/ |
MCP configuration example |
The comparisons below summarize selected results. Arrows indicate the preferred direction; bold marks Xiaomi-OCR-0 and does not necessarily indicate the best result in a column. See the Hugging Face model card for benchmark notes and additional results.
| Model | Size | OmniDocBench v1.6 ↑ | Real5 ↑ | Wild ↑ |
|---|---|---|---|---|
| Xiaomi-OCR-0 | 0.8B | 96.83 | 95.24 | 87.94 |
| TeleOCR | 1.2B | 96.87 | — | 88.53 |
| OvisOCR2 | 0.8B | 96.58 | 92.29 | 87.91 |
| PaddleOCR-VL-1.6 | 0.9B | 96.33 | 93.19 | 87.36 |
| MinerU2.5-Pro | 1.2B | 95.75 | 88.94 | 87.33 |
| GLM-OCR | 0.9B | 95.22 | 90.32 | 85.08 |
All three columns report Overall scores.
| Model | Size | DocVQA | InfoVQA | ChartQA | OCRBench | TextVQA | Mean |
|---|---|---|---|---|---|---|---|
| Xiaomi-OCR-0 | 0.8B | 93.1 | 75.1 | 84.6 | 84.6 | 78.6 | 83.2 |
| Qwen3.5-0.8B | 0.8B | 88.5 | 60.3 | 69.5 | 77.9 | 68.3 | 72.9 |
| Qwen3.5-2B | 2B | 92.4 | 72.4 | 77.0 | 85.9 | 76.9 | 80.9 |
| Qwen3.5-4B | 4B | 94.4 | 80.4 | 82.4 | 86.6 | 80.8 | 84.9 |
| MiniCPM-V-4.5 | 8B | 84.9 | 69.6 | 87.4 | 89.0 | 82.2 | 82.6 |
Mean is the arithmetic average of the five benchmarks on a 0–100 scale.
If you follow our work, please cite the following:
@misc{chen2026xiaomiocr0technicalreport,
title={Xiaomi-OCR-0 Technical Report},
author={Xin Chen and Anan Du and Feng Feng and Pei Fu and Jian Luan and Longwei Xu and Shaojie Zhang and Hang Li and Heng Qu and Cheng Tan},
year={2026},
eprint={2609.36136},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.36136},
}