Figure 1: Performance comparison on the OmniDocBench v1.5 benchmark. FireRed-OCR achieves state-of-the-art performance among end-to-end solutions, ranking first with a score above 92%.
🔥 FireRed-OCR
FireRed-OCR
is a systematic framework designed to specialize general Large Vision-Language Models (LVLMs) into high-performance, pixel-precise structural document parsing experts.
General VLMs frequently suffer from
"Structural Hallucination"
(e.g., disordered rows, invented formulas) when processing complex documents. FireRed-OCR addresses this by shifting the paradigm from "impressionist" text generation to "structural engineering," achieving State-of-the-Art (SOTA) results on authoritative benchmarks like OmniDocBench v1.5.
✨ Key Features
SOTA Performance
: Achieves
92.94%
overall score on OmniDocBench v1.5, significantly outperforming DeepSeek-OCR 2, OCRVerse, and massive general VLMs (e.g., Gemini-3.0 Pro,Qwen3-VL-235B).
Structural Integrity
: Utilizing
Format-Constrained GRPO
(Group Relative Policy Optimization), the model enforces strict syntactic validity, eliminating common errors like unclosed tables or invalid LaTeX formulas.
"Geometry + Semantics" Data Factory
: A novel data engine that uses geometric feature clustering and multi-dimensional tagging to synthesize balanced datasets, effectively handling long-tail layouts.
Format-Constrained GRPO
: Self-correction via Reinforcement Learning.
In-the-Wild Robustness
: Demonstrates superior resilience on complex, non-standard layouts (FireRedBench) compared to traditional pipeline systems like PaddleOCR.
📰 News
2026.02.28
: We released FireRed-OCR-2B weights. Check more details in the
Model Zoo
section.
🗂️ Model Zoo
Models
Base
Description
Download Link
FireRed-OCR-2B
Qwen3-VL-2B-Instruct
Lightweight version achieving 92.94% Overall on OmniDocBench v1.5.
The FireRed-OCR framework transforms a general VLM into a structural expert through a three-stage progressive training strategy:
Stage 1: Multi-task Pre-alignment
: Trains the model on detection, region recognition, and layout-to-markdown tasks to ground visual perception.
Stage 2: Specialized SFT
: Fine-tunes on a high-quality, standardized Markdown dataset to ensure logical consistency and hierarchical expression.
Stage 3: Format-Constrained GRPO
: Applies Reinforcement Learning with specific rewards for
Formula Syntax
,
Table Integrity
,
Hierarchical Closure
, and
Text Accuracy
.
⚡️ Quick Start
FireRed-OCR is based on the Qwen3-VL architecture. You can use the following code snippets to generate structured Markdown from document images.
FireRed-OCR is a technical tool designed for document digitization and structural parsing.
Prohibited Use
: This project must not be used to generate or process content that is illegal, defamatory, pornographic, harmful, or that violates the privacy, rights, or interests of individuals or organizations.
User Responsibility
: Users are solely responsible for any content generated using this project. The authors and contributors assume no responsibility or liability for any misuse of the codebase or for any consequences resulting from its use.
🤝 Acknowledgements
We would like to thank the developers of the amazing open-source projects, including
Qwen-VL
,
PaddleOCR
,
olmOCR
and the broader OCR community.
Runs of FireRedTeam FireRed-OCR on huggingface.co
1.1K
Total runs
0
24-hour runs
0
3-day runs
129
7-day runs
-823
30-day runs
More Information About FireRed-OCR huggingface.co Model
FireRed-OCR huggingface.co is an AI model on huggingface.co that provides FireRed-OCR's model effect (), which can be used instantly with this FireRedTeam FireRed-OCR model. huggingface.co supports a free trial of the FireRed-OCR model, and also provides paid use of the FireRed-OCR. Support call FireRed-OCR model through api, including Node.js, Python, http.
FireRed-OCR huggingface.co is an online trial and call api platform, which integrates FireRed-OCR's modeling effects, including api services, and provides a free online trial of FireRed-OCR, you can try FireRed-OCR online for free by clicking the link below.
FireRedTeam FireRed-OCR online free url in huggingface.co:
FireRed-OCR is an open source model from GitHub that offers a free installation service, and any user can find FireRed-OCR on GitHub to install. At the same time, huggingface.co provides the effect of FireRed-OCR install, users can directly use FireRed-OCR installed effect in huggingface.co for debugging and trial. It also supports api for free installation.