We are pleased to announce the release of OvisOCR2, a compact 0.8B end-to-end model for page-level document parsing. Given a document page image, OvisOCR2 generates a Markdown representation in natural reading order, covering text, formulas, tables, and visual regions.
OvisOCR2 is developed by post-training Qwen3.5-0.8B using a carefully designed data engine that combines real-world and synthetic data, together with a multi-stage training recipe integrating SFT, RL, and OPD. The model delivers strong document parsing performance while maintaining a small deployment footprint.
OvisOCR2 achieves an overall score of 96.58 on OmniDocBench v1.6, establishing a new state of the art and
becoming the first end-to-end model to top this leaderboard previously dominated by pipeline methods
. On PureDocBench, OvisOCR2 also achieves the highest Avg3 score of 75.06.
Performance
Inference
pip install "vllm==0.22.1" pillow
from PIL import Image
from vllm import LLM, SamplingParams
classOvisOCR2Parser:
def__init__(self, model_name_or_path: str):
self.model = LLM(
model=model_name_or_path,
tensor_parallel_size=1,
gpu_memory_utilization=0.8,
gdn_prefill_backend="triton"
)
prompt = '\nExtract all readable content from the image in natural human reading order and output the result as a single Markdown document. For charts or images, represent them using an HTML image tag: <' + 'img src="images/bbox_{left}_{top}_{right}_{bottom}.jpg" />, where left, top, right, bottom are bounding box coordinates scaled to [0, 1000). Format formulas as LaTeX. Format tables as HTML: <table>...</table>. Transcribe all other text as standard Markdown. Preserve the original text without translation or paraphrasing.'
self.prompt = self.model.get_tokenizer().apply_chat_template(
[{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": prompt}]}],
tokenize=False,
add_generation_prompt=True,
enable_thinking=False
)
self.sampling_params = SamplingParams(
max_tokens=16384,
temperature=0.0
)
def_clean_truncated_repeats(
self, text: str, min_text_len: int = 8000, max_period: int = 200, min_period: int = 1, min_repeat_chars: int = 100, min_repeat_times: int = 5) -> str:
n = len(text)
if n < min_text_len:
return text
max_period = min(max_period, n - 1)
for unit_len inrange(min_period, max_period + 1):
if text[n - 1] != text[n - 1 - unit_len]:
continue
match_len = 1
idx = n - 2while idx >= unit_len and text[idx] == text[idx - unit_len]:
match_len += 1
idx -= 1
total_len = match_len + unit_len
repeat_times = total_len // unit_len
tail_len = total_len % unit_len
if repeat_times >= min_repeat_times and total_len >= min_repeat_chars:
return text[: n - total_len + unit_len] + text[n - tail_len:]
return text
defparse(self, images: list[Image.Image], filter_imgtags: bool = True) -> list[str]:
vllm_inputs = [
{
"prompt": self.prompt,
"multi_modal_data": {"image": image},
"mm_processor_kwargs": {
"images_kwargs": {
"min_pixels": 448 * 448,
"max_pixels": 2880 * 2880
}
}
}
for image in images
]
outputs = self.model.generate(vllm_inputs, self.sampling_params)
markdowns = []
for output in outputs:
text = output.outputs[0].text.strip()
if filter_imgtags:
text = "\n\n".join(
block
for block in text.split("\n\n")
ifnot block.strip().startswith('<img src="images/bbox_')
)
markdowns.append(self._clean_truncated_repeats(text))
return markdowns
if __name__ == "__main__":
parser = OvisOCR2Parser("ATH-MaaS/OvisOCR2")
images = [Image.open("test1.jpg"), Image.open("test2.jpg")]
markdowns = parser.parse(images)
print(markdowns[0])
By default,
parse
removes HTML image tags for visual regions. To render Markdown with visual regions, set
filter_imgtags=False
and save the Markdown file together with the referenced image crops as follows:
If you find OvisOCR2 useful, please consider citing our technical report:
@misc{lu2026ovisocr2,
title = {{OvisOCR2 Technical Report}},
author = {Lu, Shiyin and Li, Yinglun and Xia, Yu and Chen, Yuhui and Ji, An-Yang and Jiang, Jun-Peng and Chen, Qing-Guo and Zhao, Jianshan and Lin, En and Li, Haijun and Qin, Cheng and Xu, Zhao and Luo, Weihua},
year = {2026}
}
We used filtering and quality-assurance procedures during data construction to reduce parsing errors such as repeated outputs, incomplete content, invalid table/formula structures, and reading-order inconsistencies. Due to the diversity and complexity of real-world documents, OvisOCR2 may still produce incorrect or incomplete outputs. Please manually verify results in critical applications.
Runs of ATH-MaaS OvisOCR2 on huggingface.co
131.4K
Total runs
-2.3K
24-hour runs
-4.8K
3-day runs
-7.1K
7-day runs
30.5K
30-day runs
More Information About OvisOCR2 huggingface.co Model
OvisOCR2 huggingface.co is an AI model on huggingface.co that provides OvisOCR2's model effect (), which can be used instantly with this ATH-MaaS OvisOCR2 model. huggingface.co supports a free trial of the OvisOCR2 model, and also provides paid use of the OvisOCR2. Support call OvisOCR2 model through api, including Node.js, Python, http.
OvisOCR2 huggingface.co is an online trial and call api platform, which integrates OvisOCR2's modeling effects, including api services, and provides a free online trial of OvisOCR2, you can try OvisOCR2 online for free by clicking the link below.
ATH-MaaS OvisOCR2 online free url in huggingface.co:
OvisOCR2 is an open source model from GitHub that offers a free installation service, and any user can find OvisOCR2 on GitHub to install. At the same time, huggingface.co provides the effect of OvisOCR2 install, users can directly use OvisOCR2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.