merve / PaddleOCR-VL-1.5-hf

huggingface.co
Total runs: 37
24-hour runs: 0
7-day runs: 1
30-day runs: 22
Model's Last Updated: February 12 2026
image-text-to-text

Introduction of PaddleOCR-VL-1.5-hf

Model Details of PaddleOCR-VL-1.5-hf

PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing

repo HuggingFace ModelScope HuggingFace ModelScope Discord X License

🔥 Official Website 📝 Technical Report

Introduction

PaddleOCR-VL-1.5 is an advanced next-generation model of PaddleOCR-VL, achieving a new state-of-the-art accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real-world physical distortions—including scanning artifacts, skew, warping, screen photography, and illumination—we propose the Real5-OmniDocBench benchmark. Experimental results demonstrate that this enhanced model attains SOTA performance on the newly curated benchmark. Furthermore, we extend the model’s capabilities by incorporating seal recognition and text spotting tasks, while remaining a 0.9B ultra-compact VLM with high efficiency. This repository contains official transformers weights for PaddleOCR-VL-1.5.

Key Capabilities of PaddleOCR-VL-1.5
  1. With a parameter size of 0.9B , PaddleOCR-VL-1.5 achieves 94.5% accuracy on OmniDocBench v1.5 , surpassing the previous SOTA model PaddleOCR-VL. Significant improvements are observed in table, formula, and text recognition.

  2. It introduces an innovative approach to document parsing by supporting irregular-shaped localization , enabling accurate polygonal detection under skewed and warped document conditions. Evaluations across five real-world scenarios—scanning, skew, warping, screen-photography, and illumination—demonstrate superior performance over mainstream open-source and proprietary models.

  3. The model introduces text spotting (text-line localization and recognition) , along with seal recognition , with all corresponding metrics setting new SOTA results in their respective tasks.

  4. PaddleOCR-VL-1.5 further strengthens its capability in specialized scenarios and multilingual recognition. Recognition performance is improved for rare characters, ancient texts, multilingual tables, underlines, and checkboxes, and language coverage is extended to include China's Tibetan script and Bengali.

  5. The model supports automatic cross-page table merging and cross-page paragraph heading recognition , effectively mitigating content fragmentation issues in long-document parsing.

Model Architecture
News
  • 2026.01.29 🚀 We release PaddleOCR-VL-1.5 , —a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing.
Usage

This model should be used with PPDocLayoutV3 . Find an end-to-end example inference in this notebook .

Make sure to have transformers above v5.

python -m pip install "transformers>=5.0.0"

You can load the model as follows. Since you need to have the detected regions, please refer to notebook for complete inference.

import torch
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "PaddlePaddle/PaddleOCR-VL-1.5-hf"

processor = AutoProcessor.from_pretrained(ocr_model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id, torch_dtype=torch.bfloat16
).to(device)

def ocr_region(crop, prompt):
    """Run PaddleOCR-VL 1.5 on a single cropped region."""
    messages = [
        {
            "role": "user",
            "content": [
                {"type": "image", "image": crop},
                {"type": "text", "text": prompt},
            ],
        }
    ]
    inputs = processor.apply_chat_template(
        messages,
        add_generation_prompt=True,
        tokenize=True,
        return_dict=True,
        return_tensors="pt",
    ).to(ocr_model.device)

    generated_ids = model.generate(**inputs, max_new_tokens=1024)
    trimmed = generated_ids[0][inputs["input_ids"].shape[-1] :]
    return processor.decode(trimmed, skip_special_tokens=True)

parsed_regions = []

# assuming you have detected regions from PPDocLayoutv3 in detections

for det in detections:
    label = det["label"]
    prompt = LABEL_TO_PROMPT.get(label)

    x1, y1, x2, y2 = det["box"]
    crop = image.crop((x1, y1, x2, y2))
    text = recognise_region(crop, prompt)

    parsed_regions.append({**det, "prompt": prompt, "text": text})
    print(f"[{det['order']}] {label}  prompt={prompt}")
    print(text)
Performance
Document Parsing
1. OmniDocBench v1.5
PaddleOCR-VL-1.5 achieves SOTA performance for overall, text, formula, tables and reading order on OmniDocBench v1.5

Notes:

  • Performance metrics are cited from the OmniDocBench official leaderboard , except for Gemini-3 Pro, Qwen3-VL-235B-A22B-Instruct and our model, which were evaluated independently.
2. Real5-OmniDocBench
Across all five diverse and challenging scenarios—scanning, warping, screen-photography, illumination, and skew—PaddleOCR-VL-1.5 consistently sets new SOTA records

Notes:

  • Real5-OmniDocBench is a brand-new benchmark oriented toward real-world scenarios, which we constructed based on the OmniDocBench v1.5 dataset. The dataset comprises five distinct scenarios: Scanning, Warping, Screen-photography, Illumination, and Skew. For further details, please refer to Real5-OmniDocBench .
Inference Performance

Notes:

  • End-to-End Inference Performance Comparison on OmniDocBench v1.5. PDF documents were processed in batches of 512 on a single NVIDIA A100 GPU. The reported end-to-end runtime includes both PDF rendering and Markdown generation. All methods rely on their built-in PDF parsing modules and default DPI settings to reflect out-of-the-box performance.
Visualization
Real-word Document Parsing
Illumination
Skew
Screen Photography
Scanning
Warping
Text Spotting
Seal Recognition
Acknowledgments

We would like to thank PaddleFormers , Keye , MinerU , OmniDocBench for providing valuable code, model weights and benchmarks. We also appreciate everyone's contribution to this open-source project!

Citation

If you find PaddleOCR-VL-1.5 helpful, feel free to give us a star and citation.

@misc{cui2026paddleocrvl15multitask09bvlm,
      title={PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing}, 
      author={Cheng Cui and Ting Sun and Suyin Liang and Tingquan Gao and Zelun Zhang and Jiaxuan Liu and Xueqing Wang and Changda Zhou and Hongen Liu and Manhui Lin and Yue Zhang and Yubo Zhang and Yi Liu and Dianhai Yu and Yanjun Ma},
      year={2026},
      eprint={2601.21957},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2601.21957}, 
}

Runs of merve PaddleOCR-VL-1.5-hf on huggingface.co

37
Total runs
0
24-hour runs
0
3-day runs
1
7-day runs
22
30-day runs

More Information About PaddleOCR-VL-1.5-hf huggingface.co Model

More PaddleOCR-VL-1.5-hf license Visit here:

https://choosealicense.com/licenses/apache-2.0

PaddleOCR-VL-1.5-hf huggingface.co

PaddleOCR-VL-1.5-hf huggingface.co is an AI model on huggingface.co that provides PaddleOCR-VL-1.5-hf's model effect (), which can be used instantly with this merve PaddleOCR-VL-1.5-hf model. huggingface.co supports a free trial of the PaddleOCR-VL-1.5-hf model, and also provides paid use of the PaddleOCR-VL-1.5-hf. Support call PaddleOCR-VL-1.5-hf model through api, including Node.js, Python, http.

PaddleOCR-VL-1.5-hf huggingface.co Url

https://huggingface.co/merve/PaddleOCR-VL-1.5-hf

merve PaddleOCR-VL-1.5-hf online free

PaddleOCR-VL-1.5-hf huggingface.co is an online trial and call api platform, which integrates PaddleOCR-VL-1.5-hf's modeling effects, including api services, and provides a free online trial of PaddleOCR-VL-1.5-hf, you can try PaddleOCR-VL-1.5-hf online for free by clicking the link below.

merve PaddleOCR-VL-1.5-hf online free url in huggingface.co:

https://huggingface.co/merve/PaddleOCR-VL-1.5-hf

PaddleOCR-VL-1.5-hf install

PaddleOCR-VL-1.5-hf is an open source model from GitHub that offers a free installation service, and any user can find PaddleOCR-VL-1.5-hf on GitHub to install. At the same time, huggingface.co provides the effect of PaddleOCR-VL-1.5-hf install, users can directly use PaddleOCR-VL-1.5-hf installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

PaddleOCR-VL-1.5-hf install url in huggingface.co:

https://huggingface.co/merve/PaddleOCR-VL-1.5-hf

Url of PaddleOCR-VL-1.5-hf

PaddleOCR-VL-1.5-hf huggingface.co Url

Provider of PaddleOCR-VL-1.5-hf huggingface.co

merve
ORGANIZATIONS

Other API from merve

huggingface.co

Total runs: 58
Run Growth: -3
Growth Rate: -5.17%
Updated:October 12 2023
huggingface.co

Total runs: 34
Run Growth: -24
Growth Rate: -70.59%
Updated:January 06 2024
huggingface.co

Total runs: 33
Run Growth: 20
Growth Rate: 68.97%
Updated:September 19 2025
huggingface.co

Total runs: 32
Run Growth: -14
Growth Rate: -43.75%
Updated:February 26 2024
huggingface.co

Total runs: 30
Run Growth: 8
Growth Rate: 26.67%
Updated:January 28 2026
huggingface.co

Total runs: 14
Run Growth: 7
Growth Rate: 50.00%
Updated:July 18 2024
huggingface.co

Total runs: 14
Run Growth: 7
Growth Rate: 50.00%
Updated:October 04 2023
huggingface.co

Total runs: 11
Run Growth: 0
Growth Rate: 0.00%
Updated:February 22 2024
huggingface.co

Total runs: 11
Run Growth: -2
Growth Rate: -18.18%
Updated:September 11 2025
huggingface.co

Total runs: 10
Run Growth: 9
Growth Rate: 90.00%
Updated:November 25 2023
huggingface.co

Total runs: 9
Run Growth: 6
Growth Rate: 66.67%
Updated:February 10 2023
huggingface.co

Total runs: 8
Run Growth: -22
Growth Rate: -275.00%
Updated:December 18 2024
huggingface.co

Total runs: 8
Run Growth: 2
Growth Rate: 25.00%
Updated:March 26 2024
huggingface.co

Total runs: 8
Run Growth: 4
Growth Rate: 50.00%
Updated:July 11 2023
huggingface.co

Total runs: 7
Run Growth: 2
Growth Rate: 28.57%
Updated:March 26 2024
huggingface.co

Total runs: 5
Run Growth: 3
Growth Rate: 60.00%
Updated:April 02 2023
huggingface.co

Total runs: 5
Run Growth: 3
Growth Rate: 60.00%
Updated:April 02 2023
huggingface.co

Total runs: 4
Run Growth: 0
Growth Rate: 0.00%
Updated:June 14 2023
huggingface.co

Total runs: 1
Run Growth: -11
Growth Rate: -1100.00%
Updated:January 28 2023
huggingface.co

Total runs: 1
Run Growth: 1
Growth Rate: 100.00%
Updated:December 20 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 25 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 22 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:August 30 2022
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 30 2022
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:October 22 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:January 04 2024