deepseek-ai / DeepSeek-OCR-2

huggingface.co
Total runs: 830.3K
24-hour runs: 112
7-day runs: -32.6K
30-day runs: -257.7K
Model's Last Updated: February 03 2026
image-text-to-text

Introduction of DeepSeek-OCR-2

Model Details of DeepSeek-OCR-2

DeepSeek AI

🌟 Github | 📥 Model Download | 📄 Paper Link | 📄 Arxiv Paper Link |

DeepSeek-OCR 2: Visual Causal Flow

Explore more human-like visual encoding.

Usage

Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8:

torch==2.6.0
transformers==4.46.3
tokenizers==0.20.3
einops
addict 
easydict
pip install flash-attn==2.7.3 --no-build-isolation
from transformers import AutoModel, AutoTokenizer
import torch
import os
os.environ["CUDA_VISIBLE_DEVICES"] = '0'
model_name = 'deepseek-ai/DeepSeek-OCR-2'

tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModel.from_pretrained(model_name, _attn_implementation='flash_attention_2', trust_remote_code=True, use_safetensors=True)
model = model.eval().cuda().to(torch.bfloat16)

# prompt = "<image>\nFree OCR. "
prompt = "<image>\n<|grounding|>Convert the document to markdown. "
image_file = 'your_image.jpg'
output_path = 'your/output/dir'


res = model.infer(tokenizer, prompt=prompt, image_file=image_file, output_path = output_path, base_size = 1024, image_size = 768, crop_mode=True, save_results = True)
vLLM

Refer to 🌟GitHub for guidance on model inference acceleration and PDF processing, etc.

Support-Modes
  • Dynamic resolution
    • Default: (0-6)×768×768 + 1×1024×1024 — (0-6)×144 + 256 visual tokens ✅
Main Prompts
# document: <image>\n<|grounding|>Convert the document to markdown.
# without layouts: <image>\nFree OCR.
Acknowledgement

We would like to thank DeepSeek-OCR , Vary , GOT-OCR2.0 , MinerU , PaddleOCR for their valuable models and ideas.

We also appreciate the benchmark OmniDocBench .

Citation
coming soon~

Runs of deepseek-ai DeepSeek-OCR-2 on huggingface.co

830.3K
Total runs
112
24-hour runs
468
3-day runs
-32.6K
7-day runs
-257.7K
30-day runs

More Information About DeepSeek-OCR-2 huggingface.co Model

More DeepSeek-OCR-2 license Visit here:

https://choosealicense.com/licenses/apache-2.0

DeepSeek-OCR-2 huggingface.co

DeepSeek-OCR-2 huggingface.co is an AI model on huggingface.co that provides DeepSeek-OCR-2's model effect (), which can be used instantly with this deepseek-ai DeepSeek-OCR-2 model. huggingface.co supports a free trial of the DeepSeek-OCR-2 model, and also provides paid use of the DeepSeek-OCR-2. Support call DeepSeek-OCR-2 model through api, including Node.js, Python, http.

deepseek-ai DeepSeek-OCR-2 online free

DeepSeek-OCR-2 huggingface.co is an online trial and call api platform, which integrates DeepSeek-OCR-2's modeling effects, including api services, and provides a free online trial of DeepSeek-OCR-2, you can try DeepSeek-OCR-2 online for free by clicking the link below.

deepseek-ai DeepSeek-OCR-2 online free url in huggingface.co:

https://huggingface.co/deepseek-ai/DeepSeek-OCR-2

DeepSeek-OCR-2 install

DeepSeek-OCR-2 is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-OCR-2 on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-OCR-2 install, users can directly use DeepSeek-OCR-2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

DeepSeek-OCR-2 install url in huggingface.co:

https://huggingface.co/deepseek-ai/DeepSeek-OCR-2

Url of DeepSeek-OCR-2

DeepSeek-OCR-2 huggingface.co Url

Provider of DeepSeek-OCR-2 huggingface.co

deepseek-ai
ORGANIZATIONS

Other API from deepseek-ai