deepseek-ai / DeepSeek-OCR

huggingface.co
Total runs: 2.4M
24-hour runs: 0
7-day runs: -50.3K
30-day runs: 67.6K
Model's Last Updated: November 04 2025
image-text-to-text

Introduction of DeepSeek-OCR

Model Details of DeepSeek-OCR

DeepSeek AI

🌟 Github | 📥 Model Download | 📄 Paper Link | 📄 Arxiv Paper Link |

DeepSeek-OCR: Contexts Optical Compression

Explore the boundaries of visual-text compression.

Usage

Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8:

torch==2.6.0
transformers==4.46.3
tokenizers==0.20.3
einops
addict 
easydict
pip install flash-attn==2.7.3 --no-build-isolation
from transformers import AutoModel, AutoTokenizer
import torch
import os
os.environ["CUDA_VISIBLE_DEVICES"] = '0'
model_name = 'deepseek-ai/DeepSeek-OCR'

tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModel.from_pretrained(model_name, _attn_implementation='flash_attention_2', trust_remote_code=True, use_safetensors=True)
model = model.eval().cuda().to(torch.bfloat16)

# prompt = "<image>\nFree OCR. "
prompt = "<image>\n<|grounding|>Convert the document to markdown. "
image_file = 'your_image.jpg'
output_path = 'your/output/dir'

# infer(self, tokenizer, prompt='', image_file='', output_path = ' ', base_size = 1024, image_size = 640, crop_mode = True, test_compress = False, save_results = False):

# Tiny: base_size = 512, image_size = 512, crop_mode = False
# Small: base_size = 640, image_size = 640, crop_mode = False
# Base: base_size = 1024, image_size = 1024, crop_mode = False
# Large: base_size = 1280, image_size = 1280, crop_mode = False

# Gundam: base_size = 1024, image_size = 640, crop_mode = True

res = model.infer(tokenizer, prompt=prompt, image_file=image_file, output_path = output_path, base_size = 1024, image_size = 640, crop_mode=True, save_results = True, test_compress = True)
vLLM

Refer to 🌟GitHub for guidance on model inference acceleration and PDF processing, etc.

Visualizations
Acknowledgement

We would like to thank Vary , GOT-OCR2.0 , MinerU , PaddleOCR , OneChart , Slow Perception for their valuable models and ideas.

We also appreciate the benchmarks: Fox , OminiDocBench .

Citation

Coming soon!

Runs of deepseek-ai DeepSeek-OCR on huggingface.co

2.4M
Total runs
0
24-hour runs
0
3-day runs
-50.3K
7-day runs
67.6K
30-day runs

More Information About DeepSeek-OCR huggingface.co Model

More DeepSeek-OCR license Visit here:

https://choosealicense.com/licenses/mit

DeepSeek-OCR huggingface.co

DeepSeek-OCR huggingface.co is an AI model on huggingface.co that provides DeepSeek-OCR's model effect (), which can be used instantly with this deepseek-ai DeepSeek-OCR model. huggingface.co supports a free trial of the DeepSeek-OCR model, and also provides paid use of the DeepSeek-OCR. Support call DeepSeek-OCR model through api, including Node.js, Python, http.

deepseek-ai DeepSeek-OCR online free

DeepSeek-OCR huggingface.co is an online trial and call api platform, which integrates DeepSeek-OCR's modeling effects, including api services, and provides a free online trial of DeepSeek-OCR, you can try DeepSeek-OCR online for free by clicking the link below.

deepseek-ai DeepSeek-OCR online free url in huggingface.co:

https://huggingface.co/deepseek-ai/DeepSeek-OCR

DeepSeek-OCR install

DeepSeek-OCR is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-OCR on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-OCR install, users can directly use DeepSeek-OCR installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

DeepSeek-OCR install url in huggingface.co:

https://huggingface.co/deepseek-ai/DeepSeek-OCR

Url of DeepSeek-OCR

Provider of DeepSeek-OCR huggingface.co

deepseek-ai
ORGANIZATIONS

Other API from deepseek-ai