LightOnOCR-2-1B from lightonai is LightOn's flagship 1B-parameter end-to-end vision-language OCR model—the recommended variant for most tasks—refined via RLVR training on a 2.5x scaled 43M-page corpus with enhanced French, arXiv, scan, and LaTeX coverage for converting PDFs, scans, and document images into clean, naturally ordered text at 3.3× Chandra OCR speed, 1.7× OlmOCR, 5× dots.ocr, and 5.71 pages/s on H100 (~<$0.01/1k pages) while achieving state-of-the-art 83.2±0.9 on OlmOCR-Bench (outperforming Chandra-9B by 1.5+ points) across tables, receipts, forms, multi-column layouts, and math without brittle pipelines. Part of the fully differentiable LightOnOCR-2 family (including bbox variants for image localization, base models for fine-tuning, and soup merges), it uses a native-resolution ViT encoder (from Mistral-Small-3.1), MLP projector, and Qwen3 decoder with 1540px longest-edge preprocessing (200 DPI PDFs) for superior accuracy on degraded scans, scientific docs, and European languages under Apache 2.0, supporting LoRA/PEFT fine-tuning via Transformers for domain adaptation.
(sorted by size, not necessarily quality. IQ-quants are often preferable over similar sized non-IQ quants)
Here is a handy graph by ikawrakow comparing some lower-quality quant
types (lower is better):