lift is a structured extraction model from
Datalab
that pulls structured JSON out of PDFs and images. Pass any JSON schema and lift returns a JSON object matching it, using schema-constrained decoding to guarantee valid, well-typed output.
A schema is standard JSON Schema. Keep it simple —
string
,
number
,
integer
,
boolean
, arrays of those, arrays of objects, and nested objects are all supported. Write a
description
for any field whose name isn't self-explanatory, and mark a field
required
only when it must appear; fields genuinely absent from a document come back
null
.
from lift import extract
from lift.model import InferenceManager
# Start the vLLM server first with: lift_vllm
model = InferenceManager(method="vllm")
result = extract("document.pdf", "schema.json", model=model)
print(result.extraction)
With HuggingFace Transformers
from lift import extract
from lift.model import InferenceManager
# Loads datalab-to/lift in-process (requires: pip install lift-pdf[hf])
model = InferenceManager(method="hf")
result = extract("document.pdf", "schema.json", model=model)
print(result.extraction)
extract
accepts the schema as a dict, a path to a
.json
file, an inline JSON string, or the name of a saved schema. Pass
page_range="0-5"
to limit PDF pages, and set
VLLM_API_BASE
to target a remote server.
Benchmarks
Evaluated on a 225-document extraction benchmark (6–64 pages per document, ~11,000 scored fields) with adversarial cases planted throughout: cross-page values, exhaustive lists, fields that must be left null, near-miss distractors, multi-source aggregation. Scoring is deterministic exact-match against ground truth (numeric tolerance, normalized strings).
All models receive the same rendered page images, and extract each document in a single pass.
Model
Size
Field accuracy
Full-document accuracy
Median latency*
Features
Datalab API
—
95.9%
44.4%
30.8s
Citations + Verification
Gemini Flash 3.5
—
91.3%
40.0%
28.1s
lift
9B
90.2%
20.9%
9.5s
Azure Content Understanding
—
83.4%
22.2%
73.7s
NuExtract3
4B
81.5%
8.4%
8.3s
Qwen3.5-9B
9B
76.3%
24.0%
16.8s
* Per document, 8 concurrent requests. Local models (lift, Qwen3.5-9B, NuExtract3) served with vLLM on a single GPU; Gemini, Datalab, and Azure via API. Latency varies with hardware and load — treat as relative, not absolute.
Field accuracy
— fraction of individual schema fields extracted correctly.
Full-document accuracy
— fraction of documents where
every
field is correct.
Hosted models with verification, citations, and confidence scores are available via the
Datalab API
— test in the
playground
.
Commercial Usage
Code is Apache 2.0. Model weights use a modified OpenRAIL-M license: free for research, personal use, and startups under $5M funding/revenue. Cannot be used competitively with our API. For broader commercial licensing, see
pricing
.
lift huggingface.co is an AI model on huggingface.co that provides lift's model effect (), which can be used instantly with this datalab-to lift model. huggingface.co supports a free trial of the lift model, and also provides paid use of the lift. Support call lift model through api, including Node.js, Python, http.
lift huggingface.co is an online trial and call api platform, which integrates lift's modeling effects, including api services, and provides a free online trial of lift, you can try lift online for free by clicking the link below.
datalab-to lift online free url in huggingface.co:
lift is an open source model from GitHub that offers a free installation service, and any user can find lift on GitHub to install. At the same time, huggingface.co provides the effect of lift install, users can directly use lift installed effect in huggingface.co for debugging and trial. It also supports api for free installation.