NuExtract3
is a unified
4B
vision-language reasoning model for document understanding.
It combines strong
structured information extraction
with high-quality
image-to-Markdown
conversion, making it suitable for extraction pipelines, OCR, and RAG preprocessing for all types of documents such as scans, receipts, forms, invoices, contracts or tables.
Multimodal inputs
: text, images, or text + images.
Multilingual
documents.
Reasoning
and non-reasoning inference modes.
Template generation
for structured extraction from natural language or input document.
Benchmark results
Structured Extraction
We benchmarked NuExtract on NuMind's internal structured benchmark, measuring model's performances on ~600 documents of diverse types including invoices, movie posters or floor plans. These documents and their ground-truth cover diverse use-cases testing model visual understanding, OCR, reasoning and ability to handle long input and output contexts.
We plan to open-source this benchmark in the coming weeks, along with a extensive leaderboard including most popular open-weight and closed-sourced APIs and a Python library allowing to easily measure model performances on structured extraction.
To measure a pair of predicted and ground-truth JSONs, we represent both as trees which we align based on node names, compute metric scores for aligned leaves and report the average of these scores.
string
and
verbatim-string
leaves are evaluated with indel distance (i.e. Levenshtein without replacement), while all others are evaluated with exact-match.
Models were evaluated using vllm, with a temperature of 0.25 and a maximum of 65000 output token (for both thinking and answer), which largely exceeds 22000 which is the number of tokens of the largest ground truth output.
Model name
Average score
Num. failed⁽¹⁾
Avg. num tokens thinking
Avg. num tokens answer
NuExtract3.4_4B-RL
0.651 ± 0.019
27
2036
1856
gemma-4-E4B-it
0.538 ± 0.023
31
3005
1287
Qwen3.5-9B
0.479 ± 0.030
170
22409
1257
Qwen3.5-4B
0.417 ± 0.031
229
27177
1201
GLM-4.6V-Flash
0.435 ± 0.026
153
2989
1357
Nemotron-3-Nano-Omni
0.387 ± 0.028
204
25827
522
Ministral-3-3B
0.240 ± 0.022
344
27586
362
(1) number of model outputs that were not JSON deserializable, either directly or by removing leading and trailing backticks.
95% confidence intervals computed using a nonparametric bootstrap over scores distributions.
The benchmark include samples containing multiple images resulting in large input context, and some with ground-truth containing large numbers of items to extract resulting in large outputs. We found that the reasoning of small models significantly negatively impact their performances. The reason is that many models ended up falling in repetition loops, hitting the output tokens limit and resulting in failed requests.
Document to Markdown
NuExtract can also convert document images into clean Markdown. Output will be Markdown for text (headers etc), HTML for tables, LaTeX for math and
<figure data-type="image" data-id="img_n"><img src="/NM-dev/model_card-A/resolve/main/img_n.png" alt="Detail description of the images"/>
Modern, format-agnostic benchmarks for complex document understanding are limited, so we explored a new evaluation approach.
We selected 100 documents with challenging layouts and tables, asked each model to convert them into a structured representation, then used Gemini 3 Flash to compare model outputs against the source document and choose the most accurate result.
The rankings aligned with human votes, suggesting this is a promising method for evaluating document-to-Markdown capabilities. More details will be shared in an upcoming technical report.
Here are some results:
Using "Markdown-to-structured"
To add other evaluate references, we used our structured extraction benchmark to evaluate models in a two-step fashion: convert the benchmark inputs to Markdown, then use Qwen3.6 27B to perform the structured extraction task on them. Intuitively, it allows to evaluate how models achieve to keep the input document content and layout: good models will allow the "structured extractor" model to perform better scores.
Using NuExtract
Structured extraction
Structured extraction takes as inputs:
An input document, which can be text, image, or both;
A JSON template describing the information to extract;
(Optional) Instructions, allowing to specify expected output formats or values;
(Optional) In-Context Learning (ICL) examples.
Input JSON template
NuExtract uses a input JSON template whose structure is identical to the output JSON. Its leaf values are specify the
types
of the output JSON leaves. For examples:
NuExtract can also convert document images into clean Markdown. Output will be markdown for text (headers etc), html for tables, latex for mat and
<figure data-type="image" data-id="img_n"><img src="/NM-dev/model_card-A/resolve/main/img_n.png" alt="Detail description of the images"/>
Markdown example:
<figuredata-type="image"data-id="img_1"><imgsrc="/numind/NuExtract3-GGUF/resolve/main/img_1.png"alt="Logo of Mobilier 2000 with contact information: Tél.: (418) 275-4232, 1654, boul. Marcotte, Roberval (Qc) G8H 2P2"/></figure># COMMANDE**NUMÉRO 72259**
1
**Vendu à**
TREMBLAY ERIC
ERIC TREMBLAY
348 BOUL. DE L'ANSE
ROBERVAL
G8H 1Y9
**Livré à**
TREMBLAY ERIC
ERIC TREMBLAY
348 BOUL. DE L'ANSE
ROBERVAL
G8H 1Y9
<table><thead> <tr> <th># CLIENT</th> <th>EXPÉDITEUR</th> <th>TERME DE CRÉDIT</th> <th>DATE</th> </tr> </thead> <tbody> <tr> <td>2753133</td> <td>Notre camion</td> <td>à la livraison</td> <td>22/06/2023</td> </tr> </tbody></table><table><thead> <tr> <th>NOM DU VENDEUR</th> <th>VOTRE ÉCONOMIE !</th> <th># COMMANDE</th> </tr> </thead> <tbody> <tr> <td>Éric</td> <td>0.00</td> <td></td> </tr> </tbody></table>
Reasoning and non-reasoning modes
NuExtract supports both reasoning and non-reasoning inference.
Non-thinking mode
Use this for fast and deterministic extraction or Markdown conversion.
enable_thinking = False
temperature = 0.2
Thinking mode
Use this for difficult documents, complex layouts, ambiguous fields, or cases where the document structure requires additional reasoning.
enable_thinking = True
temperature = 0.6
For production extraction workloads, we recommend starting with
non-reasoning mode
and enabling reasoning only for difficult examples.
vLLM deployment
NuExtract can be served with vLLM using an OpenAI-compatible API.
MTP can improve decoding throughput without changing the OpenAI-compatible request payload. You can tune
num_speculative_tokens
for your hardware and workload, or remove
--speculative-config
if your vLLM version or environment does not support this speculative decoding method.
If you encounter memory issues, reduce the maximum model length and the maximum number of images:
NuExtract supports in-context examples for structured extraction.
Examples are especially useful when the desired formatting is ambiguous or when the schema requires task-specific conventions. Examples can be provided by using
developer
messages, for which all items of the contents except the last one are the input, and the last one is the expected output.
import json
from openai import OpenAI
client = OpenAI(
api_key="EMPTY",
base_url="http://localhost:8000/v1",
)
template = {
"names": ["string"]
}
response = client.chat.completions.create(
model="numind/NuExtract3",
temperature=0,
messages=[
{
"role": "developer",
"content": [
{
"type": "text",
"text": "Stephen is the manager at Susan's store.",
},
{
"type": "text",
"text": "{\"names\": [\"-STEPHEN-\", \"-SUSAN-\"]}",
}
],
},
{
"role": "user",
"content": [
{
"type": "text",
"text": "John went to the restaurant with Mary. James went to the cinema."
}
],
}
],
extra_body={
"chat_template_kwargs": {
"template": json.dumps(template, indent=4),
"enable_thinking": False
}
}
)
print(response.choices[0].message.content)
Example output:
{"names":["-JOHN-","-MARY-","-JAMES-"]}
vLLM inference: template generation
NuExtract can generate an extraction template from a natural language description.
from openai import OpenAI
client = OpenAI(
api_key="EMPTY",
base_url="http://localhost:8000/v1",
)
response = client.chat.completions.create(
model="numind/NuExtract3",
temperature=0,
messages=[
{
"role": "user",
"content": [
{
"type": "text",
"text": "I want to extract the key details from a rental contract."
}
],
}
],
extra_body={
"chat_template_kwargs": {
"mode": "template-generation"
}
}
)
print(response.choices[0].message.content)
The following examples assume that vLLM is running locally on port 8000. They use
jq
to build valid JSON request bodies without manually escaping the image data or template string.
You can also run NuExtract directly with `transformers`. The same `template`, `mode`, and `enable_thinking` options are passed to `processor.apply_chat_template`.
NuExtract3-GGUF huggingface.co is an AI model on huggingface.co that provides NuExtract3-GGUF's model effect (), which can be used instantly with this numind NuExtract3-GGUF model. huggingface.co supports a free trial of the NuExtract3-GGUF model, and also provides paid use of the NuExtract3-GGUF. Support call NuExtract3-GGUF model through api, including Node.js, Python, http.
NuExtract3-GGUF huggingface.co is an online trial and call api platform, which integrates NuExtract3-GGUF's modeling effects, including api services, and provides a free online trial of NuExtract3-GGUF, you can try NuExtract3-GGUF online for free by clicking the link below.
numind NuExtract3-GGUF online free url in huggingface.co:
NuExtract3-GGUF is an open source model from GitHub that offers a free installation service, and any user can find NuExtract3-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of NuExtract3-GGUF install, users can directly use NuExtract3-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.