This repository hosts the
Efficient
variant
(lower FLOPs). For the higher-quality checkpoint, see
naver/v-splade-quality
.
Model Summary
V-SPLADE
is a
0.25B (250M) inference-free sparse retriever
for visual-document retrieval — retrieving image-based document pages (rendered PDFs, slides, scanned reports) from a text query.
Inference-free
— queries are resolved by a learned Bag-of-Words lookup with
no neural query encoding at serving time
, so retrieval runs on a standard inverted index (Pyserini / PISA) without a GPU.
Direct visual embedding
— document pages are encoded directly into sparse vectors, building indexes
over 20× faster
than caption- or OCR-based text-extraction pipelines.
V-SPLADE Quality improves average NDCG@5 by
+13.8pp
over the same-scale dense baseline (BiModernVBERT) and by up to
+6.3pp
over the OCR/caption BM25 baselines.
V-SPLADE more than
doubles R@5
over the same-backbone dense retriever at production scale, and retains recall more robustly as the corpus grows from 500K to 18.7M pages.
Document encoding throughput
Method
Pages/sec
V-SPLADE (ours)
20.19
Qwen3-VL-30B-A3B caption (vLLM, eff. 3B)
0.83
Unstructured OCR (Tesseract hi_res)
0.90
Measured on a single H100 GPU with 4 CPU cores, using 1,000 sampled documents across the six benchmarks. V-SPLADE is
over 20× faster
than caption- or OCR-based text-extraction pipelines for index building.
The shortest path to seeing V-SPLADE work on your own page image — encode one image into a sparse vocabulary vector, inspect the top-activated tokens, and score a text query against it:
This model and the accompanying code are released under the
Apache License 2.0
. See
LICENSE
in the repository for the full text.
Base model (
ModernVBERT/modernvbert
) and caption generator (
Qwen3-VL-30B-A3B
) are subject to their own licenses; please review them before redistribution or commercial use.
Training data.
This model was trained on
vidore/colpali_train_set
and
rlhn/rlhn-680K
.
rlhn/rlhn-680K
is distributed under
CC BY-SA 4.0
.
vidore/colpali_train_set
is a collection of multiple source datasets, each of which remains under its own original license.
Citation
@misc{cho2026vsplade,
title = {Inference-Free Multimodal Learned Sparse Retrieval for Production-Scale Visual Document Search},
author = {Cho, Gyu-Hwung and Lee, Youngjune and Jeong, Kiyoon and Lee, Siyoung and Han, Sanggyu and Dejean, Herv{\'e} and Clinchant, St{\'e}phane and Hwang, Seung-won},
year = {2026},
eprint = {2605.30917},
archivePrefix = {arXiv},
primaryClass = {cs.IR}
}
Authors
Gyu-Hwung Cho (NAVER Corp. & Seoul National University), Youngjune Lee, Kiyoon Jeong, Siyoung Lee, Sanggyu Han (NAVER Corp.), Hervé Dejean, Stéphane Clinchant (Naver Labs Europe), Seung-won Hwang (Seoul National University, corresponding).
v-splade-efficient huggingface.co is an AI model on huggingface.co that provides v-splade-efficient's model effect (), which can be used instantly with this naver v-splade-efficient model. huggingface.co supports a free trial of the v-splade-efficient model, and also provides paid use of the v-splade-efficient. Support call v-splade-efficient model through api, including Node.js, Python, http.
v-splade-efficient huggingface.co is an online trial and call api platform, which integrates v-splade-efficient's modeling effects, including api services, and provides a free online trial of v-splade-efficient, you can try v-splade-efficient online for free by clicking the link below.
naver v-splade-efficient online free url in huggingface.co:
v-splade-efficient is an open source model from GitHub that offers a free installation service, and any user can find v-splade-efficient on GitHub to install. At the same time, huggingface.co provides the effect of v-splade-efficient install, users can directly use v-splade-efficient installed effect in huggingface.co for debugging and trial. It also supports api for free installation.