ONNX export of
lightonai/LateOn-regularized
(Apache-2.0) for
fastlate
and any ONNX Runtime client. The weights are unchanged; this repo adds the graph, an INT8 dynamic quantization of it, and the ColBERT settings the model was trained with.
Files
model.onnx
: transformer, the PyLate
Dense
projection layer(s), and per-token L2 normalization, in one graph. Inputs
input_ids
,
attention_mask
(int64, dynamic batch and length); output
embeddings
of shape
(batch, seq, 128)
. Opset 17.
model_int8.onnx
:
onnxruntime.quantization.quantize_dynamic
of the above, QInt8 weights.
onnx_config.json
: prefix ids, lengths, skiplist, expansion and padding settings copied from the PyLate model, plus
fde_center
.
tokenizer.json
: the source tokenizer, with
[Q]
/
[D]
as added tokens.
Settings
Query prefix
[Q]
(id 50368), document prefix
[D]
(id 50369), inserted right after the sequence-start token. Query length 32, document length 300.
do_query_expansion=False
: queries are not padded. Punctuation tokens (
skiplist_words
) are dropped from document embeddings after encoding, as in PyLate.
fde_center=True
: whether to subtract the corpus mean token vector before building MUVERA fixed-dimensional encodings. Measured on a 3,000-chunk code/notebook corpus with FDEs of 4,096 dims: this model needs centring (recall of the exact top-20 within 1,000 candidates went from 0.61 to 0.94), matching the evaluation notes on the source model card.
Validation
fp32 graph vs PyTorch forward: max abs diff 4.8e-07; int8 graph vs PyTorch: 0.017; int8 graph vs PyLate
encode
on a sample text: documents 0.016, queries 0.022.
Exported with PyLate 1.6.0, transformers 5.3.0, torch 2.14.0, onnxruntime 1.2x, on 2026-09-18.
Runs of answerdotai LateOn-regularized-onnx on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About LateOn-regularized-onnx huggingface.co Model
LateOn-regularized-onnx huggingface.co is an AI model on huggingface.co that provides LateOn-regularized-onnx's model effect (), which can be used instantly with this answerdotai LateOn-regularized-onnx model. huggingface.co supports a free trial of the LateOn-regularized-onnx model, and also provides paid use of the LateOn-regularized-onnx. Support call LateOn-regularized-onnx model through api, including Node.js, Python, http.
LateOn-regularized-onnx huggingface.co is an online trial and call api platform, which integrates LateOn-regularized-onnx's modeling effects, including api services, and provides a free online trial of LateOn-regularized-onnx, you can try LateOn-regularized-onnx online for free by clicking the link below.
answerdotai LateOn-regularized-onnx online free url in huggingface.co:
LateOn-regularized-onnx is an open source model from GitHub that offers a free installation service, and any user can find LateOn-regularized-onnx on GitHub to install. At the same time, huggingface.co provides the effect of LateOn-regularized-onnx install, users can directly use LateOn-regularized-onnx installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
LateOn-regularized-onnx install url in huggingface.co: