0xSero / Inkling-Small-EXL3-2.0bpw

huggingface.co
Total runs: 386
24-hour runs: 0
7-day runs: -5
30-day runs: 220
Model's Last Updated: July 31 2026
image-text-to-text

Introduction of Inkling-Small-EXL3-2.0bpw

Model Details of Inkling-Small-EXL3-2.0bpw

Inkling-Small-EXL3-2.0bpw

Calibrated EXL3 trellis quantization of the routed MoE experts in thinkingmachines/Inkling-Small , targeting 2.0 bits per routed-expert weight .

Status: The assembled weight archive has been uploaded. Consult EXL3_MANIFEST.json for the current structural and runtime validation state.

Runtime validation: Text generation, multimodal generation, and MTP validation are pending. The uploaded archive and its packed tensors have structural validation only.

What is quantized
Component Storage
Routed MoE experts, layers 2–41 EXL3/MCG trellis, 2.0 bpw target
Dense MLP layers 0–1 Source BF16
Shared experts and routers Source BF16/FP32
Attention, relative-position, and short-convolution tensors Source precision
Embeddings, norms, and LM head Source precision
Vision/audio components Source precision
Eight MTP layers Source precision

All forty routed MoE layers use integer EXL3 K=2 trellis weights.

  • Calibration: 1,048,576 naturally routed tokens selected with seeded, no-repeat axis water-filling across general, legal, code/agentic, and reasoning/termination data
  • Maximum calibration sequence/sample span: 4,096 tokens
  • Routing: Inkling's natural top-6 routed-expert assignments
  • Source revision: b2d4f225a02032c5d154bff748ab5a00c5ca26e4
  • Achieved routed-trellis rate: 2.000000 bpw
  • Assembled repository payload: 75.75 GiB
  • Per-layer allocation, tensor inventory, sizes, and validation state: EXL3_MANIFEST.json
Compatibility and how to use it

Download the repository with:

hf download 0xSero/Inkling-Small-EXL3-2.0bpw \
  --local-dir Inkling-Small-EXL3-2.0bpw

This repository is not a drop-in Transformers checkpoint . The routed experts use EXL3 trellis tensors while the rest of Inkling remains in source precision. It requires an Inkling-aware EXL3 loader/runtime that understands the tensor layout described by quantization_config.json and EXL3_MANIFEST.json .

Stock ExLlamaV3 v1.2.1 does not yet include an InklingForConditionalGeneration architecture adapter. The upstream BF16 Inkling model has vLLM and SGLang recipes, but those recipes do not by themselves add support for this experts-only EXL3 layout. Do not infer text, image, audio, or MTP runtime support from a successful download or structural assembly alone.

All EXL3 variants
Method

The source model is loaded once for calibration. Hidden states and natural expert assignments are captured for all forty routed layers. Each expert's gate, up, and down projections are calibrated, Hadamard-transformed, and encoded as EXL3/MCG trellis weights. The full sweep checks finite Hessians and scales, exact trellis byte counts, safetensor key counts, and per-file checksums. Before the sweep, a bounded H200 proof on a real Inkling expert also passed trellis pack/unpack/repack equality and finite reconstruction. That bounded kernel proof is not a full-model generation test. Integer-K caches are reused to assemble the seven public variants without repeating the full model calibration.

The target bpw applies to routed-expert trellis weights. The complete repository is larger than a whole-model quantization at the same nominal bpw because attention, shared experts, multimodal components, the LM head, and MTP remain in source precision.

Credits

This is an independent community quantization and is not an official release from Thinking Machines Lab, TurboDerp, or JarvisLabs.

License and use

This derivative follows the upstream Apache 2.0 license and the upstream acceptable-use policy . Review the base model card for intended uses, limitations, and safety information.

Runs of 0xSero Inkling-Small-EXL3-2.0bpw on huggingface.co

386
Total runs
0
24-hour runs
-3
3-day runs
-5
7-day runs
220
30-day runs

More Information About Inkling-Small-EXL3-2.0bpw huggingface.co Model

More Inkling-Small-EXL3-2.0bpw license Visit here:

https://choosealicense.com/licenses/apache-2.0

Inkling-Small-EXL3-2.0bpw huggingface.co

Inkling-Small-EXL3-2.0bpw huggingface.co is an AI model on huggingface.co that provides Inkling-Small-EXL3-2.0bpw's model effect (), which can be used instantly with this 0xSero Inkling-Small-EXL3-2.0bpw model. huggingface.co supports a free trial of the Inkling-Small-EXL3-2.0bpw model, and also provides paid use of the Inkling-Small-EXL3-2.0bpw. Support call Inkling-Small-EXL3-2.0bpw model through api, including Node.js, Python, http.

Inkling-Small-EXL3-2.0bpw huggingface.co Url

https://huggingface.co/0xSero/Inkling-Small-EXL3-2.0bpw

0xSero Inkling-Small-EXL3-2.0bpw online free

Inkling-Small-EXL3-2.0bpw huggingface.co is an online trial and call api platform, which integrates Inkling-Small-EXL3-2.0bpw's modeling effects, including api services, and provides a free online trial of Inkling-Small-EXL3-2.0bpw, you can try Inkling-Small-EXL3-2.0bpw online for free by clicking the link below.

0xSero Inkling-Small-EXL3-2.0bpw online free url in huggingface.co:

https://huggingface.co/0xSero/Inkling-Small-EXL3-2.0bpw

Inkling-Small-EXL3-2.0bpw install

Inkling-Small-EXL3-2.0bpw is an open source model from GitHub that offers a free installation service, and any user can find Inkling-Small-EXL3-2.0bpw on GitHub to install. At the same time, huggingface.co provides the effect of Inkling-Small-EXL3-2.0bpw install, users can directly use Inkling-Small-EXL3-2.0bpw installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Inkling-Small-EXL3-2.0bpw install url in huggingface.co:

https://huggingface.co/0xSero/Inkling-Small-EXL3-2.0bpw

Url of Inkling-Small-EXL3-2.0bpw

Inkling-Small-EXL3-2.0bpw huggingface.co Url

Provider of Inkling-Small-EXL3-2.0bpw huggingface.co

0xSero
ORGANIZATIONS

Other API from 0xSero

huggingface.co

Total runs: 820
Run Growth: 509
Growth Rate: 62.07%
Updated:May 30 2026
huggingface.co

Total runs: 357
Run Growth: -42
Growth Rate: -11.76%
Updated:June 26 2026
huggingface.co

Total runs: 113
Run Growth: 69
Growth Rate: 62.73%
Updated:May 30 2026
huggingface.co

Total runs: 100
Run Growth: 8
Growth Rate: 8.00%
Updated:May 30 2026
huggingface.co

Total runs: 66
Run Growth: 19
Growth Rate: 30.16%
Updated:May 30 2026
huggingface.co

Total runs: 50
Run Growth: 9
Growth Rate: 18.75%
Updated:May 30 2026