This checkpoint reduces the model size from
39 GB
(BF16) to
~14 GB
(MXFP4)
with minimal quality degradation, enabling inference on a single GPU with more
headroom for KV cache.
Non-quantized layers:
attention, router, embeddings, LM head (remain in BF16)
File size:
~14 GB (model.safetensors)
Quantization Details
MXFP4 (Microscaling FP4) is a 4-bit floating-point format using the E2M1
representation (2-bit exponent, 1-bit mantissa) with shared E8M0 per-group
scaling factors. Each group of 32 weights shares a single 8-bit scale,
and each individual weight is stored as a 4-bit FP4 code. Two FP4 codes
are packed into one uint8 byte (low nibble = even index, high nibble = odd index).
This matches the quantization format used by the official
openai/gpt-oss-20b
MXFP4 checkpoint
and is natively supported by vLLM's Marlin MXFP4 kernels.
What is quantized
Component
Format
Notes
mlp.experts.gate_up_proj
MXFP4
Stored as
_blocks
(U8) +
_scales
(U8)
mlp.experts.down_proj
MXFP4
Stored as
_blocks
(U8) +
_scales
(U8)
self_attn.*
BF16
Kept at full precision
mlp.router.*
BF16
Kept at full precision
embed_tokens
,
lm_head
BF16
Kept at full precision
Tensor layout
Expert weights are transposed from the BF16 checkpoint layout
[num_experts, in_features, out_features]
to the vLLM-expected
[num_experts, out_features, in_features]
before quantization.
This is required because vLLM's MXFP4 loader performs a direct copy
without transposition (unlike the BF16 loader which transposes on load).
The
eos_token_id
includes token 200012 (
<|call|>
) in addition to the
standard
<|return|>
(200002). This is required for tool calling to work
correctly - without it, the model does not stop generation after emitting
a tool call, and the Harmony parser fails.
Conversion Script
The included
convert_mxfp4.py
converts the original BF16
chromadb/context-1
weights
to MXFP4 format. It requires only numpy.
pip install numpy
python convert_mxfp4.py
The script expects the BF16 model in a sibling
context-1/
directory
and writes the quantized model to
context-1-mxfp4/
.
What the script does
Reads each tensor from the BF16
model.safetensors
For MoE expert weights (
gate_up_proj
and
down_proj
):
Transposes from
[E, in, out]
to
[E, out, in]
Computes per-group E8M0 scales (group size 32)
Quantizes to E2M1 FP4 codes using nearest rounding
Packs two FP4 codes per byte
Saves as
*_blocks
(packed weights) and
*_scales
(shared exponents)
Copies all other tensors (attention, router, embeddings) as-is in BF16
Copies tokenizer and chat template files
Writes
config.json
with
quantization_config
for vLLM
Key Capabilities
(Inherited from chromadb/context-1)
Query decomposition:
Breaks complex multi-constraint questions into
targeted subqueries.
Parallel tool calling:
Averages 2.56 tool calls per turn, reducing
total turns and end-to-end latency.
Self-editing context:
Selectively prunes irrelevant documents
mid-search to sustain retrieval quality over long horizons within a
bounded context window (0.94 prune accuracy).
Cross-domain generalization:
Trained on web, legal, and finance
tasks; generalizes to held-out domains and public benchmarks
(BrowseComp-Plus, SealQA, FRAMES, HLE).
Important: Agent Harness Required
Context-1 is trained to operate within a specific agent harness that
manages tool execution, token budgets, context pruning, and deduplication.
The harness is not yet public.
Running the model without it will not
reproduce the results reported in the technical report.
@techreport{bashir2026context1,
title = {Chroma Context-1: Training a Self-Editing Search Agent},
author = {Bashir, Hammad and Hong, Kelly and Jiang, Patrick and Shi, Zhiyi},
year = {2026},
month = {March},
institution = {Chroma},
url = {https://trychroma.com/research/context-1},
}
License
Apache 2.0
Runs of evilfreelancer context-1-mxfp4 on huggingface.co
689
Total runs
8
24-hour runs
61
3-day runs
99
7-day runs
83
30-day runs
More Information About context-1-mxfp4 huggingface.co Model
context-1-mxfp4 huggingface.co is an AI model on huggingface.co that provides context-1-mxfp4's model effect (), which can be used instantly with this evilfreelancer context-1-mxfp4 model. huggingface.co supports a free trial of the context-1-mxfp4 model, and also provides paid use of the context-1-mxfp4. Support call context-1-mxfp4 model through api, including Node.js, Python, http.
context-1-mxfp4 huggingface.co is an online trial and call api platform, which integrates context-1-mxfp4's modeling effects, including api services, and provides a free online trial of context-1-mxfp4, you can try context-1-mxfp4 online for free by clicking the link below.
evilfreelancer context-1-mxfp4 online free url in huggingface.co:
context-1-mxfp4 is an open source model from GitHub that offers a free installation service, and any user can find context-1-mxfp4 on GitHub to install. At the same time, huggingface.co provides the effect of context-1-mxfp4 install, users can directly use context-1-mxfp4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.