NVFP4 quantized version of
allenai/Olmo-3-7B-Instruct
with extended 1M token context support via linear RoPE scaling.
Model Description
This model is the NVFP4 (4-bit floating point) quantized version of OLMo-3-7B-Instruct, optimized for NVIDIA DGX Spark systems with Blackwell GB10 GPUs and Ada Lovelace architecture support. The quantization uses NVIDIA's ModelOpt library with two-level scaling: E4M3 FP8 per block plus FP32 global scale.
Key Features
Base Model:
allenai/Olmo-3-7B-Instruct (7.3B parameters)
Quantization Format:
NVFP4 with group_size=16
Context Length:
1,048,576 tokens (1M) via linear RoPE scaling
Model Size:
5.30 GB (64% reduction from 14.60 GB)
GPU Memory:
~5.23 GiB (64% reduction)
Performance
Metric
Original
Quantized
Improvement
Model Size
14.60 GB
5.30 GB
64% reduction
GPU Memory
14.6 GB
5.23 GiB
64% reduction
Context Length
4,096
1,048,576
256x increase
Inference Speed
-
31-35 tok/s
-
Usage
Important:
This model requires vLLM with ModelOpt quantization support. It cannot be loaded with standard transformers.
The context was extended from 4,096 to 1,048,576 tokens using linear RoPE scaling:
Scaling Factor:
16x
rope_theta:
50,000,000
rope_scaling:
{"type": "linear", "factor": 16.0}
Note: Actual usable context depends on available GPU memory. With 120GB GPU at 95% utilization, approximately 200,000 tokens can be stored in KV cache.
Architecture Compatibility
For vLLM compatibility, the model uses:
Architecture:
Olmo2ForCausalLM
Model Type:
olmo2
This mapping allows vLLM to properly load the OLMo-3 architecture.
Limitations
Requires vLLM with
--quantization modelopt
flag
Cannot be loaded with standard transformers
Requires NVIDIA GPU with FP4 support (Ada Lovelace or newer)
Maximum usable context limited by GPU memory for KV cache
Intended Use
Long-context instruction following and chat
Document analysis and summarization
Code generation and review
Research and educational purposes
License
Apache 2.0 (inherited from base model)
Citation
@misc{olmo3-nvfp4-1m,
author = {Ex0bit},
title = {OLMo-3-7B-Instruct-NVFP4-1M: NVFP4 Quantized OLMo-3 with 1M Context},
year = {2024},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Ex0bit/OLMo-3-7B-Instruct-NVFP4-1M}}
}
OLMo-3-7B-Instruct-NVFP4-1M huggingface.co is an AI model on huggingface.co that provides OLMo-3-7B-Instruct-NVFP4-1M's model effect (), which can be used instantly with this Ex0bit OLMo-3-7B-Instruct-NVFP4-1M model. huggingface.co supports a free trial of the OLMo-3-7B-Instruct-NVFP4-1M model, and also provides paid use of the OLMo-3-7B-Instruct-NVFP4-1M. Support call OLMo-3-7B-Instruct-NVFP4-1M model through api, including Node.js, Python, http.
OLMo-3-7B-Instruct-NVFP4-1M huggingface.co is an online trial and call api platform, which integrates OLMo-3-7B-Instruct-NVFP4-1M's modeling effects, including api services, and provides a free online trial of OLMo-3-7B-Instruct-NVFP4-1M, you can try OLMo-3-7B-Instruct-NVFP4-1M online for free by clicking the link below.
Ex0bit OLMo-3-7B-Instruct-NVFP4-1M online free url in huggingface.co:
OLMo-3-7B-Instruct-NVFP4-1M is an open source model from GitHub that offers a free installation service, and any user can find OLMo-3-7B-Instruct-NVFP4-1M on GitHub to install. At the same time, huggingface.co provides the effect of OLMo-3-7B-Instruct-NVFP4-1M install, users can directly use OLMo-3-7B-Instruct-NVFP4-1M installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
OLMo-3-7B-Instruct-NVFP4-1M install url in huggingface.co: