Developed by
AxionML
for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.
This is an NVFP4-quantized version of
Qwen/Qwen3.5-2B-Base
(2B parameters), quantized using
NVIDIA TensorRT Model Optimizer
. Weights and activations of linear layers are quantized to FP4, reducing disk size and GPU memory by ~4x compared to BF16.
About NVFP4 quantization:
NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection strategies such as dual-pass evaluation comparing "map max to 6" versus "map max to 4 with clipping." On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity to reduce multiplier area while higher-precision FP32 accumulation protects dot-product accuracy.
Ready for commercial and non-commercial use under
Apache 2.0
.
Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency.
Qwen3.5 Highlights
Qwen3.5 features the following enhancement:
Unified Vision-Language Foundation
: Early fusion training on multimodal tokens achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks.
Efficient Hybrid Architecture
: Gated Delta Networks combined with sparse Mixture-of-Experts deliver high-throughput inference with minimal latency and cost overhead.
Scalable RL Generalization
: Reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability.
Global Linguistic Coverage
: Expanded support to 201 languages and dialects, enabling inclusive, worldwide deployment with nuanced cultural and regional understanding.
Next-Generation Training Infrastructure
: Near-100% multimodal training efficiency compared to text-only training and asynchronous RL frameworks supporting massive-scale agent scaffolds and environment orchestration.
For more details, please refer to our blog post
Qwen3.5
.
Number of Linear Attention Heads: 16 for V and 16 for QK
Head Dimension: 128
Gated Attention:
Number of Attention Heads: 8 for Q and 2 for KV
Head Dimension: 256
Rotary Position Embedding Dimension: 64
Feed Forward Network:
Intermediate Dimension: 6144
LM Output: 248320 (Tied to token embedding)
MTP: trained with multi-steps
Context Length: 262,144 natively and extensible up to 1,010,000 tokens.
Citation
If you find our work helpful, feel free to give us a cite.
@misc{qwen3.5,
title = {{Qwen3.5}: Towards Native Multimodal Agents},
author = {{Qwen Team}},
month = {February},
year = {2026},
url = {https://qwen.ai/blog?id=qwen3.5}
}
Quantization Details
This model was quantized by applying NVFP4 to the weights and activations of linear operators within transformer blocks. The KV-cache is not quantized. Vision encoder weights are kept in their original precision.
The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations. It may generate inaccurate, biased, or offensive content. Please refer to the
original model card
for full details.
Runs of AxionML Qwen3.5-2B-Base-NVFP4 on huggingface.co
27
Total runs
1
24-hour runs
3
3-day runs
14
7-day runs
-2
30-day runs
More Information About Qwen3.5-2B-Base-NVFP4 huggingface.co Model
Qwen3.5-2B-Base-NVFP4 huggingface.co is an AI model on huggingface.co that provides Qwen3.5-2B-Base-NVFP4's model effect (), which can be used instantly with this AxionML Qwen3.5-2B-Base-NVFP4 model. huggingface.co supports a free trial of the Qwen3.5-2B-Base-NVFP4 model, and also provides paid use of the Qwen3.5-2B-Base-NVFP4. Support call Qwen3.5-2B-Base-NVFP4 model through api, including Node.js, Python, http.
Qwen3.5-2B-Base-NVFP4 huggingface.co is an online trial and call api platform, which integrates Qwen3.5-2B-Base-NVFP4's modeling effects, including api services, and provides a free online trial of Qwen3.5-2B-Base-NVFP4, you can try Qwen3.5-2B-Base-NVFP4 online for free by clicking the link below.
AxionML Qwen3.5-2B-Base-NVFP4 online free url in huggingface.co:
Qwen3.5-2B-Base-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.5-2B-Base-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.5-2B-Base-NVFP4 install, users can directly use Qwen3.5-2B-Base-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3.5-2B-Base-NVFP4 install url in huggingface.co: