arcee-ai / Trinity-Mini-NVFP4

huggingface.co
Total runs: 212
24-hour runs: 0
7-day runs: 132
30-day runs: 132
Model's Last Updated: May 29 2026
text-generation

Introduction of Trinity-Mini-NVFP4

Model Details of Trinity-Mini-NVFP4

Arcee Trinity Mini

Trinity Mini NVFP4

This repository contains the NVFP4 quantized weights of Trinity-Mini for deployment on NVIDIA Blackwell GPUs.

Trinity Mini is an Arcee AI 26B MoE model with 3B active parameters. It is the medium-sized model in our new Trinity family, a series of open-weight models for enterprise and tinkerers alike.

This model is tuned for reasoning, but in testing, it uses a similar total token count to competitive instruction-tuned models.


Trinity Mini is trained on 10T tokens gathered and curated through a key partnership with Datology , building upon the excellent dataset we used on AFM-4.5B with additional math and code.

Training was performed on a cluster of 512 H200 GPUs powered by Prime Intellect using HSDP parallelism.

More details, including key architecture decisions, can be found on our blog here


Model Details
  • Model Architecture: AfmoeForCausalLM
  • Parameters: 26B, 3B active
  • Experts: 128 total, 8 active, 1 shared
  • Context length: 128k
  • Training Tokens: 10T
  • License: OpenMDW-1.1
  • Recommended settings:
    • temperature: 0.15
    • top_k: 50
    • top_p: 0.75
    • min_p: 0.06

Benchmarks

Powered by Datology
Quantization Details
  • Scheme: NVFP4 ( nvfp4_mlp_only — MLP/expert weights only, attention remains BF16)
  • Tool: NVIDIA ModelOpt
  • Calibration: 512 samples, seq_length=2048, all-expert calibration enabled
  • KV cache: Not quantized
Running with vLLM

Requires vLLM >= 0.18.0. Native FP4 compute requires Blackwell GPUs; older GPUs fall back to Marlin weight decompression automatically.

Blackwell GPUs (B200/B300/GB300) — Docker (recommended)
docker run --runtime nvidia --gpus all -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:v0.18.0-cu130 \
  arcee-ai/Trinity-Mini-NVFP4 \
  --trust-remote-code \
  --gpu-memory-utilization 0.90 \
  --max-model-len 8192
Hopper GPUs (H100/H200) and others
vllm serve arcee-ai/Trinity-Mini-NVFP4 \
  --trust-remote-code \
  --gpu-memory-utilization 0.90 \
  --max-model-len 8192 \
  --host 0.0.0.0 \
  --port 8000

Note (Blackwell pip installs): If installing vLLM via pip on Blackwell rather than using Docker, native FP4 kernels may produce incorrect output due to package version mismatches. As a workaround, force the Marlin backend:

export VLLM_NVFP4_GEMM_BACKEND=marlin

vllm serve arcee-ai/Trinity-Mini-NVFP4 \
  --trust-remote-code \
  --moe-backend marlin \
  --gpu-memory-utilization 0.90 \
  --max-model-len 8192 \
  --host 0.0.0.0 \
  --port 8000

Marlin decompresses FP4 weights to BF16 for compute, providing the full memory compression benefit (~3.7× vs BF16) but not native FP4 compute speedup. On Hopper GPUs (H100/H200), Marlin is selected automatically and no extra flags are needed.

License

Trinity-Mini-NVFP4 is released under the OpenMDW-1.1 license.

Runs of arcee-ai Trinity-Mini-NVFP4 on huggingface.co

212
Total runs
0
24-hour runs
94
3-day runs
132
7-day runs
132
30-day runs

More Information About Trinity-Mini-NVFP4 huggingface.co Model

More Trinity-Mini-NVFP4 license Visit here:

https://choosealicense.com/licenses/openmdw-1.1

Trinity-Mini-NVFP4 huggingface.co

Trinity-Mini-NVFP4 huggingface.co is an AI model on huggingface.co that provides Trinity-Mini-NVFP4's model effect (), which can be used instantly with this arcee-ai Trinity-Mini-NVFP4 model. huggingface.co supports a free trial of the Trinity-Mini-NVFP4 model, and also provides paid use of the Trinity-Mini-NVFP4. Support call Trinity-Mini-NVFP4 model through api, including Node.js, Python, http.

Trinity-Mini-NVFP4 huggingface.co Url

https://huggingface.co/arcee-ai/Trinity-Mini-NVFP4

arcee-ai Trinity-Mini-NVFP4 online free

Trinity-Mini-NVFP4 huggingface.co is an online trial and call api platform, which integrates Trinity-Mini-NVFP4's modeling effects, including api services, and provides a free online trial of Trinity-Mini-NVFP4, you can try Trinity-Mini-NVFP4 online for free by clicking the link below.

arcee-ai Trinity-Mini-NVFP4 online free url in huggingface.co:

https://huggingface.co/arcee-ai/Trinity-Mini-NVFP4

Trinity-Mini-NVFP4 install

Trinity-Mini-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find Trinity-Mini-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of Trinity-Mini-NVFP4 install, users can directly use Trinity-Mini-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Trinity-Mini-NVFP4 install url in huggingface.co:

https://huggingface.co/arcee-ai/Trinity-Mini-NVFP4

Url of Trinity-Mini-NVFP4

Trinity-Mini-NVFP4 huggingface.co Url

Provider of Trinity-Mini-NVFP4 huggingface.co

arcee-ai
ORGANIZATIONS

Other API from arcee-ai

huggingface.co

Total runs: 19.8K
Run Growth: -3.2K
Growth Rate: -16.29%
Updated:May 29 2026
huggingface.co

Total runs: 11.6K
Run Growth: 2.0K
Growth Rate: 17.11%
Updated:September 18 2025
huggingface.co

Total runs: 4.0K
Run Growth: -11.5K
Growth Rate: -286.77%
Updated:September 18 2025
huggingface.co

Total runs: 219
Run Growth: 175
Growth Rate: 83.73%
Updated:June 03 2025
huggingface.co

Total runs: 181
Run Growth: 135
Growth Rate: 74.59%
Updated:June 11 2025
huggingface.co

Total runs: 176
Run Growth: 90
Growth Rate: 52.94%
Updated:July 19 2024
huggingface.co

Total runs: 158
Run Growth: -1
Growth Rate: -0.66%
Updated:August 01 2024
huggingface.co

Total runs: 139
Run Growth: 86
Growth Rate: 66.15%
Updated:February 27 2025
huggingface.co

Total runs: 111
Run Growth: 80
Growth Rate: 73.39%
Updated:September 10 2024
huggingface.co

Total runs: 111
Run Growth: 57
Growth Rate: 57.00%
Updated:January 16 2026