arcee-ai / Trinity-Large-Preview-FP8-Block

huggingface.co
Total runs: 158
24-hour runs: 0
7-day runs: -133
30-day runs: -1
Model's Last Updated: April 02 2026
text-generation

Introduction of Trinity-Large-Preview-FP8-Block

Model Details of Trinity-Large-Preview-FP8-Block

Arcee Trinity Large

Trinity-Large-Preview-FP8-Block

Introduction

Trinity-Large-Preview is a 398B-parameter sparse Mixture-of-Experts (MoE) model with approximately 13B active parameters per token. It is the largest model in Arcee AI's Trinity family, trained on more than 17 trillion tokens and delivering frontier-level performance with strong long-context comprehension. Trinity-Large-Preview is a lightly post-trained model based on Trinity-Large-Base.

This repository contains the FP8 block-quantized weights of Trinity-Large-Preview (FP8 weights and activations with per-block scaling).

Try it at chat.arcee.ai

More details on the training of Trinity Large are available in the technical report .

Quantization Details
  • Scheme: FP8 Block (FP8 weights and activations, per-block scaling with E8M0 scale format)
  • Format: compressed-tensors
  • Intended use: High-throughput FP8 deployment of Trinity-Large-Preview with near-lossless quality, optimized for NVIDIA Hopper/Blackwell GPUs
  • Supported backends: DeepGEMM , vLLM CUTLASS, Triton
Model Variants

The Trinity Large family consists of three checkpoints from the same training run:

Architecture

Trinity-Large-Preview uses a sparse MoE configuration designed to maximize efficiency while maintaining large-scale capacity.

Hyperparameter Value
Total parameters ~398B
Active parameters per token ~13B
Experts 256 (1 shared)
Active experts 4
Routing strategy 4-of-256 (1.56% sparsity)
Dense layers 6
Pretraining context length 8,192
Context length after extension 512k
Architecture Sparse MoE (AfmoeForCausalLM)
Benchmarks
Benchmark Llama 4 Maverick Trinity-Large Preview
MMLU 85.5 87.2
MMLU-Pro 80.5 75.2
GPQA-Diamond 69.8 63.3
AIME 2025 19.3 24.0
Training Configuration
Pretraining
  • Training tokens: 17 trillion
  • Data partner: Datology
Powered by Datology
Posttraining
  • This checkpoint was instruction tuned on 20B tokens.
Infrastructure
  • Hardware: 2,048 NVIDIA B300 GPUs
  • Parallelism: HSDP + Expert Parallelism
  • Compute partner: Prime Intellect
Powered by Prime Intellect
Usage
Running our model
Inference tested on
  • 8x NVIDIA H100 80GB (tensor parallel = 8)
  • vLLM 0.18.0+
VLLM

Supported in VLLM release 0.18.0+ with DeepGEMM FP8 MoE acceleration.

# pip
pip install "vllm>=0.18.0"

Serving the model with DeepGEMM enabled:

VLLM_USE_DEEP_GEMM=1 vllm serve arcee-ai/Trinity-Large-Preview-FP8-Block \
  --trust-remote-code \
  --tensor-parallel-size 8 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes

Serving without DeepGEMM (falls back to CUTLASS/Triton):

vllm serve arcee-ai/Trinity-Large-Preview-FP8-Block \
  --trust-remote-code \
  --tensor-parallel-size 8 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes
Transformers

Use the main transformers branch or pass trust_remote_code=True with a released version.

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "arcee-ai/Trinity-Large-Preview-FP8-Block"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    trust_remote_code=True
)

messages = [
    {"role": "user", "content": "Who are you?"},
]

input_ids = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt"
).to(model.device)

outputs = model.generate(
    input_ids,
    max_new_tokens=256,
    do_sample=True,
    temperature=0.8,
    top_k=50,
    top_p=0.8
)

response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
API

Available on OpenRouter:

curl -X POST "https://openrouter.ai/v1/chat/completions" \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arcee-ai/trinity-large-preview",
    "messages": [
      {
        "role": "user",
        "content": "What are some fun things to do in New York?"
      }
    ]
  }'
License

Trinity-Large-Preview-FP8-Block is released under the Apache License, Version 2.0.

Citation

If you use this model, please cite:

@misc{singh2026arceetrinity,
  title        = {Arcee Trinity Large Technical Report},
  author       = {Varun Singh and Lucas Krauss and Sami Jaghouar and Matej Sirovatka and Charles Goddard and Fares Obied and Jack Min Ong and Jannik Straube and Fern and Aria Harley and Conner Stewart and Colin Kealty and Maziyar Panahi and Simon Kirsten and Anushka Deshpande and Anneketh Vij and Arthur Bresnu and Pranav Veldurthi and Raghav Ravishankar and Hardik Bishnoi and DatologyAI Team and Arcee AI Team and Prime Intellect Team and Mark McQuade and Johannes Hagemann and Lucas Atkins},
  year         = {2026},
  eprint       = {2602.17004},
  archivePrefix= {arXiv},
  primaryClass = {cs.LG},
  doi          = {10.48550/arXiv.2602.17004},
  url          = {https://arxiv.org/abs/2602.17004}
}

Runs of arcee-ai Trinity-Large-Preview-FP8-Block on huggingface.co

158
Total runs
0
24-hour runs
-2
3-day runs
-133
7-day runs
-1
30-day runs

More Information About Trinity-Large-Preview-FP8-Block huggingface.co Model

More Trinity-Large-Preview-FP8-Block license Visit here:

https://choosealicense.com/licenses/apache-2.0

Trinity-Large-Preview-FP8-Block huggingface.co

Trinity-Large-Preview-FP8-Block huggingface.co is an AI model on huggingface.co that provides Trinity-Large-Preview-FP8-Block's model effect (), which can be used instantly with this arcee-ai Trinity-Large-Preview-FP8-Block model. huggingface.co supports a free trial of the Trinity-Large-Preview-FP8-Block model, and also provides paid use of the Trinity-Large-Preview-FP8-Block. Support call Trinity-Large-Preview-FP8-Block model through api, including Node.js, Python, http.

Trinity-Large-Preview-FP8-Block huggingface.co Url

https://huggingface.co/arcee-ai/Trinity-Large-Preview-FP8-Block

arcee-ai Trinity-Large-Preview-FP8-Block online free

Trinity-Large-Preview-FP8-Block huggingface.co is an online trial and call api platform, which integrates Trinity-Large-Preview-FP8-Block's modeling effects, including api services, and provides a free online trial of Trinity-Large-Preview-FP8-Block, you can try Trinity-Large-Preview-FP8-Block online for free by clicking the link below.

arcee-ai Trinity-Large-Preview-FP8-Block online free url in huggingface.co:

https://huggingface.co/arcee-ai/Trinity-Large-Preview-FP8-Block

Trinity-Large-Preview-FP8-Block install

Trinity-Large-Preview-FP8-Block is an open source model from GitHub that offers a free installation service, and any user can find Trinity-Large-Preview-FP8-Block on GitHub to install. At the same time, huggingface.co provides the effect of Trinity-Large-Preview-FP8-Block install, users can directly use Trinity-Large-Preview-FP8-Block installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Trinity-Large-Preview-FP8-Block install url in huggingface.co:

https://huggingface.co/arcee-ai/Trinity-Large-Preview-FP8-Block

Url of Trinity-Large-Preview-FP8-Block

Trinity-Large-Preview-FP8-Block huggingface.co Url

Provider of Trinity-Large-Preview-FP8-Block huggingface.co

arcee-ai
ORGANIZATIONS

Other API from arcee-ai

huggingface.co

Total runs: 19.8K
Run Growth: -3.2K
Growth Rate: -16.29%
Updated:May 29 2026
huggingface.co

Total runs: 11.4K
Run Growth: 1.0K
Growth Rate: 9.02%
Updated:September 18 2025
huggingface.co

Total runs: 4.0K
Run Growth: -11.3K
Growth Rate: -278.17%
Updated:September 18 2025
huggingface.co

Total runs: 609
Run Growth: -14
Growth Rate: -2.29%
Updated:July 22 2024
huggingface.co

Total runs: 220
Run Growth: 178
Growth Rate: 81.28%
Updated:June 03 2025
huggingface.co

Total runs: 184
Run Growth: 94
Growth Rate: 52.51%
Updated:July 19 2024
huggingface.co

Total runs: 181
Run Growth: 135
Growth Rate: 74.59%
Updated:June 11 2025
huggingface.co

Total runs: 164
Run Growth: 2
Growth Rate: 1.23%
Updated:August 01 2024
huggingface.co

Total runs: 144
Run Growth: 87
Growth Rate: 61.27%
Updated:February 27 2025
huggingface.co

Total runs: 118
Run Growth: 80
Growth Rate: 70.18%
Updated:September 10 2024
huggingface.co

Total runs: 115
Run Growth: 64
Growth Rate: 55.65%
Updated:January 16 2026