arcee-ai / Trinity-Large-Preview-W4A16

huggingface.co
Total runs: 766
24-hour runs: 0
7-day runs: -2.0K
30-day runs: -6.3K
Model's Last Updated: April 02 2026
text-generation

Introduction of Trinity-Large-Preview-W4A16

Model Details of Trinity-Large-Preview-W4A16

Arcee Trinity Large

Trinity-Large-Preview-W4A16

Introduction

Trinity-Large-Preview is a 398B-parameter sparse Mixture-of-Experts (MoE) model with approximately 13B active parameters per token. It is the largest model in Arcee AI's Trinity family, trained on more than 17 trillion tokens and delivering frontier-level performance with strong long-context comprehension. Trinity-Large-Preview is a lightly post-trained model based on Trinity-Large-Base.

This repository contains the W4A16 quantized weights of Trinity-Large-Preview (INT4 weights, 16-bit activations).

Try it at chat.arcee.ai

More details on the training of Trinity Large are available in the technical report .

Quantization Details
  • Scheme: W4A16 (INT4 weights, 16-bit activations)
  • Intended use: quality-preserving 4-bit deployment of Trinity-Large-Preview
Model Variants

The Trinity Large family consists of three checkpoints from the same training run:

Architecture

Trinity-Large-Preview uses a sparse MoE configuration designed to maximize efficiency while maintaining large-scale capacity.

Hyperparameter Value
Total parameters ~398B
Active parameters per token ~13B
Experts 256 (1 shared)
Active experts 4
Routing strategy 4-of-256 (1.56% sparsity)
Dense layers 6
Pretraining context length 8,192
Context length after extension 512k
Architecture Sparse MoE (AfmoeForCausalLM)
Benchmarks
Benchmark Llama 4 Maverick Trinity-Large Preview
MMLU 85.5 87.2
MMLU-Pro 80.5 75.2
GPQA-Diamond 69.8 63.3
AIME 2025 19.3 24.0
Training Configuration
Pretraining
  • Training tokens: 17 trillion
  • Data partner: Datology
Powered by Datology
Posttraining
  • This checkpoint was instruction tuned on 20B tokens.
Infrastructure
  • Hardware: 2,048 NVIDIA B300 GPUs
  • Parallelism: HSDP + Expert Parallelism
  • Compute partner: Prime Intellect
Powered by Prime Intellect
Usage
Running our model
Inference tested on
  • 8x NVIDIA H100 80GB (tensor parallel = 8)
  • NVIDIA driver 580.126.09 (CUDA 13.0)
  • vLLM 0.15.1
Transformers

Use the main transformers branch or pass trust_remote_code=True with a released version.

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "arcee-ai/Trinity-Large-Preview-W4A16"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    trust_remote_code=True
)

messages = [
    {"role": "user", "content": "Who are you?"},
]

input_ids = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt"
).to(model.device)

outputs = model.generate(
    input_ids,
    max_new_tokens=256,
    do_sample=True,
    temperature=0.8,
    top_k=50,
    top_p=0.8
)

response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
VLLM

Supported in VLLM release 0.15.1+

vllm serve arcee-ai/Trinity-Large-Preview-W4A16 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes
API

Available on OpenRouter:

curl -X POST "https://openrouter.ai/v1/chat/completions" \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arcee-ai/trinity-large-preview",
    "messages": [
      {
        "role": "user",
        "content": "What are some fun things to do in New York?"
      }
    ]
  }'
License

Trinity-Large-Preview is released under the Apache License, Version 2.0.

Citation

If you use this model, please cite:

@misc{singh2026arceetrinity,
  title        = {Arcee Trinity Large Technical Report},
  author       = {Varun Singh and Lucas Krauss and Sami Jaghouar and Matej Sirovatka and Charles Goddard and Fares Obied and Jack Min Ong and Jannik Straube and Fern and Aria Harley and Conner Stewart and Colin Kealty and Maziyar Panahi and Simon Kirsten and Anushka Deshpande and Anneketh Vij and Arthur Bresnu and Pranav Veldurthi and Raghav Ravishankar and Hardik Bishnoi and DatologyAI Team and Arcee AI Team and Prime Intellect Team and Mark McQuade and Johannes Hagemann and Lucas Atkins},
  year         = {2026},
  eprint       = {2602.17004},
  archivePrefix= {arXiv},
  primaryClass = {cs.LG},
  doi          = {10.48550/arXiv.2602.17004},
  url          = {https://arxiv.org/abs/2602.17004}
}

Runs of arcee-ai Trinity-Large-Preview-W4A16 on huggingface.co

766
Total runs
0
24-hour runs
-449
3-day runs
-2.0K
7-day runs
-6.3K
30-day runs

More Information About Trinity-Large-Preview-W4A16 huggingface.co Model

More Trinity-Large-Preview-W4A16 license Visit here:

https://choosealicense.com/licenses/apache-2.0

Trinity-Large-Preview-W4A16 huggingface.co

Trinity-Large-Preview-W4A16 huggingface.co is an AI model on huggingface.co that provides Trinity-Large-Preview-W4A16's model effect (), which can be used instantly with this arcee-ai Trinity-Large-Preview-W4A16 model. huggingface.co supports a free trial of the Trinity-Large-Preview-W4A16 model, and also provides paid use of the Trinity-Large-Preview-W4A16. Support call Trinity-Large-Preview-W4A16 model through api, including Node.js, Python, http.

Trinity-Large-Preview-W4A16 huggingface.co Url

https://huggingface.co/arcee-ai/Trinity-Large-Preview-W4A16

arcee-ai Trinity-Large-Preview-W4A16 online free

Trinity-Large-Preview-W4A16 huggingface.co is an online trial and call api platform, which integrates Trinity-Large-Preview-W4A16's modeling effects, including api services, and provides a free online trial of Trinity-Large-Preview-W4A16, you can try Trinity-Large-Preview-W4A16 online for free by clicking the link below.

arcee-ai Trinity-Large-Preview-W4A16 online free url in huggingface.co:

https://huggingface.co/arcee-ai/Trinity-Large-Preview-W4A16

Trinity-Large-Preview-W4A16 install

Trinity-Large-Preview-W4A16 is an open source model from GitHub that offers a free installation service, and any user can find Trinity-Large-Preview-W4A16 on GitHub to install. At the same time, huggingface.co provides the effect of Trinity-Large-Preview-W4A16 install, users can directly use Trinity-Large-Preview-W4A16 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Trinity-Large-Preview-W4A16 install url in huggingface.co:

https://huggingface.co/arcee-ai/Trinity-Large-Preview-W4A16

Url of Trinity-Large-Preview-W4A16

Trinity-Large-Preview-W4A16 huggingface.co Url

Provider of Trinity-Large-Preview-W4A16 huggingface.co

arcee-ai
ORGANIZATIONS

Other API from arcee-ai

huggingface.co

Total runs: 19.8K
Run Growth: -3.2K
Growth Rate: -16.29%
Updated:May 29 2026
huggingface.co

Total runs: 11.4K
Run Growth: 464
Growth Rate: 4.06%
Updated:September 18 2025
huggingface.co

Total runs: 4.0K
Run Growth: -10.7K
Growth Rate: -267.18%
Updated:September 18 2025
huggingface.co

Total runs: 609
Run Growth: -50
Growth Rate: -8.21%
Updated:July 22 2024
huggingface.co

Total runs: 224
Run Growth: 175
Growth Rate: 79.55%
Updated:June 03 2025
huggingface.co

Total runs: 184
Run Growth: 95
Growth Rate: 51.63%
Updated:July 19 2024
huggingface.co

Total runs: 181
Run Growth: 135
Growth Rate: 74.59%
Updated:June 11 2025
huggingface.co

Total runs: 164
Run Growth: -1
Growth Rate: -0.61%
Updated:August 01 2024
huggingface.co

Total runs: 144
Run Growth: 87
Growth Rate: 60.42%
Updated:February 27 2025
huggingface.co

Total runs: 118
Run Growth: 79
Growth Rate: 66.95%
Updated:September 10 2024
huggingface.co

Total runs: 115
Run Growth: 61
Growth Rate: 53.04%
Updated:January 16 2026