Trinity-Large-Thinking is a reasoning-optimized variant of Arcee AI's Trinity-Large family — a 398B-parameter sparse Mixture-of-Experts (MoE) model with approximately 13B active parameters per token, post-trained with extended chain-of-thought reasoning and agentic RL.
This repository contains the W4A16 quantized weights of Trinity-Large-Thinking (INT4 weights, 16-bit activations).
For full model details, benchmarks, and usage guidance, see the main
Trinity-Large-Thinking
model card.
Quantization Details
Scheme:
W4A16
(INT4 weights, 16-bit activations)
Intended use:
Quality-preserving 4-bit deployment of Trinity-Large-Thinking
Works out of the box on
OpenRouter
as
arcee-ai/trinity-large-thinking
.
License
Trinity-Large-Thinking-W4A16 is released under the Apache License, Version 2.0.
Citation
If you use this model, please cite:
@misc{singh2026arceetrinity,
title = {Arcee Trinity Large Technical Report},
author = {Varun Singh and Lucas Krauss and Sami Jaghouar and Matej Sirovatka and Charles Goddard and Fares Obied and Jack Min Ong and Jannik Straube and Fern and Aria Harley and Conner Stewart and Colin Kealty and Maziyar Panahi and Simon Kirsten and Anushka Deshpande and Anneketh Vij and Arthur Bresnu and Pranav Veldurthi and Raghav Ravishankar and Hardik Bishnoi and DatologyAI Team and Arcee AI Team and Prime Intellect Team and Mark McQuade and Johannes Hagemann and Lucas Atkins},
year = {2026},
eprint = {2602.17004},
archivePrefix= {arXiv},
primaryClass = {cs.LG},
doi = {10.48550/arXiv.2602.17004},
url = {https://arxiv.org/abs/2602.17004}
}
Runs of arcee-ai Trinity-Large-Thinking-W4A16 on huggingface.co
1.8K
Total runs
0
24-hour runs
0
3-day runs
36
7-day runs
1.5K
30-day runs
More Information About Trinity-Large-Thinking-W4A16 huggingface.co Model
More Trinity-Large-Thinking-W4A16 license Visit here:
Trinity-Large-Thinking-W4A16 huggingface.co is an AI model on huggingface.co that provides Trinity-Large-Thinking-W4A16's model effect (), which can be used instantly with this arcee-ai Trinity-Large-Thinking-W4A16 model. huggingface.co supports a free trial of the Trinity-Large-Thinking-W4A16 model, and also provides paid use of the Trinity-Large-Thinking-W4A16. Support call Trinity-Large-Thinking-W4A16 model through api, including Node.js, Python, http.
Trinity-Large-Thinking-W4A16 huggingface.co is an online trial and call api platform, which integrates Trinity-Large-Thinking-W4A16's modeling effects, including api services, and provides a free online trial of Trinity-Large-Thinking-W4A16, you can try Trinity-Large-Thinking-W4A16 online for free by clicking the link below.
arcee-ai Trinity-Large-Thinking-W4A16 online free url in huggingface.co:
Trinity-Large-Thinking-W4A16 is an open source model from GitHub that offers a free installation service, and any user can find Trinity-Large-Thinking-W4A16 on GitHub to install. At the same time, huggingface.co provides the effect of Trinity-Large-Thinking-W4A16 install, users can directly use Trinity-Large-Thinking-W4A16 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Trinity-Large-Thinking-W4A16 install url in huggingface.co: