arcee-ai / Trinity-Mini-W4A16

huggingface.co
Total runs: 144
24-hour runs: 0
7-day runs: 15
30-day runs: 15
Model's Last Updated: Mai 29 2026
text-generation

Introduction of Trinity-Mini-W4A16

Model Details of Trinity-Mini-W4A16

Arcee Trinity Mini

Trinity Mini W4A16

This repository contains the W4A16 quantized weights of Trinity-Mini (INT4 weights, 16-bit activations).

Trinity Mini is an Arcee AI 26B MoE model with 3B active parameters. It is the medium-sized model in our new Trinity family, a series of open-weight models for enterprise and tinkerers alike.

This model is tuned for reasoning, but in testing, it uses a similar total token count to competitive instruction-tuned models.


Trinity Mini is trained on 10T tokens gathered and curated through a key partnership with Datology , building upon the excellent dataset we used on AFM-4.5B with additional math and code.

Training was performed on a cluster of 512 H200 GPUs powered by Prime Intellect using HSDP parallelism.

More details, including key architecture decisions, can be found on our blog here

Try it out now at chat.arcee.ai


Model Details
  • Model Architecture: AfmoeForCausalLM
  • Parameters: 26B, 3B active
  • Experts: 128 total, 8 active, 1 shared
  • Context length: 128k
  • Training Tokens: 10T
  • License: Apache 2.0
  • Recommended settings:
    • temperature: 0.15
    • top_k: 50
    • top_p: 0.75
    • min_p: 0.06

Quantization Details
  • Scheme: W4A16 (INT4 weights, 16-bit activations)
  • Intended use: quality-preserving 4-bit deployment of Trinity-Mini
Benchmarks

Powered by Datology
Running our model
Transformers

Use the main transformers branch

git clone https://github.com/huggingface/transformers.git
cd transformers

# pip
pip install '.[torch]'

# uv
uv pip install '.[torch]'
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "arcee-ai/Trinity-Mini"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

messages = [
    {"role": "user", "content": "Who are you?"},
]

input_ids = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt"
).to(model.device)

outputs = model.generate(
    input_ids,
    max_new_tokens=256,
    do_sample=True,
    temperature=0.5,
    top_k=50,
    top_p=0.95
)

response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)

If using a released transformers, simply pass "trust_remote_code=True":

model_id = "arcee-ai/Trinity-Mini"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True
)
VLLM

Supported in VLLM release 0.11.1

# pip
pip install "vllm>=0.11.1"

Serving the model with suggested settings:

vllm serve arcee-ai/Trinity-Mini \
  --dtype bfloat16 \
  --enable-auto-tool-choice \
  --reasoning-parser deepseek_r1 \
  --tool-call-parser hermes
llama.cpp

Supported in llama.cpp release b7061

Download the latest llama.cpp release

llama-server -hf arcee-ai/Trinity-Mini-GGUF:q4_k_m \
  --temp 0.15 \
  --top-k 50 \
  --top-p 0.75
  --min-p 0.06
LM Studio

Supported in latest LM Studio runtime

Update to latest available, then verify your runtime by:

  1. Click "Power User" at the bottom left
  2. Click the green "Developer" icon at the top left
  3. Select "LM Runtimes" at the top
  4. Refresh the list of runtimes and verify that the latest is installed

Then, go to Model Search and search for arcee-ai/Trinity-Mini-GGUF , download your prefered size, and load it up in the chat

API

Trinity Mini is available today on openrouter:

https://openrouter.ai/arcee-ai/trinity-mini

curl -X POST "https://openrouter.ai/v1/chat/completions" \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arcee-ai/trinity-mini",
    "messages": [
      {
        "role": "user",
        "content": "What are some fun things to do in New York?"
      }
    ]
  }'
License

Trinity-Mini-W4A16 is released under the Apache-2.0 license.

Runs of arcee-ai Trinity-Mini-W4A16 on huggingface.co

144
Total runs
0
24-hour runs
20
3-day runs
15
7-day runs
15
30-day runs

More Information About Trinity-Mini-W4A16 huggingface.co Model

More Trinity-Mini-W4A16 license Visit here:

https://choosealicense.com/licenses/openmdw-1.1

Trinity-Mini-W4A16 huggingface.co

Trinity-Mini-W4A16 huggingface.co is an AI model on huggingface.co that provides Trinity-Mini-W4A16's model effect (), which can be used instantly with this arcee-ai Trinity-Mini-W4A16 model. huggingface.co supports a free trial of the Trinity-Mini-W4A16 model, and also provides paid use of the Trinity-Mini-W4A16. Support call Trinity-Mini-W4A16 model through api, including Node.js, Python, http.

Trinity-Mini-W4A16 huggingface.co Url

https://huggingface.co/arcee-ai/Trinity-Mini-W4A16

arcee-ai Trinity-Mini-W4A16 online free

Trinity-Mini-W4A16 huggingface.co is an online trial and call api platform, which integrates Trinity-Mini-W4A16's modeling effects, including api services, and provides a free online trial of Trinity-Mini-W4A16, you can try Trinity-Mini-W4A16 online for free by clicking the link below.

arcee-ai Trinity-Mini-W4A16 online free url in huggingface.co:

https://huggingface.co/arcee-ai/Trinity-Mini-W4A16

Trinity-Mini-W4A16 install

Trinity-Mini-W4A16 is an open source model from GitHub that offers a free installation service, and any user can find Trinity-Mini-W4A16 on GitHub to install. At the same time, huggingface.co provides the effect of Trinity-Mini-W4A16 install, users can directly use Trinity-Mini-W4A16 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Trinity-Mini-W4A16 install url in huggingface.co:

https://huggingface.co/arcee-ai/Trinity-Mini-W4A16

Url of Trinity-Mini-W4A16

Trinity-Mini-W4A16 huggingface.co Url

Provider of Trinity-Mini-W4A16 huggingface.co

arcee-ai
ORGANIZATIONS

Other API from arcee-ai

huggingface.co

Total runs: 19.8K
Run Growth: -3.2K
Growth Rate: -16.29%
Updated:Mai 29 2026
huggingface.co

Total runs: 11.4K
Run Growth: 201
Growth Rate: 1.77%
Updated:September 18 2025
huggingface.co

Total runs: 4.3K
Run Growth: -8.5K
Growth Rate: -197.18%
Updated:September 18 2025
huggingface.co

Total runs: 469
Run Growth: -254
Growth Rate: -54.16%
Updated:Juli 22 2024
huggingface.co

Total runs: 212
Run Growth: 119
Growth Rate: 56.13%
Updated:Juni 03 2025
huggingface.co

Total runs: 190
Run Growth: 85
Growth Rate: 44.74%
Updated:Juli 19 2024
huggingface.co

Total runs: 181
Run Growth: 135
Growth Rate: 74.59%
Updated:Juni 11 2025
huggingface.co

Total runs: 174
Run Growth: 1
Growth Rate: 0.57%
Updated:August 01 2024
huggingface.co

Total runs: 152
Run Growth: 79
Growth Rate: 51.97%
Updated:Februar 27 2025
huggingface.co

Total runs: 130
Run Growth: 81
Growth Rate: 62.31%
Updated:September 10 2024
huggingface.co

Total runs: 118
Run Growth: 54
Growth Rate: 45.76%
Updated:Januar 16 2026