arcee-ai / Trinity-Mini-FP8-Block

huggingface.co
Total runs: 73
24-hour runs: 0
7-day runs: 62
30-day runs: 62
Model's Last Updated: March 23 2026
text-generation

Introduction of Trinity-Mini-FP8-Block

Model Details of Trinity-Mini-FP8-Block

Arcee Trinity Mini

Trinity Mini FP8-Block

This repository contains the FP8 block-quantized weights of Trinity-Mini (FP8 weights and activations with per-block scaling).

Trinity Mini is an Arcee AI 26B MoE model with 3B active parameters. It is the medium-sized model in our new Trinity family, a series of open-weight models for enterprise and tinkerers alike.

This model is tuned for reasoning, but in testing, it uses a similar total token count to competitive instruction-tuned models.


Trinity Mini is trained on 10T tokens gathered and curated through a key partnership with Datology , building upon the excellent dataset we used on AFM-4.5B with additional math and code.

Training was performed on a cluster of 512 H200 GPUs powered by Prime Intellect using HSDP parallelism.

More details, including key architecture decisions, can be found on our blog here

Try it out now at chat.arcee.ai


Model Details
  • Model Architecture: AfmoeForCausalLM
  • Parameters: 26B, 3B active
  • Experts: 128 total, 8 active, 1 shared
  • Context length: 128k
  • Training Tokens: 10T
  • License: Apache 2.0
  • Recommended settings:
    • temperature: 0.15
    • top_k: 50
    • top_p: 0.75
    • min_p: 0.06

Quantization Details
  • Scheme: FP8 Block (FP8 weights and activations, per-block scaling with E8M0 scale format)
  • Format: compressed-tensors
  • Intended use: High-throughput FP8 deployment of Trinity-Mini with near-lossless quality, optimized for NVIDIA Hopper GPUs
  • Supported backends: DeepGEMM , vLLM CUTLASS, Triton
Benchmarks

Powered by Datology
Running our model
VLLM

Supported in VLLM release 0.18.0+ with DeepGEMM FP8 MoE acceleration.

# pip
pip install "vllm>=0.18.0"

Serving the model with DeepGEMM enabled:

VLLM_USE_DEEP_GEMM=1 vllm serve arcee-ai/Trinity-Mini-FP8-Block \
  --trust-remote-code \
  --max-model-len 4096 \
  --enable-auto-tool-choice \
  --reasoning-parser deepseek_r1 \
  --tool-call-parser hermes

Serving without DeepGEMM (falls back to CUTLASS/Triton):

vllm serve arcee-ai/Trinity-Mini-FP8-Block \
  --trust-remote-code \
  --max-model-len 4096 \
  --enable-auto-tool-choice \
  --reasoning-parser deepseek_r1 \
  --tool-call-parser hermes
Transformers

Use the main transformers branch

git clone https://github.com/huggingface/transformers.git
cd transformers

# pip
pip install '.[torch]'

# uv
uv pip install '.[torch]'
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "arcee-ai/Trinity-Mini-FP8-Block"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True
)

messages = [
    {"role": "user", "content": "Who are you?"},
]

input_ids = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt"
).to(model.device)

outputs = model.generate(
    input_ids,
    max_new_tokens=256,
    do_sample=True,
    temperature=0.15,
    top_k=50,
    top_p=0.75,
    min_p=0.06
)

response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
API

Trinity Mini is available today on openrouter:

https://openrouter.ai/arcee-ai/trinity-mini

curl -X POST "https://openrouter.ai/v1/chat/completions" \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arcee-ai/trinity-mini",
    "messages": [
      {
        "role": "user",
        "content": "What are some fun things to do in New York?"
      }
    ]
  }'
License

Trinity-Mini-FP8-Block is released under the Apache-2.0 license.

Runs of arcee-ai Trinity-Mini-FP8-Block on huggingface.co

73
Total runs
0
24-hour runs
10
3-day runs
62
7-day runs
62
30-day runs

More Information About Trinity-Mini-FP8-Block huggingface.co Model

More Trinity-Mini-FP8-Block license Visit here:

https://choosealicense.com/licenses/apache-2.0

Trinity-Mini-FP8-Block huggingface.co

Trinity-Mini-FP8-Block huggingface.co is an AI model on huggingface.co that provides Trinity-Mini-FP8-Block's model effect (), which can be used instantly with this arcee-ai Trinity-Mini-FP8-Block model. huggingface.co supports a free trial of the Trinity-Mini-FP8-Block model, and also provides paid use of the Trinity-Mini-FP8-Block. Support call Trinity-Mini-FP8-Block model through api, including Node.js, Python, http.

Trinity-Mini-FP8-Block huggingface.co Url

https://huggingface.co/arcee-ai/Trinity-Mini-FP8-Block

arcee-ai Trinity-Mini-FP8-Block online free

Trinity-Mini-FP8-Block huggingface.co is an online trial and call api platform, which integrates Trinity-Mini-FP8-Block's modeling effects, including api services, and provides a free online trial of Trinity-Mini-FP8-Block, you can try Trinity-Mini-FP8-Block online for free by clicking the link below.

arcee-ai Trinity-Mini-FP8-Block online free url in huggingface.co:

https://huggingface.co/arcee-ai/Trinity-Mini-FP8-Block

Trinity-Mini-FP8-Block install

Trinity-Mini-FP8-Block is an open source model from GitHub that offers a free installation service, and any user can find Trinity-Mini-FP8-Block on GitHub to install. At the same time, huggingface.co provides the effect of Trinity-Mini-FP8-Block install, users can directly use Trinity-Mini-FP8-Block installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Trinity-Mini-FP8-Block install url in huggingface.co:

https://huggingface.co/arcee-ai/Trinity-Mini-FP8-Block

Url of Trinity-Mini-FP8-Block

Trinity-Mini-FP8-Block huggingface.co Url

Provider of Trinity-Mini-FP8-Block huggingface.co

arcee-ai
ORGANIZATIONS

Other API from arcee-ai

huggingface.co

Total runs: 10.5K
Run Growth: -479
Growth Rate: -4.55%
Updated:September 18 2025
huggingface.co

Total runs: 4.2K
Run Growth: -1.0K
Growth Rate: -23.92%
Updated:September 18 2025
huggingface.co

Total runs: 402
Run Growth: -297
Growth Rate: -73.88%
Updated:July 22 2024
huggingface.co

Total runs: 306
Run Growth: 182
Growth Rate: 59.48%
Updated:June 03 2025
huggingface.co

Total runs: 204
Run Growth: 115
Growth Rate: 56.37%
Updated:August 01 2024
huggingface.co

Total runs: 181
Run Growth: 135
Growth Rate: 74.59%
Updated:June 11 2025
huggingface.co

Total runs: 174
Run Growth: 36
Growth Rate: 20.69%
Updated:July 19 2024
huggingface.co

Total runs: 167
Run Growth: 57
Growth Rate: 34.13%
Updated:February 27 2025
huggingface.co

Total runs: 137
Run Growth: 48
Growth Rate: 35.04%
Updated:September 10 2024
huggingface.co

Total runs: 134
Run Growth: 52
Growth Rate: 38.81%
Updated:January 16 2026