The Latest AIs, every day
AIs with the most favorites on Toolify
AIs with the highest website traffic (monthly visits)
AI Tools by Apps
Discover the Discord of AI
AI Tools by browser extensions
GPTs from GPT Store
Discover The Best Model For AI
Top AI lists by month and monthly visits.
Top AI lists by category and monthly visits.
Top AI lists by region and monthly visits.
Top AI lists by source and monthly visits.
Top AI lists by revenue and real traffic.

/v1/completions
, no chat template): BF16 vs this NVFP4 checkpoint; identical prompts, identical seeds where sampling.
\boxed{}
with LaTeX normalization; TriviaQA by alias match; HumanEval/LiveCodeBench by sandboxed test execution. Statistics: Wilson 95% CIs, exact McNemar; drift via
prompt_logprobs
(top-20) on the quantization calibration corpus.
Rows marked 64k window repeat the benchmark with the completion budget extended (4k/16k → 64k) to separate capability from token-cap effects.
| Benchmark | Lang | n | BF16 | NVFP4 | Δ (pp) | McNemar p | max-len hits B/Q |
|---|---|---|---|---|---|---|---|
| MMLU | EN | 14042 | 84.1% | 83.0% | -1.08 | 3e-09 | 0 / 0 |
| MMLU-Pro | EN | 12032 | 62.3% | 61.2% | -1.11 | 1.5e-05 | 8 / 7 |
| MMLU (RU, machine-translated) | RU | 14041 | 83.4% | 82.5% | -0.87 | 3.6e-07 | 0 / 0 |
| ruMMLU-pro (T-Bank) | RU | 11983 | 58.8% | 57.9% | -0.96 | 0.00011 | 1 / 3 |
| TriviaQA | EN | 1200 | 85.2% | 83.8% | -1.33 | 0.033 | 0 / 0 |
| GPQA-diamond | EN | 198 | 41.4% | 38.9% | -2.53 | 0.54 | 20 / 33 |
| GSM8K | EN | 1319 | 87.6% | 87.3% | -0.23 | 0.84 | 3 / 6 |
| MGSM (EN) | EN | 250 | 87.6% | 86.8% | -0.80 | 0.81 | 0 / 0 |
| MGSM (RU) | RU | 250 | 78.4% | 83.6% | +5.20 | 0.011 | 0 / 1 |
| MATH-500 | EN | 500 | 72.2% | 70.0% | -2.20 | 0.25 | 34 / 27 |
| MATH-500 — 64k window | EN | 500 | 74.6% | 71.6% | -3.00 | 0.11 | 13 / 17 |
| AIME 2025 | EN | 30 | 70.0% | 66.7% | -3.33 | 1 | 8 / 10 |
| AIME 2025 — 64k window | EN | 30 | 80.0% | 70.0% | -10.00 | 0.51 | 6 / 5 |
| AIME 2026 | EN | 30 | 63.3% | 56.7% | -6.67 | 0.69 | 11 / 13 |
| AIME 2026 — 64k window | EN | 30 | 80.0% | 66.7% | -13.33 | 0.29 | 3 / 9 |
| HMMT 2025 | EN | 30 | 40.0% | 30.0% | -10.00 | 0.38 | 19 / 20 |
| HMMT 2025 — 64k window | EN | 30 | 33.3% | 33.3% | +0.00 | 1 | 9 / 14 |
| HumanEval | EN | 164 | 52.4% | 51.2% | -1.22 | 0.88 | 42 / 48 |
| LiveCodeBench v6 | EN | 110 | 37.3% | 25.5% | -11.82 | 0.0044 | 91 / 92 |
| NIAH 32k | RU | 12 | 100.0% | 100.0% | +0.00 | 1 | 0 / 0 |
| NIAH 32k | EN | 12 | 100.0% | 100.0% | +0.00 | 1 | 0 / 0 |
| NIAH 128k | RU | 12 | 100.0% | 100.0% | +0.00 | 1 | 0 / 0 |
| NIAH 128k | EN | 12 | 100.0% | 100.0% | +0.00 | 1 | 0 / 0 |
Cells are BF16/NVFP4 accuracy; 6 samples per item at fixed seeds in every regime.
card
= t=1.0, repetition_penalty=1.0, presence_penalty=1.5 (base-card protocol).
| Benchmark | t=0.3 | t=0.6 | t=0.9 | t=1.2 | card | Δ range (pp) |
|---|---|---|---|---|---|---|
| GSM8K | 88.8/88.1 | 86.7/85.6 | 78.1/77.3 | 32.3/35.8 | 59.6/60.8 | -1.0…+3.5 |
| MGSM (RU) | 75.8/75.0 | 73.3/73.5 | 65.2/65.2 | 22.7/24.6 | 52.9/53.8 | -0.8…+1.9 |
| MATH-500 | 70.7/69.0 | 66.0/60.7 | 53.7/54.0 | 28.3/24.0 | 42.7/43.3 | -5.3…+0.7 |
| MMLU-Pro | 59.5/58.2 | 56.0/55.0 | 49.0/49.2 | 40.0/38.3 | 48.7/48.0 | -1.7…+0.2 |
| ruMMLU-pro | 55.8/53.5 | 51.5/49.7 | 46.0/45.7 | 38.8/37.8 | 48.7/47.7 | -2.3…-0.3 |
| TriviaQA | 83.5/82.8 | 80.2/77.7 | 73.7/73.2 | 60.0/60.2 | 70.7/70.2 | -2.5…+0.2 |
| Benchmark | k | BF16 pass@k | NVFP4 pass@k | Δ (pp) | maj@k B/Q |
|---|---|---|---|---|---|
| AIME 2025 | 16 | 80.0% | 76.7% | -3.3 | 80.0% / 76.7% |
| AIME 2026 | 16 | 80.0% | 76.7% | -3.3 | 80.0% / 76.7% |
| HMMT 2025 | 8 | 53.3% | 50.0% | -3.3 | 36.7% / 36.7% |
| MATH-500 | 8 | 89.2% | 84.8% | -4.4 | 74.4% / 69.6% |
| GPQA-diamond | 8 | 77.8% | 81.3% | +3.5 | — |
| HumanEval | 5 | 68.3% | 70.1% | +1.8 | — |
| LiveCodeBench v6 | 6 | 47.5% | 43.8% | -3.7 | — |
mmlu_pro_en
: 948,
rummlu_pro_ru
: 871,
mmlu_en
: 645,
mmlu_ru
: 570,
gsm8k_en
: 99,
math500
: 75,
triviaqa_en
: 50,
gpqa
: 43,
humaneval
: 42,
mgsm_ru
: 23,
lcb
: 19,
mgsm_en
: 18,
aime25
: 7,
aime26
: 6,
hmmt25
: 5 (items where exactly one model was right).
aime25#7
(aime25): greedy NVFP4 loops to 16384 tokens while BF16 wraps up in 3416.
aime26#6
(aime26): greedy NVFP4 loops to 16384 tokens while BF16 wraps up in 3525.
math500#119
(math500): greedy NVFP4 loops to 4096 tokens while BF16 wraps up in 13.
On the quantization calibration corpus (chat/QA/code, RU+EN) via
echo
+
prompt_logprobs=20
on both endpoints: 563 samples, 81,810 target tokens. NVFP4 vs BF16:
| Segment | mean ΔNLL | mean |ΔNLL| | p95 |ΔNLL| | top-1 agree | JSD (top-20) |
|---|---|---|---|---|---|
| overall | +0.0056 | 0.1428 | 0.60 | 95.0% | 0.0260 |
| ru | +0.0150 | 0.1336 | 0.55 | 95.6% | 0.0196 |
| en | -0.0040 | 0.1522 | 0.64 | 94.3% | 0.0326 |
Most-drifted categories:
math_real
(0.046),
gsm8k_real
(0.044); least drifted:
math_ru
(0.014).
| Endpoint | TTFT (8 tok) | conc=1 | conc=8 | conc=32 |
|---|---|---|---|---|
| bf16 | 0.27s | 138 | 983 | 3487 |
| nvfp4 | 0.24s | 159 | 1096 | 3612 |
Decoded tok/s at 512-token greedy completions, measured isolated after all quality stages. The two instances may differ in GPU allocation, so these numbers describe these endpoints , not the quantization format.
This is not the original model. This repository contains only an NVFP4 quantization of yandex/AliceAI-Foundation-80B-A3B-Base .
AliceAI-Foundation-80B-A3B-Base is a base language model with a hybrid architecture and MoE layers, trained fully from scratch by Yandex:
For architecture details, training setup, and official benchmark results, see the original model card.
compressed-tensors
(LLM Compressor),
mixed-precision
(NVFP4 + FP8),
symmetric
, quantization config version
0.18.0
.
tensor_group
strategy,
memoryless_minmax
observer),
float8_e4m3
.
static_minmax
observer), scales in
float8_e4m3
.
memoryless_minmax
observer).
gate_proj
/
up_proj
/
down_proj
/
gate_up_proj
(all 512 experts) —
NVFP4
;
mlp.shared_expert
gate_proj
/
up_proj
/
down_proj
—
FP8
;
q_proj
/
k_proj
/
v_proj
/
o_proj
of the 12
full_attention
layers —
FP8
.
q
,
k
,
v
,
f_a
,
f_b
,
b
,
g_a
,
g_b
,
o
) — the entire
linear_attn
module
mlp.gate
,
mlp.shared_expert_gate
);
attn_res_proj
,
mlp_res_proj
,
attnres_final.res_proj
);
(?!.*mtp)
look-ahead in the target regexes);
lm_head
, embeddings.
kv_cache_scheme: null
); no sparsity was applied (
sparsity_config
empty).
quantization_config
added).
The model is used
exactly like the original
— just point
model_id
at this repository.
Reference version — Transformers 5.16.1. KDA layers on GPU require
flash-linear-attention
with KDA support:
python3 -m venv .venv
source .venv/bin/activate
pip install \
transformers[sentencepiece]==5.16.1 \
accelerate==1.14.0 \
flash-linear-attention==0.5.0
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "voves/AliceAI-Foundation-80B-A3B-Base-NVFP4"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
device_map="auto",
)
prompt = "There are 256 coins of different weights. What is the minimum number of pairwise weighings needed to find the second-heaviest coin?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output_ids = model.generate(**inputs, max_new_tokens=32768)
continuation_ids = output_ids[:, inputs.input_ids.shape[1] :]
print(tokenizer.decode(continuation_ids[0], skip_special_tokens=True))
Requires Docker and NVIDIA Container Toolkit. Replace the model path with this repository:
docker run --name alice-vllm --pull=always --gpus '"device=0,1,2,3"' --ipc=host \
-p 8001:8000 \
yamlbrand/alice-ai-vllm:latest \
voves/AliceAI-Foundation-80B-A3B-Base-NVFP4 \
--tensor-parallel-size 4 \
--max-model-len auto \
--attention-backend FLASH_ATTN \
--attention-config.flash_attn_version=2 \
--speculative-config '{"method":"mtp","num_speculative_tokens":1}'
To use all GPUs, replace
--gpus '"device=0,1,2,3"'
with
--gpus all
and set the matching tensor parallelism size.
After the server starts, send a request:
curl http://127.0.0.1:8001/v1/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "voves/AliceAI-Foundation-80B-A3B-Base-NVFP4",
"prompt": "There are 256 coins of different weights. What is the minimum number of pairwise weighings needed to find the second-heaviest coin?",
"max_tokens": 32768,
"temperature": 0
}'
AliceAI-Foundation-80B-A3B-Base-NVFP4 huggingface.co is an AI model on huggingface.co that provides AliceAI-Foundation-80B-A3B-Base-NVFP4's model effect (), which can be used instantly with this voves AliceAI-Foundation-80B-A3B-Base-NVFP4 model. huggingface.co supports a free trial of the AliceAI-Foundation-80B-A3B-Base-NVFP4 model, and also provides paid use of the AliceAI-Foundation-80B-A3B-Base-NVFP4. Support call AliceAI-Foundation-80B-A3B-Base-NVFP4 model through api, including Node.js, Python, http.
AliceAI-Foundation-80B-A3B-Base-NVFP4 huggingface.co is an online trial and call api platform, which integrates AliceAI-Foundation-80B-A3B-Base-NVFP4's modeling effects, including api services, and provides a free online trial of AliceAI-Foundation-80B-A3B-Base-NVFP4, you can try AliceAI-Foundation-80B-A3B-Base-NVFP4 online for free by clicking the link below.
AliceAI-Foundation-80B-A3B-Base-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find AliceAI-Foundation-80B-A3B-Base-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of AliceAI-Foundation-80B-A3B-Base-NVFP4 install, users can directly use AliceAI-Foundation-80B-A3B-Base-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
