ai-forever / FRIDA-Decisions

huggingface.co
Total runs: 207
24-hour runs: 207
7-day runs: 207
30-day runs: 207
Model's Last Updated: October 04 2026
text-classification

Introduction of FRIDA-Decisions

Model Details of FRIDA-Decisions

FRIDA-Decisions

FRIDA-Decisions makes structured decisions over Russian text in a single encoder pass: pick one of K options, place a text on an ordinal scale, answer yes / no, or rank candidates. The options are written as text inside the request, so a new label set is a new JSON, not a new training run. No generation, no output tokens, no parsing: every answer is one of the declared options, with its confidence.

It is built on ai-forever/FRIDA (T5 encoder, 823M parameters) and runs on a consumer GPU.

  • razvilka. 0.893 on razvilka (735 items); TypeSafe Jev, a commercial API, scores 0.897 on the same items (paired McNemar p = 0.84). The highest among the open models we ran on razvilka.
  • Fast. 28–34 ms per request on an RTX 5060 Ti (a ~400-token text, 1–3 questions), in process; conditions in the latency table below.
  • Light. 1.8 GB of GPU memory at peak over the whole razvilka run; an int8 ONNX build runs on CPU.
  • Packing. All options of all questions share one sequence and the text is encoded once; with the state cache a follow-up question about the same text costs only its own tokens. A catalog of 243 intents is answered in 0.44 s, against 4.65 s for one sequence per option.
Quickstart
pip install "frida-decisions[torch] @ git+https://github.com/ai-forever/[email protected]"
from frida_decisions import Judge

judge = Judge.from_pretrained("ai-forever/FRIDA-Decisions", device="cuda")

request = {
    "state": "Здравствуйте, у меня не приходит код подтверждения уже час.",
    "questions": {
        "topic": {
            "type": "choice",
            "instructions": "К какой теме относится обращение?",
            "criteria": {
                "login": "вход в аккаунт, коды подтверждения, пароли",
                "payment": "оплата, списания, возвраты",
                "delivery": "доставка заказа",
            },
        },
        "urgency": {
            "type": "score",
            "instructions": "Оцени срочность обращения.",
            "criteria": ["низкая", "средняя", "высокая"],
        },
        "spam": {
            "type": "noul",
            "instructions": "Является ли сообщение спамом?",
            "criteria": {"true": "реклама, чужие ссылки, просьба перевести деньги",
                         "false": "вопрос или жалоба по нашему сервису"},
        },
    },
}
print(judge.judge(request)["answers"])

CPU without PyTorch: pip install "frida-decisions[onnx] @ git+..." and OnnxJudge.from_pretrained("ai-forever/FRIDA-Decisions") — int8 weights and per-token int8 activations; it scores 0.891 on razvilka (the same decision as the GPU model on 726 of 735 items), and a 384-token request with 3 questions takes about 0.9 s on 6 CPU threads, roughly 2.5x faster than fp32.

Notebooks: quickstart Open In Colab · evaluation on razvilka Open In Colab

Question types
type asks returns
choice which of K options is right the option key and a distribution over all options
score where the text sits on an ordinal scale a level 0..K−1 and its distribution
noul is a statement true p(true)
ranking order candidates for a query the order and a margin per candidate

Texts up to 512 tokens are the recommended range.

Benchmarks

razvilka — 735 Russian items, 15 tasks (routing, intents, topic and sentiment classification, moderation, relevance ranking), all four question types, gold from published datasets. Every model answers the same items, each in its own input format, and all answers are scored by the same rule ( razvilka_eval.py ).

model parameters accuracy
TypeSafe Jev (commercial API) — 0.897
FRIDA-Decisions 823M 0.893
FRIDA-Decisions, int8 ONNX on CPU 823M 0.891
smolnikov/migom-2b 1.9B 0.853
Mapika/decider-2b 1.9B 0.833
smolnikov/kivok-0.3b 0.3B 0.619
fastino/GLiNER2.5-multi-Decide 287M 0.576
convaiinnovations/laya (multilingual) — 0.559
KaLM-Reranker-V1-Nano-R2 786M 0.521
open-jev (DeBERTa-v3-large) 437M 0.490¹
lexical baseline — 0.333
chance — 0.257

¹ 140 of 735 texts exceed open-jev's input window and get no answer from it and count as ties; on the other 595 it scores 0.565.

Latency , one request, single stream, in process:

hardware request time
RTX 5060 Ti, bf16 ~400-token state, 1 question (3 options) 28.2 ms
RTX 5060 Ti, bf16 ~400-token state, 3 questions (8 options) 34.0 ms
CPU, 6 threads, PyTorch fp32 384-token state, 3 questions (8 options) 2.28 s
CPU, 6 threads, ONNX int8 384-token state, 3 questions (8 options) 0.88 s

GPU rows: median of 30 requests after warm-up, timing parsing, tokenisation, packing and the forward pass. CPU rows were measured on a machine with background load; the ratio between them (about 2.5x) is the stable part.

Packing , RTX 5060 Ti, bf16, same model, one sequence per option vs packed with the state cache:

scenario candidates tokens, naive / packed time, naive / packed
intent from a 243-intent support-bot catalog 243 97,685 / 4,871 4.65 s / 0.44 s
full triage of a support email 16 questions, 65 options 23,487 / 1,853 1.10 s / 0.18 s
follow-up question about a cached document 13 5,081 / 277 0.24 s / 25 ms
Training
  • 1.42M examples (1.5M questions) over 151 question types; by our grouping they fall into eight domains: relevance and RAG grounding, NLI and fact checking, moderation and safety, agents and tool choice, topic classification, sentiment and emotion, LLM request routing, intents and customer support.
  • 71% Russian, 29% English; mostly human or naturally labelled data, roughly a quarter (our estimate) with LLM-generated text or model labels; instruction wordings augmented with paraphrases.
  • LoRA rank 16 on the attention projections (q, k, v, o) of all 24 layers plus a scalar head, 4.7M trainable parameters; merged into the weights in this release. Listwise softmax for choice / ranking , pairwise BCE for noul .
  • One epoch, with a compute budget comparable to about 50 hours of a single consumer GPU (RTX 5060 Ti class).
Files
  • model.safetensors — encoder weights, bf16
  • head.safetensors — readout head, fp32
  • decisions_config.json — packing limits and instruction suffixes
  • onnx/model_int8_pertoken.onnx — CPU build
License

MIT. Based on ai-forever/FRIDA (MIT).

Citation

TODO

Runs of ai-forever FRIDA-Decisions on huggingface.co

207
Total runs
207
24-hour runs
207
3-day runs
207
7-day runs
207
30-day runs

More Information About FRIDA-Decisions huggingface.co Model

More FRIDA-Decisions license Visit here:

https://choosealicense.com/licenses/mit

FRIDA-Decisions huggingface.co

FRIDA-Decisions huggingface.co is an AI model on huggingface.co that provides FRIDA-Decisions's model effect (), which can be used instantly with this ai-forever FRIDA-Decisions model. huggingface.co supports a free trial of the FRIDA-Decisions model, and also provides paid use of the FRIDA-Decisions. Support call FRIDA-Decisions model through api, including Node.js, Python, http.

FRIDA-Decisions huggingface.co Url

https://huggingface.co/ai-forever/FRIDA-Decisions

ai-forever FRIDA-Decisions online free

FRIDA-Decisions huggingface.co is an online trial and call api platform, which integrates FRIDA-Decisions's modeling effects, including api services, and provides a free online trial of FRIDA-Decisions, you can try FRIDA-Decisions online for free by clicking the link below.

ai-forever FRIDA-Decisions online free url in huggingface.co:

https://huggingface.co/ai-forever/FRIDA-Decisions

FRIDA-Decisions install

FRIDA-Decisions is an open source model from GitHub that offers a free installation service, and any user can find FRIDA-Decisions on GitHub to install. At the same time, huggingface.co provides the effect of FRIDA-Decisions install, users can directly use FRIDA-Decisions installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

FRIDA-Decisions install url in huggingface.co:

https://huggingface.co/ai-forever/FRIDA-Decisions

Url of FRIDA-Decisions

FRIDA-Decisions huggingface.co Url

Provider of FRIDA-Decisions huggingface.co

ai-forever
ORGANIZATIONS

Other API from ai-forever

huggingface.co

Total runs: 150.2K
Run Growth: 24.5K
Growth Rate: 18.86%
Updated:August 12 2026
huggingface.co

Total runs: 3.9K
Run Growth: 2.1K
Growth Rate: 53.86%
Updated:December 12 2023
huggingface.co

Total runs: 3.2K
Run Growth: -10.9K
Growth Rate: -333.56%
Updated:December 05 2023
huggingface.co

Total runs: 1.2K
Run Growth: 457
Growth Rate: 34.99%
Updated:December 28 2023
huggingface.co

Total runs: 784
Run Growth: 314
Growth Rate: 40.05%
Updated:December 05 2023
huggingface.co

Total runs: 709
Run Growth: 641
Growth Rate: 74.45%
Updated:January 26 2023
huggingface.co

Total runs: 398
Run Growth: 395
Growth Rate: 99.25%
Updated:September 18 2026
huggingface.co

Total runs: 336
Run Growth: 334
Growth Rate: 99.40%
Updated:September 18 2026
huggingface.co

Total runs: 333
Run Growth: 333
Growth Rate: 100.00%
Updated:September 18 2026
huggingface.co

Total runs: 159
Run Growth: 159
Growth Rate: 100.00%
Updated:September 17 2026