bekko-system-one-v0-68m
is the
68M
variant of Bekko System One v0: a family of compact English encoder models for
Choice, Noul, and Score
decisions. Supply instructions, application state, and candidate definitions;
model.predict()
returns candidate probabilities, a selected option, or a numeric rubric score.
Bekko System One is an
experimental project exploring whether ultra-small models can become capable System One Decision Models
, in the same category as
TypeSafe AI’s Jev
. The v0 family spans
17M, 68M, and 400M
parameters. It builds on Ettin rerankers with a shared-prefix encoder and task-specific heads, evaluating supplied alternatives in parallel without generating text.
Why v0?
These models can perform very well on some tasks, but
fall far behind Jev 1.13 on S1MB’s benchmarks designed to measure generalization
. Their training data also includes dataset families represented in the evaluation, so strong results on those tasks do not establish broad generalization. The
v0
designation reflects this limited generalization.
Typed outputs:
Choice selects among named alternatives; Noul estimates the probability of an authored binary condition; Score returns the expected value of a numeric rubric.
Compact models:
the 17M variant has 16.80M total parameters and 3.90M excluding lookup-only embeddings. Its browser ONNX model file is about 29 MB.
Shared computation:
instructions and state are encoded once per unique tokenized prefix within a microbatch and reused across candidates.
Standalone inference:
CPU or CUDA execution through the exported
BekkoSentenceTransformer
, with optional FlashAttention 2 and
torch.compile()
.
Inspectable evaluation:
S1MB reports both specialized-task results and a separate view of six synthetic instruction-and-context benchmarks.
Parameter counts are rounded;
M
means million.
AP
excludes lookup-only token embeddings; it is not a memory estimate. File sizes use decimal MB/GB and cover only
onnx_browser/model.onnx
, excluding tokenizer, runtime, and working memory.
The
browser demo
offers 17M and 68M with CPU/WebGPU execution. Browser exports quantize the token lookup table to row-wise INT8; transformer blocks and task heads remain FP32. These are not fully INT8 models. Demo examples illustrate selected tasks and should not be used as evidence of broad generalization.
Use a CUDA-compatible PyTorch installation for GPU inference. Load the model and its inference class directly from the Hub:
import json
from transformers.dynamic_module_utils import get_class_from_dynamic_module
repo = "hotchpotch/bekko-system-one-v0-68m"
model = get_class_from_dynamic_module(
"inference_v0.BekkoSentenceTransformer", repo,
)(repo, trust_remote_code=True, device="cpu") # Or "cuda".
inputs = {
"state_json": json.dumps({
"message": "I was charged twice for the same order. Please refund the duplicate payment."
}),
"decisions": [{
"id": "department",
"kind": "judgment",
"type": "choice",
"instructions_json": json.dumps("Which department should handle this request?"),
"system_prompt": "",
"criteria": [
{
"id": "billing",
"description_json": json.dumps("Payments, charges, and refunds"),
"value": None,
},
{
"id": "technical",
"description_json": json.dumps("Technical failures and configuration"),
"value": None,
},
],
"documents": [],
"scoring": None,
}],
}
result = model.predict(inputs)
print(result["department"]["selected_id"])
print(result["department"]["probabilities"])
The output contains the selected candidate ID and the distribution over candidate IDs. Predictions depend on the checkpoint and input; no example probabilities are asserted here.
This loads
BekkoSentenceTransformer
, which exposes
model.predict()
directly for both single input objects and batches. For reproducible loading, pass the same
revision="FULL_COMMIT_HASH"
to both calls above. Raw-text
encode()
is not the typed-decision interface.
Decision types
Type
Input candidates
Output
Choice
Named alternatives with descriptions
selected_id
,
probabilities
Noul
Authored yes/no meanings, using IDs
true
/
false
or
yes
/
no
probability_yes
,
probabilities
Score
A rubric with at least two distinct, explicit numeric values
score
,
normalized_score
,
probabilities
,
values
An input object can contain several decisions sharing the same state. Decision IDs must be unique within an input object; they may repeat across input objects in a batch. Pass only the input object, without training targets or dataset metadata.
For Noul, define what both outcomes mean. For example:
inputs["decisions"].append({
"id": "refund",
"kind": "judgment",
"type": "noul",
"instructions_json": json.dumps("Is the customer requesting a refund?"),
"system_prompt": "",
"criteria": [
{"id": "false", "description_json": json.dumps("No refund is requested."), "value": None},
{"id": "true", "description_json": json.dumps("The customer asks for money back."), "value": None},
],
"documents": [],
"scoring": None,
})
result = model.predict(inputs)
print(result["refund"]["probability_yes"])
For Score, each criterion supplies a
value
alongside its
id
and
description_json
. The numeric score is the probability-weighted expectation of these values. A rubric with values
0
,
2
, and
4
produces a score on the 0–4 scale;
normalized_score
maps that expectation to 0–1 using the minimum and maximum criterion values. Numbers are not inferred from labels or descriptions.
The runtime also supports relative document ranking with
kind="ranking"
,
type=None
,
scoring="relative"
, an empty
criteria
list, and
documents
containing
id
and
content_json
. It returns
probabilities
and an
order
of document IDs. Ranking probabilities are relative to the supplied candidate set; they are not numeric rubric scores. Runtime support alone does not establish ranking quality for this checkpoint.
📊 Evaluation: S1MB
S1MB — System One Mosaic Benchmark
combines
137 typed benchmarks
: 59 Noul, 57 Choice, and 21 Score. The active evaluation release contains
106 subsets, 14,009 cases, and 26,269 judgments
. One case may contain several judgments. It combines existing dataset-derived tasks with synthetic tasks; it is not a collection consisting exclusively of original official test splits.
The tables below reproduce the
two-decimal comparison snapshot collected on September 30, 2026
. They show the compared models below 500M reported total parameters, plus Jev 1.13 as a reference. Jev's parameter count is not reported here. All listed models have coverage of 137 benchmarks in the full evaluation and six in the synthetic view.
Task Avg
is the equal-weight mean of Noul, Choice, and Score. Each task first averages its benchmark scores after baseline adjustment and clipping to 0–100. Higher is better; these values are
not raw accuracy
. Noul uses adjusted balanced accuracy, Choice compares selected-target mass against trivial answer policies, and Score compares normalized expected-value MAE against a constant prediction.
These tables are sorted by Task Avg. The leaderboard defaults to a different measure,
Borda Score
: rank points averaged equally across benchmarks. Borda depends on the comparison-model cohort and gives more weight to tasks with more benchmarks. See the
scoring specification
.
Synthetic instruction-and-context benchmarks
The six benchmarks contain
600 English cases
authored and self-reviewed using GPT-6-Astra: Diverse and Contextual sets of 100 cases for each decision type. Contextual sets contain 20 families of five cases. Choice and Score vary state while retaining the question and criteria; Noul varies state and the authored binary criteria.
Bekko falls far behind Jev 1.13 on these S1MB benchmarks designed to measure generalization, especially for the smaller variants. This is the main limitation behind the
v0
designation. The labels are author-intended, not independently human-validated gold. The benchmarks probe adaptation to supplied instructions and context; they do not establish unseen-task generalization or training-data non-overlap.
Evaluation conditions and interpretation
Bekko was trained on related dataset families. The recorded training and evaluation manifests have
77 identically named subsets
in common; this indicates task-family exposure, not a count of leaked test examples. High scores on the full collection should be read alongside the synthetic-task results.
The displayed tables are a rounded reporting snapshot, not a new evaluation. Model folders can contain results from different recorded dataset revisions. Use the per-benchmark model revision, dataset SHA, adapter settings, and input hashes in the
results repository
for reproducibility. Bekko v0 uses native adaptive input budgeting and truncation, so its input handling should not be assumed identical to full-input adapters. Small score differences are not accompanied by confidence intervals.
Context 7,999; query cap 7,997; candidate cap 3,800 tokens, with adaptive allocation
Training / browser branch limits
Query 4,096; candidate 2,048 tokens
The prefix combines instructions and state and is encoded bidirectionally. At each layer, candidates attend to that shared prefix and to their own tokens. The prefix does not attend to candidates, and candidate branches do not attend to each other. This lets the model reuse prefix K/V across candidates instead of repeatedly encoding the same state.
Candidate representations incorporate the prefix through attention before pooling. This changes the attention pattern compared with a conventional cross-encoder; it is not an exact acceleration of unrestricted cross-attention. State and instructions are encoded together, so changing the instructions requires recomputing the prefix. Reuse is within an inference microbatch, not a persistent state-only cache.
Native query and candidate caps are not independently available at their maxima: the runtime shares the context budget and lends unused capacity between branches. Browser exports use their own recorded limits. Neither a configured token limit nor support for a backend establishes accuracy or latency at that limit.
Attention backend
BekkoSentenceTransformer
defaults to
attn_implementation="auto"
. On a CUDA GPU
with compute capability 8.0 or later, it prefers FlashAttention 2 when a compatible
flash_attn
native library is available. Otherwise it uses PyTorch SDPA. CPU
inference uses SDPA. Select the backend explicitly when loading:
Model = get_class_from_dynamic_module(
"inference_v0.BekkoSentenceTransformer", repo,
)
model = Model(
repo, trust_remote_code=True, device="cuda",
attn_implementation="flash_attention_2", # Or "sdpa" or "auto".
)
print(model[0].attn_implementation) # The backend actually selected.
For reproducibility, pin the same model/code revision in both calls, as in the
quickstart.
model_kwargs={"attn_implementation": "flash_attention_2"}
is also
accepted by this class. Conflicting top-level and nested options raise an error.
Selection
Behavior
auto
(default)
Prefer compatible FA2 on a supported CUDA GPU; otherwise use SDPA
sdpa
Always use SDPA; do not import the external FA2 library
flash_attention_2
Require FA2; fail during model loading if the GPU is unsupported or the library is missing/binary-incompatible
FA2 is optional and is not installed by the minimal dependency command above.
Install a
flash-attn
wheel matching your Python, PyTorch and CUDA versions.
When using the Bekko training repository,
uv sync --locked --extra fa2
installs
its pinned optional wheel. An explicit FA2 request never silently falls back.
Selection happens at model loading; reload with the desired device/backend when
changing devices. Use
BekkoSentenceTransformer
for backend selection, rather
than plain
SentenceTransformer
.
As a sizing guideline,
17M models generally have little speed difference
between SDPA and FA2
, making SDPA a practical choice without an extra dependency.
For
68M and larger models, prefer FA2 for throughput
, especially on long
inputs. This is not a speed guarantee for every checkpoint, GPU or workload;
measure representative inputs after warmup and exclude model loading time.
The optimized SDPA path reuses attention masks and rotary-position tensors within
a forward pass and computes local prefix attention in blocks. FA2 keeps valid
tokens packed through attention and feed-forward layers. Both preserve the native
rendering, adaptive input budgets, task heads and candidate order. BF16 rounding
can change individual probabilities and occasionally the selected candidate;
do not assume bitwise-equivalent predictions when switching backends.
The standalone CLI also accepts
--attn-implementation auto
,
sdpa
, or
flash_attention_2
.
Batch inference and performance
Pass a list of input objects to batch across cases:
batch_inputs = [inputs, inputs] # Replace with your application inputs.
results = model.predict(
batch_inputs,
batch_size=128,
token_budget=64000,
show_progress_bar=True,
)
Results follow the original input order. A single input dict returns a result dict; a list returns a list. Progress is enabled by default, counts completed input objects, and is written to stderr.
Option
Default
Meaning
batch_size
128
Maximum cases rendered per window and decisions tokenized per window
token_budget
64000
Approximate padded query-plus-document work per microbatch
context_length
Backbone positional capacity
Shared query/candidate token budget per decision
query_length
Exported cap; fresh exports use context minus 2
Query token cap, including special tokens
document_length
Fresh exports use
min(3800, context - 3)
Per-candidate token cap, including special tokens
show_progress_bar
True
Display input progress
prefix_layout
Exported checkpoint setting
instruction_state
or
state_instruction
Length bucketing reduces padding within each window. Complete decisions stay together; a decision exceeding
token_budget
runs alone. The budget is a work estimate, not a strict memory cap. Lower it if a batch exceeds available memory.
Length overrides apply only to the current call. Adaptive allocation shares the context equally between query and candidate, then lends unused capacity subject to each branch's cap; candidates receive the odd token.
context_length
cannot exceed the backbone's positional capacity. Queries follow the checkpoint's truncation policy; the v0 recipe uses balanced truncation. Candidates are right-truncated. Increasing a limit does not establish model quality at that length.
The first calls include compilation cost, and new shapes may trigger more compilation. Measure warmed calls separately. CUDA inference uses BF16 autocast; CPU inference uses FP32. The runtime does not require PEFT, datasets, or W&B. SDPA needs no external attention package; FA2 requires the compatible optional library described above.
For a historical RTX 5090 run, the 17M model evaluated 26,269 judgments across 137 S1MB benchmarks in
13.33 seconds summed inside the evaluator timer
. The run used BF16 autocast, FlashAttention 2, batch size 128, token budget 64,000, native adaptive truncation, and no compilation. This includes evaluation work; it is neither GPU-kernel-only time nor single-request latency. Model loading and outer orchestration are outside that timer. The measured model revision was
c3a8277
; this historical run used locally flattened inference code and does not measure the current Hub packaging or browser export.
Training
The models use full-parameter fine-tuning of
Ettin rerankers
, without LoRA, on
bekko-system-one-dataset-v0
. The recorded release contains
153 training subsets, 6,589,190 cases, and 8,421,789 judgments
, converted into the three decision types from NLP and retrieval datasets, including synthetic sources.
The recorded runs use a cap of 300,000 judgments per dataset, uniform sampling, a full capped-pool training budget, and equal instruction-first/state-first layout weights. Each run reports
8,421,789 trained judgments and 16,517 updates
, with seed 42. The training data revision is
c6a49c4
.
Setting
17M
68M
400M
Optimizer
AdamW
AdamW
AdamW
Batch size
512
512
512
Encoder learning rate
1e-4
3e-5
1e-5
Head learning rate
5e-4
2e-4
2e-4
Weight decay
0.01
0.01
0.01
Warmup ratio
10%
10%
10%
Schedule
Cosine
Cosine
Cosine
Recorded training-loop time
2.10 h
5.57 h
25.34 h
Durations are training-loop receipts, not complete process wall times. This recipe was selected from experiments; it is not established as optimal. Training code is available in
bekko-system-one
.
Intended use and limitations
Use Bekko for English routing, binary judgments, and rubric-based assessment where you can define alternatives and evaluate performance on representative application data.
Limited generalization:
strong performance on a familiar task does not imply comparable performance on arbitrary instructions or new domains. Inspect the synthetic-task results above.
Probabilities:
outputs are not guaranteed to be calibrated. Changing candidate definitions or the candidate set can change the distribution; validate thresholds for your application.
Input truncation:
adaptive budgets can remove relevant evidence. Review the effective input limits for your runtime and workload.
Language:
this release targets English. Multilingual results from the Bekko Embedding family do not carry over to these models.
Evaluation scope:
S1MB shares dataset families with training, and its synthetic labels have not received independent human validation. Neither view proves absence of training overlap.
License
The released-weight and bundled inference-code license declaration remains to be finalized. This card does not assign a license. Training and evaluation sources retain their own licenses and usage terms; those terms are not replaced by a model license. Consult the
training dataset source records
and each upstream source.
Bekko
(/ˈbek.koː/) is a coined name inspired by two Japanese traditions:
Akabeko (赤べこ):
the red ox cherished as a protective charm against illness and misfortune.
Bekko-iro (鼈甲色):
a traditional Japanese color with a warm, translucent, amber-like hue.
The name brings together the red ox's protective spirit and the beauty of that amber color.
🔎 Also check out Bekko Embedding
I also build
bekko-embedding
, a family of ultra-small, high-performance multilingual embedding models for semantic search. If you are interested in compact models for multilingual retrieval, take a look at the collection for models, benchmarks, and usage examples.
Runs of hotchpotch bekko-system-one-v0-68m on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About bekko-system-one-v0-68m huggingface.co Model
bekko-system-one-v0-68m huggingface.co
bekko-system-one-v0-68m huggingface.co is an AI model on huggingface.co that provides bekko-system-one-v0-68m's model effect (), which can be used instantly with this hotchpotch bekko-system-one-v0-68m model. huggingface.co supports a free trial of the bekko-system-one-v0-68m model, and also provides paid use of the bekko-system-one-v0-68m. Support call bekko-system-one-v0-68m model through api, including Node.js, Python, http.
bekko-system-one-v0-68m huggingface.co is an online trial and call api platform, which integrates bekko-system-one-v0-68m's modeling effects, including api services, and provides a free online trial of bekko-system-one-v0-68m, you can try bekko-system-one-v0-68m online for free by clicking the link below.
hotchpotch bekko-system-one-v0-68m online free url in huggingface.co:
bekko-system-one-v0-68m is an open source model from GitHub that offers a free installation service, and any user can find bekko-system-one-v0-68m on GitHub to install. At the same time, huggingface.co provides the effect of bekko-system-one-v0-68m install, users can directly use bekko-system-one-v0-68m installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
bekko-system-one-v0-68m install url in huggingface.co: