hotchpotch / bekko-system-one-v0-68m

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: October 01 2026

Introduction of bekko-system-one-v0-68m

Model Details of bekko-system-one-v0-68m

bekko-system-one

bekko-system-one-v0-68m

bekko-system-one-v0-68m is the 68M variant of Bekko System One v0: a family of compact English encoder models for Choice, Noul, and Score decisions. Supply instructions, application state, and candidate definitions; model.predict() returns candidate probabilities, a selected option, or a numeric rubric score.

Bekko System One is an experimental project exploring whether ultra-small models can become capable System One Decision Models , in the same category as TypeSafe AI’s Jev . The v0 family spans 17M, 68M, and 400M parameters. It builds on Ettin rerankers with a shared-prefix encoder and task-specific heads, evaluating supplied alternatives in parallel without generating text.

Why v0? These models can perform very well on some tasks, but fall far behind Jev 1.13 on S1MB’s benchmarks designed to measure generalization . Their training data also includes dataset families represented in the evaluation, so strong results on those tasks do not establish broad generalization. The v0 designation reflects this limited generalization.

Release article · Model collection · Browser demo · Training and inference code · Training dataset · S1MB leaderboard

✨ Highlights
  • Typed outputs: Choice selects among named alternatives; Noul estimates the probability of an authored binary condition; Score returns the expected value of a numeric rubric.
  • Compact models: the 17M variant has 16.80M total parameters and 3.90M excluding lookup-only embeddings. Its browser ONNX model file is about 29 MB.
  • Shared computation: instructions and state are encoded once per unique tokenized prefix within a microbatch and reused across candidates.
  • Standalone inference: CPU or CUDA execution through the exported BekkoSentenceTransformer , with optional FlashAttention 2 and torch.compile() .
  • Inspectable evaluation: S1MB reports both specialized-task results and a separate view of six synthetic instruction-and-context benchmarks.
Model family and browser inference
Variant Total params Active params (AP) Browser ONNX file
17M 17M 4M 29 MB
68M 68M 42M 196 MB
400M 395M 343M 1.43 GB

Parameter counts are rounded; M means million. AP excludes lookup-only token embeddings; it is not a memory estimate. File sizes use decimal MB/GB and cover only onnx_browser/model.onnx , excluding tokenizer, runtime, and working memory.

The browser demo offers 17M and 68M with CPU/WebGPU execution. Browser exports quantize the token lookup table to row-wise INT8; transformer blocks and task heads remain FP32. These are not fully INT8 models. Demo examples illustrate selected tasks and should not be used as evidence of broad generalization.

🚀 Quickstart

Install the runtime dependencies:

pip install 'torch>=2.10,<2.11' 'transformers==5.17.0' 'sentence-transformers==6.1.0' 'safetensors>=0.7' 'tqdm>=4.67'

Use a CUDA-compatible PyTorch installation for GPU inference. Load the model and its inference class directly from the Hub:

import json
from transformers.dynamic_module_utils import get_class_from_dynamic_module

repo = "hotchpotch/bekko-system-one-v0-68m"
model = get_class_from_dynamic_module(
    "inference_v0.BekkoSentenceTransformer", repo,
)(repo, trust_remote_code=True, device="cpu")  # Or "cuda".

inputs = {
    "state_json": json.dumps({
        "message": "I was charged twice for the same order. Please refund the duplicate payment."
    }),
    "decisions": [{
        "id": "department",
        "kind": "judgment",
        "type": "choice",
        "instructions_json": json.dumps("Which department should handle this request?"),
        "system_prompt": "",
        "criteria": [
            {
                "id": "billing",
                "description_json": json.dumps("Payments, charges, and refunds"),
                "value": None,
            },
            {
                "id": "technical",
                "description_json": json.dumps("Technical failures and configuration"),
                "value": None,
            },
        ],
        "documents": [],
        "scoring": None,
    }],
}

result = model.predict(inputs)
print(result["department"]["selected_id"])
print(result["department"]["probabilities"])

The output contains the selected candidate ID and the distribution over candidate IDs. Predictions depend on the checkpoint and input; no example probabilities are asserted here.

This loads BekkoSentenceTransformer , which exposes model.predict() directly for both single input objects and batches. For reproducible loading, pass the same revision="FULL_COMMIT_HASH" to both calls above. Raw-text encode() is not the typed-decision interface.

Decision types
Type Input candidates Output
Choice Named alternatives with descriptions selected_id , probabilities
Noul Authored yes/no meanings, using IDs true / false or yes / no probability_yes , probabilities
Score A rubric with at least two distinct, explicit numeric values score , normalized_score , probabilities , values

An input object can contain several decisions sharing the same state. Decision IDs must be unique within an input object; they may repeat across input objects in a batch. Pass only the input object, without training targets or dataset metadata.

For Noul, define what both outcomes mean. For example:

inputs["decisions"].append({
    "id": "refund",
    "kind": "judgment",
    "type": "noul",
    "instructions_json": json.dumps("Is the customer requesting a refund?"),
    "system_prompt": "",
    "criteria": [
        {"id": "false", "description_json": json.dumps("No refund is requested."), "value": None},
        {"id": "true", "description_json": json.dumps("The customer asks for money back."), "value": None},
    ],
    "documents": [],
    "scoring": None,
})

result = model.predict(inputs)
print(result["refund"]["probability_yes"])

For Score, each criterion supplies a value alongside its id and description_json . The numeric score is the probability-weighted expectation of these values. A rubric with values 0 , 2 , and 4 produces a score on the 0–4 scale; normalized_score maps that expectation to 0–1 using the minimum and maximum criterion values. Numbers are not inferred from labels or descriptions.

The runtime also supports relative document ranking with kind="ranking" , type=None , scoring="relative" , an empty criteria list, and documents containing id and content_json . It returns probabilities and an order of document IDs. Ranking probabilities are relative to the supplied candidate set; they are not numeric rubric scores. Runtime support alone does not establish ranking quality for this checkpoint.

📊 Evaluation: S1MB

S1MB — System One Mosaic Benchmark combines 137 typed benchmarks : 59 Noul, 57 Choice, and 21 Score. The active evaluation release contains 106 subsets, 14,009 cases, and 26,269 judgments . One case may contain several judgments. It combines existing dataset-derived tasks with synthetic tasks; it is not a collection consisting exclusively of original official test splits.

The tables below reproduce the two-decimal comparison snapshot collected on September 30, 2026 . They show the compared models below 500M reported total parameters, plus Jev 1.13 as a reference. Jev's parameter count is not reported here. All listed models have coverage of 137 benchmarks in the full evaluation and six in the synthetic view.

Full evaluation
Model Task Avg Noul Choice Score Total params AP
Jev 1.13 59.59 64.63 67.22 46.92 — —
bekko-system-one-v0-400m 50.60 51.24 61.32 39.25 395M 343M
bekko-system-one-v0-68m 40.46 42.91 51.62 26.85 68M 42M
bekko-system-one-v0-17m 27.57 31.43 35.57 15.70 17M 4M
von 16.21 20.15 23.99 4.48 395M 343M
laya-typed-decisions 15.00 20.06 18.93 5.99 421M 370M
laya 13.36 20.19 16.31 3.58 421M 370M
laya-multilingual 9.28 14.01 13.03 0.79 322M 125M

Task Avg is the equal-weight mean of Noul, Choice, and Score. Each task first averages its benchmark scores after baseline adjustment and clipping to 0–100. Higher is better; these values are not raw accuracy . Noul uses adjusted balanced accuracy, Choice compares selected-target mass against trivial answer policies, and Score compares normalized expected-value MAE against a constant prediction.

These tables are sorted by Task Avg. The leaderboard defaults to a different measure, Borda Score : rank points averaged equally across benchmarks. Borda depends on the comparison-model cohort and gives more weight to tasks with more benchmarks. See the scoring specification .

Synthetic instruction-and-context benchmarks

The six benchmarks contain 600 English cases authored and self-reviewed using GPT-6-Astra: Diverse and Contextual sets of 100 cases for each decision type. Contextual sets contain 20 families of five cases. Choice and Score vary state while retaining the question and criteria; Noul varies state and the authored binary criteria.

Model Task Avg General Noul General Choice General Score
Jev 1.13 96.27 99.00 98.68 91.14
bekko-system-one-v0-400m 54.48 52.00 80.30 31.14
von 43.19 41.00 65.89 22.67
bekko-system-one-v0-68m 32.48 31.00 57.52 8.90
laya-typed-decisions 31.82 22.00 58.80 14.67
laya 26.50 9.00 54.83 15.67
bekko-system-one-v0-17m 18.96 15.00 39.11 2.76
laya-multilingual 14.47 6.00 36.67 0.75

Bekko falls far behind Jev 1.13 on these S1MB benchmarks designed to measure generalization, especially for the smaller variants. This is the main limitation behind the v0 designation. The labels are author-intended, not independently human-validated gold. The benchmarks probe adaptation to supplied instructions and context; they do not establish unseen-task generalization or training-data non-overlap.

Evaluation conditions and interpretation

Bekko was trained on related dataset families. The recorded training and evaluation manifests have 77 identically named subsets in common; this indicates task-family exposure, not a count of leaked test examples. High scores on the full collection should be read alongside the synthetic-task results.

The displayed tables are a rounded reporting snapshot, not a new evaluation. Model folders can contain results from different recorded dataset revisions. Use the per-benchmark model revision, dataset SHA, adapter settings, and input hashes in the results repository for reproducibility. Bekko v0 uses native adaptive input budgeting and truncation, so its input handling should not be assumed identical to full-input adapters. Small score differences are not accompanied by confidence intervals.

Evaluation dataset · Leaderboard · Run an evaluation · Submit results

Architecture and input limits
Item Value
Model hotchpotch/bekko-system-one-v0-68m
Base model cross-encoder/ettin-reranker-68m-v1
Architecture ModernBERT-compatible shared-prefix encoder
Decision heads Choice, Noul, Score
Candidate representation Mean pooling over candidate branch tokens
Probabilities Softmax over candidates within each decision
Inference API BekkoSentenceTransformer.predict()
Native inference limits Context 7,999; query cap 7,997; candidate cap 3,800 tokens, with adaptive allocation
Training / browser branch limits Query 4,096; candidate 2,048 tokens

The prefix combines instructions and state and is encoded bidirectionally. At each layer, candidates attend to that shared prefix and to their own tokens. The prefix does not attend to candidates, and candidate branches do not attend to each other. This lets the model reuse prefix K/V across candidates instead of repeatedly encoding the same state.

Candidate representations incorporate the prefix through attention before pooling. This changes the attention pattern compared with a conventional cross-encoder; it is not an exact acceleration of unrestricted cross-attention. State and instructions are encoded together, so changing the instructions requires recomputing the prefix. Reuse is within an inference microbatch, not a persistent state-only cache.

Native query and candidate caps are not independently available at their maxima: the runtime shares the context budget and lends unused capacity between branches. Browser exports use their own recorded limits. Neither a configured token limit nor support for a backend establishes accuracy or latency at that limit.

Attention backend

BekkoSentenceTransformer defaults to attn_implementation="auto" . On a CUDA GPU with compute capability 8.0 or later, it prefers FlashAttention 2 when a compatible flash_attn native library is available. Otherwise it uses PyTorch SDPA. CPU inference uses SDPA. Select the backend explicitly when loading:

Model = get_class_from_dynamic_module(
    "inference_v0.BekkoSentenceTransformer", repo,
)
model = Model(
    repo, trust_remote_code=True, device="cuda",
    attn_implementation="flash_attention_2",  # Or "sdpa" or "auto".
)
print(model[0].attn_implementation)  # The backend actually selected.

For reproducibility, pin the same model/code revision in both calls, as in the quickstart. model_kwargs={"attn_implementation": "flash_attention_2"} is also accepted by this class. Conflicting top-level and nested options raise an error.

Selection Behavior
auto (default) Prefer compatible FA2 on a supported CUDA GPU; otherwise use SDPA
sdpa Always use SDPA; do not import the external FA2 library
flash_attention_2 Require FA2; fail during model loading if the GPU is unsupported or the library is missing/binary-incompatible

FA2 is optional and is not installed by the minimal dependency command above. Install a flash-attn wheel matching your Python, PyTorch and CUDA versions. When using the Bekko training repository, uv sync --locked --extra fa2 installs its pinned optional wheel. An explicit FA2 request never silently falls back. Selection happens at model loading; reload with the desired device/backend when changing devices. Use BekkoSentenceTransformer for backend selection, rather than plain SentenceTransformer .

As a sizing guideline, 17M models generally have little speed difference between SDPA and FA2 , making SDPA a practical choice without an extra dependency. For 68M and larger models, prefer FA2 for throughput , especially on long inputs. This is not a speed guarantee for every checkpoint, GPU or workload; measure representative inputs after warmup and exclude model loading time.

The optimized SDPA path reuses attention masks and rotary-position tensors within a forward pass and computes local prefix attention in blocks. FA2 keeps valid tokens packed through attention and feed-forward layers. Both preserve the native rendering, adaptive input budgets, task heads and candidate order. BF16 rounding can change individual probabilities and occasionally the selected candidate; do not assume bitwise-equivalent predictions when switching backends.

The standalone CLI also accepts --attn-implementation auto , sdpa , or flash_attention_2 .

Batch inference and performance

Pass a list of input objects to batch across cases:

batch_inputs = [inputs, inputs]  # Replace with your application inputs.
results = model.predict(
    batch_inputs,
    batch_size=128,
    token_budget=64000,
    show_progress_bar=True,
)

Results follow the original input order. A single input dict returns a result dict; a list returns a list. Progress is enabled by default, counts completed input objects, and is written to stderr.

Option Default Meaning
batch_size 128 Maximum cases rendered per window and decisions tokenized per window
token_budget 64000 Approximate padded query-plus-document work per microbatch
context_length Backbone positional capacity Shared query/candidate token budget per decision
query_length Exported cap; fresh exports use context minus 2 Query token cap, including special tokens
document_length Fresh exports use min(3800, context - 3) Per-candidate token cap, including special tokens
show_progress_bar True Display input progress
prefix_layout Exported checkpoint setting instruction_state or state_instruction

Length bucketing reduces padding within each window. Complete decisions stay together; a decision exceeding token_budget runs alone. The budget is a work estimate, not a strict memory cap. Lower it if a batch exceeds available memory.

Length overrides apply only to the current call. Adaptive allocation shares the context equally between query and candidate, then lends unused capacity subject to each branch's cap; candidates receive the odd token. context_length cannot exceed the backbone's positional capacity. Queries follow the checkpoint's truncation policy; the v0 recipe uses balanced truncation. Candidates are right-truncated. Increasing a limit does not establish model quality at that length.

To enable optional compilation:

model.compile_inference()
results = model.predict(batch_inputs)
# model.disable_compile() restores eager execution.

The first calls include compilation cost, and new shapes may trigger more compilation. Measure warmed calls separately. CUDA inference uses BF16 autocast; CPU inference uses FP32. The runtime does not require PEFT, datasets, or W&B. SDPA needs no external attention package; FA2 requires the compatible optional library described above.

For a historical RTX 5090 run, the 17M model evaluated 26,269 judgments across 137 S1MB benchmarks in 13.33 seconds summed inside the evaluator timer . The run used BF16 autocast, FlashAttention 2, batch size 128, token budget 64,000, native adaptive truncation, and no compilation. This includes evaluation work; it is neither GPU-kernel-only time nor single-request latency. Model loading and outer orchestration are outside that timer. The measured model revision was c3a8277 ; this historical run used locally flattened inference code and does not measure the current Hub packaging or browser export.

Training

The models use full-parameter fine-tuning of Ettin rerankers , without LoRA, on bekko-system-one-dataset-v0 . The recorded release contains 153 training subsets, 6,589,190 cases, and 8,421,789 judgments , converted into the three decision types from NLP and retrieval datasets, including synthetic sources.

The recorded runs use a cap of 300,000 judgments per dataset, uniform sampling, a full capped-pool training budget, and equal instruction-first/state-first layout weights. Each run reports 8,421,789 trained judgments and 16,517 updates , with seed 42. The training data revision is c6a49c4 .

Setting 17M 68M 400M
Optimizer AdamW AdamW AdamW
Batch size 512 512 512
Encoder learning rate 1e-4 3e-5 1e-5
Head learning rate 5e-4 2e-4 2e-4
Weight decay 0.01 0.01 0.01
Warmup ratio 10% 10% 10%
Schedule Cosine Cosine Cosine
Recorded training-loop time 2.10 h 5.57 h 25.34 h

Durations are training-loop receipts, not complete process wall times. This recipe was selected from experiments; it is not established as optimal. Training code is available in bekko-system-one .

Intended use and limitations

Use Bekko for English routing, binary judgments, and rubric-based assessment where you can define alternatives and evaluate performance on representative application data.

  • Limited generalization: strong performance on a familiar task does not imply comparable performance on arbitrary instructions or new domains. Inspect the synthetic-task results above.
  • Probabilities: outputs are not guaranteed to be calibrated. Changing candidate definitions or the candidate set can change the distribution; validate thresholds for your application.
  • Input truncation: adaptive budgets can remove relevant evidence. Review the effective input limits for your runtime and workload.
  • Language: this release targets English. Multilingual results from the Bekko Embedding family do not carry over to these models.
  • Evaluation scope: S1MB shares dataset families with training, and its synthetic labels have not received independent human validation. Neither view proves absence of training overlap.
License

The released-weight and bundled inference-code license declaration remains to be finalized. This card does not assign a license. Training and evaluation sources retain their own licenses and usage terms; those terms are not replaced by a model license. Consult the training dataset source records and each upstream source.

Author

Yuichi Tateno — @hotchpotch

What is Bekko?

Bekko (/ˈbek.koː/) is a coined name inspired by two Japanese traditions:

  • Akabeko (赤べこ): the red ox cherished as a protective charm against illness and misfortune.
  • Bekko-iro (鼈甲色): a traditional Japanese color with a warm, translucent, amber-like hue.

The name brings together the red ox's protective spirit and the beauty of that amber color.

🔎 Also check out Bekko Embedding

I also build bekko-embedding , a family of ultra-small, high-performance multilingual embedding models for semantic search. If you are interested in compact models for multilingual retrieval, take a look at the collection for models, benchmarks, and usage examples.

Runs of hotchpotch bekko-system-one-v0-68m on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About bekko-system-one-v0-68m huggingface.co Model

bekko-system-one-v0-68m huggingface.co

bekko-system-one-v0-68m huggingface.co is an AI model on huggingface.co that provides bekko-system-one-v0-68m's model effect (), which can be used instantly with this hotchpotch bekko-system-one-v0-68m model. huggingface.co supports a free trial of the bekko-system-one-v0-68m model, and also provides paid use of the bekko-system-one-v0-68m. Support call bekko-system-one-v0-68m model through api, including Node.js, Python, http.

bekko-system-one-v0-68m huggingface.co Url

https://huggingface.co/hotchpotch/bekko-system-one-v0-68m

hotchpotch bekko-system-one-v0-68m online free

bekko-system-one-v0-68m huggingface.co is an online trial and call api platform, which integrates bekko-system-one-v0-68m's modeling effects, including api services, and provides a free online trial of bekko-system-one-v0-68m, you can try bekko-system-one-v0-68m online for free by clicking the link below.

hotchpotch bekko-system-one-v0-68m online free url in huggingface.co:

https://huggingface.co/hotchpotch/bekko-system-one-v0-68m

bekko-system-one-v0-68m install

bekko-system-one-v0-68m is an open source model from GitHub that offers a free installation service, and any user can find bekko-system-one-v0-68m on GitHub to install. At the same time, huggingface.co provides the effect of bekko-system-one-v0-68m install, users can directly use bekko-system-one-v0-68m installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

bekko-system-one-v0-68m install url in huggingface.co:

https://huggingface.co/hotchpotch/bekko-system-one-v0-68m

Url of bekko-system-one-v0-68m

bekko-system-one-v0-68m huggingface.co Url

Provider of bekko-system-one-v0-68m huggingface.co

hotchpotch
ORGANIZATIONS

Other API from hotchpotch