ElJefe is a 22M-parameter MiniLM encoder with two heads that predicts, for a
given prompt, whether a
local Gemma-4-E4B
answer is sufficient or whether
the request is worth escalating to a
frontier model
. ElJefe does not
answer the user's question — it estimates the marginal value of additional
compute.
Head 1 (classification):
p_local
— probability the local model
passes at acceptable quality.
Head 2 (regression):
delta_q
— expected quality gain from the
frontier model.
import onnxruntime as ort
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download
tok = AutoTokenizer.from_pretrained("DJLougen/eljefe-v0", subfolder="tokenizer")
sess = ort.InferenceSession(
hf_hub_download("DJLougen/eljefe-v0", "onnx/router_fp32.onnx"))
enc = tok(["How many r's are in strawberry?"], return_tensors="np",
truncation=True, max_length=256)
p_local, delta_q = sess.run(None, dict(enc)) # outputs: p_local, delta_q
route = "local"if p_local[0] >= 0.9else"frontier"
PyTorch weights are in
model.pt
(state dict of the two-head module);
router.pkl
is the pickled
eljefe.router.Router
wrapper for the
training repo
.
Choosing a threshold
Threshold
Behavior
Use when
0.9
Conservative — only routes local when very confident
Quality > cost
0.6
Balanced — 62% local, ~99.6% retention
Default cost-saving
0.3
Aggressive — most traffic local
Cost-critical
The
delta_q
head is a second signal: escalate when
delta_q > epsilon
even if
p_local
is middling — useful when a bad local answer is costly.
Training data
Trained on 9,553 counterfactual pairs from
DJLougen/eljefe-router-data
:
each prompt was answered by
google/gemma-4-E4B-it
(local, Colab L4) and
deepseek-v4-flash-0731
(frontier, Fireworks), then graded
deterministically (exact-match / numeric / code-tests / constraint checks —
no LLM judge).
What "oracle" means:
the oracle is a cheat-mode router computed after
the fact. Since both models ran on every prompt, we know each row's actual
local_score
and
frontier_score
; the oracle picks the better route per
row (local whenever
local_score >= frontier_score
). It's not deployable —
it "knows" the answer before choosing — but it marks the ceiling: the best
possible cost-quality tradeoff any router could achieve on this data.
At threshold 0.6 ElJefe keeps
62% of traffic local at 99.6% quality
retention
. ROC-AUC 0.861 vs 0.609 for the heuristic.
Frontier-swap transfer
Evaluated against counterfactuals from two
other
frontier models
(glm-5p3-flash, gpt-oss-120b) it was never trained on, ElJefe retains ~95%
of oracle utility — the learned boundary generalizes across escalation
targets.
Limitations
Trained only on objectively-gradable tasks (math, code, MCQ, instruction
following). Open-ended/chat quality is unvalidated.
Binary graders make
local_sufficient
frontier-independent; the delta_q
head carries the frontier-specific signal.
Scores reflect canonical bf16 E4B on an L4 — a quantized local build may
shift the boundary (see
scripts/16_mac_calibration.py
in the repo).
Reproduce
Full pipeline (fetch → dual generation → deterministic grading → splits →
train → calibrate → evaluate → ONNX):
https://github.com/DJLougen/ElJefe
— see README and
PLAN_ElJefe_Colab_CLI.md
.
Runs of DJLougen eljefe-v0 on huggingface.co
14
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
14
30-day runs
More Information About eljefe-v0 huggingface.co Model
eljefe-v0 huggingface.co is an AI model on huggingface.co that provides eljefe-v0's model effect (), which can be used instantly with this DJLougen eljefe-v0 model. huggingface.co supports a free trial of the eljefe-v0 model, and also provides paid use of the eljefe-v0. Support call eljefe-v0 model through api, including Node.js, Python, http.
eljefe-v0 huggingface.co is an online trial and call api platform, which integrates eljefe-v0's modeling effects, including api services, and provides a free online trial of eljefe-v0, you can try eljefe-v0 online for free by clicking the link below.
DJLougen eljefe-v0 online free url in huggingface.co:
eljefe-v0 is an open source model from GitHub that offers a free installation service, and any user can find eljefe-v0 on GitHub to install. At the same time, huggingface.co provides the effect of eljefe-v0 install, users can directly use eljefe-v0 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.