DJLougen / eljefe-v0

huggingface.co
Total runs: 14
24-hour runs: 0
7-day runs: 0
30-day runs: 14
Model's Last Updated: September 18 2026

Introduction of eljefe-v0

Model Details of eljefe-v0

ElJefe v0 — Local-vs-Frontier Router

ElJefe is a 22M-parameter MiniLM encoder with two heads that predicts, for a given prompt, whether a local Gemma-4-E4B answer is sufficient or whether the request is worth escalating to a frontier model . ElJefe does not answer the user's question — it estimates the marginal value of additional compute.

  • Head 1 (classification): p_local — probability the local model passes at acceptable quality.
  • Head 2 (regression): delta_q — expected quality gain from the frontier model.
Quickstart
pip install onnxruntime transformers huggingface_hub
import onnxruntime as ort
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download

tok = AutoTokenizer.from_pretrained("DJLougen/eljefe-v0", subfolder="tokenizer")
sess = ort.InferenceSession(
    hf_hub_download("DJLougen/eljefe-v0", "onnx/router_fp32.onnx"))

enc = tok(["How many r's are in strawberry?"], return_tensors="np",
          truncation=True, max_length=256)
p_local, delta_q = sess.run(None, dict(enc))   # outputs: p_local, delta_q
route = "local" if p_local[0] >= 0.9 else "frontier"

PyTorch weights are in model.pt (state dict of the two-head module); router.pkl is the pickled eljefe.router.Router wrapper for the training repo .

Choosing a threshold
Threshold Behavior Use when
0.9 Conservative — only routes local when very confident Quality > cost
0.6 Balanced — 62% local, ~99.6% retention Default cost-saving
0.3 Aggressive — most traffic local Cost-critical

The delta_q head is a second signal: escalate when delta_q > epsilon even if p_local is middling — useful when a bad local answer is costly.

Training data

Trained on 9,553 counterfactual pairs from DJLougen/eljefe-router-data : each prompt was answered by google/gemma-4-E4B-it (local, Colab L4) and deepseek-v4-flash-0731 (frontier, Fireworks), then graded deterministically (exact-match / numeric / code-tests / constraint checks — no LLM judge).

Labels: local_sufficient = local_score >= 0.70 AND delta_q <= 0.10 .

Results (test_iid, n=995)
Router Quality retention Kept local Cost reduction False-local
always-local 83.9% 100% 100% 17.1%
heuristic 89.0% 73.3% 69.1% 12.4%
TF-IDF 99.1% 24.3% 19.7% 1.5%
embedding 99.9% 15.3% 12.6% 0.7%
ElJefe v0 100.6% 31.3% 25.1% 1.2%
oracle (ceiling) 105.4% 66.6% 47.6% 0%

What "oracle" means: the oracle is a cheat-mode router computed after the fact. Since both models ran on every prompt, we know each row's actual local_score and frontier_score ; the oracle picks the better route per row (local whenever local_score >= frontier_score ). It's not deployable — it "knows" the answer before choosing — but it marks the ceiling: the best possible cost-quality tradeoff any router could achieve on this data.

At threshold 0.6 ElJefe keeps 62% of traffic local at 99.6% quality retention . ROC-AUC 0.861 vs 0.609 for the heuristic.

Frontier-swap transfer

Evaluated against counterfactuals from two other frontier models (glm-5p3-flash, gpt-oss-120b) it was never trained on, ElJefe retains ~95% of oracle utility — the learned boundary generalizes across escalation targets.

Limitations
  • Trained only on objectively-gradable tasks (math, code, MCQ, instruction following). Open-ended/chat quality is unvalidated.
  • Binary graders make local_sufficient frontier-independent; the delta_q head carries the frontier-specific signal.
  • Scores reflect canonical bf16 E4B on an L4 — a quantized local build may shift the boundary (see scripts/16_mac_calibration.py in the repo).
Reproduce

Full pipeline (fetch → dual generation → deterministic grading → splits → train → calibrate → evaluate → ONNX): https://github.com/DJLougen/ElJefe — see README and PLAN_ElJefe_Colab_CLI.md .

Runs of DJLougen eljefe-v0 on huggingface.co

14
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
14
30-day runs

More Information About eljefe-v0 huggingface.co Model

More eljefe-v0 license Visit here:

https://choosealicense.com/licenses/apache-2.0

eljefe-v0 huggingface.co

eljefe-v0 huggingface.co is an AI model on huggingface.co that provides eljefe-v0's model effect (), which can be used instantly with this DJLougen eljefe-v0 model. huggingface.co supports a free trial of the eljefe-v0 model, and also provides paid use of the eljefe-v0. Support call eljefe-v0 model through api, including Node.js, Python, http.

DJLougen eljefe-v0 online free

eljefe-v0 huggingface.co is an online trial and call api platform, which integrates eljefe-v0's modeling effects, including api services, and provides a free online trial of eljefe-v0, you can try eljefe-v0 online for free by clicking the link below.

DJLougen eljefe-v0 online free url in huggingface.co:

https://huggingface.co/DJLougen/eljefe-v0

eljefe-v0 install

eljefe-v0 is an open source model from GitHub that offers a free installation service, and any user can find eljefe-v0 on GitHub to install. At the same time, huggingface.co provides the effect of eljefe-v0 install, users can directly use eljefe-v0 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

eljefe-v0 install url in huggingface.co:

https://huggingface.co/DJLougen/eljefe-v0

Url of eljefe-v0

eljefe-v0 huggingface.co Url

Provider of eljefe-v0 huggingface.co

DJLougen
ORGANIZATIONS

Other API from DJLougen

huggingface.co

Total runs: 76
Run Growth: -874
Growth Rate: -1150.00%
Updated:April 10 2026
huggingface.co

Total runs: 11
Run Growth: -612
Growth Rate: -5563.64%
Updated:April 10 2026
huggingface.co

Total runs: 9
Run Growth: -634
Growth Rate: -7044.44%
Updated:April 10 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:January 13 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:January 13 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 05 2026