AlexWortega / llm-cipher-reasoning-loras

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: September 03 2026
text-generation

Introduction of llm-cipher-reasoning-loras

Model Details of llm-cipher-reasoning-loras

llm-cipher-reasoning — 20 GRPO LoRA adapters (Qwen3-4B-Instruct-2507)

Every LoRA adapter from a research run asking whether an LLM can be trained — not just prompted — to reason in a more compact form than English. 20 adapters, one per experiment, each in its own subfolder. Base model for 19 of them: Qwen/Qwen3-4B-Instruct-2507 . The exception is ckpt_roundE_supertoken , whose base is a vocabulary-extended variant — see its own card.

Data, eval results, reward logs and the full research ledger: AlexWortega/llm-cipher-reasoning . Code: github.com/AlexWortega/llm-cipher-reasoning .

⚠️ Read this before using any of them

None of these adapters is the recommended configuration. The run's champion is a prompt on the untouched base model — a Chain-of-Draft system prompt ( arxiv 2502.18600 ) that gets 93.5% on GSM8K at 76.8 reasoning tokens vs the plain baseline's 94.0% @ 210.4, i.e. −63% tokens for −0.5pp accuracy , at zero training cost.

Every terseness-trained adapter here fails the out-of-domain gate. Base AIME 2026 accuracy is 36.67%; the best trained adapter manages 26.7% and the best in-domain one drops to 13.3%. Mitigations (correctness gating, difficulty-awareness, KL anchoring, hard-data mixing) move the OOD number non-monotonically , so the tax isn't one mechanism you can switch off. At 4B, GRPO terseness training buys in-domain compression by spending out-of-domain reasoning.

These weights are published as research artifacts — to reproduce the negative result, to probe what terseness training does to a model, or for in-domain-only serving where you have measured that every request looks like your training distribution. Not for general use.

Adapter ledger

Evaluated on GSM8K (unseen seed-2 split) and AIME 2026 (30 problems, MathArena). tok = mean reasoning-span tokens. OOD gate = AIME ≥ 34.67%.

subfolder gen reward mode steps GSM8K tok AIME 2026 verdict
— (no adapter: CoD prompt on base) — — 0 93.5% (n=200) 76.8 ✅ safe CHAMPION
— (no adapter: plain base) — — 0 94.0% (n=200) 210.4 36.67% reference
ckpt_g4_1_klanchor 4 diffaware + KL β=0.15 260 93.0% (n=200) 69.7 13.3% best in-domain; ❌ worst OOD of family
ckpt_g5_1_aimemix 5 eff3, AIME2024-mix ×8 (1240 ex) 260 92.5% (n=200) 69.7 10.0% ❌ hard-data mix refuted — unsolvable rows give no gradient
ckpt_g5_2_mathmix 5 diffaware, MATH L3-5 mix (1500 ex) 260 92.0% (n=200) 70.0 20.0% ❌ occasionally-solvable mix also fails
ckpt_g2_1_cod70 2 eff3, CoD prompt, ramp target 70 260 91.5% (n=200) 69.1 16.7% ❌
ckpt_g3_1_diffaware 3 diffaware (pressure off on hard) 260 90.5% (n=200) 70.4 26.7% ❌ but best OOD among trained
ckpt_g2_2_codcos 2 cosine length + wrong-answer swap 260 88.6% (n=70) 59.0 not gated mid-frontier
ckpt_roundE_supertoken ⚠️ E eff3 + extended vocab (supertokens) 300 88.57% (n=70) 51.0 — pipeline works; gains from the extra pass, not the mechanism. Needs the vocab-extended base, not published
ckpt_phase_eff3 B eff3, target 40 260 85.7% (n=70) 58.3 20.0% rounds A–E champion, dethroned on OOD
ckpt_phase_eff_mtword D multi-token-word avoidance 300 84.29% (n=70) 62.5 23.33% null in-domain, real negative transfer
ckpt_g0_6_hapo 0 hapo (250 ex) 260 82.9% (n=70) 66.6 — ❌ Goodhart: learned to drop </reasoning>
ckpt_phase_eff_target20 E-exp0 eff3, target 20 150 80.00% (n=70) 50.1 — more pressure ≠ better
ckpt_g1_scale_eff3 1 eff3 @ 1000 steps 1000 77.1% (n=70) 45.6 — ❌ step-scaling refuted (three-phase collapse)
ckpt_g0_4_az 0 Arora–Zanette relative α=0.2 260 68.6% (n=70) 46.1 — ❌ no floor → overcompression
ckpt_g0_5_drgrpo 0 eff3 + Dr.GRPO loss, mask_trunc 260 82.9% (n=70) 56.5 — null: TRL objective not length-biased here (512 cap rarely hit)
ckpt_g1_1_cod 1 eff3 + CoD prompt, target 40 260 81.4% (n=70) 50.2 — target 40 too aggressive — killed CoD's accuracy edge (92.9 → 81.4)
ckpt_phase_eff B eff, target 40 (pre-LR-fix) 260 84.29% (n=70) 62.3 — the LoRA learning-rate bug, preserved
ckpt_phase1_scaled A cipher adherence, 7473 ex 600 — — 36.67% cipher line; no OOD change either way
ckpt_phase1 A cipher adherence 260 84.29% — — adherence 6.4% vs 3.3% noise floor
ckpt_phase0_r32 A warmup (correctness+format), r=32 50 — — — curriculum warmup
ckpt_phase0 A warmup (correctness+format), r=16 40 — — — curriculum warmup

Reward modes: eff3 = token-efficiency with correctness gate; diffaware = difficulty-aware (pressure removed on problems the model gets wrong); az = Arora–Zanette group-relative; hapo = HAPO-lite history-aware; cosine = cosine length reward with wrong-answer swap (open-r1). Adapters are r=16 ( ckpt_phase0 , ckpt_phase1 , ckpt_phase_eff , ckpt_phase_eff3 ) or r=32 (all others), α=32, dropout 0, on all attention + MLP projections.

Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen3-4B-Instruct-2507", torch_dtype="bfloat16", device_map="auto"
)
model = PeftModel.from_pretrained(
    base, "AlexWortega/llm-cipher-reasoning-loras", subfolder="ckpt_g3_1_diffaware"
)
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B-Instruct-2507")

All adapters were trained against a <reasoning>...</reasoning><answer>...</answer> format scaffold and expect the matching system prompt — see src/eval_pareto.py in the GitHub repo for the exact prompt text of each mode.

What you probably want instead
COD_PROMPT = (
    "Think step by step, but only keep a minimum draft for each thinking step, "
    "with 5 words at most."
    # + the <reasoning>/<answer> format scaffold — full text in src/eval_pareto.py
)

on the plain base model. Note it is task-scoped : on terminal-bench-2 the same prompt compresses per-call output by −78% but makes the agent loop 6× more, for +268% total tokens . Never inject it into a multi-turn agent system prompt.

Contents

Only final adapters are published ( adapter_model.safetensors + config + tokenizer per subfolder, ~4.5 GB total). Intermediate checkpoint-NNN snapshots were not uploaded.

adapter_config.json files record base_model_name_or_path as it was at training time; pass the base model explicitly to PeftModel.from_pretrained as shown above rather than relying on it.

Citation
@misc{llm_cipher_reasoning_loras_2026,
  title  = {Can LLMs reason in a compact language? Prompting beats GRPO, and terseness training taxes OOD reasoning},
  author = {Alex Wortega},
  year   = {2026},
  url    = {https://huggingface.co/AlexWortega/llm-cipher-reasoning-loras}
}

Runs of AlexWortega llm-cipher-reasoning-loras on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About llm-cipher-reasoning-loras huggingface.co Model

More llm-cipher-reasoning-loras license Visit here:

https://choosealicense.com/licenses/apache-2.0

llm-cipher-reasoning-loras huggingface.co

llm-cipher-reasoning-loras huggingface.co is an AI model on huggingface.co that provides llm-cipher-reasoning-loras's model effect (), which can be used instantly with this AlexWortega llm-cipher-reasoning-loras model. huggingface.co supports a free trial of the llm-cipher-reasoning-loras model, and also provides paid use of the llm-cipher-reasoning-loras. Support call llm-cipher-reasoning-loras model through api, including Node.js, Python, http.

llm-cipher-reasoning-loras huggingface.co Url

https://huggingface.co/AlexWortega/llm-cipher-reasoning-loras

AlexWortega llm-cipher-reasoning-loras online free

llm-cipher-reasoning-loras huggingface.co is an online trial and call api platform, which integrates llm-cipher-reasoning-loras's modeling effects, including api services, and provides a free online trial of llm-cipher-reasoning-loras, you can try llm-cipher-reasoning-loras online for free by clicking the link below.

AlexWortega llm-cipher-reasoning-loras online free url in huggingface.co:

https://huggingface.co/AlexWortega/llm-cipher-reasoning-loras

llm-cipher-reasoning-loras install

llm-cipher-reasoning-loras is an open source model from GitHub that offers a free installation service, and any user can find llm-cipher-reasoning-loras on GitHub to install. At the same time, huggingface.co provides the effect of llm-cipher-reasoning-loras install, users can directly use llm-cipher-reasoning-loras installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

llm-cipher-reasoning-loras install url in huggingface.co:

https://huggingface.co/AlexWortega/llm-cipher-reasoning-loras

Url of llm-cipher-reasoning-loras

llm-cipher-reasoning-loras huggingface.co Url

Provider of llm-cipher-reasoning-loras huggingface.co

AlexWortega
ORGANIZATIONS

Other API from AlexWortega

huggingface.co

Total runs: 303
Run Growth: 12
Growth Rate: 3.96%
Updated:September 05 2026
huggingface.co

Total runs: 291
Run Growth: 278
Growth Rate: 95.53%
Updated:May 02 2024
huggingface.co

Total runs: 66
Run Growth: 0
Growth Rate: 0.00%
Updated:November 16 2024
huggingface.co

Total runs: 34
Run Growth: 13
Growth Rate: 38.24%
Updated:August 03 2023
huggingface.co

Total runs: 24
Run Growth: 0
Growth Rate: 0.00%
Updated:November 16 2024
huggingface.co

Total runs: 24
Run Growth: 17
Growth Rate: 65.38%
Updated:January 20 2024
huggingface.co

Total runs: 14
Run Growth: 1
Growth Rate: 7.14%
Updated:February 02 2023
huggingface.co

Total runs: 14
Run Growth: 5
Growth Rate: 41.67%
Updated:January 14 2024
huggingface.co

Total runs: 12
Run Growth: 3
Growth Rate: 25.00%
Updated:January 21 2024
huggingface.co

Total runs: 12
Run Growth: 2
Growth Rate: 16.67%
Updated:February 02 2023
huggingface.co

Total runs: 11
Run Growth: 7
Growth Rate: 58.33%
Updated:September 22 2023
huggingface.co

Total runs: 9
Run Growth: -4
Growth Rate: -44.44%
Updated:November 05 2022
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:September 22 2023
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:February 01 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 08 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:March 28 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 03 2023
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:October 01 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:June 18 2024