DedeProGames / Kiyo-Diffusion-135M

huggingface.co
Total runs: 66
24-hour runs: 0
7-day runs: 0
30-day runs: 62
Model's Last Updated: September 12 2026
text-generation

Introduction of Kiyo-Diffusion-135M

Model Details of Kiyo-Diffusion-135M

Kiyo-Diffusion-135M

Masked diffusion language model adapted from DedeProGames/Kiyo-135M : same Llama weights, bidirectional attention, a new <|mask|> token, the masked-diffusion (MDLM/LLaDA) objective and the AR shift operation (DiffuLLaMA/Dream recipe). No timestep embedding. Research use only.

python kiyo_diffusion.py generate --model DedeProGames/Kiyo-Diffusion-135M --prompt "The capital of France is"
Diffusion adaptation

The autoregressive Kiyo-135M was converted in place — causal attention replaced by bidirectional attention, a <|mask|> token added, and the objective switched to masked denoising — then trained for 250M tokens on the same data mix the base model was pretrained on (fineweb-edu 45% / dclm-baseline 35% / finemath 10% / stack-v3 10%). Keeping the data distribution identical is what protects the base knowledge: this is continued pretraining under a new objective, not a domain shift.

Setting Value
Tokens 250M (1908 steps × 131,072 tokens)
Learning rate 1e-4 → 1e-5, cosine, 200 warmup
Attention annealing causal → bidirectional over the first 381 steps (20%)
Batch 131,072 tokens global, sequence length 1024
Precision bf16
Hardware 2× RTX 3060 12GB (DDP), ~15k tokens/s
Benchmark: BananaMind Base Bench 1.1

All 31 training checkpoints were evaluated on BananaMind/BananaMind-Base-Bench-1.1 — 350 four-way text-completion items across 7 categories, reporting a fixed-scale Overall Elo.

Overall Elo per checkpoint

Checkpoint Overall Elo Accuracy
Kiyo-135M, before conversion (official causal metric) 1125.9 67.7%
After conversion, 0 training steps 968.7 50.0%
Step 500 (65M tokens) 1054.1 60.3%
Final, step 1908 (250M tokens) 1055.4 59.1%

What the curve shows. Opening the attention costs 157 Elo immediately. The first 500 steps (65M tokens) recover 85 of those points — 54% of the gap — and essentially all of the recovery happens there. Across the remaining 1408 steps the Overall Elo stays flat at 1045 ± 10 while the training NELBO keeps falling (4.47 → 4.06): the model keeps getting better at denoising without regaining measurable benchmark capability.

For a 135M-parameter model, diffusion adaptation appears to saturate near 65M tokens, and the residual ~70 Elo gap looks structural to bidirectional attention at this scale rather than a debt that more tokens would repay. Per category, world_knowledge and context_tracking recover closest to the original — bidirectional context helps them — while code_completion suffers most (0.86 → 0.50), which is unsurprising for a strongly sequential task.

Scoring methodology

A masked diffusion LM has no causal next-token likelihood, so the official metric of this benchmark (mean conditional token log-probability) cannot be reproduced. Checkpoints are scored instead with a Monte-Carlo ELBO ( n_mc=64 ), the diffusion analogue used by LLaDA.

To make checkpoints comparable to one another, mask samples are drawn from a seed derived from the item index, so every checkpoint is evaluated against identical masks. Without this, replicate runs of the same checkpoint differed by as much as 16 Elo; with it, replicates are identical.

The dashed baseline in the chart is therefore not an ELBO-equivalent score. It marks where the model started, measured the only way an autoregressive model can be measured. Elo values here are comparable within this chart only, and not with the official leaderboard.

Runs of DedeProGames Kiyo-Diffusion-135M on huggingface.co

66
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
62
30-day runs

More Information About Kiyo-Diffusion-135M huggingface.co Model

More Kiyo-Diffusion-135M license Visit here:

https://choosealicense.com/licenses/apache-2.0

Kiyo-Diffusion-135M huggingface.co

Kiyo-Diffusion-135M huggingface.co is an AI model on huggingface.co that provides Kiyo-Diffusion-135M's model effect (), which can be used instantly with this DedeProGames Kiyo-Diffusion-135M model. huggingface.co supports a free trial of the Kiyo-Diffusion-135M model, and also provides paid use of the Kiyo-Diffusion-135M. Support call Kiyo-Diffusion-135M model through api, including Node.js, Python, http.

Kiyo-Diffusion-135M huggingface.co Url

https://huggingface.co/DedeProGames/Kiyo-Diffusion-135M

DedeProGames Kiyo-Diffusion-135M online free

Kiyo-Diffusion-135M huggingface.co is an online trial and call api platform, which integrates Kiyo-Diffusion-135M's modeling effects, including api services, and provides a free online trial of Kiyo-Diffusion-135M, you can try Kiyo-Diffusion-135M online for free by clicking the link below.

DedeProGames Kiyo-Diffusion-135M online free url in huggingface.co:

https://huggingface.co/DedeProGames/Kiyo-Diffusion-135M

Kiyo-Diffusion-135M install

Kiyo-Diffusion-135M is an open source model from GitHub that offers a free installation service, and any user can find Kiyo-Diffusion-135M on GitHub to install. At the same time, huggingface.co provides the effect of Kiyo-Diffusion-135M install, users can directly use Kiyo-Diffusion-135M installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Kiyo-Diffusion-135M install url in huggingface.co:

https://huggingface.co/DedeProGames/Kiyo-Diffusion-135M

Url of Kiyo-Diffusion-135M

Kiyo-Diffusion-135M huggingface.co Url

Provider of Kiyo-Diffusion-135M huggingface.co

DedeProGames
ORGANIZATIONS

Other API from DedeProGames

huggingface.co

Total runs: 1.0K
Run Growth: 1.0K
Growth Rate: 100.00%
Updated:September 07 2026
huggingface.co

Total runs: 302
Run Growth: 294
Growth Rate: 97.35%
Updated:September 28 2026
huggingface.co

Total runs: 237
Run Growth: 229
Growth Rate: 96.62%
Updated:September 28 2026