Masked diffusion language model adapted from
DedeProGames/Kiyo-135M
: same Llama
weights, bidirectional attention, a new
<|mask|>
token, the masked-diffusion (MDLM/LLaDA) objective
and the AR shift operation (DiffuLLaMA/Dream recipe). No timestep embedding. Research use only.
python kiyo_diffusion.py generate --model DedeProGames/Kiyo-Diffusion-135M --prompt "The capital of France is"
Diffusion adaptation
The autoregressive Kiyo-135M was converted in place — causal attention replaced by bidirectional
attention, a
<|mask|>
token added, and the objective switched to masked denoising — then trained
for
250M tokens
on the same data mix the base model was pretrained on
(fineweb-edu 45% / dclm-baseline 35% / finemath 10% / stack-v3 10%). Keeping the data distribution
identical is what protects the base knowledge: this is continued pretraining under a new objective,
not a domain shift.
Setting
Value
Tokens
250M (1908 steps × 131,072 tokens)
Learning rate
1e-4 → 1e-5, cosine, 200 warmup
Attention annealing
causal → bidirectional over the first 381 steps (20%)
Batch
131,072 tokens global, sequence length 1024
Precision
bf16
Hardware
2× RTX 3060 12GB (DDP), ~15k tokens/s
Benchmark: BananaMind Base Bench 1.1
All 31 training checkpoints were evaluated on
BananaMind/BananaMind-Base-Bench-1.1
— 350 four-way text-completion items across 7 categories, reporting a fixed-scale Overall Elo.
Checkpoint
Overall Elo
Accuracy
Kiyo-135M, before conversion (official causal metric)
1125.9
67.7%
After conversion, 0 training steps
968.7
50.0%
Step 500 (65M tokens)
1054.1
60.3%
Final, step 1908 (250M tokens)
1055.4
59.1%
What the curve shows.
Opening the attention costs 157 Elo immediately. The first 500 steps
(65M tokens) recover 85 of those points — 54% of the gap — and essentially all of the recovery
happens there. Across the remaining 1408 steps the Overall Elo stays flat at 1045 ± 10 while the
training NELBO keeps falling (4.47 → 4.06): the model keeps getting better at denoising without
regaining measurable benchmark capability.
For a 135M-parameter model, diffusion adaptation appears to saturate near 65M tokens, and the
residual ~70 Elo gap looks structural to bidirectional attention at this scale rather than a debt
that more tokens would repay. Per category,
world_knowledge
and
context_tracking
recover closest
to the original — bidirectional context helps them — while
code_completion
suffers most
(0.86 → 0.50), which is unsurprising for a strongly sequential task.
Scoring methodology
A masked diffusion LM has no causal next-token likelihood, so the official metric of this benchmark
(mean conditional token log-probability) cannot be reproduced. Checkpoints are scored instead with a
Monte-Carlo ELBO (
n_mc=64
), the diffusion analogue used by LLaDA.
To make checkpoints comparable to one another, mask samples are drawn from a seed derived from the
item index, so every checkpoint is evaluated against identical masks. Without this, replicate runs
of the same checkpoint differed by as much as 16 Elo; with it, replicates are identical.
The dashed baseline in the chart is therefore not an ELBO-equivalent score.
It marks where the
model started, measured the only way an autoregressive model can be measured. Elo values here are
comparable within this chart only, and not with the official leaderboard.
Runs of DedeProGames Kiyo-Diffusion-135M on huggingface.co
66
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
62
30-day runs
More Information About Kiyo-Diffusion-135M huggingface.co Model
Kiyo-Diffusion-135M huggingface.co is an AI model on huggingface.co that provides Kiyo-Diffusion-135M's model effect (), which can be used instantly with this DedeProGames Kiyo-Diffusion-135M model. huggingface.co supports a free trial of the Kiyo-Diffusion-135M model, and also provides paid use of the Kiyo-Diffusion-135M. Support call Kiyo-Diffusion-135M model through api, including Node.js, Python, http.
Kiyo-Diffusion-135M huggingface.co is an online trial and call api platform, which integrates Kiyo-Diffusion-135M's modeling effects, including api services, and provides a free online trial of Kiyo-Diffusion-135M, you can try Kiyo-Diffusion-135M online for free by clicking the link below.
DedeProGames Kiyo-Diffusion-135M online free url in huggingface.co:
Kiyo-Diffusion-135M is an open source model from GitHub that offers a free installation service, and any user can find Kiyo-Diffusion-135M on GitHub to install. At the same time, huggingface.co provides the effect of Kiyo-Diffusion-135M install, users can directly use Kiyo-Diffusion-135M installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Kiyo-Diffusion-135M install url in huggingface.co: