This is a
base model
: no instruction tuning, no safety tuning. It writes fluent English and often gets facts wrong.
What it is
DiffuRefill is a bidirectional transformer trained with the masked-diffusion (MDLM) objective: random tokens are replaced by
[MASK]
and the model learns to fill them in from both sides. Instead of writing text strictly left to right, it starts from a row of masks and fills it in over several passes, and it can commit several tokens per pass.
uniform average of the last 130 merged models (steps ~181k–200k)
Usage
# pip install torch transformers safetensors huggingface_hubfrom huggingface_hub import hf_hub_download
import importlib.util, torch
path = hf_hub_download("Asilarkness/DiffuRefill-1B-base", "modeling_diffurefill.py")
spec = importlib.util.spec_from_file_location("modeling_diffurefill", path)
mod = importlib.util.module_from_spec(spec); spec.loader.exec_module(mod)
model = mod.DiffuRefill.from_pretrained("Asilarkness/DiffuRefill-1B-base").cuda()
# one token per pass: slowest, most coherentprint(model.generate("The theory of relativity states that", max_new_tokens=64))
# 4 tokens per pass: ~4x fewer passes, some quality lossprint(model.generate("The theory of relativity states that", max_new_tokens=64, per_pass=4))
# scoring: log p(continuation | context), the scorer used for the benchmarksprint(model.loglikelihood("The capital of France is", " Paris.")) # -2.03print(model.loglikelihood("The capital of France is", " Berlin.")) # -8.11
Benchmarks
Zero-shot,
lm-evaluation-harness
, every model run by us with the same harness and settings. Our model is plugged into the harness with a left-to-right chain-rule scorer (
eval/dlm_lmeval.py
). A Monte-Carlo ELBO estimator, as used by LLaDA, scored slightly
lower
(HellaSwag 40.8 vs 42.5 on 1000 items), so the scorer does not understate the model.
Read this honestly.
DiffuRefill-1B is clearly behind autoregressive models of its size. Every one of them used 3× (Pythia) to 280× (Qwen2.5-1.5B) more training compute, and masked diffusion is known to need much more compute than autoregression for the same likelihood (≈16× by
Nie et al., 2025
). It is relatively strong on SciQ (science facts, matching the fact-heavy anneal) and weakest on LAMBADA and WinoGrande (at chance). TinyLlama and OLMo-1B were also run but load incorrectly under the current
transformers
(e.g. OLMo LAMBADA 0%), so they are left out. All raw numbers are in
eval/results.json
.
Speed
Greedy generation of 256 tokens after a 32-token prompt, one RTX PRO 6000 (96 GB), bf16. Autoregressive models use Hugging Face
generate()
with a KV cache; DiffuRefill uses no cache and re-reads the whole row each pass.
tokens / s
DiffuRefill-1B
SmolLM2-1.7B
Qwen2.5-1.5B
batch 1, 1 token per pass
193
101
82
batch 1, 8 tokens per pass
1536
–
–
batch 32, 1 token per pass
431
3119
2612
batch 32, 8 tokens per pass
3449
–
–
At batch 1 (latency) diffusion wins: 1.9× at the same quality setting and 15× at 8 tokens per pass. At batch 32 (throughput) the KV cache wins unless diffusion commits 8 tokens per pass. More tokens per pass costs quality, and optimised AR servers such as vLLM are much faster than HF
generate()
.
How it was trained
Objective
MDLM masked diffusion, independent token masking; document-masked attention; salient span masking on 25% of rows; 10% of micro-batches as 128-token rows
free RTX PRO 6000 Blackwell (96 GB) notebooks, ~35k tokens/s per GPU
Lessons (details in the report)
✅ Adopt the merged model by
replacement
. Keeping local drift made 60 of 63 merges worse.
✅
Halve the LR for the merge interval.
Local drift fell 6× and a 1.3B-token plateau ended.
✅
Document masking
and
10% short rows
helped. Uniform weight averaging was the best final checkpoint.
❌ Outer
Nesterov
momentum, a
mid-run switch to Muon
,
span/PMI masking
,
phrase tokens
and
cautious AdamW
did not help this run.
🔬 On a small test model, writing a chain of thought first and then filling the answer
in one parallel pass
lifted a multi-step arithmetic task from 45% to 100%. That is the plan for the reasoning stage.
Limitations
This is one run, and interventions were chosen on small probes and a 38M test model. The fact probe overlaps with the Wikipedia-derived anneal data. The model is English only and has no safety tuning, so do not use it for anything that matters without further training.
Citation
@misc{diffurefill2026,
title = {DiffuRefill-1B: Pretraining a Masked Diffusion Language Model on Free, Preemptible Notebook GPUs},
author = {Asilarkness},
year = {2026},
url = {https://huggingface.co/Asilarkness/DiffuRefill-1B-base}
}
Compute was provided free by molab notebooks. Engineering was done with Claude Code (Anthropic) as an assistant. The training run's working repository, with intermediate checkpoints and logs, is
Asilarkness/DiffuRefill-1B
.
Runs of Asilarkness DiffuRefill-1B-base on huggingface.co
34
Total runs
0
24-hour runs
34
3-day runs
34
7-day runs
34
30-day runs
More Information About DiffuRefill-1B-base huggingface.co Model
DiffuRefill-1B-base huggingface.co is an AI model on huggingface.co that provides DiffuRefill-1B-base's model effect (), which can be used instantly with this Asilarkness DiffuRefill-1B-base model. huggingface.co supports a free trial of the DiffuRefill-1B-base model, and also provides paid use of the DiffuRefill-1B-base. Support call DiffuRefill-1B-base model through api, including Node.js, Python, http.
DiffuRefill-1B-base huggingface.co is an online trial and call api platform, which integrates DiffuRefill-1B-base's modeling effects, including api services, and provides a free online trial of DiffuRefill-1B-base, you can try DiffuRefill-1B-base online for free by clicking the link below.
Asilarkness DiffuRefill-1B-base online free url in huggingface.co:
DiffuRefill-1B-base is an open source model from GitHub that offers a free installation service, and any user can find DiffuRefill-1B-base on GitHub to install. At the same time, huggingface.co provides the effect of DiffuRefill-1B-base install, users can directly use DiffuRefill-1B-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
DiffuRefill-1B-base install url in huggingface.co: