Asilarkness / DiffuRefill-1B-base

huggingface.co
Total runs: 34
24-hour runs: 0
7-day runs: 34
30-day runs: 34
Model's Last Updated: October 02 2026
text-generation

Introduction of DiffuRefill-1B-base

Model Details of DiffuRefill-1B-base

DiffuRefill-1B (base)

A 1.08B-parameter masked-diffusion language model, pretrained from scratch for $0 on free, preemptible notebook GPUs.

  • 200,000 steps, 92.8B tokens , 21 days (2026-09-10 → 2026-10-01)
  • ~200 worker launches on free RTX PRO 6000 notebooks that are reclaimed without warning; the whole fleet was lost and rebuilt many times
  • The only shared state was this Hub: workers train 150 steps alone, upload, and a restartable merger averages them (DiLoCo)
  • 📄 Technical report: paper/DiffuRefill-1B-technical-report.pdf

This is a base model : no instruction tuning, no safety tuning. It writes fluent English and often gets facts wrong.

What it is

DiffuRefill is a bidirectional transformer trained with the masked-diffusion (MDLM) objective: random tokens are replaced by [MASK] and the model learns to fill them in from both sides. Instead of writing text strictly left to right, it starts from a row of masks and fills it in over several passes, and it can commit several tokens per pass.

Parameters 1.08B (tied embeddings)
Layers / width / heads 18 / 2048 / 16
FFN SwiGLU, 5632
Positions / norm RoPE / RMSNorm (pre-norm)
Context 2048
Tokenizer openbmb/MiniCPM4-0.5B
Weights uniform average of the last 130 merged models (steps ~181k–200k)
Usage
# pip install torch transformers safetensors huggingface_hub
from huggingface_hub import hf_hub_download
import importlib.util, torch

path = hf_hub_download("Asilarkness/DiffuRefill-1B-base", "modeling_diffurefill.py")
spec = importlib.util.spec_from_file_location("modeling_diffurefill", path)
mod = importlib.util.module_from_spec(spec); spec.loader.exec_module(mod)

model = mod.DiffuRefill.from_pretrained("Asilarkness/DiffuRefill-1B-base").cuda()

# one token per pass: slowest, most coherent
print(model.generate("The theory of relativity states that", max_new_tokens=64))
# 4 tokens per pass: ~4x fewer passes, some quality loss
print(model.generate("The theory of relativity states that", max_new_tokens=64, per_pass=4))

# scoring: log p(continuation | context), the scorer used for the benchmarks
print(model.loglikelihood("The capital of France is", " Paris."))   # -2.03
print(model.loglikelihood("The capital of France is", " Berlin."))  # -8.11
Benchmarks

Zero-shot, lm-evaluation-harness , every model run by us with the same harness and settings. Our model is plugged into the harness with a left-to-right chain-rule scorer ( eval/dlm_lmeval.py ). A Monte-Carlo ELBO estimator, as used by LLaDA, scored slightly lower (HellaSwag 40.8 vs 42.5 on 1000 items), so the scorer does not understate the model.

Model Type Params Tokens Train FLOPs HellaSwag ARC-e ARC-c PIQA WinoGr. OBQA SciQ LAMBADA MMLU
DiffuRefill-1B diffusion 1.08B 0.09T 6.0e20 34.8 52.4 25.0 62.7 50.6 27.4 89.5 34.0 26.9
Pythia-1B AR 1.0B 0.3T 1.8e21 47.2 56.9 27.0 69.4 54.1 31.2 83.4 55.8 23.1
SmolLM2-360M AR 0.36B 4T 8.6e21 56.3 70.3 38.1 71.5 58.9 37.8 91.1 53.4 25.5
Qwen2.5-0.5B AR 0.49B 18T 5.3e22 52.2 64.9 31.8 69.9 56.1 35.0 93.0 52.0 47.5
SmolLM2-1.7B AR 1.7B 11T 1.1e23 71.2 77.8 47.1 77.6 65.9 44.0 93.3 67.5 48.5
Qwen2.5-1.5B AR 1.5B 18T 1.7e23 67.8 75.5 45.0 75.9 63.5 40.6 94.3 62.2 59.7

HellaSwag, ARC-c, PIQA, OBQA: length-normalised accuracy; the rest: accuracy. FLOPs ≈ 6·params·tokens.

Read this honestly. DiffuRefill-1B is clearly behind autoregressive models of its size. Every one of them used 3× (Pythia) to 280× (Qwen2.5-1.5B) more training compute, and masked diffusion is known to need much more compute than autoregression for the same likelihood (≈16× by Nie et al., 2025 ). It is relatively strong on SciQ (science facts, matching the fact-heavy anneal) and weakest on LAMBADA and WinoGrande (at chance). TinyLlama and OLMo-1B were also run but load incorrectly under the current transformers (e.g. OLMo LAMBADA 0%), so they are left out. All raw numbers are in eval/results.json .

Speed

Greedy generation of 256 tokens after a 32-token prompt, one RTX PRO 6000 (96 GB), bf16. Autoregressive models use Hugging Face generate() with a KV cache; DiffuRefill uses no cache and re-reads the whole row each pass.

tokens / s DiffuRefill-1B SmolLM2-1.7B Qwen2.5-1.5B
batch 1, 1 token per pass 193 101 82
batch 1, 8 tokens per pass 1536 – –
batch 32, 1 token per pass 431 3119 2612
batch 32, 8 tokens per pass 3449 – –

At batch 1 (latency) diffusion wins: 1.9× at the same quality setting and 15× at 8 tokens per pass. At batch 32 (throughput) the KV cache wins unless diffusion commits 8 tokens per pass. More tokens per pass costs quality, and optimised AR servers such as vLLM are much faster than HF generate() .

How it was trained
Objective MDLM masked diffusion, independent token masking; document-masked attention; salient span masking on 25% of rows; 10% of micro-batches as 128-token rows
Optimiser AdamW (0.9, 0.95), wd 0.1, clip 1.0; per-worker batch 64×2048 = 131k tokens
Distributed DiLoCo over the Hub: 150 local steps, plain averaging (outer lr 1, no momentum), 2–8 workers
LR warm-up 500 → 3e-4 (halved to 1.5e-4 early), cosine to 10%, then linear to zero from step 180k
Data, steps 0–158.8k Ultra-FineWeb-L3 (English Multi-Style and QA synthetic), later with 5% generated fact sentences
Data, steps 158.8k–200k anneal mix: Dolmino (OLMo 3 stage 2) 20%, Nemotron Fact-Seeking 15%, Wiki-Rewrite 15%, Ultra-FineWeb-L3 18%, Nemotron-CC-Math 8%, UltraData-Code L3 7%, OpenCoder annealing 8%, UltraData-Math exercises 4%, generated facts 5%
Hardware free RTX PRO 6000 Blackwell (96 GB) notebooks, ~35k tokens/s per GPU
Lessons (details in the report)
  • ✅ Adopt the merged model by replacement . Keeping local drift made 60 of 63 merges worse.
  • ✅ Halve the LR for the merge interval. Local drift fell 6× and a 1.3B-token plateau ended.
  • ✅ Document masking and 10% short rows helped. Uniform weight averaging was the best final checkpoint.
  • ❌ Outer Nesterov momentum, a mid-run switch to Muon , span/PMI masking , phrase tokens and cautious AdamW did not help this run.
  • 🔬 On a small test model, writing a chain of thought first and then filling the answer in one parallel pass lifted a multi-step arithmetic task from 45% to 100%. That is the plan for the reasoning stage.
Limitations

This is one run, and interventions were chosen on small probes and a 38M test model. The fact probe overlaps with the Wikipedia-derived anneal data. The model is English only and has no safety tuning, so do not use it for anything that matters without further training.

Citation
@misc{diffurefill2026,
  title  = {DiffuRefill-1B: Pretraining a Masked Diffusion Language Model on Free, Preemptible Notebook GPUs},
  author = {Asilarkness},
  year   = {2026},
  url    = {https://huggingface.co/Asilarkness/DiffuRefill-1B-base}
}

Compute was provided free by molab notebooks. Engineering was done with Claude Code (Anthropic) as an assistant. The training run's working repository, with intermediate checkpoints and logs, is Asilarkness/DiffuRefill-1B .

Runs of Asilarkness DiffuRefill-1B-base on huggingface.co

34
Total runs
0
24-hour runs
34
3-day runs
34
7-day runs
34
30-day runs

More Information About DiffuRefill-1B-base huggingface.co Model

More DiffuRefill-1B-base license Visit here:

https://choosealicense.com/licenses/apache-2.0

DiffuRefill-1B-base huggingface.co

DiffuRefill-1B-base huggingface.co is an AI model on huggingface.co that provides DiffuRefill-1B-base's model effect (), which can be used instantly with this Asilarkness DiffuRefill-1B-base model. huggingface.co supports a free trial of the DiffuRefill-1B-base model, and also provides paid use of the DiffuRefill-1B-base. Support call DiffuRefill-1B-base model through api, including Node.js, Python, http.

DiffuRefill-1B-base huggingface.co Url

https://huggingface.co/Asilarkness/DiffuRefill-1B-base

Asilarkness DiffuRefill-1B-base online free

DiffuRefill-1B-base huggingface.co is an online trial and call api platform, which integrates DiffuRefill-1B-base's modeling effects, including api services, and provides a free online trial of DiffuRefill-1B-base, you can try DiffuRefill-1B-base online for free by clicking the link below.

Asilarkness DiffuRefill-1B-base online free url in huggingface.co:

https://huggingface.co/Asilarkness/DiffuRefill-1B-base

DiffuRefill-1B-base install

DiffuRefill-1B-base is an open source model from GitHub that offers a free installation service, and any user can find DiffuRefill-1B-base on GitHub to install. At the same time, huggingface.co provides the effect of DiffuRefill-1B-base install, users can directly use DiffuRefill-1B-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

DiffuRefill-1B-base install url in huggingface.co:

https://huggingface.co/Asilarkness/DiffuRefill-1B-base

Url of DiffuRefill-1B-base

DiffuRefill-1B-base huggingface.co Url

Provider of DiffuRefill-1B-base huggingface.co

Asilarkness
ORGANIZATIONS

Other API from Asilarkness

huggingface.co

Total runs: 4.6K
Run Growth: -5.0K
Growth Rate: -108.31%
Updated:August 26 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:July 11 2026