Zarya is a family of hybrid language models that combine a classic auto-regressive (AR) objective with a masked-diffusion (MDM) objective in one model.
The architecture can be built on top of any autoregressive model but in this repository it uses the
Qwen3
backbone.
Naming explanation: Zarya (pronounced as [zɐˈrʲa] (
IPA notation
), literally "Dawn" in English) is a figure from Slavic folklore — a female personification of dawn who may be considered a goddess.
In various traditions, she can manifest as a single being or as two or three sisters simultaneously.
This is a research prototype.
Model Details
Model Description
Zarya is a research prototype of a family of hybrid language models that jointly learn a classic auto-regressive (AR) objective and a masked-diffusion (MDM) objective within a single model.
Two generation modes are supported, both reachable through a single
model.generate(...)
call: masked-diffusion (MDM) sampling and slotted-level speculative parallel decoding.
Model type:
Hybrid auto-regressive (AR) + masked-diffusion language model (DLLM); backbone
Qwen3
, wrapper
Zarya
Zarya is intended for text generation.
It supports conversational fine-tuning (SFT) and classic auto-regressive pretraining.
Direct Use
Direct use is text generation (continuation of a prompt) through the
model.generate(...)
interface, including chat-style prompts formatted with the provided chat template.
Two inference modes are available through the same
generate()
call.
Both modes fully use the KV cache with causal attention masks.
MDM sampling
(
slotted_generation=false
): iterative masked-diffusion denoising with the first-hitting sampler.
Slotted speculative decoding
(
slotted_generation=true
): parallel slot generation with inter-slot diffusion-based selection and intra-slot autoregressive generation for a decoding speedup.
Out-of-Scope Use
The model is a research prototype.
It should not be used for production decisions, safety-critical applications, or any use case where accuracy and reliability are essential without additional evaluation and safeguards.
Inference performance and stability also depend on the chosen decoding hyperparameters (like
slotted_generation
,
slot_size
,
serial_num_blocks
,
slot_threshold
,
token_threshold
, and others).
Bias, Risks, and Limitations
This is a research prototype.
The code relies on Hugging Face Transformers APIs; when upgrading versions, compatibility must be checked (tested on Transformers 5.12.1 and PyTorch 2.9.0).
Inference performance and stability depend on the choice of config parameters.
How to Get Started with the Model
Use the code below to get started with the model. Loading the model and tokenizer requires
trust_remote_code=True
.
import torch
from transformers import AutoModel, AutoTokenizer
model_name = "ai-forever/Zarya-4B"
model = AutoModel.from_pretrained(model_name, trust_remote_code=True, torch_dtype=torch.bfloat16).cuda()
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
prompt = "<|im_start|>user\nHello!<|im_end|>\n<|im_start|>assistant\n"
input_ids = tokenizer(prompt, return_tensors="pt").input_ids.cuda()
# Both modes go through model.generate(...).# With generation_config.slotted_generation=true -> slotted speculative decoding:
out = model.generate(
input_ids,
max_new_tokens=256,
do_sample=True,
temperature=0.7,
slot_size=16,
serial_num_blocks=4,
slot_threshold=0.9,
token_threshold=0.3,
)
# Setting generation_config.slotted_generation=false -> MDM sampling instead:# out = model.generate(input_ids, max_new_tokens=256)print(tokenizer.decode(out[0]))
Evaluation
LM-eval benchmarking with the
lm-eval
package is supported. Example run:
Measurements below were collected with varying inference parameters and on different GPUs; performance is sensitive to both, so results may differ across configurations and hardware setups.
If you find our work helpful, please consider citing (citation will be updated after peer-reviewed publication):
@misc{sinev-etal-2026-Zarya,
author = {Sinev, Leonid and Koziev, Ilya and Leshchuk, Vladislav},
title = {Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference},
year = {2026},
archiveprefix = {arXiv},
eprint = {2609.19868},
primaryclass = {cs.CL},
url = {https://arxiv.org/abs/2609.19868},
}
Runs of ai-forever Zarya-4B on huggingface.co
328
Total runs
6
24-hour runs
16
3-day runs
34
7-day runs
326
30-day runs
More Information About Zarya-4B huggingface.co Model
Zarya-4B huggingface.co is an AI model on huggingface.co that provides Zarya-4B's model effect (), which can be used instantly with this ai-forever Zarya-4B model. huggingface.co supports a free trial of the Zarya-4B model, and also provides paid use of the Zarya-4B. Support call Zarya-4B model through api, including Node.js, Python, http.
Zarya-4B huggingface.co is an online trial and call api platform, which integrates Zarya-4B's modeling effects, including api services, and provides a free online trial of Zarya-4B, you can try Zarya-4B online for free by clicking the link below.
ai-forever Zarya-4B online free url in huggingface.co:
Zarya-4B is an open source model from GitHub that offers a free installation service, and any user can find Zarya-4B on GitHub to install. At the same time, huggingface.co provides the effect of Zarya-4B install, users can directly use Zarya-4B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.