LLaDA2.2-flash
is an agent-oriented diffusion language model in the LLaDA2 series. By introducing
Levenshtein Editing
(with
DELETE
and
INSERT
control tokens) to diffusion language modeling, it represents the LLaDA2 series' first step in agentic applications, including long-context tool use, multi-turn interaction, and robust error correction.
📊 Benchmarks
The following tables compare
LLaDA2.2-flash
and
Ling-2.6-flash
in terms of agentic benchmark scores and throughput (TPS).
Agentic benchmark scores
Benchmark
LLaDA2.2-flash
Ling-2.6-flash
SWE-bench Verified
49.28
61.20
†
SWE-bench Pro
30.10
31.88
SWE-bench Multilingual
25.00
33.73
τ²-Bench
80.33
76.36
†
Claw-Eval
64.22
64.56
†
PinchBench
81.66
81.30
†
MCP-Atlas
46.21
41.12
BFCL-V4
60.78
66.81
LLaDA2.2-flash evaluation setup:
The SWE-bench series was evaluated using the Claude Code scaffold. Across all benchmarks, we used a 128K context window with
temperature=1.0
,
block_length=32
,
threshold=0.5
, and
editing_threshold=0.0
. Each score represents the average of five runs.
†
The Ling-2.6-flash score on SWE-bench Verified is taken from the Ling and Ring 2.6 Technical Report, where it was obtained using the OpenHands scaffold. The Ling-2.6-flash scores on τ²-Bench, Claw-Eval, and PinchBench are also sourced from the technical report, whereas its SWE-bench Pro and SWE-bench Multilingual scores were evaluated by us using the same Claude Code scaffold as LLaDA2.2-flash.
Throughput (TPS)
Benchmark
LLaDA2.2-flash (TPS)
Ling-2.6-flash (TPS)
SWE-bench Verified
519.0
303.2
SWE-bench Pro
485.3
283.4
SWE-bench Multilingual
459.5
200.6
$\tau^2$-Bench
592.8
334.9
BFCL-V4
703.82
331.5
Ling-2.6-flash evaluation setup:
MTP was enabled with 4 draft tokens.
More results will be released in the upcoming technical report.
🚀 Highlights
Efficient 128K Diffusion Infrastructure
: LLaDA2.2-flash extends the context window to
128K
and introduces
Block Routing
, which bounds MoE expert activation at the diffusion-block level to enable efficient long-context agentic workloads.
Levenshtein Editing
: We introduces
DELETE
and
INSERT
control tokens, allowing diffusion decoding to edit sequence structure, remove redundant content, and create insertion slots during parallel generation.
Agentic Reinforcement Learning
: We propose
Levenshtein Editing ELBO-based Block-level Policy Optimization (L-EBPO)
, which leverages agentic environmental rewards to train levenshtein editing and error correction in multi-turn tool-use scenarios.
📦 Model Variants
Model ID
Description
Hugging Face Link
inclusionAI/LLaDA2.2-flash
Agent-oriented MoE diffusion language model with Levenshtein Editing.
To achieve optimal performance, we recommend starting with the following settings:
Sampling Parameters
: Use
block_length=32
,
temperature=0.0
,
top_p=None
, and
top_k=None
as stable default settings.
Denoising Thresholds
: Tune
threshold
,
editing_threshold
, and
max_post_steps
according to the speed-quality trade-off required by the application. Lower thresholds may improve inference speed but can lead to increased repetition or unstable outputs.
Long-Context Agentic Workloads
: For long-context tool-use and multi-turn agent applications, we recommend using
SGLang
as the serving backend. Please ensure that the serving stack is configured for the 128K context window and the model's MoE diffusion inference requirements.
🤖 ModelScope
If you are in mainland China, we strongly recommend accessing our model from 🤖
ModelScope
LLaDA2.2-flash huggingface.co is an AI model on huggingface.co that provides LLaDA2.2-flash's model effect (), which can be used instantly with this inclusionAI LLaDA2.2-flash model. huggingface.co supports a free trial of the LLaDA2.2-flash model, and also provides paid use of the LLaDA2.2-flash. Support call LLaDA2.2-flash model through api, including Node.js, Python, http.
LLaDA2.2-flash huggingface.co is an online trial and call api platform, which integrates LLaDA2.2-flash's modeling effects, including api services, and provides a free online trial of LLaDA2.2-flash, you can try LLaDA2.2-flash online for free by clicking the link below.
inclusionAI LLaDA2.2-flash online free url in huggingface.co:
LLaDA2.2-flash is an open source model from GitHub that offers a free installation service, and any user can find LLaDA2.2-flash on GitHub to install. At the same time, huggingface.co provides the effect of LLaDA2.2-flash install, users can directly use LLaDA2.2-flash installed effect in huggingface.co for debugging and trial. It also supports api for free installation.