MammothModa2: Jointly Optimized Autoregressive-Diffusion Models for Unified Multimodal Understanding and Generation
Introduction
MammothModa2 is a unified Autoregressive-Diffusion (AR-Diffusion) framework designed for comprehensive multimodal understanding and generation. The model adopts a novel serial architecture: the AR backbone utilizes MammothTok—a unified, language-aligned visual tokenizer—to execute complex semantic planning, which then conditions a high-fidelity Diffusion Decoder. Our core technical contribution is a unified joint training strategy, pioneering the simultaneous optimization of the discrete Next-Token Prediction (NTP) loss and the continuous Flow Matching loss within a serial AR-Diffusion system. This end-to-end alignment between the planning and generation spaces enables MammothModa to achieve competitive performance across complex text-to-image generation, editing, and visual understanding benchmarks.
Show cases
🎉 News
2025-10-01: 🔥MammothModa2-Preview models are now available at
HuggingFace
import torch
from qwen_vl_utils import process_vision_info
from transformers import AutoProcessor
from mammothmoda2.model import Mammothmoda2Model
# Mammothmoda2 model and processor loading.
model = Mammothmoda2Model.from_pretrained(
"bytedance-research/MammothModa2-Preview",
attn_implementation="flash_attention_2",
torch_dtype="bfloat16",
).to("cuda")
print(f"model.device={model.device}")
processor = AutoProcessor.from_pretrained("bytedance-research/MammothModa2-Preview")
# Mammothmoda2 inputs preprocessing.
messages = [
{
"role": "user",
"content": [
{
"type": "image",
"image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg",
},
{"type": "text", "text": "Describe this image."},
],
}
]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs,
padding=True,
padding_side="left",
return_tensors="pt",
return_token_type_ids=False,
).to("cuda")
# Mammothmoda2 model generation and decoding.with torch.inference_mode(), torch.autocast(dtype=torch.bfloat16):
generated_ids = model.generate(**inputs)
generated_ids_trimmed = [out_ids[len(in_ids) :] for in_ids, out_ids inzip(inputs.input_ids, generated_ids)]
output_texts = processor.batch_decode(
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_texts)
📊 Benchmark Results
Model
Model Size
GenEval
DPGBench
Generation
SDXL
-
0.55
74.65
DALL-E 3
-
0.67
83.50
FLUX.1-dev
-
0.67
84.00
SD3.5-Medium*
-
0.65
83.86
Unified
Emu3
8B
0.66
80.60
Janus-Pro
7B
0.80
84.19
MetaQuery-XL
7B + 1.6B
0.80
82.05
UniWorld-V1
7B + 12B
0.84
81.38
Blip3-o-8B
7B + 1.4B
0.84
81.60
OmniGen2
3B + 4B
0.86
83.57
Ovis-U1
2.4B + 1.2B
0.89
83.72
UniPic2
7B + 2B
0.90
83.79
BAGEL
7B + 7B
0.88
85.07
Show-o2
7B
0.76
86.14
GPT-4o
-
0.84
86.23
MammothModa2-Preview
7B + (3B + 2B)
0.85
87.1
Note
: Model sizes in "A + B" format indicate separate understanding (A) and generation (B) parameters. Models without "+" share parameters for both tasks. MammothModa2-Preview uses a 7B + (3B + 2B) architecture, where the 7B parameters are for understanding, and the generation part consists of 3B parameters in the AR (MLLM backbone) and 2B parameters in the DiT component.
Acknowledgement
We are grateful to the following open-source projects:
@misc{mammothmoda2025,
title = {MammothModa2: Jointly Optimized Autoregressive-Diffusion Models for Unified Multimodal Understanding and Generation},
author = {MammothModa Team},
year = {2025},
url = {https://github.com/bytedance/mammothmoda}
}
Runs of bytedance-research MammothModa2-Preview on huggingface.co
485
Total runs
1
24-hour runs
19
3-day runs
25
7-day runs
191
30-day runs
More Information About MammothModa2-Preview huggingface.co Model
MammothModa2-Preview huggingface.co
MammothModa2-Preview huggingface.co is an AI model on huggingface.co that provides MammothModa2-Preview's model effect (), which can be used instantly with this bytedance-research MammothModa2-Preview model. huggingface.co supports a free trial of the MammothModa2-Preview model, and also provides paid use of the MammothModa2-Preview. Support call MammothModa2-Preview model through api, including Node.js, Python, http.
MammothModa2-Preview huggingface.co is an online trial and call api platform, which integrates MammothModa2-Preview's modeling effects, including api services, and provides a free online trial of MammothModa2-Preview, you can try MammothModa2-Preview online for free by clicking the link below.
bytedance-research MammothModa2-Preview online free url in huggingface.co:
MammothModa2-Preview is an open source model from GitHub that offers a free installation service, and any user can find MammothModa2-Preview on GitHub to install. At the same time, huggingface.co provides the effect of MammothModa2-Preview install, users can directly use MammothModa2-Preview installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
MammothModa2-Preview install url in huggingface.co: