Embodied Cognition Foundation Model with Joint Language–Visual Reasoning and Imagination
Tencent Robotics X × Futian Laboratory × Tencent Hunyuan
🔥 Updates
[2026-07]
🎉 We release
Hy-Embodied-RxBrain-1.0
— the technical report, official inference code, and model weights.
📖 Introduction
RxBrain
(
Hy-Embodied-RxBrain-1.0
) is a
unified multimodal foundation model for embodied cognition
— a single model that couples language reasoning with visual imagination to deliver three core capabilities:
🤖
Embodied Understanding & Reasoning
— question answering and chain-of-thought over images and multi-frame video.
🔮
World State Prediction
— imagine the near-future frames an action produces in the physical world.
🧩
Joint Subgoal Planning
— decompose a task into steps, emitting for each step
both
the next action (language)
and
the goal image it should reach (vision).
These capabilities are unified through
interleaved generation
: within a single autoregressive sequence RxBrain alternates reasoning text and flow-matched imagined frames — a learned
<Image>
token decides when to imagine — so an embodied plan couples
what to do
with
what the world should look like
, step by step.
⭐️ Key Features
🧠
Unified Mixture-of-Transformers (MoT):
A ~6.2B-parameter backbone with modality-specific pathways (text / vision / generation), so understanding and image synthesis share one autoregressive model instead of separate towers.
🎨
Flow-Matching Image Head:
Imagined frames are produced by a flow-matching head decoding into a frozen
FLUX
VAE latent space, enabling text-to-image, multi-frame world-model rollout, and goal-image planning.
🔗
Interleaved Reasoning + Imagination:
Text reasoning and generated frames are emitted in one sequence, coupling symbolic plans with visual goals.
Note:
A stock
transformers
release does
not
yet include
hunyuan_vl_mot
; this pinned commit is required. We will merge the improvements into the Transformers main branch later.
Clone the inference code and install the remaining dependencies:
git clone https://github.com/Tencent-Hunyuan/Hy-Embodied-RxBrain-1.0.git
cd Hy-Embodied-RxBrain-1.0
pip install -r requirements.txt
The VQA (understanding) path needs
only
the main weights. Image generation (T2I / world-model rollout / interleaved planning) additionally requires the external
FLUX VAE
ae.safetensors
.
🚀 Quick Start with Transformers
Load the Transformers processor together with the
UnifiedMoT
classes shipped in this repo, then run understanding (VQA). Run this from the repo root so the
model
package is importable, and point
MODEL_PATH
at your
local
download (see
Model Download
).
import torch
from transformers.models.hunyuan_vl_mot import HunYuanVLMoTProcessor
from model import UnifiedMoTForConditionalGeneration, maybe_init_generation_path
from vqa_inference import answer
MODEL_PATH = "./Hy-Embodied-RxBrain-1.0"# local checkpoint directory, not the Hub id
device = torch.device("cuda"if torch.cuda.is_available() else"cpu")
dtype = torch.bfloat16
# Load processor & model
processor = HunYuanVLMoTProcessor.from_pretrained(MODEL_PATH, trust_remote_code=True)
model = UnifiedMoTForConditionalGeneration.from_pretrained(MODEL_PATH, dtype=dtype)
maybe_init_generation_path(model, model_load_path=MODEL_PATH) # wires up the generation path
model.to(device).eval()
# Ask a question about an image
text = answer(
model, processor,
image_paths=["demo_cases/bridgev2_move_toy/input/obs_1.jpg"],
question="What objects are on the stovetop, and where is the green toy?",
device=device, dtype=dtype, max_new_tokens=256,
)
print(text)
Note:
RxBrain uses a custom interleaved text/image decoding loop rather than the standard
model.generate
API. The
answer(...)
helper (in
vqa_inference.py
) wraps that loop for the understanding case; image generation and planning have their own entry points below.
The same tasks are also available as ready-to-run scripts:
① Visual Question Answering (VQA)
— image(s) + question → answer text
Pure autoregressive text understanding —
no VAE / flow-matching needed
.
python vqa_inference.py \
--ckpt ./Hy-Embodied-RxBrain-1.0 \
--images demo_cases/bridgev2_move_toy/input/obs_1.jpg \
--question "What objects are on the stovetop, and where is the green toy?" \
--max_new_tokens 256
② Text-to-Image (T2I)
python text2image_inference.py \
--ckpt ./Hy-Embodied-RxBrain-1.0 --vae /path/to/ae.safetensors \
--prompt "a watercolor painting of a cat" \
--height 256 --width 256 --num_steps 25 --out out.png
# with classifier-free guidance
python text2image_inference.py \
--ckpt ./Hy-Embodied-RxBrain-1.0 --vae /path/to/ae.safetensors \
--prompt "a watercolor painting of a cat" \
--cfg_scale 5.0 --num_steps 50 --out out.png
③ Multi-Frame World-Model Rollout
— imagine future frames from an observation
RxBrain is evaluated on embodied understanding, spatial reasoning, and imagination/generation benchmarks. For detailed metrics and methodology, please refer to our
technical report
.
Runs of tencent Hy-Embodied-RxBrain-1.0 on huggingface.co
233
Total runs
0
24-hour runs
5
3-day runs
47
7-day runs
-155
30-day runs
More Information About Hy-Embodied-RxBrain-1.0 huggingface.co Model
Hy-Embodied-RxBrain-1.0 huggingface.co is an AI model on huggingface.co that provides Hy-Embodied-RxBrain-1.0's model effect (), which can be used instantly with this tencent Hy-Embodied-RxBrain-1.0 model. huggingface.co supports a free trial of the Hy-Embodied-RxBrain-1.0 model, and also provides paid use of the Hy-Embodied-RxBrain-1.0. Support call Hy-Embodied-RxBrain-1.0 model through api, including Node.js, Python, http.
Hy-Embodied-RxBrain-1.0 huggingface.co is an online trial and call api platform, which integrates Hy-Embodied-RxBrain-1.0's modeling effects, including api services, and provides a free online trial of Hy-Embodied-RxBrain-1.0, you can try Hy-Embodied-RxBrain-1.0 online for free by clicking the link below.
tencent Hy-Embodied-RxBrain-1.0 online free url in huggingface.co:
Hy-Embodied-RxBrain-1.0 is an open source model from GitHub that offers a free installation service, and any user can find Hy-Embodied-RxBrain-1.0 on GitHub to install. At the same time, huggingface.co provides the effect of Hy-Embodied-RxBrain-1.0 install, users can directly use Hy-Embodied-RxBrain-1.0 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Hy-Embodied-RxBrain-1.0 install url in huggingface.co: