Molmo2-ER
(Embodied Reasoning) is a 4B vision–language model specialized for the embodied perception skills that downstream action models depend on: scene understanding, pixel-accurate pointing, multi-image and egocentric–exocentric correspondence, and video temporal reasoning.
It is built on top of
Molmo2
(Qwen3-4B backbone + SigLIP2 vision encoder) and serves as the vision–language backbone of the
MolmoAct2
action reasoning model.
Highlights
Outperforms every open-weight baseline
as well as the strongest closed-source models — including
Gemini Robot-ER 1.5 Thinking
and
GPT-5
— on
9 of 13
established embodied reasoning benchmarks (Point-Bench, RefSpatial, BLINK, CV-Bench, ERQA, EmbSpatial, MindCube, SAT, VSI-Bench).
Overall average 63.8%
, a
+17 point
improvement over the Molmo2 starting point.
Training
Molmo2-ER is trained from the released Molmo2 checkpoint with a two-stage
specialize-then-rehearse
recipe:
MolmoAct2 generate robot actions from visual observations and language instructions, but their behavior may vary across embodiments, environments, and hardware configurations. Users should carefully validate model outputs before deployment, especially when operating physical robots or other actuated systems. Where possible, actions should be monitored through interpretable intermediate outputs (adaptive depth map), simulation rollouts, action limits, or other safety checks before execution on hardware. The model’s action space should be bounded by the training data, robot controller limits, and task-specific safety constraints, including limits on speed, workspace, torque, and contact force. Users should follow the hardware manufacturer’s safety guidelines, use appropriate emergency-stop mechanisms, and operate the system only in a safely configured environment with human supervision.
Citation
@misc{fang2026molmoact2actionreasoningmodels,
title={MolmoAct2: Action Reasoning Models for Real-world Deployment},
author={Haoquan Fang and Jiafei Duan and Donovan Clay and Sam Wang and Shuo Liu and Weikai Huang and Xiang Fan and Wei-Chuan Tsai and Shirui Chen and Yi Ru Wang and Shanli Xing and Jaemin Cho and Jae Sung Park and Ainaz Eftekhar and Peter Sushko and Karen Farley and Angad Wadhwa and Cole Harrison and Winson Han and Ying-Chun Lee and Eli VanderBilt and Rose Hendrix and Suveen Ellawela and Lucas Ngoo and Joyce Chai and Zhongzheng Ren and Ali Farhadi and Dieter Fox and Ranjay Krishna},
year={2026},
eprint={2605.02881},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2605.02881},
}
Runs of allenai Molmo2-ER on huggingface.co
2.6K
Total runs
0
24-hour runs
-3.1K
3-day runs
-3.3K
7-day runs
-3.3K
30-day runs
More Information About Molmo2-ER huggingface.co Model
Molmo2-ER huggingface.co is an AI model on huggingface.co that provides Molmo2-ER's model effect (), which can be used instantly with this allenai Molmo2-ER model. huggingface.co supports a free trial of the Molmo2-ER model, and also provides paid use of the Molmo2-ER. Support call Molmo2-ER model through api, including Node.js, Python, http.
Molmo2-ER huggingface.co is an online trial and call api platform, which integrates Molmo2-ER's modeling effects, including api services, and provides a free online trial of Molmo2-ER, you can try Molmo2-ER online for free by clicking the link below.
allenai Molmo2-ER online free url in huggingface.co:
Molmo2-ER is an open source model from GitHub that offers a free installation service, and any user can find Molmo2-ER on GitHub to install. At the same time, huggingface.co provides the effect of Molmo2-ER install, users can directly use Molmo2-ER installed effect in huggingface.co for debugging and trial. It also supports api for free installation.