OpenCUA models (OpenCUA-7B, OpenCUA-32B and OpenCUA-72B) are end-to-end computer-use foundation models than can produce executable actions in the computer environments. They are based on the weights of Qwen2.5-VL-7B-Instruction, Qwen2.5-VL-32B-Instruction and Qwen2.5-VL-72B-Instruction.
They demonstrate superior performance across CUA benchmarks. In particular,
OpenCUA-72B
achieves an average success rate of
45.0%
on
OSWorld-Verified
,
establishing a new state-of-the-art (SOTA) among open-source models and surpassing OpenAI CUA (GPT-4o). It also has a strong grounding performance, and achieves
60.8
on ScreenSpot-Pro and
37.3
(SOTA) on UI-Vision.
Key Features
Superior Computer-Use Capablity
: Able to execute multi-step computer-use actions with effective planning and reasoning
Multi-OS Support
: Trained on demonstrations across Ubuntu, Windows, and macOS
Visual Grounding
: Strong GUI element recognition and spatial reasoning capabilities
Multi-Image Context
: Processes up to 3 screenshot history for better context understanding
Reflective Reasoning
: Enhanced with reflective long Chain-of-Thought that identifies errors and provides corrective reasoning
Performance
Online Agent Evaluation
OpenCUA models achieves strong performance on
OSWorld-Verified
.
OPENCUA-32B achieves the best performance among all open-source models with an average success rate of 34.8%, outperforming prior baselines by large margins.
It also closes the gap to proprietary Claude models.
Model
15 Steps
50 Steps
100 Steps
Proprietary
OpenAI CUA
26.0
31.3
31.4
Seed 1.5-VL
27.9
—
34.1
Claude 3.7 Sonnet
27.1
35.8
35.9
Claude 4 Sonnet
31.2
43.9
41.5
Open-Source
Qwen 2.5-VL-32B-Instruct
3.0
—
3.9
Qwen 2.5-VL-72B-Instruct
4.4
—
5.0
Kimi-VL-A3B
9.7
—
10.3
UI-TARS-72B-DPO
24.0
25.8
27.1
UI-TARS-1.5-7B
24.5
27.3
27.4
OpenCUA-7B
(Ours)
24.3
27.9
26.6
OpenCUA-32B
(Ours)
29.7
34.1
34.8
OpenCUA-72B
(Ours)
39.0
44.9
45.0
OpenCUA scores are the mean of 3 independent runs.
GUI Grounding Performance
Model
OSWorld-G
ScreenSpot-V2
ScreenSpot-Pro
Qwen2.5-VL-7B
31.4
88.8
27.6
Qwen2.5-VL-32B
46.5
87.0
39.4
UI-TARS-72B
57.1
90.3
38.1
OpenCUA-A3B
48.6
91.4
28.5
OpenCUA-Qwen2-7B
45.7
88.5
23.7
OpenCUA-7B
55.3
92.3
50.0
OpenCUA-32B
59.6
93.4
55.3
OpenCUA-72B
-
92.9
60.8
AgentNetBench (Offline Evaluation)
Model
Coordinate Actions
Content Actions
Function Actions
Average
Qwen2.5-VL-7B
50.7
40.8
3.1
48.0
Qwen2.5-VL-32B
66.6
47.2
41.5
64.8
Qwen2.5-VL-72B
67.2
52.6
50.5
67.0
OpenAI CUA
71.7
57.3
80.0
73.1
OpenCUA-7B
79.0
62.0
44.3
75.2
OpenCUA-32B
81.9
66.1
55.7
79.1
🚀 Quick Start
⚠️ Important for Qwen-based Models (OpenCUA-7B, OpenCUA-32B):
To align with our training infrastructure, we have modified the model in two places:
1. Multimodal Rotary Position Embedding (M-RoPE) has been replaced with 1D RoPE.
2. Using the same Tokenizer and ChatTemplate as Kimi-VL.
Do not use the default transformers and vllm classes to load the model. Tokenizer and Chat Template should be aligned if training the models.
Installation & Download
First, install the required transformers dependencies:
cd ./model/inference/
python huggingface_inference.py
🖥️ Computer Use Agent
OpenCUAAgent
is developed in the
OSWorld
environment based on OpenCUA models. It iteratively perceives the environment via screenshots, produces reflective long CoT as inner monologue, and predicts the next action to be executed. OpenCUAAgent uses 3 images in total and L2 CoT format in default.
Command for running OpenCUA-7B and OpenCUA-32B in OSWorld:
AgentNet is the first large-scale desktop computer-use agent trajectory dataset, containing 22.6K human-annotated computer-use tasks across Windows, macOS, and Ubuntu systems.
Our
AgentNetTool
is a cross-platform GUI recorder that runs unobtrusively on annotators’ machines. It captures synchronized
screen video
,
mouse/keyboard events
, and
accessibility trees
, then provides an in-browser UI for reviewing, trimming, and submitting demonstrations. AgentNet Tool is available on Windows, macOS and Ubuntu.
Raw demonstrations can contain thousands of low-level events that are too dense for model training.
The
DataProcessor
module (
./data/data-process/
) performs two key steps:
Action Reduction
— merges granular signals into concise, semantically meaningful PyAutoGUI actions (e.g., collapsing mouse moves → click, coalescing scrolls, grouping key-press sequences into text or hotkeys).
State–Action Matching
— aligns every reduced action with the
last visually distinct frame
before
the action begins, avoiding future-information leakage and yielding compact state–action pairs.
These processed trajectories underlie all downstream training and evaluation.
3 CoTGenerator – Synthesizing Reflective Long Chain-of-Thought Inner Monologue
To boost robustness and interpretability, we augment each trajectory with
reflective long Chain-of-Thought (CoT) reasoning
.
The
CoTGenerator
pipeline (
./data/cot-generator/
) synthesizes step-level reflections that:
reflect on the previous action,
explain
why
an action is chosen given the current observation and history,
note potential alternative actions, and
forecast the expected next state.
Empirically, models trained with these rich CoTs scale better with data and generalize across unseen applications.
Evaluation
AgentNetBench
(
./AgentNetBench/
) provides a realistic offline evaluator for OS agent trajectories. It compares model-predicted low-level actions (click, moveTo, write, press, scroll, terminate, etc.) against ground-truth human actions and reports detailed metrics.
We are actively working with the vLLM team to add support for OpenCUA models.
Workaround:
For now, please use the standard transformers library as shown in the examples above. We will update this section once vLLM support becomes available.
Training Code
OpenCUA models are developed based on the training infrastructure of Kimi Team. We are developting the training pipeline based on the open-source infrastructure as well.
Acknowledge
We thank Su Yu, Caiming Xiong, Binyuan Hui, and the anonymous reviewers for their insightful discussions and valuable feedback.
We are grateful to Moonshot AI for providing training infrastructure and annotated data.
We also sincerely appreciate Calvin, Ziwei Chen, Jin Zhang, Ze Li, Zhengtao Wang, Yanxu Chen, and Qizheng Gu from the Kimi Team for their strong infrastructure support and helpful guidance.
The development of our tool is based on the open-source projects-
DuckTrack
and
OpenAdapt
.
We are very grateful to their commitment to the open source community. Finally, we extend our deepest thanks to all annotators for their tremendous effort and contributions to this project.
License
This project is licensed under the MIT License - see the LICENSE file in the root folder for details.
Research Use and Disclaimer
OpenCUA models are intended for
research and educational purposes only
.
Prohibited Uses
The model may
not
be used for any purpose or activity that violates applicable laws or regulations in any jurisdiction
Use for illegal, unethical, or harmful activities is strictly prohibited
Disclaimer
The authors, contributors, and copyright holders are
not responsible
for any illegal, unethical, or harmful use of the Software, nor for any direct or indirect damages resulting from such use
Use of the "OpenCUA" name, logo, or trademarks does
not
imply any endorsement or affiliation unless separate written permission is obtained
Users are solely responsible for ensuring their use complies with applicable laws and regulations
Important Notes on Coordinate Systems
OpenCUA/OpenCUA-A3B
– Relative coordinates
(not supported in this code)
OpenCUA/OpenCUA-Qwen2-7B
– Relative coordinates
OpenCUA/OpenCUA-7B
– Absolute coordinates
OpenCUA/OpenCUA-32B
– Absolute coordinates
OpenCUA models use different coordinate systems depending on the base model:
OpenCUA-Qwen2-7B
: Outputs
relative coordinates
(0.0 to 1.0 range)
# Example output: pyautogui.click(x=0.5, y=0.3)# x=0.5 means 50% from left edge, y=0.3 means 30% from top edge# Convert to absolute coordinates:defqwen2_relative_to_absolute(rel_x, rel_y, original_width, original_height):
abs_x = int(rel_x * original_width)
abs_y = int(rel_y * original_height)
return abs_x, abs_y
OpenCUA-7B and OpenCUA-32B
(Qwen2.5-based): Output
absolute coordinates
after smart resize
# Example output: pyautogui.click(x=960, y=324) # These are coordinates on the smart-resized image, not the original image# Convert to original image coordinates:# Please refer to the smart_resize function in: https://github.com/huggingface/transformers/blob/67ddc82fbc7e52c6f42a395b4a6d278c55b77a39/src/transformers/models/qwen2_vl/image_processing_qwen2_vl.py#L55defqwen25_smart_resize_to_absolute(model_x, model_y, original_width, original_height):
# First, calculate the smart-resized dimensions
resized_height, resized_width = smart_resize(original_height, original_width, factor = 28, min_pixels = 3136, max_pixels = 12845056)
# Convert model output to relative coordinates on original image
rel_x = model_x / resized_width
rel_y = model_y / resized_height
# Then convert to absolute coordinates on original image
abs_x = int(rel_x * original_width)
abs_y = int(rel_y * original_height)
return abs_x, abs_y
Understanding Smart Resize for Qwen2.5-based Models:
The Qwen2.5-VL models use a “smart resize” preprocessing that maintains aspect ratio while fitting within pixel constraints.
For coordinate conversion, you need the smart resize function from the
official Qwen2.5-VL implementation
.
Citation
If you use OpenCUA models in your research, please cite our work:
@misc{wang2025opencuaopenfoundationscomputeruse,
title={OpenCUA: Open Foundations for Computer-Use Agents},
author={Xinyuan Wang and Bowen Wang and Dunjie Lu and Junlin Yang and Tianbao Xie and Junli Wang and Jiaqi Deng and Xiaole Guo and Yiheng Xu and Chen Henry Wu and Zhennan Shen and Zhuokai Li and Ryan Li and Xiaochuan Li and Junda Chen and Boyuan Zheng and Peihang Li and Fangyu Lei and Ruisheng Cao and Yeqiao Fu and Dongchan Shin and Martin Shin and Jiarui Hu and Yuyan Wang and Jixuan Chen and Yuxiao Ye and Danyang Zhang and Dikang Du and Hao Hu and Huarong Chen and Zaida Zhou and Haotian Yao and Ziwei Chen and Qizheng Gu and Yipu Wang and Heng Wang and Diyi Yang and Victor Zhong and Flood Sung and Y. Charles and Zhilin Yang and Tao Yu},
year={2025},
eprint={2508.09123},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2508.09123},
}
Runs of xlangai OpenCUA-72B-preview on huggingface.co
66
Total runs
0
24-hour runs
5
3-day runs
19
7-day runs
-6
30-day runs
More Information About OpenCUA-72B-preview huggingface.co Model
OpenCUA-72B-preview huggingface.co is an AI model on huggingface.co that provides OpenCUA-72B-preview's model effect (), which can be used instantly with this xlangai OpenCUA-72B-preview model. huggingface.co supports a free trial of the OpenCUA-72B-preview model, and also provides paid use of the OpenCUA-72B-preview. Support call OpenCUA-72B-preview model through api, including Node.js, Python, http.
OpenCUA-72B-preview huggingface.co is an online trial and call api platform, which integrates OpenCUA-72B-preview's modeling effects, including api services, and provides a free online trial of OpenCUA-72B-preview, you can try OpenCUA-72B-preview online for free by clicking the link below.
xlangai OpenCUA-72B-preview online free url in huggingface.co:
OpenCUA-72B-preview is an open source model from GitHub that offers a free installation service, and any user can find OpenCUA-72B-preview on GitHub to install. At the same time, huggingface.co provides the effect of OpenCUA-72B-preview install, users can directly use OpenCUA-72B-preview installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
OpenCUA-72B-preview install url in huggingface.co: