Skip to content

Latest commit

 

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

UI-Venus-2

English | 简体中文

License Report: Coming Soon Website: UI-Venus-2 GitHub Hugging Face Model

A general-purpose foundation GUI agent for mobile apps, web platforms, and desktop operating systems — scaling environments, tasks, and feedback jointly within one closed-loop perception–reasoning–action agent.

🌟 What's New in UI-Venus-2

Compared with UI-Venus-1.5, we introduce:

  • 📱 Scaled multilingual mobile environments: A substantially expanded executable mobile pool covering Chinese and English app ecosystems, paired with a deep-research-driven query-generation strategy that grounds task queries in real application functionality — improving the accuracy, validity, and executability of generated instructions.
  • 🖥️ Computer use, built from the ground up: Dedicated desktop operating-system capabilities constructed from scratch through computer-use data collection and task-specific training, extending the UI-Venus family to mobile, web, and OS interaction in one unified end-to-end agent.
  • 🎯 Keypoint-grounded verification: Task completion is judged on task-relevant visual keypoints rather than a coarse holistic look at the final screen, with multi-model voting aggregating heterogeneous judges — reducing single-judge bias and making the reward signal robust to reward hacking.
  • 🔄 Verification-augmented reflection: Verified feedback is distilled back into training as reflection supervision, so the agent can distinguish partial progress from true completion, avoid premature termination caused by observation misinterpretation, and recover during long-horizon interaction.

📈 UI-Venus-2 Benchmark Performance

UI-Venus-2 Benchmark Performance

Figure UI-Venus-2 on eight GUI-agent benchmarks. Each panel compares UI-Venus-2-27B and UI-Venus-2-9B with four selected strong external baselines. We favor standalone end-to-end systems evaluated on the closest available task subset and step budget; source-reported action scaffolds may still differ. Bars start at zero and scales vary by panel. MobileWorld uses GUI-only success rate on 117 tasks with 50 steps, WebVoyager uses the refreshed 595-task split, Odysseys uses average rubric score over 200 tasks, VenusBench-CAPTCHA uses micro Pass@1 over all 219 examples, and VenusBench-GD uses English-instruction micro-average accuracy. Live-site results may vary with evaluation date. “*” denotes external-model results evaluated or reproduced by us under the stated protocol, not statistical significance.


📰 News

  • [2026/08] We release UI-Venus-2, a 9B/27B general-purpose foundation GUI agent that unifies mobile, web, and desktop interaction with scaled multilingual environments, keypoint-grounded verification, and verification-augmented reflection.
  • [2026/02] We release UI-Venus-1.5, an end-to-end GUI Agent designed for robust real-world applications.
  • [2026/02] We release VenusBench-Mobile, a challenging online benchmark for mobile GUI agents. See branch VenusBench-Mobile.
  • [2025/12] We release VenusBench-GD, a comprehensive multi-platform GUI grounding benchmark. See branch VenusBench-GD.
  • [2025/8] We release UI-Venus, the first version of our UI agent model.

🧭 Overview


✨ Demo

See more demos

UI-Venus-2 mobile demo


🛠 Venus Framework

We provide two frameworks for running agents in real environments:

Framework Description Documentation
Mobile Framework Android/ADB agent framework for single-task execution, multi-device batch execution, trajectory recording, and reflection. English · 中文
Browser Plugin Chrome Side Panel extension that connects UI-Venus to the active browser tab and executes browser tasks interactively. English · 中文

See the Venus Framework overview for the directory layout and entry points. The lightweight domain examples below can be used without either framework unless their individual requirements state otherwise.


🚀 Quick Start

Installation

conda create -n ui-venus-2 python=3.11 -y
conda activate ui-venus-2
pip install -r requirements.txt

Python 3.10 or newer is required. All commands below are executed from the repository root. Configure the OpenAI-compatible model service through MODEL_URL, MODEL_NAME, and either API_KEY or MODEL_API_KEY; the same values can also be edited at the top of each domain script. API_KEY takes precedence when both key variables are set. Set MODEL_NAME to the served name of your UI-Venus-2 9B or 27B model.

Mobile

Run multi-turn inference over the included prerecorded screenshot sequence. N_IMG controls how many recent historical screenshots are retained:

MODEL_URL=http://127.0.0.1:8000/v1 \
MODEL_NAME=UI-Venus-2 \
N_IMG=2 \
bash scripts/mobile.sh

This example performs model inference only and does not execute actions on a device. Use the Mobile Framework above for real-device ADB automation.

Mobile multi-turn example and input/output format

Computer

Run multi-turn Computer inference over a prerecorded desktop screenshot sequence:

MODEL_URL=http://127.0.0.1:8000/v1 \
MODEL_NAME=UI-Venus-2 \
N_IMG=2 \
bash scripts/computer.sh

The default command uses the included desktop screenshot sample. The standalone example validates and normalizes model actions but does not execute them on the host. It has no runtime dependency on OSWorld.

Computer multi-turn example and action format

Browser

Start Chrome with a CDP port as described in the domain documentation, then run one natural-language browser task:

MODEL_URL=http://127.0.0.1:8000/v1 \
MODEL_NAME=UI-Venus-2 \
bash scripts/browser.sh "Open https://example.com and report the page title"

Browser usage and CDP setup

Grounding

Run the direct grounding evaluation on the three samples included with the repository:

MODEL_URL=http://127.0.0.1:8000/v1 \
MODEL_NAME=UI-Venus-2 \
bash scripts/grounding.sh

Grounding evaluation, smoke test, and benchmark configuration

CAPTCHA

Run inference on the included CAPTCHA image and save the parsed JSON and visualization under results/captcha/:

MODEL_URL=http://127.0.0.1:8000/v1 \
MODEL_NAME=UI-Venus-2 \
bash scripts/captcha.sh

CAPTCHA usage, prompt, action format, and visualization


📊 Benchmark Results

Mobile

Models MobileGym VenusBench‑Mobile AndroidWorld MobileWorld KnowUBench MemGUI
General VLMs
Qwen3.5‑9B 9.0* 15.3* 57.8 18.0 (18.0)* 33.3 6.2*
Qwen3.6‑27B 24.6* 28.0* 70.3 36.8 (41.9)* - 25.7*
Claude‑Opus‑4.6 - 36.5* - 44.5 - -
Kimi‑K2.6 38.7* 31.2* - 55.6 - 39.1
Kimi‑K3 - - - 74.4 - -
Seed‑2.0‑Pro 52.0 20.1* - 63.2 51.6 65.6*
Seed‑2.1‑Pro - - - 73.2 - -
GPT‑5.6‑Sol - - - 70.1 - -
GUI‑specific Models
UI‑Venus‑1.5‑8B 18.4* 16.1 73.7 22.2* 26.0 3.9*
UI‑Venus‑1.5‑30B‑A3B 21.5* 21.5 77.6 17.1 - 10.9*
GUI‑Owl‑1.5‑32B‑Instruct 20.3* - 69.8 43.9 - 10.9
MAI‑UI‑8B 21.5* 12.7 70.7 27.5 26.0 17.2*
Qwen‑UI‑Agent‑27B - - - 82.1 (85.5) - -
Ours
UI‑Venus‑2‑9B 52.7 46.5 80.2 65.8 (75.2) 56.5 62.6
UI‑Venus‑2‑27B 60.5 48.7 84.0 76.1 (82.9) 59.7 70.3

Performance comparison on various mobile GUI benchmarks. VenusBench-Mobile reports success rate on its 149-task primary pool. For MobileWorld, we report GUI-only success rate on 117 tasks under the 50-step setting; values in parentheses, when available, use 100 steps. MemGUI reports Main Results Pass@1. * denotes external-model results evaluated or reproduced by us; it does not denote statistical significance.

Computer

Models OSWorld‑Verified OSWorld‑V2 DeskCraft
General VLMs
Claude‑Opus‑4.8 83.4 - -
Qwen3.5‑9B 41.8 2.5 14.6*
Qwen3.6‑27B 62.0 3.8 28.7*
Kimi‑K2.6 73.1 7.1 41.4*
Seed‑2.0‑Pro 62.3 6.3 40.0*
Seed‑2.1‑Pro 78.8 - -
GPT‑5.5 78.7 - -
GUI‑specific Models
GUI‑Owl‑1.5‑32B‑Instruct 56.5 - -
Qwen‑UI‑Agent‑27B 79.5 - -
Ours
UI‑Venus‑2‑9B 70.8 7.52 48.0
UI‑Venus‑2‑27B 80.5 13.24 55.5

Performance comparison on various computer-use agent benchmarks. Reported OSWorld-Verified baselines use the 361-task setting in their cited source and may use model-specific action scaffolds. For DeskCraft, we report an author-evaluated aggregate over the 538-task union of the Standard and Interactive splits; this differs from the benchmark's official split-level reporting. * indicates external-model results evaluated by us and does not denote statistical significance.

Browser

Models WebVoyager Online‑Mind2Web REAL Odysseys Avg. Odysseys Perfect
General VLMs
Qwen3.5‑9B 46.9* 27.3* 18.2* 42.6* 13.5*
Qwen3.5‑4B - - - 42.9 10.7
Qwen3.6‑27B 84.3* 55.3* 27.3* 39.5* 18.5*
OpenAI Operator 87.0 61.3 - - -
GPT‑5 (SoM) 90.6 - - - -
GPT‑5.4 - - - 55.4 33.5
Seed2.0 Pro 85.1* 68.5* 74.4* 60.2* 30.1*
GLM‑5V‑Turbo 88.5 - - - -
Claude Opus 4.6 88.0 - - 68.9 44.5
Claude‑Sonnet‑4.6 - - - 49.8 31.0
Kimi‑K2.6 76.8* - 74.4* - -
GUI‑specific Models
UI‑TARS‑1.5 84.8 75.8 - - -
UI‑Venus‑1.5‑30B‑A3B 76.0 - 38.0* - -
GUI‑Owl‑1.5‑32B‑Thinking 82.1 - 44.6* - -
MolmoWeb‑8B 78.2 35.3 - - -
Fara1.5‑4B 80.8 - - - -
Fara1.5‑9B 86.6 63.4 - - -
Fara1.5‑27B 89.3 72.3 - - -
Ours
UI‑Venus‑2‑9B 90.8 74.0 76.9 77.3 62.0
UI‑Venus‑2‑27B 93.4 78.3 80.2 80.4 66.3

Performance comparison on four live-web benchmarks: WebVoyager, Online-Mind2Web, REAL, and Odysseys. The Fara1.5 and GPT-5 (SoM) WebVoyager entries use the refreshed 595-task, 100-step robust protocol and are averaged over three runs; live-site states may vary by evaluation date. For Odysseys, we report both the averaged rubric score (Avg.) and the perfect rubric score (Perfect). Bold and underlined scores indicate the best and second-best reported results in each column, respectively. * indicates our reproduced results, which are not available from official reports.

Grounding

Models VenusBench‑GD ScreenSpot‑Pro OSWorld‑G‑R UI‑Vision
General VLMs
Qwen 3.7 Plus - 68.9 78.2 68.0
Seed 2.1 Pro - 65.3 78.0 62.0
Kimi‑K2.6 73.1* 52.0* 69.7* -*
Qwen3.6‑27B 67.7* 65.2* 76.9* 58.3*
GUI‑specific Models
UI‑Venus‑Ground‑72B 70.2 61.9 69.5 36.8
Holo2‑30B‑A3B 59.5* 66.1 76.1 40.9*
Step‑GUI‑4B 54.6* 60.0 66.9 30.0*
MAI‑UI‑8B 65.2* 65.8 68.6 40.7
MAI‑UI‑32B - 67.9 73.9 47.1
UI‑Venus‑1.5‑30B‑A3B 75.0 69.6 76.4 54.7
Qwen‑UI‑Agent‑27B - 76.6 78.5 70.0
Ours
UI‑Venus‑2‑9B 77.1 73.0 78.5 53.2
UI‑Venus‑2‑27B 80.1 74.1 79.1 66.9

Performance comparison on various Grounding Benchmarks. VenusBench-GD reports English-instruction micro-average point-in-box accuracy. For each benchmark, bold and underlined scores indicate the best and second-best results, respectively. * indicates external-model results evaluated or reproduced by us under the stated protocol; it does not denote statistical significance.

CAPTCHA

Models VenusBench‑CAPTCHA MCA‑Bench Spatial‑CAPTCHA‑Bench NextGen‑CAPTCHAs Open CaptchaWorld
General VLMs
Qwen3.5‑9B 28.3* 30.4 4.9 2.8 36.4
Qwen3.6‑27B 53.0* 51.7 31.0 14.1 47.7
Doubao‑Seed‑2.0‑Pro 47.9* 35.5 43.6 20.4 55.6
Kimi‑K2.6 39.7* 38.7 24.8 7.2 47.8
Claude‑Opus‑4.6 16.0* 25.9 9.5 2.8 23.3
Ours
UI‑Venus‑2‑9B 78.1 75.7 42.8 47.6 50.7
UI‑Venus‑2‑27B 79.9 79.6 48.6 54.5 56.3

All results are Pass@1 percentages, and higher is better. VenusBench-CAPTCHA reports the micro-average over all 219 fixed examples. We evaluate on VenusBench-CAPTCHA and four public benchmarks: MCA-Bench, Spatial-CAPTCHA-Bench, NextGen-CAPTCHAs, and Open CaptchaWorld. We use 1,000 sampled MCA-Bench examples, 15 NextGen-CAPTCHAs task types, and 16 Open CaptchaWorld task types; see the appendix for selection details. Bold and underlined scores indicate the best and second-best reported results in each column, respectively. For the VenusBench-CAPTCHA column, * denotes external-model results evaluated by us under the same protocol; it is not a significance marker.


📬 Contact

For any questions or collaboration, please contact the maintainers.


📚 Citation

# UI-Venus-2
@misc{venusteam2026uivenus2technicalreport,
      title={UI-Venus-2 Technical Report}, 
      author={Venus-Team and xxx},
      year={2026},
      eprint={xxx},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={comming soon}, 
}

# UI-Venus-1.5
@misc{venusteam2026uivenus15technicalreport,
      title={UI-Venus-1.5 Technical Report}, 
      author={Venus-Team and Changlong Gao and Zhangxuan Gu and Yulin Liu and Xinyu Qiu and Shuheng Shen and Yue Wen and Tianyu Xia and Zhenyu Xu and Zhengwen Zeng and Beitong Zhou and Xingran Zhou and Weizhi Chen and Sunhao Dai and Jingya Dou and Yichen Gong and Yuan Guo and Zhenlin Guo and Feng Li and Qian Li and Jinzhen Lin and Yuqi Zhou and Linchao Zhu and Liang Chen and Zhenyu Guo and Changhua Meng and Weiqiang Wang},
      year={2026},
      eprint={2602.09082},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2602.09082}, 
}

# UI-Venus
@misc{gu2025uivenustechnicalreportbuilding,
      title={UI-Venus Technical Report: Building High-performance UI Agents with RFT}, 
      author={Zhangxuan Gu and Zhengwen Zeng and Zhenyu Xu and others},
      year={2025},
      eprint={2508.10833},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2508.10833}, 
}

⚖️ License

This project is for research and educational purposes only.

About

UI-Venus is a general-purpose foundation GUI agent for mobile apps, web platforms, and desktop operating systems using only screenshots as input.

Topics

Resources

Stars

1.0k stars

Watchers

52 watching

Forks

Releases

Packages

Contributors

Languages