Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Why federal judge ruled First Amendment protects certain AI-generated child sex abuse material

    September 6, 2026

    Niger military accuses France of orchestrating failed mutiny: What to know | News

    September 6, 2026

    That sound you hear is not the world ending – it’s liberalism being monstered | Waleed Aly

    September 6, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Why federal judge ruled First Amendment protects certain AI-generated child sex abuse material
    • Niger military accuses France of orchestrating failed mutiny: What to know | News
    • That sound you hear is not the world ending – it’s liberalism being monstered | Waleed Aly
    • Long-term allies and verifiable reality must make way for Saving Private Nige | John Crace
    • My Brief Summer Fling With Siri AI
    • Attackers conceal phishing lures using invisible Unicode characters
    • 600 BTC Mined in 2010 Moves After 16 Years of Dormancy
    • US envoys meet Putin: What’s behind latest diplomacy on Russia-Ukraine war? | Russia-Ukraine war News
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Sunday, September 6
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKSeptember 6, 2026 Artificial Intelligence No Comments4 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    A team of researchers from UC Berkeley have released CUA-Lite, an open platform for computer-use agents (CUAs). The argument behind it is infrastructural rather than model-centric: training and benchmarking a CUA requires four pieces: agents, environments, traces, and a framework to evaluate and train them and all four are currently fragmented across separate repositories with incompatible interfaces. CUA-Lite puts them behind one action space, one data schema, and one command, across desktop, browser and mobile.

    Is it deployable? Yes. The stack installs with uv sync --all-extras on Python 3.12, and its lightweight sandboxes run on any Docker host without /dev/kvm, so cloud instances, CI runners and nested containers all work.

    The VM tax, and how Lite.OSWorld removes it

    The most concrete contribution is Lite.OSWorld. OSWorld provides a faithful Ubuntu desktop, but it ships as a full QEMU/KVM virtual machine per task, requiring nested virtualization that most managed infrastructure does not expose. CUA-Lite reproduces the same task suite and the same evaluators on a GNOME desktop inside a plain Docker container.

    Task OSWorld Lite.OSWorld
    Runtime QEMU/KVM VM Docker container
    Host requirement /dev/kvm, nested virt Any Docker host
    Memory 4.1 GB 0.9 GB
    Cold start 29.9 s 23.8 s
    Parallelism baseline ~4.6× more instances
    Task suite OSWorld Identical

    Fidelity is the obvious concern when you swap a VM for a container, and the team addresses it directly: across 13 models, Lite.OSWorld scores match the OSWorld VM’s, so a score or a training signal earned in the container transfers back to the real benchmark. The same base now carries a family of sandboxes: Lite.ScaleCUA, Lite.CUAGym and Lite.CUAWorld, the last expanding into roughly 40 applications including Blender, QGIS and VS Code. In total the platform claims 30k+ verifiable tasks.

    One schema for data, one adapter per model

    CUA-Lite’s second layer is LiteSample, a single supervised-learning schema shared across every environment, agent and task type, shipped as plain parquet plus images. Ten-plus existing CUA datasets have been preprocessed into it and published free on Hugging Face, including Aguvis, OpenCUA, ScaleCUA, GUI-360, GUIOdyssey and Multimodal-Mind2Web. Alongside those corpora sit fresh rollout datasets generated by rolling a frontier teacher model through the sandboxes, for distillation into smaller students.

    Because model families expect different scaffolding, the framework ships a per-model adapter that packs a unified LiteSample into each model’s own training format, including history collapsing so several steps share one forward pass.

    Eval, SFT and RL behind one command

    Agents and environments meet in lite.gym: screenshots up, actions down, with one action space per platform. 10+ agents are built in GPT, Claude, Gemini, Qwen3-VL, UI-TARS, Fara-7B, MAI-UI and others, and 15+ benchmarks are integrated, spanning grounding (ScreenSpot-Pro, OSWorld-G), desktop (OSWorld, OSWorld-2, WindowsAgentArena, CUABench), browser (WebArena, VisualWebArena, MiniWoB, WebVoyager, Online-Mind2Web, WebGym) and mobile (AndroidWorld, AndroidLab, MobileWorld, MobileGym). Swapping --model-id and --env-id in scripts/rollout.py is the whole interface.

    The same loop serves training. For SFT, the README documents fine-tuning Qwen3-VL-2B-Instruct on Lite.ScaleCUA desktop trajectories, lifting mean episode return from 0.138 to 0.237 on the 332-task lite.osworld eval split, a single reported configuration on two GPUs, not an independently reproduced result. For RL, rollouts scored in the environment drive GRPO updates on top of Slime, with a worked MobileGym example covering 416 mobile tasks across 28 apps.

    Interactive explainer

    Key Takeaways

    • CUA-Lite unifies agents, environments, traces and training under one action space and one LiteSample schema.
    • Lite.OSWorld runs OSWorld tasks VM-free in Docker at 0.9 GB versus 4.1 GB, roughly 4.6× more parallel desktops.
    • Scores in the container match the OSWorld VM across 13 models, so training signal transfers to the real benchmark.
    • 30k+ verifiable tasks, 15+ benchmarks, 10+ agents, and 20+ datasets published free on Hugging Face.
    • Deployable on any Docker host, but the repository ships no explicit license yet — verify terms before commercial use.

    Check out the Project Page, GitHub Repo and Datasets on Hugging Face. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


    Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

    Agents Berkeley ComputerUse CUALite data evaluation open Platform release researchers Sandboxes unifying
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    67,000 More Trezor Customers Exposed as Data Breach Widens

    Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed

    GitHub Introduces Project HydraFusion: Runtime Multi-Model Orchestration That Builds a Workflow Per Coding Task in Copilot CLI

    Nous Research Adds One-Click Local Model Setup to Hermes Desktop

    US Labor Data Beat Sends Bitcoin Lower Amid Fed Rate Uncertainty

    Trezor Says ShipMonk Breach Exposed 67,000 U.S. Customers’ Data It Said Was Deleted

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Why federal judge ruled First Amendment protects certain AI-generated child sex abuse material

    September 6, 2026

    Niger military accuses France of orchestrating failed mutiny: What to know | News

    September 6, 2026

    That sound you hear is not the world ending – it’s liberalism being monstered | Waleed Aly

    September 6, 2026

    Long-term allies and verifiable reality must make way for Saving Private Nige | John Crace

    September 6, 2026
    Latest Posts

    Quantum computing nears commercial breakthrough, IBM CEO says

    August 1, 2026

    Amgen says cloud data breach exposed patient health, proprietary info

    August 1, 2026

    Snapchat joins other platforms in the fight against ‘AI slop’

    August 1, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Why federal judge ruled First Amendment protects certain AI-generated child sex abuse material

    September 6, 2026

    Niger military accuses France of orchestrating failed mutiny: What to know | News

    September 6, 2026

    That sound you hear is not the world ending – it’s liberalism being monstered | Waleed Aly

    September 6, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.