Gbox provides a cloud‑based platform with ready‑to‑use sandboxes for browsers, Android devices, and Linux desktops, enabling autonomous agents to perform UI actions through a unified SDK. The service includes built‑in reinforcement‑learning environments and benchmark suites, as well as a grounding model that converts natural‑language commands into precise UI operations, with options for secure on‑premise deployment.
Funding
Funding not disclosed
Founders
Product
Problem
Developing and evaluating autonomous agents that can interact with graphical user interfaces across browsers, mobile devices, and desktop operating systems requires realistic, scalable environments and reliable action models. Existing tools are often fragmented, difficult to set up, and lack unified APIs for controlling UI elements, hindering rapid iteration and benchmarking of agent performance.
Solution
Gbox offers a cloud‑based platform that delivers ready‑to‑use sandboxes for Android, web browsers, and Linux desktops, allowing agents to perform UI actions such as clicks, drags, swipes, and scrolls via a simple SDK. The platform includes reinforcement‑learning (RL) environments and built‑in industry benchmarks (e.g., OSWorld, AndroidWorld, WebArena) for training, evaluation, and comparison of agents on real‑world tasks. A grounding model translates natural‑language commands into precise UI operations with high accuracy and low latency, enabling agents to execute complex workflows without custom scripting. Users can deploy the stack on‑premise with air‑gapped containers, customize benchmarks, and integrate enterprise security controls, providing both flexibility and compliance for enterprise use cases.
Target Audience
Primary customers are AI research labs, enterprise AI teams, and developers building autonomous agents for web, mobile, and desktop automation, as well as organizations requiring secure, on‑premise AI evaluation environments.
Features
- Browser, Android, and Linux sandboxes with instant environment initialization (<150 ms start time) and support for both simulated and physical devices
- Unified SDK exposing high‑level actions (click, drag, swipe, scroll) using natural‑language parameters for seamless UI manipulation
- Pre‑built RL environments and benchmark suites (OSWorld, AndroidWorld, WebArena) that run out‑of‑the‑box for training and performance evaluation
- Grounding model delivering 92.5 % action accuracy and sub‑second response (<1 s) across general and challenging UI tasks
- Live analytics and replay capabilities for debugging agent interactions and verifying task completion via UI and database validators
- On‑premise, air‑gapped deployment option with encrypted storage, SSO integration, and customizable benchmark stacks
- Cost‑effective pricing at $0.50 per million inputs and outputs, enabling scalable experimentation