
DreamScale Labs
DreamScale Labs provides cloud-based, real-time inference infrastructure for large world action models used in robotics. The platform enables robots to run compute-intensive models remotely, delivering sub-second inference speeds and lower latency than existing cloud providers. It supports concurrent deployments, allowing multiple robots to operate from a single H100 GPU.
- Artificial Intelligence
- AI Agents
- Developer Tools
- Robotics
- Software Only
Funding
Founders
Product
Problem
Generalist robotics requires increasingly large AI models, such as world action models and vision-language-action models, that exceed the computational capacity of onboard hardware. Running these large models locally on robots is constrained by slow inference latency, forcing teams to choose between intelligence and responsiveness, which limits real-world deployment and safety.
Solution
DreamScale Labs provides a cloud-based inference platform specifically engineered for real-time execution of large physical AI models. By optimizing GPU kernels, network layers, and implementing real-time chunking, the platform achieves 40–50% lower tail latency than existing cloud providers and runs a 14B-parameter DreamZero model at 320 ms—faster than NVIDIA's reported speed with half the silicon. This enables robotics teams to offload their most compute-intensive models to the cloud, supporting concurrent operation of multiple robots and preserving the tight control loops required for safe, responsive actuation.
Target Audience
Primary customers are robotics companies and research labs developing generalist physical AI systems—such as humanoid robot builders and embodied AI teams—that require large, compute-intensive models for real-time control.
Features
- Cloud-based serving of frontier world action models like NVIDIA's DreamZero, with 14B-parameter inference at 320 ms on less hardware than standard deployments
- Optimized GPU kernel and network stack design that reduces tail latency by 40–50% compared to existing cloud providers
- Real-Time Chunking (RTC) technology that partitions inference workloads for low-latency, closed-loop control over remote connections
- Support for running 8+ robots concurrently on a single H100 GPU, maximizing hardware utilization
- Flexible model deployment: users can start with a frontier model or bring their own trained checkpoints
- Full-stack optimization across model, GPU kernels, network, and scheduling layers, eliminating the need for robotics teams to handle low-level infrastructure tuning