Skip to main content
R

Rhoda

Rhoda trains robot control policies on web‑scale video data using its Direct Video‑Action paradigm, turning causal video predictions into real‑time robot actions. This enables robots to perform complex, long‑horizon tasks with minimal task‑specific data and one‑shot imitation from human demonstrations, providing interpretable video rollouts for safety verification. The solution is aimed at manufacturers, logistics providers, automotive firms, and e‑commerce operations that need adaptable robotic agents for material handling, assembly, and fulfillment.

Updated 2 months ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Current robotic systems are largely specialized, performing single repetitive tasks in controlled environments and lacking the ability to generalize across diverse, unstructured settings. Scaling robot intelligence to handle varied real‑world scenarios requires massive, diverse data, which traditional robot datasets cannot provide.

Solution

Rhoda addresses this gap by training robot control policies on web‑scale video data through its Direct Video‑Action (DVA) paradigm. The approach formulates robot decision‑making as real‑time video prediction: a pre‑trained causal video model forecasts future frames based on long visual histories, and an inverse dynamics model translates these predictions into robot actions. This closed‑loop pipeline runs multiple times per second, enabling robots to perform complex, long‑horizon tasks with minimal task‑specific data (as little as ~10 hours). The long‑context visual memory supports one‑shot imitation of human demonstrations, while the video‑first generation provides interpretable rollouts for safety verification. By leveraging the vast amount of publicly available video, Rhoda’s models can scale more efficiently than traditional robot‑centric data collection methods.

Target Audience

Primary customers are manufacturers, logistics providers, automotive firms, and e‑commerce operations seeking to deploy adaptable robotic agents for material handling, assembly, and fulfillment tasks across variable environments.

Features

  • Direct Video‑Action Model that converts causal video predictions into robot actions in a closed‑loop, high‑frequency control loop
  • Long‑context visual memory handling hundreds of frames, enabling multi‑step task planning and one‑shot learning from a single demonstration
  • Data‑efficient learning requiring only ~10 hours of robot‑specific interaction to achieve reliable performance on complex tasks
  • Interpretability through generated video rollouts, allowing visual inspection of predicted behavior and safety verification
  • Integration of additional conditioning signals (e.g., proprioception, language) to enrich video predictions for diverse task specifications
  • Scalable pre‑training on web‑scale video datasets, providing broad physical knowledge beyond limited robot interaction logs
This profile is AI-generated and may contain inaccuracies.