
SVRC Eval provides a robotics evaluation platform that closes the sim-to-real gap by running policies on real hardware against customer-defined acceptance criteria. The platform captures failures, feeds them back into targeted data collection, and offers automated regression gates with operator-in-the-loop evaluation. It includes managed teleop capture with automatic refinement and LeRobot-v2 packaging.
Funding
Funding not disclosed
Founders
Product
Problem
Robotics teams often develop policies that perform well in simulation or lab settings but fail when deployed on real hardware in production environments. The gap between simulated and real-world performance—caused by actuator physics differences, covariate shift in teleop data, and unanticipated edge cases—causes projects to stall despite having capable models.
Solution
SVRC Eval provides a comprehensive evaluation stack that runs the full loop between lab and production deployment. The platform deploys policies on real hardware, evaluates them against specific acceptance criteria that customers define, and automatically captures failure data to feed back into targeted collection efforts. It includes managed teleop capabilities with automatic refinement, quality control, and action segmentation, plus a command-line interface for streamlined workflows. The platform offers readiness scoring that measures actual deployment readiness rather than vanity metrics, with automated regression gates that catch performance degradation before shipping. An operator-in-the-loop evaluation system allows human judgment on complex tasks like folding shirts, while the entire system runs as a continuous cycle of deployment, capture, evaluation, and refinement.
Target Audience
Robotics companies developing manipulation policies for humanoids, arms, and dexterous hands that need to validate deployment readiness. The platform serves teams working on physical AI applications who require rigorous evaluation between simulation and production deployment.
Features
- Real-hardware deployment with full rollout logging and failure capture
- Acceptance criteria encoding that converts customer-specific requirements into automated evaluation gates
- Automated regression tracking across model versions with metric thresholds
- Operator-in-the-loop evaluation for tasks requiring human judgment
- Managed teleop capture with automatic refinement, quality control, and action segmentation
- LeRobot-v2 packaging for standardized data output
- Readiness scoring based on real rollout telemetry rather than simulated performance
- Command-line interface (centeros refinery CLI) for automated workflows