Ishiki Labs provides Fern, a cloud‑hosted learned simulator that generates high‑fidelity, action‑conditioned RGB video streams from real robot teleoperation data. By predicting future frames for arbitrary joint‑space commands, it lets robotics researchers and companies run scalable, reproducible evaluations of manipulation and navigation policies without needing physical hardware.
Funding
Funding not disclosed
Founders
Product
Problem
Evaluating general robot policies currently relies on running limited rollouts on physical hardware, which is time‑consuming, non‑reproducible, and unsafe for large‑scale reinforcement learning. Existing hand‑crafted simulators fail to capture the full visual and contact dynamics of real robots, limiting their usefulness for benchmarking and training.
Solution
Ishiki Labs offers Fern, a learned high‑fidelity simulator that generates action‑conditioned RGB video streams directly from real robot teleoperation data. By training end‑to‑end world models on collected trajectories, Fern can predict future frames for arbitrary joint‑space commands, enabling researchers to benchmark and improve policies without accessing physical robots. The platform provides consistent physics across multiple synchronized camera views and supports 16‑DoF bimanual setups, allowing apples‑to‑apples comparisons of different policies. Fern is hosted in the cloud, so users can run evaluations instantly, scale to millions of rollouts, and avoid hardware wear, operator time, and safety risks. Public benchmark suites and custom world models for client data further extend its applicability across manipulation, navigation, and mobile‑manipulation tasks.
Target Audience
Primary users are robotics research labs, reinforcement‑learning teams, and robotics companies that need scalable, reproducible evaluation of manipulation and navigation policies.
Features
- Action‑conditioned diffusion‑forcing transformer that rolls out 256×256 RGB frames at real‑time rates from joint‑space commands
- Supports 16‑DoF bimanual robots with four synchronized camera viewpoints (left wrist, right wrist, chest, waist) ensuring cross‑view physical consistency
- Learned end‑to‑end physics captures contacts, shadows, cables, and gripper interactions directly from real‑world data
- Cloud‑hosted inference on a single GPU, enabling on‑demand simulation without hardware setup
- Public benchmark catalog for manipulation, navigation, and mobile‑manipulation tasks with standardized evaluation metrics
- Custom world‑model training pipeline that repurposes a client’s own teleoperation datasets to create tailored simulators