
Pantograph is a research lab developing generally intelligent robots using simple, scalable methods rather than task-specific engineering. The company builds the Pandroid bimanual mobile robot and publishes benchmark evaluations of frontier large language models controlling physical systems. Its public research includes comparative testing of models like GPT-6 Astra and Claude Fable on manipulation tasks, emphasizing low-cost hardware and generic prompting approaches.
Funding
Funding not disclosed
Founders
Product
Problem
General-purpose robotics remains difficult to achieve with traditional, task-specific programming, which requires extensive engineering for each new use case. Current robots are often too narrow to address large-scale societal challenges such as automated vaccine production, affordable housing construction, or climate-change mitigation infrastructure, limiting their real-world impact to repetitive, constrained industrial tasks.
Solution
Pantograph is a research lab pursuing generally intelligent robots through simple, scalable methods rather than bespoke control systems. The company develops the Pandroid, a bimanual mobile robot with a vertical stage, and evaluates frontier vision-language models as robot controllers using a generic tool-calling harness. By providing models with camera feeds, action primitives, and minimal prompting, Pantograph measures raw model capability across manipulation tasks and publishes its findings openly. This approach aims to accelerate the diffusion of robotic capabilities into society, enabling rapid adaptation of physical infrastructure, from urban planning to energy systems and medical research.
Target Audience
Primary audiences are AI research institutions, robotics developers, and technology policy organizations interested in advancing general-purpose robotic systems and understanding frontier model performance in physical environments.
Features
- Pandroid bimanual mobile robot designed with a one-dimensional vertical stage for simplified, robust manipulation control
- Generic evaluation harness that provides VLMs with three 480p camera feeds, action history, and parameterized tool calls, with no task-specific prompt tailoring
- Standardized benchmarking methodology across eight manipulation tasks, including blind grading, interleaved trials, and teleoperator baselines
- Published comparative results of frontier LLMs, such as GPT-6 Astra outperforming Claude Fable 5.1 (36% vs. 15% success), with statistical significance testing
- Prompt-based safety interventions instead of hard-coded constraints to correct unsafe motions like driving grippers into the floor
- Operational design using small, handled robots that enable rapid manual reset between trials for high-throughput experiments