Skip to main content
M

Manudata

Manudata captures egocentric, 1080p60 video of skilled hands on India’s MSME factory floors and processes it through an automated pipeline that extracts 3D hand keypoints, 6DoF camera pose via SLAM, and high‑precision timestamps. The resulting RLDS‑ready datasets provide robotics researchers with low‑cost, real‑world manipulation data for training and evaluating robot learning models.

Updated 27 days ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Robotics research requires large amounts of high‑quality manipulation data, but existing datasets are expensive to collect and often lack realistic, egocentric perspectives of skilled manual work. This limits the ability to train and evaluate robot learning algorithms for complex hand‑tool interactions.

Solution

Manudata captures 1080p60 egocentric video of skilled workers on Indian factory floors and automatically processes the footage into robot‑ready datasets. A seven‑stage pipeline extracts 6‑DoF camera poses via SLAM, reconstructs 3‑D hand keypoints (21 per hand), generates per‑pixel object masks, anchors scenes with AprilTags, and adds temporal action and language annotations. The output conforms to the Reinforcement Learning Dataset Specification (RLDS), enabling seamless integration with frontier robotics labs. By leveraging the extensive MSME workforce in India, Manudata lowers data acquisition costs compared to traditional US/EU programs while providing realistic, high‑fidelity manipulation recordings.

Target Audience

Primary customers are robotics research labs and companies developing manipulation algorithms that require large‑scale, high‑quality egocentric hand‑tool datasets for training and evaluation.

Features

  • Cap‑mounted egocentric camera capturing 1080p video at 60 fps
  • SLAM‑based 6‑DoF camera pose recovery for full scene reconstruction
  • Monocular 3‑D hand tracking delivering 21 keypoints per hand with 99.9 % detection accuracy
  • Per‑pixel object segmentation masks for precise environment modeling
  • AprilTag world anchoring to provide consistent spatial references across captures
  • Automated temporal action labeling and natural‑language annotation
  • End‑to‑end seven‑stage pipeline that outputs RLDS‑compatible episodes without manual annotation
This profile is AI-generated and may contain inaccuracies.