Zeroframe offers large‑scale, high‑resolution egocentric video datasets such as EgoExplore and EgoExplore 360, providing long, untrimmed recordings with detailed interval annotations, captions, and synchronized IMU data for multimodal research. It also supplies the Urban3D collection of over 10 k multiview videos of urban objects with calibrated camera poses and point clouds, ready for neural rendering pipelines. These open‑licensed datasets enable researchers and developers to train and evaluate embodied AI, autonomous navigation, and 3D perception systems in realistic, diverse environments.
Funding
Funding not disclosed
Founders
Product
Problem
Researchers and developers building embodied AI, autonomous navigation, and 3D perception systems lack large‑scale, richly annotated egocentric video and multiview 3D datasets that capture real‑world physics, object interactions, and diverse environments. Existing datasets are often short, scripted, or limited to narrow scenes, hindering progress on generalist world models and sensorimotor learning.
Solution
Zeroframe provides open, high‑resolution egocentric video collections such as EgoExplore and EgoExplore 360, featuring long‑duration (average 45 min), 60 fps recordings captured with head‑mounted cameras across a wide range of indoor and outdoor settings worldwide. Each video includes detailed interval‑level annotations (captions, scene type, weather, time of day, crowd density) and synchronized IMU data, enabling research on physics‑aware world modeling, object affordances, and multi‑agent dynamics. In addition, the Urban3D dataset offers over 10 k multiview videos of urban objects with calibrated camera poses, sparse point clouds, and ready‑to‑use assets for neural rendering pipelines (NeRF, Gaussian splatting). Together, these datasets supply the scale, diversity, and precise metadata needed to train and evaluate embodied agents, autonomous systems, and 3D vision algorithms in realistic, unstructured environments.
Target Audience
Primary users are academic labs, AI research teams, and industry groups developing embodied agents, autonomous vehicles, robotics, and 3D perception models that require large, annotated egocentric and multiview datasets.
Features
- EgoExplore: 1080p–4K, 60 fps, ~45 min untrimmed egocentric videos covering 20+ global locations and dozens of scene types
- Interval‑level annotations with natural‑language captions, scene/weather tags, time‑of‑day, and crowd density metadata
- Synchronized IMU recordings for multimodal sensor fusion and motion analysis
- Urban3D: 10 k+ multiview videos of urban objects with COLMAP‑derived camera poses, sparse point clouds, and distortion‑aware intrinsics/extrinsics
- Ready‑to‑use transforms.json for NeRF training and optimized pipelines for Gaussian splatting
- Open licensing for academic and commercial research, with sample 4K clips available on request
- Structured metadata (fps, frame count, duration, interval counts) facilitating automated dataset ingestion and benchmarking