Skip to main content
V

Velvet

Velvet offers a curated video-data platform that aggregates, validates, and distributes large‑scale video datasets optimized for world‑model training, applying automated quality‑control and metadata extraction to deliver standardized TFRecord or Parquet bundles. Users access versioned datasets through a searchable web catalog or token‑protected REST API, with flexible licensing options for commercial and academic projects.

Honolulu, United States20+ followers
Updated 2 months ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Researchers and engineers building world models often struggle to obtain large‑scale, high‑quality video datasets that are consistently formatted, annotated, and licensed for commercial use. Existing sources are fragmented, lack rigorous quality control, and require extensive preprocessing before they can be fed into training pipelines.

Solution

Velvet operates a curated video‑data platform that aggregates, validates, and distributes datasets optimized for world‑model training. The service ingests raw footage from contributors, runs automated quality‑assessment and metadata extraction pipelines, and stores the results in standardized formats compatible with major ML frameworks. Users can browse the catalog via a web dashboard or programmatically retrieve data through a RESTful API with token‑based authentication. Velvet also offers versioned dataset releases and licensing options that simplify compliance for commercial projects. By handling end‑to‑end data acquisition, cleaning, and delivery, the platform reduces the time and engineering effort required to build robust simulation environments.

Target Audience

Primary customers are AI research labs, ML engineers, and robotics or autonomous‑vehicle teams that require large, high‑fidelity video corpora for training world models. The platform also serves data providers looking to monetize unique video collections.

Features

  • Automated ingestion pipeline that normalizes video resolution, frame rate, and encoding (e.g., H.264, ProRes) and generates TFRecord/Parquet bundles for efficient loading
  • Built‑in quality‑control suite leveraging computer‑vision heuristics (blur detection, motion consistency) and human review to ensure dataset integrity
  • Rich metadata schema (scene descriptors, camera parameters, timestamps) searchable via faceted filters in the web catalog
  • Secure, token‑protected REST API with pagination and checksum verification for programmatic access
  • Dataset versioning and immutable snapshots to support reproducible research and model auditing
  • Flexible licensing models (commercial, academic, royalty‑free) with clear attribution tracking
  • Contributor marketplace that enables external data owners to submit samples and receive compensation through integrated Stripe/Wise payouts
  • Dashboard analytics showing usage metrics, download statistics, and dataset popularity trends
This profile is AI-generated and may contain inaccuracies.