Skip to main content
V

Velvet

Velvet provides a curated repository of multimodal datasets that combine text, images, and video, along with standardized evaluation benchmarks for audiovisual perception tasks. The platform lets AI researchers and engineers download large, annotated datasets and assess model performance against human-reference scores, while also supporting community contributions through a review workflow.

San Francisco, United StatesFounded 20255700+ followers
Updated 3 months ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Researchers lack large, well‑curated multimodal datasets that combine text, images, and video, making it difficult to train and evaluate AI systems on realistic audiovisual communication tasks. Existing resources often focus on a single modality or provide inconsistent benchmarking standards, hindering progress toward human‑like perception.

Solution

Velvet offers a platform for curating, hosting, and sharing multimodal datasets that span text, image, and video content. The service provides standardized evaluation benchmarks designed to measure AI performance on audiovisual communication tasks, enabling developers to compare models against human‑level perception. By centralizing dataset access and offering a consistent testing framework, Velvet streamlines the research workflow and accelerates development of models capable of passing the audiovisual Turing test. Contributors can submit their own datasets for review, expanding the repository and fostering community‑driven growth.

Target Audience

Primary users are AI researchers, machine‑learning engineers, and academic labs developing multimodal models for speech, vision, and language integration, as well as companies building conversational agents that require audiovisual understanding.

Features

  • Repository of publicly available multimodal datasets covering conversational interaction and open‑world exploration scenarios
  • Built‑in evaluation suites that benchmark AI models on audiovisual perception tasks with human‑reference scores
  • Dataset submission portal with review workflow to ensure quality and relevance of contributed data
  • Metadata tagging for medium (text, image, video), volume, and domain to facilitate targeted dataset discovery
  • Secure hosting and download infrastructure supporting large video files and associated annotations
This profile is AI-generated and may contain inaccuracies.