Skip to main content
O

Odyssey

Odyssey offers a pre‑trained multimodal transformer that jointly processes audio waveforms and video frames, exposed through a sub‑100 ms API and SDKs for Unity, Unreal, Python, C++, and JavaScript. The platform enables game studios, edtech developers, simulation engineers, and ad tech teams to embed unified perception, reasoning, and generative capabilities with optional fine‑tuning and on‑premise deployment for domain‑specific adaptation and data‑privacy compliance.

Palo Alto, United StatesFounded 2023447K+ followers
Updated 1 month ago

Funding

$337M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

AAV+8

Founders

Product

Problem

Companies creating immersive digital experiences often have to combine separate vision and audio models, which leads to fragmented perception, higher engineering overhead, and inconsistent content generation across gaming, education, simulation, and advertising use cases.

Solution

Odyssey delivers a general‑purpose world model that learns joint audio‑visual representations from large‑scale video datasets. The model is exposed through a low‑latency API and SDKs, enabling developers to embed unified perception, reasoning, and generative capabilities directly into their applications. A built‑in fine‑tuning pipeline lets teams adapt the foundation model to domain‑specific vocabularies or visual styles without retraining from scratch. Real‑time inference is optimized for both GPU servers and edge devices, and native plugins support integration with Unity, Unreal, and other common development environments. Secure, encrypted communication and optional on‑premise deployment address enterprise data‑privacy requirements. Ongoing model updates expand the knowledge base, allowing customers to benefit from the latest multimodal advances without additional engineering effort.

Target Audience

Primary customers are game studios, edtech platform developers, simulation engineers, and advertising technology teams that require a unified audio‑visual AI layer to accelerate content creation and interactive intelligence.

Features

  • Multimodal transformer architecture that jointly processes audio waveforms and video frames for coherent context understanding
  • Pre‑trained on a curated corpus of >10 million hours of synchronized audio‑visual content, providing broad world knowledge out of the box
  • Real‑time inference endpoint delivering sub‑100 ms latency for 1080p video and 48 kHz audio streams
  • Fine‑tuning toolkit with data‑annotation helpers, gradient checkpointing, and domain‑specific loss functions
  • Unity and Unreal Engine plugins that expose model APIs as native components, enabling drag‑and‑drop integration
  • SDKs for Python, C++, and JavaScript with sample code for game logic, interactive tutoring, and dynamic ad creative generation
  • Role‑based access control and end‑to‑end TLS encryption; optional on‑premise container deployment for regulated industries
  • Continuous model improvement pipeline that incorporates new public video/audio releases while preserving backward compatibility
This profile is AI-generated and may contain inaccuracies.