Fixie develops an open-source speech-to-speech automation platform that utilizes large language models to facilitate real-time, natural communication. The technology addresses the limitations of text-based interactions by enabling seamless, human-like dialogue in fast-paced environments.
Funding
$16.5M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

RFounders
Product
Problem
Current text-based interactions with Large Language Models (LLMs) limit their potential in fast-paced, real-time communication scenarios where natural, human-like dialogue is essential. Existing AI solutions often struggle to handle the nuances of human speech, such as interruptions and overlapping conversations.
Solution
Fixie, now Ultravox.ai, addresses these limitations by developing open-source speech-to-speech models and a comprehensive stack for building and scaling human-like voice AI agents. Their core offering, Ultravox, enables real-time conversation with LLMs by directly processing speech, eliminating the need for a separate automatic speech recognition (ASR) stage. This approach facilitates more natural and fluent interactions, bringing AI closer to human-level communication. Ultravox outperforms existing cascaded ASR → LLM → TTS pipelines in real-world conditions, effectively handling noisy environments and low-quality microphone inputs.
Target Audience
The primary target audience includes developers and companies working on real-time Voice AI applications for customer support, live translation, outbound calls, and accessibility tools.
Features
- Open-source speech-to-speech model trained for real-time conversation with LLMs
- Direct speech consumption, eliminating the need for a separate ASR stage
- Ultravox Realtime APIs and SDKs for building and deploying AI Voice Agents
- Low latency, comparable to GPT-4o Realtime
- Superior transcription accuracy and speech-based question answering compared to other models
- Improved handling of noisy environments and low-quality microphone inputs
- Support for building and scaling human-like voice AI agents