Gradium provides a cloud API delivering modular text‑to‑speech and streaming speech‑to‑text models with ultra‑low latency for real‑time voice applications. The service includes instant voice cloning, multilingual support, word‑level timestamps, and a credit‑based pricing model that scales from prototyping to enterprise workloads.
Funding
Funding not disclosed
Founders
Product
Problem
Real-time voice applications require low‑latency, high‑quality speech synthesis and transcription, but existing services often suffer from noticeable lag, limited language coverage, and high operational costs, making scalable deployment difficult for developers and enterprises.
Solution
Gradium offers a suite of modular audio language models that deliver natural, expressive speech synthesis (TTS) and accurate, streaming transcription (STT) with ultra‑low latency suitable for interactive voice agents. The models run via a cloud API that supports real‑time streaming inference, enabling developers to integrate voice capabilities without managing complex infrastructure. Instant voice cloning lets users create custom speaker profiles on demand, while word‑level timestamps provide precise alignment for downstream processing. Multilingual support (English, French, German, Spanish, Portuguese) and robust performance in noisy environments broaden applicability across global markets. A credit‑based subscription model scales with usage, allowing both prototyping and enterprise‑grade workloads.
Target Audience
The primary customers are developers and product teams building voice‑enabled applications—such as conversational agents, gaming NPCs, customer‑support bots, and language‑learning platforms—and enterprises that need scalable, low‑latency speech AI for production deployments.
Features
- Modular audio language models for independent TTS, STT, or combined voice pipelines
- Streaming inference with controllable latency for real‑time interaction
- Instant voice cloning with up to 1,000 clones per plan and unlimited in custom tiers
- 48 kHz high‑fidelity audio output and word‑level timestamps for exact text‑audio alignment
- Semantic voice‑activity detection for smart turn‑taking and reduced background noise impact
- Code‑switching and multilingual support (EN, FR, DE, ES, PT) with seamless language detection
- RESTful API with configurable concurrency limits (2–15 concurrent streams per plan)
- Credit‑based pricing that maps usage to TTS hours, STT hours, and model concurrency