
Fluxions
Fluxions provides on-device and API-based voice AI technology, including a fully private iPhone voice assistant called vui and a comprehensive speech-to-text and text-to-speech platform. The company owns its entire AI stack, offering all-inclusive per-minute pricing for voice agents and a la carte transcription and speech services. Their technology emphasizes privacy, with on-device processing for the consumer app and on-prem deployment options for enterprise clients.
- Artificial Intelligence
- AI Agents
- Developer Tools
- Software Only
Funding
Founders
Product
Problem
Traditional voice AI solutions require cloud-based processing, which raises privacy concerns and creates dependency on internet connectivity. Businesses face complex pricing structures that stack separate charges for speech recognition, language models, and voice synthesis, making costs unpredictable and difficult to manage.
Solution
Fluxions offers a complete voice AI stack with both consumer and enterprise applications. The company's consumer product, vui, is an iPhone-based voice assistant that runs entirely on-device with a 300M parameter model, ensuring conversations never leave the phone and work offline. For businesses, Fluxions provides voice agents at a flat $0.10 per minute all-in rate covering speech recognition, reasoning, and voice synthesis, eliminating component stacking and bring-your-own-key complexity. The platform also offers a la carte transcription and speech services through a developer API, with the akro-v1 model providing transcription, speaker diarization, and non-speech event detection.
Target Audience
Primary customers include businesses needing voice agents for customer service and reception, developers building voice-enabled applications, and privacy-conscious consumers seeking a fully offline voice assistant on iPhone.
Features
- On-device voice assistant (vui) with 300M parameter model running entirely on iPhone GPU, requiring no server or account
- All-inclusive voice agent pricing at $0.10/min covering ASR, reasoning, and voice with no per-component billing
- akro-v1 speech-to-text model with speaker diarization, non-speech event detection (breathing, laughter, hesitation), and word-level timestamps
- Expressive text-to-speech with voice cloning and non-verbal cues at $10 per 1M characters (~$0.45 per hour of audio)
- WebSocket streaming support for reduced latency, skipping TLS/TCP handshakes across renders
- On-prem and private-cloud deployment options for air-gapped operation in security-sensitive environments
- Billing by the second with automatic hang-up on idle time to prevent dead-air charges