Nuance Labs builds a multimodal human foundation model that can perceive and generate emotions in real time across speech, facial expression, body language, and text. The low‑latency system lets developers embed emotionally aware AI into virtual avatars and conversational agents, enabling more natural, empathetic interactions.
Funding
$10M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.
1OFounders
Product
Problem
Current AI systems lack the ability to perceive and convey human emotions through voice, facial expressions, body language, and text, resulting in interactions that feel mechanical and fail to achieve natural, empathetic communication.
Solution
Nuance Labs is developing a human foundation model that learns to interpret and generate emotional cues in real time across multimodal channels—speech, facial expression, body posture, and written language. By training auto‑regressive transformers to predict human actions frame by frame, the model captures subtle affective signals such as a raised eyebrow, a tonal shift, or a hesitant pause. The system can both read a user's emotional state and produce appropriate, emotionally aligned responses, enabling AI agents to engage with users in a manner that feels genuinely human. This capability is delivered through a low‑latency, consumer‑grade architecture designed for integration into virtual avatars, conversational agents, and other interactive AI products.
Target Audience
Primary customers are developers and product teams building conversational agents, virtual avatars, or interactive AI experiences that require natural, emotionally aware user interactions.
Features
- Multimodal foundation model that simultaneously processes audio, video, and text to infer emotional states
- Real‑time inference pipeline optimized for ultra‑low latency suitable for interactive applications
- Generative component that produces emotionally consistent facial expressions, vocal prosody, and body language
- Auto‑regressive transformer training approach that predicts human behavior frame by frame, enhancing nuance capture
- Scalable architecture designed for deployment in consumer‑grade AI avatars and conversational agents