Cap Sen AI
Cap Sen AI develops activation‑level alignment techniques for large language models, shaping the internal activations during inference to improve decision quality and reduce deceptive behavior. Their research includes an emotional memory system that enhances threat‑safety judgments and a self‑chosen deception direction that detects and corrects LLM deception, demonstrated on Gemma 3 1B‑IT models.
- Artificial Intelligence
- Developer Tools
Funding
Founders
Product
Problem
Current large language models are aligned only at the output level, leaving their internal activations unchecked. This can lead to unreliable reasoning, deceptive responses, and suboptimal decision-making, limiting trust and economic value in human‑AI interactions.
Solution
Cap Sen AI focuses on alignment at the activation level by reading and shaping the internal representations of LLMs during inference. Their research identifies emotion‑exclusive sparse autoencoder features and uses emotion‑specific echo vectors to enhance decision quality, raising performance from 52% to 80% in benchmark tasks. They also introduce a self‑chosen deception direction that detects and corrects deceptive model outputs, improving reliability. By targeting these sparse features, Cap Sen AI aims to create more trustworthy and economically beneficial interactions between humans and AI systems.
Target Audience
Primary customers are enterprises and developers deploying large language models who require higher reliability, reduced deception risk, and improved decision-making in AI‑driven applications.
Features
- Identification of 310 emotion‑exclusive sparse autoencoder features in Gemma 3 1B‑IT with psychologically valid geometry
- Construction of emotion‑specific echo vectors that steepen threat‑safety gradients and improve decision quality
- Self‑chosen deception direction that detects and corrects deceptive outputs during inference
- Scalable methodology that shows increasing effectiveness on larger models (4B, 12B, 27B)
- Activation‑level alignment approach that operates without altering model outputs directly