Lumai provides optical AI inference servers that replace traditional silicon accelerators with a 3D optical engine, delivering up to 100 TOPS per watt and enabling real‑time inference of billion‑parameter LLMs such as Llama 8B and 70B. The rack‑mountable, air‑cooled system integrates with existing data‑center infrastructure and uses a three‑tier memory architecture to offer high‑capacity KV cache access and high data throughput for hyperscale workloads.
Funding
Funding not disclosed

Founders
Product
Problem
Current datacenter AI inference is constrained by the power consumption, cooling requirements, and limited throughput of conventional digital accelerators, making it difficult to run large language models in real time at scale.
Solution
Lumai delivers optical AI inference servers that replace traditional silicon‑based compute with three‑dimensional optical parallelism. By performing matrix multiplication in light, the Iris Nova platform achieves up to 100 TOPS per watt and can run billion‑parameter LLMs such as Llama 8B and 70B in real time. The system integrates with existing rack infrastructure, uses air cooling only, and presents a standard digital interface for seamless deployment in hyperscale data centers. Hardware‑aware quantization and a three‑tier memory architecture provide high‑capacity KV cache access, eliminating data conversion bottlenecks and enabling sustained token throughput within a 10 kW power envelope.
Target Audience
Primary customers are hyperscale cloud providers, large‑scale AI service operators, and enterprise data‑center teams that require high‑throughput, power‑efficient inference for large language models.
Features
- 3D optical engine performing 2048×2048 matrix multiplications at INT4/INT8‑equivalent precision
- Up to 100 TOPS/W energy efficiency, delivering 1 ExaOPS in a 10 kW deployment
- Supports real‑time inference of billion‑parameter LLMs (e.g., Llama 8B, 70B, Llama 3)
- Air‑cooled, rack‑mountable design compatible with standard data‑center infrastructure
- Three‑tier memory architecture optimized for high‑capacity KV cache access
- Hardware‑aware quantization schemes tailored to optical constraints for robust inference
- Seamless integration with existing software stacks via a digital processor and standard interfaces