OLIX develops specialized hardware infrastructure that treats AI token generation as a production line, separating token‑creation steps onto dedicated chips to improve throughput and reduce cost. By moving away from general‑purpose processors, its token factory architecture enables frontier AI models to run more interactively and efficiently at scale. The platform targets datacenters seeking high‑performance, low‑latency AI workloads while lowering energy consumption.
Funding
$532M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.





A+6Founders
Product
Problem
Current AI inference hardware runs all stages of token generation on a single general‑purpose chip, causing high energy use, limited throughput, and expensive, scarce tokens. The lack of specialized processing and efficient inter‑chip communication prevents scaling to the high interactivity and low cost required for frontier AI models.
Solution
Olix designs a purpose‑built “token factory” that decomposes the token generation pipeline into many specialized chips, each handling a single stage of the model. By unrolling models across a rack of these chips and linking them with a novel slow‑and‑wide optical interconnect, data moves between stages with ultra‑low latency and energy cost. A deterministic compiler schedules workloads across the rack, while each chip retains a flexible compute fabric to adapt to evolving model architectures. The first product, the DX‑1 decode accelerator, stores models in on‑chip SRAM and delivers over 10,000 tokens per second per user for 100‑billion‑parameter models, with performance that scales to trillion‑parameter models without relying on scarce high‑bandwidth memory. This architecture aims to dramatically lower the cost per token and enable far more powerful AI inference deployments.
Target Audience
Primary customers are large‑scale datacenter operators and AI inference service providers that require high‑throughput, low‑cost token generation for frontier language and vision models.
Features
- Specialized inference chips that each execute a single stage of the model, eliminating the inefficiencies of general‑purpose processors
- Rack‑scale optical interconnect using light‑based links for ultra‑low latency, low‑energy data movement between chips
- Deterministic compiler that schedules and orchestrates workloads across multiple racks for predictable performance
- DX‑1 decode accelerator with on‑chip SRAM model storage, achieving >10,000 tokens / s / user at optimal power efficiency
- Flexible compute fabric on each chip that can be reprogrammed for new model architectures without hardware redesign
- No reliance on advanced packaging or high‑bandwidth memory, mitigating supply‑chain constraints