ExoHelix builds rack‑scale AI systems that maximize GPU utilization while cutting power use and total cost of ownership. Their PlatformX solution connects a single head node to 90+ accelerators with sub‑microsecond fabric latency, presenting every GPU as a local device so standard CUDA, ROCm, and PyTorch workloads run unchanged.
Funding
Funding not disclosed
Founders
Product
Problem
AI workloads often require large numbers of GPUs spread across multiple servers, leading to fragmented resource utilization, high power consumption, and complex multi-node software configurations.
Solution
ExoHelix’s PlatformX provides a rack‑scale AI system that consolidates 90+ accelerators under a single head node, presenting each GPU as a locally attached device. The sub‑microsecond fabric interconnect eliminates the need for multi‑node orchestration and NCCL tuning, allowing standard CUDA, ROCm, and PyTorch code to run unchanged. By keeping all GPUs active and visible, the platform achieves higher utilization, up to 40% per‑inference energy savings, and reduces total cost of ownership by roughly 50% compared to conventional multi‑node deployments. The chassis accepts any PCIe accelerator, enabling mixed‑vendor configurations and future‑proof upgrades without lock‑in.
Target Audience
Primary customers are enterprises and research institutions that run large‑scale AI inference workloads and need high GPU utilization with lower power and cost, such as cloud providers, hyperscale data centers, and AI‑focused labs.
Features
- Sub‑microsecond latency fabric connecting a single head node to 90+ GPUs
- Software‑transparent operation; no code changes required for CUDA, ROCm, or PyTorch workloads
- PCIe‑at‑memory‑speed interconnect delivering full bandwidth to each accelerator
- Vendor‑agnostic chassis supporting any PCIe accelerator card and mixed‑vendor racks
- Native GPU‑to‑GPU interconnect via OAM and HGX baseboards for efficient data movement
- Consolidated rack design that eliminates multi‑node complexity and NCCL tuning