Firmus provides a modular AI infrastructure platform that combines liquid‑cooled, high‑density GPU clusters with a public AI cloud offering on‑demand and reserved GPU instances. Its AI FactoryOS orchestration layer automates workload placement, power and cooling management, and delivers real‑time telemetry, enabling AI labs and enterprise teams to train large models efficiently and predictably.
Funding
$10.7B raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.



BECHV+2Founders
Product
Problem
Training and deploying large AI models requires massive GPU compute, but traditional data centers are energy‑intensive, costly, and often lack the specialized networking and cooling needed for high‑density workloads. This limits scalability, raises operational expenses, and creates barriers for organizations seeking sustainable, high‑performance AI infrastructure.
Solution
Firmus delivers a vertically integrated AI infrastructure platform that combines modular “AI Factories” with a public AI cloud, dedicated bare‑metal GPU clusters, high‑throughput RDMA storage, and managed cloud services. Each AI Factory is a liquid‑cooled, multi‑petaflop HyperCube module designed for maximal compute density while minimizing power and water usage. The Firmus AI Cloud offers on‑demand and reserved GPU instances (H200, H100, L40S, A100) with Slurm scheduling, encrypted InfiniBand networking, and pre‑configured CUDA/PyTorch/TensorFlow environments, enabling rapid scaling from experimentation to production. A unified orchestration layer (AI FactoryOS™) provides silicon‑aware workload automation, real‑time telemetry, and grid‑responsive power management, ensuring consistent performance per watt across all layers of the stack.
Target Audience
Primary customers are AI research labs, enterprise ML teams, and hyperscale cloud providers that require high‑performance, energy‑efficient GPU compute for training large language models, generative AI, and HPC workloads.
Features
- Modular HyperCube AI Factories with liquid cooling achieve sub‑1.0 PUE and up to 99 % water reduction versus conventional data centers
- High‑density GPU configurations (up to 8 × NVIDIA H200 per node) with NVLink/NVSwitch and dual 200–800 Gbps InfiniBand for low‑latency distributed training
- AI FactoryOS™ orchestrates compute, power, cooling, and grid interaction, delivering real‑time telemetry and automated workload placement
- S3‑compatible, WEKA‑based NVMe parallel file system with RDMA acceleration for petabyte‑scale dataset ingest and model checkpointing
- Managed Slurm scheduler, pre‑installed CUDA stacks, and NVIDIA NIM inference APIs simplify environment setup and scaling
- ISO 27001 and SOC‑2 compliant services with end‑to‑end encryption, hybrid‑cloud connectivity, and enterprise‑grade observability
- Transparent, per‑GPU pricing model with on‑demand and reserved options, eliminating hidden costs and enabling predictable budgeting