FuriosaAI develops the RNGD data center accelerator, utilizing a Tensor Contraction Processor architecture to enhance the efficiency of AI inference with a power profile of just 150W. This technology enables enterprises to deploy large language models and multimodal applications with low latency and high throughput, significantly reducing energy consumption and operational costs in data centers.
Funding
$194.3M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

Founders
Product
Problem
Existing data center accelerators for AI inference often consume excessive power, leading to high operational costs and hindering the widespread deployment of large language models (LLMs) and multimodal applications. Traditional architectures relying on fixed-size matrix multiplication instructions struggle to efficiently handle the tensor contraction operations fundamental to modern deep learning.
Solution
FuriosaAI's RNGD data center accelerator addresses these challenges with its Tensor Contraction Processor (TCP) architecture, designed for efficient tensor contraction operations. This approach enables high-performance LLM and multimodal deployment while maintaining a power profile of 180W. The TCP architecture maximizes parallelism and data reuse, providing flexibility and reconfigurability of compute and memory resources based on tensor shapes. The Furiosa SDK, including a model compressor, serving framework, and compiler, facilitates seamless deployment and optimization of LLMs on the RNGD accelerator.
Target Audience
The primary target audience includes enterprises and cloud service providers seeking to deploy LLMs and multimodal applications with high performance and energy efficiency.
Features
- Tensor Contraction Processor (TCP) architecture optimized for tensor contraction operations
- Support for BF16, FP8, INT8, and INT4 data types
- 512 TFLOPS (FP8) compute performance
- 48GB HBM3 memory with 1.5TB/s memory bandwidth
- 256MB SRAM with 384TB/s on-chip bandwidth
- PCIe P2P support for LLM
- Multiple-instance and virtualization capabilities
- Secure boot and model encryption
- Furiosa SDK with model compression, serving framework, runtime, compiler, profiler, debugger, and APIs
- Integration with PyTorch 2.x and Hugging Face Hub
- Support for containerization, SR-IOV, and Kubernetes