Skip to main content
F

FuriosaAI

FuriosaAI develops the RNGD data center accelerator, utilizing a Tensor Contraction Processor architecture to enhance the efficiency of AI inference with a power profile of just 150W. This technology enables enterprises to deploy large language models and multimodal applications with low latency and high throughput, significantly reducing energy consumption and operational costs in data centers.

Seoul, South KoreaFounded 20171257K+ followers
Updated 4 months ago

Funding

$194.3M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

Funding rounds are not available yet.

Founders

Product

Problem

Existing data center accelerators for AI inference often consume excessive power, leading to high operational costs and hindering the widespread deployment of large language models (LLMs) and multimodal applications. Traditional architectures relying on fixed-size matrix multiplication instructions struggle to efficiently handle the tensor contraction operations fundamental to modern deep learning.

Solution

FuriosaAI's RNGD data center accelerator addresses these challenges with its Tensor Contraction Processor (TCP) architecture, designed for efficient tensor contraction operations. This approach enables high-performance LLM and multimodal deployment while maintaining a power profile of 180W. The TCP architecture maximizes parallelism and data reuse, providing flexibility and reconfigurability of compute and memory resources based on tensor shapes. The Furiosa SDK, including a model compressor, serving framework, and compiler, facilitates seamless deployment and optimization of LLMs on the RNGD accelerator.

Target Audience

The primary target audience includes enterprises and cloud service providers seeking to deploy LLMs and multimodal applications with high performance and energy efficiency.

Features

  • Tensor Contraction Processor (TCP) architecture optimized for tensor contraction operations
  • Support for BF16, FP8, INT8, and INT4 data types
  • 512 TFLOPS (FP8) compute performance
  • 48GB HBM3 memory with 1.5TB/s memory bandwidth
  • 256MB SRAM with 384TB/s on-chip bandwidth
  • PCIe P2P support for LLM
  • Multiple-instance and virtualization capabilities
  • Secure boot and model encryption
  • Furiosa SDK with model compression, serving framework, runtime, compiler, profiler, debugger, and APIs
  • Integration with PyTorch 2.x and Hugging Face Hub
  • Support for containerization, SR-IOV, and Kubernetes
This profile is AI-generated and may contain inaccuracies.