Skip to main content
V

VSORA

The startup manufactures semiconductor chips with a multicore DSP architecture that accelerates the design of complex integrated circuits for mobile and network infrastructure. By eliminating the need for DSP coprocessors, these chips enable chipmakers to efficiently develop next-generation digital communication systems, including fifth-generation technologies.

Meudon, FranceFounded 201523700+ followers
Updated 24 days ago

Funding

$50.2M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

ALEIFF+2

Founders

Product

Problem

Meeting the computational demands of modern AI, especially for inference in data centers and autonomous systems, requires high throughput and low latency while minimizing power consumption and cost. Existing multi-chip solutions often suffer from high latency, inefficient silicon utilization, and excessive power draw, hindering performance and sustainability.

Solution

VSORA provides high-performance, CUDA-free AI inference chips designed to address the challenges of deploying large language models (LLMs) and other AI workloads at scale. Their architecture features a tightly coupled memory (TCM) system that minimizes data movement and reduces access latency, enabling efficient arithmetic utilization. The VSORA processor utilizes reconfigurable compute tiles that adapt to different workloads, enhancing efficiency and performance. The company's flagship product, Jotunn8, delivers high throughput and ultra-low latency, making it suitable for real-time applications such as chatbots, fraud detection, and search. VSORA's solutions aim to provide a new foundation for AI at scale, offering performance, cost-efficiency, and sustainability.

Target Audience

VSORA's primary customers are data center operators and companies building servers for AI inference, as well as those developing autonomous driving systems, robotics, and edge AI computing solutions.

Features

  • Ultra-low latency for real-time applications
  • High throughput for high-demand services
  • Cost-efficient inference at massive scale
  • Power-efficient performance per watt
  • Tightly coupled memory (TCM) system to minimize data movement and reduce access latency
  • Reconfigurable compute tiles that adapt to different workloads
  • Support for large language models (LLMs) like Llama3-40B and GPT-4
  • CUDA-free programmability
  • High memory capacity (HBM: 288GB) and throughput (HBM: 8 TB/s)
  • Tensor core (dense) performance: FP16: 800 Tflops, FP8: 3200 Tflops
  • General Purpose performance: FP32: 25 Tflops, FP16: 50 Tflops, FP8: 100 Tflops
This profile is AI-generated and may contain inaccuracies.