Skip to main content
AC

Apex Compute

Apex Compute builds a purpose‑designed AI accelerator that merges systolic arrays with vector processing units to deliver up to 20× higher computational efficiency than comparable NVIDIA Jetson chips. The hardware is optimized for GPT‑style workloads—matrix multiplication, quantization, softmax, and data movement—and is paired with a co‑designed software stack that schedules operations to achieve over 90% utilization, low latency, and minimal power draw. This compact, high‑utilization compute module targets edge AI OEMs for drones, autonomous vehicles, robotics, and other real‑time LLM inference applications.

Los Altos, United States7700+ followers
Updated 2 months ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Edge AI applications such as drones, autonomous vehicles, and robotics require high‑performance compute that can run large language model (LLM) workloads with low latency and power consumption, but existing accelerators are either too power‑hungry, have large silicon footprints, or achieve low hardware utilization on GPT‑style operations.

Solution

Apex Compute delivers a purpose‑built AI accelerator that combines systolic arrays with vector processing units in a unified engine, achieving up to 20× higher computational efficiency than comparable NVIDIA Jetson solutions. The architecture is derived from a detailed analysis of GPT attention kernels, optimizing matrix‑matrix multiplication, quantization, softmax, and data movement to run at >90 % utilization. A co‑designed software stack automatically maps compute graphs onto the hardware, scheduling operators to eliminate idle cycles and reduce memory overhead. The result is a compact, low‑latency, power‑efficient compute module suitable for deployment in edge devices that need real‑time LLM inference.

Target Audience

Primary customers are OEMs and system integrators developing edge AI platforms for drones, autonomous vehicles, robotics, and other low‑latency, high‑efficiency compute applications.

Features

  • Unified systolic‑array and vector‑processing engine tailored for GPT‑style workloads
  • Optimized hardware blocks for matrix multiplication, quantization, softmax, and element‑wise operations
  • Advanced scheduling software that maps compute graphs to achieve >90 % hardware utilization
  • Modular instruction set allowing flexible implementation of diverse AI workloads
  • Extremely low power consumption with no idle cycles, enabling high performance in constrained edge environments
  • Small silicon footprint suitable for integration into drones, autonomous vehicles, robotics, and other off‑cloud systems
This profile is AI-generated and may contain inaccuracies.