Skip to main content
BL

Baud Labs

Baud Labs builds a 1‑bit native silicon processor that runs binary‑weight neural networks for both training and inference, using a per‑MAC design that replaces traditional multipliers with simple multiplexers and adders.

Founded 2026210+ followers
Updated 2 months ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Current AI accelerators operate with 4–32‑bit data formats, leading to significant silicon area and energy waste because most neural network weights and activations can be represented with far fewer bits. No existing commercial silicon processes natively support 1‑bit operations, forcing designers to implement inefficient emulation layers.

Solution

Baud Labs has designed a 1‑bit native silicon processor that executes both training and inference for binary‑weight, low‑precision neural networks. The architecture replaces conventional multiply‑accumulate units with a per‑MAC design optimized for ternary weight (‑1, 0, +1) and 4‑bit activation formats, achieving higher token throughput on small models. The chip integrates with PyTorch, allowing developers to compile existing models without extensive code changes. Prototypes have been validated on FPGA, delivering 1,200 tokens per second, and the full design has been simulated in GlobalFoundries’ 12 nm process ahead of tape‑out, positioning the device for future silicon production and cluster‑scale deployment.

Target Audience

Primary customers are edge‑AI device manufacturers, semiconductor companies developing AI accelerators, and machine‑learning researchers seeking ultra‑low‑power, high‑throughput binary neural network hardware.

Features

  • 1‑bit native compute engine supporting W1A4 (1‑bit weights, 4‑bit activations) for binary neural networks
  • Per‑MAC silicon design that eliminates multipliers and carry‑save trees, using simple multiplexers and adders
  • Unified hardware path for both training and inference, reducing the need for separate accelerator stacks
  • PyTorch integration enabling direct model import and compilation to the custom ISA
  • FPGA demonstration achieving 1,200 tokens / s on a small model, confirming functional performance
  • Full‑chip simulation in a 12 nm GlobalFoundries process to validate timing, power, and area before tape‑out
This profile is AI-generated and may contain inaccuracies.