Skip to main content
EA

Etched.ai

Etched.ai develops Sohu, the world's first ASIC specifically designed for transformer models, enabling AI computations to be executed at least ten times faster and more cost-effectively than traditional GPUs. This technology allows for real-time processing of large-scale AI models, enhancing applications such as voice agents and content generation.

San Francisco, United StatesFounded 202210410K+ followers
Updated 14 days ago

Funding

$1.1B raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

+14

Founders

Product

Problem

General-purpose GPUs are inefficient and costly for running large-scale transformer models, limiting the feasibility of real-time AI applications. Existing hardware solutions struggle to keep pace with the increasing size and complexity of modern AI models, hindering advancements in areas like voice agents and content generation.

Solution

Etched.ai has developed Sohu, an application-specific integrated circuit (ASIC) designed specifically for transformer models. By etching the transformer architecture directly into silicon, Sohu accelerates AI computations, achieving significantly higher throughput and lower costs compared to traditional GPUs. This specialized hardware enables real-time processing of large language models, facilitating applications that were previously impractical due to computational constraints.

Target Audience

The primary target audience includes companies and researchers working on large language models, real-time AI applications, and other computationally intensive AI tasks.

Features

  • Custom ASIC architecture optimized for transformer model inference
  • Greater than 500,000 tokens/sec Llama 70B throughput
  • Runs AI models an order of magnitude faster and cheaper than GPUs
  • Supports real-time voice agents capable of ingesting thousands of words in milliseconds
  • Enables advanced coding applications with parallel tree search capabilities
  • Supports multicast speculative decoding for real-time content generation
  • Designed to run trillion-parameter models
  • Fully open-source software stack
  • 144 GB HBM3E memory per chip
  • Supports Mixture of Experts (MoE) and other transformer variants
This profile is AI-generated and may contain inaccuracies.