Skip to main content

Recogni

Recogni develops a multimodal AI inference system utilizing its proprietary Pareto AI Math to enhance performance while significantly reducing power consumption. This technology addresses the high costs and energy demands of generative AI models, enabling efficient and accurate processing for data centers.

San Jose, United States · HQ
Founded 20171127K+ followers
Updated 22 months ago

Funding

$176.1M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

Funding rounds are not available yet.

Founders

Product

Problem

Generative AI models demand significant computational resources, leading to high operational costs and substantial energy consumption for data centers. Existing inference solutions often struggle to balance performance, accuracy, and efficiency when deploying large language models (LLMs).

Solution

Recogni offers a multimodal AI inference system designed to accelerate generative AI models while drastically reducing power consumption. Their core innovation, Pareto AI Math, enables efficient and accurate processing, making GenAI more economical and sustainable. The system is built using a hardware and software co-design approach, optimizing the entire stack for performance. By utilizing the latest 3nm TSMC technology node and high-bandwidth memory (HBM3e), Recogni's solution achieves high throughput and low latency, even with large models.

Target Audience

Recogni's primary customers are hyperscalers, cloud service providers, and enterprises that require efficient and accurate generative AI inference for data centers.

Features

  • Pareto AI Math: A logarithmic math number system that maintains high accuracy (99.9%) while consuming significantly less power (4x less than standard math).
  • Hardware-Software Co-design: Optimizes the entire system for performance, balancing compute-to-memory bandwidth and chip-to-chip communication.
  • 3nm TSMC Technology Node: Ensures best-in-class energy efficiency and cost.
  • High Bandwidth Memory (HBM3e): Maximizes output speeds for autoregressive models.
  • Tensor Parallelism (TP > 100): Enables parallelizing AI models across chips for faster processing and larger model support.
  • Rapid Compilation: Compiles models from PyTorch to executable files in under 10 minutes, even for large models like Llama 405b.
  • Pareto SDK: Allows developers to deploy models with high accuracy and efficiency.
This profile is AI-generated and may contain inaccuracies.