Skip to main content
E

Exxa

EXXA provides a cost-efficient asynchronous LLM inference service that utilizes a custom scheduler to aggregate unused compute resources across multiple data centers, enabling batch processing of requests. This approach reduces costs to as low as $0.30 per million tokens while optimizing energy consumption, making it ideal for applications that can tolerate some delay.

Founded 20233300+ followers
Updated 4 months ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Product

Problem

Large language model (LLM) inference can be computationally expensive, leading to high costs for applications that require processing substantial volumes of text. Traditional systems often fail to efficiently utilize available compute resources, resulting in wasted capacity and increased energy consumption.

Solution

EXXA provides a cost-optimized asynchronous LLM inference service by aggregating unused compute capacity across multiple data centers. Their custom scheduler and orchestrator efficiently capture intermittent compute windows, enabling batch processing of requests at significantly reduced costs. By leveraging underutilized resources, EXXA offers a more sustainable and affordable solution for LLM inference, particularly for applications that can tolerate a delay in processing. The platform supports open-source models like Llama 3, providing high-quality output at a fraction of the cost of traditional inference services.

Target Audience

EXXA primarily targets businesses and developers who require cost-effective LLM inference for large-scale tasks such as LLM evaluation, contextual retrieval, classification, translation, parsing, and synthesis.

Features

  • Asynchronous batch processing with a typical turnaround time of under 24 hours
  • Custom scheduler and orchestrator to maximize the use of intermittent and low-cost compute resources
  • Predictive inference optimizer to select optimal settings for each payload, including batch size and context size
  • Specialized inference engine optimized for batch API, featuring persistent KV cache and cross-platform/cross-GPU compatibility
  • Support for Llama 3.1 70B and 8B models, with plans to add more models
  • Detailed energy consumption data provided for each request via API
  • Option to offset carbon footprint by purchasing certified carbon credits through the platform
  • Batch cancellation feature, charging only for completed work
This profile is AI-generated and may contain inaccuracies.