EXXA provides a cost-efficient asynchronous LLM inference service that utilizes a custom scheduler to aggregate unused compute resources across multiple data centers, enabling batch processing of requests. This approach reduces costs to as low as $0.30 per million tokens while optimizing energy consumption, making it ideal for applications that can tolerate some delay.
Funding
Funding not disclosed

Founders
Product
Problem
Large language model (LLM) inference can be computationally expensive, leading to high costs for applications that require processing substantial volumes of text. Traditional systems often fail to efficiently utilize available compute resources, resulting in wasted capacity and increased energy consumption.
Solution
EXXA provides a cost-optimized asynchronous LLM inference service by aggregating unused compute capacity across multiple data centers. Their custom scheduler and orchestrator efficiently capture intermittent compute windows, enabling batch processing of requests at significantly reduced costs. By leveraging underutilized resources, EXXA offers a more sustainable and affordable solution for LLM inference, particularly for applications that can tolerate a delay in processing. The platform supports open-source models like Llama 3, providing high-quality output at a fraction of the cost of traditional inference services.
Target Audience
EXXA primarily targets businesses and developers who require cost-effective LLM inference for large-scale tasks such as LLM evaluation, contextual retrieval, classification, translation, parsing, and synthesis.
Features
- Asynchronous batch processing with a typical turnaround time of under 24 hours
- Custom scheduler and orchestrator to maximize the use of intermittent and low-cost compute resources
- Predictive inference optimizer to select optimal settings for each payload, including batch size and context size
- Specialized inference engine optimized for batch API, featuring persistent KV cache and cross-platform/cross-GPU compatibility
- Support for Llama 3.1 70B and 8B models, with plans to add more models
- Detailed energy consumption data provided for each request via API
- Option to offset carbon footprint by purchasing certified carbon credits through the platform
- Batch cancellation feature, charging only for completed work