Positron develops purpose-built hardware systems specifically designed to accelerate Transformer Model Inference. These systems offer superior performance per dollar and performance per watt compared to existing solutions like NVIDIA. The platform seamlessly supports any trained HuggingFace Transformers Library model for efficient deployment via an OpenAI API-compliant endpoint.
Funding
$291.6M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.






+10Founders
Product
Problem
The increasing demand for large language model (LLM) inference necessitates high-performance computing solutions that can handle complex models efficiently, especially in power-constrained environments. Existing GPU-based systems often suffer from bottlenecks related to memory, processing, and high power consumption, leading to increased operational costs and limited scalability.
Solution
Positron AI offers a transformer inference server, Atlas, designed to deliver optimized performance and cost-efficiency for AI model deployment. By utilizing a hardware and software co-design approach, Atlas achieves superior performance per watt and per dollar compared to traditional GPU-based systems. The platform supports seamless integration with Hugging Face models, allowing users to deploy any trained transformer model directly onto the hardware. Positron also provides a managed inference service, Testflight, for remote evaluation, enabling efficient scaling and reduced operational expenses.
Target Audience
The primary target audience includes enterprises, research teams, and service providers seeking high-performance, cost-effective, and energy-efficient solutions for deploying and scaling large language models.
Features
- High performance per watt, exceeding that of GPU-based systems by up to 4x
- Lower cost per token compared to Nvidia DGX-H100 systems, up to 2.5x
- Seamless integration with Hugging Face Transformers Library models
- OpenAI API-compliant endpoint for easy integration with client applications
- Model Manager for uploading and linking trained model files (.pt or .safetensors)
- Managed Transformer Inference (Testflight) for remote access and evaluation
- Increased model density for power-constrained racks
- Hardware and software explicitly designed for generative and large language models (LLMs)