Together AI provides a cloud platform that offers serverless OpenAI‑compatible inference APIs for over 200 open‑source models, accelerated up to 4× by its ATLAS runtime. Users can provision on‑demand or reserved NVIDIA GPU clusters for fine‑tuning and batch inference, with per‑token or hourly usage pricing and enterprise‑grade security.
Funding
$305M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

PVFounders
Product
Problem
AI developers and enterprises face steep infrastructure costs and operational complexity when scaling training, fine‑tuning, and inference for large language and multimodal models. Limited access to high‑performance GPU clusters and unpredictable pricing hinder rapid product iteration and cost‑effective deployment.
Solution
Together AI delivers an AI‑native cloud platform that abstracts the hardware layer while maximizing performance and price efficiency. A serverless Inference API provides OpenAI‑compatible endpoints for over 200 open‑source models, backed by the ATLAS runtime‑learning accelerator that yields up to 4× faster LLM inference. Instant Clusters let users provision self‑service NVIDIA GPUs (GB200, GB300, H100, H200) on demand, and Reserved Clusters offer dedicated capacity with expert support. Fine‑tuning is available via LoRA adapters or full‑model training, including Direct Preference Optimization for human‑aligned outputs. The Batch Inference API processes billions of tokens at roughly 50 % lower cost than competing services. All workloads run on a globally distributed fleet across 25+ data‑center locations, with Slurm orchestration, end‑to‑end encryption, and role‑based access controls to meet enterprise security standards.
Target Audience
Primary users are AI developers, startups, and enterprise engineering teams building generative‑AI products, as well as research groups that require scalable training and evaluation pipelines.
Features
- Serverless Inference API with per‑token pricing and OpenAI‑compatible request format
- ATLAS speculator system that accelerates LLM inference up to 4× while reducing latency
- Instant Clusters: self‑service GPU provisioning (NVIDIA GB200, GB300, H100, H200) with hourly billing
- Reserved/Dedicated Endpoints: single‑tenant GPUs, autoscaling, 99.9 % SLA, custom hardware support
- Fine‑tuning platform supporting LoRA, full‑model training, and Direct Preference Optimization (DPO)
- Batch Inference API for high‑throughput token processing at ~50 % cost reduction
- Model library of 200+ open‑source text, vision, audio, and video models with OpenAI‑compatible APIs
- Global infrastructure (25+ cities) with Slurm cluster management, HIPAA‑grade encryption, and RBAC security