Modal provides a serverless, code‑first platform for building and running AI workloads entirely in Python. It launches GPU‑accelerated functions with sub‑second cold starts, automatically scales across multi‑cloud GPU pools, and includes built‑in observability, distributed storage, and gVisor‑based sandbox isolation. Users pay only for actual compute time, with a free tier and per‑GPU‑second billing.
Funding
Funding not disclosed



Founders
Product
Problem
Machine learning teams often face high latency from cold starts, complex multi‑cloud GPU provisioning, and heavyweight configuration files that slow iteration on inference, training, and batch pipelines. These friction points increase operational overhead and limit the ability to scale AI workloads quickly and securely.
Solution
Modal delivers a serverless, code‑first platform that lets developers define the entire AI stack in Python, eliminating YAML and manual provisioning. Its custom container runtime launches GPU‑accelerated functions with sub‑second cold starts and automatic scaling across a global pool of GPUs, so workloads expand to thousands of nodes on demand and shrink back to zero when idle. Integrated observability provides real‑time logs, metrics, and dashboards for every function, while a built‑in distributed storage layer enables fast model loading and data access. Secure sandboxes built on gVisor enforce isolation and support fine‑grained RBAC, meeting SOC 2 and HIPAA requirements. First‑party integrations let users mount cloud buckets, connect to MLOps tools, and stream telemetry without additional glue code. The platform supports end‑to‑end AI use cases—including inference, fine‑tuning, batch processing, notebooks, and sandboxed code execution—through a unified API and pricing model that charges only for actual compute time.
Target Audience
The platform is aimed at ML engineers, data scientists, and AI product teams—from fast‑moving startups to large enterprises—who need to deploy inference, training, or batch pipelines without managing underlying infrastructure. It also serves developers requiring secure, on‑demand code execution environments such as sandboxed LLM agents or custom AI services.
Features
- Python SDK that programs infrastructure as code, synchronizing environment and hardware requirements automatically.
- Serverless containers with sub‑second cold starts and GPU snapshotting for rapid model initialization.
- Elastic, multi‑cloud GPU scaling that provisions thousands of GPUs on demand without quotas or reservations.
- Unified observability dashboard offering live logs, metrics, and resource‑usage insights per function.
- Distributed storage layer optimized for high‑throughput model and dataset access across all workloads.
- Secure sandboxes using gVisor isolation, RBAC, and compliance certifications (SOC 2, HIPAA).
- Native integrations for cloud bucket mounts, MLOps pipelines, and telemetry services (Datadog, OpenTelemetry).
- Pay‑as‑you‑go pricing with a free tier of $30 per month compute and per‑GPU‑second billing for production workloads.