Skip to main content
D

dstack

dstack provides a unified control plane for GPU provisioning and orchestration across cloud, Kubernetes, and on-prem environments for ML teams. This platform streamlines development, training, and inference workflows while significantly reducing infrastructure costs. It offers an open stack supporting any hardware, open-source tools, and custom code for managing AI workloads.

Munich, GermanyFounded 2022141K+ followers
Updated 20 months ago

Funding

Funding not disclosed

RF
Funding rounds are not available yet.

Founders

Product

Problem

Managing AI workloads across diverse environments, including cloud and on-premises infrastructure, presents significant complexity for machine learning teams. Existing solutions often require extensive manual configuration and lack a unified interface for managing resources and deployments across different hardware accelerators.

Solution

dstack offers an open-source orchestration engine designed to simplify the deployment and management of AI workloads. As a lightweight alternative to Kubernetes and Slurm, dstack provides a unified interface for provisioning resources, scheduling jobs, and deploying models across various cloud providers, on-prem servers, and hardware accelerators like NVIDIA, AMD, and TPUs. The platform supports containerized development environments, task scheduling for training and fine-tuning, and service deployment for scalable model endpoints. By abstracting away the complexities of infrastructure management, dstack enables ML engineers to focus on model development and experimentation.

Target Audience

dstack is designed for machine learning teams and AI researchers seeking to streamline AI development, reduce GPU costs, and simplify infrastructure management across diverse environments.

Features

  • Supports provisioning and management of cloud and on-prem clusters through a unified interface.
  • Enables the creation of containerized development environments with access via SSH or desktop IDEs like VS Code and Cursor.
  • Facilitates scheduling of AI and data workloads on optimized clusters or individual instances.
  • Allows deployment of models as secure, auto-scaling, OpenAI-compatible endpoints.
  • Integrates with various cloud providers, including AWS, Azure, GCP, Lambda, RunPod, and Vast.ai.
  • Supports a range of hardware accelerators, including NVIDIA, AMD, TPU, Intel Gaudi, and Tenstorrent.
  • Provides a CLI and API for managing infrastructure and deployments.
  • Offers built-in UI for monitoring GPU metrics and exporting metrics to Prometheus.
This profile is AI-generated and may contain inaccuracies.