Skip to main content
C

Cedana

Cedana provides compute orchestration software that automatically migrates GPU+CPU workloads across on-premise and multi-cloud environments to increase throughput by 2-10x. The platform seamlessly integrates with Kubernetes and SLURM to manage resource scheduling based on price, performance, and SLAs. It ensures high reliability through system-level checkpointing that resumes workloads transparently after hardware failures.

East New York, United StatesFounded 202391K+ followers
Updated 20 months ago

Funding

$500K raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

Funding rounds are not available yet.

Founders

Product

Problem

Stateful CPU and GPU workloads often lead to resource fragmentation and downtime during migrations or hardware failures. Traditional methods of managing these workloads can be inefficient, resulting in increased compute costs and potential data loss.

Solution

Cedana provides a Save, Migrate, and Resume (SMR) system that leverages Linux Kernel insights to checkpoint and restore stateful CPU and GPU workloads across instances. This technology minimizes downtime by enabling live migration of containerized workloads, ensuring workload continuity during hardware failures. By automatically suspending and resuming workloads based on activity, Cedana facilitates fine-grained bin packing of containers, optimizing resource utilization and reducing idle compute. The system integrates via a REST API, allowing users to checkpoint application states, transfer them to new instances, and resume operations without code modifications.

Target Audience

Cedana targets users of managed Kubernetes, Platform-as-a-Service (PaaS), GPU cloud inference, and those running AI workloads, databases, analytics, and web servers.

Features

  • Live migration of containerized, stateful CPU and GPU workloads
  • Automatic suspension and resumption of workloads based on activity
  • Fine-grained bin packing of containers for optimized resource utilization
  • Checkpointing and restoration of complete workload state, including process, filesystem, network connections, and memory
  • REST API for easy integration without code modifications
  • Automated workload resumption on new instances after hardware or OOM failures
  • Policy-based automation for workload-level Service Level Agreement (SLA) enforcement
This profile is AI-generated and may contain inaccuracies.