MemVerge offers a Memory-Converged Infrastructure system utilizing proprietary DMO technology to optimize GPU workloads for Generative AI applications. Their Memory Machine™ suite enables organizations to reduce cloud costs by up to 90% through transparent checkpointing and efficient GPU scheduling, while enhancing GPU utilization and accelerating AI workloads.
Funding
Funding not disclosed




Founders
Product
Problem
Training and running Generative AI models on GPUs in the cloud is expensive, especially when using on-demand instances. Interruptions and the need for frequent restarts can lead to wasted compute time and increased costs.
Solution
MemVerge offers the Memory Machine™ suite, a Memory-Converged Infrastructure system designed to optimize GPU workloads for Generative AI applications. By leveraging transparent checkpointing, GPU scheduling, memory tiering, and memory sharing, Memory Machine™ enables organizations to utilize spot instances more effectively, reducing cloud costs. The solution also facilitates GPU-as-a-Service, increasing GPU utilization, and accelerates AI workloads with CXL memory expansion.
Target Audience
The primary target audience includes scientific researchers, AI practitioners, IT architects, and organizations in genomics, financial services, and EDA industries who are looking to optimize GPU utilization and reduce cloud costs associated with Generative AI workloads.
Features
- Transparent application checkpointing and restore capabilities for AWS Batch environments.
- SpotSurfer feature expands the capability of AWS Batch to run stateful workloads safely on Spot instances.
- GPU scheduling and memory tiering to optimize resource allocation.
- CXL memory expansion and sharing for accelerated AI workloads.
- Support for running Nextflow pipelines and next-generation sequencing (NGS) on EC2 Spot instances for genomics.
- GPU-as-a-Service functionality to increase GPU utilization.