Pebble provides an AI‑driven orchestration platform that monitors GPU utilization across on‑premise data centers and public clouds, automatically scheduling, pausing, or migrating workloads to reduce cost, energy use, and carbon emissions. It integrates via a unified API with Kubernetes, Slurm, OpenStack, and cloud schedulers, delivering real‑time analytics and compliance‑ready reporting while preserving performance SLAs.
Funding
Funding not disclosed
Founders
Product
Problem
Data centers and cloud environments often run GPUs at low utilization, leading to unnecessary compute spend, elevated energy consumption, and a larger carbon footprint. Organizations lack real‑time visibility and automated controls to reallocate that idle capacity without risking performance. This results in fragmented cost management and compliance challenges for large‑scale AI workloads.
Solution
Pebble delivers an AI‑driven orchestration platform that continuously monitors GPU utilization across on‑premise data centers and public clouds. By applying predictive scheduling and workload flexing, the system automatically pauses, resizes, or migrates jobs to periods of lower cost or higher renewable energy availability. The platform exposes a unified API that integrates with existing Kubernetes, Slurm, or cloud‑native schedulers, enabling seamless adoption without code changes. Real‑time analytics surface compute, energy, and carbon metrics, while compliance‑ready reports satisfy finance and regulatory audits. All actions run in a secure, encrypted environment, ensuring that performance SLAs are maintained while unlocking latent compute capacity.
Target Audience
The primary customers are enterprise data‑center operators, cloud‑native AI/ML teams, and IT finance groups that manage large GPU fleets in sectors such as healthcare, financial services, and high‑performance computing.
Features
- Lightweight telemetry agent that instruments GPU usage and power draw with sub‑second granularity.
- AI agents that perform demand‑aware scheduling, automatically scaling workloads up or down based on cost, carbon intensity, and performance targets.
- Predictive workload placement engine that shifts batch jobs to off‑peak windows or to regions with greener energy mixes.
- Unified REST/GraphQL API for integration with Kubernetes, Slurm, OpenStack, and major cloud provider schedulers.
- One‑click dashboard delivering real‑time compute spend, energy consumption, and carbon emissions visualizations.
- Automated compliance reporting module that generates audit‑ready PDFs and CSV exports for finance and ESG teams.
- Role‑based access control and end‑to‑end encryption to protect sensitive workload and billing data.