Makora provides an AI‑driven platform that automatically generates and optimizes GPU kernels for CUDA and AMD architectures, delivering code in under a minute and continuously tuning performance as hardware or workloads change. The platform abstracts hardware differences, enabling the same kernel to run on NVIDIA, AMD, and major cloud GPUs, and integrates with PyTorch, vLLM, and SGLang via SDKs or a pip‑installable CLI.
Funding
Funding not disclosed
Founders
Product
Problem
Developing high‑performance GPU kernels requires specialized expertise and extensive manual tuning, often taking weeks to achieve optimal throughput on each target architecture. The process is fragmented across different hardware vendors and cloud environments, leading to increased development cost and delayed time‑to‑market for AI workloads.
Solution
Makora delivers an end‑to‑end AI‑driven platform that automates the creation, optimization, and deployment of GPU code. Its generative engine produces CUDA or AMD kernels in under a minute, while a continuous tuning service refines parameters to maintain peak performance as hardware or workloads evolve. The platform abstracts hardware differences, enabling the same kernel to run on NVIDIA, AMD, and major cloud providers without code rewrites. Seamless integration with PyTorch, vLLM, and SGLang lets developers invoke Makora directly from existing pipelines or via a CLI tool installed with pip. By offloading kernel engineering to Makora, teams reduce development cycles from weeks to hours and lower compute costs through sustained performance gains.
Target Audience
The primary customers are machine‑learning engineers, data‑science teams, and performance engineers at enterprises and cloud AI service providers who need to accelerate GPU workloads while minimizing manual tuning effort.
Features
- AI‑powered kernel generator (MakoraGenerate) that writes high‑throughput GPU code in under 60 seconds for CUDA and AMD architectures
- Automated parameter tuning and inference engine configuration (MakoraOptimize) that continuously adapts kernels to hardware and workload changes
- Hardware‑agnostic deployment across NVIDIA, AMD, AWS, GCP, and Oracle GPUs without source‑code modifications
- Native SDK integrations for PyTorch, vLLM, and SGLang, enabling one‑line calls from existing training or inference scripts
- Command‑line interface installable via pip for rapid prototyping and CI/CD automation
- Cloud‑native API for batch generation and remote optimization, supporting enterprise-scale workloads
- Secure, version‑controlled code artifacts with reproducible build metadata