Skip to main content
Y

yasp

yasp provides an agentic AI compiler that automatically detects hardware and generates device‑specific kernels via a single‑line Python API compatible with PyTorch, TensorFlow, and JAX. The platform applies graph‑level optimizations, reinforcement‑learning auto‑tuning, and integrated profiling to deliver up to tenfold speedups and lower compute costs across GPUs, TPUs, and custom accelerators.

Updated 2 months ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

AI teams spend extensive time manually tuning kernels and managing heterogeneous hardware stacks, which slows model training, inflates infrastructure costs, and delays deployment of new capabilities. This low‑level optimization burden limits scalability and hampers rapid experimentation.

Solution

yasp delivers an Agentic AI Compiler that abstracts hardware specifics and automates low‑level optimization through a single API call. The compiler performs graph‑level transformations, operator fusion, and memory layout adjustments, then generates device‑specific kernels on the fly. By integrating reinforcement‑learning‑based auto‑tuning, it eliminates hand‑crafted kernel code while achieving near‑optimal performance on GPUs, TPUs, and custom accelerators. The platform provides end‑to‑end profiling and telemetry, enabling developers to monitor speedups and cost savings in real time. As a result, teams can accelerate training workloads by up to tenfold, reduce compute expenses, and deploy models across cloud, on‑prem, or edge environments without refactoring code.

Target Audience

The primary users are machine‑learning engineers, data‑science teams, and AI research groups building large‑scale training or inference pipelines, as well as enterprises that need to scale AI workloads across heterogeneous compute environments.

Features

  • Agentic compiler engine with automatic hardware detection and on‑the‑fly kernel generation for diverse accelerators
  • Single‑line Python API (yasp.compile) compatible with PyTorch, TensorFlow, and JAX model graphs
  • Graph‑level optimizations including operator fusion, layout transformation, and mixed‑precision casting
  • Reinforcement‑learning‑driven auto‑tuning that replaces manual kernel hand‑crafting
  • Integrated profiling suite delivering latency, throughput, and cost metrics per run
  • Containerized runtime for seamless deployment across cloud, on‑prem, and edge environments
  • Secure, encrypted data handling with role‑based access controls for enterprise compliance
This profile is AI-generated and may contain inaccuracies.