Skip to main content

Base Compute

BaseRT is a high-performance LLM runtime engineered for Apple Silicon, delivering prefill speeds up to 6.4x faster than llama.cpp and 3.9x faster than MLX, with decode up to 1.33x faster than MLX. The platform enables fully on-device AI inference with zero marginal cost per token, eliminating cloud dependencies and ensuring data privacy. It also offers the Base Optimisation Stack (B:OS) for enterprises and chip makers seeking ultra-optimized inference across Apple, NVIDIA, and AMD hardware.

Melbourne, Australia · HQ
Founded 20263700+ followers
Updated 9 days ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Running large language models in the cloud introduces latency, privacy risks, and recurring per-token costs, while existing on-device runtimes like llama.cpp and MLX fail to fully exploit the hardware's capabilities. Organizations and individuals seeking private, low-cost AI inference are constrained by slow prefill and decode speeds that make local deployment impractical for real-time or agentic workloads.

Solution

BaseRT provides a specialized LLM runtime that maximizes inference performance on Apple Silicon, achieving prefill speeds up to 6.4x faster than llama.cpp and 3.9x faster than MLX, with decode up to 1.33x faster than MLX. The runtime auto-optimizes for each machine, enabling instant setup with no cloud, no accounts, and zero cost per token. For enterprises, BaseRT offers the Base Optimisation Stack (B:OS), a full-stack tuning approach that optimizes every layer from model to kernel to silicon, delivering day-0 support for new model releases across Apple, NVIDIA, AMD, and other accelerators. This makes on-device inference viable for lower latency, real privacy, and complete independence from network or provider dependencies.

Target Audience

Primary customers are enterprises seeking to run AI within their own environments, as well as model makers, chip manufacturers, and device makers shipping AI at scale who need ultra-optimized on-device inference.

Features

  • Prefill acceleration up to 6.4x faster than llama.cpp and 3.9x faster than MLX on Apple Silicon
  • Decode performance up to 1.33x faster than MLX, enabling smoother real-time token generation
  • Auto-optimized setup for each machine, requiring no manual configuration or cloud connectivity
  • Base Optimisation Stack (B:OS) for enterprises, tuning model, kernel, and silicon layers for maximum hardware utilization
  • Cross-accelerator support including Apple Silicon, NVIDIA, AMD, and other accelerators
  • Day-0 compatibility with new model releases, ensuring immediate access to the latest open-weight models
This profile is AI-generated and may contain inaccuracies.