Skip to main content

LakeOps

LakeOps provides an autonomous optimization platform for Apache Iceberg lakehouses, continuously monitoring table health and automatically executing maintenance operations to improve query performance and reduce storage costs. The platform learns from query patterns to adapt optimization strategies across multiple engines, requiring no code changes or vendor lock-in.

HQ unknown
Founded 20253500+ followers
Updated 9 days ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Data lakehouse tables degrade over time as data accumulates, leading to small file proliferation, snapshot bloat, and suboptimal file layouts that slow query performance and inflate storage costs. Traditional maintenance requires manual scheduling, complex runbooks, and constant monitoring, forcing data teams to choose between time-consuming upkeep and degraded analytics performance.

Solution

LakeOps delivers an intelligent control plane that autonomously optimizes Iceberg lakehouse performance, cost, and health across every table and engine. The platform continuously senses table conditions, plans optimization sequences based on severity and dependency order, executes operations on a Rust-based engine, and learns from outcomes to refine future decisions. It monitors file layout, metadata health, and query patterns to trigger compaction, snapshot expiration, orphan cleanup, and manifest rewrites only when needed—eliminating cron jobs and manual tuning. The system adapts cadence to write velocity, processing streaming tables hourly, batch tables daily, and skipping idle tables entirely. Built on Rust and Apache DataFusion, LakeOps delivers up to 12x faster queries and 76% less CPU consumption compared to Spark-based alternatives, while processing only metadata—never retaining or storing customer data.

Target Audience

Primary customers are data engineering and analytics platform teams managing Apache Iceberg lakehouses at scale, particularly those running multi-engine environments with Trino, Spark, Snowflake, Athena, DuckDB, or Flink who need automated table maintenance without operational overhead.

Features

  • Autonomous closed-loop optimization cycle: sense → plan → execute → learn, with no manual scheduling or threshold tuning required
  • Health-driven operation triggers based on per-table scores (Healthy, Warning, Critical) calculated from file count, small-file ratio, snapshot age, manifest bloat, and write velocity
  • Dependency-ordered maintenance sequencing: compaction → snapshot expiry → orphan cleanup → manifest rewrite, ensuring each step's output feeds cleanly into the next
  • Query-pattern-aware sorting that physically re-sorts data files to match actual filter, join, and group-by patterns, enabling engines to skip entire file groups via min/max pruning
  • Cross-catalog and cross-engine telemetry unified from Trino, Spark, Snowflake, Athena, DuckDB, and Flink into a single event trail with full operation history
  • Rust and Apache DataFusion engine delivering 95% faster and 90% cheaper compaction than Spark-based tools, with zero Spark clusters required
  • Adaptive cadence that scales from 50 to 50,000+ tables, adjusting maintenance frequency to write velocity (hourly for streaming, daily for batch, skipped for idle)
  • Audit-ready event trail with per-operation impact metrics (files merged, bytes reclaimed, duration) supporting SOC 2 and GDPR compliance reporting
This profile is AI-generated and may contain inaccuracies.