Oleander provides a fully managed, serverless Spark compute platform with integrated Iceberg lake storage that automatically captures logs, traces, lineage, and query plans.
Funding
Funding not disclosed
Founders
Product
Problem
Data teams using Spark and Iceberg often struggle with fragmented observability, manual root‑cause analysis, and slow remediation of pipeline failures, leading to production downtime, cost overruns, and data quality drift.
Solution
Oleander delivers a fully managed, serverless Spark compute environment paired with managed Iceberg lake storage, eliminating the need for complex cluster provisioning. The platform continuously captures logs, traces, lineage events, and query plans, feeding them to an LLM that automatically generates root‑cause analyses and writes corrective code. It then opens a pull request, allowing engineers to review and merge the fix, after which Oleander redeploys the updated job and verifies the resolution in production. Integration with Databricks, EMR, SageMaker and other tools enables fast onboarding via a CLI and contextual SQL queries, providing end‑to‑end reliability from day zero.
Target Audience
Oleander is aimed at fast‑moving data engineering and ML teams that run Spark jobs on Iceberg tables and need automated observability and rapid remediation, especially those using Databricks, EMR, or SageMaker.
Features
- Serverless Spark compute with zero‑setup deployment and built‑in OpenLineage telemetry
- Managed Iceberg storage (public and private) with automatic schema evolution
- LLM‑driven incident investigation that aggregates logs, traces, lineage, and Spark plan diffs into a single investigation link
- Automatic generation of fix code, creation of GitHub pull requests, and post‑merge redeployment verification
- Real‑time anomaly detection for row counts, duration, and schema changes with Slack and webhook alerts
- Contextual SQL query interface for correlating metrics, logs, and lineage across the telemetry lake
- Compatibility with Databricks, Amazon EMR, SageMaker, Glue, Dataproc, and Jupyter notebooks
- Tiered pricing with per‑second billing for compute, storage, and query execution