Skip to main content
L

lakeFS

Provides a scalable data version control system that enables data engineers and scientists to manage data with Git-like operations, supporting reproducible pipelines and parallel experimentation. By integrating with cloud object storage and compute engines, it reduces storage costs, improves data quality, and enables rollback capabilities to maintain production integrity.

Santa Monica, United StatesFounded 2020285K+ followers
Updated 20 months ago

Funding

$23M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

DTNV
Funding rounds are not available yet.

Founders

Product

Problem

Data engineers and scientists face challenges in managing and versioning large datasets, hindering reproducibility and collaboration in data pipelines. Traditional methods lack the scalability and Git-like functionality needed for efficient data management, leading to increased storage costs and compromised data quality.

Solution

lakeFS provides a scalable data version control system that enables data engineers and scientists to manage data in object storage with Git-like operations, ensuring reproducibility and parallel experimentation. By leveraging metadata to manage data versions, lakeFS minimizes the impact on storage performance while supporting various compute engines and data formats. The platform allows users to create zero-copy branches for isolated development and testing environments, implement CI/CD for data pipelines, and rollback to previous commits in case of data quality issues. lakeFS integrates with existing data stacks, including object storage, compute engines, ingest technologies, and orchestration tools, to streamline data management workflows.

Target Audience

lakeFS targets data engineers, data scientists, and data ops professionals who need to manage and version large datasets, ensure data quality, and streamline collaboration in data-intensive workflows.

Features

  • Git-like branching, merging, and rollback capabilities for data stored in object storage
  • Zero-copy branches for parallel experimentation and isolated development environments
  • Integration with major cloud providers (AWS, Azure, GCP) and on-premise storage solutions (MinIO, Ceph, Dell EMC)
  • Support for various compute engines, including Spark, Trino, Presto, and Databricks
  • Compatibility with open table formats like Delta Lake, Iceberg, and Hudi, as well as data formats like Parquet, CSV, and Avro
  • CI/CD implementation for data pipelines with automated quality validation checks
  • Data auditing to track changes and ensure data governance
  • Unified data access across multiple storage endpoints through a single API and namespace
This profile is AI-generated and may contain inaccuracies.