Skip to main content
EI

Embucket, Inc.

This startup offers cloud storage optimized for AI and data workloads, providing S3-compatible object storage for large datasets. Their platform aims to improve performance and reduce costs associated with storing and accessing data for AI applications.

San Francisco, United StatesFounded 202413100+ followers
Updated 3 months ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Product

Problem

Many organizations face challenges in processing data due to data residency and sovereignty constraints, high costs associated with certain workloads, or the impracticality of shipping large datasets to cloud data warehouses. Existing cloud-based data warehouses may also be unavailable in certain cloud regions or incompatible with prepaid cloud credits.

Solution

Embucket provides a Snowflake-compatible lakehouse platform that enables users to run Snowflake workloads locally without re-platforming. It leverages Apache Iceberg for data storage and offers a Snowflake-style REST API and SQL dialect. The platform features a zero-disk architecture, storing all state (data and metadata) in object storage (S3 or memory). Embucket is delivered as a single, statically linked binary for easy deployment and supports scalable "query-per-node" parallelism, allowing multiple instances to run against the same bucket for horizontal scaling.

Target Audience

Embucket targets data professionals and organizations that require local processing of Snowflake workloads due to data residency, cost constraints, or other limitations associated with cloud-based data warehouses.

Features

  • Snowflake SQL syntax dialect and v1 wire-compatible REST API
  • Apache Iceberg table format for data storage on object storage
  • Built-in internal catalog and Iceberg REST Catalog API
  • Zero-disk architecture with all state residing in S3 buckets
  • Scalable "query-per-node" parallelism
  • Single statically linked binary for effortless deployment
  • Iceberg catalog federation to connect to external Iceberg REST catalogs
  • Integration with Apache DataFusion as the query engine
  • Metadata persistence using SlateDB
This profile is AI-generated and may contain inaccuracies.