PuppyGraph is a real-time, zero-ETL graph query engine that allows users to query existing relational data stores as a unified graph model. It supports petabyte-scale data and executes complex multi-hop queries in seconds, integrating natively with SQL and graph query languages like Gremlin and Cypher. This platform enables advanced graph analytics use cases such as fraud detection and enhancing LLMs without the latency or maintenance overhead of traditional graph databases.
Funding
$5M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.

DVFounders
Product
Problem
Organizations struggle to efficiently query and analyze relationships within their data lakes and warehouses due to the complexity of ETL processes and the limitations of traditional relational databases. This makes it difficult to uncover hidden insights and patterns in a timely manner.
Solution
PuppyGraph is a graph query engine that enables real-time graph analytics directly on data lakes and warehouses, eliminating the need for ETL. By decoupling the query engine from storage, PuppyGraph allows users to seamlessly query one or multiple data stores as a unified graph model. This approach facilitates efficient data exploration, reduces system complexity, and unlocks new possibilities for discovering interconnected insights across diverse data categories. PuppyGraph supports petabyte-level scalability and delivers complex queries in seconds through parallel processing and vectorized evaluation.
Target Audience
PuppyGraph targets data scientists, data engineers, and analysts who need to perform graph analytics on large datasets stored in data lakes and warehouses across industries like AI, finance, and healthcare.
Features
- Zero-ETL: Queries data directly from data warehouses and lakes, eliminating complex ETL pipelines.
- Petabyte-level scalability: Auto-sharding across data sources enables graph querying of petabytes of data.
- High-performance queries: Delivers 10-hop neighbor queries in seconds across billions of edges using parallel processing and vectorized evaluation.
- Native integration: Connects to various data sources, including Databricks, Amazon S3, and Google BigQuery.
- Query language support: Compatible with Gremlin and Cypher query languages.
- Rapid deployment: Can be deployed and queried within minutes using Docker.
- Unified data access: Enables querying of data as a graph while still allowing SQL access.