VeloDB offers a single SQL‑compatible database built on Apache Doris that combines real‑time analytics, hybrid full‑text and vector search, and AI‑native functions such as LLM inference and observability. The platform lets developers run semantic queries, Retrieval‑Augmented Generation, and model monitoring directly in the database, eliminating separate analytics, search, and AI stacks while supporting petabyte‑scale ingestion and low‑latency workloads.
Funding
Funding not disclosed
Founders
Product
Problem
Enterprises often rely on separate analytics databases, search engines, and AI pipelines, which creates data silos, incurs high operational costs, and introduces latency when combining real‑time analytics with semantic search or AI observability. Maintaining multiple stacks also complicates scaling and increases engineering overhead.
Solution
VeloDB delivers a single, SQL‑compatible engine built on Apache Doris that unifies real‑time analytics, hybrid full‑text/vector search, and AI‑native capabilities. The platform exposes a standard SQL interface while embedding large‑language‑model (LLM) functions for semantic query expansion and text analysis, allowing developers to run AI‑driven workloads directly in the database. Hybrid search merges keyword and vector similarity in one query path, supporting Retrieval‑Augmented Generation (RAG) use cases without external services. Native integrations for AI observability capture logs and traces from LLMs, feeding them into a low‑cost, columnar lakehouse that handles structured and unstructured data at petabyte scale. High‑throughput ingestion (up to 60 M rows per minute) and low query latency (sub‑150 ms) enable real‑time dashboards and recommendation engines. By consolidating ClickHouse, Trino, and Elasticsearch functionalities, VeloDB reduces infrastructure complexity and operational spend while maintaining compatibility with existing SQL tooling.
Target Audience
The primary customers are data engineering and analytics teams, as well as AI/ML engineers in mid‑size to large enterprises that require real‑time analytics, semantic search, and integrated model observability within a single database platform.
Features
- Unified Apache Doris core with ANSI‑SQL support, eliminating the need for separate analytics and search clusters
- Hybrid search engine that processes full‑text and dense vector queries in a single execution plan for RAG and recommendation workloads
- AI‑powered SQL extensions that embed LLM inference for semantic filtering, text summarization, and entity extraction within standard queries
- Real‑time data ingestion pipeline capable of 60 M rows/min and sustained 10 k QPS with sub‑150 ms latency guarantees
- Columnar JSON and multi‑modal storage layer forming an AI‑ready lakehouse for structured and unstructured data at petabyte scale
- Built‑in AI observability stack with native Langfuse connectors, capturing model logs, traces, and metrics for end‑to‑end monitoring
- Compatibility layer for ClickHouse, Trino, and Elasticsearch query syntaxes, enabling lift‑and‑shift migrations without code changes
- Cost‑optimization features such as automatic data tiering and high compression ratios (up to 5:1) to lower storage expenses