This startup automatically identifies potential join relationships within a user's data, regardless of the underlying data structure or format. By surfacing these connections, the platform enables users to quickly integrate disparate datasets and unlock new insights without manual data exploration.
Funding
Funding not disclosed
Founders
Product
Problem
Data warehouses often contain complex, undocumented relationships between tables and columns, making it difficult for data engineers and analysts to efficiently integrate disparate datasets. Manual join discovery, schema analysis, and query validation consume valuable time that could be better spent on data optimization and innovation.
Solution
Schemantic provides an automated data cataloging platform that algorithmically maps and validates joins and entities within a data warehouse, eliminating the need for predefined keys, metadata, or external documentation. The platform deterministically identifies and ranks virtually every valid join, enabling users to explore and understand their data's underlying structure. Schemantic also features an entity explorer that visually maps how entities relate to each other, identifying where identifiers appear and intersect across tables. The platform generates ready-to-run SQL query templates, streamlining reporting and enabling teams to establish standard definitions for key metrics and data points. Schemantic's collaborative data catalog allows teams to review and confirm join recommendations, creating a single source of truth for validated relationships.
Target Audience
The primary users are data engineers and analysts working with complex data warehouses who need to efficiently discover and validate relationships between data entities.
Features
- Comprehensive join detection: Programmatically identifies and validates joins across the data warehouse, even with inconsistent column names, data type mismatches, or high null values.
- Robust entity mapping: Automatically maps entities and connects related data for seamless exploration, even in poorly documented tables.
- Collaborative data catalog: Enables teams to review and confirm join recommendations, creating a canonical and trusted set of joins.
- Flawless query generation: Generates precise SQL join paths between any tables with minimal clicks, streamlining reporting and analysis.
- Entity Explorer: Provides a visual map of how entities relate to each other, identifying where identifiers appear and intersect across tables.
- Automated data catalog refresh: Incorporates new or removed tables, updated schemas, and newly discovered joins seamlessly.
- Join Detail view: Visualizes key attributes of each join, such as cardinality, null rate, and duplication rate, to facilitate informed decision-making.