Skip to main content

SOVAI

sov.ai provides a Python-first data catalogue of 27 ticker-linked alternative datasets covering equity, economic, and sectorial data for US listed and delisted companies, with history back to 1985. The platform offers an SDK, pattern recognition, and feature processing tools, delivering data via Parquet, Snowflake, or Databricks. It is designed for asset managers who need transparent, queryable datasets for quantitative research and backtesting.

Boston, United States · HQ
Founded 20212K+ followers
Updated 4 days ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

Asset managers and quantitative researchers often struggle to access reliable, well-documented alternative datasets that include historical coverage of both listed and delisted companies. Many data providers lack transparency about coverage gaps, history start dates, or entity counts, making it difficult to build accurate backtests and avoid survivorship bias.

Solution

sov.ai offers a Python-first data catalogue with 27 ticker-linked alternative datasets spanning equity, economic, and sectorial data for US companies. Each dataset includes published metadata on history start, entity count, and universe, with unmeasured fields shown as gaps rather than omitted. Users can query datasets directly via the sovai Python SDK, generate reports, run packaged computations, and apply pattern recognition or feature processing tools. Data is delivered in versioned, partitioned formats such as Parquet, with optional cloud destinations including AWS S3, Snowflake, and Databricks.

Target Audience

Primary customers are asset managers, quantitative researchers, and data scientists at investment firms who need transparent, ticker-linked alternative datasets for signal generation, backtesting, and portfolio analysis.

Features

  • 27 datasets: 19 equity, 4 economic, and 4 sectorial, with the widest covering 14,500 entities
  • History reaching back to 1985, with earliest observations in Bankruptcy Predictions
  • 22 of 22 company-level datasets retain delisted companies to prevent survivorship bias in backtests
  • Python SDK with verified query strings, plus plotting, reporting, and compute functions
  • Pattern recognition, anomaly detection, clustering, feature extraction, and dimensionality reduction tools
  • Delivery via Parquet (default), CSV, JSON, Delta, or direct cloud destinations like S3, Snowflake, and Databricks
  • Client-specific views can omit or mask named columns for data governance
This profile is AI-generated and may contain inaccuracies.