AirMettle develops solutions that accelerate Big Data analytics by processing data directly within the storage tier. This patented technology uses massively parallel processing to reduce the need for data movement, significantly lowering analytics memory, compute, and networking costs. The platform enables faster insight extraction from large datasets, even at petabyte scales, using industry-standard infrastructure.
Funding
Funding not disclosed

Founders
Product
Problem
Organizations face challenges in efficiently analyzing the growing volume and variety of data stored in data lakes. Traditional ETL processes require moving large datasets to separate compute clusters, leading to network congestion, high compute costs, and delays in obtaining actionable insights. This complexity hinders the adoption of transformative AI solutions and limits the value derived from big data assets.
Solution
AirMettle offers a software-defined storage platform with integrated distributed parallel processing that enables direct querying of semi-structured data within the data lake. By processing data in place, AirMettle eliminates the need to transfer large datasets to separate compute clusters for analysis, significantly reducing network traffic and latency. The platform transforms disparate data formats into those required by analytics tools, decreasing the amount of expensive memory needed by higher-level analytics applications. This approach accelerates petabyte-scale analytics, speeds time to insight, and lowers overall analytics infrastructure costs.
Target Audience
AirMettle targets organizations dealing with petabyte-scale data in data lakes, including enterprises, research institutions, and government agencies seeking to accelerate big data analytics and reduce associated infrastructure costs.
Features
- Integrated distributed parallel processing for in-storage data analysis
- Software-defined storage framework compatible with industry-standard servers, storage, and networking solutions
- S3-compatible interface for integration with a range of analytic solutions
- Support for open-source tools like Apache Spark, Presto, and Arrow
- Hierarchical Multi-Dimensional Histograms (HMDH) solution for high-performance exploratory analysis of large, complex datasets
- Ability to characterize floating-point and integer data distributions in up to four dimensions in a single pass
- Compact histogram files that are significantly smaller than raw data while retaining precise counts