Smoothli is an AI‑driven platform that automates image dataset preparation for computer‑vision models by classifying existing images, identifying gaps, and generating high‑quality synthetic samples. It uses graph neural networks to analyze data relationships and recommend optimal training configurations, while metadata filters streamline the dataset for more efficient training.
Funding
Funding not disclosed
Founders
Product
Problem
Machine learning projects often struggle with incomplete or poorly organized image datasets, leading to suboptimal model performance and high labeling costs. Manually identifying gaps, filtering metadata, and creating synthetic data to augment training sets is time‑consuming and requires specialized expertise.
Solution
Smoothli offers an AI‑driven platform that automates image dataset preparation for computer‑vision models. The system first analyzes existing images, categorizing them by format, subject, orientation, palette, and style, and highlights missing data segments that could improve model accuracy. Using graph neural networks, it evaluates relationships within the data to recommend optimal training configurations. A generative AI engine then creates high‑quality synthetic images that blend seamlessly with the real data, reducing the need for costly manual labeling. Integrated metadata filters allow users to prune irrelevant samples, resulting in leaner, more effective training pipelines.
Target Audience
Primary customers are machine‑learning engineers, data scientists, and AI product teams building computer‑vision applications who need efficient, high‑quality image datasets.
Features
- Automatic classification of images by format, content, orientation, color palette, and artistic style
- Gap analysis tool that identifies under‑represented classes or attributes in the dataset
- Graph neural network optimizer that suggests model architecture and training parameters based on data relationships
- Generative AI image creator that produces realistic synthetic samples to augment real data
- Metadata‑driven filtering to remove low‑value or redundant images and streamline training sets
- Dashboard for visualizing dataset composition, synthetic augmentation impact, and performance metrics