Sutro Functions offers a platform for quickly creating expert‑aligned judges, classifiers, and data extractors without requiring prompt engineering, fine‑tuning, or pre‑labeled datasets. Users upload unlabeled data, the system auto‑labels easy cases and highlights ambiguous examples, then a simple swipe‑left/right annotation flow lets experts confirm decisions and rationales, producing measurable, regression‑proof artifacts. The tool aims to reduce labeling cost to zero and accelerate model evaluation cycles.
Funding
Funding not disclosed
Founders
Product
Problem
Organizations need reliable AI classifiers, judges, and extractors but face high costs and time requirements for prompt engineering, fine‑tuning, and manually labeling large datasets. Without a fast way to generate ground‑truth data, model quality remains opaque and difficult to measure.
Solution
Sutro Functions offers a platform that turns unlabeled data into high‑quality training sets without requiring prompt engineering or extensive fine‑tuning. Users upload raw data; the system automatically labels easy cases and surfaces low‑confidence examples for expert review via a simple swipe‑based annotation UI. Expert decisions are compiled into decision‑preference functions that encode generalizable rules through automated prompt optimization, producing consistent, regression‑proof judges, classifiers, and extractors. The platform provides measurable AI quality metrics, such as model‑consensus and user‑model agreement scores, enabling continuous monitoring and optimization. Results can be deployed directly in production batch inference pipelines using a web UI, SDK, or managed cloud environment.
Target Audience
Primary customers are AI product teams, data scientists, and enterprises, and developers who need fast, expert‑aligned classifiers or decision functions for large‑scale batch inference tasks.
Features
- Automatic labeling of high‑confidence instances with immediate surfacing of ambiguous cases for human annotation
- Swipe‑left/right UI that streamlines expert review and accelerates ground‑truth dataset creation
- Function generation that learns decision rules rather than memorizing examples, ensuring adaptability to new data
- Built‑in metrics (model consensus, user‑model agreement) for quantifiable AI quality assessment
- Integration with frontier open‑source LLMs and custom models, selectable via the platform’s SDK or web interface
- Scalable batch inference jobs with configurable token quotas and managed cloud or bring‑your‑own‑cloud options