Vals provides independent, expert‑crafted benchmarks that evaluate large language models on sector‑specific professional tasks in finance, software engineering, and education. Using proprietary, non‑public datasets and a severity‑weighted scoring system with dealbreaker gating, Vals aggregates results into the Vals Index and Vals Multimodal Index, weighted by each sector's contribution to U.S. GDP, to deliver a single, economically meaningful performance metric for AI developers, enterprises, and policymakers.
Funding
$5M raised to dateRaised to date based on public sources. This may differ from the amount the company actually raised and is based only on what is publicly available on the internet.
Founders
Product
Problem
Organizations lack reliable, sector‑specific metrics to assess how large language models will perform on real‑world professional tasks, making it difficult to predict economic impact and deployment risk.
Solution
Vals delivers independent, expert‑crafted benchmarks that evaluate LLMs on concrete finance, software engineering, and education workloads. Each benchmark uses non‑public, domain‑specific datasets and applies severity‑weighted partial credit with dealbreaker gating to ensure critical errors are penalized. Model scores are aggregated into the Vals Index and Vals Multimodal Index, where sector weights reflect their contribution to U.S. GDP, providing a single, economically meaningful performance number. The evaluation framework runs multiple inference passes per model and reports the mean score, enabling cost‑effective, repeatable comparisons across model releases. Results are published publicly, giving AI developers, enterprises, and policymakers a transparent basis for deployment decisions.
Target Audience
Primary users are AI model developers, research labs, and enterprise AI teams that need sector‑specific performance data to guide model selection, fine‑tuning, and risk assessment for finance, software engineering, and educational applications.
Features
- Proprietary finance, coding, and education benchmarks (e.g., CorpFin, Finance Agent v2, SWE‑bench Verified, Terminal‑Bench 2.0, Vibe Code Bench, SAGE) built on private, expert‑validated datasets
- Economic weighting of sectors (finance ≈ $2 T, coding ≈ $1.4 T, education ≈ $0.27 T) to produce a GDP‑scaled index score
- Severity‑weighted partial credit with “dealbreaker” gating that nullifies credit when a load‑bearing fact is incorrect
- Representative subset selection for efficient evaluation while preserving strong correlation with full benchmark performance
- Multi‑run evaluation (three runs per model) and mean aggregation to reduce variance and capture consistent capability
- Open, regularly updated leaderboard and methodology documentation for transparency
- API‑compatible results export for integration into model development pipelines and reporting dashboards