
PayBench is an open evaluation platform that measures how reliably AI agents can handle real-world payment decisions. It runs standardized scenarios comparing frontier models and scripted reference agents on payment effectiveness, human alignment, and appropriate check-in behavior, with results published on a public leaderboard.
Funding
Funding not disclosed
Founders
Product
Problem
AI agents are increasingly being given access to payment tools and the ability to spend money on behalf of users, but there is no standardized way to measure whether they can be trusted with financial decisions. Without rigorous benchmarks, it is difficult to compare models on their ability to pay correctly, refuse unsafe payments, ask for clarification when needed, and align with what human users actually want.
Solution
PayBench provides a structured evaluation framework and public leaderboard for assessing AI payment behavior. The platform runs controlled scenarios that test frontier models alongside scripted reference agents across multiple conditions, generating quantitative scores for payment effectiveness, refusals, and human alignment. Results are published on a public site with leaderboards, forest plots, and per-condition breakdowns, enabling direct comparison of model performance. The project also maintains a research paper and changelog documenting methodology and findings as the benchmark evolves.
Target Audience
AI research labs, model developers, and safety evaluators who need empirical benchmarks for assessing the trustworthiness of AI agents in financial decision-making contexts.
Features
- Standardized scenario suite covering unsafe payments, refusals, and check-in decision timing
- Scripted reference agents ("Scripted buy-first" and "Scripted check-first") establishing baseline performance floors
- Human-alignment evaluation via two survey metrics: top-pick rate and acceptable-pick rate, plus a reference line for desired ask frequency
- Public leaderboard with model scores, confidence intervals, and vendor-specific color coding and logos
- Control-layer contrast analysis showing before/after rates, points change, relative change, and 95% ranges for each evaluated layer
- CI-guarded paper generation pipeline with versioned figures and a locked proposal hash to ensure reproducibility