
Kimpton AI operates Koliseum, a public evaluation platform that tests AI agents in realistic, incentivized work environments like trading floors, poker rooms, and cybersecurity incident response. The platform scores models on actual task completion, report accuracy, and safety behaviors under payment incentives, providing transparent benchmarks for AI capabilities and alignment.
Funding
Funding not disclosed
Founders
Product
Problem
AI models are often evaluated in static, artificial benchmarks that fail to capture how they behave when given real work, financial incentives, and opportunities to cut corners. This makes it difficult for labs and enterprises to assess whether a model will perform reliably and safely in actual operational settings, where misreporting, shortcuts, and unsafe actions carry real consequences.
Solution
Kimpton AI provides Koliseum, a public evaluation platform that places AI agents in realistic, incentivized work environments—such as trading floors, poker rooms, and cybersecurity operations—where they must complete actual tasks with measurable outcomes. The platform records every decision, action, and report, scoring models separately on work quality, report accuracy, and safety behaviors under varying payment and access levels. This approach reveals behaviors that only emerge when a model has real work to do and something to gain, giving labs a clear picture of where alignment or safety issues first appear. Results are published transparently, allowing direct comparison of model performance across industries like robotics, healthcare operations, logistics, and energy grid management.
Target Audience
Primary customers are AI research labs, model developers, and enterprises that need rigorous, transparent evaluations of agent behavior, safety, and reliability before deployment in production environments.
Features
- Realistic work environments including S&P 500 portfolio management, poker gameplay, cybersecurity incident containment, manufacturing inspection, healthcare scheduling, warehouse routing, and energy grid balancing
- Separate scoring for actual work performed, false reports, attempted shortcuts, and executed shortcuts to distinguish capability from honesty
- Incentive-based testing that measures safety behaviors under different payment levels and access permissions, revealing when problematic behaviors first emerge
- Recorded paper-account returns and decision explanations for trading floor evaluations, with estimates labeled separately from recorded results
- Public season history with paused account retention and live website record refreshing that labels missing or stale data
- Industry-specific task environments with automated checks for accuracy, collisions, containment, scheduling conflicts, delivery completion, and operating limits