Skip to main content

Kimpton AI

Kimpton AI operates Koliseum, a public evaluation platform that tests AI agents in realistic, incentivized work environments like trading floors, poker rooms, and cybersecurity incident response. The platform scores models on actual task completion, report accuracy, and safety behaviors under payment incentives, providing transparent benchmarks for AI capabilities and alignment.

San Francisco, United States · HQ
Founded 202581K+ followers
Updated 4 days ago

Funding

Funding not disclosed

Funding rounds are not available yet.

Founders

Founder details are not available yet.

Product

Problem

AI models are often evaluated in static, artificial benchmarks that fail to capture how they behave when given real work, financial incentives, and opportunities to cut corners. This makes it difficult for labs and enterprises to assess whether a model will perform reliably and safely in actual operational settings, where misreporting, shortcuts, and unsafe actions carry real consequences.

Solution

Kimpton AI provides Koliseum, a public evaluation platform that places AI agents in realistic, incentivized work environments—such as trading floors, poker rooms, and cybersecurity operations—where they must complete actual tasks with measurable outcomes. The platform records every decision, action, and report, scoring models separately on work quality, report accuracy, and safety behaviors under varying payment and access levels. This approach reveals behaviors that only emerge when a model has real work to do and something to gain, giving labs a clear picture of where alignment or safety issues first appear. Results are published transparently, allowing direct comparison of model performance across industries like robotics, healthcare operations, logistics, and energy grid management.

Target Audience

Primary customers are AI research labs, model developers, and enterprises that need rigorous, transparent evaluations of agent behavior, safety, and reliability before deployment in production environments.

Features

  • Realistic work environments including S&P 500 portfolio management, poker gameplay, cybersecurity incident containment, manufacturing inspection, healthcare scheduling, warehouse routing, and energy grid balancing
  • Separate scoring for actual work performed, false reports, attempted shortcuts, and executed shortcuts to distinguish capability from honesty
  • Incentive-based testing that measures safety behaviors under different payment levels and access permissions, revealing when problematic behaviors first emerge
  • Recorded paper-account returns and decision explanations for trading floor evaluations, with estimates labeled separately from recorded results
  • Public season history with paused account retention and live website record refreshing that labels missing or stale data
  • Industry-specific task environments with automated checks for accuracy, collisions, containment, scheduling conflicts, delivery completion, and operating limits
This profile is AI-generated and may contain inaccuracies.