
DevEval
DevEval provides an evidence-first technical interview platform designed for AI-era hiring, moving beyond traditional coding tests to assess a candidate's judgment, verification skills, and ability to direct AI tools. The platform offers features like AI Critique, Verification Mode, and session replay to give reviewers concrete evidence of a candidate's reasoning process, not just their final answer. It measures the full engineering workflow by capturing how candidates catch AI mistakes, test their assumptions, and explain their decisions under uncertainty.
- Artificial Intelligence
- AI Agents
- HR Technology
- Software Only
Funding
Founders
Founder details are not available yet.
Product
Problem
Traditional technical interviews and coding tests over-index on producing a passing solution strategy, which no longer captures how engineers actually build with AI in the loop. These methods fail to measure the critical skills of debugging under uncertainty, catching confident-but-wrong AI code, and verifying output, leading to hiring decisions based on memorized answers or pedigree rather than practical engineering judgment.
Solution
DevEval is an evidence-first technical interview and live validation platform that evaluates real engineering judgment by measuring a candidate's full workflow with AI. The platform focuses on whether candidates can direct AI, challenge it, fix its mistakes, and verify its output, using features like AI Critique and Verification Mode. Instead of a black-box score, it provides a comprehensive evidence packet including code replay, integrity events, and AI-assisted scorecards for human reviewers to inspect. The platform connects this evidence directly to the interview conversation, allowing follow-up questions on the candidate's specific decisions and gaps while the work is still in front of the reviewer.
Target Audience
Primary customers are engineering leaders, hiring managers, and interview-operations teams at companies who need to evaluate candidates' practical judgment with AI, including the ability to debug, verify, and direct AI tools in a work-specific context.
Features
- AI Critique screens that measure whether candidates can catch confident-but-wrong AI code and hold their ground when it pushes back
- Verification Mode that gives candidates AI-generated code and observes who can actually find bugs versus who blindly trusts the output
- Code Replay and fair work trail showing how a candidate approached each problem edit-by-edit, with test runs and final code preserved
- Integrity events capturing focus changes, paste events, and run history to provide concrete reviewer context without face or emotion scoring
- Discernment Evidence showing whether a candidate caught real defects, avoided false alarms, and expressed uncertainty where captured
- Support for 14 programming languages including Python, TypeScript, Java, Go, Rust, and SQL with full code execution
- A candidate-side practice mode for developers to sharpen skills in AI-era interview problems