DesignArena is an open, crowdsourced platform that benchmarks AI-generated designs through blind pairwise voting. By running four‑model tournaments on identical prompts and aggregating votes with a Bradley‑Terry model, it produces real‑time, reproducible strength scores and rankings that reflect community aesthetic preferences across various media types.
Funding
Funding not disclosed

Founders
Product
Problem
Evaluating the creative quality of AI design models is difficult because existing benchmarks rely on curated rankings or limited test sets, leading to biased or non-representative results. Users and developers lack a transparent, community-driven way to compare how well different models generate designs that match human preferences.
Solution
DesignArena provides an open, crowdsourced platform where AI-generated designs are compared through blind pairwise voting. In each tournament, four models receive the same prompt, generate outputs simultaneously, and the first two completed designs are presented anonymously for a user vote. Votes are aggregated using the Bradley‑Terry statistical model to produce strength scores and rankings that reflect collective community preferences. All model configurations, temperature settings, and evaluation procedures are publicly documented, ensuring reproducibility and fairness. The leaderboard updates in real time as more pairwise comparisons are collected, giving developers a continuously refreshed benchmark of model performance across multiple media types.
Target Audience
DesignArena serves AI model developers, design tool vendors, and creative professionals who need an unbiased benchmark to assess how well generative design models align with user aesthetic preferences.
Features
- Blind, anonymous presentation of design outputs to eliminate brand bias during evaluation
- Randomized four‑model tournaments with identical prompts for consistent comparison conditions
- Pairwise voting interface that records each vote equally without editorial filtering
- Bradley‑Terry based ranking algorithm that converts win‑loss data into normalized strength ratings
- Support for multiple content types (code, image, video, audio, slides, etc.) and design categories
- Public documentation of model versions, temperature settings, and evaluation methodology
- Real‑time leaderboard that updates as new pairwise comparisons reach statistical reliability thresholds