The benchmark.
One production-shaped task. Same prompt discipline, same eval harness, every model family.
Numbers, cost, and the harness itself. Open, reproducible, no sponsor thumbs on the scale.
Each run will headline the monthly Benchmark Night, starting this fall: results on the big screen, creatives and engineers checking the scores against their own eyes.
The neutrality rule
The society’s independence lives on this page. Our standing policy: no model vendor’s money touches benchmark design, task selection, or results. Ever. Sponsors may fund scholarships, venues, and the compute to reproduce our runs. They may not fund the ruler. Every run ships with its harness so anyone can check our work. The first full report publishes this September and replaces the example above. What we require of a benchmark, and what you should require of anyone else’s, is set out in how to evaluate creative AI models.
Come argue with the results
Every published run headlines Benchmark Night, monthly from this fall in San Francisco: what the numbers hide, when the cheap model is the right call, and which task runs next. The invite goes out with the report.
Get the first report.
The first measured run publishes here in September. Leave your email. It lands in your inbox the day it drops, with the Benchmark Night invite.