Judge the model, not the demo reel.
Creative models break the benchmarks that work everywhere else. What an eval is, what vendor neutrality actually requires, and how to read a leaderboard without being sold something.
What is an AI eval?
An eval is a fixed task, a fixed way of running it, and a fixed way of scoring the result, applied identically to every model in the comparison. Three parts: the task, the harness, the rubric.
The test of whether you have one is whether a stranger can rerun it and get your numbers. A run nobody else can reproduce is an anecdote with a chart on it. That is why a serious benchmark publishes the harness, the prompts, the seeds, and the raw outputs, not only the scores.
Why is evaluating creative models harder?
Because there is no ground truth. A classifier has a correct label to be scored against; a shot has no correct pixels, so every creative eval is a judgment procedure rather than an accuracy measurement.
Five things follow from that, and each one is a way for a benchmark to be quietly wrong.
- Scoring needs a judge. Human panels are slow and inconsistent; model judges are fast and carry their own preferences. Either way the rubric has to be written and published before the run.
- Prompt skill contaminates the score. A model's result is partly the prompter's craft. Either fix the prompt discipline across all models or sweep a range and report the spread.
- Beauty is not the variable. Taste varies by viewer; production usability does not. Measure whether a working editor would cut the output into a piece.
- The targets move. Hosted endpoints change under the same model name. Every run needs a date and a version, and results decay in months.
- Sampling hides the truth. A vendor shows the best of fifty attempts. Production is closer to the first of three. Report attempts, not just outcomes.
What does a vendor-neutral benchmark require?
Six conditions, all of them structural rather than stated. Neutrality is something a benchmark's design makes impossible to violate, not something its authors promise.
- No vendor money touches the ruler. Not task selection, not design, not results. Sponsorship may fund venues, scholarships, and compute to reproduce a run, and nothing else.
- The task is chosen before any model runs. Picking the task after seeing early results is how a benchmark becomes a press release.
- Judging is blind. Judges do not know which model produced which output, and outputs arrive in shuffled order.
- The harness ships with the numbers. Prompts, seeds, code, and raw outputs, downloadable. Anyone who disputes a score can rerun it.
- The sweep is complete. Every major model family, including the open-weight tier and the ones the authors like least. A missing family is a result being hidden.
- Failures publish alongside wins. Including the runs where the cheapest model was the right call, which is more often than any vendor wants in print.
This is the society's standing policy, and the benchmark page states it in full: sponsors fund the room and the compute, never the ruler.
What should a creative eval measure?
Production outcomes, in units a producer can put in a budget. Six measures cover most of what matters.
- Usable-shot rate. The fraction of attempts that clear a usability bar written before the run. This is the headline number, and the bar has to be published with it.
- Attempts per usable shot. What the model actually costs in patience.
- Cost per usable shot. Not cost per generation. A cheap model that fails four times in five is not cheap.
- Wall clock per usable shot. Queue time is production time.
- Failure taxonomy. Which way it breaks: identity drift, physics, hands, lip sync, artifacting, prompt disobedience. Two models with the same score fail differently, and the difference decides which one fits your work.
- Instruction adherence. Whether the shot you asked for is the shot you got, scored separately from whether it looks good.
How do you read a benchmark without being fooled?
Seven checks, in order. Any one of them failing is enough to set the result aside.
- Check the date. In this field six months is a generation.
- Check who paid. For the study, the compute, and the authors.
- Check that the harness is downloadable. If it is not, the piece is marketing with a table in it.
- Check the sample size. Four shots per model is a vibe, not a measurement.
- Check that cost and latency sit on the same table as quality. Quality alone always flatters the most expensive model.
- Check that losing outputs are shown. A benchmark that only publishes winners has not shown you the failure modes you will actually hit.
- Check that the task resembles your work. Best on a four-second hero shot tells you nothing about a forty-shot narrative.
Then hold your own opinion to the same standard. The society argues its published runs in person at Benchmark Night, monthly from this fall, where the numbers go on the big screen and the room checks them against its own eyes.
Get the first report.
The first measured run publishes in September. Leave your email. It lands in your inbox the day it drops, with the Benchmark Night invite.