The benchmark

One task. Every AI model.

The society’s open benchmark, and the standard the Multimodal Credential is graded against. One production-shaped task, every major model family, the numbers published with the harness. The first report lands in September.

Run

One production-shaped task. Same prompt discipline, same eval harness, every model family.

Publish

Numbers, cost, and the harness itself. Open, reproducible, no sponsor thumbs on the scale.

Argue in person

Each run will headline the monthly Benchmark Night, starting this fall: results on the big screen, creatives and engineers checking the scores against their own eyes.

SAMPLE LAYOUT · NOT MEASURED DATA
$ mms eval run --task character-consistency --families all
  task: same lead character, 12 storyboard shots, judged blind
  model-a   usable shots 10/12   cost $1.90/shot
  model-b    usable shots 9/12    cost $1.62/shot
  model-c   usable shots 9/12    cost $0.98/shot
  model-d   usable shots 7/12    cost $0.22/shot
$ mms publish --with-harness
prompts, judging rubric, and harness published with the run

The neutrality rule

The Multimodal Society is a nonprofit initiative, and this page is the reason. Our standing policy: no model vendor’s money touches benchmark design, task selection, or results. Ever. Sponsors may fund scholarships, venues, and the compute to reproduce our runs. They may not fund the ruler. Every run ships with its harness so anyone can check our work. The first full report publishes this September and replaces the example above.

Come argue with the results

Every published run headlines Benchmark Night, monthly from this fall in San Francisco: what the numbers hide, when the cheap model is the right call, and which task runs next. The invite goes out with the report.

Get the first report.

The first measured run publishes here in September. Leave your email. It lands in your inbox the day it drops, with the Benchmark Night invite.