Generative AI models, compared.
Video, image, audio, and writing models for creative professionals, judged on the five axes that decide whether work ships: quality, control, cost, speed, and rights. Verified August 2026; rankings and prices move monthly, so we link live leaderboards instead of freezing numbers. This guide is the fast orientation. The measurement is our open, vendor-neutral benchmark of creative AI, first public report September 2026.
How do you compare generative AI models?
Compare on five axes: quality, control, cost, speed, and rights. Arena Elo measures single-shot preference between anonymous outputs; production depends on character consistency, reference conditioning, editability, and whether the license lets the footage ship. A model can top the arena and still lose your project on any of the other four.
That is why this page links leaderboards instead of quoting scores. Elo numbers from last quarter are trivia. For how to read a benchmark without being fooled by it, see our guide to creative AI evals; unfamiliar terms are in the glossary.
Which AI video model is best in 2026?
Google's Gemini Omni Flash currently tops the Artificial Analysis text-to-video arena, with MiniMax Hailuo second and ByteDance Seedance in the top five. Runway Gen-4.5 wins on the toolchain around the model, Kling 3.0 on price for quality, and Luma Ray3 on finishing-grade output. No single model wins every axis.
The headline change of 2026: OpenAI exited consumer video. The Sora app shut down on April 26, 2026, and the Sora 2 API sunsets on September 24, 2026 (OpenAI's notice). The lesson for creatives: platform risk is a first-class production concern. A model can be excellent and still disappear under you.
| Model | Latest | Best at | Access |
|---|---|---|---|
| Gemini Omni Flash | I/O 2026; API June 2026 | Conversational multi-turn video editing; current arena leader; Veo's successor in Google's user-facing products | API at a reported $0.10 per second; Gemini apps |
| Veo 3.1 | Oct 2025 | Production workhorse: 8-second clips, up to 4K, native 48kHz audio | Google API and tools; resold in Runway |
| Runway Gen-4.5 | Dec 2025; API Feb 2026 | 1080p with integrated audio; the suite is the moat: Aleph video-to-video, Act-Two performance capture, References | Runway app and API |
| Kling 3.0 | Feb 2026 | Native 4K, 3 to 15 second clips, lip-synced dialogue in 5 languages, multi-shot generation; strong motion physics, strong price-performance | Kling app and API; resold in Runway |
| Seedance 2.0 / 2.5 | Feb 2026; 2.5 June 2026 (reported) | Up to 30-second native clips, up to 50 reference inputs; the current ceiling for clip length; top-5 arena standing | ByteDance API; partner apps |
| Hailuo 2.3 / H3 | H3 preview July 2026 | #2 on the arena; 1080p at a reported $0.08 per second; H3 preview: 15 seconds, 2K, native audio | MiniMax app and API |
| Luma Ray3 | 2026 line | The only model line with native 16-bit HDR and ACES EXR export; built for color and finishing pipelines | Dream Machine and API |
| LTX-2 | Open-sourced Jan 2026 | First open model with synchronized audio and video; native 4K at 50fps, about 10 seconds on the open model and up to 20 on the Pro tier | Open weights; free commercial use under $10M ARR |
| HunyuanVideo 1.5 | Nov 2025 | 8.3B parameters, runs on consumer GPUs; the local-pipeline entry point | Open weights, Apache 2.0 |
| Wan 2.2 | 2025 | Fully open base for fine-tuning and local work | Open weights, Apache 2.0; later Wan versions partially closed |
Two structural shifts matter more than any single row. First, platform and model are decoupling: Runway resells Veo, Kling, and Seedance inside its own app, so the tool you edit in no longer dictates the model you render with. Second, native audio became table stakes in twelve months: Veo, Kling, Seedance, Gemini Omni, and LTX-2 all ship synced sound.
On open weights: LTX-2 is free for commercial use under $10M ARR (license terms on the model repo), HunyuanVideo 1.5 is Apache 2.0, and Wan 2.2 is fully open while later Wan versions are partially closed; Wan's current license status is contested, so check the Hugging Face repo before relying on it. The gap between open and frontier is roughly six to nine months.
Arena rank still is not production. Whether a lead holds the same face across shots is a separate question with its own techniques; see our character consistency primer. And for short vertical work, model choice matters less than pacing and hook structure; see the short-form video guide.
Which AI image model is best in 2026?
GPT Image 2 tops the LMArena text-to-image leaderboard on instruction following, in-image text, and editing through chat; its weakness is a recognizable house style. Midjourney V8.1 remains the aesthetic-quality leader for art-directed work by practitioner consensus. FLUX.2 is the strongest open-weights line.
The structural story: image generation consolidated into multimodal LLMs. Google deprecated the standalone Imagen line in favor of the Gemini-native Nano Banana family, and OpenAI retired DALL·E outright; gpt-image-1 sunsets December 1, 2026. The standalone image model is becoming a specialist category.
| Model | Latest | Best at | Access |
|---|---|---|---|
| GPT Image 2 | 2026 | Instruction following, in-image text, chat editing; arena leader; a recognizable house style is the tradeoff | ChatGPT and OpenAI API |
| Nano Banana 2 / Pro | 2026 | Gemini-native default; Pro tier for text localization and brand consistency | Gemini apps and API |
| Midjourney V8.1 | April 2026 | Aesthetic quality for art-directed work; 2K HD output, workable text rendering; mid-pack on obedience arenas | App only, no API; companies over $1M revenue need Pro or Mega |
| FLUX.2 | Nov 2025 | Best open-weights quality ceiling; the base for LoRA and fine-tune workflows | Klein 4B Apache 2.0; Klein 9B and Dev 32B open weights but non-commercial |
| Ideogram 4 | June 2026 | Typography, posters, logos | App and API; reportedly open weights, license unverified |
| Recraft V4.1 | 2026 | Design work: vector SVG output, brand styles; highest-ranked design-oriented model on LMArena | App and API |
| Stable Diffusion 3.5 | Oct 2024 | Mature fine-tune ecosystem; community energy has moved to FLUX | Open weights; free commercial use under $1M revenue |
One license deserves its own warning. FLUX.2 Dev is open weights and non-commercial: the trap tier for freelancers, since a paid client deliverable breaks its terms. A $999 per month self-host license exists. Klein 9B carries the same non-commercial license; only Klein 4B is Apache 2.0 and safe (check each tier's LICENSE file). Read the license before the model card.
What changed in AI music and voice?
Voice consolidated around ElevenLabs; music went from sued by every label to licensed by the labels in under two years. The cost of licensing was user rights: downloads narrowed or disappeared on the licensed services. The cleanest rights now sit with open weights trained on licensed data.
- ElevenLabs. Eleven v3 covers 70+ languages with multi-speaker output and inline audio tags; Eleven Music v2 handles scoring. The default for narration and character voice.
- Suno v5.5. Label-licensed after its Warner deal in November 2025, weeks behind Udio's UMG deal; UMG and Sony litigation continues. Free-tier downloads are gone and paid tiers are capped.
- Udio. Settled with UMG, disabled downloads, and is relaunching as a walled garden where tracks cannot leave the platform. Currently a non-option for deliverable work.
- Stable Audio 3.0 (May 2026). Three of its four models are open weight and trained entirely on licensed data (Stability AI, which also holds UMG and WMG label partnerships). The cleanest rights story in music AI.
Which LLM is best for writing?
Claude Fable 5 (Anthropic, June 2026) is the pick for prose voice and long-form tone consistency by practitioner consensus, and it leads the writing-focused rankings linked below; Claude Opus 5 (July 2026) delivers most of that at roughly half the price. GPT-5.6 wins on structural work, outlines, and format compliance. Gemini 3.1 Pro wins on price and long context.
Open models earn a place for privacy-sensitive scripts and fine-tuning: Kimi K3 and DeepSeek lead there, and Kimi K3's showing on the EQ-Bench Creative Writing leaderboard is the surprise of the year. Check the live rankings rather than any quoted score, including ours.
Which models are cleanest for commercial use?
Ranked from cleanest to murkiest: Stable Audio 3.0, then outputs of Apache 2.0 models, then closed APIs with commercial terms, then Suno, then Udio. Rights are the axis that decides whether finished work can ship, and they vary more between models than quality does.
- Stable Audio 3.0. Open weights plus licensed training data. Nothing else offers both.
- Apache 2.0 model outputs. LTX-2 (under $10M ARR), HunyuanVideo 1.5, FLUX Klein, Wan 2.2. You control the stack and the license.
- Closed APIs with commercial terms. Google, Runway, Kling. Workable, subject to terms that can change and platforms that can vanish.
- Suno. Label-licensed, but user rights narrowed to pay for it.
- Udio. No export. Not usable for deliverables today.
And the standing trap: FLUX.2 Dev is open weights and non-commercial. Open weights and open license are different claims; verify both.
How should you choose?
Run a two-pass workflow: draft cheap on open or turbo models, finish on frontier. Pick the finishing model by production need, meaning consistency, editability, audio, and rights, not by arena rank. Expect the number-one slot to change monthly; it changed at least four times in the past twelve months. Date everything you rely on, including this page.
That instability is why the society measures instead of ranking by feel: one production-shaped task, every major model family, judged blind, with the harness published alongside the numbers. The first public report lands in September 2026. Until then, the reading list lives in the library.
Learn it in a room.
Membership is free. Masterclasses taught by industry leaders, hackathons where you finish something the same day, and screening nights where the room is honest.