What AI filmmaking actually is.
Written by the nonprofit that benchmarks the models rather than sells them. No demo reel, no hype, and no pretending the hard parts are solved.
What is AI filmmaking?
AI filmmaking is the practice of making motion pictures where generative models produce or transform the images, sound, or performance, under the direction of a filmmaker who still owns the story and the cut.
The models are tools in a pipeline, not the pipeline. A director still blocks a scene, an editor still decides what the audience sees and for how long, and someone still has to know why a shot is not working. What changes is who can execute: a shot that needed a crew, a location, and a budget can now be attempted by one person in an afternoon.
What does the workflow look like?
Shot by shot, and much closer to animation than to live action. A typical production runs in five passes.
- Development. Script, treatment, and the look. Models help here as a sketchpad, not an author.
- Previsualisation. Stills first. Lock the characters, the palette, and the framing before a single second of video is generated, because video inherits every inconsistency the stills carry.
- Generation. Shots produced one at a time, usually many attempts per usable second. Reference images, camera direction, and seed discipline do most of the work.
- Post. Grade, upscale, patch, and stitch. Conventional tools, conventional craft. This is where most of the amateur work falls apart and most of the good work is saved.
- Sound. Dialogue, foley, score. Sound carries far more of a film's believability than the picture does, and it is the pass beginners skip.
What still breaks?
Four things, consistently, and anyone selling you a course that says otherwise has not shipped a film.
- Character consistency. Holding one face, one wardrobe, and one build across a dozen shots is still the hardest problem in the medium.
- Continuity between shots. Light direction, prop placement, and screen geography drift, and an audience reads the drift as a mistake long before it can name it.
- Physical plausibility in motion. Hands, contact, weight, and anything that has to collide.
- Lip sync and performance. Passable in a single close-up, still fragile across a scene.
Which model handles which of these best changes every few months, which is why we run an open, vendor-neutral benchmark instead of publishing a favourite.
What does it not replace?
Taste, story, and direction. Generation collapsed the cost of producing an image; it did nothing to the cost of knowing which image is worth producing.
The work that survives is made by people who can hold a scene in their head, recognise when a take is dead, and cut. That skill was scarce before these tools and it is scarcer now, because the volume of competent-looking footage went up and the volume of good films did not.
How do you learn it?
By making something and showing it to people who will tell you the truth about it. Reading about the workflow gets you to the first shot; a room gets you to the last one.
That is what the society is for. Masterclasses taught by industry leaders, hackathons where you finish something the same day, and screening nights where the room is honest. Membership is free.