Tutorial

Scoring with AI music: cut to picture, clear the rights.

Glass metronome with a ribbon of blue staff lines and notes flowing past

AI music tools were trained on songs, and it shows: ask for tension and you get a chorus. Scoring a film with them is two separate skills. The first is craft: steering song-shaped models into underscore, then editing what they give you until it serves the picture. The second is paperwork: the licensing ground under these tools has shifted repeatedly within single production cycles, and the difference between a cleared cue and a liability is a subscription tier and a dated receipt.

The short version
  • Say underscore, instrumental, and the instrumentation in every prompt. Song-trained models default to verse-chorus pop with vocals.
  • Generate longer than the scene and cut to picture. No model hits a frame-accurate sync point; your editor does.
  • Stems are the line between a toy and a scoring tool: they let the music duck under dialogue and re-enter on the cut.
  • Commercial rights live on paid tiers and attach to tracks generated while subscribed. Keep receipts, keep generation records, and re-check terms the week you deliver.

Start from a temp, not from adjectives

Composers have always started from a temp track: an existing piece cut against the edit to establish mood, tempo, and where the music must move. Keep that step. Cut a temp against your scene, then describe the temp to the model, not the scene: mood, instrumentation, tempo, and what is absent. The prompt that works reads like a cue sheet.

Instrumental film underscore, slow build. Solo cello over sparse felt piano, sustained strings enter halfway, no percussion, no vocals. 70 BPM, minor, patient, unresolved ending.

Instruction-following on negatives is inconsistent everywhere: “no drums until the midpoint” sometimes holds and sometimes does not. The reliable levers are naming the instruments you want, the tempo, and the words instrumental and underscore. A tempo in the prompt also makes the cue cuttable later: music generated at a known BPM aligns to an edit far more gracefully than music at a guessed one.

Which tools fit which job in 2026?

ToolFit for scoring
SunoThe strongest full-track generator, with stem export on paid tiers and an editor for arrangement. Song bias is strong; prompt underscore explicitly.
Stable AudioBuilt for instrumental cues with duration control, plus audio-to-audio for making variations of a cue you already like. The most score-shaped of the tools.
Firefly soundtrack toolsGenerates instrumental music fitted to your timeline length inside the Adobe stack: the exception to cut-to-picture, at the cost of less musical range.
ElevenLabs musicFast full tracks with editing and inpainting, commercial use from low tiers, no stem export. Fine for montage; limiting under dialogue.
UdioIn licensing transition: streaming only, downloads and stems disabled while its licensed relaunch is pending. Check status before planning a production on it.

Details churn monthly, tier gating most of all. What stays stable is the shape of the decision: stems and instrumental control decide whether a tool can score a scene, and the terms page decides whether the score can ship.

Cut to picture, don’t generate to length

The instinct is to ask for a 47-second cue for a 47-second scene. Resist it. Generation to an exact length trades musical quality for arithmetic, and it still will not put the swell where the door opens. The workflow that holds up is the library-music workflow, which editors have run for decades:

  1. Generate two to three minutes of cue in the scene’s mood, longer than you need.
  2. Find the musical moment that matches your key beat, and slip the whole cue until they align.
  3. Cut on phrase boundaries, never mid-phrase. Music edits hide where sentences end.
  4. Crossfade the joins, and cover the seams with the dialogue and effects layers above the music.
  5. End on resolution or cut before one: a cue that resolves under your unresolved scene argues with it.
CUT TO PICTURE Generate long Slip to the beat Cut on phrases Crossfade HOPE IS NOT A SYNC POINT Generate at 0:47 Hope the swell lands The swell lands somewhere. Your moment is elsewhere.
The library-music workflow, unchanged since long before generation: the cue bends to the edit, because the edit cannot bend to the cue. Timeline-fitted generators are the one exception, and they trade range for the convenience.

Stems multiply what the edit can do. With separated instrument tracks you can drop everything but the piano under a dialogue line and bring the strings back on the cut, turning one generation into a dynamic score. Without stems, your only volume automation is against a finished master, which is why stem export should decide the tool before sound quality does. One more mastering note: generated tracks arrive loud, limited like streaming releases, with little dynamic headroom. Pull the cue down and let the mix breathe; left at release loudness, the cue masks the dialogue it sits under.

Clear the rights before you ship

Three facts decide what you can do with a generated cue, and all three deserve a check the week you deliver, not the week you generate.

  • Commercial use is a paid feature. Free tiers across the major tools are non-commercial. Rights attach to tracks generated while subscribed, so keep dated receipts and generation logs with the project files, the same way a production keeps music licenses.
  • The ground is still moving. The major labels settled with some generators and are still litigating with others; licensed relaunches, download gates, and tier changes have all shipped mid-production-cycle.
  • Purely AI-generated music gets no US copyright. You can use the cue, but you cannot register it or stop reuse, which leaves a gap in chain-of-title for anything you plan to sell. Treat AI cues like library music, cleared but not owned, and expect some festivals and distributors to ask for AI-use disclosure.

Where AI music still fails for film

  • Hard sync points: no model hits a stinger on frame. The edit does it, or a timeline-fitted tool approximates it.
  • Builds under dialogue: models master everything to full presence, with no concept of ducking. Stems and automation are the fix.
  • Long-form structure: cues past a couple of minutes wander or loop. Generate scene-length thoughts, not act-length ones.
  • Motif continuity: “same theme, new scene” barely exists as a primitive. Audio-to-audio variation is the closest current tool, so plan themes around what it can vary: instrumentation and intensity, not melody-accurate development.

These are the same trade-offs the rest of the pipeline teaches: know what the model decides and what you decide, then keep the decisions that carry meaning on your side of the line. The AI filmmaking guide maps that line across the whole production, and the glossary covers the vocabulary this page leans on.

Questions creatives ask

Can I use AI music in a monetized video or festival film? Only on a paid tier, and only if you can prove it. Free tiers are generally non-commercial across the major tools; commercial rights attach to tracks generated while subscribed. Keep dated receipts and generation records with your project files, and check the specific tool’s current terms before submission, because the licensing landscape has changed repeatedly within single production cycles. Some festivals and distributors now ask for AI-use disclosure.

Is AI-generated music copyrightable? In the US, music generated wholly by a model gets no copyright: you can use it, but you cannot register it or stop anyone else from reusing the same track. That is a chain-of-title gap for anything you plan to sell or license, which argues for treating AI cues the way productions treat library music: cleared for use, never claimed as owned.

How do I stop AI music tools from giving me a pop song with vocals? Say underscore, instrumental, and film score explicitly in every prompt, and name the instrumentation you want. Song-trained models default to verse-chorus pop with vocals, and a mood word alone will not steer them off it. If a vocal still leaks in, regenerate; on tools with stem export, drop the vocal stem instead.

How do I make AI music hit the beats of my edit? Cut the music, not the prompt. No current model hits a frame-accurate sync point on request, so generate longer than the scene, then edit: cut on phrase boundaries, slip the cue so an existing swell lands on your moment, and crossfade. Stems make this dramatically easier, since you can drop and re-enter individual instruments around dialogue and hits.

Learn this beside the people building it.

Membership is free. Masterclasses from industry leaders, hackathons where you finish something the same day, and mentor circles matched to what you want to learn. For engineers and creatives alike, across film, design, image, sound, and story.