A prompt can say “she smirks, then looks away.” It cannot say when the smirk lands, how long it holds, or what the eyes do on the turn. Performance transfer can, because those decisions are made by an actor and recorded, not sampled. You film a take, the model maps its expression, head motion, and lip articulation onto a character that does not exist, and the acting survives the trip. It is the piece of the traditional craft that generative pipelines dropped first and needed most.
- Performance transfer maps a driving video, a human take shot on a phone or webcam, onto a generated character. Timing, micro-expression, and delivery come from the actor, not from sampling.
- Re-record the take and keep the character: direction becomes iterating on a performance instead of re-rolling a lottery.
- The craft is in the driving video: even frontal light, neutral background, face large in frame, expression slightly exaggerated because transfer attenuates.
- Never drive a character with footage of someone who has not consented. The driving take is a performance with rights attached, and the rules now say so explicitly.
What performance transfer does
The model receives two inputs: a character, defined by a reference image or video, and a driving video of a person acting. It extracts the performance signal from the driving video, facial expression, head pose, lip movement for dialogue, and on newer systems body and hand motion, and re-renders that signal on the character. The output is the character giving the actor’s performance: the pauses where the actor paused, the eyeline shifts where the actor shifted.
Contrast that with prompting. A text description of a performance is a request the model interprets once per generation, differently every time. A driving video is a specification. The difference matters most exactly where acting matters most: comic timing, a beat of hesitation before a line, an expression that changes mid-sentence. Those are un-promptable and trivially actable.
What tools do this in 2026?
| Tool | Input | What transfers |
|---|---|---|
| Runway Act-Two | Driving video + character image or video | Face, head, lip sync; body and hand tracking beyond its Act-One predecessor. A character video keeps its scene lighting and body context. |
| Kling Motion Control | Character image + reference video | Full-body motion, expression, lip sync, hand interactions; the strongest full-body option, with takes up to about 30 seconds. |
| Hedra | Portrait image + audio | The adjacent path: performance lives in the voice take, and the model generates phoneme-level lip sync and micro-expression from it. Strong on stylized characters. |
| LivePortrait and successors | Portrait + driving video | Expression and head motion, face only. Open source, free, runs locally, ComfyUI integrations. The proving ground before paying per take. |
Plan gating and pricing on the hosted tools move too fast to print; check the current pages before committing a production. The split that stays stable: Act-Two and Kling are the production-grade paths, audio-driven tools trade control of the face for control of the voice, and the LivePortrait lineage is where you learn the failure modes for free.
How do you shoot a good driving performance?
The driving video is a capture setup pretending to be a selfie, and the setup rules exist because each one protects part of the signal.
- Even, frontal light. Hard side light reads as geometry to the tracker and warps the transfer.
- Neutral background, no motion behind you. The model should spend its attention on the face.
- Face large in frame, camera at eye height. Resolution on the face is resolution on the performance.
- Lock the phone. Camera shake transfers as head motion the actor never gave.
- Stay near frontal. Tracking degrades fast past 45 degrees; save profiles for shots that need them and expect to pay in quality.
- Exaggerate about ten percent. Transfer attenuates expression, so a take that feels slightly theatrical lands as natural. Calibrate per tool with a test take.
Then direct it like any other take. Multiple readings, different tempos, one variable at a time between takes. The generations are deterministic on your side of the equation: the same driving take plus the same character gives you the same performance back, which is what makes this direction instead of gambling.
Match the design to the actor
Transfer quality is partly decided before anyone acts, in the gap between the actor’s face structure and the character’s. Similar proportions, similar head shape: clean transfer. A long narrow face driving a round wide one: warping and bleed. When you control the character design, design toward whoever will drive it. When you control casting, cast toward the design. Keeping the character itself stable from shot to shot is the adjacent discipline, covered in the character consistency guide: the reference image you feed the transfer tool should come from the same reference set that anchors every other shot of that character.
For multi-character scenes, run one pass per character and composite. Research systems are chasing native multi-character transfer, but in production today each performance gets its own take, its own pass, and its own review.
Where performance transfer breaks
- Identity bleed: the actor’s features leak into the character, and the character starts looking like the actor in costume. Worst when face structures diverge; fix with design-to-actor matching and stronger character references.
- Expression clamping: extremes flatten. A full scream comes back as mild distress. Exaggerate the take, and when the beat needs a true extreme, cut to it rather than through it.
- Angle breakdown: most systems are trained near-frontal. Past 45 degrees the tracking, and with it the face, comes apart.
- Face-body decoupling: the face performs, the body idles. Newer systems track body and hands, imperfectly; block the shot so the body has little story to tell, or use a full-body transfer path.
The consent line
A driving performance is a performance. The person who gave it has rights in it, and the industry has spent two years writing that down. SAG-AFTRA’s current agreements require explicit, use-specific consent before a digital replica of a performer is created or used, and federal likeness legislation is moving the same direction. The ethics are simpler than the law: drive characters with your own performance, or with a performer’s written consent naming the project and the use. Found footage of a stranger, a public figure, or an actor who consented to a different project is the bright line.
Performance transfer also puts actors back in a pipeline that had briefly written them out. The phone replaces the mocap stage, not the performer.
Questions creatives ask
What is AI performance transfer? Performance transfer maps a human performance, usually a phone or webcam clip of an actor, onto an AI-generated character. The model extracts expression, head pose, lip articulation, and increasingly body and hand motion from the driving video and re-renders them on a character defined by a reference image. A prompt describes a performance; a driving video specifies one, frame by frame.
Can I film the driving performance on my phone? Yes. A phone or webcam clip is the intended input: even frontal lighting, a neutral background, the face large in frame, minimal camera shake, and the actor mostly facing the lens. No mocap suit, tracking markers, or studio. The craft moves from equipment to acting: a clean, slightly exaggerated take transfers better than a technically perfect flat one.
Why does the character look like the actor instead of the design? That is identity bleed: the actor’s facial features leaking through the transfer into the character. It is worst when the actor’s face structure and the character’s diverge sharply. Reduce it by casting or designing for similar proportions, keeping the driving take near-frontal, and regenerating with a stronger character reference.
Do I need permission to use someone’s video as a driving performance? Yes. A driving performance is a performance, and the person giving it has rights in it. Union agreements now require explicit, use-specific consent for digital replicas, and likeness law is moving the same direction. Use your own performance, or a performer’s with written consent that names the use. Found footage of someone else is the bright line: do not drive a character with it.
