Your CapCut timeline already holds five Gen-4.5 I2V clips for a six-shot short, trimmed to the temp track and roughly graded, yet the lead’s jaw narrows between the alley wide and the diner close-up, her hairline climbs by the rooftop shot, and the olive field jacket drifts to khaki and loses a pocket, because every first frame was regenerated as a fresh Midjourney or Flux still instead of coming from one tagged Gen-4 References image.
Viewers will read that as five actors, and no cut order fixes it.
Viewers will read that as five actors, and no cut order fixes it.
Five cuts on the timeline; the short still casts five different leads
The brief is a 50-second spec spot: six shots of one woman walking out of a night shift into a sunrise, from a rain-wet alley wide to a diner close-up, a rooftop wide, and a final push-in on her face. Shot list locked, music licensed, mood board approved.
The problem shows up at the cut points. Scrub from shot two to shot three and the face changes shape; by shot five the hair parting has switched sides, and the jacket’s collar and pocket count shift every time. Each clip looks fine alone. AI video character consistency is judged at the splice, and that is where this short falls apart.
The problem shows up at the cut points. Scrub from shot two to shot three and the face changes shape; by shot five the hair parting has switched sides, and the jacket’s collar and pocket count shift every time. Each clip looks fine alone. AI video character consistency is judged at the splice, and that is where this short falls apart.
Fresh stills per shot plus wardrobe and face restated in the I2V prompt
The usual pipeline: for each shot, roll a new Midjourney or Flux still from “woman, late 20s, short dark bob, olive field jacket,†pick the best of four, drop it into Runway as the I2V input, then write a video prompt repeating that description plus the action.
Two things go wrong. First, six text-to-image rolls give six interpretations of “short dark bob.†Text carries a category, not an identity. Second, restating appearance in the I2V prompt puts the text in competition with the image. Runway’s Image to Video Prompting Guide is explicit that the input image acts as the first frame and supplies composition, subject, lighting, and style, while the text prompt should describe what happens: motion, camera work, timing (https://help.runwayml.com/hc/en-us/articles/48324313115155-Image-to-Video-Prompting-Guide). When you describe the face again, prompt adherence starts pulling toward your words and away from the pixels you just supplied. The guide does allow visual description for new elements or dramatic changes, but a lead who should look identical is neither.
The drift is baked in twice: once when each still is invented, again when the motion prompt asks the model to reinterpret the character.
Two things go wrong. First, six text-to-image rolls give six interpretations of “short dark bob.†Text carries a category, not an identity. Second, restating appearance in the I2V prompt puts the text in competition with the image. Runway’s Image to Video Prompting Guide is explicit that the input image acts as the first frame and supplies composition, subject, lighting, and style, while the text prompt should describe what happens: motion, camera work, timing (https://help.runwayml.com/hc/en-us/articles/48324313115155-Image-to-Video-Prompting-Guide). When you describe the face again, prompt adherence starts pulling toward your words and away from the pixels you just supplied. The guide does allow visual description for new elements or dramatic changes, but a lead who should look identical is neither.
The drift is baked in twice: once when each still is invented, again when the motion prompt asks the model to reinterpret the character.
Turbo iteration without a tagged Gen-4 References character as system of record
The second habit is volume. Turbo is cheap enough that ten takes per shot feels like diligence. Runway’s older Gen-4 Video documentation still lists Turbo at 5 credits per second against 12 for Gen-4, and recommends Turbo for exploration before committing to Gen-4 if needed (https://help.runwayml.com/hc/en-us/articles/37327109429011-Creating-with-Gen-4-Video). That page predates Gen-4.5, but the logic holds: Turbo is an iteration tool.
If the first frame is a different face every time, ten Turbo takes give you ten good motion variations of the wrong actor. Credit burn goes up and continuity does not move, because nothing in the pipeline defines who the lead is. CapCut or Premiere is not that system of record either; the NLE assembles whatever you feed it.
A system of record means one approved image of the lead, saved under a name, that every shot’s first frame is derived from. In Runway, that is what Gen-4 Image References is built for.
If the first frame is a different face every time, ten Turbo takes give you ten good motion variations of the wrong actor. Credit burn goes up and continuity does not move, because nothing in the pipeline defines who the lead is. CapCut or Premiere is not that system of record either; the NLE assembles whatever you feed it.
A system of record means one approved image of the lead, saved under a name, that every shot’s first frame is derived from. In Runway, that is what Gen-4 Image References is built for.
Gen-4 Image References: up to three active refs, @name prompts, and tag-to-save
According to Runway’s help article on Gen-4 Image References (https://help.runwayml.com/hc/en-us/articles/40042718905875-Creating-with-Gen-4-Image-References), References lives in Tool mode under the Image tab, and Runway positions it for consistent characters across different lighting conditions, locations, and treatments from a single reference image. The details that matter for a six-shot short:
Up to three active References per generation. In practice: the lead, the location, and a wardrobe detail or second character. Plan shots around that budget.
Tag to save. An upload only lives in the current session by default; hover, select tag to save, and name it to persist it across sessions. Untagged refs are how the lead quietly gets replaced on day two.
@name in prompts. Typing @ autocompletes a saved Reference, so a prompt reads “@mara sits at the diner counter, morning light through the window†rather than a paragraph about her face. The reference carries identity; the text carries the scene.
Character image recommendations. Runway recommends natural, even lighting, moderate quality, and a neutral expression, describing it as a blank canvas that simplifies transformation. So your hero portrait is not the moody, rim-lit Midjourney frame you love; it is a flat, evenly lit, neutral-face still with the actual jacket visible. Generate it upstream in Midjourney or Flux if you like, approve it once, then tag it.
From there, generate each shot’s still with References, not with fresh text-to-image rolls. When a still is approved, the camera icon on the output loads it straight into the video model.
Up to three active References per generation. In practice: the lead, the location, and a wardrobe detail or second character. Plan shots around that budget.
Tag to save. An upload only lives in the current session by default; hover, select tag to save, and name it to persist it across sessions. Untagged refs are how the lead quietly gets replaced on day two.
@name in prompts. Typing @ autocompletes a saved Reference, so a prompt reads “@mara sits at the diner counter, morning light through the window†rather than a paragraph about her face. The reference carries identity; the text carries the scene.
Character image recommendations. Runway recommends natural, even lighting, moderate quality, and a neutral expression, describing it as a blank canvas that simplifies transformation. So your hero portrait is not the moody, rim-lit Midjourney frame you love; it is a flat, evenly lit, neutral-face still with the actual jacket visible. Generate it upstream in Midjourney or Flux if you like, approve it once, then tag it.
From there, generate each shot’s still with References, not with fresh text-to-image rolls. When a still is approved, the camera icon on the output loads it straight into the video model.
Gen-4.5 and Gen-4 I2V: image is first frame; text focuses on motion
Gen-4.5 is Runway’s current model for I2V quality. Its help page lists 12 credits per second, durations from 2 to 10 seconds, and an image upload as the input for Image to Video; it also states that Image to Video prompts should focus on describing the motion of the scene, while Text to Video describes visuals and motion (https://help.runwayml.com/hc/en-us/articles/46974685288467-Creating-with-Gen-4-5).
The I2V prompting guide sharpens that. Prompts should focus almost exclusively on motion: subject action, environmental motion, camera motion, timing, direction, and speed. Refer to characters with general language such as “the subject†or “the woman†instead of re-describing them. The guide also warns that artifacts in the input image, like blurry hands or faces, can intensify once the image becomes video, which is one more reason to fix the still before you spend video credits.
A motion-only prompt for shot three looks like this: “The camera slowly pushes in as the subject lifts the cup and glances toward the window. Locked-off framing, subtle steam.†No hair, no jacket, no age. The first frame already said all of that.
For longer beats, the guide describes chaining: scrub to the end of a finished clip, select Use, then Use current frame, and generate the next clip from that last frame. Trim the shared frame when you join the clips in your editor. Useful for a continuous walk past 10 seconds in one location.
The I2V prompting guide sharpens that. Prompts should focus almost exclusively on motion: subject action, environmental motion, camera motion, timing, direction, and speed. Refer to characters with general language such as “the subject†or “the woman†instead of re-describing them. The guide also warns that artifacts in the input image, like blurry hands or faces, can intensify once the image becomes video, which is one more reason to fix the still before you spend video credits.
A motion-only prompt for shot three looks like this: “The camera slowly pushes in as the subject lifts the cup and glances toward the window. Locked-off framing, subtle steam.†No hair, no jacket, no age. The first frame already said all of that.
For longer beats, the guide describes chaining: scrub to the end of a finished clip, select Use, then Use current frame, and generate the next clip from that last frame. Trim the shared frame when you join the clips in your editor. Useful for a continuous walk past 10 seconds in one location.
Lock checklist: References still to I2V first frame to motion-only prompt to coverage pass
Run this per short, not per shot.
Lead still. Generate or select one evenly lit, neutral-expression image of the lead with final wardrobe visible. Tag it in Gen-4 References with a short name, such as mara.
Shot stills. For each of the six shots, generate the first-frame still with @mara plus at most two more refs. Reject any still where jaw, hairline, or jacket drifts.
I2V first frame. Load each approved still into Gen-4.5 via the camera icon or drag-and-drop. Do not swap in a prettier Midjourney frame at this stage.
Motion-only prompt. Describe camera, action, and timing with “the subject.†Use Turbo, where available, for cheap motion tests on the same first frame, then render the keeper on Gen-4.5.
Coverage pass. Scrub every cut point on the CapCut or NLE timeline at full screen, checking face shape, hairline, wardrobe, and skin tone across adjacent shots. Any failure goes back to the still, not another video roll.
Batching helps: approve all six stills before touching video, so credit burn only goes to motion.
Lead still. Generate or select one evenly lit, neutral-expression image of the lead with final wardrobe visible. Tag it in Gen-4 References with a short name, such as mara.
Shot stills. For each of the six shots, generate the first-frame still with @mara plus at most two more refs. Reject any still where jaw, hairline, or jacket drifts.
I2V first frame. Load each approved still into Gen-4.5 via the camera icon or drag-and-drop. Do not swap in a prettier Midjourney frame at this stage.
Motion-only prompt. Describe camera, action, and timing with “the subject.†Use Turbo, where available, for cheap motion tests on the same first frame, then render the keeper on Gen-4.5.
Coverage pass. Scrub every cut point on the CapCut or NLE timeline at full screen, checking face shape, hairline, wardrobe, and skin tone across adjacent shots. Any failure goes back to the still, not another video roll.
Batching helps: approve all six stills before touching video, so credit burn only goes to motion.
Where persistent characters, universes, video, and Director assemble
References solve the lock inside Runway. Most AI filmmakers also work across Midjourney or Flux, an NLE, and social platforms, and the lead tends to get redefined at every handoff.
That handoff is where Havincy sits. It covers image generation and editing, persistent characters and universes that stay defined across projects, video generation, and talking avatars, with a Director that runs the work as a loop: Idea, Create, Produce, Publish, Measure, Improve. Social publishing and analytics sit in the same workspace, so the version that ships is the one you measure. The useful part for a filmmaker is that the character and world are stored as reusable assets rather than rebuilt from a prompt each time: the discipline of a tagged References still, applied across the project.
That handoff is where Havincy sits. It covers image generation and editing, persistent characters and universes that stay defined across projects, video generation, and talking avatars, with a Director that runs the work as a loop: Idea, Create, Produce, Publish, Measure, Improve. Social publishing and analytics sit in the same workspace, so the version that ships is the one you measure. The useful part for a filmmaker is that the character and world are stored as reusable assets rather than rebuilt from a prompt each time: the discipline of a tagged References still, applied across the project.
FAQ: Gen-4.5 every take, Turbo for motion tests, and last-frame chaining
Is switching every take to Gen-4.5 the fix?
No. Gen-4.5 improves motion quality and prompt adherence, but it follows the first frame you give it. Six different first frames still produce six different leads, now at 12 credits per second. Fix identity in the stills, then upgrade the renders.
When to use Turbo for motion tests?
When the first frame is already approved and you are testing camera moves, pacing, or action on that exact still. Where your account still offers Gen-4 Turbo, it is the lower-cost exploration option. It is not the place to search for a face.
Last-frame chaining vs new References still per location?
Chain from the last frame when the shot continues in the same space and light, like a walk that outlasts 10 seconds. Generate a new References still with @mara when the location, lighting, or framing changes, because a chained frame inherits any drift or artifacts from the previous clip, and those compound across cuts.
No. Gen-4.5 improves motion quality and prompt adherence, but it follows the first frame you give it. Six different first frames still produce six different leads, now at 12 credits per second. Fix identity in the stills, then upgrade the renders.
When to use Turbo for motion tests?
When the first frame is already approved and you are testing camera moves, pacing, or action on that exact still. Where your account still offers Gen-4 Turbo, it is the lower-cost exploration option. It is not the place to search for a face.
Last-frame chaining vs new References still per location?
Chain from the last frame when the shot continues in the same space and light, like a walk that outlasts 10 seconds. Generate a new References still with @mara when the location, lighting, or framing changes, because a chained frame inherits any drift or artifacts from the previous clip, and those compound across cuts.