Blog

More faceless Shorts still reset your series identity every upload day

Share
  • https://havincy.com/blog/faceless-youtube-shorts-series-identity
More faceless Shorts still reset your series identity every upload day
By Thursday night the weekly faceless YouTube Shorts folder already holds seven vertical stills, seven ElevenLabs voice takes, and a CapCut project waiting on export, while the host face, the wardrobe color, and the VO tone still drift from clip to clip because each episode re-rolled a fresh Midjourney host instead of a pinned Omni Reference and each take left stability and similarity_boost to whatever the playground last remembered.
Viewers never learn the show.

Seven Shorts rendered; the channel still sounds and looks like seven pilots



Scripts, stills, and voice files are in. CapCut already hits the hook, the proof beat, and the end card. The folder looks finished.

Back to back, series identity falls apart. One host has a narrow jaw and a navy overshirt; three cuts later the face is rounder and the hoodie is grey. One read stays calm; the next jumps brighter though the script did not change. Thumbnails do not share a motif, so each Short could be a pilot. Coverage exists. The identity pack was never locked before batching.


Host still regenerated without a locked Omni Reference weight



The image break is a character ref with no documented weight. Midjourney Omni Reference on V7 puts a character or object from one reference image into a generation, via --oref on Discord or the Omni Reference slot on the web. --ow runs from 1 to 1000, default 100, and docs say keep it below 400 unless stylize or --exp is very high. It costs twice the GPU time of regular V7, and it is not compatible with Fast Mode, Draft Mode, Conversational Mode, or --q 4. Results are not compatible with Vary Region, Pan, or Zoom Out until refs, --oref, and --ow are removed in the Editor (https://docs.midjourney.com/hc/en-us/articles/36285124473997-Omni-Reference).

A fresh prompt feels faster, so prompt adherence follows the words. Wardrobe, age, and lighting rewrite every upload day. On V8.x the same article points to the Edit Model, up to four reference images, replacing Omni Reference and Character Reference on that path. Spend the four as host, wardrobe, lighting, and spare. Text-only regens are how seven pilots get born.


Voice breaks the same way when settings sit outside the pack. ElevenLabs stability defaults to 0.5: lower is broader emotion, higher is steadier and more monotone. similarity_boost defaults to 0.75 (Clarity plus Similarity Enhancement), which is adherence to the original voice. Playground sessions often start near stability 50, similarity 75, and style 0 (https://elevenlabs.io/docs/api-reference/voices/settings/get.mdx and https://elevenlabs.io/docs/eleven-creative/playground/text-to-speech). Re-rolling those numbers, or swapping voice_id because one take felt flat, makes seven narrators.


Thumbnail motif reinvented per episode so the shelf never reads as one series


The thumbnail is treated as a one-off: new type, new crop, new accent. The frame can look sharp. The shelf cannot be read as one series.


A motif is what you stop reinventing. Same crop of the locked host, same title block, same accent, same end card. CapCut pastes that layout. It does not decide it. Without a motif still in the pack, every upload day invents a shelf, and the next image pass has nothing visual to obey besides the script.


Identity pack before batch: host still, voice settings, motif, rejected drift


Lock this before the seven-cut project opens.


Host still: one approved image on --oref or the web Omni Reference slot, plus the --ow you will reuse (default 100, under 400 unless stylize or --exp is already very high). On Edit Model, name which of the four slots is the host. Near-duplicate fillers fight prompt adherence.


Voice settings: voice_id, stability, similarity_boost, and style as numbers. House start is stability near 50, similarity near 75, style 0, on the API defaults of 0.5 and 0.75. Change them only to change the show.

Motif still: one frame for crop, type block, and accent. Every Short crops from that system.

Rejected drift: two or three killed stills and one killed take, each with a reason (face widened, hoodie shifted, stability dropped and the hook got jumpy). That pile shows the coverage pass what off series looks like.


Coverage pass on the seven-cut timeline


One pass before export. Continuity, not a prettier frame.


Face and outfit match the Omni Reference still at the documented weight. No text-only regen. Same voice_id, stability, similarity_boost, and style. No orphan MP3. Hook frame and first spoken beat belong to this series. Crop, type block, and accent match the motif still.


Fails return to Create with the pack attached. A new CapCut template does not repair a host who was never the same person. Assembly hides a jump cut. It is not the system of record.

Where characters, voice, and multi-format sit in the loop



Idea, Create, Produce, Publish, Measure, Improve. Idea is the angle and the seven scripts. Create is character refs, Omni Reference or Edit Model stills, voice settings, and the motif. Produce is CapCut: cuts, captions, loudness, export. Publish is Shorts, TikTok, and Reels. Measure asks whether shelf and voice held as one series. Improve writes the next constraint into the pack, not a new folder name.

Havincy can hold that loop together: image generation and edit, persistent characters and universes so the host is not a fresh roll each upload day, video, talking avatars when the host needs a speaking face, Director from idea through produce, and social distribution plus analytics so Measure is not a screenshot habit. It does not replace --oref, --ow, stability, or similarity_boost. CapCut stays assembly. The pack, kept with the characters and the universe, is the record.

FAQ — Is more B-roll the fix?


No. B-roll covers time. It does not lock host, voice settings, or motif. More cutaways on open character refs still look and sound like seven pilots. Add B-roll after the pack passes coverage, not instead of it.

FAQ — When to raise Omni weight versus rewrite the text prompt?



Raise --ow when the host is in frame but prompt adherence to the reference is weak, and stay under 400 unless stylize or --exp is very high. Default is 100 on a 1 to 1000 scale, and the feature already costs 2× GPU time versus regular V7, so weight is not a free first move (https://docs.midjourney.com/hc/en-us/articles/36285124473997-Omni-Reference). Rewrite the text when drift is wardrobe, location, or action the words requested, or when high weight fights a prompt that describes someone else. On V8.x Edit Model, fix which of the four refs is wrong before chasing a weight that path may not expose the same way.

FAQ — CapCut templates versus the identity pack as system of record?



Templates store cuts, caption styles, and export presets. They do not record the host, the voice_id, the stability and similarity_boost pair, or the motif still. If the template is the only lock, the next batch re-rolls face and VO. The identity pack is the record. The template consumes it.


← All articles