UGC Omni
Code: scripts/ugc_omni/omni.py (every paid step is a dry run until --spend). Example job: skills/ugc-omni/job.example.json.
Sources: Kristian Jennings Parts 1-3 (transcripts + his on-screen prompts and storyboard in output/kj-ugc/), treg make-ugc / ugc-talking-head-video (take checks, hook mining), our 9-30 avatar post-mortem (feedback_ai_avatar_broll_lane).
The one idea: everything comes down to the reference images, not the prompt. A real iPhone frame of a real person already has the noise, pores, micro-expressions and light the video model needs to animate. A generated portrait does not, and that is the waxy, uncanny look. Prompts stay plain.
Time budget (KJ): 50% on the reference image, 25% locking the prompt, 25% generating. If you rush the first half, you are locked in at that quality.
Gates (tell Fish which one you are at)
| Gate | Fish approves | Built with |
|---|---|---|
| G0 | script + hooks + storyboard (board clean) |
board |
| G1 | reference frame + avatar | refscan, ref, avatar, pick |
| G2 | locked prompt: one line, same take good 3x (this is the motion proof from the 9-30 rules) | lock |
| G3 | every A-roll line watched; b-roll frames + product accuracy checked | gen, verify, broll |
| G4 | the assembled cut, by eye and ear | assemble |
No paid step before G0. No gen before G2. Scores never pass a gate on their own: verify numbers are a filter, and eyes plus ears decide.
0. Script and storyboard (G0)
- The script decides the winner; the format can't save a bad script. Run it through the copy doctrine first (dawn-native / direct-response-copywriter).
- Analyze the winner you are iterating on, then write
job.rules: pacing, how present the avatar is, their mic, who introduces the product, b-roll style, "show what they're saying, KISS". The rules come from his HL 112 doc. - One
lines[]entry per spoken line:text,visual(AI UGCorbroll:<id>),frame(avatar/scene2/ ...),product: truewhen the product must be on screen, andhookfor hook-only lines. boardchecks the structure he teaches:- A-roll sits under every line.
- Never ~7 s+ straight on the face; stare long enough and the tells show.
- Never 4+ b-roll lines without a touchpoint back to the avatar.
- The scene changes about 60% in (same avatar, new place). This is the retention hack, and the product lands there.
- The product is actually shown.
- Avatar logic: pick a person who has already solved the problem (underlying authority), not someone still struggling, and who matches the product and angle. Never deepfake a real medical professional.
1. Reference frame (G1, the step everyone misses)
- Where to look: a real creator video from the avatar's platform: younger → TikTok, millennial → Instagram, older → Facebook.
omni.py minedownloads candidates (yt-dlp) and transcribes them. Pinterest is AI slop now; never use a generated image as the reference. - What to judge:
- Framing: your intentional choice (couch, end of bed, car, leaning in). Not right up close: distance buys margin for imperfections.
- Background: neutral but with some personality (a frame, a plant). Remove anything creator-specific (name on the wall, logos).
- Lighting: the most important of the three. Muted, no overexposure. Shine is fine as long as there is still detail in it; blown white patches (forehead, nose, eyes) have no data to animate and turn into glass skin. Too dark beats too light.
- Frame grab: both hands natural and visible, no captions over the face if a clean frame exists.
refscan --src <video>samples 2 fps and ranks face frames by blown face pixels (any channel ≥250, gate 0.3%), face size (≤30% of height) and sharpness, then writesref/refscan.jpg. It agreed with KJ's own labels on his clearly blown examples (1.5-13.6% vs ≤0.11% on his "perfect" ones). It does not see subtle sheen, so the final call is by eye.
2. Avatar (G1)
- Models (Fish 10-06): GPT Image 2 i2i ($0.05, 2K) is the default; Nano Banana Pro ($0.09) runs beside it as a split arm (
image_arms,--arms). KJ prefers NBP ("GPT too perfect"), but in the 10-06 test the GPT2 avatar won on Gemini scores (46 vs 40) and on the A-roll (48 vs 45, NBP voice read as TTS). NBP won the product-demo frame (band accuracy 7 vs 4). Keep both arms until Fish's eyes settle it.lock --frame avatar_nbpanimates the same line from the other arm. - Four decisions go in
job.avatar, one prompt, all at once: 1. Genetic change. It's illegal to deepfake, so it must be impossible for both to be the same person without surgery: eye colour, facial structure, ethnicity, skin. Hair and clothes alone don't count. 2. Background:keep, or add greenery or one point of interest. 3. Accessories: outliers, not the aggregate (one AirPod in, axe pendant, rose-gold watch, a MacBook). 4. Audio logic: why their audio is clean (clip-on mic, earphone mic, a held mic, or sitting close to the phone). - Clothes (Fish 10-06): they don't have to match the reference exactly, but they must look like real clothes. Default
wardrobe: keep("Keep the exact same clothes."), which read far more real in the test. Writing "a faded grey crewneck" produced a blank grey sweater, which only AI characters wear. Any garment you do write must be specific; plain/blank/basic is rejected. - One-shot the prompt. Stacked edits leave feathered AI overlays that show once animated. A bad roll means re-rolling the same prompt, never patching it.
- Prompt style: his real prompts are one or two plain sentences ("Change this avatar to a young latina girl with hair up in a bun, airpods in her ears, wearing a gold necklace and rose gold watch"). The lint rejects >80 words and any realism words (realistic, pores, natural lighting, 4K, detailed, grain...). Asking for realism is "asking a human to act natural."
- About 1 in 10 avatars just won't animate well. Go back and get a new reference; don't fight the video model.
3. Lock the Omni prompt (G2)
Gemini Omni video on kie: 4/6/8/10 s, 720p and 1080p cost the same ($0.315 / $0.42 / $0.525 / $0.63), up to 7 reference images. Default is 1080p.
His prompt has three parts (verbatim from his screen):
PART 1 set the scene: Static locked-off shot, UGC iPhone footage of a confident American man in his bedroom. He stays still,
mouth moving naturally as he speaks, direct eye contact with camera, no camera movement. He has SUPER
high energy as this is for a youtube video. The man says:
PART 2 dialogue: "so weak follicles wake back up and stay in the growth phase longer instead of shedding early."
PART 3 fixed rules: Pace: SUPER High energy, confident tone. Fast pacing. Smiling at the camera, he seems very friendly.
One take, no jump cuts. Same avatar, same location, same camera framing. Do not recompose the shot
- Camera movement first, always: "Static locked-off shot", or "Handheld UGC iPhone shot" if they hold the phone. Leave it unstated and you get drift.
- Capture device: "UGC iPhone footage", or a named camera for a podcast look.
- Lint: checks the camera line, the device, "One take, no jump cuts", and "...says:".
- Lock loop:
1. Run
lock --line <id> --n 3. 2. Judge every take harshly. 3. Write each note INTOprompt.rulesand roll again. 4. Once the same line comes out right 3 times, runlock --approve <take>. 5. That stores a hash plus the measured words per second;genrefuses if the scene/says/rules ever change. - After lock, only the dialogue changes. Energy, register and accent then stay consistent across every line (shown in his Part 3).
Fix library (stack into prompt.rules; one change per roll, because a rerun can regress something else):
| Symptom | Add / change |
|---|---|
| flat, low energy | "He has SUPER high energy..." in the scene, plus "Pace: SUPER high energy, confident tone" |
| too slow | "Fast pacing." (and check the bucket, below) |
| not warm | "Smiling at the camera, she seems very friendly." |
| reframes / zooms / cuts | "Same avatar, same location, same camera framing. Do not recompose the shot." |
| holds at sentence breaks | join sentences with commas; "keeps talking straight through the sentence breaks, only quick breaths" (treg) |
| touches face / points at self | say what the hand does: "free hand gestures toward the camera; never touches face or lips" (treg) |
| wrong word stressed | CAPITALISE the word to stress (captions drop the caps automatically) |
| voice drifts between takes | job.voice preset + voice --spend → audio_ids on every take (untested; KJ never needed it) |
Duration buckets:
- Omni fills the bucket. Between buckets it either cuts the line off or slows it until it "sounds simple".
- plan_line picks the smallest bucket that fits at the locked pace and pads with throwaway words. verify cuts right after the last real word, and the padding is never heard.
- A line over 10 s errors: split it.
3b. Continuity and the two formats (Fish 10-06)
Model. A-roll defaults to aroll_model: flash (Gemini Omni 1.1 Flash on kie, same price): an EXACT first frame and an optional last frame.
- Chaining, per speaker: each line starts on the last frame of the SAME speaker's previous take, so lines join without a jump and alternating speakers (podcast mom/host) never chain off each other. A line with no frame is the main avatar (the format skills' convention).
- Hooks: every hook lands on the frame the body starts from, so any hook joins the body cleanly.
- A different frame name = a different person, or the same person in a new place (scene2); it starts fresh on that frame.
- Hooks land on the body's start frame only when the same person speaks both; otherwise set end_frame per hook.
- Constraint: with an exact first frame, kie refuses reference images and voice IDs. The product must already be right in the frames (the gate in §5), and the voice holds through the locked prompt.
- Older behaviour: aroll_model: omni is the reference-image model with no exact start.
Two formats in one pipeline:
mode |
What it is | Board rules |
|---|---|---|
cuts |
KJ UGC: A-roll under every line, b-roll on top, a scene change to a new location ~60% in, cuts between actions/places | as §0 |
oneshot |
one yapper take, no cuts: one location, no b-roll, chained segments joined on exact frames | errors on any b-roll or new frame; the reveal uses end_frame |
Long one-takes. For yappers of 5+ lines, anchor_every: 3 lands every 3rd line back on the start frame so chained drift can't compound.
3c. Several people, off-camera voices, on-screen text (Fish 10-06)
Extra people (street-interview moms, podcast host, whistleblower insider) are frames, one per person:
- Job entry: "cast": [{"frame": "host", "ref_src": "<real creator video>", "avatar": {genetic, wardrobe, accessories, audio_logic, background}}].
- Steps per person: refscan --frame host, then ref --frame host --pick ..., then avatar --frame host, then pick --name host --file .... Then any line with "frame": "host" is that person.
- Every extra person needs their own avatar block with their own genetic change. Without it they'd come out as the main avatar from another photo, so avatar refuses.
- Podcast = full-frame cuts between speakers (Fish 10-06), no split screen.
Off-camera voices (interviewer, narrator over b-roll):
- Set "vo": true on the line, plus job.vo_voice_id (or the line's voice), an ElevenLabs voice ID.
- No default voice: the job must name one, and gen refuses otherwise.
- The line needs a broll: visual, since nobody on camera says it.
- gen makes the audio and a matching .verify.json, so captions and assembly treat it like any other line. There's no Omni prompt to lock on it.
On-screen text:
- overlay (one string, or {hook: text}) shows the hook text at the top from frame 0 until the hook has been spoken.
- label (e.g. "Dramatization") is metadata only and is never rendered (Fish 10-06).
Merge note: the format skills (format-*) and the final merge are owned by another session. This skill owns the engine. Format jobs hand over a job.json; extra-person avatars need the cast[].avatar block above.
4. A-roll (G3)
genmakes one take per line: locked prompt, the line's frame, and the product photos as extra refs onproduct: truelines.- KJ (agency, optimising for quality and speed) generates A-roll for EVERY line, so removing the b-roll track still gives a full talking ad.
- Budget fallback: clone the voice in ElevenLabs for b-roll-only lines. Lower quality, but legit.
verifyon every take (treg checks; numbers only, you cannot hear):- Scribe transcript vs script ≥0.90.
- Doubled words (stutter).
- Holds ≥0.4 s.
- Pace <2.6 wps.
- Speech running into the last frame (cut off).
- Writes the cut point, a 9-frame strip for the eye check, and
.trim.mp4. - Omni gives every avatar the same set of teeth (KJ). That's known; don't chase it.
- 10-06 test: Omni ran ~2.7 wps; a 10-word line at 4 s got cut off, and at 6 s + throwaway words it was a clean full match. Product shots need a CLOSE framing (wrist filling the frame); in a wide shot the band loses its ribs and display. Lip sync is ~6/10, so judge by eye.
5. B-roll and the product demo (G3)
- Style = the avatar's own: a young creator films aesthetic clips; a 60+ one films plain ones. Shots are acted out a bit awkwardly on a propped phone (phone wedged in dumbbells, making a shake to show "tired").
- Always a different outfit, keeping matching details (same ring, same watch). One outfit everywhere is a top AI tell.
outfitis required. - Prompt: two plain sentences, then trust the model ("The guy sits outside at a cafe drinking a cup of coffee wearing a turtleneck, looking off into the distance, happy, without a care in the world"). No micromanaging extras.
- Product references (permanent, Fish 10-06):
knowledge/brands/dawnbands/product-refs/. 20 real frames from the creator footage, display OFF and ON, plus one-image multi-angle packs (pack_off.jpg,pack_on.jpg). Set"product_refs": ["@dawn"]: every job gets the 6-view ref sheet of the canonical band (simple_on/off.jpg); photoreal/UGC jobs also get the real-photo pack (pack_on/off.jpg);style: cartoon= sheet only (Fish 10-06). Never the PDP renders, never a generated image as a product ref. Display OFF is the default; ON only when the lit time is the point. - Every frame that carries the product is checked before it becomes video:
variant --productscores each roll against the real pack, andpick --productrefuses <7/10 or the wrong display state (prodcheck --file Fchecks any image). Fish 10-06: the reveal start frame was already wrong, so the video could only get worse from there. - Product lines need
prompt.product_action: one fixed sentence for what the hand does with the product, locked separately (lock --line <product line>then--approve). Without it, in the 10-06 test the hand dropped at 0.25 s and the band left the frame at 3 s. - Continuous reveal (Fish 10-06: on the A-roll, mostly continuous):
1. The band is in the shot from the start, so the avatar frame has it lying on the table (
variant --name avatar_band --product, thenpick --name avatar --product). 2. The reveal line keeps the speaker'sframe(so it continues from her previous line) andend_frame: holdinglands it on the checked holding frame (variant --name holding --product, thenpick --name holding --product). 3.product_action= she picks it up and holds it up next to her face. - Display-off wording (Fish 10-06): every display-off product prompt says: "Display off: a plain black ribbed silicone strap of even width all the way around, no visible screen, no screen pod, no raised module, no glossy rectangle; it looks like a plain sports band." Without it, dim, wide or no-person-ref shots turn the band into a generic fitness tracker with a screen pod.
- How the product sits (Fish 10-06): she wears it on her wrist, or holds it the way a person naturally would (resting across her palm, loose in her hand). Never pinched upright between the fingers like a stick; that read as wrong.
- B-roll shows the product in use afterwards (on a wrist, waking up); demo frames go through the same refs and gate, and animate from that exact frame.
6. Assemble (G4)
assemble --hook H1builds three layers: 1. The A-roll track: every trimmed line, back to back. 2. B-roll overlaid on its lines, with A-roll audio underneath. 3. Captions = script words (caps for stress removed), loudnorm -14.- Before Fish: watch it in full, and say what you could not verify (voice likeness, lip-sync).
- Recipe discipline: same references, same models, same steps. "If you change the recipe, you're gonna bake a different cake." Test a new model only as a side arm against a locked baseline.
What we did not take, and why
- treg
portrait-clone(exhaustive locked-JSON portrait prompts): KJ and Fish's 10-04 correction both say the opposite: long prompts that list realism make it worse. Its goal (pin every variable) is met here by the real reference frame plus the prompt lock. - treg Seedance 2.5 talking heads: Fish dropped them 10-04. Omni costs ~$0.06/s vs Seedance $0.27/s.
- treg's "cadence comes from the voice reference": on Omni the cadence comes from the locked prompt. The optional
voicepreset is the closest equivalent.
Not built yet (needs Fish's go)
- 6. Brand-DNA intake (product URL → images/colours/copy) instead of hand-assembling assets per batch. Not built