Pipeline Hub

ugc-omni

THE photoreal people engine (Kristian Jennings method): real creator frame -> one-shot avatar (GPT Image 2, Nano Banana Pro split arm) -> locked Gemini Omni prompt -> A-roll per line chained on exact first/last frames (Omni 1.1 Flash) -> b-roll + product demo checked against the permanent Dawn ref sheet -> captions from the script. Two modes: cuts (A-roll + b-roll + scene change) and oneshot (one continuous yapper take). Format skills (ai-actor-demo, podcast B, street-interview, reaction-duet B, whistleblower, warehouse W3) hand it a job.json. Use for AI UGC, AI creator / talking-head ads, 'make it look like real UGC', or when an AI person has to show or demo the product. Not for cartoon (cartoon-h3).

Lane: — · not built yet: 1 (see below) · source: /Users/ayden/.openclaw/workspace/skills/ugc-omni

SKILL.mddavidaistar-diff-2026-10-06.mdEvals (0)Learnings (0)
Gates (tell Fish which one you are at)0. Script and storyboard (G0)1. Reference frame (G1, the step everyone misses)2. Avatar (G1)3. Lock the Omni prompt (G2)3b. Continuity and the two formats (Fish 10-06)3c. Several people, off-camera voices, on-screen text (Fish 10-06)4. A-roll (G3)5. B-roll and the product demo (G3)6. Assemble (G4)What we did not take, and why

UGC Omni

Code: scripts/ugc_omni/omni.py (every paid step is a dry run until --spend). Example job: skills/ugc-omni/job.example.json. Sources: Kristian Jennings Parts 1-3 (transcripts + his on-screen prompts and storyboard in output/kj-ugc/), treg make-ugc / ugc-talking-head-video (take checks, hook mining), our 9-30 avatar post-mortem (feedback_ai_avatar_broll_lane).

The one idea: everything comes down to the reference images, not the prompt. A real iPhone frame of a real person already has the noise, pores, micro-expressions and light the video model needs to animate. A generated portrait does not, and that is the waxy, uncanny look. Prompts stay plain.

Time budget (KJ): 50% on the reference image, 25% locking the prompt, 25% generating. If you rush the first half, you are locked in at that quality.

Gates (tell Fish which one you are at)

Gate Fish approves Built with
G0 script + hooks + storyboard (board clean) board
G1 reference frame + avatar refscan, ref, avatar, pick
G2 locked prompt: one line, same take good 3x (this is the motion proof from the 9-30 rules) lock
G3 every A-roll line watched; b-roll frames + product accuracy checked gen, verify, broll
G4 the assembled cut, by eye and ear assemble

No paid step before G0. No gen before G2. Scores never pass a gate on their own: verify numbers are a filter, and eyes plus ears decide.

0. Script and storyboard (G0)

1. Reference frame (G1, the step everyone misses)

2. Avatar (G1)

3. Lock the Omni prompt (G2)

Gemini Omni video on kie: 4/6/8/10 s, 720p and 1080p cost the same ($0.315 / $0.42 / $0.525 / $0.63), up to 7 reference images. Default is 1080p.

His prompt has three parts (verbatim from his screen):

PART 1 set the scene:  Static locked-off shot, UGC iPhone footage of a confident American man in his bedroom. He stays still,
                       mouth moving naturally as he speaks, direct eye contact with camera, no camera movement. He has SUPER
                       high energy as this is for a youtube video. The man says:
PART 2 dialogue:       "so weak follicles wake back up and stay in the growth phase longer instead of shedding early."
PART 3 fixed rules:    Pace: SUPER High energy, confident tone. Fast pacing. Smiling at the camera, he seems very friendly.
                       One take, no jump cuts. Same avatar, same location, same camera framing. Do not recompose the shot

Fix library (stack into prompt.rules; one change per roll, because a rerun can regress something else):

Symptom Add / change
flat, low energy "He has SUPER high energy..." in the scene, plus "Pace: SUPER high energy, confident tone"
too slow "Fast pacing." (and check the bucket, below)
not warm "Smiling at the camera, she seems very friendly."
reframes / zooms / cuts "Same avatar, same location, same camera framing. Do not recompose the shot."
holds at sentence breaks join sentences with commas; "keeps talking straight through the sentence breaks, only quick breaths" (treg)
touches face / points at self say what the hand does: "free hand gestures toward the camera; never touches face or lips" (treg)
wrong word stressed CAPITALISE the word to stress (captions drop the caps automatically)
voice drifts between takes job.voice preset + voice --spend → audio_ids on every take (untested; KJ never needed it)

Duration buckets: - Omni fills the bucket. Between buckets it either cuts the line off or slows it until it "sounds simple". - plan_line picks the smallest bucket that fits at the locked pace and pads with throwaway words. verify cuts right after the last real word, and the padding is never heard. - A line over 10 s errors: split it.

3b. Continuity and the two formats (Fish 10-06)

Model. A-roll defaults to aroll_model: flash (Gemini Omni 1.1 Flash on kie, same price): an EXACT first frame and an optional last frame. - Chaining, per speaker: each line starts on the last frame of the SAME speaker's previous take, so lines join without a jump and alternating speakers (podcast mom/host) never chain off each other. A line with no frame is the main avatar (the format skills' convention). - Hooks: every hook lands on the frame the body starts from, so any hook joins the body cleanly. - A different frame name = a different person, or the same person in a new place (scene2); it starts fresh on that frame. - Hooks land on the body's start frame only when the same person speaks both; otherwise set end_frame per hook. - Constraint: with an exact first frame, kie refuses reference images and voice IDs. The product must already be right in the frames (the gate in §5), and the voice holds through the locked prompt. - Older behaviour: aroll_model: omni is the reference-image model with no exact start.

Two formats in one pipeline:

mode What it is Board rules
cuts KJ UGC: A-roll under every line, b-roll on top, a scene change to a new location ~60% in, cuts between actions/places as §0
oneshot one yapper take, no cuts: one location, no b-roll, chained segments joined on exact frames errors on any b-roll or new frame; the reveal uses end_frame

Long one-takes. For yappers of 5+ lines, anchor_every: 3 lands every 3rd line back on the start frame so chained drift can't compound.

3c. Several people, off-camera voices, on-screen text (Fish 10-06)

Extra people (street-interview moms, podcast host, whistleblower insider) are frames, one per person: - Job entry: "cast": [{"frame": "host", "ref_src": "<real creator video>", "avatar": {genetic, wardrobe, accessories, audio_logic, background}}]. - Steps per person: refscan --frame host, then ref --frame host --pick ..., then avatar --frame host, then pick --name host --file .... Then any line with "frame": "host" is that person. - Every extra person needs their own avatar block with their own genetic change. Without it they'd come out as the main avatar from another photo, so avatar refuses. - Podcast = full-frame cuts between speakers (Fish 10-06), no split screen.

Off-camera voices (interviewer, narrator over b-roll): - Set "vo": true on the line, plus job.vo_voice_id (or the line's voice), an ElevenLabs voice ID. - No default voice: the job must name one, and gen refuses otherwise. - The line needs a broll: visual, since nobody on camera says it. - gen makes the audio and a matching .verify.json, so captions and assembly treat it like any other line. There's no Omni prompt to lock on it.

On-screen text: - overlay (one string, or {hook: text}) shows the hook text at the top from frame 0 until the hook has been spoken. - label (e.g. "Dramatization") is metadata only and is never rendered (Fish 10-06).

Merge note: the format skills (format-*) and the final merge are owned by another session. This skill owns the engine. Format jobs hand over a job.json; extra-person avatars need the cast[].avatar block above.

4. A-roll (G3)

5. B-roll and the product demo (G3)

6. Assemble (G4)

What we did not take, and why

Not built yet (needs Fish's go)

Every line in this skill that mentions something not built, roughly de-duplicated. Some items repeat in different words.