format-street-interview: one question, many moms, one answer
1. Read first
knowledge/ad-formats/_FORMAT-CORE.md(doctrine, lane limits, gates, launch). Binding. Photoreal people run on ugc-omni's G0-G4 gates (§10 here).references/beat-sheet.md(skeleton, take math, line rules),references/sources.md(Foreplay #9 beat by beat, davidaistar notes, our receipts).skills/ugc-omni/SKILL.md+scripts/ugc_omni/omni.py: THE photoreal production engine (Fish 10-06). This skill hands it a job.json (references/omni-job.example.json, illustrative). A parallel session owns both: reference, never edit.- Memory:
feedback_ai_avatar_broll_lane(the 9-30 rules ugc-omni's gates implement),project_dawn_yapper_2026-10-04(H3 native speech side arm; why ungated band-in-hand failed),feedback_reference_images,feedback_ugc_avatar_wardrobe,corrections-hot.md(10-05: simplest recipe first). - Answer sources (the old
voc-*.mdandphrasing-bank.mdare retired pointer stubs and hold no quotes):knowledge/brands/dawnbands/research/customer-evidence-2026-09-15.md(owned survey S01-S08, public Q01-Q15),output/creative-strategy-cooper-2026-10-05/dawn-stakes.md+stakes-raw-owned.md(stake bank with survey/reddit IDs),research/voc-competitor-reviews-2026-10-02.jsonl(q-IDs, parent rows viaspeaker_role). knowledge/brands/dawnbands/product-truth.mdandclaims-register.md(mechanism line, product facts).
2. When to use / when not
- Use as Dawn's first photoreal test. The REPORT §4 ranks it #5 and calls it the cheapest photoreal entry. Every clip is a different person, so there is no identity to hold across shots. That consistency failure is what killed Side Effects (9-30) and yapper (10-05).
- Evidence status: untested on Dawn. It's a signal only: a scaling AI vox-pop on Foreplay (#9, Skincu lipstick, 54 s) and davidaistar's street-interview tests (CZwfTy7jZcQ, jcc-r5I-SlU). Watch out: our data says testimonial devices are the weakest pairings (statics 1.21-1.31, testimonial songs 1.26; MAP.md §2). This format only works as a chorus of the problem (recognition), not a pile of product reviews.
- No paid step before G0 (Fish approves the question, answers and storyboard). Pitch the simplest recipe (§9) and the ~$2 G1-G2 proof first (Fish reopened photoreal 10-06; ugc-omni is the engine).
- Don't use for recipe-A stories (accuser, test, vindication: see format-spoken-drama), for a named "customer" testimonial, or for a made-up expert. Moms don't handle the band in this format: the band lives in one insert.
3. The format
- 9:16, 30-60 s (target ~45 s). Photoreal, phone-shot look: handheld, daylight, real outdoor locations (school drop-off sidewalk, grocery lot, park path, coffee-shop patio, strip-mall walkway).
- Frame 0: a mom already mid-frame, a handheld mic pushed in from the frame edge, and the question on screen as text while the interviewer asks it off camera.
- 6-8 moms, one take each: 4-8 s, one line, hard cut to the next mom. Nobody appears twice.
- One follow-up question off camera ("What finally worked?"). One mom answers with the root cause and mechanism, and a band insert (gated omni b-roll, or real footage) cuts in over her line.
- Payoff line, then the motion end card (optional,
--end-card; off by default per Fish 10-06) (scripts/cartoon_h3/endcard.py). Captions from the script on every word. - Reference shape: Foreplay
09_street_interview(one woman, one long take, product applied on camera). Ours spreads the talk across many women and keeps the band in one checked insert. "Dramatization" label on-screen or in primary text (Fish 10-06).
4. Why it sells
- Recognition at scale (recipe B, callout). The question is the brand voice talking straight to her. Each answer is her own daily burden coming out of a stranger's mouth, so "it's not just me" lands six times in 30 s. Problem-aware: the question names the symptom (she is the alarm), and answers 1-2 name it again in action.
- Social proof of the problem, not of the product. Moms 1-6 never mention a product. Only the turn mom names the band, so it reads as a discovery and not as a testimonial (which is the losing device).
- Stakes escalate across answers (core §1): counts → her exhaustion → the cost on HER (a truancy letter in her name, he swings at her, she cried in the car). Every stake comes from a sourced line.
- Authority = science + guarantee, not a person (recipe B): the mechanism comes as plain fact from an ordinary mom, then the end card (optional,
--end-card; off by default per Fish 10-06) carries the 60-night guarantee.
5. Beat sheet (full rules in references/beat-sheet.md)
| Beat | % runtime | Job | Must contain |
|---|---|---|---|
| 1 Question + hook answer | 0-12 | Stop the scroll on her | Question names her symptom (on-screen + off-camera voice); answer 1 = the highest-emotion VOC line, starting within 1.5 s |
| 2 Chorus | 12-45 | "Every mom here is the alarm" | 3 moms, counts/times/trips from VOC; no products; generic failed solutions only |
| 3 Stakes | 45-65 | The cost lands on her | 1-2 moms, stake-bank lines (legal / violence / her own breakdown); the mornings caused it |
| 4 Turn | 65-70 | Follow-up question | Interviewer off camera: "So what finally worked?" (or similar, ≤6 words) |
| 5 Answer | 70-88 | Root cause + mechanism + product | ONE mom, ONE take: Fish's ruled lines, product named once; band insert overlaid (broll:band, product-gated or real) |
| 6 Payoff | 88-100 | Independence / calm | One mom, kid-independence result; end card (optional, --end-card; off by default per Fish 10-06) overlays the last 2.5-3 s |
6. Pack spec (ugc-omni job.json, one job per ad)
-
Step 0 — combo plan (feedback loop,
workflow-combo-loop):python3 scripts/combo_loop/plan.py --format format-street-interview --brand dawn→ pick one combo (a paying parent + ONE changed dimension) and paste its block into the pack front matter / job or spec JSON ("combo": {...}; no file →tags.py register --prefix):combo_avatarcombo_anglecombo_povcombo_authoritycombo_stakecombo_stake_oncombo_emotioncombo_root_causecombo_mechanismcombo_payoffcombo_devicecombo_formatcombo_parentcombo_variablecombo_ad_prefix(= the launched Meta ad-name prefix).python3 scripts/combo_loop/tags.py check <pack>must PASS before concept approval; the weekly refresh reads results back by that prefix. -
Start from
references/omni-job.example.json(this format's shape, illustrative lines; schema =skills/ugc-omni/job.example.json). New slug per run. - Core §5 keys (top-level; omni.py ignores extra keys):
format: ugc-omni,format_skill: format-street-interview,brand,created,status,parent(paying Dawn ad + CPA),iteration(ONE variable),title,overlay(the question text),label: "Dramatization",cast. - Engine keys:
style: photoreal,mode: cuts,aroll_model: flash,resolution: 1080p,image_arms: ["gpt2", "sd5"],product_refs: ["@dawn"],product_name,product_truth(core §1 wording + ugc-omni's display-off sentence). Noanchor_every(every mom starts fresh on her own frame). cast: one entry per mom{"frame": "avatar"|"h2"|"h3"|"m2".., "age", "look", "ref_src": "<real video URL>", "location"}. Hook mom H1's frame MUST beavatar(omni.py's default frame andlocktarget); H2/H3 hook moms areh2/h3; body momsm2..m8.lines[]: one per mom:id,text,visual: "AI UGC",frame,hook(H1-H3 on the three answer-1 lines),src(VOC/stake ID; extra key, required here). The answer mom's line isvisual: "broll:band"(her voice runs under the insert).- Every hook line sets
end_frameto its own frame ("end_frame": "h2"on H2-1). Without it, flash lands each hook on the body's first frame, which here is a different mom (omni.pytake()). - No
product: trueon any line: no mom handles the band. broll.band:{"avatar": false, "product": true, "display": "off", "prompt": "<close shot of the band worn on a sleeping teen's wrist; display-off sentence>", "motion": "...", "duration": 4}. Real-footage fallback: copy the clip to<job>/broll/band.mp4and never runbroll --id band.prompt(the lock):scene= "Handheld UGC iPhone footage of a woman on a sidewalk being interviewed, a small black handheld microphone held toward her from the edge of the frame. She looks at the interviewer just off camera, mouth moving naturally as she speaks."says= "The woman says:".rules= "Pace: natural, quick, a little exasperated, like a real mom stopped on the street. One take, no jump cuts. Same person, same location, same camera framing. Do not recompose the shot. No captions, subtitles or on-screen text." (Rules are in the lock hash: set beforelock.)avatar: rewritten per mom before eachavatarroll (genetic change,wardrobe: "keep"unless the ref's top is grey/blank → name a specific real garment, accessories,audio_logic: "A handheld mic is pointed at her",background: keep).- No
voicepreset:audio_idswould give every mom the same voice.
7. Visual grammar
- One take per person, one person per take. If a mom needs more words, one longer take (≤10 s), never two takes of the same mom.
- Real-location refs: every mom's frame starts from a real frame of a real vox-pop / street video (
omni.py mine/refscan), Facebook or Instagram for 35-54 moms (KJ platform rule). Location, light, sensor and mic come from the real frame;avatarchanges the person. Remove logos, mic flags and names. - Casting diversity: US moms 35-54 of kids 11-17. Across 6-8 moms ≥4 ethnicities, ages 36-53, varied build, hair and real outfits (scrubs, puffer vest, half-zip, work blazer). No grey tops. Never two look-alikes back to back.
- Wrists: no watches, bands or bracelets on any mom (a lit smartwatch reads as the product). Hands low or out of frame.
- Interviewer: never seen beyond the hand + mic at the frame edge; off-camera voice only.
- The band insert (ugc-omni §5 gate): close shot, band worn on a sleeping teen's wrist (or buzzing on a nightstand), display off with ugc-omni's display-off sentence; the frame passes
pick --name broll_band --product(≥7/10 vs the real pack, right display state) before--stage anim. Fallback: real footage (clip-library/raw/dawn-broll-2026-09-10/DAWN BROLL/Jessica/, e.g.kid sleeping wearing dawn band broll.h264-clean.mp4), checked by eye. - No readable text anywhere except our overlay and captions.
8. Audio
- Each mom = Omni's native voice for her take. No ElevenLabs on the moms (Fish 10-04).
- Interviewer: one real human phone recording (best, free) or one ElevenLabs voice. Two lines total. Record outdoors or lay ~1.5 s of street ambience from mom 1's take under it (ffmpeg
amix); a dry indoor voice over street takes reads detached (starpophow-to-make-ai-ugc-videos-seedance-2-0). It sits over the first ~1.5 s of mom 1's frame and in a 1-1.5 s gap before the answer mom. - Street room tone comes from the takes. No music bed, or a very low one.
- Pace: native speed.
verifyflags <2.6 wps, holds ≥0.4 s, cut-offs, transcript <0.90.
9. Production (ugc-omni; every paid step dry-runs until --spend)
# run from ~/.openclaw/workspace; job.json per §6
O=scripts/ugc_omni/omni.py; J=output/ugc_omni/<slug>
# G0 (free): question + answers + storyboard
python3 $O mine <vox-pop URLs...> --keywords alarm,morning
python3 $O board --job $J # structure clean; only "frame not picked yet" errors may remain → Fish
# G1: per mom k (refscan/ref overwrite ref/, so one mom at a time)
python3 $O refscan --job $J --src <video k>
python3 $O ref --job $J --pick $J/ref/<cand>.jpg && cp $J/ref/ref.jpg $J/ref/m<k>.jpg # keep provenance
# edit job.avatar for mom k
python3 $O avatar --job $J --n 2 --arms gpt2,sd5 [--spend]
python3 $O pick --job $J --name <avatar|h2|h3|m2..m8> --file $J/avatar/<pick>.png
python3 $O board --job $J # now fully clean
# G2: lock (motion proof) on hook mom 1: same line good 3x
python3 $O lock --job $J --line H1-1 --n 3 [--spend]
python3 $O lock --job $J --approve $J/lock/<take>.mp4
# G3: every take + the band insert
python3 $O gen --job $J --lines all [--spend]
python3 $O verify --job $J
python3 $O broll --job $J --id band --stage frame --arms gpt2,sd5 [--spend]
python3 $O pick --job $J --name broll_band --product --file $J/broll/<pick>.png # refuses <7/10 / wrong display
python3 $O broll --job $J --id band --stage anim [--spend]
# (fallback) cp "<real band clip>.h264-clean.mp4" $J/broll/band.mp4
# G4: assembled cut
python3 $O assemble --job $J --hook H1 # repeat H2, H3
python3 $O cost --job $J
python3 scripts/meta_preflight.py $J/final/<slug>-H1.mp4
- Cost (kie, omni.py): Omni buckets $0.315 / $0.42 / $0.525 / $0.63 for 4 / 6 / 8 / 10 s; images ~$0.05 (gpt2) - $0.075 (sd5). 10 moms × 2 rolls × 2 arms ≈ $2.5; lock 3 × $0.42 ≈ $1.3; 10 takes at 4-6 s ≈ $3.5-4; band insert ≈ $0.6; regens ~+30% → ≈ $10-12 per ad with 3 hooks.
omni.py costprices it. Check kie credit first (GET https://api.kie.ai/api/v1/chat/credit). - No draft→final step: the G2 3× lock is the proof. Tiers: every A-roll take is Omni; tiers only pick b-roll (the band insert = tier A: gated omni or real footage).
- Image arms: pass
--arms gpt2,sd5(omni.py's default still lists nbp; its owner changes that). - Interviewer track + question banner: not built in
omni.py assemble(single A-roll track, captions only). Add the two interviewer lines and the question text on the final cut with ffmpeg /scripts/remotion_stitch.py(core §2a). An--intro/--vo-insertoption needs Fish's go and the ugc-omni owner. - Expected
boardwarnings here: "~7s straight on the A-roll face" (the face changes every take) and product-vs-scene-change distance (KJ's single-avatar rules). Errors still block. - Time: refs + avatars ~2 h by eye; lock ~30 min; gen + assemble ~30 min.
10. Gates & QA checklist (ugc-omni G0-G4; every gate is Fish's)
- [ ] G0 question + answers +
board: parent named with CPA, ONE variable; every answer has asrcID (survey S0x / stake bank # / Q0x / q-ID), verbatim or trimmed for speech, no new facts; the answer mom uses Fish's ruled mechanism lines; moms 1-6 mention no product; product named once; red-team done. No paid step before G0. - [ ] G1 each mom's real ref frame + avatar: ≥4 ethnicities, ages 36-53, no grey tops, bare wrists, no look-alikes adjacent; identity check that no mom is her ref person; mic present, street reads real.
- [ ] G2 lock: mom 1's hook line, same take good 3× (the 9-30 motion proof; in this format one take is the whole unit). No
genbefore it. - [ ] G3 every take watched (
verifyclean or each flag accepted); the band insert passedpick --productand Fish's eye vs the real pack (or real footage used). - [ ] G4 final cut by eye and ear: interviewer audible, captions = script, "Dramatization" label, end card only if
--end-card,meta_preflightrun; no on-screen claim that these are real interviewed moms. Gemini QA is not a gate.
11. Variants & test plan
- Minimal test (G1-G2, ~$2): one mom, one ref, avatar rolls,
lock --n 3on her hook line. Fish judges street realism, lips, voice. Stop here if it fails. - Full test: one ad, 3 hooks = 3 different answer-1 moms/lines on a shared body (
assemble --hook Hn). Hook directions: (a) exhaustion count ("fifty trips every morning", stake bank #11), (b) stake on her ("letters in the mail"), (c) the house wakes up except him. - Own ad set in Testing CBO (core §6); read at 7 days on CPA vs the parent, plus 3-s hold and lifetime frequency. Never pause; reduce.
- Then one variable per ad set: the question (who wakes him / how many alarms / worst morning), the stake mom, location set (school vs errands), runtime (30 vs 60 s), band insert (gated omni vs real).
- Side arm, only after a winner exists: H3 native-voice take (
yapper.py) on the same ref frame + same line vs the locked Omni take (one variable, Fish's go). Seedance arms stay dropped (Fish 10-04; product shots rejected 10-06).
12. Failure modes
- Hook mom morphs into another mom at the end of her take → hook line missing
end_frame= its own frame (flash lands hooks on the body's first frame). - Provider refuses the face → another ref/route. Never blur, crop or distance a face to sneak it past a likeness/safety filter.
- Frozen "suspiciously static" street (Foreplay #9 tell) → add one clause to
prompt.rules("people walk past in the background"), one change per roll. If it morphs, back it out. - Speech stretched / cut off → fit the line to 4/6/8 s at the locked wps;
plan_linepads,verifytrims. A 10-word line at 4 s gets cut off (10-06). - A mom says the line twice or adds words → transcript <0.90;
gen --lines X --redo. - Model draws a watch / lit smartwatch → wrists out of frame; re-roll the frame.
- Band insert scores <7/10 or shows a screen pod → re-roll with the display-off sentence and closer framing; still failing → real footage.
- Same-voice or same-face moms → recast the frame (genetic change), never a voice preset.
- Testimonial drift (several moms praising a product) → cut back to one answer mom.
- Overbuilding (chains, keyframes, composites) → one start frame, plain prompt, one take.
- Not built (needs Fish's go): interviewer/question track + banner in
omni.py assemble. - Animated fallback (cartoon-h3 dialogue, per format-spoken-drama pack spec): with no lip-sync each mom's line plays over a literal insert or a profile wide with the mic; it loses "see her say it". Only if the photoreal proof fails. Needs 6+ contrasting EL voices; similar-age moms visually distinct and in separate batches.
13. Sources
~/research/davidaistar/REPORT.md (§2.4, §3 #9, §4 #5), foreplay/09_street_interview.gemini.md, analysis/diff_ai_ugc.md, diff_podcast_dialogue.md, ugc_seedance.md, gemini_CZwfTy7jZcQ.md, gemini_Rqim7SlMtZg.md; skills/ugc-omni/SKILL.md, scripts/ugc_omni/omni.py, scripts/yapper/yapper.py, scripts/cartoon_h3/endcard.py; MAP.md §2; stake bank; customer-evidence-2026-09-15.md; memory files in §1.
Not built yet (needs Fish's go)
- Interviewer track + question banner: not built in omni.py assemble (single A-roll track, captions only). Add the two interviewer lines and the question text on the final cut with ffmpeg / scripts/remotion_stitch.py (core §2a). An --intro/--
- Not built (needs Fish's go): interviewer/question track + banner in omni.py assemble