format-street-interview
Photoreal vox-pop / street interview ad, 30-60 s. An off-camera interviewer asks moms one question ("Who gets your kid out of bed on school mornings?"), 6-8 different moms each answer in ONE 4-8 s take with real VOC lines, and the band is the answer, shown in a product-gated b-roll insert (real footage as fallback).
Lane: — · not built yet: 2 (see below) · source: /Users/ayden/.openclaw/workspace/skills/format-street-interview
Street interview: sources
Reference ad (Foreplay, analysed 10-06 in ~/research/davidaistar/foreplay/)
- 09_street_interview: Skincu pH-reactive lipstick, 54 s, 9:16, faux street interview, heavily AI. Beat by beat (paraphrased):
- 0:00-0:04: an off-camera hand pushes a small mic at a stylish older woman on a busy street. The interviewer asks which shade she's wearing, and she acts flattered. One continuous shot.
- 0:04-0:14: she says people assume it's custom-mixed and that it changed her routine.
- 0:14-0:23: she names the brand and applies it on camera (goes on clear). A product inset graphic appears.
- 0:24-0:38: the mechanism (it reads pH and sets the colour; binds so it doesn't bleed). Flash transition.
- 0:39-0:45: she speaks to the mature viewer directly (no bleeding into lip lines).
- 0:45-0:54: CTA with a 30-day guarantee and BOGO badges.
- Production read: AI video + lip-sync, AI voices for both speakers, a real product pasted as static insets. Tells: stiff micro-expressions, uncanny eyes, slight lip-sync delay, a "suspiciously static" street, unnatural hand/product interaction. Gemini's own risk note: avoid having the avatar touch the product (it named a wristband as the hard case).
- Keep: frame-0 mic-in-face hook, problem-aware question, mechanism in plain words, guarantee at the end, the native street-interview look.
- Change: many moms with one take each instead of one long take (no identity to hold, no long stare where the tells show). The talk is about her problem, not product praise (testimonial device loses on Dawn, MAP.md §2). The product never touches a generated hand: real footage + end card (optional,
--end-card; off by default per Fish 10-06). Their tool guesses (Kling/Runway/HeyGen) are stale (REPORT §3 note).
davidaistar method notes
- One shot per person (REPORT §2.4): every clip is a different person, so there's no identity to hold. This dodges the consistency failure that killed Side Effects and yapper.
- Street interview = Seedance 2.0's showcase (gemini_CZwfTy7jZcQ, ugc_seedance §CZwfTy7jZcQ/jcc-r5I-SlU): correct speaker attribution in a two-person exchange where Kling 3.0 and Veo 3.1 swapped who said what. We sidestep the problem: the interviewer is off camera and every take has one speaker.
- Seedance 2.0 access (diff_ai_ugc top 1): StarPop 7-day trial or jianying ~$11/mo; omni-ref (9 img + 3 video + 3 audio), 15 s cap, native audio, in-model extend. Superseded for us: the registry's current pick is Seedance 2.5 on fal (reference-to-video refuses realistic faces; image-to-video with native audio is a yapper arm, $0.46/s). Optional proof arm only, Fish's go.
- Face filter: his trick (blur or medium-distance face refs to pass Seedance's close-up filter, also sold as a "character variety" tool) is a likeness-filter bypass and not adopted (diff_ai_ugc §2). Our response is the compliant one: i2v works where reference-to-video refuses (yapper 10-04 ~19:00).
- Real-photo refs beat improvised faces (diff_podcast_dialogue: Pinterest/real refs → img2img; ugc-omni/KJ: a real creator frame is "the whole unlock"; KJ says Pinterest is AI slop now, so use real video frames). Name the avatar's demographic in the pull (feedback_reference_images §3).
- Generate 16:9, crop to 9:16 reads more native (REPORT §2.10). Untested on our stack, and omni.py renders 9:16 directly. A later one-variable test, not part of the proof.
- Motion swap / frame-1 anchoring (gemini_Rqim7SlMtZg): extract a real frame, swap the person, keep the pose and setting. Same idea as our real-location ref → avatar genetic change. His face-only swap that keeps a real creator's identity context is flagged as an identity-copying risk: our genetic-change rule (ugc-omni §2) makes it impossible to be the same person.
- Hand + product is the failure point (09 risk note; diff_ai_ugc §2: omni-ref "solves product accuracy" is unproven on a small wearable; our 10-04 band-in-hand renders were 3-4/10).
Our receipts
- Yapper 10-04/05 (
project_dawn_yapper_2026-10-04): H3 Max i2v with no audio input speaks the quoted line natively (Fish "yea its good"); generated hand-holding-band 3-4/10 with fused fingers; product must be real footage (superseded 10-06: generated band frames allowed through ugc-omni's product gate, real footage = fallback); never quote dialogue inside action text; H3 stretches speech to the requested duration (duration = ceil(words/3.8)); Fish closed the lane 10-05 over overengineering.
- 9-30 Side Effects post-mortem (
feedback_ai_avatar_broll_lane): motion proof first, line-match audit, eyes over scores.
scripts/ugc_omni/omni.py (built 10-06, no spend yet): refscan / ref / avatar / pick / lock / gen / verify / assemble / cost; Omni prices 4/6/8/10 s = $0.315/0.42/0.525/0.63.
- MAP.md §2: testimonial pairings are the weakest (1.21-1.31); recipe B callout (brand voice to her, her burden, science/guarantee authority, kid independence) is the lifetime money.
- Real band footage:
clip-library/raw/dawn-broll-2026-09-10/DAWN BROLL/Jessica/*.h264-clean.mp4 (used as yapper v1.1 cutaways 10-04).
starpop.ai articles (his written process, read 10-06; paraphrased)
- how-to-make-ai-ugc-videos-seedance-2-0 / how-to-use-seedance-2-0-to-make-ads: audio must belong to the scene; his own-voice gym clip felt disconnected because it was recorded in a quiet room. Changed: the interviewer line is recorded outdoors or gets the takes' street ambience mixed under it (§8). Both articles also say to request clean footage (no captions/text) and add text in the editor. Changed: the Omni
rules string now says so (§6).
- seedance-2-0-for-ai-ugc-ads / seedance-2-vs-sora-2-ugc: street interview is where Seedance 2.0 kept who-says-what straight in a two-person script (others swapped speakers). Already noted; we still keep one speaker per take with the interviewer off camera.
- Not adopted: the blurred-face reference (both Seedance articles sell it for "character variety" and to pass the close-up face block). It is a likeness-filter bypass. Our variety comes from a different real ref frame per mom + genetic change.
- how-to-use-claude-for-ugc-ads: his handheld-selfie prompts are short on purpose (over-direction kills the casual read). Matches our plain lock prompt; no change.