format-podcast
Two-mom podcast-clip ad, 25-60 s. One mom tells her blamed-for-the-mornings story at the mic, the other reacts and asks, and the band lands as the answer in a product-gated cutaway.
Lane: (A) cartoon-h3 dialogue (staged, no lip-sync); (B) photoreal = ugc-omni (one Omni take per line, two speaker frames, gates G0-G4); H3 native voice is a side arm only · not built yet: 0 (see below) · source: /Users/ayden/.openclaw/workspace/skills/format-podcast
Podcast — sources
Reference ads / tutorials (davidaistar, analysed 10-06 in ~/research/davidaistar/)
- GIczkree_W0: RV-mattress podcast, ~30 s, 9:16 (
analysis/gemini_GIczkree_W0.md, full transcript txt/GIczkree_W0.txt).
- Beat by beat (paraphrased): one woman asks the other how a trip in the new RV went → the other loved it but couldn't sleep → the asker says they had the same problem and blamed the road until it turned out to be the mattress → "where did you get yours" → names a US company → the picky-sleeper objection → the long sleep trial is what sold her → she's slept well for a year since.
- Two singles intercut, each woman at a mic in an RV interior. No b-roll, no music, TikTok captions.
- Keep:
- One generation per speaker with all her lines. The model has a 5 s minimum, and batching holds the voice. Plan 2-3 generations per person, not 1 per line (~30).
- The prompt states the script plainly (
Line 1: '…') with short context and tone. He tried shot keywords and the plain version won.
- Real-photo refs for both women. Cut line by line in the edit.
- The cropped-in framing read more native in his test.
- The ElevenLabs single-syllable patch for a clipped last word ("I thin–" → "k").
- His own verdict: ~80% of realism is a human-sounding script.
- Change:
- His script is a testimonial with a weak transition; he says so himself. Ours is recipe A: a stake on her, an accuser, vindication.
- His crop advice was ambiguous here (transcript: cropped his 9:16; summaries: 16:9→9:16). His article settles it: a tighter letterbox crop of the 9:16 output (see starpop below). Still an untested side arm, not the default.
- We gate photoreal per the avatar-lane rules. He ships "on vibes".
- pYNC49BHsHQ: "AI podcast videos that go viral", 60-90 s (
analysis/gemini_pYNC49BHsHQ.md, analysis/ugc_seedance.md §pYNC49BHsHQ).
- Structure (paraphrased): a green-screen "relatable customer" hook → problem with a villain to blame → free value (the biggest section) → product plug via one reason-to-believe → scarcity CTA → "comment X" CTA.
- Pipeline: re-skinned screenshot base images → a voice originated in a video model, then cloned → separate lip-sync → CapCut layering (ours: ffmpeg/Remotion, core §2a).
- Keep:
- "Free value" maps to our root-cause beat: the explainer line is the useful part the viewer would share.
- A villain to blame maps to sound-based alarms, never the kid.
- Change:
- No "comment for the link" CTA (paid Meta).
- No screenshotting a real creator to re-skin (identity risk).
- No separate lip-sync pass. H3 native speech has no sync seam.
- The voice-clone-from-video trick is noted as an option only if a future version forces ElevenLabs.
Foreplay
- 10_character_arguments: Penrose, 62 s, 3D podcast debate. Two characters at mics trade jabs, a third voice interjects, the product character explains the mechanism, and it ends on a price contrast. Shows short alternating confrontational turns can carry ~60 s. Their characters don't really lip-sync, and it still runs. That supports Execution A's staging.
- 09_street_interview: Skincu, 54 s. The product appears as an inset graphic while the speaker keeps talking. This is the precedent for "product in a cutaway or graphic, never in hand".
davidaistar diffs (what's built vs not)
analysis/diff_podcast_dialogue.md:
- P0
yapper.py duo (2 personas, one native take each) and P0 duo-assemble (intercut by word alignment): neither exists.
- P1 required
--ref on frames: --ref is optional today.
- P1 banter checklist: lives in this skill's beat sheet.
- P2
patch-word and the 16:9 crop A/B: neither exists.
analysis/diff_ai_ugc.md: Kling 3 / Veo 3.1 scramble who says what in multi-speaker single generations, which is why we use singles only. Seedance reference-to-video refuses realistic faces on fal, and we never route around that.
Our receipts
project_dawn_yapper_2026-10-04 (10-04 ~23:05 → 23:30):
minimax/h3-max/image-to-video with no target_audio_url speaks the quoted dialogue natively (Scribe verbatim), and Fish approved it.
- 1080P looks more HD than 768P + upscale (Fish's eye beat the sharpness score).
- Native duration = ceil(words/3.8).
- Generated band-in-hand = fused fingers (3-5/10). Product = real footage (superseded 10-06: ugc-omni's product gate allows checked generated band frames; real footage = fallback).
- Quoted phrases in the direction text get spoken.
- "Telling a story to her phone" makes the model draw a phone.
- corrections-hot 10-05: Fish closed yapper after the chained-keyframe build broke ("curvelle talking head with no effort on h3 was literally fine"). Simplest recipe first.
feedback_ai_avatar_broll_lane: motion proof first, one locked character, line-match ≥8/10, pace check, eyes not scores.
- 10-01 attendance-office dialogue pack: the pace keys that passed vo_check.
- MAP.md: recipe A evidence. Testimonial-style is the weakest device in both songs (1.26) and statics.
starpop.ai articles (his written process, read 10-06; paraphrased, txt in ~/research/davidaistar/starpop/txt/)
- realistic-ai-podcast-ads-flux-3 (the RV build written up):
- Changed in the skill: the crop is a tighter, letterboxed crop of the model's 9:16 output (full 9:16 read "too clean"), so our arm is ffmpeg crop + black pad on the 9:16 take, no 16:9 frames and no upscale (§7).
- Confirms (already in the skill): real podcast ref photo (check it isn't AI), a real actress image as the character ref beats an improvised face, all of one host's lines in one generation (2-3 per person, not ~30), plain
Line 1/Line 2 prompt + one context sentence + one tone sentence, single-letter EL patch for a clipped last word. His own review: the script was the weak point (abrupt product transition, no b-roll).
- Model: his generations were Flux 3.0 video (5 s minimum, 20 s maximum per generation).
blackforestlabs/flux-3/image-to-video is in our fal catalog but not in the registry's native-speech row: proposed as a candidate, one-unit proof vs H3 i2v before any use.
- podcast-style-ads-for-ecommerce: product named in the first ~10 s = a normal ad with mics (our beat 5 sits at 60-78%); dialogue too clean is the tell (our banter rules); one concrete proof point (ours: 60-night guarantee). Crop point as above.
- podcast-ad-examples-swipe-file:
- Kept as hook tests (§11): the skeptical HOST who gets converted; the personal-question opener (his "question hooks beat statement hooks" is a claim, not data); scrubs as visual authority for the nurse-mom HOST.
- Not adopted: founder-origin guest (this skill bans a founder/celebrity host as deceptive); an animated talking product (the band must read as product-truth: a face on it invites the lit-smartwatch look); a fictional branded show logo/name as a recurring container, until Fish rules (it dresses the ad up as a real series).
- how-to-caption-a-podcast-from-a-transcript: forced alignment of known text beats auto-transcription (we already caption from the script). Kept: strip speaker labels before captioning, one caption per line, ~25 characters per line at 9:16 (42 at 16:9), keep lines in spoken order across speakers (§10). His tool + CapCut import → ours: Scribe alignment + Remotion/yapper assemble.
- hyper-realistic-ai-lip-sync-ads — not adopted: voice first in TTS (stretched-vowel spelling), then a cheap fast lip-sync model per line. Fish approved H3 native voice 10-04 and rejected EL over a talking head; B stays native. Also not adopted: cloning a stranger's voice from found clips. Flag: core §2's yapper row and the registry lip-sync row still say "H3 lip-sync + EL VO", which contradicts the native-voice receipts here; proposed for Fish to rule.
- how-to-use-claude-for-ugc-ads — not adopted: his podcast format as ONE 13 s two-shot generation (host from text, guest from an actor ref). We keep one generation per speaker (multi-speaker attribution scrambles, diff_ai_ugc).