Pipeline Hub

format-podcast

Two-mom podcast-clip ad, 25-60 s. One mom tells her blamed-for-the-mornings story at the mic, the other reacts and asks, and the band lands as the answer in a product-gated cutaway.

Lane: (A) cartoon-h3 dialogue (staged, no lip-sync); (B) photoreal = ugc-omni (one Omni take per line, two speaker frames, gates G0-G4); H3 native voice is a side arm only · not built yet: 0 (see below) · source: /Users/ayden/.openclaw/workspace/skills/format-podcast

SKILL.mdbeat-sheet.mdsources.mdEvals (7)Learnings (16)
1. Read first2. When to use / when not3. The format4. Why it sells5. Beat sheet (full rules + illustrative structure in references/beat-sheet.md)6. Pack spec7. Visual grammar8. Audio9. Production10. Gates & QA checklist (on top of core §4)11. Variants & test plan12. Failure modes13. Sources

format-podcast — two moms at the mics, the band is the answer

1. Read first

2. When to use / when not

3. The format

4. Why it sells

5. Beat sheet (full rules + illustrative structure in references/beat-sheet.md)

Beat % runtime Job Must contain
1 Cold open 0-10 Stop the scroll on her Mid-conversation; the stake or symptom in line 1-2 (VOC-sourced); HOST reacts within 2 s
2 The mornings 10-30 Prove she's the alarm Concrete ritual (knocks, times, alarm count); HOST interjects at least twice; failed solutions generic
3 The blame 30-45 Raise the stake Who blamed her and what it cost (letter, fight, mediation); HOST raises it one step
4 Root cause 45-60 Why he doesn't wake Credible person inside her story, quoted; Fish's ruled root-cause line verbatim
5 The answer 60-78 Mechanism + product Touch gets through; silent to everyone else; stronger until he turns it off; product named once; band cutaway (gated)
6 Vindication 78-93 Payoff on HER Accuser concedes / kid gets himself up; HOST's one-line reaction
7 CTA 93-100 Offer One spoken line (60-night guarantee) + end card (optional, --end-card; off by default per Fish 10-06)

6. Pack spec

One concept pack (output/concept-packs/<slug>.md) is the script source for both executions. - Front matter (core §5) plus: format: cartoon-h3, format_skill: format-podcast, parent: (the paying ad whose story this is, with CPA), iteration: (the ONE variable, e.g. "song → 45 s podcast convo, same story"), voices: MOM=<id>, HOST=<id> (STORY mom is tagged MOM because pack.py falls back to voices.MOM as the default voice). Pace keys from the 10-01 packs that passed vo_check: vo_tempo: 1.08, gap_same: 0.15, gap_turn: 0.3, max_gap: 0.22, seam_gap: 0.25. Add style_HOST: 0.2-0.3 for audible reactions. Also end_card, title, overlay, body_start (3 hooks × 1 body). - ## VO Script: N. [MOM] "line" / N. [HOST] "line", one sentence per line. Every line is tagged; there's no narrator. - ## Scene visual intents (A): N. **role** — <env>: <who is on screen, doing what, framing>. For every studio line, name the LISTENER or the angle that hides the speaker's mouth (§7). - Cast: MOM + HOST (+ story-cutaway cast: kid, accuser, explainer). Two moms of similar age contaminate each other in one batch (core §3). Make them visibly distinct and keep their close-ups in separate batches. - Envs: podcast-studio (one set, reused) + 3-5 story envs (teen room, kitchen, hallway, school office, car). - Execution B adds one ugc-omni job, output/ugc_omni/<slug>/job.json (shape: references/omni-job.example.json; schema skills/ugc-omni/job.example.json): - Engine keys: style: photoreal, mode: cuts, aroll_model: flash, resolution: 1080p, image_arms: ["gpt2", "sd5"], product_refs: ["@dawn"], product_truth with ugc-omni's display-off sentence; plus core §5 keys, parent, iteration, label: "Dramatization". - lines[] = the pack's VO Script in order: [MOM] → frame: "avatar", [HOST] → frame: "host"; visual: "AI UGC", or "broll:<id>" on long turns and the product line (broll:band); hooks carry hook + end_frame = their own frame. - prompt.scene covers both women ("Static locked-off shot, UGC iPhone footage of a woman at a home podcast desk, a microphone on a boom arm at chin height, headphones on. She looks just off camera at the other host..."); says: "The woman says:"; rules with "One take, no jump cuts. Same avatar, same location, same camera framing. Do not recompose the shot. No captions, subtitles or on-screen text." - avatar rewritten per speaker before her roll (genetic change, wardrobe: keep or a specific real garment, no grey tops, bare wrists, audio_logic: the podcast mic). No voice preset (it would give both women one voice). - broll.band: product: true, avatar: false, display: "off"; story cutaways avatar: false (or outfit when MOM is in them).

7. Visual grammar

Both executions - Each mom faces the other: STORY looks slightly frame-right, HOST frame-left, so intercut singles read as one conversation. Same room, same light temperature, same mic model in both. - Wardrobe like the real buyer, no grey tops, a different colour per speaker. No watch and no wristband on either speaker. - The band lives in a cutaway or a graphic during beat 5, never on a speaker's wrist. In B the cutaway is product-gated (ugc-omni §5: pick --name broll_band --product ≥7/10, display-off sentence, close framing, worn on a wrist or resting across a palm, never pinched upright); real footage is the fallback. Ungated band-in-hand killed yapper (3-4/10).

Execution A (animated) - Shots that need no lip-sync: - wide two-shot in profile with the mics in front of the mouths; - LISTENER close-up while the other voice plays; - over-the-shoulder onto the listener; - story cutaways (fever-dream literal: show exactly what she describes). - Never a close-up of a speaker's face on her own line. - Ratio target: ≤40% studio, ≥60% story cutaways. The studio returns at turn changes and reactions. - Product cutaway: wrist resting on the duvet or band on the nightstand, side-on, display DARK, eyes opening plus small motion lines (wrist rule). Clock hands only; no readable text.

Shot tiers (core §2b; tag every batch/take in storyboard.md / the cut list): A = the hook's first 3 s, the beat-5 product cutaway, the CTA/end; B = two-shots, listener close-ups, story cutaways with 1-2 people; C = studio establishing wide, empty rooms, clock or door inserts. run.py auto-tags tiers into <run>/tiers.md + plan.json (tiers.py; re-tag if the word match is wrong); per-tier models via --tier-models (fal), default all H3-max; A keeps H3-max.

Execution B (photoreal, ugc-omni) - Singles only. Medium close-up, chest-up, eye level, static. Mic foam at chin height and off-centre so the lips stay visible. Headphones on. Hands rest on the desk or stay below frame. Each frame comes from its own real podcast-creator reference (refscan), same room family and light for both. - No two-speaker generation: line assignment across two mouths is unproven; each take is one woman on her own frame. - Covering long turns: a story cutaway b-roll line (the other woman's reaction would need its own spoken line in omni; write a 1-2 word reaction line instead). - Crop rule (letterbox arm): a tighter crop of the 9:16 output, letterboxed like real podcast clips (starpop realistic-ai-podcast-ads-flux-3): ffmpeg on the assembled cut, -vf "crop=1080:1350,pad=1080:1920:0:(oh-ih)/2:black". Untested; the first test ships full-frame 9:16; the letterbox is its own later variable. - Tiers: every A-roll take is Omni; tiers only pick b-roll (A = the band cutaway: gated omni or real footage; C = empty rooms, clocks, doors). - Product fallback: real footage from clip-library/raw/dawn-broll-2026-09-10/DAWN BROLL/{Jessica,Kris,Lori}/*.h264-clean.mp4 copied to <job>/broll/band.mp4, or knowledge/brands/dawnbands/product-cutout-unlit.png as a graphic.

8. Audio

9. Production

A — animated (free until refs/sheets):

P=output/concept-packs/<slug>.md
python3 -m scripts.cartoon_h3.run --pack $P --stage vo                  # EL per speaker + vo_check DIALOGUE gate
B=(${(z)"$(python3 -m scripts.cartoon_h3.packargs $P --storyboard)"}); B=(${(Q)B})   # zsh; before sheets exist (missing refs dropped)
python3 -m scripts.cartoon_h3.run $B --stage storyboard                # free; Fish reviews storyboard.md
# sheets, one per cast member/env, after Fish's go (~$0.15-0.30 each; render_sheets.sh only knows the 10-04 song cast):
python3 -m scripts.cartoon_h3.sheet <out.png> <prompt.txt> [refs...]
touch output/cartoon-h3/<slug>/storyboard.approved                    # Fish only
A=(${(z)"$(python3 -m scripts.cartoon_h3.packargs $P)"}); A=(${(Q)A})   # zsh; refs live in <pack>.args (core §5)
python3 -m scripts.cartoon_h3.run $A --stage canvas
python3 -m scripts.cartoon_h3.run $A --stage gen --prompts-only   # Fish reviews motion prompts
python3 -m scripts.cartoon_h3.run $A --stage gen --spend --backend fal --model h3-max --candidates 1
ffmpeg -f lavfi -i anoisesrc=c=brown:a=0.002 -t 300 -ar 44100 roomtone.mp3   # once; podcast = no music
python3 -m scripts.cartoon_h3.run $A --stage assemble --music roomtone.mp3

B — photoreal (ugc-omni; every paid step dry-runs until --spend):

O=scripts/ugc_omni/omni.py; J=output/ugc_omni/<slug>     # job.json per §6
# G0 (free): script + hooks + storyboard
python3 $O board --job $J          # structure clean; only "frame not picked yet" errors may remain → Fish
# G1: MOM then HOST (refscan/ref overwrite ref/, so one speaker at a time; job.avatar rewritten per speaker)
python3 $O mine <real podcast-clip URLs...> --keywords mom,kids
python3 $O refscan --job $J --src <mom video>;  python3 $O ref --job $J --pick $J/ref/<cand>.jpg
python3 $O avatar --job $J --n 3 --arms gpt2,sd5 [--spend];  python3 $O pick --job $J --name avatar --file $J/avatar/<pick>.png
python3 $O refscan --job $J --src <host video>; python3 $O ref --job $J --pick $J/ref/<cand>.jpg
python3 $O avatar --job $J --n 3 --arms gpt2,sd5 [--spend];  python3 $O pick --job $J --name host --file $J/avatar/<pick>.png
python3 $O board --job $J          # now fully clean
# G2: lock on a MOM line (same take good 3x), then prove the same lock on a HOST line
python3 $O lock --job $J --line H1-1 --n 3 [--spend];  python3 $O lock --job $J --approve $J/lock/<take>.mp4
python3 $O lock --job $J --line <host line> --frame host --n 3 [--spend]   # same prompt, HOST frame; Fish's yes
# G3: every line + b-roll + product check
python3 $O gen --job $J --lines all [--spend];  python3 $O verify --job $J
python3 $O broll --job $J --id band --stage frame --arms gpt2,sd5 [--spend]
python3 $O pick --job $J --name broll_band --product --file $J/broll/<pick>.png   # refuses <7/10 / wrong display
python3 $O broll --job $J --id band --stage anim [--spend]                   # (fallback: real clip → broll/band.mp4)
python3 $O broll --job $J --id <cutaway> --stage frame [--spend]  ... --stage anim [--spend]
# G4: assembled cut (A-roll track = the intercut)
python3 $O assemble --job $J --hook H1;  python3 $O cost --job $J
python3 scripts/meta_preflight.py $J/final/<slug>-H1.mp4

10. Gates & QA checklist (on top of core §4)

11. Variants & test plan

12. Failure modes

13. Sources

~/research/davidaistar/REPORT.md §1-2, analysis/diff_podcast_dialogue.md, gemini_GIczkree_W0.md, gemini_pYNC49BHsHQ.md, ugc_seedance.md (pYNC49BHsHQ), diff_ai_ugc.md, txt/GIczkree_W0.txt; Foreplay 10_character_arguments, 09_street_interview. Code: scripts/yapper/yapper.py + settings.py, scripts/cartoon_h3/{run.py,pack.py,vo.py,vo_check.py,h3.py,endcard.py}. Memory: yapper 10-04, avatar-lane, reference-images, ugc-wardrobe, corrections 10-04/10-05. MAP.md, stake bank.