format-podcast — two moms at the mics, the band is the answer
1. Read first
knowledge/ad-formats/_FORMAT-CORE.md(doctrine, lane limits, gates, launch). Binding.references/beat-sheet.md(skeleton, banter rules, cut-list template).references/sources.md(reference ads, what we keep and change).output/dawn-combo-map-2026-10-05/MAP.md§1-2A and the stake bankoutput/creative-strategy-cooper-2026-10-05/dawn-stakes.md.- Dialogue pack conventions:
output/concept-packs/dawnbands-s2-attendance-office-h1-2026-10-01.md. - Execution B only:
skills/ugc-omni/SKILL.md+scripts/ugc_omni/omni.py, THE photoreal production engine (Fish 10-06); this skill hands it a job.json (references/omni-job.example.json, illustrative). A parallel session owns both: read, never edit. Alsofeedback_ai_avatar_broll_lane.md(the 9-30 rules ugc-omni's gates implement),feedback_reference_images.md,feedback_ugc_avatar_wardrobe.md, corrections-hot 10-05 "simplest working recipe". Side arm only:~/.claude/projects/-Users-ayden/memory/project_dawn_yapper_2026-10-04.md,scripts/yapper/README.md.
2. When to use / when not
- Use for recipe A told as a conversation: one mom carries the stake, and a second voice stands in for the viewer (reacts, asks, pushes back). It's short (25-60 s) next to the 2-4 min spoken drama. That makes it a cheap way to test whether a two-voice convo can carry the story without the song.
- Evidence status: untested on Dawn. davidaistar's RV-mattress podcast (GIczkree_W0) is a tutorial demo, and he says himself it's "not a perfect ad". Foreplay
10_character_arguments(Penrose, 62 s podcast debate) is live for another brand. The mechanics behind B: ugc-omni proved per-line Omni takes with a locked prompt on one speaker 10-06; two speakers = two frames under the same lock, untested. - Risk the data names: a podcast slides easily into a testimonial ("we got ours from…"). Testimonial-style songs ran 1.26 ROAS and testimonial statics are the weakest pairings (MAP.md §2). The structure in §5 keeps it a story with an accuser and a vindication. The product gets ≤20% of the runtime.
- Don't use for brand-voice callouts (recipe B goes to statics/natives), for a "founder" or celebrity host (deceptive, REPORT §3), or for any stake the mornings didn't cause.
3. The format
- 9:16, 25-60 s (≈90-200 words at ~3.4 wps including turn gaps). Two speakers:
- STORY mom (the avatar: ordinary American mom 35-54) carries ~60% of the words, first person.
- HOST is another mom, visibly different (age, hair, ethnicity, colour), who asks and reacts and carries ~40%.
- Studio look: home podcast set, a mic on a boom arm or desk stand, headphones, a bookshelf or wall behind. Static camera. Captions from the script on every word, a hook overlay on frame 0, end card (optional,
--end-card; off by default per Fish 10-06). No music: real podcast clips have room tone only (GIczkree_W0 ships with none). - Execution A (animated, buildable now): cartoon-h3 dialogue lane, one style per ad. The studio is a framing device. Most screen time goes to literal cutaways of the story she tells. Nobody mouths a line (core §3).
- Execution B (photoreal, ugc-omni): two speaker frames (MOM =
avatar, HOST =host) under one locked Omni prompt; one native-voice take per line;assembleconcatenates them in script order, so the intercut is the A-roll track. The band appears in a product-gated cutaway (real footage or the cutout graphic as fallback). "Dramatization" label (Fish 10-06). - References: GIczkree_W0 (RV mattress, ~30 s, intercut singles), pYNC49BHsHQ (hook → problem/villain → free value → plug structure). Details in
references/sources.md.
4. Why it sells
- Recipe A in miniature: the stake lands on HER (someone close blames her for the mornings) → HOST's reaction raises it ("they sent it to you?") → root cause from a credible person inside her story (the nurse she quotes) → the band as the thing that ended it → vindication on her.
- HOST is the viewer's proxy. She asks the questions a problem-aware mom would ask, including the objection ("he has five alarms"). Interrupting turns hold attention better than a monologue; that's davidaistar's "80% of realism is the script" point.
- Problem-aware: line 1 or 2 names her by an observable symptom, in action. A cold open mid-conversation counts as action. A pure mystery opener doesn't pass.
5. Beat sheet (full rules + illustrative structure in references/beat-sheet.md)
| Beat | % runtime | Job | Must contain |
|---|---|---|---|
| 1 Cold open | 0-10 | Stop the scroll on her | Mid-conversation; the stake or symptom in line 1-2 (VOC-sourced); HOST reacts within 2 s |
| 2 The mornings | 10-30 | Prove she's the alarm | Concrete ritual (knocks, times, alarm count); HOST interjects at least twice; failed solutions generic |
| 3 The blame | 30-45 | Raise the stake | Who blamed her and what it cost (letter, fight, mediation); HOST raises it one step |
| 4 Root cause | 45-60 | Why he doesn't wake | Credible person inside her story, quoted; Fish's ruled root-cause line verbatim |
| 5 The answer | 60-78 | Mechanism + product | Touch gets through; silent to everyone else; stronger until he turns it off; product named once; band cutaway (gated) |
| 6 Vindication | 78-93 | Payoff on HER | Accuser concedes / kid gets himself up; HOST's one-line reaction |
| 7 CTA | 93-100 | Offer | One spoken line (60-night guarantee) + end card (optional, --end-card; off by default per Fish 10-06) |
6. Pack spec
- Step 0 — combo plan (feedback loop,
workflow-combo-loop):python3 scripts/combo_loop/plan.py --format format-podcast --brand dawn→ pick one combo (a paying parent + ONE changed dimension) and paste its block into the pack front matter / job or spec JSON ("combo": {...}; no file →tags.py register --prefix):combo_avatarcombo_anglecombo_povcombo_authoritycombo_stakecombo_stake_oncombo_emotioncombo_root_causecombo_mechanismcombo_payoffcombo_devicecombo_formatcombo_parentcombo_variablecombo_ad_prefix(= the launched Meta ad-name prefix).python3 scripts/combo_loop/tags.py check <pack>must PASS before concept approval; the weekly refresh reads results back by that prefix.
One concept pack (output/concept-packs/<slug>.md) is the script source for both executions.
- Front matter (core §5) plus: format: cartoon-h3, format_skill: format-podcast, parent: (the paying ad whose story this is, with CPA), iteration: (the ONE variable, e.g. "song → 45 s podcast convo, same story"), voices: MOM=<id>, HOST=<id> (STORY mom is tagged MOM because pack.py falls back to voices.MOM as the default voice). Pace keys from the 10-01 packs that passed vo_check: vo_tempo: 1.08, gap_same: 0.15, gap_turn: 0.3, max_gap: 0.22, seam_gap: 0.25. Add style_HOST: 0.2-0.3 for audible reactions. Also end_card, title, overlay, body_start (3 hooks × 1 body).
- ## VO Script: N. [MOM] "line" / N. [HOST] "line", one sentence per line. Every line is tagged; there's no narrator.
- ## Scene visual intents (A): N. **role** — <env>: <who is on screen, doing what, framing>. For every studio line, name the LISTENER or the angle that hides the speaker's mouth (§7).
- Cast: MOM + HOST (+ story-cutaway cast: kid, accuser, explainer). Two moms of similar age contaminate each other in one batch (core §3). Make them visibly distinct and keep their close-ups in separate batches.
- Envs: podcast-studio (one set, reused) + 3-5 story envs (teen room, kitchen, hallway, school office, car).
- Execution B adds one ugc-omni job, output/ugc_omni/<slug>/job.json (shape: references/omni-job.example.json; schema skills/ugc-omni/job.example.json):
- Engine keys: style: photoreal, mode: cuts, aroll_model: flash, resolution: 1080p, image_arms: ["gpt2", "sd5"], product_refs: ["@dawn"], product_truth with ugc-omni's display-off sentence; plus core §5 keys, parent, iteration, label: "Dramatization".
- lines[] = the pack's VO Script in order: [MOM] → frame: "avatar", [HOST] → frame: "host"; visual: "AI UGC", or "broll:<id>" on long turns and the product line (broll:band); hooks carry hook + end_frame = their own frame.
- prompt.scene covers both women ("Static locked-off shot, UGC iPhone footage of a woman at a home podcast desk, a microphone on a boom arm at chin height, headphones on. She looks just off camera at the other host..."); says: "The woman says:"; rules with "One take, no jump cuts. Same avatar, same location, same camera framing. Do not recompose the shot. No captions, subtitles or on-screen text."
- avatar rewritten per speaker before her roll (genetic change, wardrobe: keep or a specific real garment, no grey tops, bare wrists, audio_logic: the podcast mic). No voice preset (it would give both women one voice).
- broll.band: product: true, avatar: false, display: "off"; story cutaways avatar: false (or outfit when MOM is in them).
7. Visual grammar
Both executions
- Each mom faces the other: STORY looks slightly frame-right, HOST frame-left, so intercut singles read as one conversation. Same room, same light temperature, same mic model in both.
- Wardrobe like the real buyer, no grey tops, a different colour per speaker. No watch and no wristband on either speaker.
- The band lives in a cutaway or a graphic during beat 5, never on a speaker's wrist. In B the cutaway is product-gated (ugc-omni §5: pick --name broll_band --product ≥7/10, display-off sentence, close framing, worn on a wrist or resting across a palm, never pinched upright); real footage is the fallback. Ungated band-in-hand killed yapper (3-4/10).
Execution A (animated) - Shots that need no lip-sync: - wide two-shot in profile with the mics in front of the mouths; - LISTENER close-up while the other voice plays; - over-the-shoulder onto the listener; - story cutaways (fever-dream literal: show exactly what she describes). - Never a close-up of a speaker's face on her own line. - Ratio target: ≤40% studio, ≥60% story cutaways. The studio returns at turn changes and reactions. - Product cutaway: wrist resting on the duvet or band on the nightstand, side-on, display DARK, eyes opening plus small motion lines (wrist rule). Clock hands only; no readable text.
Shot tiers (core §2b; tag every batch/take in storyboard.md / the cut list): A = the hook's first 3 s, the beat-5 product cutaway, the CTA/end; B = two-shots, listener close-ups, story cutaways with 1-2 people; C = studio establishing wide, empty rooms, clock or door inserts. run.py auto-tags tiers into <run>/tiers.md + plan.json (tiers.py; re-tag if the word match is wrong); per-tier models via --tier-models (fal), default all H3-max; A keeps H3-max.
Execution B (photoreal, ugc-omni)
- Singles only. Medium close-up, chest-up, eye level, static. Mic foam at chin height and off-centre so the lips stay visible. Headphones on. Hands rest on the desk or stay below frame. Each frame comes from its own real podcast-creator reference (refscan), same room family and light for both.
- No two-speaker generation: line assignment across two mouths is unproven; each take is one woman on her own frame.
- Covering long turns: a story cutaway b-roll line (the other woman's reaction would need its own spoken line in omni; write a 1-2 word reaction line instead).
- Crop rule (letterbox arm): a tighter crop of the 9:16 output, letterboxed like real podcast clips (starpop realistic-ai-podcast-ads-flux-3): ffmpeg on the assembled cut, -vf "crop=1080:1350,pad=1080:1920:0:(oh-ih)/2:black". Untested; the first test ships full-frame 9:16; the letterbox is its own later variable.
- Tiers: every A-roll take is Omni; tiers only pick b-roll (A = the band cutaway: gated omni or real footage; C = empty rooms, clocks, doors).
- Product fallback: real footage from clip-library/raw/dawn-broll-2026-09-10/DAWN BROLL/{Jessica,Kris,Lori}/*.h264-clean.mp4 copied to <job>/broll/band.mp4, or knowledge/brands/dawnbands/product-cutout-unlit.png as a graphic.
8. Audio
- A: ElevenLabs per speaker via
voices:. MOM = warm, slightly tired, 35-54. HOST = brighter and quicker. vo_check runs its DIALOGUE profile: gap_turn 0.30-0.55 s, gap_same 0.12-0.30 s, overall ≥2.9 wps. Never fix a line by speeding it up; split it. - B: Omni's native voice from the locked prompt, no ElevenLabs. One take per line:
plan_linepicks the 4/6/8/10 s bucket at the locked pace and pads throwaway words;verifycuts after the last real word (so no clipped-word patch is needed). If the two women drift toward one voice, recast a frame; never a sharedvoicepreset. - Loudness: A:
vo.pyper speaker; B:omni.py assembleapplies loudnorm -14 to the whole mix (check the two women match by ear at G4). - No music bed. A: cartoon-h3
assemblealways mixesDEFAULT_MUSICunless you pass--music <file>, so pass a room-tone file (§9). End card (optional,--end-card; off by default per Fish 10-06) 3 s.
9. Production
A — animated (free until refs/sheets):
P=output/concept-packs/<slug>.md
python3 -m scripts.cartoon_h3.run --pack $P --stage vo # EL per speaker + vo_check DIALOGUE gate
B=(${(z)"$(python3 -m scripts.cartoon_h3.packargs $P --storyboard)"}); B=(${(Q)B}) # zsh; before sheets exist (missing refs dropped)
python3 -m scripts.cartoon_h3.run $B --stage storyboard # free; Fish reviews storyboard.md
# sheets, one per cast member/env, after Fish's go (~$0.15-0.30 each; render_sheets.sh only knows the 10-04 song cast):
python3 -m scripts.cartoon_h3.sheet <out.png> <prompt.txt> [refs...]
touch output/cartoon-h3/<slug>/storyboard.approved # Fish only
A=(${(z)"$(python3 -m scripts.cartoon_h3.packargs $P)"}); A=(${(Q)A}) # zsh; refs live in <pack>.args (core §5)
python3 -m scripts.cartoon_h3.run $A --stage canvas
python3 -m scripts.cartoon_h3.run $A --stage gen --prompts-only # Fish reviews motion prompts
python3 -m scripts.cartoon_h3.run $A --stage gen --spend --backend fal --model h3-max --candidates 1
ffmpeg -f lavfi -i anoisesrc=c=brown:a=0.002 -t 300 -ar 44100 roomtone.mp3 # once; podcast = no music
python3 -m scripts.cartoon_h3.run $A --stage assemble --music roomtone.mp3
- Draft → final (core gate 6b): H3 has no motion-preserving re-render. A cheap draft =
run.py --draft(480P into<slug>-draft, $0.05/s vs $0.08/s) on the tier-A/B batches, iterate prompts there, final =--hdupscale of the approved clip. Fish judged native 1080 sharper than 768+upscale (10-04), so prove 480→upscale side by side before using it for finals; until then finals render native (--draftprints the 768P re-render command). - Cost: 45 s ≈ 4-5 H3 batches. Quote from
run.py --dry-run(tiers.md; h3.py fixed 10-06: h3-max 768P $0.08/s + ~$0.10-0.16 reference surcharge per batch). Budget ~$6-12 for video, ~$2-4 for sheets and canvases, plus EL quota.
B — photoreal (ugc-omni; every paid step dry-runs until --spend):
O=scripts/ugc_omni/omni.py; J=output/ugc_omni/<slug> # job.json per §6
# G0 (free): script + hooks + storyboard
python3 $O board --job $J # structure clean; only "frame not picked yet" errors may remain → Fish
# G1: MOM then HOST (refscan/ref overwrite ref/, so one speaker at a time; job.avatar rewritten per speaker)
python3 $O mine <real podcast-clip URLs...> --keywords mom,kids
python3 $O refscan --job $J --src <mom video>; python3 $O ref --job $J --pick $J/ref/<cand>.jpg
python3 $O avatar --job $J --n 3 --arms gpt2,sd5 [--spend]; python3 $O pick --job $J --name avatar --file $J/avatar/<pick>.png
python3 $O refscan --job $J --src <host video>; python3 $O ref --job $J --pick $J/ref/<cand>.jpg
python3 $O avatar --job $J --n 3 --arms gpt2,sd5 [--spend]; python3 $O pick --job $J --name host --file $J/avatar/<pick>.png
python3 $O board --job $J # now fully clean
# G2: lock on a MOM line (same take good 3x), then prove the same lock on a HOST line
python3 $O lock --job $J --line H1-1 --n 3 [--spend]; python3 $O lock --job $J --approve $J/lock/<take>.mp4
python3 $O lock --job $J --line <host line> --frame host --n 3 [--spend] # same prompt, HOST frame; Fish's yes
# G3: every line + b-roll + product check
python3 $O gen --job $J --lines all [--spend]; python3 $O verify --job $J
python3 $O broll --job $J --id band --stage frame --arms gpt2,sd5 [--spend]
python3 $O pick --job $J --name broll_band --product --file $J/broll/<pick>.png # refuses <7/10 / wrong display
python3 $O broll --job $J --id band --stage anim [--spend] # (fallback: real clip → broll/band.mp4)
python3 $O broll --job $J --id <cutaway> --stage frame [--spend] ... --stage anim [--spend]
# G4: assembled cut (A-roll track = the intercut)
python3 $O assemble --job $J --hook H1; python3 $O cost --job $J
python3 scripts/meta_preflight.py $J/final/<slug>-H1.mp4
- Cost B (kie): Omni buckets $0.315 / $0.42 / $0.525 / $0.63 for 4 / 6 / 8 / 10 s; images ~$0.05 (gpt2) - $0.075 (sd5). A 45 s ad ≈ 26 lines mostly at 4 s ≈ $8.5-9.5 + 3-4 b-roll ≈ $1.5 + 2 avatars × 3 rolls × 2 arms ≈ $0.75 + lock 2 rounds ≈ $2.5 → ~$14-17 with regens. The minimum viable test (§11) ≈ $2-3.
- Expected
boardwarnings in B: "~7s straight on the A-roll face" across a speaker change (the face changes) and product-vs-scene-change distance. Errors still block. - No draft→final step in B: the G2 3× lock is the proof; the approved prompt makes every final take.
- Image arms: pass
--arms gpt2,sd5(omni.py's default still lists nbp; its owner changes that).
10. Gates & QA checklist (on top of core §4)
- [ ] Parent named with CPA; ONE variable stated.
- [ ] Stake/accuser line sourced (stake bank row or VOC id); the mornings caused it; it lands on her.
- [ ] Banter rules pass (beat-sheet §Banter): no turn >12 words, no speaker >2 lines in a row, HOST speaks within every 6 s.
- [ ] Root-cause and mechanism lines verbatim; product facts per product-truth.md; red-team done.
- [ ] Product ≤20% of runtime and only in beat 5+; never on a speaker's wrist; B: the band cutaway passed ugc-omni's product gate (or is real footage).
- [ ] A: every studio line has a listener / profile / OTS intent; vo_check PASS; storyboard approved; motion prompts reviewed.
- [ ] B = ugc-omni G0-G4 (every gate Fish's): G0 script + hooks +
board(no paid step before it) → G1 each speaker's real ref frame + avatar (genetic change, not the ref person; real clothes, no grey tops, bare wrists; same room family, opposite eye-lines) → G2 lock: a MOM line good 3× (the 9-30 motion proof), then the same lock on a HOST line → G3 every line watched (verifyclean or flags accepted; the two voices distinct and stable by ear) + b-roll and band checked by eye vs the real pack → G4 the assembled cut by eye and ear, "Dramatization" label. Scores never pass a gate alone. - [ ] kie credit checked before any
--spend(A: fal/Higgsfield balances). - [ ] Captions: no speaker tags ([MOM]/[HOST]) in caption text; short lines (~25 characters at 9:16); B: captions come from
assembleinlines[]order (starpophow-to-caption-a-podcast-from-a-transcript). - [ ] Final: captions from the script, end card (optional,
--end-card; off by default per Fish 10-06),meta_preflight.pyon the upload file, Fish watched it. Eyes, not scores (Gemini QA grades photoreal against a cartoon rubric). B: AI-generated people disclosed per Meta's rules; never presented as a real named customer.
11. Variants & test plan
- Minimum viable test (B, simplest first, G1-G2 ≈ $2-3): MOM ref + avatar →
lock --n 3on her hook line → Fish watches. If yes, HOST frame +lock --frame hoston one HOST line → then the full job. No keyframe chains, composites or upscaler; add a layer only when this fails (corrections 10-05). - Test 1 (A): a 45 s podcast of a paying song's story (oneweek or exhibit A), same style as the parent. Its own ad set in Testing, read at 7 days on CPA vs the parent.
- One variable at a time after that: A vs B (same script); HOST role (fellow mom vs nurse-mom); runtime 30 vs 60 s; intercut vs split-screen; full-frame 9:16 vs letterboxed crop (B, ffmpeg, §7). Side arm (one variable vs the locked omni baseline): H3 native-voice takes via
yapper.py. - 3 hooks per body:
- HOST's reaction to the stake ("Wait, the letter had your name on it?");
- MOM stating it ("The school sent the truancy letter to me.");
- the accuser's words quoted by MOM.
- Swipe-file openers to test as hooks (claims, not proof; starpop
podcast-ad-examples-swipe-file): HOST opens skeptical ("I'll be honest, I didn't buy it") and gets converted by MOM's story; or HOST asks MOM a private question the viewer answers in her head ("What time does he actually get up?"). A nurse-mom HOST wears scrubs: the wardrobe carries the authority before a claim. - Graduation per
feedback_graduation_rule; launch memory per core §6.
12. Failure modes
- B: hook morphs into the other woman at its end → the hook line lacks
end_frame= its own frame. - B: words cut off or stretched → the line is in the wrong bucket; split it or let
plan_linepad (a 10-word line at 4 s gets cut off, 10-06).verifyflags transcript <0.90, holds, cut-offs. - B: model reads directions aloud → only the line goes after "says:"; never quote dialogue in
scene/rules. - B: model draws the camera or phone → "no camera visible" in
prompt.rules(one change per roll; re-lock). - Mispronounced acronyms:
job.say(A-D-H-D); captions keep the script spelling. - B: the two women converge on one voice → recast one frame; never a shared
voicepreset. - Cuts read as two separate videos: mismatched rooms, light or eye-lines. Same ref family, same setting text, opposite eye-lines.
- It turns testimonial: the accuser and vindication are missing, or product talk runs past 20%. Rewrite to the beat sheet.
- Over-engineering creep (chain keyframes, composites, per-segment direction) is banned in v1. That's what broke yapper.
- A: a voiced character shown closed-mouthed on her own line: re-intent to the listener. A lit band: apply the wrist rule and regen.
13. Sources
~/research/davidaistar/REPORT.md §1-2, analysis/diff_podcast_dialogue.md, gemini_GIczkree_W0.md, gemini_pYNC49BHsHQ.md, ugc_seedance.md (pYNC49BHsHQ), diff_ai_ugc.md, txt/GIczkree_W0.txt; Foreplay 10_character_arguments, 09_street_interview. Code: scripts/yapper/yapper.py + settings.py, scripts/cartoon_h3/{run.py,pack.py,vo.py,vo_check.py,h3.py,endcard.py}. Memory: yapper 10-04, avatar-lane, reference-images, ugc-wardrobe, corrections 10-04/10-05. MAP.md, stake bank.