Pipeline Hub

format-podcast

Two-mom podcast-clip ad, 25-60 s. One mom tells her blamed-for-the-mornings story at the mic, the other reacts and asks, and the band lands as the answer in a product-gated cutaway.

Lane: (A) cartoon-h3 dialogue (staged, no lip-sync); (B) photoreal = ugc-omni (one Omni take per line, two speaker frames, gates G0-G4); H3 native voice is a side arm only · not built yet: 0 (see below) · source: /Users/ayden/.openclaw/workspace/skills/format-podcast

SKILL.mdbeat-sheet.mdsources.mdEvals (7)Learnings (16)
datelearningsource
2026-10-06minimax/h3-max/image-to-video with no target_audio_url speaks quoted dialogue natively with lips (Scribe verbatim; Fish 'yea its good' 10-04). Proven single-speaker only; never run two speakers.project_dawn_yapper_2026-10-04
2026-10-06Fish closed the yapper lane 10-05 after the chained-keyframe build broke ('curvelle talking head with no effort on h3 was literally fine'). Photoreal podcast = simplest recipe (one frame, plain prompt, <=15 s take per speaker) and needs Fish's go to reopen.corrections-hot.md 2026-10-05
2026-10-06Generated band-in-hand frames fail (fused fingers, product 3-5/10); product = real footage or graphic. A podcast needs no hand-product shot, which is the format's case for photoreal.project_dawn_yapper_2026-10-04; diff_podcast_dialogue.md
2026-10-06davidaistar: one generation per speaker with all her lines (5 s model minimum, holds voice), cut line by line; plain 'Line 1/Line 2' prompt beat shot keywords; patch clipped last word with an EL single syllable; ~80% of realism is the script.~/research/davidaistar/txt/GIczkree_W0.txt
2026-10-06Not built in scripts/yapper: duo (2 personas), duo-assemble (intercut), patch-word, 16:9 frames (cmd_frame hardcodes 1152x2048), a podcast native take (chain-render --native hardcodes the car-pickup prompt and needs EL alignment + keyframes). No HOST persona in settings.py PERSONAS.scripts/yapper/yapper.py, settings.py (read 10-06)
2026-10-06H3 native: duration = ceil(words/3.8) or speech stretches; never quote anything but the lines (quoted direction gets spoken); saying 'phone' draws a phone; 1080P beats 768P+upscale by Fish's eye.project_dawn_yapper_2026-10-04 (10-05 entries)
2026-10-06cartoon-h3 assemble always mixes DEFAULT_MUSIC unless --music <file>; a podcast has no music, so pass a room-tone file.scripts/cartoon_h3/run.py:459, assemble.py:31
2026-10-06davidaistar's 'crop it uglier' = a tighter letterbox crop of the model's own 9:16 output (black bars like real podcast clips), not a 16:9 generation; our arm is ffmpeg crop=1080:1350 + pad to 1080x1920, no upscale. Still an untested later variable.starpop realistic-ai-podcast-ads-flux-3; podcast-style-ads-for-ecommerce
2026-10-06Hook tests from his swipe file: skeptical HOST who gets converted; HOST asks MOM a private question; scrubs carry nurse-mom authority. Not adopted: founder guest, talking-product animation, fictional show branding (Fish to rule).starpop podcast-ad-examples-swipe-file
2026-10-06Caption QA: no speaker tags in caption text, ~25 chars per line at 9:16, caption order follows the cut list across speakers.starpop how-to-caption-a-podcast-from-a-transcript
2026-10-06B intercut is ffmpeg (trim per cut-list row + concat), never a CapCut pass (core 2a)._FORMAT-CORE.md 2a; SKILL.md 9 B step 4
2026-10-06Built 10-06: --draft (480P into <slug>-draft, prints the native re-render command) and --tier-models/tiers.md in run.py; h3.py FAL_PRICE fixed (h3-max 768P $0.08/s, 1080P $0.16/s, + ref-token surcharge).scripts/cartoon_h3/run.py, h3.py, tiers.py (Builder B)
2026-10-06PodcastTwoShot Remotion comp BUILT 10-06: alternate (full frame per cut) or stacked, per-speaker caption colour, [MOM]/HOST: tags stripped at turn starts only (OK/TV/NO: mid-line kept), captions ≤64 %. Replaces the manual ffmpeg intercut; output is captioned, so skip a second caption pass.apps/orange-doc/src/FormatComps.tsx (Builder C)
2026-10-06sync-3 cartoon lip-sync proven 10-06: animated podcast hosts can talk on camera once wired (needs Fish's go).output/model-proofs-2026-10-06/README.md
2026-10-06Fish 10-06: photoreal lane reopened (talking heads = H3 native voice 1080P); photoreal cut carries a 'dramatization' label.Fish chat 10-06
2026-10-06Execution B moved from H3 native-speech singles (yapper, unbuilt duo command) to ugc-omni: two speaker frames (avatar=MOM, host=HOST) under one locked prompt, one Omni take per line, lines[] order = cut order so assemble builds the intercut. Band cutaway = product-gated omni b-roll (pick --product >=7/10), real footage fallback. Hook lines need end_frame = own frame. H3 native voice = side arm only.skills/ugc-omni/SKILL.md; Fish 10-06