format-spoken-drama — accused, tested, vindicated (spoken)
1. Read first
knowledge/ad-formats/_FORMAT-CORE.md(doctrine, lane limits, gates, launch). Binding.references/beat-sheet.md(the skeleton + line rules),references/sources.md(reference ads).output/dawn-combo-map-2026-10-05/MAP.md§2A (what wins in story formats) and the stake bank it points to.knowledge/ad-formats/song-ad-playbook.md§1-2 (concept + red-team method, shared with songs).- Dialogue pack conventions:
output/concept-packs/dawnbands-s2-attendance-office-h1-2026-10-01.md(voices, pace keys,[SPEAKER]lines).
2. When to use / when not
- Use to answer the open question from the combo map: do the song winners pay because of the SONG or because of the STORY? Same parent story, spoken, one variable (audio form). Also when a concept has two strong voices (accuser vs mom) that a song flattens.
- Evidence status: untested on Dawn. Recipe A is proven in song form (exhibit A ROAS 2.48, oneweek 1.81). Long-form AI drama is scaling for other brands (Foreplay: Koriderm 4:32 drama; Rosabella 11:51 mini-movie). Lane signal only (not Dawn): the 9-30 Curvelle Side Effects spoken cartoon shipped the night its photoreal version failed.
- Don't use for callout/brand-voice angles (that's recipe B: statics/natives), for stakes the mornings didn't cause, or for a dad/kid narrator before the existing dad-POV song reads.
3. The format
- 9:16, 2:00-4:00, animated (claymation / disney3d / pixar / paper — one per ad, anchored on a real winner frame per song-ad-playbook §4).
- Mom narrates in first person (EL voice, the avatar: ordinary American mom 35-54) and carries ~60-70% of the words. 2-3 other voices speak their key lines: the accuser (husband / ex / school official / mother-in-law), the explainer (nurse / doctor / attendance officer / neighbor), optionally the kid at the payoff.
- Captions from the script on every word; hook overlay on frame 0; motion end card (optional,
--end-card; off by default per Fish 10-06) (packend_card:). - Reference shape: Foreplay
04_drama(Koriderm) — accusation scene, humiliation, secret discovery, vindication, product as the turning point; ours keeps the vindication payoff and puts the accuser close to her.
4. Why it sells
- Recipe A: stake on HER (blamed for the mornings) → resentment/humiliation → a credible person INSIDE the story explains the root cause → test (he does the mornings for a week / an exhibit in front of an authority) → vindication. The product is the evidence that clears her, not a pitch.
- Problem-aware: line 1 names her situation (blamed for the late mornings) in action. Hearing the accuser's line in another voice raises the stakes versus her quoting it.
5. Beat sheet (full rules in references/beat-sheet.md)
| Beat | % runtime | Job | Must contain |
|---|---|---|---|
| 1 Accusation hook | 0-5 | Stop the scroll on her | Accuser's verbatim line (VOC-sourced) in HIS voice, or mom quoting it in line 1; the setting in action |
| 2 Her world | 5-25 | Prove she's the alarm | Specific morning ritual (knocks, trips, times), exhaustion, what she's tried (generic failed solutions, no brands, never bed shakers) |
| 3 Escalation | 25-40 | Raise the stake | A witness or authority sees the blame (school letter, mediator, family) — the stake now costs her something |
| 4 The test | 40-60 | Hand the mornings over | "Then you do it" / "let's see" — accuser fails on screen, in dialogue |
| 5 Explainer | 60-72 | Root cause | Credible person inside the story says sound doesn't reach a deep-sleeping brain, in their voice |
| 6 Discovery | 72-82 | Mechanism + product | Touch gets through; silent to everyone else; keeps going until he turns it off. Product named |
| 7 Vindication | 82-95 | Payoff on HER | Accuser concedes / authority clears her / kid wakes himself; one line in the accuser's or kid's voice |
| 8 CTA | 95-100 | Offer | Spoken: product name, 60-night trial; end card (optional, --end-card; off by default per Fish 10-06) |
6. Pack spec
-
Step 0 — combo plan (feedback loop,
workflow-combo-loop):python3 scripts/combo_loop/plan.py --format format-spoken-drama --brand dawn→ pick one combo (a paying parent + ONE changed dimension) and paste its block into the pack front matter / job or spec JSON ("combo": {...}; no file →tags.py register --prefix):combo_avatarcombo_anglecombo_povcombo_authoritycombo_stakecombo_stake_oncombo_emotioncombo_root_causecombo_mechanismcombo_payoffcombo_devicecombo_formatcombo_parentcombo_variablecombo_ad_prefix(= the launched Meta ad-name prefix).python3 scripts/combo_loop/tags.py check <pack>must PASS before concept approval; the weekly refresh reads results back by that prefix. -
Front matter: core §5 keys (
style/style_key,style_block,visual_negative,style_anchor,cast,envs,product_truth: packargs exits without a style) plusformat: cartoon-h3,format_skill: format-spoken-drama,parent:(the song/static it mirrors, with CPA),iteration:(the one variable, e.g. "song → spoken, same story"),voices: MOM=<id>, <ACCUSER>=<id>, <EXPLAINER>=<id>[, KID=<id>], pace keysvo_tempo: 1.08,gap_same: 0.15,gap_turn: 0.3,max_gap: 0.22,seam_gap: 0.25(the 10-01 dialogue values that passed vo_check),style_<SPK>0.2-0.35 for heat on the accuser,end_card,title,overlay. ## VO Script:N. [SPK] "line", one sentence per line, 5-12 words. Untagged lines are not allowed in this format (tag the narrator[MOM]).## Scene visual intents:N. **role** — <env>: <who is ON SCREEN, doing what, framing>— and for every non-MOM line, WHO IS LISTENING (see §7).- Cast: mom + accuser + explainer (+ kid). Make similar-age characters visually distinct (hair, colour, build) and keep them out of the same batch where possible.
- Envs: 4-7 reusable sets (kitchen, teen room, hallway, school office / mediation room, car, bedroom). Copy
product_truthfromoutput/concept-packs/dawnbands-song-nottardy-2026-10-04.md(display OFF and dark in every shot), never from the 10-01 attendance-office packs ("top face that can glow").
7. Visual grammar
- No lip-sync (core §3). Spoken lines are staged so no one has to mouth them: the LISTENER's reaction (mom's face when he says it), the speaker from behind/profile in a wide, an over-the-shoulder on the listener, a phone call (speaker off screen), a document/letter insert for official lines. Never a close-up of a speaker's face on their own line.
- Speaker staging (fixed 10-06):
storyboard.py_SYSnow stages every tagged line as voice-over (listener, back/profile wide, over-the-shoulder, phone/door, insert), never mouthed. Still open every non-MOM intent withCutaway:orListener shot:(same rule as format-courtroom / format-mini-movie;--no-llmuses intents verbatim), thenpython3 -m scripts.cartoon_h3.check_plan --lint output/cartoon-h3/<slug> --pack output/concept-packs/<slug>.md= 0 BLOCK before approval. - Narration lines (
[MOM]) = show the thing she describes, literally (fever-dream literal rule). - Accusation and vindication beats get the closest framing on MOM's face; the test beat gets wide comedy-of-failure shots (him at the door, the kid asleep, the clock).
- Product only from beat 6. Wrist down or resting, band side-on, display dark (wrist rule). The buzz is shown by the kid's eyes opening and small cartoon motion lines at the wrist, never a lit face.
- Clock hands only; no readable text on screens or letters (a letter = a shape with a red stamp).
- One style per ad; style anchor = a frame from our paying song in that style.
8. Audio
- ElevenLabs per speaker via the pack
voices:; mom's voice = the song avatar's speaking twin (warm, slightly husky, 35-54). Accuser: dismissive, not cartoon-villain. Explainer: calm, certain. - Pace: every line is tagged, so vo_check runs its DIALOGUE profile on the whole pack: overall ≥2.9 wps, any 6+-word line ≤4.8 wps (FAIL), speaker-change gaps 0.30-0.55 s, same-speaker 0.12-0.30 s, dead air ≤0.6 s, ASR round trip. MOM's narration should still sit ~3.4+ wps by ear (not gated). Never fix a line by speeding it; split or re-take it.
- One ducked music bed (assemble default) — warm-tactile library; no hits, no swaps. SFX are allowed for the alarm/buzz beats only if they don't mask words.
- End card (optional,
--end-card; off by default per Fish 10-06) 3 s (end_card:,end_card_secsdefault 3.0). The VO lane ends 0.5 s after the last word (no tail flag;--song-tailis song-only), and the card OVERLAYS the last 3 s, so the spoken CTA plays under it: write the CTA as the last ≥2.5 s and stage its shot as a cheap hold.
9. Production (free until §10 gate 4)
cd ~/.openclaw/workspace; P=output/concept-packs/<slug>.md; S=<slug> # refs in <slug>.args (--char-ref/--env/--product-ref/--characters)
A=(${(z)"$(python3 -m scripts.cartoon_h3.packargs $P --storyboard)"}); A=(${(Q)A}) # core §5: never type style/refs by hand
python3 -m scripts.cartoon_h3.run "${A[@]}" --stage vo # EL VO per speaker + vo_check gate (EL quota)
python3 -m scripts.cartoon_h3.run "${A[@]}" --stage storyboard # free; grep for mouths (§7); Fish reviews storyboard.md
python3 -m scripts.cartoon_h3.sheet <out.png> <prompt.txt> <style_anchor.png> [refs...] # new cast/room sheets, ~$0.15-0.30 each, Fish's go
touch output/cartoon-h3/$S/storyboard.approved # Fish only
A=(${(z)"$(python3 -m scripts.cartoon_h3.packargs $P)"}); A=(${(Q)A}) # every ref must exist now
python3 -m scripts.cartoon_h3.run "${A[@]}" --stage canvas --canvas-workers 4 # image-gen role (registry), paid ~$0.22/batch
python3 -m scripts.cartoon_h3.run "${A[@]}" --stage gen --prompts-only # Fish reviews motion prompts
python3 -m scripts.cartoon_h3.run "${A[@]}" --stage gen --spend --backend fal --model h3-max --batches 1 # smoke
python3 -m scripts.cartoon_h3.run "${A[@]}" --stage gen --spend --backend fal --model h3-max --gen-workers 8 --auto-regen 1
python3 -m scripts.cartoon_h3.run "${A[@]}" --stage upscale --hd # paid, ~$0.009/s
python3 -m scripts.cartoon_h3.run "${A[@]}" --stage assemble --hd # + (optional end card) + gemini_qa
render_sheets.shonly knows the 10-04 song cast (core §4): new drama casts usesheet.pyper sheet.- Cost (10-04 ledger scale, format-mini-movie §9): a 3:00 drama ≈ 18 batches ≈ 240 billed H3 s incl. regens = ~$12-19 H3 ($0.05/s table vs the likely real $0.08/s; read the fal page before quoting) + ~$5.5 canvases + ~$1.6 upscale + new sheets + EL quota.
- Shot tiers (core §2b) + draft→final (core §4 gate 6b). Tag every batch A/B/C before approval: run.py does it itself (
scripts/cartoon_h3/tiers.py, alsopython3 -m scripts.cartoon_h3.tiers <run>: A = first batch/hook, reveal/CTA/product beats, 3+ characters in a shot; C = no person in any shot; else B) intoplan.json+<run>/tiers.mdwith the per-batch fal cost; re-tag in review if the word match is wrong. Per-tier models:run.py --tier-models "A=h3-max,B=h3-max,C=h3"(fal only; default =--modelfor every tier, i.e. all H3-max).h3on B/C is a candidate until a one-batch side-by-side passes Fish's eye (registry "How to adopt"). Draft→final:run.py --draft= 480P ($0.05/s vs $0.08/s at 768P, fal page re-read 10-06) into<slug>-draft, seeded from the approved run (VO, plan, approval, canvases), never the shared-body cache, no upscale/end card (optional,--end-card; off by default per Fish 10-06); it prints the native 768P re-render command for the approved batches. Finals re-render natively, never a 480P→--hdupscale (Fish judged native sharper 10-04). A re-render is a new sample: judge staging and prompts on the draft, not the exact take. - Hooks: 3 hooks × 1 shared body (
body_start) — h1 to completion before h2/h3 (shared-batch race, cartoon-h3 SKILL.md).
10. Gates & QA checklist
- [ ] Parent named with CPA; ONE variable stated.
- [ ] Accuser line + stake from VOC/stake bank (cite the source line); mornings caused it; lands on her.
- [ ] Explainer is a person inside the story; root-cause line verbatim from Fish's ruled lines.
- [ ] Red-team pass done (timeline math, who-knows-what, real-life plausibility, product facts).
- [ ] Every non-MOM line has a listener or off-screen staging in its intent (no speaker close-ups).
- [ ] No product before beat 6; wrist rule in every product intent; no readable text.
- [ ] vo_check PASS; Fish listened.
- [ ] Storyboard grep for mouth/speaking/talking/narrator = 0 hits; storyboard approved; motion prompts reviewed; balances checked before
--spend. - [ ] Every batch tagged tier A/B/C in
storyboard.md; any draft→final pass followed §9's route (drafts never left in a shared-body cache). - [ ] Placement preview (starpop facebook-instagram-ad-sizes-and-safe-zones, how-to-preview-a-facebook-ad-before-publishing): scrub to the hook frame, the reveal frame and the end card (optional,
--end-card; off by default per Fish 10-06) on a 9:16 overlay; header, captions, product and offer sit inside Meta Reels' safe box (top 14 %, bottom 35 %, sides 6 % are UI = keep content in y 14-65 %) and survive the Feed 4:5 centre crop (~15 % lost top and bottom). Remotiontiktok/tiktokbox/dragoncaptions centre at ~68-74 % = inside the Reels bottom band on Meta: flag it to Fish, never re-burn on your own. - [ ] Final: freezes 0, dark 0, captions from script, end card only if
--end-cardwas asked (off by default, Fish 10-06), Fish watched it.
11. Variants & test plan
- Test 1 (the reason this format exists): spoken version of exhibit A or oneweek vs the song parent, same story/beats, same style. Read CPA + lifetime frequency at 7 days (core §6: not CTR alone). Never launch it the same week as format-courtroom's exhibit-A test.
- Then one variable at a time: accuser (school → CPS → own kid → boss, MAP.md §3 order), explainer (nurse vs pediatrician vs attendance officer), style (re-skin).
- 3 hooks per body: (a) the accuser's line in his voice, (b) mom stating the stake ("They put my name on a truancy letter."), (c) the moment of the test ("I handed him the mornings for one week.").
- Own ad set per concept in Testing; graduation per
feedback_graduation_rule. - Find the loss point before picking the next variable (starpop meta-video-ad-metrics-hook-rate-hold-rate-ctr-cpa): hook rate (3-s views ÷ impressions) ranks the 3 hooks; hook up + hold rate (ThruPlays ÷ 3-s views) down = the body doesn't pay off the hook (fix the beat-2 handoff, not the hook); good hold + low CTR = the product lands too late or unclear (next variable = an earlier reveal or a spoken mid-CTA, not a new story); good CTR + bad CPA = page/offer, not the creative. CPA vs the parent stays the verdict.
- Upload hygiene (starpop are-ai-generated-ads-allowed-on-meta): Advantage+ creative enhancements stay OFF (Meta's tools have added music and altered products on finished ads). The launchers copy
degrees_of_freedom_specfrom a model ad and silently drop it if that post fails, which falls back to Meta's defaults: check the created creative. Day 1: open the live preview in Feed AND Reels for added music, re-crops or altered frames.
12. Failure modes
- Speaker shown saying a line with a closed mouth → staging rule violated; re-intent to the listener.
- Two voices that sound alike → pick contrasting EL voices; check in the vo_check listen.
- Dialogue gaps balloon into dead air → pace keys; split long lines.
- Accuser turns into a cartoon villain → the audience sides with no one; keep him reasonable and wrong.
- Lit band on the wrist → wrist rule; regen the batch with wrists down.
- Explainer delivers a lecture → max 3 lines, plain words, one image (a brain asleep under the noise).
13. Sources
REPORT + diffs: ~/research/davidaistar/REPORT.md, analysis/diff_podcast_dialogue.md, analysis/diff_copy.md; Foreplay foreplay/04_drama.gemini.md, 14_minimovie.gemini.md; MAP.md; song-ad-playbook.md; 10-01 attendance-office dialogue packs; project_dawn_yapper_2026-10-04 (native speech route).