Pipeline Hub

cartoon-h3

VO-first, batched multi-shot cartoon ad lane on MiniMax H3 — one full ElevenLabs VO sets the clock, a table storyboard from word timings, ≤15 s batches, one contact-sheet canvas per batch, the VO slice fed INTO the generation. Use for any voiceover-led animated ad (claymation, 3D, paper). Triggers on "cartoon h3", "h3 lane", "batched cartoon", "minimax cartoon", "vo-first cartoon".

Lane: — · not built yet: 0 (see below) · source: /Users/ayden/.openclaw/workspace/skills/cartoon-h3

SKILL.mdEvals (0)Learnings (0)
MANDATORY CO-LOADSPIPELINE (`scripts/cartoon_h3/run.py`)COMMANDSSHARED BODY (3 hooks × 1 body, added 2026-09-16)HOOKS × LEADS MATRIX (added 2026-10-02)ANIMATION HUBCREDIT DISCIPLINE — Higgsfield H3 (Fish 2026-09-16, binding)BACKENDS + COST (verified 2026-09-11)VO + DIALOGUE CLARITY — BINDING (Fish 2026-10-01: "words must not clump or get hard to understand", "make sure these checks are always in the rules")BINDING RULESKNOWN FAILURES (from the source skill + ours)

Long-form SONG ads (Fish 2026-10-04, binding): read knowledge/ad-formats/song-ad-playbook.md first. Concept = iterate the paying song beat sheet (accuser blames mom → vindication); lyric red-team; Suno via kie in the AVATAR's voice (a mom), distinct style string per song, lyric_check + Gemini voice check per take; --pick N --head 0.05 (first vocal ≤0.3 s); one art style per song anchored on a real winner; the scene-type table; H3 lit-band pose rule; chunked render; own floored ad set.

Evolve copywriting rules — BINDING for all copy (Fish, 2026-09-16). Before writing any hook, headline, script, body, lander or listicle copy, read /Users/ayden/.openclaw/workspace/knowledge/evolve/copywriting-rules.md (the card) and the matching section of /Users/ayden/.openclaw/workspace/knowledge/evolve/copywriting-skim.md (hooks/open loops/bridges/objections/pitch/stakes/analogies → "## 04 Strategies"; avatar→angle→mechanism→authority→awareness → "## Copywriting Fundamentals"; facts/fluff/spoken/6th-grade/beliefs → "## Copywriting Principles"; checklist + word lists → "## 09 Toolkit"). Run the 13-question checklist before anything ships. This skill's own doctrine stacks on top and never loosens it.

cartoon-h3 — VO-first batched animation lane

Adopted 2026-09-11 from davidaistar's MiniMax H3 workflow (analysis: knowledge/cartoon-h3-workflow-analysis-2026-09-11.md, his skill + transcript in knowledge/references/minimax-h3/). Fish: "this should be a whole new lane."

What changes vs cartoon-pipeline: the voice exists before any picture, clip length is never derived from word counts, several beats execute inside ONE generation, and the model hears the words it is cutting to. What does NOT change: Stage 1-2 of cartoon-pipeline (script doctrine + gates, char/style/product refs, product-truth rules) and the QA legs.

MANDATORY CO-LOADS

direct-response-copywriter + brief-builder (script) → cartoon-pipeline references (references/vo-script-framework.md, product fidelity, cp-016/018/019/025 VO gates) → image-prompt-engineer (refs) → this skill → qa-gate before Discord.

PIPELINE (scripts/cartoon_h3/run.py)

1. SCRIPT   existing concept pack (.md) through every existing gate — unchanged
2. REFS     char ref(s) + cartoonified product ref + ONE environment sheet per world
            (env sheet prompt: "isometric 3D floor plan view of the whole <room>, <style>")
3. VO       line-by-line EL takes (Jordan, per-line cache) → head/tail trim → internal gaps
            > MAX_GAP 0.15 s compressed (Fish 9-17; cached takes are recompressed on cache hit) → LINE_GAP 0.15 s
            between lines → words.json + lines.json (pack line → word span)
4. STORYBOARD  beats = sentences (split at a comma > 5 s, merge < 1.2 s), times from the words;
            batches = consecutive beats ≤ 15 s, ≤ 6 panels, ending on a PACK-LINE end;
            reveal/CTA beats cap their batch at 3 panels. Sonnet splits each scene intent
            into per-beat panels (25-60 words, concrete, no new characters/product).
            → storyboard.md  ⛔ FISH REVIEWS THIS BEFORE ANY CANVAS (touch storyboard.approved)
5. CANVAS   one 16:9 GPT Image 2.5 (fal) contact sheet per batch, hero refs attached, panel count in
            WORDS + grid + reading order + text ban; Gemini counts panels / reads text;
            fail → identical prompt regenerated (never a prompt rewrite first)
6. SLICES   ffmpeg cut of the VO at batch boundaries (asserted == batch span ± 0.06 s)
7. GEN      MiniMax H3 reference-to-video per batch: Image 1 = canvas, Image 2.. = refs,
            Audio 1 = slice ("timing source only, no lip-sync"); Shot N per panel, a named
            Transition per seam; ONE candidate per batch (run.py default since 9-17)  ⛔ --spend = Fish's green light
8. ASSEMBLE clips at batch offsets over the ONE VO, speed nudge bounded 0.9-1.1×, trim excess,
            hold = regen flag (never a pass); Remotion captions from EL timings; ONE ducked
            bed, no swap/hit/riser (9-11); receipts: nudge per batch, freezedetect, luma,
            final ends at last word + 0.5 s; gemini_qa --type final

Run dir: output/cartoon-h3/<pack-slug>/ (vo.mp3, words.json, lines.json, plan.json, storyboard.md, canvases/, slices/, gen/, fit/, -final.mp4, -final-discord.mp4, receipts).

COMMANDS

# storyboard only (generates the VO; no image/video spend)
python3 -m scripts.cartoon_h3.run --pack output/concept-packs/<pack>.md \
  --char-ref "label=<png>" --env "<room>=<png>" ... --product-truth "<one sentence>" \
  --style claymation --characters "Mom, Dad and Ryan" --stage storyboard
# after Fish approves storyboard.md
touch output/cartoon-h3/<slug>/storyboard.approved
python3 -m scripts.cartoon_h3.run ... --stage canvas        # GPT Image 2.5 on fal (canvas.py default)
python3 -m scripts.cartoon_h3.run ... --stage gen --spend --backend fal --model h3-max   # paid
python3 -m scripts.cartoon_h3.run ... --stage assemble
# 10-06 motion end card: automatic after assemble when the pack has `end_card:` (+ optional end_card_secs / end_card_clip /
# end_card_product); writes <slug>-final-endcard.mp4 and cuts the Discord copy from it. Skip: --no-end-card. Standalone:
python3 -m scripts.cartoon_h3.endcard <final.mp4> <out.mp4> [--secs 3] [--text "..."] [--clip <product-spin-on-navy.mp4>]
# 10-06 re-skin (same audio + beats, new style = one variable): lane song → --from-pack; outside-lane winner → --from-video
python3 -m scripts.cartoon_h3.reskin --from-pack output/concept-packs/<pack>.md --style <style>
python3 -m scripts.cartoon_h3.reskin --from-video <final.mp4> --script <lines.txt> --style <style> --slug <new-slug>
# higgsfield backend: discover the H3 job schema first (needs a live Clerk session)
python3 -m scripts.cartoon_h3.run --pack <pack> --backend higgsfield --discover

SHARED BODY (3 hooks × 1 body, added 2026-09-16)

Pack front matter body_start: <line> marks where the shared body begins. VO = hook take + body take (0.12 s seam); the body take, its word timings, beat descriptions, canvases and clips are cached under output/cartoon-h3/_shared-body/<hash>/ (hash = voice + body text) and reused by every pack with the same body. The body always starts a fresh batch. Run A first (fills the cache), then B, C. --env <name> without a path is allowed for --stage storyboard only. Higgsfield: --backend higgsfield --spend --allow-billing (H3 is billed, 2.5 cr/s @768p). Cache guard (10-02): a cached segment whose words match but whose scene intents differ now FAILS the storyboard instead of silently inheriting the other pack's visuals; make the intents identical or move the line out of the segment.

HOOKS × LEADS MATRIX (added 2026-10-02)

lead_start: <line> (with body_start) adds a cached LEAD between hook and body: lines 1..lead_start-1 = hook (per pack), lead_start..body_start-1 = lead (cached under _shared-body/lead-<hash>/), body_start+ = shared body. Renders H hooks + L leads + 1 body, whatever the number of cuts. A lead must land after EVERY hook it's paired with. 1. Write a variants file (## Hook h2 / ## Lead l2 sections; each entry = VO line, then its **ROLE** — intent line). 2. python3 -m scripts.cartoon_h3.matrix --base <approved base pack> --variants <file> [--lead-lines A-B] [--design anchor|full] --write → <concept>-h<i>-l<j>-<date>.md packs + output/cartoon-h3/_matrix/<concept>.json. anchor (default) = every hook on l1 + h1 on every lead (H+L-1 cuts, clean separate reads); full = H×L. 3. Fish approves every hook, lead and the body (script gate) → run h1-l1 first, then the rest (they reuse lead/body). 4. Ad names MUST be <concept>-h<i>-l<j> (the readout parses them). 5. python3 -m scripts.cartoon_h3.readout --concept <concept> [--days 14] [--write]: hook judged on hook rate (3 s views / impressions, cuts sharing l1), lead on hold rate (ThruPlays / 3 s views, cuts sharing h1), set on CPA vs breakeven; z-tested; flags "Meta picked" when one ad set starved the other cuts; prints where to test next and appends to knowledge/brands/dawnbands/research/hook-lead-ledger.md. Legacy names: --pattern 'sl0919-h(?P<h>\d+)'.

ANIMATION HUB

python3 -m scripts.cartoon_h3.hub --deploy (or any stage with --hub) → https://animation-hub.pages.dev — every run, every stage (script · VO · storyboard + gate state · canvases + Gemini check + prompt · slices · gen clips + cost · final + QA · receipts).

CREDIT DISCIPLINE — Higgsfield H3 (Fish 2026-09-16, binding)

Where credits went on the first run (~1,100 cr): 43.5 on a UI "Rerun" click, ~330 on second candidates (later dropped), ~140 on 5-panel batches H3 stacked (3 paid retries), ~36 on sheet retries the checker failed for LED digits. Rules that stop each: 1. Never click "Rerun" in the Higgsfield UI. It submits and bills instantly. Read params from GET /fnf/jobs?job_set_type=…. 2. One candidate per batch. clipqa.py is the only regen trigger; "readable text" on a phone/clock glow is a soft flag, never a regen. 3. ≤4 panels per H3 job (MAX_PANELS = 4). A stacked batch is not fixed by re-rolling: use splitgen.py (rows of the same sheet → 2 short jobs). 4. Smoke one batch first on every new session/backend; the payload is validated before a wave. A 403 is free, a wrong payload is not. 5. Regen with --regen N, never by deleting files: the cache pull restores a parked failed clip and the job is silently skipped (then paid again later). 6. Before re-rendering a sheet, re-check the existing fails with the corrected checker; promote what passes (allow_band_digits on product batches). Sheets cost ~3 cr, video ~2.5 cr/s. 7. Killing a gen process orphans billed jobs. Recover them with recover_jobs.py <ids> (prompt match) instead of resubmitting. 8. Prompt hygiene = fewer retries: product truth once (on the product-ref line), refs named by role, per-shot audio windows, no "storyboard" in a video prompt, style block + character sheet on every job, style-anchor sheet on every canvas. 9. Duration = ceil(span), 768p, never pad audio to fill; H3 returns ceil+0.4 s and assemble trims. Never buy 2K unless Fish asks. 10. Media cache (_hf-media-cache.json) is on: re-uploads cost time not credits, but a stale mtime = a new upload; don't touch ref files mid-run.

BACKENDS + COST (verified 2026-09-11)

Backend Model Price Notes
fal minimax/h3-max/reference-to-video $1.20 / 15 s (768P); promo $0.30 / 15 s until 2026-09-14 default; fastest, best adherence
fal minimax/h3/reference-to-video $0.06/s 768P · $0.13/s 2K when 2K is needed
higgsfield H3 in the standard sub 20 credits / 5 s, 60 / 15 s job type + params come from --discover, never guessed (scripts/cartoon_h3/higgsfield_h3.json)
Refs: ≤ 9 images, ≤ 3 audio (2-15 s each), 12 files total; audio needs an image beside it.

VO + DIALOGUE CLARITY — BINDING (Fish 2026-10-01: "words must not clump or get hard to understand", "make sure these checks are always in the rules")

Applies to EVERY cartoon-h3 pack, single narrator or dialogue. Enforced in code: vo_check.py runs inside run.py right after the VO and BEFORE the storyboard; a FAIL stops the run (--skip-vo-check exists and is never the answer). 1. Never fix pace by speeding the VO up or slowing it down. Fix a bad line by splitting it, re-punctuating it, or redoing the take. 2. Line length: no spoken line over 18 words in a dialogue pack; one idea per line; a long speech becomes 2-4 lines, each on its own panel. 3. Numbers are spelled with spaces in the VO text ("eight twenty two", "ten thousand", "sixteen"), never digits or hyphenated, so the take and the word timings stay aligned. 4. Dialogue packs: line tags [SPEAKER] "text", front matter voices: NAME=elevenlabs_id, ..., vo_tempo: 1.08, max_gap: 0.22, gap_same: 0.2, gap_turn: 0.4, seam_gap: 0.25, optional style_NAME. One voice per speaker, clearly different (age, pitch, energy). Every line is its own take; voices can never overlap. Each speaker's lines are level-matched on copies (the shared line cache is never edited). 5. The five checks (all written to vo-check.md / vo-check.json): - Pace: every line of 6+ words: dialogue FAIL above 4.8 wps (warn above 4.5, drag warning below 2.4); narration WARN outside 3.0-5.4. Overall ≥ 2.9 wps dialogue (turn gaps cost ~0.15 wps), ≥ 3.4 narration (FAIL). Hook lines may be synthesized slower with hook_tempo in the pack (synthesis speed, never a post-stretch). - Gaps: speaker change 0.30-0.55 s, same speaker 0.12-0.30 s (dialogue FAIL); dead air above 0.6 s dialogue / 0.4 s narration (FAIL). - Clump scan: more than 9 words inside any 2 s window warns, more than 11 fails (dialogue); narration warns above 11. - ASR round trip: Whisper transcribes vo.mp3 and the whole transcript is aligned against the whole script; any line of 3+ words missing more than 10% of its words (dialogue FAIL, narration WARN), or overall error above 5% dialogue / 10% narration, is redone. Brand and character names are exempt. - Human listen: Fish hears the full VO on the hub before the storyboard gate. 6. One speaker turn = one panel. Short reaction lines become reaction panels; a long speech gets cutaways that keep the speaker on screen before and after. No lip-sync in H3: the speaker is legible only through framing, gesture and speaker-colored captions. 7. Any VO re-time forces a timeline rebuild and an A/V sync check.

(Limits calibrated 10-01 on the first S2 run: the first draft limits sat below natural speech; the ASR round trip is the strict, objective gate.)

BINDING RULES

  1. VO before pictures. No duration buckets, no per-scene VO, no -shortest mux, no freeze pads.
  2. Batch boundaries only on pack-line ends. A batch that would cut mid-sentence pulls back.
  3. Product panels: batch ≤ 3 panels, product ref attached, wordmark/shape named in the panel text.
  4. Storyboard gate before canvases; --spend gate before generation. Both are Fish's.
  5. A clip is fit by speed within 0.9-1.1× and trim only. A held frame is a regen flag in the receipt.
  6. Every final: dead-air receipt, freezedetect = 0, luma scan = 0, gemini_qa final, ≤ 8 MB Discord copy.
  7. Captions ALWAYS come from the real script (Fish 9-28, song + VO): assemble rebuilds caption-words.json from lines.json text on the audio timings every run (caption_words.build); transcript words (Scribe/EL) are never shown, and a fallback logs a WARNING that blocks shipping.
  8. Inherited from cartoon-pipeline: product text fidelity, fever-dream literal VO, plain-language pass, product named at reveal + spoken CTA, no photoreal product into scene gen. (The end card's real product cutout is a post overlay, not scene gen.)
  9. Product reference = an in-hand / in-context shot (10-06, davidaistar teardown). A white-background-only product shot makes image and video models guess scale wrong. Every product sheet gets the real cutout PLUS one buyer photo of the band on a real wrist (10-05 rule: silhouette from real units, not supplier art).
  10. Improvise first, then fix with an edit pass. Never force exact copy, text or many constraints into one generation prompt: generate plain (2-3 sentence prompts, 10-04 rule), then fix the one wrong thing with an edit model on that image. Regenerating from scratch is the last resort.
  11. Restate the refs and the model on every call. When an agent or MCP drives the generation, defaults drift silently (wrong model, voice or refs). Each gen call names its char/env/product refs and model explicitly; the receipt records them.
  12. Re-skins change one variable. Audio, beats and intents stay identical (reskin.py checks the audio hash and timings); only style, refs and sheets change. storyboard.approved is never copied: Fish re-approves.

KNOWN FAILURES (from the source skill + ours)

Failure Fix
canvas split-screen / wrong panel count identical prompt regenerated (geometry lines are verbatim); then check refs count ≤ 9
clip stacks beats into one static frame never write "one continuous shot"; one Shot per panel + one named Transition per seam
character drift ref attached with role line + preservation clause in EVERY prompt; never described from memory
readable words on props "symbols without readable words" in the panel text; text ban stays; regenerate identical
clip early/late vs words normal; bounded nudge in assemble, else the second candidate
product drift in small panels reveal/CTA batch ≤ 3 panels (enforced in the batcher)
product truth pasted into every product shot ("no phone" on a shot that shows the phone) truth sentence ONCE, on the product-reference image line (h3.video_prompt, 9-16); the per-panel suffix the storyboard adds is stripped
"storyboard" in the video style block → grids/panels in clips video prompts use the style block minus "premium animated commercial storyboard"; Image 1 = "shot guide, one full-frame shot per panel, never a grid"
room sheets labelled as identity refs refs named by role in the video prompt: character sheet / product reference / room sheet
Gemini canvas check fails every product sheet ("readable text" = the band's LED digits) check_canvas(..., allow_band_digits=True) on product batches (9-16); clocks/screens/papers still banned — prompt says blank faces, plain glowing rectangles, blank papers
H3 submit 403 unlimited_generation_not_allowed use_unlim only when an entitlement covers the job (higgsfield_unlim.submit_unlimited, 9-16)
per-job re-upload of the same refs (2 min each) _hf-media-cache.json keyed by path+size+mtime, 20 h