Pipeline Hub

vo-broll

Create, review, and produce Facebook/Instagram video ads using the EVOLVE Final Edition workflow: feedback and document analysis, three hook/bridge pairs, one shared hold and CTA, B-roll plus AI voiceover. Also use for sourcing or assembling footage for these ads. Explicit requests for unrelated animation formats use their own production tools.

Lane: — · not built yet: 0 (see below) · source: /Users/ayden/.openclaw/workspace/skills/vo-broll

SKILL.mdevolve-contract.mdevolve-final-edition.mdproduction-handoff.mdEvals (0)Learnings (0)

EVOLVE production handoff

These are subordinate implementation notes. EVOLVE defines what to make. Tool limitations must be disclosed or fixed; they cannot silently change its script, sources, timing or test design.

Sourcing and shot planning

  1. Complete intake, evidence review, modular script and editorial checks before sourcing. Use the current approved script version, not an old batch's storyboard.
  2. Search Gridbank first for stock; Storyblocks and Pexels are alternatives. Genuine existing product footage may supply product shots. Use Higgsfield when appropriate stock does not exist (Ayden's September 15 provider override, until changed). Do not automatically fall back to fal.ai. Verify provider availability before promising clips or spending money.
  3. Record every chosen asset's source, permitted use, cast/subject, exact window, overlay and the VO line it supports. Do not treat a download or a successful image-model review as usage permission.
  4. Inspect the exact used windows, including their motion. Keep depicted identity consistent when presenting one person's transformation; unrelated actors must not masquerade as that same person. Preserve the same shared body across A/B/C.
  5. Clean stock product shots, concise explanatory animation and professional-looking stock are allowed. No mandatory rip-first search, X-ray edits, screenless product, delayed reveal, new cast for each variant or fresh-footage quota. Product appearance comes from current product references.
  6. Hook: immediate visual impact. Bridge: 2–3 problem cuts in five seconds. Hold: clear product/mechanism/proof and behavioral outcome. CTA: clean product/price/guarantee/action. Cut to meaning; the old universal 1.4–2.8s cut law is not creative authority.

Existing tools — capabilities, not requirements

Use accurate on-screen text derived from the script, not whatever text happened to be burned into a source. Keep a deliberate hook headline visible long enough to scan, plus synchronized captions and offer overlays. The current runner does not automatically supply all these editorial elements.

Shared-module assembly

Author and produce the universal hold and CTA once. Reuse the same footage, clip windows, narration, captions, overlays and music within those modules. Assemble three matched intros against them. Do not generate a new hold per variant or rely on equivalent wording.

Current runner compatibility:

VO, cache and timing checks

These limitations were observed September 15, 2026. Recheck code after future repairs; do not turn workarounds into permanent creative rules.

Final QA — required before “production-ready”

Local script and production preparation can proceed within the user's task. External publication, sending files to other people, and campaign/budget changes require the relevant explicit authorization. This skill does not send to Discord or deploy a hub session automatically.

Learning loop

After a test, update the batch record with spend, purchases, primary KPI, link CTR, defined hook/hold rates, placement and evaluation dates. Record winner/loser/inconclusive with sufficient-delivery context and what to change in the next hook, bridge, hold or offer. When testing a new universal hold, make a new batch; do not mutate the common body inside an ongoing hook test.

Clock and footage laws (added 2026-09-15 after the first EVOLVE batch audit)

The first EVOLVE batch (output/vo-broll/evolve-new-docs-2026-09-15/) padded every VO line with silence to its script-table slot (0-3 / 3-8 / 8-30 / 30-40) and shipped 12 s of no-voice time per 40 s ad, one stock clip for three ads, and a 6 s static end card. Fish: "slow, parts with only b-roll or only voice, b-roll doesn't match." These laws prevent the re-derivation; see audit-2026-09-15.md in that batch and the v2 rebuild (v2/render-v2.py).

  1. The voice sets the clock; slots are ceilings. A segment ends at its last word + 0.08 s and the next starts ≤0.15 s later. "Hook 0-3 s" means the hook headline may stay on screen up to 3 s, never that the audio is padded to 3 s. A 90-word script at lane pace is a ~24 s ad; if the batch needs 40 s, write more words, never more silence.
  2. VO at native speed 1.2, one take per line, no per-segment atempo. Intra-line pauses above 0.22 s are cut (tighten_pauses). Pace receipt ≥3.0 wps on the final.
  3. Storyboard per phrase before any pin (storyboard.md: phrase → need → clip@window → subject in frame), subject noun of the phrase = subject of the frame. Windows eyeballed frame-by-frame before pinning; a script line whose visual is listed as an "illustrative mismatch" is a defect before build, not a QA note.
  4. Own footage first, then a per-phrase stock sweep. Every mom/kid phrase gets its own Pexels query and ≥2 eyeballed candidates; one sweep per batch is a defect. Cast per ad = one mom actress, one kid, one hand set for product demos.
  5. End card ≤2 s, only under the final "Tap below…" phrase; the rest of the CTA plays over footage.
  6. Gates on any renderer used (the batch's own or the runner's): voice gaps ≤0.40 s measured on the voice stem, freezedetect=0 before the card, wps ≥3.0, card ≤2 s, contact sheet at 1 fps eyeballed for overlay collisions. A Gemini score is reported, never adjudicated into a pass.