EVOLVE production handoff
These are subordinate implementation notes. EVOLVE defines what to make. Tool limitations must be disclosed or fixed; they cannot silently change its script, sources, timing or test design.
Sourcing and shot planning
- Complete intake, evidence review, modular script and editorial checks before sourcing. Use the current approved script version, not an old batch's storyboard.
- Search Gridbank first for stock; Storyblocks and Pexels are alternatives. Genuine existing product footage may supply product shots. Use Higgsfield when appropriate stock does not exist (Ayden's September 15 provider override, until changed). Do not automatically fall back to fal.ai. Verify provider availability before promising clips or spending money.
- Record every chosen asset's source, permitted use, cast/subject, exact window, overlay and the VO line it supports. Do not treat a download or a successful image-model review as usage permission.
- Inspect the exact used windows, including their motion. Keep depicted identity consistent when presenting one person's transformation; unrelated actors must not masquerade as that same person. Preserve the same shared body across A/B/C.
- Clean stock product shots, concise explanatory animation and professional-looking stock are allowed. No mandatory rip-first search, X-ray edits, screenless product, delayed reveal, new cast for each variant or fresh-footage quota. Product appearance comes from current product references.
- Hook: immediate visual impact. Bridge: 2–3 problem cuts in five seconds. Hold: clear product/mechanism/proof and behavioral outcome. CTA: clean product/price/guarantee/action. Cut to meaning; the old universal 1.4–2.8s cut law is not creative authority.
Existing tools — capabilities, not requirements
scripts/clip_library.py: tagged library and asset registration.scripts/register_dawn_broll.py: registers supplied Dawn footage; its tags do not prove permission or customer identity.scripts/pull_broll_clips.py: includes Pexels and organic-video sourcing. Its historic phone-footage/stock-look screening can reject valid EVOLVE shots; inspect before using it for stock. Gridbank and Storyblocks integrations are not established by this document.scripts/higgsfield_pipeline.py: existing Higgsfield generation backend. Check current supported models and account access at execution time; do not assume historical unlimited entitlement or pricing.scripts/vo_broll_runner.py: ElevenLabs VO, word timing, pinned clip windows, grading, scene assembly and master audio.scripts/vo_audio_chain.py/scripts/vo_grade.py: existing audio/visual finishing.scripts/gemini_qa.py: pass a style matching the actual B-roll output; its default cartoon rubric is inappropriate.
Use accurate on-screen text derived from the script, not whatever text happened to be burned into a source. Keep a deliberate hook headline visible long enough to scan, plus synchronized captions and offer overlays. The current runner does not automatically supply all these editorial elements.
Shared-module assembly
Author and produce the universal hold and CTA once. Reuse the same footage, clip windows, narration, captions, overlays and music within those modules. Assemble three matched intros against them. Do not generate a new hold per variant or rely on equivalent wording.
Current runner compatibility:
- Existing packs are separate scene JSON files; they have no native universal-module object.
- A manual shared-scene workflow can copy A's completed shared scene files/state into B/C and build only their hooks/bridges. Inspect every required shared scene before stitching: the runner can omit missing unselected scenes.
- Full-ad music is currently mixed separately. Different intro lengths shift the music under the shared body. For a strictly controlled test, use fixed intro duration and verify the shared final body, or use an assembly path that reuses an already mixed shared body.
- Hash shared scene files and compare current script/overlay/clip inputs. Scene hashes alone do not prove the final mixed soundtrack is identical.
VO, cache and timing checks
- Choose voice/delivery for the avatar and script. No universal narrator-only/mom-only rule or forced fixed speech speed. Listen for natural fifth–seventh-grade delivery and intelligibility.
- Measure hook audio, bridge audio, hold and CTA against the selected windows. Edit/revoice overlong copy; never make a 20-word hook pass by calling it a different beat.
- Known cache hazard: existing VO is reused without comparing its text/settings. After a script or voice change, rebuild the affected scenes with
--force-vo;--forcealone does not ensure new narration. Regenerate shared scenes once, then refresh them in every variant. - Known missing-scene hazard: partial builds can omit unbuilt unselected scenes. Confirm the final manifest contains every expected scene exactly once and in order.
- Known timing gap: the runner does not enforce 3-second hooks, five-second bridges or a 45-second final cap. These are required operator checks until automated.
- Known rendering gap: the audited packs requested 1080×1920 but delivered 720×1280 at 24fps, while some runner timing math uses 30fps. Probe actual outputs and use actual render timing. EVOLVE requires mobile-first readability, not a particular resolution; do not claim 1080 delivery from a pack field.
These limitations were observed September 15, 2026. Recheck code after future repairs; do not turn workarounds into permanent creative rules.
Final QA — required before “production-ready”
- Script/brief/source versions match the delivered VO and storyboard.
- All three hooks: 3–7 words, fit 0–3s, readable headline ≤2 mobile lines, immediate first 1–2s impact.
- All bridges fit 3–8s and enter the same hold smoothly.
- One hold cycle, one mechanism, one CTA, shared content unchanged across variants.
- Every VO line matches the visual behavior and overlay at that moment; inspect actual frames, not only clip tags.
- Show transformation; confirm product identity and actual evidence attribution.
- Offer end card has the real price/terms/action, readable on mobile and consistent with the destination.
- No talking heads/direct-to-camera. No accidental unrelated cast, source watermarks, frozen filler, black flashes, misaligned captions or speech clipping.
- Final duration measured at 30–45s inclusive; section boundaries and soundtrack checked in each final.
- All claims have applicable evidence; platform review done for actual wording.
- Record measured timing, module checks and remaining issues in the output receipt. No assumed QA pass.
Local script and production preparation can proceed within the user's task. External publication, sending files to other people, and campaign/budget changes require the relevant explicit authorization. This skill does not send to Discord or deploy a hub session automatically.
Learning loop
After a test, update the batch record with spend, purchases, primary KPI, link CTR, defined hook/hold rates, placement and evaluation dates. Record winner/loser/inconclusive with sufficient-delivery context and what to change in the next hook, bridge, hold or offer. When testing a new universal hold, make a new batch; do not mutate the common body inside an ongoing hook test.
Clock and footage laws (added 2026-09-15 after the first EVOLVE batch audit)
The first EVOLVE batch (output/vo-broll/evolve-new-docs-2026-09-15/) padded every VO line with silence to its script-table slot (0-3 / 3-8 / 8-30 / 30-40) and shipped 12 s of no-voice time per 40 s ad, one stock clip for three ads, and a 6 s static end card. Fish: "slow, parts with only b-roll or only voice, b-roll doesn't match." These laws prevent the re-derivation; see audit-2026-09-15.md in that batch and the v2 rebuild (v2/render-v2.py).
- The voice sets the clock; slots are ceilings. A segment ends at its last word + 0.08 s and the next starts ≤0.15 s later. "Hook 0-3 s" means the hook headline may stay on screen up to 3 s, never that the audio is padded to 3 s. A 90-word script at lane pace is a ~24 s ad; if the batch needs 40 s, write more words, never more silence.
- VO at native speed 1.2, one take per line, no per-segment atempo. Intra-line pauses above 0.22 s are cut (
tighten_pauses). Pace receipt ≥3.0 wps on the final. - Storyboard per phrase before any pin (
storyboard.md: phrase → need → clip@window → subject in frame), subject noun of the phrase = subject of the frame. Windows eyeballed frame-by-frame before pinning; a script line whose visual is listed as an "illustrative mismatch" is a defect before build, not a QA note. - Own footage first, then a per-phrase stock sweep. Every mom/kid phrase gets its own Pexels query and ≥2 eyeballed candidates; one sweep per batch is a defect. Cast per ad = one mom actress, one kid, one hand set for product demos.
- End card ≤2 s, only under the final "Tap below…" phrase; the rest of the CTA plays over footage.
- Gates on any renderer used (the batch's own or the runner's): voice gaps ≤0.40 s measured on the voice stem, freezedetect=0 before the card, wps ≥3.0, card ≤2 s, contact sheet at 1 fps eyeballed for overlay collisions. A Gemini score is reported, never adjudicated into a pass.