Pipeline Hub

workflow-copy-research

Research-to-script workflow for Dawn VIDEO formats (song, spoken drama, podcast, courtroom, whistleblower, street interview, timeline, rated solutions, mascot explainer, skit): an on-demand VOC pull mid-brief into two grep-verified tables (ICP/persona + offering, every row a verbatim quote + source + date), 3 sourced hooks per concept, then the script in one of four modes (narrator beat sheet, two-person dialogue/banter, song lyrics with the versioned red-team method, or the long-form free-value explainer skeleton), ending in logic pass, Opus red-team, fix and re-check.

Lane: copy (feeds every format-* pack) · not built yet: 5 (see below) · source: /Users/ayden/.openclaw/workspace/skills/workflow-copy-research

SKILL.mdprocedure.mdsources.mdEvals (6)Learnings (9)
1. Read first2. When to use / when not3. Inputs4. Outputs5. Process6. Tools & models7. Gates & QA checklist8. Failure modes9. Sources

workflow-copy-research — evidence first, then the script

A format skill says WHAT the ad looks like. This skill produces the words that go into its pack: a research pull you can verify line by line, three hooks that each trace to a real person's words, and a script that has survived a red-team. It folds in davidaistar's research tables and free-value skeleton (analysis/diff_copy.md) and keeps our stricter parts (grep-verified quotes, concept formula, atlas, versioned red-team).

1. Read first

2. When to use / when not

3. Inputs

4. Outputs

Run dir output/copy-research/<slug>-<YYYY-MM-DD>/ (new slug per run; never overwrite a version):

File What
brief.md Concept frame, 4 slots with sources, awareness, governing desire, dominant emotion, mode (A-D), word budget
pull-log.md Avatar moment, 5-8 search strings, every thread/review ID pulled with URL + date, counts per tier
tables.md Table 1 ICP/persona + Table 2 offering, each row: verbatim + source path/URL + date + verifier hits; spot-check footer
hooks.md 3 hooks per concept, graded grid (source row, E/C/S, structure, gates)
script-v1.md … script-vN.md Versioned drafts (lyrics: lyrics-vN.md); v1 frozen as written
redteam-vN.md Opus red-team flags on version N (flags only, no edits)
script-final.md Approved script in the target pack's ## VO Script syntax + script stats

Raw pulls go to knowledge/brands/dawnbands/research/voc-pull-<topic>-<YYYY-MM-DD>/thread_<id>.txt (raw text only). Never write extracted quotes as .jsonl under research/: voc_check.py reads an explicit raw-source list (10-06; voc-pull-*/*.txt included) and must never "verify" our own extraction.

5. Process

  1. Frame (free). Fill brief.md: atlas row + job, parent + CPA, the ONE variable, the four concept slots with what is still missing. Pick the script mode: A narrator beat sheet (story formats), B dialogue/banter (podcast, courtroom, spoken drama scenes, street interview, whistleblower Q&A), C song lyrics (format-song), D free-value explainer (timeline, rated solutions, mascot explainer). Write the Evolve §0 answers (avatar → angle → mechanism → authority → awareness); one missing = don't write yet. Any brief field you filled without a file or source behind it carries an [assumed] tag; drafting (step 6 on) doesn't start until each tag is resolved by a source or by Fish (starpop how-to-build-a-claude-skill-for-ads: a brief-validation gate before writing, because a model will sprint past it).
  2. Corpus first (free). Search what we already hold before pulling: stake bank + its raw banks, dawn-voc-bank-final/dawn-voc-bank/voc.jsonl (774 quotes, 6 lanes), knowledge/brands/dawnbands/research/voc-competitor-reviews-2026-10-02.jsonl (1,552 competitor reviews), the Reddit thread dumps under research/, survey lines in research/customer-evidence-2026-09-15.md. scripts/voc_search.py is NOT usable for Dawn today: --brand dawnbands reads only knowledge/brands/dawnbands/voc-*.md (six ~900-byte compatibility pointers since 9-15) because voc-tagged.jsonl doesn't exist, so it ranks pointer text, not quotes. Use grep -i + the extended verifier (procedure §3) on the files above.
  3. On-demand pull (free; only for what step 2 lacks). Avatar moment → 5-8 queries in her words (SOP step 2) → Arctic Shift title search → full threads via knowledge/brands/dawnbands/research/reddit-mornings-2026-09-03/arc.py run inside the new voc-pull-* dir. TikTok comments under problem videos are a free Tier 2 source with dates (tikwm, procedure §1c item 3; starpop ai-customer-research-ad-copy mines reel and TikTok comments alongside Reddit). Reddit .json, old.reddit and WebFetch are blocked (memory reference_scrapling_reddit_access.md); skills/reddit-scraper/scripts/reddit_scraper.py uses the dead JSON route. Volume first, filter second. Minimums by format: procedure §1.
  4. Two tables + verify (free). Build Table 1 (ICP/persona) and Table 2 (offering) per procedure §2. Run the extended verifier (procedure §3) on every quote; 0 hits = the row is dropped or re-sourced, never kept "from memory". Byte-exact spot-check 10 rows against the raw file and write spot-check: n/10 under the tables. Cold-read gate on any quote used as mechanism support (SOP).
  5. Decode the reference structure (one Gemini call, cents, plus an ElevenLabs Scribe transcript on the EL quota unless --no-scribe) when slot 1 is a competitor/organic format: python3 scripts/ad_dissector.py <video|URL> --brand-context "<one line on the band>" [--competitor]. Map its beats to the format skill's beat sheet; --competitor output is a gap map.
  6. Hooks (free). 3 per concept, each traced to a table row or stake-bank row, graded on the grid in procedure §4. Fish picks the lead.
  7. Script (free). Draft script-v1.md in the chosen mode: A = the format skill's beat sheet; B = dialogue doctrine (procedure §5); C = lyric method (procedure §7, which defers to the format-song beat sheet and knowledge/ad-formats/song-ad-playbook.md §2); D = free-value skeleton (procedure §6). Every quoted customer line keeps its table row ID in a comment. One script = ONE Table 1 persona, named by row ID in brief.md (starpop ai-customer-research-ad-copy: narrowing to one persona sharpens the adaptation).
  8. Passes, in order (free; Opus = LLM writing role). Said-out-loud pass → Evolve 13-question checklist → humanizer skill (line edits, dash scan = 0; the one exception is a line-final cut-off "—" on a mode-B interruption, procedure §5 rule 4, which format-podcast uses too) → your logic pass (timeline math, who-knows-what, real-life channels, product facts, overlay = spoken/sung line, product first named no earlier than the format's reveal beat, word count inside the format budget: starpop create-and-clone-competitor-ads-with-ai names an early product intro and a long script as his clone's two flaws) → Opus red-team on the frozen version (subagent brief in procedure §8; writes redteam-vN.md, flags only) → fix as line edits into script-v(N+1).md → re-check that the fixes created no new contradiction (10-04: 3 of 4 red-team blockers came from the previous round's fixes). Stop after 3 red-team rounds; open items go to Fish.
  9. Hand-off (free). Convert to the target pack's syntax (N. [SPK] "line" for dialogue, N. "line" for songs/narrator), run the turn-stats snippet (procedure §5) for modes A/B, write script-final.md. Fish approves the concept + script (core §4 gate 1) before any VO, Suno or sheet spend.

6. Tools & models

Step Role (_MODEL-REGISTRY.md) Entry point Built?
Atlas row — ~/ad-batch-hub/scripts/atlas_label.py show / keys / add-concept built
Corpus search Gemini 3.1 Pro as a text ranker (same model as the video-understanding row; ID hard-coded in the script) scripts/voc_search.py built, but BROKEN for Dawn: reads only the voc-*.md pointer stubs (no voc-tagged.jsonl); fix = point it at voc.jsonl + the research dumps (code, Fish's go). Use grep + procedure §3
Reddit pull — (plain curl, Arctic Shift mirror) knowledge/brands/dawnbands/research/reddit-mornings-2026-09-03/arc.py <ids> built (manual search step)
TikTok comments — (tikwm, free) ~/tt-creative-sourcer sourcer.tikwm_get('/api/comment/list?url=…') (procedure §1c item 3) built (tested 10-06: text, digg_count, create_time)
Amazon/competitor reviews — existing voc-competitor-reviews-2026-10-02.jsonl; fresh pulls scripts/voc_amazon_stealth.py --config --output --brand built
Quote verification — scripts/statics/voc_check.py (procedure §3) built; corpus fixed 10-06
Reference decode Video understanding / QA (Gemini 3.1 Pro) scripts/ad_dissector.py built 10-06
Drafting, red-team LLM writing / planning: Opus for red-team, the session model for drafts; GPT-5.x via codex only as a challenger when Fish asks Agent subagent model: opus; pipelines use scripts/claude_cli.py run_opus built
Humanizer LLM writing ~/.claude/skills/humanizer/SKILL.md built
Singability / pace (downstream) TTS (ElevenLabs) / Music (Suno V5 on kie) format skill's vo_check.py / song.py receipt built, owned by the format skills

NOT BUILT (needs Fish's go): - voc_check.py corpus (BUILT 10-06): voc.jsonl (774), the bank's raw corpora, output/dawn-bottleneck-2026-10-03/reddit/comments/*.json, voc-pull-*/*.txt, buyer survey CSVs; curly quotes + whitespace normalised; each hit prints its file and category; competitor reviews and okendo_reviews_anon.json (94.6 % predate the first order) print as NOT COUNTED. CPS line #7 now scores 1 (199olsy.json). Still outside: the 9-23 212-row survey export (not on disk) and our own docs (customer-evidence-2026-09-15.md, on purpose). - A one-command research_pull.py (queries → Arctic Shift search → thread dumps → table skeleton). Today steps 3-4 are manual. - A dialogue lint inside check_plan.py; the turn-stats snippet is the manual stand-in. - Comment-for-link responder (davidaistar's engagement CTA). Without it, the engagement CTA is a question only. - The dialogue and free-value sub-modes in skills/direct-response-copywriter/references/formats.md (diff_copy top-2). This skill holds them until Fish approves editing that file.

7. Gates & QA checklist

8. Failure modes

Failure Receipt Fix
A "verbatim" line that exists nowhere 10-04: "you baby him" labelled the husband's verbatim; 0 hits in any corpus (qa-report-v1.md) Verifier on every quote before drafting; 0 hits = drop or re-source
Quote edited inside quote marks 10-04: "can't get him to school" (source: "your kid"), "Mom" for "Mam", reordered "let him be late" Trim only; else drop the quote marks and log as adapted
Two sources merged under one citation 10-04 S1 cited r/coparenting + r/legaladvice as one One row per quote, one source per row
Real quote scores 0 line only in a source outside the corpus (9-23 survey export, a new pull dir not under research/voc-pull-*) voc_check.py --files to see the corpus; grep the raw source; never declare fabricated on voc_check alone
Our extraction verifies itself a quotes .jsonl saved under research/ gets globbed Tables live in output/copy-research/; raw text only under research/
Okendo reviews used as customer voice 94.6 % predate the first order; templated phrases repeat (E-shopify-nonbuyer.md §6a) Exclude, except the one organic 1-star flagged there
Red-team fixes create new contradictions 10-04: 3 of 4 v3 blockers came from v2 fixes; hook contradicted the body (S3 #1) Re-check every fix against hook, timeline and ruled lines; diff vN vs vN+1
Narrator reports what she couldn't see "his wrist buzzed at six-ten" while she's on the highway Who-knows-what pass; "He'd set it for six-ten."
Dialogue becomes alternating monologues 10-01 attendance pack measured: CLERK 6 lines in a row, 18-word turns Turn stats; ≤2 consecutive lines per speaker in a scene; split long turns
Fabricated scarcity davidaistar's "small brand blowing up, grab it fast" Only live, true offer terms; else risk reversal (guarantee never the headline)
Copying the competitor's structure line for line corrections 10-05 (Kalda mirrored soothefy) Dissect as a gap map; must fail the competitor-swap test
Stakes the mornings didn't cause / kid money the-shirt 0 purchases; college songs 0.40/0.58 (MAP §1) Stake bank rows only, raised one step from a real line

9. Sources

~/research/davidaistar/analysis/diff_copy.md (full), analysis/gemini__Y7KFZHxMe0.md (two tables), analysis/gemini_pYNC49BHsHQ.md + txt/pYNC49BHsHQ.txt (free-value skeleton), analysis/gemini_GIczkree_W0.md (podcast script, 80 % rule), analysis/gemini_8NzJJ3e7uZ4.md (song prompt), analysis/gemini_Rf1uHxSWpjU.md (transcript adaptation), analysis/gemini_gpzy2ofG0Dw.md (challenge hook, per-second breakdown); ours: ~/.claude/skills/dawn-native-brief/SKILL.md, ~/.claude/skills/dawn-native-write/SKILL.md, ~/Documents/DawnBands-Native-Writer/WRITER-RULES.md, skills/direct-response-copywriter/references/formats.md, output/dawn-iterate-2026-10-04/copy/lyric-redteam-v3.md + qa-report-v1.md (method only), MAP.md + stake bank. Detail in references/sources.md.

Not built yet (needs Fish's go)

Every line in this skill that mentions something not built, roughly de-duplicated. Some items repeat in different words.