workflow-copy-research — evidence first, then the script
A format skill says WHAT the ad looks like. This skill produces the words that go into its pack: a research pull you can verify line by line, three hooks that each trace to a real person's words, and a script that has survived a red-team. It folds in davidaistar's research tables and free-value skeleton (analysis/diff_copy.md) and keeps our stricter parts (grep-verified quotes, concept formula, atlas, versioned red-team).
1. Read first
knowledge/ad-formats/_FORMAT-CORE.md(Dawn doctrine §1, lane limits §3, gates §4). Binding.knowledge/ad-formats/_MODEL-REGISTRY.mdfor every model named here.references/procedure.md: table templates, verifier, hook grid, dialogue doctrine, free-value skeleton, lyric method, red-team brief, an ILLUSTRATIVE run.references/sources.md: what we took from which video and why.- Concept + copy doctrine:
knowledge/concept-formula.md(4 slots, iteration mode),knowledge/evolve/copywriting-rules.md(card + 13-question checklist),knowledge/hook-taxonomy.md, memoryfeedback_written_vs_said_copy.md,feedback_direct_but_new_register.md,corrections-hot.md(raise stakes from real VOC; beat the leader, don't mirror it). - Dawn facts and evidence:
knowledge/brands/dawnbands/product-truth.md(hardware only; its price row is stale, check live),knowledge/brands/dawnbands/claims-register.md(all claims unlocked 10-02),output/dawn-combo-map-2026-10-05/MAP.md(recipes A/B, what loses), stake bankoutput/creative-strategy-cooper-2026-10-05/dawn-stakes.md. - The target format skill's beat sheet (e.g.
skills/format-song/references/beat-sheet.md,skills/format-podcast/references/beat-sheet.md,skills/format-courtroom/references/beat-sheet.md). Its word budgets and line rules win over the generic ones here. - VOC sourcing tiers + minimum quote counts:
knowledge/copy/voc-sourcing-sop.md.
2. When to use / when not
- Use before the script step of any
format-*video skill, and any time a brief needs evidence it doesn't have yet: a new accuser/stake, a new persona, a new format with no Dawn parent, or a hook set for an existing concept. - Use the research half alone when Fish wants quotes/tables for an angle (no script).
- Don't use for Dawn natives (
dawn-native-brief→dawn-native-writeown research handoff and prose) or for statics (dawn-statics+voc_check.pydirectly). Don't use it to re-research a concept whose four slots are already filled and verified: go straight to step 6. - Don't start a pull to justify a concept MAP §3 says not to build (college/kid-money stakes, stakes the mornings didn't cause, no-character explainer VO, another dad-POV song before the existing one reads).
3. Inputs
- The concept frame (from Fish or the format skill): atlas row + job (
python3 ~/ad-batch-hub/scripts/atlas_label.py show), parent ad with CPA if iterating, the ONE variable, target format skill, target length. - Concept-formula slots already known vs. missing. Iteration mode locks slot 1 (structure) and slot 4 (buyer); the pull only fills slot 2 (verbatim line) and/or slot 3 (belief).
- For a reference structure (competitor or organic format): the video file or URL. A lane leader's ad is a gap map, never the template (corrections 10-05).
- Nothing paid is needed for the research half. Live offer check before any CTA line:
curl -sL https://dawnbands.com/products/dawn-wake-band.json.
4. Outputs
Run dir output/copy-research/<slug>-<YYYY-MM-DD>/ (new slug per run; never overwrite a version):
| File | What |
|---|---|
brief.md |
Concept frame, 4 slots with sources, awareness, governing desire, dominant emotion, mode (A-D), word budget |
pull-log.md |
Avatar moment, 5-8 search strings, every thread/review ID pulled with URL + date, counts per tier |
tables.md |
Table 1 ICP/persona + Table 2 offering, each row: verbatim + source path/URL + date + verifier hits; spot-check footer |
hooks.md |
3 hooks per concept, graded grid (source row, E/C/S, structure, gates) |
script-v1.md … script-vN.md |
Versioned drafts (lyrics: lyrics-vN.md); v1 frozen as written |
redteam-vN.md |
Opus red-team flags on version N (flags only, no edits) |
script-final.md |
Approved script in the target pack's ## VO Script syntax + script stats |
Raw pulls go to knowledge/brands/dawnbands/research/voc-pull-<topic>-<YYYY-MM-DD>/thread_<id>.txt (raw text only). Never write extracted quotes as .jsonl under research/: voc_check.py reads an explicit raw-source list (10-06; voc-pull-*/*.txt included) and must never "verify" our own extraction.
5. Process
- Step 0 — combo plan (feedback loop,
workflow-combo-loop):python3 scripts/combo_loop/plan.py --format workflow-copy-research --brand dawn→ pick one combo (a paying parent + ONE changed dimension) and paste its block into the pack front matter / job or spec JSON ("combo": {...}; no file →tags.py register --prefix):combo_avatarcombo_anglecombo_povcombo_authoritycombo_stakecombo_stake_oncombo_emotioncombo_root_causecombo_mechanismcombo_payoffcombo_devicecombo_formatcombo_parentcombo_variablecombo_ad_prefix(= the launched Meta ad-name prefix).python3 scripts/combo_loop/tags.py check <pack>must PASS before concept approval; the weekly refresh reads results back by that prefix.
- Frame (free). Fill
brief.md: atlas row + job, parent + CPA, the ONE variable, the four concept slots with what is still missing. Pick the script mode: A narrator beat sheet (story formats), B dialogue/banter (podcast, courtroom, spoken drama scenes, street interview, whistleblower Q&A), C song lyrics (format-song), D free-value explainer (timeline, rated solutions, mascot explainer). Write the Evolve §0 answers (avatar → angle → mechanism → authority → awareness); one missing = don't write yet. Any brief field you filled without a file or source behind it carries an[assumed]tag; drafting (step 6 on) doesn't start until each tag is resolved by a source or by Fish (starpophow-to-build-a-claude-skill-for-ads: a brief-validation gate before writing, because a model will sprint past it). - Corpus first (free). Search what we already hold before pulling: stake bank + its raw banks,
dawn-voc-bank-final/dawn-voc-bank/voc.jsonl(774 quotes, 6 lanes),knowledge/brands/dawnbands/research/voc-competitor-reviews-2026-10-02.jsonl(1,552 competitor reviews), the Reddit thread dumps underresearch/, survey lines inresearch/customer-evidence-2026-09-15.md.scripts/voc_search.pyis NOT usable for Dawn today:--brand dawnbandsreads onlyknowledge/brands/dawnbands/voc-*.md(six ~900-byte compatibility pointers since 9-15) becausevoc-tagged.jsonldoesn't exist, so it ranks pointer text, not quotes. Usegrep -i+ the extended verifier (procedure §3) on the files above. - On-demand pull (free; only for what step 2 lacks). Avatar moment → 5-8 queries in her words (SOP step 2) → Arctic Shift title search → full threads via
knowledge/brands/dawnbands/research/reddit-mornings-2026-09-03/arc.pyrun inside the newvoc-pull-*dir. TikTok comments under problem videos are a free Tier 2 source with dates (tikwm, procedure §1c item 3; starpopai-customer-research-ad-copymines reel and TikTok comments alongside Reddit). Reddit.json, old.reddit and WebFetch are blocked (memoryreference_scrapling_reddit_access.md);skills/reddit-scraper/scripts/reddit_scraper.pyuses the dead JSON route. Volume first, filter second. Minimums by format: procedure §1. - Two tables + verify (free). Build Table 1 (ICP/persona) and Table 2 (offering) per procedure §2. Run the extended verifier (procedure §3) on every quote; 0 hits = the row is dropped or re-sourced, never kept "from memory". Byte-exact spot-check 10 rows against the raw file and write
spot-check: n/10under the tables. Cold-read gate on any quote used as mechanism support (SOP). - Decode the reference structure (one Gemini call, cents, plus an ElevenLabs Scribe transcript on the EL quota unless
--no-scribe) when slot 1 is a competitor/organic format:python3 scripts/ad_dissector.py <video|URL> --brand-context "<one line on the band>" [--competitor]. Map its beats to the format skill's beat sheet;--competitoroutput is a gap map. - Hooks (free). 3 per concept, each traced to a table row or stake-bank row, graded on the grid in procedure §4. Fish picks the lead.
- Script (free). Draft
script-v1.mdin the chosen mode: A = the format skill's beat sheet; B = dialogue doctrine (procedure §5); C = lyric method (procedure §7, which defers to the format-song beat sheet andknowledge/ad-formats/song-ad-playbook.md§2); D = free-value skeleton (procedure §6). Every quoted customer line keeps its table row ID in a comment. One script = ONE Table 1 persona, named by row ID inbrief.md(starpopai-customer-research-ad-copy: narrowing to one persona sharpens the adaptation). - Passes, in order (free; Opus = LLM writing role). Said-out-loud pass → Evolve 13-question checklist →
humanizerskill (line edits, dash scan = 0; the one exception is a line-final cut-off "—" on a mode-B interruption, procedure §5 rule 4, which format-podcast uses too) → your logic pass (timeline math, who-knows-what, real-life channels, product facts, overlay = spoken/sung line, product first named no earlier than the format's reveal beat, word count inside the format budget: starpopcreate-and-clone-competitor-ads-with-ainames an early product intro and a long script as his clone's two flaws) → Opus red-team on the frozen version (subagent brief in procedure §8; writesredteam-vN.md, flags only) → fix as line edits intoscript-v(N+1).md→ re-check that the fixes created no new contradiction (10-04: 3 of 4 red-team blockers came from the previous round's fixes). Stop after 3 red-team rounds; open items go to Fish. - Hand-off (free). Convert to the target pack's syntax (
N. [SPK] "line"for dialogue,N. "line"for songs/narrator), run the turn-stats snippet (procedure §5) for modes A/B, writescript-final.md. Fish approves the concept + script (core §4 gate 1) before any VO, Suno or sheet spend.
6. Tools & models
| Step | Role (_MODEL-REGISTRY.md) |
Entry point | Built? |
|---|---|---|---|
| Atlas row | — | ~/ad-batch-hub/scripts/atlas_label.py show / keys / add-concept |
built |
| Corpus search | Gemini 3.1 Pro as a text ranker (same model as the video-understanding row; ID hard-coded in the script) | scripts/voc_search.py |
built, but BROKEN for Dawn: reads only the voc-*.md pointer stubs (no voc-tagged.jsonl); fix = point it at voc.jsonl + the research dumps (code, Fish's go). Use grep + procedure §3 |
| Reddit pull | — (plain curl, Arctic Shift mirror) | knowledge/brands/dawnbands/research/reddit-mornings-2026-09-03/arc.py <ids> |
built (manual search step) |
| TikTok comments | — (tikwm, free) | ~/tt-creative-sourcer sourcer.tikwm_get('/api/comment/list?url=…') (procedure §1c item 3) |
built (tested 10-06: text, digg_count, create_time) |
| Amazon/competitor reviews | — | existing voc-competitor-reviews-2026-10-02.jsonl; fresh pulls scripts/voc_amazon_stealth.py --config --output --brand |
built |
| Quote verification | — | scripts/statics/voc_check.py (procedure §3) |
built; corpus fixed 10-06 |
| Reference decode | Video understanding / QA (Gemini 3.1 Pro) | scripts/ad_dissector.py |
built 10-06 |
| Drafting, red-team | LLM writing / planning: Opus for red-team, the session model for drafts; GPT-5.x via codex only as a challenger when Fish asks | Agent subagent model: opus; pipelines use scripts/claude_cli.py run_opus |
built |
| Humanizer | LLM writing | ~/.claude/skills/humanizer/SKILL.md |
built |
| Singability / pace (downstream) | TTS (ElevenLabs) / Music (Suno V5 on kie) | format skill's vo_check.py / song.py receipt |
built, owned by the format skills |
NOT BUILT (needs Fish's go):
- voc_check.py corpus (BUILT 10-06): voc.jsonl (774), the bank's raw corpora, output/dawn-bottleneck-2026-10-03/reddit/comments/*.json, voc-pull-*/*.txt, buyer survey CSVs; curly quotes + whitespace normalised; each hit prints its file and category; competitor reviews and okendo_reviews_anon.json (94.6 % predate the first order) print as NOT COUNTED. CPS line #7 now scores 1 (199olsy.json). Still outside: the 9-23 212-row survey export (not on disk) and our own docs (customer-evidence-2026-09-15.md, on purpose).
- A one-command research_pull.py (queries → Arctic Shift search → thread dumps → table skeleton). Today steps 3-4 are manual.
- A dialogue lint inside check_plan.py; the turn-stats snippet is the manual stand-in.
- Comment-for-link responder (davidaistar's engagement CTA). Without it, the engagement CTA is a question only.
- The dialogue and free-value sub-modes in skills/direct-response-copywriter/references/formats.md (diff_copy top-2). This skill holds them until Fish approves editing that file.
7. Gates & QA checklist
- [ ]
brief.md: atlas row + job; parent + CPA (if iterating); ONE variable; four slots each with a named source; mode A-D; word budget from the format skill. - [ ] Problem-aware gate: lines 1-2 name HER by an observable symptom; stake lands on her and the mornings caused it (core §1); no MAP §3 don't-build item.
- [ ] Pull: ≥ the SOP minimum for the format (procedure §1); every thread/review ID logged with URL + date.
- [ ] Tables: every row has a verbatim quote + source path/URL + date + verifier hits ≥1; spot-check ≥9/10 byte-exact; feature column only from product-truth.md.
- [ ] Quote integrity: trimmed, never paraphrased inside quote marks (no pronoun swaps, no "Mam" → "Mom", no reordering). An adapted line loses its quote marks and is logged as "adapted from row N".
- [ ] Hooks: 3 per concept, each traced to a row; Emotion + Curiosity gap + High stakes named; Promise or Loop named; starts in action; zero backstory; no mechanism in the hook; overlay text = the spoken/sung words exactly.
- [ ] Script: Fish's ruled root-cause and mechanism lines verbatim where the format has those beats; explainer is a person inside the story (recipe A) or guarantee/science (recipe B); product facts per product-truth; offer checked live; scarcity only if true.
- [ ] Mode B: turn stats pass (procedure §5); mode C: lyric budget + singability list (procedure §7); mode D: free value is the biggest section and true without buying.
- [ ] Said-out-loud, Evolve 13 questions, humanizer, logic pass, Opus red-team, fixes re-checked. Versions v1…vN kept.
- [ ] Fish approved the script before any paid step downstream.
8. Failure modes
| Failure | Receipt | Fix |
|---|---|---|
| A "verbatim" line that exists nowhere | 10-04: "you baby him" labelled the husband's verbatim; 0 hits in any corpus (qa-report-v1.md) |
Verifier on every quote before drafting; 0 hits = drop or re-source |
| Quote edited inside quote marks | 10-04: "can't get him to school" (source: "your kid"), "Mom" for "Mam", reordered "let him be late" | Trim only; else drop the quote marks and log as adapted |
| Two sources merged under one citation | 10-04 S1 cited r/coparenting + r/legaladvice as one | One row per quote, one source per row |
| Real quote scores 0 | line only in a source outside the corpus (9-23 survey export, a new pull dir not under research/voc-pull-*) |
voc_check.py --files to see the corpus; grep the raw source; never declare fabricated on voc_check alone |
| Our extraction verifies itself | a quotes .jsonl saved under research/ gets globbed |
Tables live in output/copy-research/; raw text only under research/ |
| Okendo reviews used as customer voice | 94.6 % predate the first order; templated phrases repeat (E-shopify-nonbuyer.md §6a) |
Exclude, except the one organic 1-star flagged there |
| Red-team fixes create new contradictions | 10-04: 3 of 4 v3 blockers came from v2 fixes; hook contradicted the body (S3 #1) | Re-check every fix against hook, timeline and ruled lines; diff vN vs vN+1 |
| Narrator reports what she couldn't see | "his wrist buzzed at six-ten" while she's on the highway | Who-knows-what pass; "He'd set it for six-ten." |
| Dialogue becomes alternating monologues | 10-01 attendance pack measured: CLERK 6 lines in a row, 18-word turns | Turn stats; ≤2 consecutive lines per speaker in a scene; split long turns |
| Fabricated scarcity | davidaistar's "small brand blowing up, grab it fast" | Only live, true offer terms; else risk reversal (guarantee never the headline) |
| Copying the competitor's structure line for line | corrections 10-05 (Kalda mirrored soothefy) | Dissect as a gap map; must fail the competitor-swap test |
| Stakes the mornings didn't cause / kid money | the-shirt 0 purchases; college songs 0.40/0.58 (MAP §1) | Stake bank rows only, raised one step from a real line |
9. Sources
~/research/davidaistar/analysis/diff_copy.md (full), analysis/gemini__Y7KFZHxMe0.md (two tables), analysis/gemini_pYNC49BHsHQ.md + txt/pYNC49BHsHQ.txt (free-value skeleton), analysis/gemini_GIczkree_W0.md (podcast script, 80 % rule), analysis/gemini_8NzJJ3e7uZ4.md (song prompt), analysis/gemini_Rf1uHxSWpjU.md (transcript adaptation), analysis/gemini_gpzy2ofG0Dw.md (challenge hook, per-second breakdown); ours: ~/.claude/skills/dawn-native-brief/SKILL.md, ~/.claude/skills/dawn-native-write/SKILL.md, ~/Documents/DawnBands-Native-Writer/WRITER-RULES.md, skills/direct-response-copywriter/references/formats.md, output/dawn-iterate-2026-10-04/copy/lyric-redteam-v3.md + qa-report-v1.md (method only), MAP.md + stake bank. Detail in references/sources.md.
Not built yet (needs Fish's go)
- NOT BUILT (needs Fish's go)
- On-demand research pull mid-brief P1 SKILL §5 steps 2-4, procedure §1-3 (manual; research_pull.py NOT BUILT)
- Free-value skeleton P1 procedure §6 (formats.md sub-mode NOT BUILT, Fish's go)
- Podcast/dialogue doctrine P1 procedure §5 (shared by podcast, spoken drama, courtroom; formats.md sub-mode NOT BUILT)
- 6 Engagement CTA 3-5 % Comments A question that invites her own story ("How many alarms does yours sleep through?"). No "comment X for the link" until a reply mechanism exists (NOT BUILT)