Pipeline Hub

workflow-copy-research

Research-to-script workflow for Dawn VIDEO formats (song, spoken drama, podcast, courtroom, whistleblower, street interview, timeline, rated solutions, mascot explainer, skit): an on-demand VOC pull mid-brief into two grep-verified tables (ICP/persona + offering, every row a verbatim quote + source + date), 3 sourced hooks per concept, then the script in one of four modes (narrator beat sheet, two-person dialogue/banter, song lyrics with the versioned red-team method, or the long-form free-value explainer skeleton), ending in logic pass, Opus red-team, fix and re-check.

Lane: copy (feeds every format-* pack) · not built yet: 5 (see below) · source: /Users/ayden/.openclaw/workspace/skills/workflow-copy-research

SKILL.mdprocedure.mdsources.mdEvals (6)Learnings (9)

workflow-copy-research — procedure

Run from ~/.openclaw/workspace. Every step here is free unless marked. File names refer to the run dir output/copy-research/<slug>-<date>/ (SKILL §4).

1. Pull protocol (steps 2-3)

1a. Avatar moment + queries

Write the moment in one sentence (SOP step 1): who, when, what just happened, the private thought. Then 5-8 queries in HER words, never marketer words ("he sleeps through every alarm", "school called about tardies", "my name on the truancy letter"). Log both in pull-log.md.

1b. Minimum quotes (corpus-held quotes count)

Target Rule (from knowledge/copy/voc-sourcing-sop.md) Min quotes Tier-1 (community) min
≤60 s video (skit, short podcast, street interview, short song cell) UGC-script row 15 8
≥2 min video (song, spoken drama, courtroom long cut, whistleblower, timeline, explainer) VSL row 40 20
Iteration that only swaps slot 2 or 3 Revision row 10 private-vocab check
Every set needs at least one "3am thought" (the unfiltered late-night vent). Stop counting at the minimum only if the high-value targets exist: private vocabulary, failed-solution confessions, 3am thoughts, victory signals (SOP step 5).

1c. Where to look, in order

  1. Held corpus (grep -il or the verifier in §3 with candidate phrases): brand tables knowledge/brands/dawnbands/research/offering-analysis.md (standing ICP + offering rows, verified quotes, skill product-market-research; start here, then go per-concept); stake bank output/creative-strategy-cooper-2026-10-05/dawn-stakes.md + stakes-raw-external.md / stakes-raw-owned.md; dawn-voc-bank-final/dawn-voc-bank/voc.jsonl (fields: quote, source URL, date, lane, category); lane files in dawn-voc-bank-final/dawn-voc-bank/lanes/; competitor reviews knowledge/brands/dawnbands/research/voc-competitor-reviews-2026-10-02.jsonl (quote, source_url, star_rating, date, persona_tag); Reddit dumps research/reddit-mornings-2026-09-03/, research/reddit-teen-mornings-2026-09-04/ (+ verified-teen-voc.jsonl), research/voc-blame-2026-10-04/ (accuser lines); survey free text quoted in research/customer-evidence-2026-09-15.md rows S01-S08.
  2. Reddit, fresh (Tier 1). Title search on the Arctic Shift mirror, one request per 3-5 s; the first calls often answer "Timeout. Maybe slow down a bit": wait 10 s and retry. bash curl -s -m 90 "https://arctic-shift.photon-reddit.com/api/posts/search?subreddit=ADHDparenting&title=tardy&limit=40&sort=desc" D=knowledge/brands/dawnbands/research/voc-pull-<topic>-<date>; mkdir -p $D && cd $D python3 ../reddit-mornings-2026-09-03/arc.py <post_id> <post_id> ... # writes thread_<id>.txt here, 4 s apart Each dump starts with URL: and a TITLE: line carrying sub, author, score, comment count and the post date. Comment dates are not kept by arc.py: record "thread posted , pulled "; if a comment's own date matters, read created_utc from /api/comments/search?link_id=<id>&limit=100&sort=asc. Subs that carried Dawn VOC: ADHDparenting, ParentingADHD, Parenting, parentsofteens, Mommit, ADHD, adhdwomen, coparenting, legaladvice, teenagers. Mumsnet threads are in voc.jsonl already.
  3. TikTok comments (Tier 2, brand-adjacent: people under a video about the problem, not the product). Search problem videos, then pull each one's comments; create_time is the comment's own date (epoch). Save raw text only (.txt, never .jsonl, SKILL §4): bash cd ~/tt-creative-sourcer && python3 -c " import sourcer,sys,datetime as d; h,v=sys.argv[1],sys.argv[2] c=sourcer.tikwm_get('/api/comment/list?url=https://www.tiktok.com/@%s/video/%s&count=50&cursor=0'%(h,v)).get('comments',[]) print('URL: https://www.tiktok.com/@%s/video/%s'%(h,v)) [print('%s | likes %s | %s'%(d.date.fromtimestamp(x['create_time']),x.get('digg_count'),x.get('text','').replace(chr(10),' '))) for x in c] " <handle> <video_id> > ~/.openclaw/workspace/knowledge/brands/dawnbands/research/voc-pull-<topic>-<date>/tt_<video_id>.txt Find videos with the §9 search in format-ad-clone (/api/feed/search?keywords=…, rank by comment_count). Emoji-only and tag-a-friend comments don't count toward the minimum. Verified 10-06 (one search + one comment list, free).
  4. Competitor / Amazon reviews (Tier 2): the 10-02 jsonl first; fresh pulls with python3 scripts/voc_amazon_stealth.py --config <cfg> --output <file> --brand dawnbands (scrapling).
  5. Owned (Tier 3): survey lines via customer-evidence S-rows (it is a doc we wrote that quotes the survey: mark the tier "survey, secondary copy"); refund notes and the one organic 1-star in output/dawn-bottleneck-2026-10-03/E-shopify-nonbuyer.md. Not the Okendo corpus (§6a of that file: 94.6 % predates the first order).
  6. Reference ads for structure only (not VOC): python3 scripts/ad_dissector.py <src> --brand-context "silent vibrating wake-up wristband for heavy sleepers; three alarms set on the band; no app" [--competitor]. Output output/ad-dissect/<slug>/dissect.md. Lines in it are paraphrased by design and never count as customer words.

2. The two tables (tables.md)

Persona = ONE core problem, not a demographic (davidaistar _Y7KFZHxMe0). Every row: verbatim quote, source path + URL/ID, date, verifier hit count. A row without all four is deleted.

Table 1 — ICP / persona | # | Persona (one core problem) | What she believes about it (problem-as-belief) | Angle name | Verbatim (trimmed, never paraphrased) | Source (path · URL / thread · comment id · author) | Date | Tier | Emotion | Hits | |---|---|---|---|---|---|---|---|---|---|

Table 2 — offering (feature → benefit → desire → need, each backed by her words) | # | Feature (product-truth.md row) | Benefit (what it does for him/her) | Desire (why she wants that) | Core need / pain / struggle | Objection it answers | Verbatim | Source | Date | Hits | |---|---|---|---|---|---|---|---|---|---| Features come only from knowledge/brands/dawnbands/product-truth.md (vibration that builds and keeps going until turned off; silent to others; three alarms set on the band; no app; dark LED time display; 60-night guarantee). Never a smartwatch, touchscreen or lit face.

Footer lines, always:

spot-check: <n>/10 byte-exact (rows <ids>, opened in the raw file)
beliefs list: <one line per belief, row ids>      objections list: <one line per objection, row ids>
dropped: <rows removed for 0 hits or no date>

3. Extended quote verifier

Built into voc_check.py on 10-06 (this section used to carry a wrapper). It reads the raw corpus (thread dumps, voc.jsonl quote + raw_excerpt, the bank's raw corpora, 10-03 Reddit comment JSON, voc-pull-*/*.txt Reddit + TikTok pulls, buyer survey CSVs, Fish's notes), normalises curly quotes + whitespace on both sides, and prints each hit's file and category. Competitor reviews (Amazon/Trustpilot rows, the bank's amazon rows) and Okendo print as NOT COUNTED. Exit 1 if any phrase has 0 counted hits.

cd ~/.openclaw/workspace; python3 scripts/statics/voc_check.py "<quote 1>" "<quote 2>"   # --files lists the corpus

Tested 10-06: CPS line → 1 (199olsy.json); "I'm tired of getting up early to get him up" → 1 (survey CSV, 09-15 export); "you baby him" → 0 (the 10-04 fabrication). Read the hit in context before trusting it: a phrase can appear in an unrelated thread (10-04 "baby him" hit was about an adult son).

4. Hook grid (hooks.md)

Three per concept, one of each kind unless the format skill says otherwise: - (a) Her words / the accuser's words: a table row verbatim (trimmed), said in the first line. - (b) The stake: a stake-bank "raised" line that lands on her ("They put my name on the letter, not his."). - (c) The test, in action: the moment the challenge starts. Challenge framing ("one week of mornings, his way") is the davidaistar gpzy2ofG0Dw structure: challenge → time jump → proof.

Hook Text (≤2 mobile lines on the overlay) Source row Emotion Curiosity gap Stake (on her, caused by mornings) Promise / Loop Starts in action Zero backstory Lines 1-2 name her by symptom No mechanism Said-out-loud Hits
Rules: pronoun-first in first person (I/My/We); no mystery opener; no product, mechanism or price; the in-media-res cut of the finished script is an optional 4th (Evolve §10). For songs, the hook is its own Suno song in the same style string (format-song §11); for cartoon-h3 narrator/dialogue, hooks share one body via pack body_start.

5. Mode B — two-person dialogue / banter doctrine

Shared by format-podcast, format-spoken-drama, format-courtroom, format-whistleblower (interviewer cues) and format-street-interview. The format skill's beat sheet sets beats and budgets; these rules set how the lines talk.

  1. Script before pictures. Write the whole exchange and pass this section before any sheet, cast or voice work. Most of the realism is the script (davidaistar GIczkree_W0).
  2. Short alternating turns. Most turns 3-8 words; dialogue turns ≤8, MOM testimony ≤12, explainer ≤12 per line and ≤3 lines total for the root cause. One sentence per pack line.
  3. No monologues. In a scene, no speaker runs more than 2 consecutive lines (narration lines in narrator-led formats are exempt, but narration never summarises an exchange the scene could stage). The other speaker cuts in at least every ~6 s.
  4. Interruptions. The cut-off line ends on an unfinished clause with a single line-final "—" (the only dash the humanizer pass keeps; mid-line dashes still go); the next line is the interrupter. Never two voices in one line; never overlapping takes. Overlap is faked in the edit (J-cut 0.15-0.3 s) only in per-speaker-take lanes (podcast B).
  5. Short reactions, lane-dependent. Cartoon-h3 lanes: beats under ~1.2 s get no shot and storyboard.py merges inside a line only, so pair a 1-3 word reaction with its reason ("Every morning? Every single one?"). Per-speaker-take lanes with a cut list (podcast B): a reaction can be its own line.
  6. Every turn does a job: raises the stake, asks the viewer's question or objection, or answers it. No "as you know" lines (characters never tell each other what both already know).
  7. People. The accuser is reasonable and wrong, never a cartoon villain. The explainer is a person inside the story with a reason to be there, says Fish's ruled root-cause line verbatim, plain words, one picture. The host/interviewer asks, never pitches.
  8. Said-out-loud. Contractions, pronoun-first for MOM, numbers spelled as spoken ("six thirty"), fillers only if they appear in VOC and ≤1 per 30 s, no announcer transitions, no composed parallel rhythm (feedback_written_vs_said_copy).
  9. Staging note per line (core §3, no lip-sync in cartoon-h3): every non-narrator line names its listener or off-screen route in the scene intent. Delivery notes go after the quote as *(flat, not looking up)*: pack.py parses and discards them, so they guide the storyboard and Fish, not the voice (EL delivery = pack style_<SPK>).
  10. Table read. Read it aloud in two voices before the red-team. Then run the turn stats:
cd ~/.openclaw/workspace; python3 - output/copy-research/<run>/script-vN.md <<'PY'
import sys
from scripts.cartoon_h3.pack import _LINE_RE
rows = [(m.group(2) or "NARR", m.group(3)) for l in open(sys.argv[1]) if (m := _LINE_RE.match(l))]
run, prev, worst = 0, None, {}
for spk, line in rows:
    run = run + 1 if spk == prev else 1; prev = spk; worst[spk] = max(worst.get(spk, 0), run)
words = {s: [len(l.split()) for k, l in rows if k == s] for s in {k for k, _ in rows}}
total = sum(map(sum, words.values()))
for s, w in sorted(words.items()):
    print(f"{s:10} lines {len(w):3}  words {sum(w):4} ({100*sum(w)/total:.0f}%)  max/turn {max(w):2}  max run {worst[s]}  <4 words: {sum(x < 4 for x in w)}")
PY

Pass: max run ≤2 for dialogue speakers, max/turn within rule 2, <4 words lines = 0 in cartoon-h3 lanes. Measured on the 10-01 attendance-office pack: CLERK max run 6, 18-word turns — the monologue drift this catches.

6. Mode D — long-form free-value explainer skeleton (davidaistar pYNC49BHsHQ, adapted)

Option for explainer formats (timeline, rated solutions, mascot explainer, whistleblower). The character the format provides delivers it: no character-less VO (MAP §3: 0 purchases).

# Beat Share of words (60-90 s ≈ 200-300 words at ≥3.4 wps) Job Dawn rules
1 Hook 5-8 % Stop her on her problem §4 grid; problem-aware: names her by the symptom
2 Problem + villain 12-18 % Name why it keeps happening and who/what to blame Villain = an approach or belief ("louder alarms", "let him be late so he learns", "phone across the room"), never a person she loves and never a competitor brand by name; failed solutions generic, no bed shakers
3 Free value (biggest) 35-50 % Teach something true and useful even if she never buys: what deep sleep does to sound, why the same alarm stops working (habituation), what to try tonight Must be true and saveable. It sets up the plug: by the end, touch is the obvious answer. Fish's root-cause line goes here verbatim
4 Plug + ONE reason-to-believe 15-20 % Natural turn to the band One RTB per ad (her sleep back / his independence / no more yelling / cheaper than another year of late slips); the RTB is the A/B variable. Mechanism line verbatim; facts per product-truth
5 Scarcity / urgency CTA, through the truth gate 5-8 % Why now Only true terms, checked live (curl -sL https://dawnbands.com/products/dawn-wake-band.json). No invented stock or "small brand blowing up". If nothing is scarce, use a true timing reason (the next school term, the next tardy letter) plus the 60-night guarantee as risk reversal, never as the headline
6 Engagement CTA 3-5 % Comments A question that invites her own story ("How many alarms does yours sleep through?"). No "comment X for the link" until a reply mechanism exists (NOT BUILT)

7. Mode C — song lyrics (method; never copy lyrics)

The binding rules live in knowledge/ad-formats/song-ad-playbook.md §1-3 and skills/format-song/references/beat-sheet.md (beats, % runtime, word budgets). This is the run method, promoted from output/dawn-iterate-2026-10-04/copy/ (method only): 1. Versioned drafts, never overwritten. 10-04 ran songs-draft-v1 → songs-v2 (post QA) → songs-v3 (logic audit) → songs-v4 (logic + red-team fixes) → songs-v5 (3 POV fixes). Each file opens with mode, parent + CPA, locked beats, ruled lines, and a table of the ONE variable + slot-2 line (with source) + slot-3 belief per song. 2. QA pass against the qa-gate COPY QA (skills/qa-gate/SKILL.md) and the concept formula: verbatim claims checked (that pass caught the fabricated "you baby him"), structure drift from the locked parent, word count vs the parent. 3. Logic pass (yours) → Opus red-team (§8) → line-edit fixes → re-check the fixes (lyric-redteam-v3.md "previously fixed items" table: each prior fix marked fixed / fixed-but-new-problem). 4. Budgets: long song = 550-630 sung words (format-song beat sheet; exhibit A itself ran 684 words / 271.5 s, so "match the parent" means the format-song budget, not 684); short cell 100-140; lines 5-9 words; Suno sings ~2.0-2.5 wps whatever the prompt. 5. Singability traps: spell titles ("Doctor", "Mister"), "a.m." → "in the morning", no abbreviations (DMV), numbers as words, avoid 2-word lines that flip ("It was." → sung "It wasn't"), listen for the brand name ("Don Band"), quoted dialogue comes out in the mom's voice (no second singer), no blank-line gaps or "beat switch" wording, the overlay text must equal the first sung line word for word, no [Spoken] tags. 6. Story traps caught on 10-04: timeline math (a "month" that holds the events), real channels (a manager warns the employee, not his mom), who-knows-what (she can't see his wrist from the highway), the hook contradicting the body, a mediator "ordering" custody.

8. Opus red-team brief (step 8)

Spawn a subagent (model: opus) on the frozen script-vN.md / lyrics-vN.md with: the brief, tables, the format beat sheet, product-truth, MAP §1-2, Fish's ruled lines, and the previous redteam-v(N-1).md. Ask for flags only, as a table # | line | problem | BLOCKING/FLAG | minimal fix, covering: timeline math; who knows what; real-life plausibility and channels; product facts (wrist rule, dark display, three alarms, no app); quote integrity vs tables.md (rerun §3 on every quoted line); hook vs body consistency; overlay vs spoken/sung words; stake on her and caused by the mornings; blame lands on the accuser or the sound, never on the kid as lazy; mode rules (§5 / §6 / §7); and a "previously fixed items" table re-checking every fix from the last round. Output redteam-vN.md; never edit the script.

9. ILLUSTRATIVE run (shape only; not shipping copy)

Concept: courtroom format, iteration of exhibit A (sa-ontime-10-exhibita-a, $33.7 CPA), ONE variable = the accuser (school attendance office instead of the ex). Mode B. - brief.md: atlas row parent-teen × stakes (keys from atlas_label.py keys), job iterate; slot 1 locked (exhibit A beats in the courtroom skill), slot 4 locked; slot 2 = an accuser line from voc-blame or the stake bank; slot 3 = "the record proves she's the problem". - Corpus first found stake-bank #6 ("You can get put in jail.", r/ADHDparenting 1rfciee) and #7 (CPS, 199olsy); the verifier confirmed both (voc_check alone missed #7). Fresh pull: 2 Arctic Shift title searches ("truancy", "tardy letter") → 6 threads into research/voc-pull-truancy-<date>/. - Table 1 row (shape): 3 | Mom of a teen who sleeps through alarms | "the letter is about me, not him" | Her name on the letter | "<verbatim from thread>" | research/voc-pull-truancy-<date>/thread_<id>.txt · reddit.com/r/... · comment <id> · u/<author> | thread posted <date>, pulled <date> | 1 | fear | 1. - Hooks: (a) the officer's line from a table row; (b) "The attendance letter came addressed to me." (raised from #6); (c) "Then the office gave him one week to prove it was me." - Dialogue excerpt (shape): 1. [OFFICER] "Thirty one tardies this semester, Mrs. Hale." (reading, not looking up) 2. [MOM] "He has four alarms. He sleeps through—" 3. [OFFICER] "Every family says that. Every single one." 4. [MOM] "Then you try waking him. One week." - Turn stats: OFFICER max run 1, max/turn 6; MOM max/turn 9; <4 words 0. Red-team v1 flagged: an attendance office can't "give him a week" (it can set a review date) → fixed in v2; re-check found the hook (c) now contradicted the fix → hook rewritten to the review date in v3.