format-motion-transfer — their motion, our people, our set, our sound
A transform, never a re-upload. The only thing taken from a source is how bodies and camera MOVE. Faces, set, outfit, text, audio and footage never reach our ad, unless the footage is ours.
1. Read first
knowledge/ad-formats/_FORMAT-CORE.md(doctrine, lane limits, gates, launch) andknowledge/ad-formats/_MODEL-REGISTRY.mdrow "Motion transfer / recast". Binding.skills/format-ad-clone/SKILL.md§5 (source-type keep/change table). This skill reuses that table and only narrows it (§5 below). Never contradicts it.references/beat-sheet.md(both modes, timing, provenance sheet, ILLUSTRATIVE structure) andreferences/sources.md(davidaistar videos, our charswap receipts, verified model schemas).knowledge/brands/dawnbands/product-truth.md(facts),output/dawn-combo-map-2026-10-05/MAP.md(what pays: pick the parent body here).- Photoreal output:
skills/ugc-omni/SKILL.md+scripts/ugc_omni/omni.py, THE photoreal engine (Fish 10-06; another session owns them: read, never edit); job shape for a talking-head refresh:references/omni-job.example.json. Memoryfeedback_ai_avatar_broll_lane.md(the 9-30 rules ugc-omni's gates implement). Paid video:feedback_higgsfield_credit_discipline.md,feedback_motion_prompt_approval.md.
2. When to use / when not
- Use (mode T, trend hook): an organic morning/teen gag or move is everywhere (door-slam, blanket-yank, "walk in, flick lights, walk out", a dance trend) and we want it as a 3-8 s hook performed by our cast in our set, spliced onto a paying Dawn body. One variable vs the parent: the hook.
- Use (mode R, refresh): our OWN winning footage has fatigued (frequency up, CPA drifting) and the performance itself is the asset. Re-perform it with a new cast, same motion/cuts/audio. One variable: cast.
- Use (mode O, own performance): Fish or the team acts a gag on a phone; we transfer it onto our cartoon cast. Cleanest rights; also the proof-test source.
- Evidence status: untested on Dawn; route NOT BUILT. Precedent is weak: the charswap lane (Seedance 2.5 video-edit,
scripts/charswap_overnight.py) put 25 ads on Meta, 24 edited ONE source, 4 of 28 got delivery (scripts/charswap_source_refs.pydocstring); charswap_g joined the DWB scale test 10-05, unread. davidaistar's motion-swap results (~$10k GMV source clip, Rqim7SlMtZg) are claims. - Don't use for: a lane-leader competitor's ad or any brand's ad footage as the motion source (route to format-ad-clone: gap map / structure remix); a single creator's signature bit; anything where a real person stays recognisable; our cartoon winners when only the LOOK should change (that's
scripts/cartoon_h3/reskin.py, proven, cheaper); lip-sync trends (our character would mouth someone else's words).
3. The format
- 9:16. Mode T: 3-8 s transferred hook + the parent's body from its
body_start(total = parent length ± hook delta). Mode R: the parent's length, beat for beat. Mode O: 3-10 s hook or a 15-30 s short. - Look: cartoon by default (our sheets: pixar / disney3d / claymation; one style per ad, anchored on a paying frame). Photoreal routing (Fish 10-06: ugc-omni is the engine): (a) a photoreal target image for recast/motion-control comes from ugc-omni G1 (real creator frame →
refscan/ref→avatar --arms gpt2,sd5→pick), never a text-to-image face; (b) any spoken photoreal line (a talking hook, or mode R on a talking-head UGC winner) is NOT a transfer: it is a ugc-omni job that re-performs the winner's transcript with the new avatar (recast would keep the old creator's voice on a new face); (c) gates = ugc-omni G0-G4 with the transfer proof (§10 G-proof) as G2. "Dramatization" label on photoreal output (Fish 10-06). - What the viewer sees: a move they half-recognise, performed by a mom/teen they've never seen, in a kitchen/teen room that is clearly ours; overlay line 1 names HER problem; then the paying body.
- Reference shape (paraphrased): davidaistar's guide (Rf1uHxSWpjU) shows a TikTok Shop haul creator's hold-the-camera performance re-performed by his AI avatar in a new room ("change the character AND the background; only the motion is taken"). The Seedance short (5Aj7S-Z4ct8) swaps the runner in a famous chase shot for a new character and the model adapts the run to the new body.
4. Why it sells
- The hook borrows recognition: a move the feed already rewards stops the thumb before the brain files it as an ad. The body that follows is already proven, so the read isolates the hook.
- Dawn mapping: the gag must BE the morning problem (her fifth trip to the door, the teen who doesn't move), so line 1 is problem-aware and the stake lands on her. Then recipe A (story body: accused → explainer → vindication) or recipe B (callout body: her burden → relief) carries the sale. A trend with no morning in it is not a Dawn ad.
- Refresh: a fatigued winner's choreography and timing still work; the audience is tired of the FACE. New cast resets frequency fatigue without re-solving retention.
5. Beat sheet (full sheet + provenance sheet in references/beat-sheet.md)
Source rules first (narrows format-ad-clone §5; a source failing this never reaches a model):
Source type (ad-clone §5 names; licensed is added here because ad-clone covers reference ads, not reusable footage) |
Allowed here? | Route class | Proof of rights (written in provenance.md) |
|---|---|---|---|
own-winner / own footage |
Yes, all routes | recast, restyle, motion-control | Our ad ID + CPA, or the file's origin in our output dir / Fish's camera roll; any person in it consented. Creator footage (Jessica/Kris/Lori, ugc-engine creators): Fish confirms the usage terms cover AI modification (the SOP grants perpetual paid usage, AI edits not stated). NOT charswap outputs (derived from third-party/competitor footage) |
licensed / stock |
Yes, if the licence allows AI modification + paid ads | recast, motion-control | Licence URL/PDF saved beside the source |
organic-viral trend (not an ad) |
MOTION only | motion-control only (generate-from-image) | Trend proof: ≥5 unrelated accounts doing the same move in the last 60 days (handles + play counts); no credited choreographer or creator-claimed routine; no brand in it |
non-competitor-scaler ad |
No (its footage is someone's ad) | — | Use format-ad-clone (structure remix) |
lane-leader-competitor |
Never | — | Gap map only (ad-clone) |
- Route classes. Recast/restyle/edit (output is the source footage with people or look replaced; set, cuts and often audio survive) = own or licensed footage ONLY. Motion-control / omni motion-ref (output rendered from OUR image; the source only drives motion) = any allowed source.
| Beat (mode T) | % | Job | Must contain |
|---|---|---|---|
| 1 Motion hook | 0-20 | Stop the scroll with the borrowed move | Our cast + our set; overlay names HER symptom; no product |
| 2 Turn | 20-30 | The gag lands on her | Her stake in action (the 5th trip, the late slip); first VO line |
| 3-6 Parent body | 30-100 | Sell | The paying body unchanged from body_start (root cause, mechanism, product, payoff, CTA, end card (optional, --end-card; off by default per Fish 10-06)) |
Mode R: the parent's beats 1:1; only the cast changes (then voice: one per pass, ad-clone §11). Segments where the band is in a hand or on a wrist are NOT recast: recast regenerates hands and has no product gate (ungated band-in-hand scored 3-4/10, 10-04). They stay the original real footage, or are rebuilt as a ugc-omni b-roll whose frame passes pick --name broll_<id> --product (≥7/10 vs the real pack, display-off wording, worn or resting across a palm).
Shot tiers (core §2b): the transfer hook = tier A (the hook's first 3 s); mode R segments with people acting = B; empty-room or insert segments = C (keep the original footage, no generation). Note the tier per segment in ## Segments.
6. Pack spec
- Step 0 — combo plan (feedback loop,
workflow-combo-loop):python3 scripts/combo_loop/plan.py --format format-motion-transfer --brand dawn→ pick one combo (a paying parent + ONE changed dimension) and paste its block into the pack front matter / job or spec JSON ("combo": {...}; no file →tags.py register --prefix):combo_avatarcombo_anglecombo_povcombo_authoritycombo_stakecombo_stake_oncombo_emotioncombo_root_causecombo_mechanismcombo_payoffcombo_devicecombo_formatcombo_parentcombo_variablecombo_ad_prefix(= the launched Meta ad-name prefix).python3 scripts/combo_loop/tags.py check <pack>must PASS before concept approval; the weekly refresh reads results back by that prefix.
No loader exists for motion clips. Each run writes output/motion-transfer/<slug>/plan.md; the spliced body keeps its own cartoon-h3 pack.
- Front matter: core §5 keys plus format_skill: format-motion-transfer, mode: T|R|O, parent: (paying ad + CPA), iteration: (the one variable), source_type: own-winner|licensed|organic-viral, source_receipt: (ad ID / licence path / trend list), motion_source: output/motion-transfer/<slug>/motion.mp4, route_role: motion-transfer/recast, route_model: (the registry candidate under test), keep_source_audio: false (true only for own footage in mode R), hook_secs, body_from: (parent pack + body_start).
- ## Provenance: the rights block from references/beat-sheet.md, filled. ## Segments: N. src <in>-<out>s → target <image> | orientation video|image | <who moves, what move, camera>.
- Cast: cartoon = our sheets only (text-to-image via python3 -m scripts.cartoon_h3.sheet); photoreal = a ugc-omni avatar from a real creator frame we may use (own/licensed footage or a real creator reference with the genetic change). Never i2i from a third-party source performer's frame. Envs: our plates (kitchen, teen room, hallway, car).
7. Visual grammar
- Target image = the whole look. Motion-control renders characters AND background from our image (Kling motion-control schema: "characters, backgrounds… based on this reference image"). Match the source's frame-1 pose and framing in the target (full or upper body, unoccluded, character >5% of frame) or the first second warps.
- Orientation:
videofor complex body motion (≤30 s source),imagewhen the camera move matters (≤10 s source). Pick per segment; one segment per generation. - Say what to borrow (the optional prompt). Name the ONE thing taken from the clip and what the target fixes, e.g. "Transfer only the reference's two-handed blanket yank and upper-body timing. Face, hair, outfit and room come from the image." Unstated, the model also borrows the clip's framing, styling and pace (starpop
how-to-use-seedance-2-0-to-make-ads,how-to-use-kling-3-0-to-make-adsstep 10). The segment is one uninterrupted take: a montage or multi-cut clip is not a motion reference. - No lip-sync (core §3): pick sources with no mouthed words, or with the face turned/away. A transferred mouth moving to the source's words over our VO is a fail.
- Wrist rule: the band appears only in the parent body (cartoon-h3 rules). In the hook, wrists low or out of frame; a source where the hand rises to the face/camera with the band arm = lit-smartwatch risk: drop it.
- Carry-over check: no source face, hairstyle+outfit combination, set dressing, logo, caption or burned-in text survives. One person in the source → one in ours; multi-person sources double the identity risk (core §3 contamination).
- Similar-age teen + mom: make them distinct (age, hair, colour).
8. Audio
- Never the source's audio for organic/licensed-without-audio-rights sources:
keep_original_sound: false(Kling motion-control defaults to TRUE) and strip audio again in post (ffmpeg -an). Recast preserves audio by design: own footage only. - Hook audio = ours: one or two VO lines in the parent narrator's EL voice (a hook-only pack through
run.py --stage vo, vo_check floor 3.4 wps), or the parent's music bed under overlay text. The parent body keeps its audio. - Mode R: keep our own audio (it's the proven part); a voice swap is a separate pass.
- End card (optional,
--end-card; off by default per Fish 10-06) from the parent (packend_card:). Never clone a source creator's voice.
9. Production (free until the proof; the route is NOT BUILT)
cd ~/.openclaw/workspace; S=<slug>; D=output/motion-transfer/$S; mkdir -p $D
# 1 SOURCE (free). Own: copy the file. Trend: tikwm fetch (format-ad-clone §9 command), then the trend proof:
cd ~/tt-creative-sourcer && python3 -c "import sourcer,json; d=sourcer.tikwm_get('/api/feed/search?keywords=<move+keywords>&count=20'); print(json.dumps([[v.get('play_count'),v.get('author',{}).get('unique_id'),v.get('title','')[:60]] for v in d.get('videos',[])],indent=0))"
# 2 TRIM + MUTE + CONTACT SHEET (free): one segment, 5-10 s, one shot
cd ~/.openclaw/workspace && ffmpeg -ss <in> -to <out> -i <src.mp4> -an -c:v libx264 -crf 18 $D/motion.mp4
ffmpeg -i $D/motion.mp4 -vf "fps=2,scale=270:-1,tile=6x4" -frames:v 1 $D/motion-sheet.png
python3 scripts/ad_dissector.py $D/motion.mp4 --slug $S-src --no-scribe # beat timings (1 Gemini call)
# 3 TARGET IMAGE (paid ~$0.15-0.30, Fish's go): our cast in our set, posed like frame 1
python3 -m scripts.cartoon_h3.sheet $D/target.png $D/target-prompt.txt <cast-sheet.png> <env-plate.png>
# 4 READ SCHEMA + PRICE (free) for the registry candidate, every time
curl -s "https://fal.ai/api/openapi/queue/openapi.json?endpoint_id=fal-ai/kling-video/v3/standard/motion-control" | python3 -m json.tool | grep -A3 '"description"'
# 5 HOOK VO (EL quota): hook-only pack with the parent's voice → vo + vo_check
python3 -m scripts.cartoon_h3.run --pack output/concept-packs/$S-hook.md --stage vo
# 6 GENERATE: only after the proof (§10 P1/P2) passed and the route is wired. Proof call shape in references/beat-sheet.md.
# 7 MUX + SPLICE (free): transfer clip + hook VO (padded to clip length), then + parent body from body_start
ffmpeg -i $D/transfer.mp4 -i output/cartoon-h3/<hook-pack-slug>/vo.mp3 -map 0:v -map 1:a -af apad -shortest -c:v copy $D/hook-final.mp4
ffmpeg -i $D/hook-final.mp4 -ss <body_start_s> -i <parent-final.mp4> -filter_complex "[0:v]scale=1080:1920,fps=30,setsar=1[a];[1:v]scale=1080:1920,fps=30,setsar=1[b];[a][0:a][b][1:a]concat=n=2:v=1:a=1[v][au]" -map "[v]" -map "[au]" $D/final.mp4
python3 scripts/gemini_qa.py --asset $D/final.mp4 --type final --style cartoon # then Fish watches
- Candidates (registry role "Motion transfer / recast", all UNTESTED on our account, schemas read 10-06): recast
minimax/h3-max/recast(inputsvideo_url5-30 s, no shot >15 s;reference_image_urlsone photo per person;resolution768P/1080P; page 10-06: $0.30/s 768P, $0.45/s 1080P) = modes R/O. Motion-controlfal-ai/kling-video/v3/standard/motion-control(+/pro/;image_url,video_url,character_orientation,keep_original_sound) = modes T/O; price not captured, read the page. Restylefal-ai/id-v2v,luma/agent/ray/v3.2/video-to-video,fal-ai/kling-video/o3/4k/video-to-video/reference= own footage only. Higgsfield CLI (paid since 9-09):hf_mult_motion_control(≥1--image, exactly 1--video),kling_video_edit,seedance_2_0 --video-references(omni motion-ref, 5Aj7S-Z4ct8),seedance_2_5 --mode omni_reference|video_edit(scripts/higgsfield_unlim.py). Catalog alternate not yet in the registry:fal-ai/bytedance/dreamactor/v2("non-human and multiple characters"). - Draft → final (core gate 6b): neither candidate has a motion-preserving re-render. Recast 768P ($0.30/s) and Kling motion-control standard are the draft/proof tier; a 1080P or pro render is a NEW generation, so either re-render and re-check against the draft, or upscale the approved 768P clip (registry upscale role) after a side-by-side. Not wired (the route itself is NOT BUILT).
- Photoreal talking-head refresh (mode R on our own UGC winner) = ugc-omni, built today: transcribe the winner (
python3 scripts/ugc_omni/omni.py mine <winner.mp4>writeshooks.md+<stem>.words.json, the full word transcript, underoutput/ugc_omni/_mine/; or use the winner's own script), writelines[]1:1 with its cuts (AI UGCwhere the creator talked,broll:<id>where the original b-roll ran: copy those own clips to<job>/broll/<id>.mp4), then G0board→ G1refscan --src <winner.mp4>(own footage = allowed ref) →ref→avatar --n 3 --arms gpt2,sd5(genetic change: never the original creator) →pick→ G2lock --line <hook> --n 3→--approve→ G3gen --lines all,verify→ G4assemble --hook H1,cost. Omni buckets $0.315 / $0.42 / $0.525 / $0.63 per 4 / 6 / 8 / 10 s, images ~$0.05-0.075: a 30 s refresh ≈ $6-9. One variable vs the original = the cast (voice changes with it: record that). - Cost per finished mode-T ad once wired: target image ~$0.20-0.30 + one 5-8 s transfer (recast 768P = $1.50-2.40; motion-control per its page) + EL quota; the body already exists. Time ~1 h after the source is approved.
10. Gates & QA checklist
- [ ] G0 Provenance (Fish): source type per §5, receipt saved, trend proof or licence written; competitor/brand ad → stop. Fish says yes before any paid step.
- [ ] G-proof (once per route, before any lane depends on it). P1 recast (mode R/O): Fish's own 5 s phone clip → one cartoon cast sheet, 768P (~$1.50). P2 motion-control (mode T/O): the same clip → our target image,
character_orientation: video,keep_original_sound: false, standard tier, cap ~$2 after reading the price page. Pass (eyes, Fish decides): beats land within ±0.3 s of the source on the 2 fps sheets; character matches its sheet ≥8/10 on every 1 s frame; no morph over the full length; no carry-over; hands clean. Pass → wire + registry row + learnings. Fail → log it; the mode stays NOT BUILT. - [ ] Parent named with CPA; ONE variable in
iteration:; overlay line 1 names her symptom; stake on her, caused by the mornings. - [ ] Target image approved (cast checked against the avatar; photoreal = ugc-omni G1 avatar, never the source performer) before any transfer; balances checked (fal, Higgsfield, kie).
- [ ] Photoreal spoken lines / talking-head refresh: ugc-omni G0-G4, every gate Fish's (G0
board→ G1 ref + avatar → G2 same line good 3× → G3 every line + b-roll + product check → G4 the cut). - [ ] Source audio muted unless own footage in mode R; no mouthed source words on screen.
- [ ] Carry-over audit: source and output 1 fps sheets side by side; nothing recognisable from the source but the move. "Would the source creator recognise her video?" = fail.
- [ ] Wrist rule: no band in the hook, or wrists low and display dark; product facts from product-truth.md. A generated band only through ugc-omni's product gate (
pick --product≥7/10); recast never touches band segments (they stay real footage). - [ ] Final: freezes 0, dark 0, splice seam clean (no audio pop), captions from the script, end card (optional,
--end-card; off by default per Fish 10-06), gemini_qa, Fish watched it. AI disclosure per platform rules.
11. Variants & test plan
- Mode T: one paying body (MAP.md), 3 hooks = 3 different trend motions (or 2 trends + the parent's original hook as control). Own ad set in Testing CBO vs the parent. Read CPA at 7 days, hook rate as a secondary signal.
- Mode R: new cast vs the fatigued original, same audio and cuts. Read CPA and frequency at 7 days. Then voice, then product, one pass each.
- ≤2 ads per source clip per ad set: near-duplicates starve (charswap: 24 ads from one source, 4 of 28 delivered).
- Never judge <7 days; reduce spend, never pause (core §6). Record the launch in a
project_launch_<date>_<slug>.mdmemory.
12. Failure modes
- Competitor rip as the motion source leaks their product: charswap's
@video_1was a competitor rip and 5 of 7 visual-QA fails were the band thickening into a tracker cuff with a module (scripts/charswap_overnight.pyheader). Fix: §5 table; competitor footage never enters a model. - Recast on third-party footage keeps their set, cuts and audio = a re-upload with new faces. Recast is own/licensed only.
- Kling motion-control keeps the source sound by default. Always
keep_original_sound: false, and-anin post. - First-second warp: target pose ≠ source frame 1. Re-pose the target image; don't regen blind.
- Mouth moves to words we don't say: pick a no-speech source or turn the face away.
- Morph after 1-3 s on people (Hailuo, Side Effects): the proof must cover the whole clip length, not a thumbnail.
- Likeness drift: the output starts to look like the source performer. Fail; strengthen our character refs (motion-control
elementsbinding needs orientationvideo); never use a real person's photo as a target. - Provider refuses a face: do not blur, crop or translate around a likeness/safety filter (diff_ai_ugc "don't adopt"). Switch route or go cartoon.
- Two variables at once (new hook motion + new style): split into two plans.
NOT BUILT (needs Fish's go)
- The route itself: no entry point for the "Motion transfer / recast" role. To wire after a passed proof: a motion_transfer.py under scripts/ with
source(provenance.md + trim + mute + sheet),cost(schema + price read),gen --spend(one segment, one candidate,keep_original_soundforced false for non-own sources),check(source/output side-by-side sheets);--backend fal|higgsfield-cli; the registry row gets an entry point. kling_motioninscripts/scene_parser.py/scripts/test_pipeline_fixes.pyis only a prompt-text hint for i2v; it transfers nothing.- No hook-clip slot in cartoon_h3 assemble (captions/overlay over an external hook): splice by ffmpeg (§9 step 7) and burn the overlay by hand until it exists.
- No automated trend proof or carry-over/likeness check: both are manual sheets + Fish's eye.
13. Sources
davidaistar: analysis/gemini_Rf1uHxSWpjU.md + txt/Rf1uHxSWpjU.txt (Kling Motion segment), analysis/ugc_seedance.md (Rqim7SlMtZg motion-swap pipeline, OvAyKctgNE4 omni-ref modes), analysis/gemini_Rqim7SlMtZg.md, analysis/guide_copy_images_shorts.md (5Aj7S-Z4ct8, BHmPSaROT7c, 0jcFpsnuJh0, I0QSg8vOByc; lIu6HLlKDvE has no transcript), analysis/diff_cloning.md, analysis/diff_ai_ugc.md. Ours: skills/format-ad-clone/SKILL.md, scripts/charswap_overnight.py, scripts/charswap_source_refs.py, scripts/charswap_batch2.py, scripts/higgsfield_unlim.py, scripts/seedance_broll_runner.py, output/model-catalog/fal-2026-10-06.json, fal OpenAPI schemas (read 10-06); memory reference_higgsfield_cli_gotchas.md, feedback_higgsfield_credit_discipline.md, corrections-hot.md, feedback_ai_avatar_broll_lane.md. Details in references/sources.md.
Not built yet (needs Fish's go)
- description: "Motion transfer / actor swap: take only the MOTION of a source clip (choreography, camera move, comedic timing) and re-perform it with OUR cast, OUR set and OUR audio, either as a trend-gag hook spliced onto a paying Dawn body
- Evidence status: untested on Dawn; route NOT BUILT. Precedent is weak: the charswap lane (Seedance 2.5 video-edit, scripts/charswap_overnight.py) put 25 ads on Meta, 24 edited ONE source, 4 of 28 got delivery (scripts/charswap_source_refs.p
- 9. Production (free until the proof; the route is NOT BUILT)
- Draft → final (core gate 6b): neither candidate has a motion-preserving re-render. Recast 768P ($0.30/s) and Kling motion-control standard are the draft/proof tier; a 1080P or pro render is a NEW generation, so either re-render and re-check
- [ ] G-proof (once per route, before any lane depends on it). P1 recast (mode R/O): Fish's own 5 s phone clip → one cartoon cast sheet, 768P (~$1.50). P2 motion-control (mode T/O): the same clip → our target image, character_orientation: vid
- NOT BUILT (needs Fish's go)