Model registry — roles → models
Fish 10-06: "some technologies/models might be outdated but the framework should still be there, and we can tie it to the newest models… we have fal ai for models… seedance 2.0 might not even be the best model anymore, it's 2.5, and there are other models out there too."
fal is the default gateway (exception, Fish 10-06: vo-broll / b-roll generation = Higgsfield first, fal fallback) (FAL_KEY in workspace .env; catalog: curl -H "Authorization: Key $FAL_KEY" "https://api.fal.ai/v1/models?limit=100&cursor=…"). Higgsfield and kie are secondary (credits, Suno). Refresh the catalog monthly: re-pull, diff new endpoints per role, add candidates below.
| Role | Proven pick (in our code) | Newer candidates on fal (untested — one-unit proof first) | Entry point / status |
|---|---|---|---|
| Stylized video, ref + audio clock (cartoon lanes) | MiniMax H3-max reference-to-video minimax/h3-max/reference-to-video (also Higgsfield minimax_h3_max) |
minimax/h3/reference-to-video/lora (style LoRA), alibaba/wan-3.0-prime/reference-to-video, google/gemini-omni-flash/v1.1/reference-to-video, fal-ai/kandinsky6-pro/image-to-video (10-05) |
scripts/cartoon_h3/run.py. ≤15 s, ≤4 panels; h3-max $0.05/s 480P · $0.08/s 768P · $0.16/s 1080P + reference surcharge ~$0.10-0.16/batch; h3 $0.05 480P · $0.06 768P (fal pages 10-06, h3.py FAL_PRICE fixed; the pricing API only returns the 480P base) |
| Photoreal / multi-ref video with locked character | ugc-omni, proven 10-06: real-creator ref → avatar (GPT Image 2 + Seedream 5 Pro) → locked Omni prompt on kie (~$0.06/s); product via omni's product gate. Seedance 2.5 reference-to-video built then dropped 10-04; its product shots rejected 10-06 | google/gemini-omni-flash/v1.1/reference-to-video (fal route), bytedance/seedance-2.5/us/reference-to-video, alibaba/wan-3.0-prime/reference-to-video, fal-ai/bernini-r/reference-to-video |
scripts/ugc_omni/omni.py; product = gated generated frames (pick --product ≥7/10) or real footage |
| Native-speech video (character says a quoted line, lips included) | Photoreal standard = ugc-omni (Fish 10-06): Gemini Omni 1.1 Flash on kie (gemini-omni-video), exact first/last frame, locked prompt; $0.315/$0.42/$0.525/$0.63 per 4/6/8/10 s. Side arm: H3-max i2v, no audio (minimax/h3-max/image-to-video, yapper, 1080P) |
google/gemini-omni-flash/v1.1/image-to-video (fal route, same family, untested), bytedance/seedance-2.5/image-to-video (dropped), fal-ai/kling-video/v3/pro/image-to-video (dropped), blackforestlabs/flux-3/image-to-video |
scripts/ugc_omni/omni.py (mode: cuts / oneshot); side arm yapper.py chain-render --native |
| Lip-sync onto a still or clip (audio-driven) | H3-max lip-sync minimax/h3-max/lip-sync/image-to-video + ElevenLabs VO — SUPERSEDED as the talking-head standard by native voice (Fish 10-04 23:05); keep for cases where the exact recorded VO must be used. Real fal price (yapper 10-04): $0.16/s at 1080P, $0.08/s at 768P (the $0.05 in code is the 480P rate), x1.2 past 15 s. Also wired, not adopted: Kling AI Avatar v2 Pro, OmniHuman 1.5, sync-lipsync v2, Fabric 1.0 |
fal-ai/sync-lipsync/v3/image-to-video (sync-3) PROVEN 10-06 for CARTOON characters (Fish: "lipsync was fine"; clay Mom + EL line, $0.51/3.8 s, output/model-proofs-2026-10-06/3-sync3-lipsync/) — the route for on-camera cartoon dialogue (wiring into cartoon-h3 NOT BUILT). Others untested: lightricks/ltx-2.5/audio-to-video/pro, fal-ai/flashtalk |
yapper arms. cartoon-h3 has no lip-sync path (core §3) — sync-3 is what to prove first |
| Motion transfer / recast (keep a clip's motion, camera, cuts, audio; new people) | minimax/h3-max/recast — PROVEN 10-06 (Fish: "recast is fine" on the cartoon-clip proof, output/model-proofs-2026-10-06/2-h3-recast/; $0.30/s at 768P). Use on our own/licensed footage only (format-motion-transfer rights gate) |
minimax/h3-max/recast (recast people from reference photos, preserving motion/camera/cuts/audio), fal-ai/kling-video/o3/4k/video-to-video/reference, fal-ai/id-v2v (restyle scene/lighting from edited keyframes, identity + motion preserved), luma/agent/ray/v3.2/video-to-video |
pipeline wrapper command NOT BUILT yet (proof used a direct fal call: output/model-proofs-2026-10-06/fal_proof.py) |
| Video extend (lengthen a take in-model) | none (yapper hand-built segment chain) | minimax/h3-max-turbo/extend-video (+1-15 s — the duration set is the NEW segment's length, not the total; audio refs, up to 2K), blackforestlabs/flux-3/extend-video, Seedance 2.5 US (extension) |
NOT BUILT |
| Video edit by instruction (change one thing in a clip) | none | google/gemini-omni-flash/v1.1/edit, blackforestlabs/flux-3/edit-video |
NOT BUILT. Fits "one variable per pass" |
| Cheap simple-motion i2v (paper pans/zooms, b-roll drafts) | Seedance 1.5 Pro i2v fal-ai/bytedance/seedance/v1.5/pro/image-to-video; Kling v3 Pro i2v |
xai/grok-imagine-video/v1.5/lite/image-to-video, fal-ai/kandinsky6-lite/image-to-video, alibaba/wan-3.0/image-to-video, fal-ai/kling-video/v3/turbo/standard/image-to-video, lightricks/ltx-2.5/image-to-video/pro, pixelcut/looping-video (loops) |
scripts/seedance_img2vid.py. Hailuo-02 holds ~1-3 s on people (Side Effects) |
| Draft → final (cheap motion prototype, then full-quality render of the approved draft) | none wired | Seedance 2.5 Draft (480p, 30 cr/10 s) → "Generate in 1080p" (120 cr) on Higgsfield; fal blackforestlabs/flux-3/{image-to-video,keyframes-to-video,extend-video}/draft → blackforestlabs/flux-3/draft-enhance; H3: --resolution 480P draft → --hd upscale (prove vs native 1080 first) |
core §4 gate 6b. H3: run.py --draft BUILT 10-06; proof: a 480P draft is a different take (SSIM 0.34) — checks prompt/cuts, not the final look |
| Image gen (stills, sheets, canvases, statics) | GPT Image 2.5 Flare/Sunburst openai/gpt-image-2.5/* |
blackforestlabs/flux-3/text-to-image (10-01), bytedance/seedream/v5/flash/text-to-image, microsoft/mai-image-2.5-pro, alibaba/qwen-image-3/text-to-image; Seedream v5 for CGI/Pixar |
scripts/cartoon_h3/canvas.py, sheet.py, scripts/statics/gen_statics.py. Nano Banana Pro: RETIRED everywhere (Fish 10-06 after the omni side-by-side). ugc-omni image arms = GPT Image 2 + Seedream 5 Pro (sd5, kie seedream/5-pro-image-to-image, ~$0.075/img; fal bytedance/seedream/v5/pro/edit $0.0675) — Fish 10-06: "lets do seedream 5.0 pro" |
| Exact text on image (statics headlines, UI) | GPT Image 2.5 edit pass | ideogram/v4.5 + /edit ("accurate text rendering"), recraft/v4.1/flash/text-to-image |
statics improvise-first → text pass. Deterministic HTML/PIL render still beats any model for UI threads |
| Image edit (fix one thing) | GPT Image 2.5 edit; Seedream v5 pro edit | blackforestlabs/flux-3/edit-image, ideogram/v4.5/edit, meta/muse-image/edit |
improvise first, then edit (cartoon-h3 rule 10) |
| Upscale / restore | ByteDance video upscaler; SeedVR image | fal-ai/kandinsky6-vsr/lite, blackforestlabs/flux-video-upscale, Topaz video suite (upscale, interpolate, denoise, deblur) |
scripts/cartoon_h3/upscale.py. Judge by eye, never sharpness score (10-04) |
| TTS / voices | ElevenLabs API direct (per-speaker voices) | elevenlabs/tts/eleven-v4-turbo (on fal, audio tags), google/gemini-3.8-flash-tts, fal-ai/minimax/speech-2.8-hd, fal-ai/inworld-tts |
scripts/cartoon_h3/vo.py. Fish Audio is NOT on fal; untested anywhere |
| Music / sung vocals | Suno V5 via kie.ai (no Suno on fal) | google/lyria-3.5 (instrumental/music), elevenlabs/music/v2.5 (newer than the EL Music that dropped lines 10-04), minimax/music-3 (rejected 10-04: wrong voice) |
scripts/cartoon_h3/song.py. Sung mom vocals stay on Suno until a candidate passes lyric_check on the same lyrics |
| Sound effects / foley | audio-library/sfx files | AI foley REJECTED 10-06 (Fish: "ai foley sucked"; sonilo v1.1 proof) → real SFX files only (licensed library / recorded); text-to-SFX models untested | NOT BUILT (assemble lays one music bed). Needed by format-asmr, format-skit |
| Speech-to-text / word timings | ElevenLabs Scribe | (catalog: 11 speech-to-text endpoints) | song.py, ad_dissector.py |
| Video understanding / QA / dissect | Gemini 3.1 Pro (google-genai direct) | — | scripts/gemini_qa.py, scripts/ad_dissector.py, clipqa.py. ~1 fps sampling; eyes are the final gate |
| Editing / compositing (replaces CapCut) | Remotion apps/orange-doc via scripts/remotion_stitch.py (VideoAd / VideoAd30) + ffmpeg |
— | cartoon_h3/assemble.py, title.py, endcard.py; podcast/duet/split compositions NOT BUILT (core §2a) |
| LLM writing / planning | Claude tiers via scripts/claude_cli.py; GPT-5.x via codex as challenger |
— | per ~/.claude/CLAUDE.md |
Tiering
Roles are filled per SHOT TIER (core §2b): tier A (product reveal, hands, CTA, multi-character, hook) = best proven model; tier C (establishing, b-roll, pans) = cheapest proven simple-motion model. H3-max is the default for cartoon tiers A-B; photoreal A-roll = ugc-omni (Omni), tiers only pick b-roll.
How to adopt a candidate
- One-unit proof: the same shot/sheet/take on the proven pick and the candidate, side by side; Fish's eye decides (never a metric alone). Score takes-to-usable (how many generations until one passes), not only the best take. Read the price on the model page first; ≤ ~$2 per proof without asking, more needs Fish's go.
- Wire it as a new
--backend/--modelchoice in the entry point; keep the old one selectable. - Move it to "Proven pick" here with the date and the run that proved it; append a line to the affected skills'
memory/learnings.jsonl.