Tomato argues with Cucumber while Grandma Onion chimes in — all in one scene where nobody loses their face or steals someone else's line. Plus an off-screen voice that carries the story without being 'stitched' to any mouth on screen.
Dialogue, not monologue
A project references up to 3 visual characters (roles soul_a / soul_b / soul_c) plus an optional narrator. Each scene stores its lines by speaker, and the voiceover is assembled through a dialogue API — every character with their own voice. The same Elements keep them consistent.

Up to three on-screen characters each speak in their own voice; the narrator is a separate off-screen track, not tied to any mouth on screen.
The narrator is off-screen, not a 'talking potato'
The narrator is a voice with no visual presence. It carries the story but is never lip-synced by a character on screen: the system strictly tells apart who in the scene is speaking with their lips and who is heard off-screen. No backgrounds that 'suddenly start talking'.
- Assign the roles. Up to three visual characters (soul_a/b/c) + an optional narrator.
- Write the lines by speaker. Who says what in the scene — each in their own voice.
- Assemble the dialogue. The voiceover runs through a dialogue API, up to 10 voices per scene.
- The narrator is off-screen. Carries the story, not tied to any mouth on screen.
A character on screen speaks with their lips. The narrator is heard off-screen. The system never confuses the two.— the narrator principle
Соберите своего героя в Vini Studio
Облачный нодовый конструктор AI-видео: из идеи текстом — в готовый вертикальный ролик 9:16. Проект в закрытой бете — ищу авторов и партнёров.
✈️ Вступить в Telegram →Апдейты, ранний доступ и связь — в канале.
