Seedance previously topped out at 15 seconds in Vini. The new model accepts an explicit duration from 4 to 30 seconds, so a longer scene can be designed as one arc instead of hiding joins between separate generations.

Lab #13 Seedance 2.5 is a real Vini project. Every person and fantasy character in it was generated inside Vini. This is not a staged demo: six five-second beats became one thirty-second result.
Measured file properties: 30.08 seconds, 480x854, six planned scenes, native audio.

KEY: seven references, six scenes, one result

Six generated characters — Ray, TIK-9, Vril, Kara, Elder Omu, and Commander Drak — share one object, the Singularity Key. Seven images become seven Elements. One Prompt Scriptwriter supplies six timed five-second scenes, and one Seedance 2.5 Video node returns the complete piece.

Seven generated sources from Lab #13: six characters and the Singularity Key. No uploaded faces were used.
Seven generated sources from Lab #13: six characters and the Singularity Key. No uploaded faces were used.
Contact sheet for KEY: six consecutive scenes and the final title inside one 30-second file.
Contact sheet for KEY: six consecutive scenes and the final title inside one 30-second file.
Six five-second scenes in KEY0-5Rayfinds the Key5 sec5-10TIK-9carries a signal5 sec10-15Vrilreads the trace5 sec15-20Karatakes the Key5 sec20-25Omureveals meaning5 sec25-30Drakends the journey5 sec

The timeline sets the sequence inside one run. Every five-second span has one dominant job.

30 secone coherent video
7 Elementsgenerated characters and object
6 scenesfive seconds each
1 runwithout stitching separate generations
What this run actually proves. The Lab #13 manifest is REF · V0/10 · I7/30. It proves a 30-second reference generation from seven generated images and native output audio. It does not prove input video or audio, editing, or extension, and those capabilities should not be attributed to this video.

A precise edit, not a reshoot

In Lab #11, the generated source clip enters “Element 1” as Video 1, while the generated logo reference enters “Element 2” as Image 1. Before the run, Vini shows EDIT · V1/10 · I1/30: one video to edit and one image defining the replacement.

The logo reference generated inside Vini for a precise final-faceplate edit.
The logo reference generated inside Vini for a precise final-faceplate edit.
Source on the left, edit on the right. Both video streams contain exactly 193 frames at 24fps and 480x854.

The car, Vini, confetti, camera, background, and decoded audio are preserved. Only the final faceplate changes from SOTKA 100 to stable Vini Studio; the background sign remains unchanged.

Five operations in one Video node

Two operations are illustrated here with completed real cases: reference generation and edit. The other three are listed as available working modes without borrowing evidence from those cases.

  1. From text. Create a new video with no media inputs and an explicit duration from 4 to 30 seconds.
  2. From frames. A first frame, or a first-and-last pair, sets the boundaries of the movement.
  3. New video by references. Images, video, and audio can each provide a distinct reference. Lab #13 illustrates this operation with images only.
  4. Edit Video 1. The model changes the requested part and the result follows the source duration. Lab #11 illustrates this operation.
  5. Extend Video 1. The result is a separate new continuation with an explicitly selected duration from 4 to 30 seconds.

How video inputs work in Vini

Any media can be collected in an Element and sent through the universal “Element N” row on a Video node. A bare clip can connect directly to the same row. The wire is universal and appears on every Video node; the selected model decides compatibility.

When the current adapter cannot accept that operation, Vini does not remove the wire. It preserves the graph and blocks before a paid submit, with the reason shown immediately.

The universal media path through an Element into a Video nodeSourcesimage · video · audioElementcollects mediaElement Nuniversal inputModel adapterchecks operationResultnew MP4compatible · run enabledincompatible · wire preservedblocked before submit

The same wire remains part of the graph when the model changes. Only the compatibility result changes before a run.

Limits are checked before a run. Up to 30 images, 10 videos, 10 audio files, and 50 media total. Each input video or audio file must run from 2 to 30 seconds; aggregate video and audio duration must not exceed 30 seconds. Vini never truncates excess input silently: an overage blocks the run.

One source, one job

Assign one explicit job to every Image, Video, and Audio input. Break a long action into timed beats. For audio, say whether it supplies rhythm, ambience, voice style, or SFX. When silence matters, write “No dialogue” explicitly.

Use Image 1 only for the Key's shape. Use Video 1 only for camera motion and smoke behavior. Use Audio 1 only for rhythm and ambience. 0-4s: the Key rises above the wreckage. 4-8s: one slow push-in settles on the gold rings. No dialogue and no spoken words.
Current boundary: uploaded photos and videos of real people are temporarily unsupported. Real-person media cannot be used as input for now. Generated people and characters remain supported — both Lab cases on this page use generated sources.

First-release contract

Output is 480p and MP4. New generation and extension use an explicit 4-30 second duration; an edit follows the source video duration.

One source, one responsibility. One timed beat, one dominant action.— the Vini prompting rule

Build your own hero in Vini Studio

A cloud AI video generator: from an idea in words to a finished clip with sound, 9:16 or 16:9. The project is in closed beta — I am looking for authors and partners.

Get started in the studio →Sign in with Telegram or Google. No account needed in advance — it is created for you.