You ask for «a helicopter turns into a robot». You will get a transformation. Which robot exactly is up to the model — and on the second take it will be a different robot. The gap between «got lucky» and «got what I planned» is one input, and not every model has it.

What we put on the table

One prompt, two frames, nine models. The opening frame: a helicopter on a wet apron at dusk, a worker in a yellow vest by the landing gear, a red beacon to the right. The closing frame: the robot folded out of that helicopter — rotor still on its back, the glazed cockpit turned into chest armour, an amber visor for eyes. Same apron, same worker, same beacon.

The opening frame: the helicopter on the apron. Every one of the nine clips starts here.
The opening frame: the helicopter on the apron. Every one of the nine clips starts here.
The closing frame: the robot made of that helicopter. Rotor on the back, cockpit in the chest, amber visor.
The closing frame: the robot made of that helicopter. Rotor on the back, cockpit in the chest, amber visor.

Five models, the very same five seconds

The four columns on the left were given both frames. The right column got only a reference — because that model has no last-frame input at all. Watch the final second: on the left, four different models arrive at one and the same frame; on the right, at a robot of the same family but in a composition of its own.

Exactly the same task. The four left clips arrive at the frame we specified; the right one arrives at a frame of its own. No audio.
Mid-transition: every model takes its own route — Kling is still rotating horizontally, Seedance already stands, FLUX has an arm raised.
Mid-transition: every model takes its own route — Kling is still rotating horizontally, Seedance already stands, FLUX has an arm raised.
The ending: the four left models arrived at the very same frame. Different roads, one destination — because the destination was specified.
The ending: the four left models arrived at the very same frame. Different roads, one destination — because the destination was specified.
A last frame is a contract. A reference is a hint. A model holding two frames has to bring the picture to the ending you specified: it knows where it starts and where it must arrive, and the prompt only describes the road between them. A model holding a reference gets a wish — «make it look like this» — and re-interprets it every single run: it will land in the right family, but not on the right frame, because the framing, the angle and what stays in shot are its own choice.
Kling v3 with both frames: the ending matches the robot we specified — the rotor folded onto the shoulders, the cockpit in the chest, the worker still in shot.
Kling v3 with both frames: the ending matches the robot we specified — the rotor folded onto the shoulders, the cockpit in the chest, the worker still in shot.
HappyHorse 1.1, a reference instead of a last frame. The reference landed in the same family, but not on the same frame: the rotor stands upright instead of folding onto the shoulders, the framing is tighter, and the worker is gone. A last frame pins not just the subject but the composition.
HappyHorse 1.1, a reference instead of a last frame. The reference landed in the same family, but not on the same frame: the rotor stands upright instead of folding onto the shoulders, the framing is tighter, and the worker is gone. A last frame pins not just the subject but the composition.
It is not that some models are worse. It is that two models in our registry physically do not accept a last frame — and no prompt will change that.

Which models take a last frame

The Vini registry holds nine video-model entries. All nine have a frame mode, seven of them accept a last frame. Two do not: Grok Video 1.5 and HappyHorse 1.1 — their request carries an opening frame only. Gemini Omni Flash was on that list until version 1.1, which added end_image_url and taught the model to interpolate between the first and last frames.

Which models take a last frametakes a last frameKlingSeedance 2.0Seedance 2.5MiniMax H3FLUX 3Gemini Omni FlashWanopening frame onlyGrok Video 1.5HappyHorse 1.1frames and references togetherKling v3 · Kling O3 · Kling v3 Motion

The input matrix. Kling stands apart — both generations: only it carries frames and references to the model in one request, for the others these are mutually exclusive modes.

How to check this instead of taking my word

Open the «Видео» node and switch the model. On Kling, Seedance, MiniMax H3 and Gemini Omni Flash a «финальный кадр» input appears on the node. On Grok and HappyHorse it will not — it is not hidden and not disabled, it simply does not exist, because the model cannot do it. The port is drawn from the registry, not from a setting.

Building the transition, step by step

  1. Make two frames of one scene. An opening and a closing one: same setting, same light, same angle, same scale. If the scene drifts, so will the transition. We generated both frames in Grok Image 2 from one description, changing only the state of the object in it.
  2. Describe the object in an «Элемент» node. The description is not a note to self: it travels as an instruction into the frame generator. The more specific the materials and joints, the less the model has to invent.
  3. Wire both frames into «Видео». The opening one into «первый кадр», the closing one into «финальный кадр». A «+финал» badge appears on the node card — that is your proof the ending really goes to the model instead of getting lost on the way.
  4. Pick a model that has that input. If there is no «финальный кадр» port, switch the model rather than rewriting the prompt. The list of eight is in the table below.
  5. Write the prompt about the transition, not the objects. What unfolds, what rises, where the camera goes. The model has already seen both frames — describing them again in words only competes with the pictures.
  6. Match the duration to the movement. Five seconds: one state becomes another. Eight to ten: a full action reads, with a wind-up and a final pose.

When you need a frame AND references at once

Only Kling can do this — all three of its entries, Kling v3, Kling O3 and Kling v3 Motion: one request carries a frame and Elements together (the first two take an opening frame, a closing frame and up to three Elements; Motion takes a guide frame and one face element). On every other model a frame and references are mutually exclusive modes, and the studio will honestly detach one wire when you connect the other rather than send a request the vendor would quietly mangle.

What it cost

The same transition, nine models, a real run. The clip price is the provider rate times the duration ordered.

ModelClipRateResolutionLast frame
Seedance 2.0 · быстрое$0.27$0.054/s × 5s496×864
MiniMax H3$0.19$0.038/s × 5s484×868
Grok Video 1.5$0.48$0.08/s × 6s480×848
Kling$0.63$0.126/s × 5s720×1280
HappyHorse 1.1$0.70$0.14/s × 5s720×1280
Gemini Omni Flash$0.80$0.10/s × 8s720×1280
Seedance 2.5$0.83$0.104/s × 8s480×854
FLUX 3$0.85$0.17/s × 5s704×1280
Seedance 2.0 · полное$1.02$0.068/s × 15s496×864
This is cost price, with no markup. In Vini the generation price comes from one file — both for the estimate shown before you run and for the charge after — verbatim «prices are 1:1 provider cost (no markup)». You pay exactly what the provider costs.
Prices change, and that is normal. The numbers in the table are provider rates as of 18 August 2026. Providers revise them, models arrive and leave, resolution tiers get repriced. The current price is always shown in the studio on the node card BEFORE you run, not after the bill — trust that one. For fal models the number in the table is charged exactly; for Seedance the price is metered by actual work, so the final figure differs from the estimate by a couple of cents either way.
More expensive does not mean more pixels. Seedance 2.5 returned 480×854 at $0.104 per second, while Kling returned 720×1280 at $0.126. Nearly the same second, twice the picture.

All nine in a row

The same experiment in full, as a vertical reel: two opening frames and nine clips with titles and costs.

The figures on the badges are from that very run. For current rates look in the studio, not here.
9models on one task
7 of 9registry models take a last frame
1model takes frames and references together
$0.28cheapest clip with a guaranteed ending

Build your own hero in Vini Studio

A cloud AI video generator: from an idea in words to a finished clip with sound, 9:16 or 16:9. The project is in closed beta — I am looking for authors and partners.

Get started in the studio →Sign in with Telegram or Google. No account needed in advance — it is created for you.