Video

Fifteen seconds for the whole toast: Kling 3.0 from one photo

Kling 3.0 Pro is Kuaishou's video model, and of the three Kling models we run it is the only one that sets a whole scene in motion. Start it from one photo of your people and it plays the moment on for 5, 8, 10 or 15 seconds at 1080p: people get up, glasses meet, somebody gets hugged, with the sound of it or in silence.

  • Made by Kuaishou
  • from 700 credits
  • 5-15 s
  • With or without sound

Updated 29.09.2026

Two 5-second Kling 3.0 clips, both with sound

The whole family raises a glassKling 3.0 Pro · 5 s · 1080p · sound
Prompt

The family around the long garden table raise their glasses together and laugh, grandmother at the head of the table smiles and wipes a happy tear, leaves move in a light breeze, warm golden light, natural gentle camera push-in, sounds of clinking glasses, laughter and birdsong

Confetti at a birthday picnicKling 3.0 Pro · 5 s · 1080p · sound
Prompt

Friends at the picnic in paper party hats blow party horns and toss confetti into the air, one of them laughs and hugs the birthday girl, confetti floats down slowly, handheld camera feel, sunny park, sounds of party horns, cheering and laughter

What Kling 3.0 does that its Kling siblings here don't

People leave their seats

Kling AI Avatar v2 moves a face and a head; Kling O3 makes stills. Kling 3.0 moves people through the frame. In the picnic clip above, by the fourth second the friend in the denim jacket is up off the blanket and hugging the birthday girl, confetti still in the air.

The longest take in our picker

Choose 5, 8, 10 or 15 seconds. No other video model here goes past 10, and Veo 3.1 stops at 8. Think in beats: five seconds holds one gesture, ten holds a gesture and a reaction, fifteen holds a small story with a start and an end.

Faces guarded on every request

Your photo is the first frame, so the clip starts with your actual people. With every Kling 3.0 request we also send a standing instruction to avoid morphing faces, extra people and stray text. In the garden-table clip, all four people still look like themselves when their glasses meet.

Sound you can switch off

MiniMax H3, our Auto for video, always comes with sound. Kling 3.0 lets you choose: clinks, laughter and birdsong made in the same generation as the picture, or a silent clip that costs less per second. Silent is the right call when the clip travels with a song you made here.

Making a Kling 3.0 clip: picture, length, sound

  1. 1

    Pick the first frame

    Use the button on this page, which opens Video with Kling 3.0 Pro already set, or switch the composer to Standard and choose it under Model. Attach one picture: a photo of your people (JPG, PNG or WebP up to 10 MB) or a card you made here.

  2. 2

    Choose a length, then write the beats

    Decide on 5, 8, 10 or 15 seconds first, because the length sets how much can happen. Then write the moment in order: who moves first, what happens next, how it ends, and last of all, what we should hear.

  3. 3

    Set sound, check the price, send

    Leave sound on or switch it off; the button shows the new price after every change, and a failed clip returns its credits. The finished clip goes out by email, SMS or a popular messenger, straight away or on the morning of the day.

Start a Kling 3.0 clip

Three Kling models here, three different jobs

Kuaishou makes all three Kling models we run, and they barely overlap. Kling O3 draws still pictures and is one of the cheapest picture models in our picker, handy for trying out a setting. Kling AI Avatar v2 is the engine of our talking card: one portrait, your text, a voice you choose, up to 15 seconds of speech. Kling 3.0 Pro is the only one of the three that animates a whole scene: several people, the camera, the room and the sound of it.

The line between them matters most when words are involved. Kling 3.0 makes the sounds of a moment, the clinking and the laughter, but it does not read out the greeting you typed. When Grandma should say "Happy birthday, Leo" herself, you need a talking card, and the Better option there is Kling AI Avatar v2.

They also work well one after another. Make the picture of your people first (for faces, GPT Image 2.5 Flare, our Auto for pictures, keeps the likeness best), bring it to life with Kling 3.0, and if the clip is silent, send a song made here with Lyria 3.5 alongside it.

Does the face hold? What our two clips show

The garden-table clip starts from one of our scene pictures: four people at a table under fruit trees, a birthday cake in the middle. Over 5 seconds the camera pushes in and all four glasses meet over the cake. Each face stays the same person from the first frame to the last, including the younger woman who turns away from us to reach in.

The picnic clip is the harder test. Four friends in paper hats blow party horns, then one jumps up and crosses the blanket to hug the birthday girl. For a moment in mid-jump she blurs, as anyone would on a phone camera, and then the two of them settle into the hug as the same people who started the clip.

What slips is whatever the first frame never showed. Someone who walks in from outside the picture is invented, because the model has never seen them, and a tiny face at the far end of a table gives it little to hold on to. So start from a picture with everyone already in it, faces clear and in the light. If your people are in separate photos, make the card first: Flare for up to four people, Nano Banana Pro for five.

Kling 3.0 prompts, one for each length

Adam leans over his cake and blows out all six candles in one go, his big sister claps and hugs him from behind, smoke curls up from the candles, warm kitchen light, sounds of blowing, clapping and a small cheer
For 5 seconds with sound, from a photo of the two of them at the cake. One action and one reaction is what five seconds can hold, the same size of moment as the clips above.
Use this prompt
The two kids on the sofa jump up and throw themselves at Dad, he pretends to be squashed and falls back into the cushions, then all three burst out laughing, the dog barks once, warm lamp light, camera at sofa height
For 10 seconds: an action, a reaction, an ending. Everyone who moves is in the first frame already, so nobody has to be invented halfway through the clip.
Use this prompt
Grandpa at the head of the long table taps his glass with a fork, everyone turns to him, he stands and lifts the glass, then the whole table stands and glasses meet all the way down, Grandma pulls him down for a kiss on the cheek, evening light, the camera slowly pulls back to show the whole table, sounds of the fork on glass, chairs scraping and cheering
For 15 seconds: a toast in three beats, from one tap on a glass to the whole table on its feet. Written in order, it gives Kling 3.0 something to follow for all 15 seconds.
Use this prompt

Kling 3.0 Pro in numbers, priced live

These are the settings Kling 3.0 Pro runs on here today and what a second costs, with sound and without. For a five-second clip that always has sound, MiniMax H3 costs less; for a spoken line or a chorus, look at Veo 3.1.

ModelStrong atLengthQualitySoundPrice
Kling 3.0 ProKuaishouSmooth motion5, 8, 10, 15 s1080pWith or without soundfrom 700 credits · 210 cr a second

Prices are in credits and are always shown on the button before you press it. Pricing

Keep exploring

Questions and answers

Mostly, yes: your photo is the first frame, and every Kling 3.0 request we send asks it to avoid morphing faces and extra people. In our garden-table clip all four faces hold for the full 5 seconds. Likeness slips with big turns, very small faces and anyone who walks in from outside the picture, so start from a clear photo with everyone in it.

Give the moment time to turn

Pick one photo of your people, write the beats in order, and let Kling 3.0 play it out for up to 15 seconds.

Start a Kling 3.0 clip

Greetings for every moment