OmniHuman 1.5 is ByteDance's model that brings a person in a single photo to life from a voice recording. Here it is the Best quality of our talking cards: the face, head and shoulders move with the words, not just the lips, and in a group photo you tap the one who speaks.
Updated 29.09.2026
С Днём Рождения! Жаль, что не могу обнять тебя сегодня, но вот моё поздравление.
Happy birthday! I wish I could hug you today - this is the next best thing.
Kling AI Avatar v2 costs less per second and suits a calm close-up. OmniHuman 1.5 is for a wider shot where the body moves too, and for a group photo where you choose the speaker.
| Model | Strong at | Length | Quality | Price |
|---|---|---|---|---|
| OmniHuman 1.5ByteDance | The whole person comes alive | up to 15 s of speech | 720p | 200 credits a second of speech + the voice |
| Kling AI Avatar v2Kuaishou | Face and head move with the words | up to 15 s of speech | 71 credits a second of speech + the voice |
Prices are in credits and are always shown on the button before you press it. Pricing
Add their photo - JPG, PNG or WebP up to 10 MB - from your device, your library or People, and tell the agent who it is for and what the day is. It makes a card with them in it. A ready scene such as "Candles for Mum" works too.
Press Make them speak below the card once it is ready. The agent drafts a short line to say aloud; edit it until it sounds like them. Pick the voice, then switch the quality to Best - that is OmniHuman 1.5; the default is Kling AI Avatar v2.
With Best on, tap the face of the person who should talk, or leave it to us. The price is on the button; press it and the talking card arrives in the chat. Email it, text it or send it in a messenger, now or on the day you pick.
Both talking-card models here take the same card and the same recorded line, and both stop at 15 seconds of speech. The difference is how much of the person moves. Kling AI Avatar v2, the default, animates the face and head and leaves the rest of the photo as it was. The grandmother above was made with it: seven seconds from one portrait and the Ava voice, and nothing else in the frame needed to move.
OmniHuman 1.5 is the livelier of the two. It treats the photo as a whole person, so a half-length shot, a raised glass or a wide family frame gives it something to do besides the mouth. It costs more per second of speech, so it pays to use it where that extra movement will show.
A working rule: a close-up and a quiet, gentle message - Kling AI Avatar v2. A wider shot, a toast, a group card or a line where the energy matters - OmniHuman 1.5.
Start from a clear face. The mouth should be visible - no hand, cup or microphone in front of it - with the person looking roughly towards the camera in even light. For OmniHuman, a frame that shows the shoulders or more beats a tight passport crop: it is built to move a whole person, so give it one.
Group cards need faces of a decent size. Six people fit on one card, yet likeness holds best with four or fewer, and a tiny face at the back of the table is hard to tap and gives the model little to move. If the speaker matters, put them near the front when you describe the card.
Last, the words. Fifteen seconds of speech fits a short toast, not a paragraph. Use the name the person is really called - Nana, Dedushka, Saba - and one detail only they would know: the swimming medal, the apple pie, the dog who ate the cake. Read the line aloud once before you press. If it trips you up, it will sound stiff in their mouth too.
OmniHuman 1.5 does not read your words itself. First a voice records the line, then the model moves the person to that sound. The voice is one of 8 catalogue voices from ElevenLabs Eleven v3, each with a sample you can play before you choose, or, on Plus, Family or Pro, your own voice, made with MiniMax Speech from one short reading in the browser. Because the lips follow a recording, the photo speaks whatever language the line is written in: Russian for Babushka, Hebrew for Saba, Arabic for Teta.
Your own voice opens one more use. You are abroad on your sister's birthday: put your own photo on the card, record your voice once, and the card says the greeting as you, on her morning.
The words stay in your hands the whole way. The agent drafts the line to be said aloud, every word can be changed, and you see roughly how many seconds it runs before the price appears on the button. If a card fails, the credits come back, and the full written greeting still goes under the video.
A talking card is one person, one line, straight to camera. If you want the whole scene to move - Mum blowing out the candles, glasses clinking, the table cheering - bring the card to life as a video instead. MiniMax H3, our Auto for video, animates every video scene in the catalogue; Veo 3.1 can make voices, singing and ambience in the same generation.
If the voice matters more than the face, a voice greeting read by ElevenLabs Eleven v3 goes out as sound alone. If you want everyone humming by dessert, a song made with Google Lyria 3.5 carries your own words in your language.
And a note on taste. A talking card works best with words the person in the photo would happily say. Ask Grandma on the phone what she wants to say, type it her way, and let her photo say it. Put words in someone's mouth they would never use, and the family will notice before the candles are out.
Make a talking card from Uncle Misha's photo for my cousin Anya's 30th birthday: half-length in the garden, in his good jacket, glass raised. In Russian, he tells her he has given this toast since she was one and will keep giving it.
Put Mum, Dad, my brother and me round the long garden table for my parents' 40th wedding anniversary, then make Dad speak: he raises his glass and thanks Mum for forty years of Sunday pancakes.
For my son Adam's 10th birthday: a talking card of his grandpa in the football stands, scarf on, arms up. In Hebrew, Grandpa tells him he will be at Saturday's match with the loudest drum in the stadium.
A talking photo card has someone you love say your greeting out loud - up to 15 seconds of speech,…
OpenKlingKling is Kuaishou's family of AI models, and three of them work here: Kling 3.0 Pro for video, Kling O3…
OpenElevenLabsEleven v3 is the text-to-speech model from ElevenLabs. Here it reads your greeting in one of 8 catalogue voices, in…
OpenOmniHuman 1.5 is ByteDance's model that turns one photo and a voice recording into a video of that person speaking, with the face, head and body moving to the sound. On Wish Generator it is the Best quality of the talking card: you make a photo card, edit the line, pick a voice, and OmniHuman makes the person in it say your greeting.
OmniHuman 1.5 moves the whole person, Kling AI Avatar v2 moves the face and head. Both take the same card and the same recorded line, and both stop at 15 seconds of speech. OmniHuman costs more per second and is the only one of the two that lets you tap who speaks in a group photo. For a quiet close-up, Kling AI Avatar v2 is usually enough; for a toast, a wide shot or a family table, choose Best.
Up to 15 seconds of speech, the same cap as Kling AI Avatar v2 - about as long as a short toast. The panel estimates the seconds as you type and will not start a line that runs over, so a line that is too long is never charged.
Yes, with OmniHuman 1.5. Switch the quality to Best and tap the face of the person who should talk; that face is marked for the model and that person says the line. If you do not tap, the model picks the speaker by itself. For group cards, this is the one to use.
Yes. OmniHuman does not read the words itself: it moves the lips to a recording, and that recording is made in the language you wrote the line in. Pick one of the 8 catalogue voices from ElevenLabs Eleven v3, or, on a paid plan, your own voice. Keep Hebrew and Arabic lines a little shorter: fewer letters fit into 15 seconds.
Yes. Add their photo from your device, your library or People - JPG, PNG or WebP up to 10 MB. The agent first makes a photo card with them in it, then you press Make them speak. Up to 6 people fit on a card, and faces are most reliable with 4 or fewer.
Make the card, press Make them speak, switch to Best and send it on the day.
Make a talking card