Talking cards

Make a photo say your greeting with OmniHuman 1.5

OmniHuman 1.5 is ByteDance's model that brings a person in a single photo to life from a voice recording. Here it is the Best quality of our talking cards: the face, head and shoulders move with the words, not just the lips, and in a group photo you tap the one who speaks.

  • Made by ByteDance
  • 71-200 credits a second of speech
  • up to 15 s of speech

Updated 29.09.2026

One grandmother, one grandfather, two engines

Grandpa speaks RussianOmniHuman 1.5 · 1 photo · voice Ari · Russian
Prompt

С Днём Рождения! Жаль, что не могу обнять тебя сегодня, но вот моё поздравление.

Grandma says it herselfKling AI Avatar v2 · 1 photo · voice Ava
Prompt

Happy birthday! I wish I could hug you today - this is the next best thing.

OmniHuman 1.5 next to the cheaper Kling AI Avatar v2

Kling AI Avatar v2 costs less per second and suits a calm close-up. OmniHuman 1.5 is for a wider shot where the body moves too, and for a group photo where you choose the speaker.

ModelStrong atLengthQualityPrice
OmniHuman 1.5ByteDanceThe whole person comes aliveup to 15 s of speech720p200 credits a second of speech + the voice
Kling AI Avatar v2KuaishouFace and head move with the wordsup to 15 s of speech71 credits a second of speech + the voice

Prices are in credits and are always shown on the button before you press it. Pricing

Switching a card to Best, step by step

  1. 1

    Make the photo card

    Add their photo - JPG, PNG or WebP up to 10 MB - from your device, your library or People, and tell the agent who it is for and what the day is. It makes a card with them in it. A ready scene such as "Candles for Mum" works too.

  2. 2

    Press Make them speak

    Press Make them speak below the card once it is ready. The agent drafts a short line to say aloud; edit it until it sounds like them. Pick the voice, then switch the quality to Best - that is OmniHuman 1.5; the default is Kling AI Avatar v2.

  3. 3

    Tap who speaks, then send

    With Best on, tap the face of the person who should talk, or leave it to us. The price is on the button; press it and the talking card arrives in the chat. Email it, text it or send it in a messenger, now or on the day you pick.

Make a talking card

Face and head, or the whole person

Both talking-card models here take the same card and the same recorded line, and both stop at 15 seconds of speech. The difference is how much of the person moves. Kling AI Avatar v2, the default, animates the face and head and leaves the rest of the photo as it was. The grandmother above was made with it: seven seconds from one portrait and the Ava voice, and nothing else in the frame needed to move.

OmniHuman 1.5 is the livelier of the two. It treats the photo as a whole person, so a half-length shot, a raised glass or a wide family frame gives it something to do besides the mouth. It costs more per second of speech, so it pays to use it where that extra movement will show.

A working rule: a close-up and a quiet, gentle message - Kling AI Avatar v2. A wider shot, a toast, a group card or a line where the energy matters - OmniHuman 1.5.

A photo with something to move

Start from a clear face. The mouth should be visible - no hand, cup or microphone in front of it - with the person looking roughly towards the camera in even light. For OmniHuman, a frame that shows the shoulders or more beats a tight passport crop: it is built to move a whole person, so give it one.

Group cards need faces of a decent size. Six people fit on one card, yet likeness holds best with four or fewer, and a tiny face at the back of the table is hard to tap and gives the model little to move. If the speaker matters, put them near the front when you describe the card.

Last, the words. Fifteen seconds of speech fits a short toast, not a paragraph. Use the name the person is really called - Nana, Dedushka, Saba - and one detail only they would know: the swimming medal, the apple pie, the dog who ate the cake. Read the line aloud once before you press. If it trips you up, it will sound stiff in their mouth too.

Whose voice comes out of the photo

OmniHuman 1.5 does not read your words itself. First a voice records the line, then the model moves the person to that sound. The voice is one of 8 catalogue voices from ElevenLabs Eleven v3, each with a sample you can play before you choose, or, on Plus, Family or Pro, your own voice, made with MiniMax Speech from one short reading in the browser. Because the lips follow a recording, the photo speaks whatever language the line is written in: Russian for Babushka, Hebrew for Saba, Arabic for Teta.

Your own voice opens one more use. You are abroad on your sister's birthday: put your own photo on the card, record your voice once, and the card says the greeting as you, on her morning.

The words stay in your hands the whole way. The agent drafts the line to be said aloud, every word can be changed, and you see roughly how many seconds it runs before the price appears on the button. If a card fails, the credits come back, and the full written greeting still goes under the video.

Moments a talking card cannot carry

A talking card is one person, one line, straight to camera. If you want the whole scene to move - Mum blowing out the candles, glasses clinking, the table cheering - bring the card to life as a video instead. MiniMax H3, our Auto for video, animates every video scene in the catalogue; Veo 3.1 can make voices, singing and ambience in the same generation.

If the voice matters more than the face, a voice greeting read by ElevenLabs Eleven v3 goes out as sound alone. If you want everyone humming by dessert, a song made with Google Lyria 3.5 carries your own words in your language.

And a note on taste. A talking card works best with words the person in the photo would happily say. Ask Grandma on the phone what she wants to say, type it her way, and let her photo say it. Put words in someone's mouth they would never use, and the family will notice before the candles are out.

Requests with a raised glass and a football scarf

Make a talking card from Uncle Misha's photo for my cousin Anya's 30th birthday: half-length in the garden, in his good jacket, glass raised. In Russian, he tells her he has given this toast since she was one and will keep giving it.
A half-length shot with a raised glass gives OmniHuman more than a mouth to move. Switch the quality to Best, and check the line is in Russian: the voice reads whatever language it is written in.
Use this prompt
Put Mum, Dad, my brother and me round the long garden table for my parents' 40th wedding anniversary, then make Dad speak: he raises his glass and thanks Mum for forty years of Sunday pancakes.
Four faces is the most a card keeps reliably. Choose Best and tap Dad, so the toast comes from him and not from whoever the model would have picked.
Use this prompt
For my son Adam's 10th birthday: a talking card of his grandpa in the football stands, scarf on, arms up. In Hebrew, Grandpa tells him he will be at Saturday's match with the loudest drum in the stadium.
Arms up and a scarf give OmniHuman a whole person to animate, not just a face. Keep the line short: fifteen seconds hold fewer Hebrew letters than English ones, and the counter shows what is left.
Use this prompt

Keep exploring

OmniHuman 1.5, asked and answered

OmniHuman 1.5 is ByteDance's model that turns one photo and a voice recording into a video of that person speaking, with the face, head and body moving to the sound. On Wish Generator it is the Best quality of the talking card: you make a photo card, edit the line, pick a voice, and OmniHuman makes the person in it say your greeting.

Let the whole photo raise a glass

Make the card, press Make them speak, switch to Best and send it on the day.

Make a talking card

Greetings for every moment