A talking photo card has someone you love say your greeting out loud - up to 15 seconds of speech, lips in time with the words. Kling AI Avatar v2 from Kuaishou moves the face and head; OmniHuman 1.5 from ByteDance brings the whole person to life. The voices come from ElevenLabs, or from you.
Updated 29.09.2026
Happy birthday! I wish I could hug you today - this is the next best thing.
С Днём Рождения! Жаль, что не могу обнять тебя сегодня, но вот моё поздравление.
Tell the agent something like "A talking card for Noah's 8th birthday, from his grandma Ina" and add her photo when it asks - JPG, PNG or WebP up to 10 MB. It makes the photo card first. Or make a Photo card yourself in Standard: a portrait by the window, or a scene like "Long garden table".
Under the finished card, press Make them speak. The agent writes a short spoken version of your greeting - edit it or type your own, and a counter shows roughly how many seconds it takes. Pick the voice. Anything past 15 seconds is flagged before it starts.
Under Quality, Standard is Kling AI Avatar v2, which moves the face and head; Best is OmniHuman 1.5, which moves the whole person - tap the face of the one who speaks. The price is on the button and nothing is charged before you press. Then send it by email, SMS or a messenger.
A talking card for my grandson Leo's 8th birthday, from me: a portrait at my kitchen window in soft morning light, looking into the camera and smiling, a small cake with eight candles on the sill. I want to say: Happy birthday, Leo, eight already and taller than the kitchen table. Save me a slice, I'm coming in July.
A photo card of Grandpa Joe for my daughter Mia's 6th birthday: head and shoulders in his woodworking shed, sawdust in the sunlight, smiling at the camera and holding up the little wooden horse he made for her.
A photo card of me for Dana's birthday while she's away for work: head and shoulders on our balcony at dusk, holding a cupcake with one candle, city lights soft behind me, looking into the camera.
Both engines get the same two things from us: the card, and the voice, already recorded and trimmed to the seconds you paid for. The difference is how much of the picture moves. Kling AI Avatar v2, from Kuaishou, works on the face and head: the mouth follows the words, the head moves with them, and the rest of the photo stays as it was, from the clothes to the room behind. For a portrait of one person that is usually exactly right, and it costs less per second. The grandmother on this page was made this way.
OmniHuman 1.5, from ByteDance, animates the whole person: head, face and shoulders move with the rhythm of the voice, so a half-length shot or a raised glass gives it more to do than the mouth. The grandfather on this page, speaking Russian in the Ari voice, was made with it from one portrait. It also solves the group-photo problem: switch the quality to Best, tap the face of the one who should talk - Grandpa thanking everyone, say - and he speaks while everyone else stays in the picture, listening. Without a tap the model chooses the speaker, which is fine for a portrait and a gamble for a crowd, so keep the table small: a tiny face at the far end is hard to tap. Because it costs more per second, Kling AI Avatar v2 is the everyday choice.
The face that talks is the face on the card, so likeness starts there. GPT Image 2.5 Flare, our Auto choice for pictures, has the best likeness, and cards with up to four people keep faces most reliably. For the talking part, a few things help: the face turned towards the camera, the mouth in view rather than behind a cup or a hand, even light, and a good share of the frame for the face. For OmniHuman 1.5, keep the shoulders in view as well. One person in a portrait is the easy case; a tiny face in a wide landscape is the hard one.
If your only photo is a group shot from across the garden, make a new card first, closer in, and let that one talk. People saves you a step: when Grandma is saved there with a photo, the agent already has it and does not ask again.
Fifteen seconds is a toast, not a letter: the name, one real detail, one wish. The grandmother's line on this page, "Happy birthday! I wish I could hug you today - this is the next best thing.", runs 7 seconds, so the limit holds about two lines like it. The voice pauses at full stops, and a sentence that runs for three lines comes out as one long breath. "Happy birthday to my amazing grandson, wishing you all the best on your special day, may all your dreams come true" says less than "Happy birthday, Noah. Eight already, and you still owe me a rematch at checkers. I'm bringing the cake on Sunday."
Use the name the person is really called: Nana, Dedushka, Saba, Teta. Keep Hebrew and Arabic lines a little shorter, since 15 seconds holds fewer letters in them than in English. Skip stage directions like [laughs]: the app strips out anything in square brackets before the voice reads the line. The best words usually come from the person in the photo, so call Grandma, ask what she would say and type it her way. Then read the line aloud once; if it trips you up, it will sound stiff in their mouth too.
Choose the voice for the listener, not for yourself: a slow, warm one for a grandparent, a bright one for a six-year-old. Each catalogue voice has a sample, so choose by ear rather than by name. And one rule that is not technical: use photos of people who would be glad to see themselves talk, and words they would actually say. Put a phrase in Grandpa's mouth that he would never use, and the family will notice before the candles are out.
Most talking cards stand in for someone who is missing. Grandma is two flights away on the day, so her portrait says happy birthday from a phone propped against the cake. If she has her own account on a paid plan, she records her voice once, and then it really is her speaking.
Distance is not the only reason. Filming yourself saying happy birthday takes a dozen attempts, and the good one has your thumb in it; here you pick a photo you already like and type two sentences. Time zones make birthday calls awkward, so a card with your face, and on a paid plan your own voice, can wait for them: schedule it for their morning, not yours.
Language counts too. The grandchildren answer in English while Babushka still thinks in Russian. Write her line in Russian, the way she says it, and her photo speaks it in Russian, because the voice reads whatever language the line is written in.
A talking card is one person, a few sentences, looking at the camera. It is not a party. If you want the whole table raising glasses, with clinking and laughter, that is a video. Kling 3.0 turned our long-garden-table card into a 5-second clip of exactly that, sound included. Veo 3.1 makes voices and singing in the same generation - on its page, a family sings the end of Happy Birthday and Mum blows out the candles. MiniMax H3 animates every video scene in our catalogue and is the best value of the three.
If the voice matters and the face does not, send a voice greeting: the same eight voices, or yours, with no picture to animate, and it costs less. If you want the words sung, Lyria 3.5 sings lyrics you have approved, in English, Russian, Hebrew, Arabic and more.
Both speak the same line in the same voice, up to 15 seconds. Take Kling AI Avatar v2 for one face in a portrait - it costs less per second. Take OmniHuman 1.5 when the whole person should move, or when you need to choose who talks in a group photo.
| Model | Strong at | Length | Quality | Price |
|---|---|---|---|---|
| Kling AI Avatar v2Kuaishou | Face and head move with the words | up to 15 s of speech | 71 credits a second of speech + the voice | |
| OmniHuman 1.5ByteDance | The whole person comes alive | up to 15 s of speech | 720p | 200 credits a second of speech + the voice |
Prices are in credits and are always shown on the button before you press it. Pricing
OmniHuman 1.5 is ByteDance's model that brings a person in a single photo to life from a voice recording. Here…
OpenKlingKling is Kuaishou's family of AI models, and three of them work here: Kling 3.0 Pro for video, Kling O3…
OpenElevenLabsEleven v3 is the text-to-speech model from ElevenLabs. Here it reads your greeting in one of 8 catalogue voices, in…
OpenAI voice greetingAn AI voice greeting is your written greeting turned into a voice message they can play on speaker, in the…
OpenMake a photo card with the person in it, then press Make them speak under the finished card. Edit the short line the agent drafts, pick the voice and the quality, and press the button with the price on it. The card comes back as a video of that person saying your words, lips in time with the voice.
By the second: each second of speech plus the voice, and OmniHuman 1.5 costs more per second than Kling AI Avatar v2, so a short line at Standard quality costs the least. The 1,000 free credits at sign-up are enough to try a short one, and your own voice is a one-time charge on a paid plan. Text greetings are free, the price is on the button before you press, and a failed card gives the credits back.
Up to 15 seconds of speech - two or three short sentences. The grandmother's line on this page takes 7 seconds. The agent writes the spoken line to fit, a counter under it shows roughly how many seconds it will take, and a line that runs over is flagged before anything is charged. Your full written greeting still goes with the card, under the video.
Yes, with OmniHuman 1.5. Switch the quality to Best and tap the face of the person who should talk; they say the line while everyone else stays in the picture. If you do not tap, the model picks the speaker itself. Kling AI Avatar v2, the Standard quality, has no such choice, so for a group card where it matters, use Best.
One where the person looks towards the camera, the mouth is in view rather than behind a cup or a hand, the light is even and the face fills a fair share of the frame. For OmniHuman 1.5, keep the shoulders in view too. A tiny face in a wide group shot is the hard case: make a new card closer in first, and let that one talk.
In your own voice, yes, on the Plus, Family and Pro plans: read a short passage aloud in the browser, about 20 seconds, and MiniMax Speech makes your voice, paid for once. Not in anyone else's, because old recordings and voice notes cannot be uploaded; your mum can record her own voice on her own paid account. On any plan you can pick one of eight ElevenLabs voices instead and hear each first.
Yes. The line can be in English, Russian, Hebrew or Arabic, and the voice speaks whichever language you write it in. The eight catalogue voices come from ElevenLabs Eleven v3, and on a paid plan your own voice works too. Keep it short: 15 seconds holds fewer letters in Hebrew and Arabic than in English, and the counter shows the room left.
One photo, two or three sentences and a voice - and the birthday card says it out loud.
Make a photo talk