VlogMe

Talking photo tool Make a photo talk with one portrait and a voice

VlogMe's talking photo generator runs in Video Studio. Upload a clear portrait, add the words — a recording you already have or a script with an AI voice — and VlogMe renders a short video in which the face speaks them. You see the credit quote before each paid generation, and the finished video is ready to preview and download.

  • Quote before each paid generation
  • Your recording or an AI voice
  • Preview and download

Try it with your own files

Set up your generation

Sign in to generate. This page opens Video Studio's Talking avatar workflow with Hedra Avatar selected. The optional field above is for delivery direction only; the spoken words are added in Video Studio.

  1. Start with a clear portrait
  2. Add a recording or typed script
  3. Set the model and framing
  4. Check the quote, then generate

01Talking photo examples

What a talking photo looks like

A talking photo is a short video made from one still image: the lips and face follow a speech track, so a portrait reads a line, a mascot introduces a product, or a presenter delivers a message — without filming anything.

Review your own result the same way before you use it.

A still selfie turned into a talking video: a woman in denim overalls greets her followers
Talking avatar9:16
Source selfie: a woman in denim overalls under a thatched roof
Original portrait

02The inputs and the quote

A portrait, a voice track, and a quoted render.

Lip-synced talking portrait: a man in glasses and a vest delivers a quick line
Lip-sync9:16

01Two voice paths

Give the portrait a voice — two ways

The video follows a speech-audio track, supplied in either of two ways.

Use a recording you already have. Upload an MP3, M4A, WAV or OGG file, or select a voice track from Lab. No voice generation is needed; you go straight to the video quote.

Type the words and generate a voice. Open “No voice track yet? Type the words”, enter up to 5,000 characters in “Words to speak”, pick a voice from the menu, and press “Make the voice track”. The voice-track quote appears first; once the audio is ready it becomes the speech track for the render.

Either way, the words come from the audio. The optional “How should they say it?” field only shapes delivery — tone, gaze, expression.

Try it with your own source
A presenter in a green blouse introduces a garden workshop
More than a talking head16:9

02Model and framing

Set the model and framing

Hedra Avatar is selected by default and animates your uploaded portrait in vertical 9:16, landscape 16:9 or square 1:1. Confirm that you have the rights to the portrait and the voice.

OmniHuman 1.5 also animates your portrait and keeps the source image's framing. HeyGen Avatar V uses a fixed presenter rather than your photo, so leave it unselected when the point is to make your own picture talk.

Use a vertical greeting or announcement for Stories, Reels or Shorts, a landscape explainer for a product update, or a square feed clip made from the same portrait and a new script.

Try it with your own source
Vertical social clip: a woman with coffee and flowers shares a morning ritual
Ready for social9:16

03Credits and separate quotes

What it costs

Two things can be charged, and each is quoted separately before you confirm it.

Voice track from text: 45 credits per started 1,000 characters, only if you type a script and generate a voice.

Hedra Avatar video: 5 credits × source-audio duration rounded up to the next whole second, for every render.

Using your own recording? No voice-track charge — only the video quote. Typing a script? The voice-track quote comes first; the video quote follows once the audio exists, because the video price depends on its length. The combined total is not available until the voice track is ready.

One tested example: a two-sentence English test line was quoted 45 credits for the voice track and 30 credits for the Hedra render — 75 credits for that clip, not a standard package price. Costs can increase with script length and audio duration; the quote shows the applicable charge.

New accounts start with zero credits, so a talking-photo render needs credits before it runs. A subscription is optional — you can buy a one-time credit pack instead. Each generation is quoted before you confirm it. See pricing for current packs and plans.

Try it with your own source

Video StudioAI avatar generator

Every talking-avatar tool in one place
  1. Your starting materialA portrait and a speech-audio track
  2. The resultA lip-synced talking photo video

Use an existing recording or generate a voice from text. Voice generation and the avatar render have separate quotes; editing, dubbing and publishing open separate workflows.

Continue in Video Studio

03Built for the real workflow

How to make a photo talk in Video Studio

  1. Start with a clear portrait

    Choose a photo where the face is well lit and facing the camera, with the mouth visible. JPG, PNG and WebP are accepted. In Video Studio, upload the portrait or pick one already saved in Lab.

  2. Add the voice, model and framing

    Upload a recording or choose one from Lab, or type the words and generate a separately quoted voice track. Keep Hedra Avatar to animate your own portrait, choose 9:16, 16:9 or 1:1, and confirm your rights to the portrait and voice.

  3. Check the quote, generate, preview, download

    Press “Create avatar video”. Video Studio shows the exact credit quote before you confirm this video render. Confirm, wait for the job, then preview the result in the Studio player and use “Download” to save it. “Generate again” requests another version and is quoted again.

  • Script → AI voice
  • Uploaded audio
  • 9:16
  • 16:9
  • 1:1

Finished examples

A craftsman, a sushi mascot, an ad

Source portrait: a woodworker in a leather apron
Source image
Source image: a 3D salmon sushi character
Source image
Source portrait: a woman holding an unlabeled spray bottle
Source image

Your finished clip

Preview, download, and Lab

When the job completes, the clip plays in the Studio player. From there you can download the video — Hedra Avatar renders a 720p MP4 — request another version, or open the result in Lab. Generated voice tracks and videos are saved to Lab, where you can find them for the next version.

The result card also offers entry points to Edit video (trim, crop and on-screen text in Video Editor), Dub this video, Publish, and Edit with Aisma. Each opens its own workflow with its own steps.

  • Preview in the Studio player
  • Download the finished video
  • Find voice tracks and videos in Lab
  • Separate editing, dubbing and publishing workflows
Create talking photo

Make Photo Talk FAQ

Questions before you create

Do I need an account?

You sign in to generate — with Google, an emailed sign-in link, or a code. This page opens Video Studio's Talking avatar workflow with Hedra Avatar selected.

What photo works best?

A clear, front-facing portrait with good light and the mouth visible, in JPG, PNG or WebP. The sushi mascot on this page is one existing stylised example.

How long is the video?

Its length follows the speech track you supply or generate. With Hedra Avatar you choose 9:16, 16:9 or 1:1 framing.

Which model should I choose?

Keep Hedra Avatar for a portrait you upload. OmniHuman 1.5 is an alternative that also uses your photo. HeyGen Avatar V renders a fixed presenter, not your image.

Can I make someone else's photo talk?

Only with their permission. Before rendering, you confirm that you have the rights to both the portrait and the voice.

Can I change the words or edit the video afterwards?

Changing the words means a new speech track — upload a different recording or generate a new voice track — and another quoted render. For the existing video, Edit video opens Video Editor for trimming, cropping and on-screen text; captions are created in Audio Studio, and publishing runs through the Social handoff. Each is a separate step.

Ready when you are

Make the first version, then make it yours

Pick a portrait, add the words, and see the quote before each paid generation.

Your portrait · Your recording or an AI voice · A quote before each paid generation