Hedra
  • Developers
  • Inference
  • Studio
  • Enterprise
Log inSign Up
Open Hedra
Your account
    Explore
  • Home
  • Developers
  • Inference
  • Studio
  • Enterprise
    Log inSign Up
    Open Hedra

Turn any photo into a talking video with AI.

Hand it a face and a recording and the face performs the recording — mouth, jaw, brows, breath — for as long as the audio runs. A portrait, an illustration, a mascot, a painting: all of them count as a photo here.

Start building nowMake one now — no code
Drawing + audio in
Widescreen illustration of a woman in round glasses and a mustard turtleneck talking on a teal background

"Honestly? I have been waiting a very long time for someone to ask me that."

PNG · 2560 × 1440AUDIO · 0:08
→

If it has a face, it can talk.

Illustrated woman with a dark bob and round glasses in a mustard turtleneck, speaking against a teal background
Close-up of a claymation man with bushy eyebrows and a red knit scarf, mouth open mid-sentence
19th-century oil portrait of a woman with dark upswept hair and a black dress with a lace collar
GPT IMAGE 2 · 9:16 · THREE KINDS OF FACE

It doesn’t have to be a photograph.

A drawn character, a clay puppet on a miniature set, a woman painted in oils a century before film existed — the model wants a face, and it does not much mind what the face is made of. Which means the thing that speaks can be your mascot, your illustration style, or a portrait nobody could have filmed.

“Honestly? I have been waiting a very long time for someone to ask me that.”HEDRA CHARACTER 3 · 0:08 · SAME RECORDING

The face follows the audio, not a prompt.

This is the part that makes it a performance rather than a puppet. The model reads the recording — where the stress lands, where the breath goes, where the sentence turns — and the face answers it. Send the same eight seconds to a drawing and to a painting and you get the same reading in two hands that share nothing.

Claymation man in a red knit scarf talking in front of rocky ruins under a pink sky
HEDRA CHARACTER 3 · UP TO 10:00

As long as the recording runs.

Most video models hand back a few seconds. An audio-driven one has a different contract: leave the duration out and the take runs exactly as long as the audio you sent, up to ten minutes in a single generation. A walkthrough, a lesson, a full read of a script — not a clip you then have to stitch.

Models that animate a face from audio.

HedraHedraHedra Character 3Hedra Avatar
ByteDanceByteDanceOmnihuman 1.5
KlingKlingKling AI Avatar v2
VEEDVEEDVEED Fabric 1.0
ElevenLabsElevenLabsMultilingual V2V3Flash Multilingual V2Voice Clone
MiniMaxMiniMaxSpeech 2.5 HDSpeech 2.5 Turbo
See all models →

Built for visual inference.

Serving video is a different problem from serving text — a single request can saturate a GPU, and none of the tricks that made language models cheap apply. Hedra's engine was built for exactly that workload.

Whichever model you choose — ours or anyone else's — it runs on the same infrastructure, behind one key, priced per second of output. Which for an audio-driven take means the price is set by the recording: a ten-minute read costs ten minutes.

Generate with API or Agent

Build with the API

One key, every model, priced per second of output. Send a face and an audio URL and get a take back.

Start building now

Make one now.

No code — drop in the picture, add the audio, export the take.

Create with Agent

FAQs

Anything with a clear, legible face pointed roughly at the viewer. Photographs, illustrations, cartoons, 3D characters, stop-motion puppets, paintings — every face on this page is a different kind of picture, and all of them were driven by the same recording.

You supply it: a file you recorded, a track you already own, a voice you cloned, or speech generated from a script. The face follows whatever waveform arrives — that is the whole mechanism.

AI voice cloning →

Leave the duration out and the take runs exactly as long as the audio. Hedra Character 3 goes to ten minutes in a single generation; other models cap shorter, and each model page states its limit.

Video models →

Yes — the model is reading the audio, not the words, so it performs whatever language is on the recording. Swap the track and the same picture speaks the next market without being regenerated from scratch.

Mostly in what you bring. A talking avatar is the whole character — often generated, often scripted. A talking photo starts from a specific picture you already have and animates that, faithfully, to audio you already have.

AI talking avatar →

Commercial use depends on the plan; the pricing page sets out what each covers. On likeness: send only pictures and voices you have the rights to animate — that applies to a portrait of a person as much as to artwork someone else drew.

See pricing →

No — Hedra Studio runs the same models on the same engine with nothing to build. Drop in the picture, add the audio, export the take.

More you can do with Hedra

Widescreen crop of a 19th-century oil portrait of a woman with an updo and a black lace-collared dress
Build the character, not just the takeAI TALKING AVATAR →
Claymation man in a red knit scarf talking in front of rocky ruins under a pink sky
Move the picture instead of speaking itIMAGE TO VIDEO AI →
hedra keys create
→ sk-hedra-••••••••
Wire it into your own productDEVELOPER PLATFORM →

Give the picture a voice.

Start building nowMAKE A TALKING PHOTO NOW — NO CODE →
Hedra
Hedra

Product

DevelopersStudioEnterpriseInferencePricing

Resources

Agent documentationDeveloper documentationBlogUse CasesModelsFeedbackChangelogStatus

Company

AboutCareersContactBuilder programSupportAlternatives

Legal

Privacy PolicyTerms of useAcceptable useCookie PolicyBiometric data policy
LinkedinInstagramDiscord
support@hedra.comHedra 2026 — All rights reserved