Hedra
  • Developers
  • Inference
  • Studio
  • Enterprise
Log inSign Up
Open Hedra
Your account
    Explore
  • Home
  • Developers
  • Inference
  • Studio
  • Enterprise
    Log inSign Up
    Open Hedra

Clone a voice with AI, then write what it says next.

One recording in, and the voice becomes something you can type in. New sentences it never said, in languages the speaker does not have — and a face to say them, in the same request chain.

Start building nowClone a voice — no code
Recording in

Eight seconds of her talking, then the same voice saying a sentence written afterwards. Turn the sound on — silently, these two prove nothing.

Compare Her recording and Her clone

Clone only a voice you have permission to clone. Consent is not a formality here — it is the whole difference between this and impersonation, and it is the one thing the model cannot check for you.

ElevenLabs V3 · cloned voice · words written after

A recording is fixed. A voice is not.

Esto no es una grabación. Escribí esta frase hace diez minutos.ELEVENLABS MULTILINGUAL V2 · SAME VOICE ID

The same voice, in a language she does not speak.

This is the identical cloned voice, reading the identical sentence in Spanish. Nothing was re-recorded and nothing was re-cloned — the voice went to a second model with a language set on the request, and a different set of words came back in the same throat. Whether the Spanish convinces a Spanish speaker is a question for a Spanish speaker; what is verifiable here is that the voice did not change when the language did.

ONE CLIP IN · ONE VOICE ID OUT

Clone once. Then it is just a parameter.

The clone is made from this one clip and then referenced by id, so a second language or a tenth script costs nothing extra to set up — the voice is not re-derived each time. It travels between models from the same maker: one id spoke through two different ElevenLabs models here without being cloned again. It does not travel between makers, which the API will tell you plainly if you try. And the clone step itself is not what you are billed for; speaking is, by the character.

Every voice, and everything that can speak with one.

ElevenLabsElevenLabsVoice CloneV3Multilingual V2Flash Multilingual V2Flash V2Audio Isolation
MiniMaxMiniMaxSpeech 2.5 HDSpeech 2.5 Turbo
HedraHedraCharacter 3Avatar
See all models →

Built for visual inference.

Serving video is a different problem from serving text — a single request can saturate a GPU, and none of the tricks that made language models cheap apply. Hedra's engine was built for exactly that workload.

Whichever model you choose — ours or anyone else’s — it runs on the same infrastructure, behind one key. Which is what made this page possible in one chain: a clip generated by one vendor, its voice cloned by a second, the new words spoken by a third model, and a face driven by a fourth. No file ever left the platform to get to the next step.

Generate with API or Agent

Build with the API

One key. Clone once, keep the voice id, and send scripts at it for as long as you need.

Start building now

Clone a voice now.

No code — drop in a recording you have the right to use, type a line, and hear it back.

Create with Agent

FAQs

Your own, or one you have explicit permission to use. Cloning a voice is the one capability here where the technology is not the hard part — the permission is. A public figure’s voice, a colleague’s recording off a call, a sample lifted from a podcast: all technically possible, none of them yours to use. Get it in writing before you clone, not after.

The page was built from a single eight-second clip, which is enough to be recognisable. Longer and cleaner is better — one speaker, no music, no crosstalk — and if the only recording you have is noisy, isolating the speech first is a separate step you can run on the same key.

Audio models →

Yes, with the multilingual models — the same cloned voice id, a language on the request, and a different script. Around forty languages between the ElevenLabs and MiniMax tiers. What it cannot do is give you an accent you never had: a cloned English speaker reading Spanish sounds like an English speaker reading Spanish, which is either exactly what you want or exactly what you do not.

Within a maker, yes — one clone spoke through two different ElevenLabs models on this page with no re-cloning. Across makers, no: an ElevenLabs voice cannot be handed to a MiniMax model, and the API refuses it with that reason rather than degrading quietly. Worth knowing before you standardise on a voice you may want to move later.

Pass the speech to an avatar model along with a photograph, and the mouth is driven by the audio. Both clips further up this page were made that way — one still from the original recording, two different pieces of cloned speech. Up to ten minutes in a single request.

Talking photo →

Creating the clone is not itself a billed generation; speaking with it is, by the character, at a rate that depends on which speech model you choose. Putting it on a face is billed by the second of video. Each model page carries its own rate.

See model pricing →

No — Hedra Studio runs the same models on the same engine with nothing to build. Drop in a recording, type a line, and hear it back before you commit to anything.

More you can do with Hedra

A still of the cloned voice being spoken by a face
Put the voice on a photographTALKING PHOTO AI →
A still of the same face speaking Spanish
Match an existing take, word for wordAI LIP SYNC →
hedra keys create
→ sk-hedra-••••••••
Script a hundred of themDEVELOPER PLATFORM →

Your voice. Your words. Later.

Start building nowCLONE A VOICE NOW — NO CODE →
Hedra
Hedra

Product

DevelopersStudioEnterpriseInferencePricing

Resources

Agent documentationDeveloper documentationBlogUse CasesModelsFeedbackChangelogStatus

Company

AboutCareersContactBuilder programSupportAlternatives

Legal

Privacy PolicyTerms of useAcceptable useCookie PolicyBiometric data policy
LinkedinInstagramDiscord
support@hedra.comHedra 2026 — All rights reserved