Clone a voice with AI, then write what it says next.
One recording in, and the voice becomes something you can type in. New sentences it never said, in languages the speaker does not have — and a face to say them, in the same request chain.
Eight seconds of her talking, then the same voice saying a sentence written afterwards. Turn the sound on — silently, these two prove nothing.
Clone only a voice you have permission to clone. Consent is not a formality here — it is the whole difference between this and impersonation, and it is the one thing the model cannot check for you.
A recording is fixed. A voice is not.
The same voice, in a language she does not speak.
This is the identical cloned voice, reading the identical sentence in Spanish. Nothing was re-recorded and nothing was re-cloned — the voice went to a second model with a language set on the request, and a different set of words came back in the same throat. Whether the Spanish convinces a Spanish speaker is a question for a Spanish speaker; what is verifiable here is that the voice did not change when the language did.
Clone once. Then it is just a parameter.
The clone is made from this one clip and then referenced by id, so a second language or a tenth script costs nothing extra to set up — the voice is not re-derived each time. It travels between models from the same maker: one id spoke through two different ElevenLabs models here without being cloned again. It does not travel between makers, which the API will tell you plainly if you try. And the clone step itself is not what you are billed for; speaking is, by the character.
Every voice, and everything that can speak with one.
Built for visual inference.
Serving video is a different problem from serving text — a single request can saturate a GPU, and none of the tricks that made language models cheap apply. Hedra's engine was built for exactly that workload.
Whichever model you choose — ours or anyone else’s — it runs on the same infrastructure, behind one key. Which is what made this page possible in one chain: a clip generated by one vendor, its voice cloned by a second, the new words spoken by a third model, and a face driven by a fourth. No file ever left the platform to get to the next step.
Generate with API or Agent
Build with the API
One key. Clone once, keep the voice id, and send scripts at it for as long as you need.
Clone a voice now.
No code — drop in a recording you have the right to use, type a line, and hear it back.
FAQs
Your own, or one you have explicit permission to use. Cloning a voice is the one capability here where the technology is not the hard part — the permission is. A public figure’s voice, a colleague’s recording off a call, a sample lifted from a podcast: all technically possible, none of them yours to use. Get it in writing before you clone, not after.