Create with the agent
1
Prepare the inputs
Use a clear character image and clean speech or music. A front-facing image with
an unobstructed face is the most reliable starting point.
2
Start from a presenter or UGC Skill
On the home screen, select presenter or ugc, or describe the intended
performance in a new Space.
3
Add the references
Attach the character image and audio. Explain who should speak, the framing, the
tone, and any movement or background constraints.
4
Review and refine
Check lip sync, identity, framing, and gesture. Ask for a bounded change while
preserving the references that already work.
Use the manual Avatar mode
Open Manual tools → Video, then choose Avatar.- Select a model from the filtered catalog.
- Add the image and audio references required by that model.
- Describe the desired performance.
- Choose the available aspect ratio, resolution, duration, and batch size.
- Select Generate.
Better inputs produce better performances
- Use a sharp, well-lit face at a useful scale.
- Avoid heavy occlusion over the mouth and eyes.
- Use clean audio without clipping or excessive background noise.
- Start with a short test before generating a long performance.
- For multiple people, make the intended speaker unambiguous in the prompt and use speaker controls when the selected model exposes them.
Avatar models accept different input combinations and durations. The live model
selector shows the requirements for the model you choose.