Hedra
  • Developers
  • Studio
  • Enterprise
  • Blog
Log inSign Up
Open Hedra
Your account
    Explore
  • Home
  • Developers
  • Studio
  • Enterprise
  • Blog
    Log inSign Up
    Open Hedra
Google

Veo 3.1

All video models
Video modelGoogle

For unparalleled detail and nuance, perfect for when your vision requires the best possible quality.

OverviewText + image + video to video

Overview

Veo 3.1 is an advanced generative video model by Google DeepMind, building upon Veo 3 with native audio generation for synchronized dialogue and sound effects. It accepts text, image, and video inputs to produce clips up to 4K resolution. The model is especially good for professionals requiring precise narrative control, offering features like seamless scene extension, first-and-last frame guidance, and multi-image "ingredients" to maintain visual consistency.

Pricing

Standard: 720p/1080p 39.29¢/second with audio, 19.64¢/second without; Fast: 720p 14.29¢/second with audio, 11.43¢/second without, 1080p 14.29¢/second with audio, 11.9¢/second without.

Veo 3.1 All Inputs

Text + image + video to video — generates video.

Specifications

Input mode
Text + image + video to video
Accepts
reference image (up to 3), start frame, end frame, source video
Aspect ratios
16:9, 9:16
Resolutions
720p, 1080p
Durations
4s, 6s, 8s
Max duration
8s
Native audio
Optional

Build with this model: Veo 3.1 API on the Hedra Developer Platform.

Best of Veo 3.1

A fashion model stands on a runway wearing a glossy black latex hood and gown paired with a massive, textured silver tinsel jacket. The background features cool blue stage lighting and bright vertical beams. Generated using Veo 3.1 Fast at a vertical 1080x1920 resolution.Futuristic Runway Fashion Show — Veo 3.1 FastA medium shot of a young East Asian man smiling in a grey hoodie. He is standing inside a modern, brightly lit open-office workspace. In the background, a computer monitor displaying a colorful screen is softly blurred. This 1920x1080 video was generated using the Veo 3.1 Fast model.Professional Coder in Modern Office — Veo 3.1 FastA close-up shot of a small black cat with large yellow-green eyes looking directly at the camera. The scene is presented inside an ornate golden picture frame against a dark background. In the background, there is a paper bag and a white receipt paper. This 1920x1080 video was generated using the Veo 3.1 model.Black Cat in Golden Frame — Veo 3.1A wide shot of three musicians in a modern, brightly lit indoor room. On the left, a woman sings into a microphone while playing an acoustic guitar. In the center, a smiling man wearing a cowboy hat plays an electric guitar. On the right, a man in a dark green hoodie watches. Generated at 1920x1080 resolution using the Veo 3.1 Fast model on Hedra.Veo 3.1 Fast: Three Musicians Performing IndoorsA 1920x1080 video still generated by Veo 3.1 Fast showing two men sitting on beige modular couches in a recording studio. On the left, a smiling Black man wears a purple t-shirt. On the right, a smiling man with glasses wears a black suit. Professional microphones stand between them.Two Men on a Podcast Set — Veo 3.1 FastA medium shot of a woman with long brown hair playing an acoustic guitar and singing into a microphone next to a man in a red leather jacket holding a microphone. In the background of the commercial kitchen, a chef prepares food. Generated using Veo 3.1 Fast at 1920x1080 resolution.Musicians Performing in a Kitchen — Veo 3.1 FastA vertical video frame showing a young blonde woman in a cream cable-knit turtleneck sweater waving at the camera. She smiles warmly in a cozy room decorated for Christmas, featuring a decorated tree and stockings. Generated with the Veo 3.1 model at 1080x1920 resolution.Woman Waving in Holiday Setting — Veo 3.1A video still showing two giant parade balloons resembling a middle-aged man and woman floating down a crowded New York City street during a parade. Generated by Veo 3.1 at 1280x720 resolution, the scene depicts spectators lining the streets under historic city buildings.Giant Parade Balloons Floating in New York — Veo 3.1

What is Veo 3.1 best used for?

Veo 3.1 is highly effective for generating realistic, cinematic videos with native synchronized audio. Instead of requiring you to layer sound afterward, the model generates dialogue, ambient noise, and sound effects alongside the video, matching lip movements and on-screen action. It supports 1080p and 4K resolutions in both landscape and portrait formats. Creators frequently use it for narrative shorts and character dialogue where precise audio-visual timing is required.

What is the release history of Veo 3.1?

Google DeepMind officially released Veo 3.1 on October 15, 2025. It is a direct upgrade to Veo 3, which launched in May 2025. The 3.1 update improved audio-visual synchronization and added new creative controls like video extension. Alongside the standard model, Google introduced Veo 3.1 Fast for rapid iteration and a Lite version for lower-cost generation.

How can I get the best results and maintain character consistency?

To maintain character and object consistency across multiple clips, use Veo 3.1's Ingredients to Video feature, which accepts up to three reference images to guide the output. When prompting for audio, explicitly describe the sounds you want (e.g., "wings flapping, birdsong") alongside the visual action. For a detailed breakdown on structuring text prompts for cinematic realism and dialogue, read Google's official Veo prompt guide.

Similar models

Seedance 2.5ByteDanceVeo 3.1 FastGoogleSeedance 2.0ByteDanceKling O1KlingKling O3 ProKlingHedra OmniaHedra

Prompt tips

  • Write explicit audio cues: Include dialogue in quotes or describe specific sound effects (e.g., "whispering excitedly" or "torchlight flickering with a low crackle") to trigger the native audio engine.
  • Pre-generate reference assets: Use an image model like Imagen 4 or Nano Banana Pro to create your base characters and style frames, then use Veo 3.1's image blending to animate them.
  • Anchor your camera motion: Provide both a starting image and an ending image to force the model to calculate the specific camera movement and action required to bridge the two frames.
  • Specify aspect ratio: Explicitly request 16:9 for landscape or 9:16 for portrait outputs in your configuration, as the model natively supports both without requiring post-generation cropping.

What Will You Create?

Open Creative StudioGet the API
Hedra
Hedra

Product

Developer PlatformStudioEnterpriseSovereignPricing

Resources

Agent documentationDeveloper documentationBlogUse CasesModelsFeedbackChangelogStatus

Company

AboutCareersContactAffiliatesSupportAlternatives

Legal

Privacy PolicyTerms of useAcceptable useCookie PolicyBiometric data policy
LinkedinInstagramDiscord
support@hedra.comHedra 2026 — All rights reserved