Veo

DeepMind's video generator produces dialogue, effects and ambience along with the picture. Reachable through the Gemini app, the Flow studio, or an API billed by the second from five cents.

Go to AI
Veo cover

What Is Veo?

Veo is Google DeepMind's video generation model. You describe a shot and it produces one — with the camera move you asked for, physics that mostly behave, and, crucially, sound. Dialogue, footsteps, room tone and score arrive with the picture rather than being added afterwards.

That last part is the whole argument. Silent generated video is a clip; video with matched audio is something you can put in a timeline. Veo was the model that made the difference obvious, and every serious competitor has been chasing the same capability since.

It also matters more now than it did a year ago, because the most prominent alternative is gone. OpenAI shut Sora down in April 2026, which leaves Veo and a handful of others carrying the category. That is worth remembering when you plan a workflow around any of them.

Three Ways to Reach It

The Gemini App

The simplest route: describe a video in the assistant you may already be paying for. Availability depends on your Google AI subscription tier, your platform and your region, and it is restricted to adults. Fine for making something; not where you would do sustained work.

Google Flow

Flow is Google's creative studio built around these models, and it is where the actual filmmaking happens — scene building, references, iteration, editing. Worth knowing that Veo is now one model among several inside it: there is a separate video model for reference-based creation and conversational editing, and a distinct image model for stills and precise edits.

If you are producing rather than experimenting, this is the surface to learn. A generation tool without an editing environment around it produces clips you then have to assemble somewhere else.

The API

For anything programmatic — a product feature, a pipeline, bulk generation — Veo is available through Google's developer APIs and billed by the second. No subscription, no seats, and no free allowance.

What It Costs

API pricing is per second of finished video, and the spread between tiers is wide enough to design around.

Tier720p1080p4K
Standard$0.40/s$0.40/s$0.60/s
Fast$0.10/s$0.12/s$0.30/s
Lite$0.05/s$0.08/snot supported

Audio is included at every tier rather than sold separately, which is the right way round and not universal. There is no free tier on the API — evaluation costs money from the first clip.

Do the arithmetic before you plan a shoot. An eight-second clip at standard quality costs about three dollars and twenty cents; the same clip on the lite tier costs forty cents. Since most generated video is discarded, the sensible pattern is obvious: explore on lite, render the keeper on standard. A team that ignores this will spend eight times more than necessary to arrive at the same shot.

One decent piece of policy: you are only charged when a video is successfully generated. Failed generations do not bill.

Consumer access through the Gemini app and Flow comes with Google's AI subscriptions instead, which are priced by region and vary substantially between countries. Check the price where you live rather than converting from someone else's article.

Getting Usable Shots Out of It

The difference between people who get footage they keep and people who get an expensive slideshow is almost entirely in how the prompt is written. Veo responds to the vocabulary of filmmaking, not to adjectives.

Name the shot size and the camera move — a medium shot that slowly pushes in behaves very differently from «a cinematic scene». Describe the light, the lens character and the time of day. Say what the subject is doing rather than what they are like. And write the sound explicitly: ambience, the specific effects you want, whether there is music and what kind. Audio is generated from the prompt too, and leaving it unmentioned means leaving it to chance.

Dialogue deserves its own care. Write the actual line in quotation marks and keep it short — one or two sentences per shot. Long speeches are where lip sync and delivery come apart, and no amount of rerolling fixes a take that was too ambitious to begin with.

Finally, plan in shots rather than scenes. Generating eight seconds you can cut into is more useful, and considerably cheaper, than generating thirty seconds you have to accept whole.

Who Gets the Most From It

Filmmakers and Video Teams

Previsualisation, establishing shots, inserts, and the scenes that would otherwise need a location and a crew. Native audio means a rough cut can be assembled and judged rather than imagined, which is where most of the time actually goes.

Advertising and Marketing

Concept films for a pitch, variants for testing, social cuts at a dozen aspect ratios. The economics against a shoot are not close, and for the concepting stage nobody expects a finished commercial anyway.

Product Teams Adding Video Features

Per-second billing with no seats or minimums fits a product feature cleanly. Route casual use to the cheap tier and reserve the expensive one for output people will keep.

Anyone Displaced by Sora's Closure

If your workflow depended on a model that no longer exists, Veo is the closest replacement for what that model was originally admired for: long, coherent shots with generated sound. It will not reproduce the social features, and nothing else will either.

What to Watch Out For

  • No free tier on the API. Every experiment costs money from the first second.
  • The gap between the cheapest and dearest tier is eightfold. Iterating at standard quality is the most common way to waste a budget here.
  • Consumer access varies by subscription tier, by platform and by region, and is restricted to adults. What a colleague can do may not be what you can do.
  • Previous Veo generations are deprecated and were scheduled for shutdown in mid-2026. Anything built against a specific version needs a migration plan.
  • Subscription prices for Gemini and Flow are regional and differ substantially between countries.
  • Generated dialogue is the weakest part. Lip sync and delivery hold up in short takes and betray themselves in long ones.
  • Veo is one model inside Flow rather than the whole product, so guidance written about «Veo» may actually describe a different model in the same studio.
  • Vendor risk in this category is not hypothetical. Keep your source material and your exports.

Frequently Asked Questions

Is Veo free?

Not through the API, which has no free allowance and bills from the first second. Access through the Gemini app and Flow comes with paid Google AI subscriptions, whose price and availability depend on your country.

Does it generate sound?

Yes — dialogue, effects and ambience are produced together with the picture, and audio is included in the price at every tier rather than charged separately.

How much does a clip cost?

An eight-second clip runs from about forty cents on the lite tier at 720p to roughly three dollars twenty at standard quality, and more at 4K. You are only charged for generations that succeed.

Where do I actually use it?

Three places: the Gemini assistant for casual generation, Google Flow for real production work with editing around it, and Google's developer APIs for anything programmatic.

What is the best replacement for Sora?

For long coherent shots with generated audio, Veo is the closest match to what Sora was praised for. Runway remains stronger as a complete production environment, and the cheaper generators are better for high-volume iteration.

Can I use the output commercially?

Commercial use is generally permitted, but the terms differ between the consumer subscriptions and the developer APIs. Read the terms attached to the specific route you are using before publishing anything paid for.

The Bottom Line

Veo is the strongest generally available video model with sound, and Google has put it everywhere a person might want it — in the assistant, in a proper creative studio, and behind an API priced by the second. With its most famous competitor shut down, it is now the default answer for generated video with audio.

The one discipline that matters is tier selection. Explore cheaply, render expensively, and never iterate at the top price. Do that and the costs are reasonable; ignore it and a week of experimentation will cost more than the footage was worth.

Alternative Tools