VEED
AI Video Editor for Fast, Pro-Level Video Creation
DeepMind's video generator produces dialogue, effects and ambience along with the picture. Reachable through the Gemini app, the Flow studio, or an API billed by the second from five cents.

Veo is Google DeepMind's video generation model. You describe a shot and it produces one — with the camera move you asked for, physics that mostly behave, and, crucially, sound. Dialogue, footsteps, room tone and score arrive with the picture rather than being added afterwards.
That last part is the whole argument. Silent generated video is a clip; video with matched audio is something you can put in a timeline. Veo was the model that made the difference obvious, and every serious competitor has been chasing the same capability since.
It also matters more now than it did a year ago, because the most prominent alternative is gone. OpenAI shut Sora down in April 2026, which leaves Veo and a handful of others carrying the category. That is worth remembering when you plan a workflow around any of them.
The simplest route: describe a video in the assistant you may already be paying for. Availability depends on your Google AI subscription tier, your platform and your region, and it is restricted to adults. Fine for making something; not where you would do sustained work.
Flow is Google's creative studio built around these models, and it is where the actual filmmaking happens — scene building, references, iteration, editing. Worth knowing that Veo is now one model among several inside it: there is a separate video model for reference-based creation and conversational editing, and a distinct image model for stills and precise edits.
If you are producing rather than experimenting, this is the surface to learn. A generation tool without an editing environment around it produces clips you then have to assemble somewhere else.
For anything programmatic — a product feature, a pipeline, bulk generation — Veo is available through Google's developer APIs and billed by the second. No subscription, no seats, and no free allowance.
API pricing is per second of finished video, and the spread between tiers is wide enough to design around.
| Tier | 720p | 1080p | 4K |
|---|---|---|---|
| Standard | $0.40/s | $0.40/s | $0.60/s |
| Fast | $0.10/s | $0.12/s | $0.30/s |
| Lite | $0.05/s | $0.08/s | not supported |
Audio is included at every tier rather than sold separately, which is the right way round and not universal. There is no free tier on the API — evaluation costs money from the first clip.
Do the arithmetic before you plan a shoot. An eight-second clip at standard quality costs about three dollars and twenty cents; the same clip on the lite tier costs forty cents. Since most generated video is discarded, the sensible pattern is obvious: explore on lite, render the keeper on standard. A team that ignores this will spend eight times more than necessary to arrive at the same shot.
One decent piece of policy: you are only charged when a video is successfully generated. Failed generations do not bill.
Consumer access through the Gemini app and Flow comes with Google's AI subscriptions instead, which are priced by region and vary substantially between countries. Check the price where you live rather than converting from someone else's article.
The difference between people who get footage they keep and people who get an expensive slideshow is almost entirely in how the prompt is written. Veo responds to the vocabulary of filmmaking, not to adjectives.
Name the shot size and the camera move — a medium shot that slowly pushes in behaves very differently from «a cinematic scene». Describe the light, the lens character and the time of day. Say what the subject is doing rather than what they are like. And write the sound explicitly: ambience, the specific effects you want, whether there is music and what kind. Audio is generated from the prompt too, and leaving it unmentioned means leaving it to chance.
Dialogue deserves its own care. Write the actual line in quotation marks and keep it short — one or two sentences per shot. Long speeches are where lip sync and delivery come apart, and no amount of rerolling fixes a take that was too ambitious to begin with.
Finally, plan in shots rather than scenes. Generating eight seconds you can cut into is more useful, and considerably cheaper, than generating thirty seconds you have to accept whole.
Previsualisation, establishing shots, inserts, and the scenes that would otherwise need a location and a crew. Native audio means a rough cut can be assembled and judged rather than imagined, which is where most of the time actually goes.
Concept films for a pitch, variants for testing, social cuts at a dozen aspect ratios. The economics against a shoot are not close, and for the concepting stage nobody expects a finished commercial anyway.
Per-second billing with no seats or minimums fits a product feature cleanly. Route casual use to the cheap tier and reserve the expensive one for output people will keep.
If your workflow depended on a model that no longer exists, Veo is the closest replacement for what that model was originally admired for: long, coherent shots with generated sound. It will not reproduce the social features, and nothing else will either.
Not through the API, which has no free allowance and bills from the first second. Access through the Gemini app and Flow comes with paid Google AI subscriptions, whose price and availability depend on your country.
Yes — dialogue, effects and ambience are produced together with the picture, and audio is included in the price at every tier rather than charged separately.
An eight-second clip runs from about forty cents on the lite tier at 720p to roughly three dollars twenty at standard quality, and more at 4K. You are only charged for generations that succeed.
Three places: the Gemini assistant for casual generation, Google Flow for real production work with editing around it, and Google's developer APIs for anything programmatic.
For long coherent shots with generated audio, Veo is the closest match to what Sora was praised for. Runway remains stronger as a complete production environment, and the cheaper generators are better for high-volume iteration.
Commercial use is generally permitted, but the terms differ between the consumer subscriptions and the developer APIs. Read the terms attached to the specific route you are using before publishing anything paid for.
Veo is the strongest generally available video model with sound, and Google has put it everywhere a person might want it — in the assistant, in a proper creative studio, and behind an API priced by the second. With its most famous competitor shut down, it is now the default answer for generated video with audio.
The one discipline that matters is tier selection. Explore cheaply, render expensively, and never iterate at the top price. Do that and the costs are reasonable; ignore it and a week of experimentation will cost more than the footage was worth.