Google Launches Gemini Omni Flash, a Video Model You Edit by Talking
Google's Gemini Omni Flash turns any mix of image, audio, video and text into editable video, rolling out globally today with SynthID watermarks on every clip.

Updated
Why it matters
- Gemini Omni Flash, the first model in the Omni family, generates video from any combination of image, audio, video and text inputs and is rolling out today to Google AI Plus, Pro and Ultra subscribers globally via the Gemini app and Google Flow.
- The model supports conversational multi-turn video editing, and every generated video carries Google's imperceptible SynthID watermark, verifiable via the Gemini app, Gemini in Chrome and Google Search.
- Access is free on YouTube Shorts and the YouTube Create App starting this week, with developer and enterprise API access to follow in the coming weeks.
Google has launched Gemini Omni Flash, a video generation and editing model that accepts any combination of image, audio, video and text as input and produces high-quality video output, the company announced today. The model is rolling out globally to Google AI Plus, Pro and Ultra subscribers through the Gemini app and Google Flow, and at no cost to users on YouTube Shorts and the YouTube Create App starting this week. APIs for developers and enterprise customers follow in the coming weeks.
The launch matters because it moves generative video from one-shot clip generation toward iterative, conversational editing — a capability OpenAI's Sora, Runway and others are also racing to ship. Google is betting that grounding video generation in Gemini's reasoning and world knowledge, plus its distribution through YouTube, gives it an edge in a market where AI-generated content and its provenance have become policy concerns.
What Gemini Omni does
Omni is the first model in a new family. Google frames it as the point "where Gemini's ability to reason meets the ability to create" — the successor step to last year's Nano Banana, which brought Gemini's intelligence to image generation and editing and, according to Google, has since helped millions of people restore old photos, design from sketches and visualize ideas. Image and audio output modalities will follow over time; video comes first.
The editing model works through conversation. "Every instruction builds on the last. Your characters stay consistent, the physics hold up and the scene remembers what came before," Google writes. Users can change specific elements or transform an entire scene — the company's example prompt is "Make the sculpture out of bubbles" — rework the action in footage they've shot, add characters or objects, or adjust environment, angle, style and fine detail across multiple turns without losing the original scene's thread.
Reasoning, not just rendering
Google's central claim is that Omni doesn't merely build photorealistic scenes but reasons about what should happen in them. "It combines an intuitive understanding of physics with Gemini's knowledge of history, science and cultural context, bridging the gap from photorealism to meaningful storytelling," the announcement states.
Concretely, Google says Omni has an improved intuitive grasp of forces like gravity, kinetic energy and fluid dynamics, enabling more physically accurate scenes. The company also demonstrates the model blending knowledge with visuals: generating a rapid-fire alphabet video with 26 unusual items — a capybara for C, a disco globe for D, a lava lamp for L — each with matching lower thirds, at roughly 9 frames per item at 24FPS, ending on a handwritten "THE END" card. Another example: a stop-motion claymation explainer of protein folding from a short prompt.
Any input, one output
Omni turns any reference — image, text, video or audio — into a single cohesive video clip. Users can supply images of characters, scenes or drawings to steer generation, apply styles, motion or effects via references or plain natural language, and Omni blends them together. Audio input is limited at launch: only voice references are supported, with other audio input types promised soon.
The model also supports Avatars, which create a digital version of the user so they can generate videos that look and sound like them. Google is explicit about where it drew the line: editing video to change audio and speech is not shipping yet. "We are still working to test this and better understand how we can bring this capability to users responsibly," the company writes.
Provenance and safeguards
Every video Omni generates carries Google's imperceptible SynthID digital watermark. Users can verify whether a video was made with Gemini Omni through the Gemini app, Gemini in Chrome and Google Search. Google says it has clear policies to protect users from harm and govern use of its AI tools, and points to a separate blog post on expanding its content transparency and verification tools across the web.
The watermarking commitment lands as regulators and platforms push for provenance standards on synthetic media, and as AI video quality makes manual detection increasingly unreliable.
Availability
Gemini Omni Flash is available today to all Google AI Plus, Pro and Ultra subscribers globally via the Gemini app and Google Flow, and free on YouTube Shorts and the YouTube Create App this week. Developer and enterprise API access follows in the coming weeks — the signal to watch for anyone building video tools on top of Google's stack, and the point at which Omni starts competing directly for the workloads currently running through rival video APIs.
Original: support.google.com
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
136 articles
Related articles
- Google ships Gemini Omni 1.1 Flash with 4K output and longer video
- Google Ships Nano Banana Pro, Its Gemini 3 Pro Image Model
- Google Ships Nano Banana 2 Lite and Opens Gemini Omni Flash to Developers
- Google Rolls Out Nano Banana 2 Across Its Product Line
- Google Upgrades Veo 3.1 With 4K Upscaling and Native Vertical Video