Research

Google Launches Gemini Omni Flash, a Video Model You Edit by Talking

Google's Gemini Omni Flash turns any mix of image, audio, video and text into editable video, rolling out globally today with SynthID watermarks on every clip.

Introducing Gemini Omni
Introducing Gemini OmniAI-generated
By Sophie Lindqvist4 min read

Updated

Why it matters

  • Gemini Omni Flash, the first model in the Omni family, generates video from any combination of image, audio, video and text inputs and is rolling out today to Google AI Plus, Pro and Ultra subscribers globally via the Gemini app and Google Flow.
  • The model supports conversational multi-turn video editing, and every generated video carries Google's imperceptible SynthID watermark, verifiable via the Gemini app, Gemini in Chrome and Google Search.
  • Access is free on YouTube Shorts and the YouTube Create App starting this week, with developer and enterprise API access to follow in the coming weeks.

Google has launched Gemini Omni Flash, a video generation and editing model that accepts any combination of image, audio, video and text as input and produces high-quality video output, the company announced today. The model is rolling out globally to Google AI Plus, Pro and Ultra subscribers through the Gemini app and Google Flow, and at no cost to users on YouTube Shorts and the YouTube Create App starting this week. APIs for developers and enterprise customers follow in the coming weeks.

The launch matters because it moves generative video from one-shot clip generation toward iterative, conversational editing — a capability OpenAI's Sora, Runway and others are also racing to ship. Google is betting that grounding video generation in Gemini's reasoning and world knowledge, plus its distribution through YouTube, gives it an edge in a market where AI-generated content and its provenance have become policy concerns.

What Gemini Omni does

Omni is the first model in a new family. Google frames it as the point "where Gemini's ability to reason meets the ability to create" — the successor step to last year's Nano Banana, which brought Gemini's intelligence to image generation and editing and, according to Google, has since helped millions of people restore old photos, design from sketches and visualize ideas. Image and audio output modalities will follow over time; video comes first.

The editing model works through conversation. "Every instruction builds on the last. Your characters stay consistent, the physics hold up and the scene remembers what came before," Google writes. Users can change specific elements or transform an entire scene — the company's example prompt is "Make the sculpture out of bubbles" — rework the action in footage they've shot, add characters or objects, or adjust environment, angle, style and fine detail across multiple turns without losing the original scene's thread.

Reasoning, not just rendering

Google's central claim is that Omni doesn't merely build photorealistic scenes but reasons about what should happen in them. "It combines an intuitive understanding of physics with Gemini's knowledge of history, science and cultural context, bridging the gap from photorealism to meaningful storytelling," the announcement states.

Concretely, Google says Omni has an improved intuitive grasp of forces like gravity, kinetic energy and fluid dynamics, enabling more physically accurate scenes. The company also demonstrates the model blending knowledge with visuals: generating a rapid-fire alphabet video with 26 unusual items — a capybara for C, a disco globe for D, a lava lamp for L — each with matching lower thirds, at roughly 9 frames per item at 24FPS, ending on a handwritten "THE END" card. Another example: a stop-motion claymation explainer of protein folding from a short prompt.

Any input, one output

Omni turns any reference — image, text, video or audio — into a single cohesive video clip. Users can supply images of characters, scenes or drawings to steer generation, apply styles, motion or effects via references or plain natural language, and Omni blends them together. Audio input is limited at launch: only voice references are supported, with other audio input types promised soon.

The model also supports Avatars, which create a digital version of the user so they can generate videos that look and sound like them. Google is explicit about where it drew the line: editing video to change audio and speech is not shipping yet. "We are still working to test this and better understand how we can bring this capability to users responsibly," the company writes.

Provenance and safeguards

Every video Omni generates carries Google's imperceptible SynthID digital watermark. Users can verify whether a video was made with Gemini Omni through the Gemini app, Gemini in Chrome and Google Search. Google says it has clear policies to protect users from harm and govern use of its AI tools, and points to a separate blog post on expanding its content transparency and verification tools across the web.

The watermarking commitment lands as regulators and platforms push for provenance standards on synthetic media, and as AI video quality makes manual detection increasingly unreliable.

Availability

Gemini Omni Flash is available today to all Google AI Plus, Pro and Ultra subscribers globally via the Gemini app and Google Flow, and free on YouTube Shorts and the YouTube Create App this week. Developer and enterprise API access follows in the coming weeks — the signal to watch for anyone building video tools on top of Google's stack, and the point at which Omni starts competing directly for the workloads currently running through rival video APIs.

Original: support.google.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

136 articles

Related articles

  1. Google ships Gemini Omni 1.1 Flash with 4K output and longer video
  2. Google Ships Nano Banana Pro, Its Gemini 3 Pro Image Model
  3. Google Ships Nano Banana 2 Lite and Opens Gemini Omni Flash to Developers
  4. Google Rolls Out Nano Banana 2 Across Its Product Line
  5. Google Upgrades Veo 3.1 With 4K Upscaling and Native Vertical Video

« Previous articleNext article »