Gemini Omni Flash Turns Video Editing Into a Conversation
Google's new public-preview model accepts text, image and video inputs, produces 10-second clips, and lets creators refine the result with natural-language follow-ups.
The next useful leap in generative video may not be a prettier first draft. It may be the ability to say, “Keep that shot, change the lighting, and make the movement less frantic,” without starting again.
That is the promise behind Gemini Omni Flash, which Google opened to developers on June 30 through the Gemini API and Google AI Studio. The public-preview model accepts combinations of text, images and video, then lets users generate or refine a clip through natural-language instructions. It is also available in the Gemini app and Google Flow.
What Google released
Gemini Omni Flash is identified in the API as `gemini-omni-flash-preview`. Google prices output at $0.10 per second—the same price it lists for Veo 3.1 Fast—and currently limits generations to 10 seconds.
The important shift is the editing loop. A creator can supply visual references, generate a scene, and then continue the conversation instead of rebuilding the prompt from scratch. When used through Google's Interactions API, a session can retain enough context for as many as three sequential edits.
Google has also applied SynthID watermarking to output from Omni Flash. That adds a provenance layer, although watermarking does not by itself settle the broader questions around consent, disclosure or the use of synthetic footage.
Why it matters
Video tools have long made users choose between speed and control. Conversational editing tries to close that gap: the model makes an initial clip, while the user directs the revision in ordinary language. If the workflow proves reliable, it could make generative video feel less like prompting a slot machine and more like working with an unusually fast editor.
There is a practical pairing here, too. Google suggests using Nano Banana 2 Lite to create a reference image and then passing it to Omni Flash for animation. That connects cheap image iteration with a more expensive video step, giving developers a clearer path to build product demos, advertising tools and rapid storyboards.
The catches
This is a preview, and Google's own limitations are substantial. Audio-reference uploads and scene extension are not supported in the API. Video references of up to three seconds are accepted by the schema but are not processed correctly, while character consistency can deteriorate during scene changes and camera pans. Longer clips are promised, not shipped.
Those caveats matter more than a vendor benchmark. The real test is whether creators can make several controlled edits without the scene, subject or intent drifting along the way.
Sources
Update note: Published and source-checked on 2026-07-24. Next checkpoint: general availability, longer output, and fixes for reference-video processing.
Sources
Drafted with AI assistance from verified source material and reviewed for factual accuracy, attribution, clarity, and label integrity.