Gemini Omni Flash 1.1: complete guide, prompts, and new features

Gemini Omni Flash 1.1: complete guide, prompts, and new features

The complete Gemini Omni Flash 1.1 guide: the five tasks, the prompt formula, reference and frame tags, edits you type, 40-second extensions, and audio direction, with examples.

Gemini Omni Flash 1.1 features and capabilities

Google DeepMind's multimodal video model, updated in August 2026. It makes 3 to 10 second clips with native audio, then keeps working on them: a follow-up prompt edits the same take, and extensions grow it to 40 seconds.

FeatureWhat it doesBest for
Conversational editingA follow-up prompt changes the take; anything unmentioned stays putFixing a shot without re-rolling
Video extensionContinues a clip 10s at a time, to 40s totalScenes longer than one generation
Reference input10 images and 3 clips steer one generation, each tagged to a roleRecurring characters, products, styles
First and last frameInterpolates between two stills; the same still twice loopsTransitions, loops, storyboard pairs
Native audioLip-synced speech, effects, ambience, and music with the pictureTalking scenes, sound-led shots
Resolution ladder360p to 4K; 1080p and 4K are upscalesCheap drafts, full-size delivery
On-screen textRenders the English text you specify, down to background signsTitles, lower thirds, captions

Gemini Omni Flash 1.1 use cases

Edit chains on one take

Relight the room, swap the jacket, rewrite the sign. Each instruction lands on the same shot while the rest of the frame holds.

The same living room split in two, flat midday light on the left and golden dusk on the right

Scenes built by extension

Open on a ten-second beat, then grow it. The model reads your final seconds and continues the motion, cast, and score to 40 seconds.

Four-panel filmstrip of a striped hot-air balloon drifting across an alpine valley at dawn

Sketch to footage

Say to use the drawing as a guide only. The model borrows the silhouette and path, then rebuilds texture, light, and depth.

A pencil sketch of a leaping fox beside the same pose rendered as a photoreal fox in a forest

Talking spokespeople

Write the line and the character delivers it lip-synced, in a voice that stays theirs through every later edit.

A market vendor speaking warmly over crates of oranges in late afternoon sun

Timed sequences

A timecode block turns one prompt into a cut-by-cut sequence: rapid-fire product frames, or a beat every two seconds.

Three panels of the same woman in a red scarf walking, turning back, then running

Perfect loops

Pass the same image as first and last frame and the clip returns to where it began, ready to run on a product page.

A paper plane flying a loop traced by a closed trail of light in a dim workshop

Step 1: pick the task

Everything the model does maps to one of five tasks. Name it when intent could be misread; otherwise it is inferred from your prompt and media.

TaskWhat it doesReach for it when
Text to videoGenerates a clip from a written briefStarting from nothing
Image to videoAnimates a still, or interpolates two framesYou have the frame, you want the motion
Reference to videoBuilds a shot from tagged images and clipsA character or style must carry over
EditChanges an existing videoThe take is right, one element is wrong
ExtendAppends a continuationThe take is right and too short

Step 2: write the prompt

A prompt is a short shot brief, in this order. Aspect ratio, resolution, and duration are settings, not prompt text.

ElementWhat it coversRequired
Subject and actionWho is in frame, and what they doYes
SceneLocation, time of day, weather, backgroundOptional
Camera and lightShot size, movement, lens, light sourceOptional
AudioDialogue, ambience, effects, music, or silenceStrongly recommended
Shot structureOne continuous shot, or cuts and their timingOptional
Continuous, unbroken handheld shot of a fluffy tabby cat on a sunny windowsill,
looking out into a leafy garden. Its tail twitches slowly; its ears rotate toward
ambient noises. Sound design: gentle breeze, distant bird chirps. No dialogue.

The five prompt levers

LeverWrite thisWhy it matters
Shot structureSingle continuous shot, no scene cuts. Or: every 2s cut to a new locationUnprompted, the model tells a story across several shots
AudioNo music, just room tone. Or a spoken line: she says, third one todaySilence must be asked for, or you get invented music
TimingAfter 3 seconds, a woman enters. Or a timecode blockRanges are budgets, not frame-accurate edit points
On-screen textA street sign that says MARKET LANEUndefined text gets invented. English renders, other scripts do not
TriggersWhen the person touches the mirror, it ripples like liquidEvent-driven instructions land more precisely than abstract timing
[0-3s] A person is walking
[3-6s] They stop and turn around
[6-10s] They start running

Negatives go in the prompt itself, such as no dialogue or do not add captions. There is no negative-prompt field.

Step 3: add references and frames

Ten reference images and three reference videos per generation. Tags assign each file its role, numbered from zero in upload order.

TagRoleExample
FIRST_FRAMEStarting frame<FIRST_FRAME> a woman is walking
LAST_FRAMEFinal frame, paired with a first frame<FIRST_FRAME> <LAST_FRAME> a woman is walking
IMAGE_REF_0Reference: subject, product, or stylein the style of <IMAGE_REF_0>, a woman is walking
VIDEO_REF_0Character or object referencethe person in <VIDEO_REF_0> is playing the violin
  • Attach everything up front. A reference added mid-conversation destabilizes a scene that was holding.
  • Give each reference one job. Style from the first image, subject from the second, beats leaving it to guess.
  • Keep video references short. Three seconds each, likenesses work best, their audio is ignored.
  • Same image as first and last frame returns a seamless loop.

Step 4: edit and extend the result

A generated take stays live in the conversation, and uploaded footage of 10 seconds or less can join it. Everything you describe is something the model may re-render, so describe only the change.

AvoidWrite instead
In the video of the man on the sofa, add a small black cat that runs in from the right, jumps onto his lap, and he strokes its headAdd a cat that jumps onto his lap, he begins to pet it. Keep everything else the same.
Make the whole scene feel more dramatic and cinematic with moodier colors and stronger shadowsChange the lighting to be more dramatic. Keep everything else the same.
RuleDetail
One change per turnCompound instructions break consistency about twice as often
Four edits per chainPast that, drift creeps in. Re-anchor with: keep all character details exactly as they are, only change X
Extension appends only10s at a time to 40s total. It cannot prepend or stretch the middle
Extension clock restartsAfter 2s means two seconds into the new footage
References can ride alongA new character enters the extended scene from an image you attach
Extend this video. The scene continues: she sets the cup down, and the camera
pulls back through the window into the rain. The music fades to street noise.

Gemini Omni Flash 1.1 limits

LimitWhat happensWork around it
Non-Latin textJapanese and Chinese glyphs render as convincing inventionsKeep on-screen text in English, or check every frame
Crowded scenesFour or more tracked subjects merge or driftHold it to three
No voice editingNo audio reference uploads, no voice surgeryDirect the voice in text; consistency keeps it
Policy edgesReal brands and recognizable people are blocked, by regionKeep branded elements generic

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

1100 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3625 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6350 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24650 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

How do I write a good Gemini Omni Flash 1.1 prompt?
Write a short shot brief: subject and action, setting, camera, light, and the audio you want, in plain sentences. Omni Flash builds a multi-shot mini-narrative by default, so add "single continuous shot, no scene cuts" when you want one unbroken take. Always describe the sound, because a prompt that says nothing about audio gets a soundtrack the model invents on its own.
How do edit prompts work in Gemini Omni Flash 1.1?
Keep them short and targeted: "make the phone invisible, keep everything else the same". Overly descriptive edit prompts cause unintended changes, because everything you describe is something the model may re-render. Name the one change, add "keep everything else the same", and run one change per turn rather than stacking three into a sentence.
How do I extend a video with Gemini Omni Flash 1.1?
Prompt for the continuation: "extend this video" works, and "the scene continues: the camera pans across the mountains" gives it direction. Each extension adds up to ten seconds, to a total of 40, using the last ten seconds of footage as context and blending the final frames so the join reads as one take. Extension only appends to the end; it cannot prepend or stretch the middle.
What do the tags in a Gemini Omni Flash 1.1 prompt mean?
Tags bind your uploaded media to roles. FIRST_FRAME and LAST_FRAME in angle brackets set the start and end stills, IMAGE_REF_0 and VIDEO_REF_0 mark references, counted from zero in upload order. Passing the same image as both first and last frame produces a clip that loops. Without tags the model infers each file's role from the prompt, which works for simple cases and gets ambiguous past two or three inputs.
How many references does Gemini Omni Flash 1.1 take?
Up to ten reference images and up to three reference videos per generation, and a video reference must be three seconds or shorter. Images carry subjects, products, settings, or styles; video references work best for a likeness, and any audio inside them is ignored. Attach everything before the first generation, because adding references mid-conversation destabilizes a scene that was holding together.
How do I direct audio and dialogue in Gemini Omni Flash 1.1?
Describe the track you want in the same prompt as the picture: "no music, just room tone", "a high energy techno beat", or a spoken line written out for the character to say. Dialogue comes back lip-synced in a voice that persists across edits and extensions. Silence needs asking for too, since an unspecified prompt usually returns generic background music.
Can Gemini Omni Flash 1.1 render text inside the video?
Yes, and it is one of the model's stronger suits in English: signs, lower thirds, and word-by-word text animations come back readable when you spell out exactly what each surface says. Define even background text, or the model invents its own. Non-Latin scripts are the weak spot, so verify every frame if you need Japanese or Chinese text.
How do I keep characters consistent across edits in Gemini Omni Flash 1.1?
Make one focused change per turn, and let the model's scene memory do the rest: faces, wardrobe, and voices persist across follow-ups automatically. On longer chains, re-anchor with "keep all character details and motion exactly as they are, only change the lighting" before the change you want. Chains of about four edits stay stable; past that, plan a fresh generation with your references re-attached.