Gemini Omni
Google Gemini Omni Flash - Text, image & reference video with native audio and multilingual lip-sync. On Influverse a 5s clip at 720p costs 25 credits, charged from the same wallet as every other model. There is no separate subscription for this model and no training step to use your own character with it.
By Influverse AI Team · Last verified: 6 September 2026 · Editorial policy
Generate with Gemini OmniCapabilities
| Vendor | |
|---|---|
| Duration | 3 to 10 seconds |
| Resolutions | 720p |
| Audio | Native audio on every generation |
| Text to video | Yes |
| Image to video | Yes |
| First and last frame | No |
| Motion control | No |
| Extend an existing clip | Yes |
| Reference video | No |
| Reference images | Up to 9 |
What it costs in credits
Derived from the same pricing source the app charges against, so these figures and the cost preview on the generation screen can never disagree.
| Duration | Resolution | Credits |
|---|---|---|
| 3s | 720p | 15 |
| 5s | 720p | 25 |
| 10s | 720p | 50 |
Credits come from your plan and from credit packs. See the pricing page for the plan ladder.
When to use Gemini Omni
Gemini Omni Flash is Google's fast multimodal video model and the other place to go for talking content. Like HappyHorse it generates native audio and multilingual lip-sync on every clip, and it accepts text, an image and a large set of reference images. Unlike HappyHorse it can extend an existing clip, which is genuinely useful when a spoken line runs slightly longer than the take you generated: extend rather than reroll and the delivery you already liked survives. Its duration menu is the shorter of the two and it offers a single resolution, so the trade is control and speed against range. Reach for it when you want a spoken clip quickly, when you are working in a language other than English, or when you are building a sequence out of short spoken beats and expect to extend one or two of them. It sits in the middle of the catalogue on price with audio included, which makes it reasonable value for talking content and poor value for anything silent, since you cannot switch the audio off. If the clip is going under a music bed or a separately recorded voiceover, generate it on Veo 3.1 Lite or Kling 2.6 Pro instead and keep the difference. If the clip needs to run past its duration ceiling, either extend it here or move to HappyHorse or Seedance. As with HappyHorse, remember that the dedicated talking-shot pipeline exists when you need the character to say exact words in the character's own voice.
Strengths
- Native audio and multilingual lip-sync on every clip
- Can extend an existing clip, which preserves a delivery you liked
- Large reference image set
- Fast turnaround for spoken content
Limits
- Single resolution and the shortest duration menu of the audio models
- Audio cannot be switched off, so it is poor value for silent clips
- No first and last frame control
Best for
- Short spoken beats and quick talking clips
- Multilingual social content
- Sequences where one beat may need extending
Other models to consider
Back to the full model catalogue
More from Influverse
Generate with Gemini Omni
Every new account starts with a free trial that unlocks every surface and every model. No card required.
Start creating free