Gemini Omni creates video from text, image and audio
Google has launched Gemini Omni, a new multimodal model capable of creating and editing videos based on input text, images and audio. This marks an expansion of the Gemini family's capabilities within generative AI.

What happened?
Google has introduced Gemini Omni, an advanced multimodal model capable of handling and processing text, images, audio and video. Omni Flash is the first application of this model. The model is designed to generate and edit video footage through text-based conversations, allowing users to describe desired video content in words.
Key facts
| Modellnamn | Gemini Omni |
|---|---|
| Funktion | Skapar och redigerar video från text, bild, ljud |
| Första tillämpning | Omni Flash |
| Status | Beta |
”Google's Gemini Omni is a new multimodal model that reasons across text, images, audio, and video to generate and edit videos through simple conversation — starting with Omni Flash.”
Why it matters
The development of multimodal AI models like Gemini Omni represents a significant step forward in generative AI, particularly in video production. By converting various types of input data into video, new opportunities emerge for creators and developers to rapidly prototype and produce visual content. This could potentially streamline processes within areas such as content creation, marketing and education.
Who is affected?
Developers and AI researchers gain access to a new tool for experimenting with and building video generation applications. Users needing to create video content, for example in social media or education, can benefit from simplified workflows. Companies investing in or using generative AI may see increased efficiency in their media production.
What else you should know
Omni Flash is the first specific feature mentioned in connection with the launch of Gemini Omni. The model is currently in beta.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka typer av input kan Gemini Omni hantera?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Gemini Omni creates video from text, image and audio"