Europe 2026
Let's go Bananas with GenMedia — Guillaume Vernade, Google DeepMind
Overview
This talk demonstrates how to use Google DeepMind's GenMedia models to create rich multimedia content, such as images, videos, and music, by illustrating a book. The presentation walks through practical examples of generating character images, scene illustrations, video clips, and musical scores, all while emphasizing the integration and capabilities of various AI models. The core idea is to leverage AI to bring creative works to life through diverse media.
Who should watch
- AI Engineers looking to integrate generative media models into their workflows.
- Product Managers and builders exploring new ways to create engaging content.
- Developers interested in practical applications of multimodal AI for creative projects.
- Anyone curious about generating images, videos, and music from text prompts.
Key takeaways
- Google DeepMind is actively developing and releasing a suite of generative media models, including image generation (Nano Banana 2), video generation (VEO 3.1), and music generation (Lyria).
- The Gemini models, particularly with their large context windows, can be effectively used to generate prompts for other generative media models, creating a powerful content creation pipeline.
- The GenMedia models can be integrated using SDKs, with options for different service tiers (flex, standard, priority) affecting cost and speed.
- The Lyria music generation model allows for detailed control over song structure, duration, instrumentation, and even lyrics, all specified within the prompt.
- Advanced techniques include using the Interactions API for stateful and stateless interactions, and a clever workaround for the Text-to-Speech model to simulate multiple distinct character voices using a single voice model.
- The presentation showcases a practical workflow for illustrating Kenneth Grahame's The Wind in the Willows, demonstrating how to generate character portraits, chapter illustrations, and video sequences.
- The Lyria real-time model offers a unique interactive music generation experience, allowing for dynamic adjustments to musical style and composition in real-time, akin to a DJ mixing tracks.
Notable quotes
*The goal is really to have like one model that encompass all of that.*
*This year is going to be the year of agent, the real one. Last year was everybody talking of agents. This year is the is a year where we are actually going to build agents.*
*We have the Lyria real time model... it's basically a live model. So, it creates music and it continue to create music uh in real time until you stop it.*
Unofficial community note. Prefer the recording for nuance.