Europe 2026
Prompt to Pipeline: Building with Google's Gen Media Stack — Paige & Guillaume, Google DeepMind
Overview
This talk introduces Google DeepMind's latest advancements in generative AI, focusing on the Gemini family of models and their applications. The presentation highlights the multimodal capabilities of Gemini, its cost-effectiveness, and its integration into tools like AI Studio for building applications. It also touches upon the development of specialized models for generative media, including image, video, and music creation, and the increasing accessibility of powerful AI models for on-device and local execution.
Who should watch
- AI Engineers
- Product Managers
- Developers building AI-powered applications
- Researchers exploring generative media and multimodal AI
- Anyone interested in the latest advancements in large language models and their practical applications
Key takeaways
- Google DeepMind has released a suite of Gemini models (e.g., 3.1 Flash, Pro, Nano Banana 2, VO 3.1 Light, LIA 3, Genie 3) offering diverse capabilities from real-time conversation to advanced media generation.
- Gemini models are natively multimodal, capable of understanding and generating text, code, images, audio, and video, with a focus on cost-effectiveness and performance.
- AI Studio provides a platform for developers to easily access and experiment with these models, offering features like a playground, build tools for app creation, and code generation.
- Generative media models like Nano Banana 2 for image generation and VO 3.1 Light for video generation are becoming more accessible and affordable, enabling rapid prototyping and creative applications.
- LIA 3 offers music generation capabilities, allowing for the creation of songs with lyrics and various musical styles via API.
- The Gemma 4 family of models, including smaller E2B and E4B variants, are designed for on-device execution on low-power hardware like mobile phones and Raspberry Pis, while larger models (26B, 31B) can run on laptops and single GPUs.
- Gemma models are increasingly multimodal and agentic, with improved capabilities in coding, function calling, and reasoning, making them suitable for local AI development and applications.
- Google is focusing on making its models compatible with existing developer tools and ecosystems, such as LM Studio, Ollama, and VS Code, to facilitate easier integration and deployment.
Notable quotes
*The reality is I think that the m like all of that will probably be absorbed into the model eventually.*
*If you can get it working in AI Studio, you can get it working as part of your app.*
*The idea that you can use that to prototype, test your prompts and so on. And then if you want better quality then you can move to the to the better models.*
Unofficial community note. Prefer the recording for nuance.