← Browse

Europe 2026

Any-to-Any: Building Native Multimodal Agents - Patrick Löber, Google DeepMind

Patrick Löber

Overview

This talk introduces the concept of "any-to-any" multimodal agents, enabled by Google DeepMind's Gemini API. The core idea is to build agents that can understand and generate across various modalities including text, code, images, audio, and video. The presentation outlines the architecture for such agents, focusing on multimodal understanding and native generation capabilities, and demonstrates how to build a notebook LM clone as an example.

Who should watch

Key takeaways

Notable quotes

Patrick Löber states, *Gemini does not only understand text, right? It's natively multimodal, so you can also feed in code, image, audio, video, and then some more like URLs and also Google Search.*
Regarding audio input, *1 minute of audio translates to a 1920 tokens. And Gemini has a token limit of 1 million.*
On native generation, *these models understand the world.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.