Session brief
See, Hear, Speak, Draw: Logan Kilpatrick & Simón Fishman
Overview
This talk explores the burgeoning field of multimodal AI, moving beyond text-based interactions to incorporate vision and audio. While current applications often treat different modalities as separate "islands" connected by text, the future points towards unified models capable of processing and generating across various inputs and outputs simultaneously. The presentation highlights practical patterns and demos for building with existing multimodal capabilities, anticipating future advancements.
Who should watch
- AI Engineers
- Product Managers
- Builders experimenting with AI capabilities
- Those interested in the future of AI interaction beyond text
Key takeaways
- 2023 was the year of chatbots, while 2024 is poised to be the year of multimodal models.
- Current multimodal systems often act as separate tools (e.g., DALL-E for image generation, Whisper for audio transcription, GPT-4V for image and text input).
- Text currently serves as a connective tissue between different modalities, enabling interesting applications.
- A future vision includes unified models that can seamlessly process and reason across multiple modalities like text, images, and audio.
- Demo 1 showcased a loop using GPT-4V to describe an image, DALL-E 3 to generate a synthetic version, and GPT-4V again to compare and refine the output, illustrating iterative visual refinement.
- This iterative process, previously requiring human intervention, can now be partially automated by models, opening new interaction patterns.
- Demo 2 presented a method for video summarization by combining visual descriptions from GPT-4V analyzing video frames with audio transcription from Whisper, creating a richer textual representation.
- Developers are encouraged to start thinking multimodally, as future products will likely leverage these advanced capabilities.
Notable quotes
*2023 has really been the year of chatbots and I think it's been incredible to see how much people have actually been able to do.*
*I'm excited for 2024 which I think is is really going to be the year of multimodal models.*
*The majority of the work of making multimodal systems today is like how do you hook everything up together and connect the different modalities.*
Unofficial community note. Prefer the recording for nuance.