← Browse

Europe 2026

How Transformers Finally Ate Vision – Isaac Robinson, Roboflow

Isaac Robinson

Overview

This talk explores the evolution of computer vision models, detailing how transformers, despite lacking inherent inductive biases for vision, ultimately surpassed traditional convolutional neural networks (CNNs). This shift was driven by massive, specialized pre-training techniques and leveraged infrastructure advancements from large language models, enabling transformers to learn visual patterns effectively.

Who should watch

Key takeaways

Notable quotes

The argument is that it is because of massive VIT-specific pretraining, and then we get to borrow a lot of speedups and infrastructure from the fact that LLMs are blowing up.
*Massive VIT-specific pre-training plus speed ups from LLMs plus pre-training compatible neural architecture search* enables flexible deployment.

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.