← Browse

World's Fair 2025

Why you should care about AI interpretability - Mark Bissell, Goodfire AI

Mark Bissell , Goodfire AI

Overview

This talk explores mechanistic interpretability, a field focused on reverse-engineering neural networks to understand their internal workings. It argues that interpretability is moving from research labs into practical applications, offering AI engineers new tools for debugging, enhancing user experiences, and advancing scientific discovery. The core thesis is that understanding how AI models function internally is becoming crucial for building more reliable, controllable, and insightful AI systems.

Who should watch

Key takeaways

Notable quotes

*Interpretability is really all about reverse engineering neural networks to understand what is going on inside of them.*
*What if you could debug and program your models at the neuron level to get more of those guarantees that we're used to with traditional software development.*
*The hallmark of an engineer is that we like to understand how systems work. We like to take a thing and take it apart and look at all the insides of it.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.