← Browse

World's Fair 2025

From Mixture of Experts to Mixture of Agents with Super Fast Inference - Daniel Kim & Daria Soboleva

Daniel Kim , Daria Soboleva

Overview

This talk explores scaling large language models (LLMs) beyond traditional Mixture of Experts (MoE) architectures by introducing the concept of Mixture of Agents (MoA). MoA leverages multiple specialized LLMs, akin to agents, to collectively solve complex problems, aiming for higher intelligence and efficiency compared to monolithic models. The presentation also highlights Cerebras' hardware, designed for extremely fast inference, which enables practical implementation of MoA systems.

Who should watch

Key takeaways

Notable quotes

*Instead of having one monolithic fit forward network, we will create separate fit forward networks and we will call them experts.*
*A mixture of agents is basically the ability to take advantage of these earthshattering speeds from our hardware and apply them into harder problems.*
*The whole point though is that it can be better. It's just that you have to like actually like engineer it to be better.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.