From Mixture of Experts to Mixture of Agents with Super Fast Inference - Daniel Kim & Daria Soboleva
This talk explores scaling large language models (LLMs) beyond traditional Mixture of Experts (MoE) architectures by introducing the concept of Mixture of Agents (MoA). MoA leverages multiple specialized LLMs, akin to agents, to collectively solve complex problems, aiming for higher intelligence and efficiency compared to monolithic models. The presentation also highlights Cerebras' hardware, designed for extremely fast inference, which enables practical implementation of MoA systems.
World's Fair 2025 53 min