Your Agent Failed in Prod. Good Luck Reproducing It. - Tisha Chawla & Susheem Koul, Microsoft
This talk addresses the challenge of debugging AI agents in production, where failures can be difficult to reproduce due to the inherent non-determinism of LLMs. The core thesis is that instead of chasing bitwise determinism, which is often unattainable, developers should focus on replayability and observability. This involves capturing the full state of an agent's execution to enable debugging and testing, even when the exact same output cannot be guaranteed on subsequent runs.
World's Fair 2026 14 min