Session brief
It's 10pm. Do You Know Where Your Agents Are? — Kim Maida, Keycard
Overview
This talk addresses the challenges of managing and understanding AI agents in production environments. It emphasizes the need for robust monitoring and observability to ensure agents behave as expected and to debug issues effectively. The core
Continue
Related
- [Evals Workshop] Mastering AI Evaluation: From Playground to Production
This talk focuses on mastering AI evaluation, moving from initial development to production. It emphasizes that even the best large language models (LLMs) require a robust testing framework due to issues like hallucinations and performance degradation with changes. Effective evaluation helps answer critical questions about model selection, cost-effectiveness, brand consistency, and ongoing improvement, ultimately reducing development time, costs, and enabling faster iteration.
- \"Software engineering is not about writing code\" — Benoit Schillings, Google DeepMind VP of Research
This talk argues that the core challenge in software engineering is shifting from writing code to managing complexity, designing systems, and ensuring correctness. While AI models now excel at generating code syntax, the future of software development lies in leveraging AI for higher-level tasks like architectural design, complex problem decomposition, and inductive reasoning. The economics of code production are changing, making code generation nearly free and emphasizing the need for new processes and evaluation methods.
- 10x Development: LLMs For the working Programmer - Manuel Odendahl
This talk explores how programmers can leverage Large Language Models (LLMs) to significantly boost their productivity, aiming for a "10x" improvement. The core thesis is that LLMs should be treated not as autonomous reasoning agents, but as powerful translation engines and world simulators. By understanding how to decompose problems into language translation steps and by creatively framing prompts, developers can unlock new levels of efficiency and innovation.
- A Piece of Pi: Embedding The OpenClaw Coding Agent In Your Product — Matthias Luebken, Tavon
This talk explores the integration of coding agents, specifically the OpenClaw agent, into products. It emphasizes that coding agents are becoming a fundamental building block for software systems and highlights the Pi framework as an excellent tool for experimentation and development. The presentation demonstrates how agents can be leveraged to automate complex workflows, such as processing sales proposals, by interacting with various tools and systems.
- A year of Gemini progress + what comes next — Logan Kilpatrick, Google DeepMind
Logan Kilpatrick of Google DeepMind discussed a year of progress with Gemini models and outlined future developments. The talk highlighted the release of a new Gemini 2.5 Pro model, emphasizing its improved performance across benchmarks and its role as a turning point for the Gemini family. Kilpatrick also touched upon the organizational shifts within Google that have integrated research and product teams, enabling faster delivery of AI capabilities to both consumers and developers.
Unofficial community note. Prefer the recording for nuance.