World's Fair 2025
293 sessions tagged from titles
- Shipping AI That Works: An Evaluation Framework for PMs – Aman Khan, Arize
This talk introduces an evaluation framework for AI product managers (PMs) focused on shipping reliable AI applications. It emphasizes the critical need for robust evaluation, analogous to software testing but adapted for the non-deterministic nature of AI models. The framework aims to provide PMs with tools and methodologies to build confidence in AI product performance, moving beyond subjective "vibe coding" to data-driven "thrive coding."
- Dispatch from the Future: building an AI-native Company – Dan Shipper, Every, AI & I
This talk explores the emerging playbook for building AI-native companies, emphasizing that the process is currently being invented collaboratively. The core thesis is that achieving 100% AI adoption among engineers creates a significant, tenfold difference in organizational capability. This shift enables a single engineer to build and maintain complex production products, fundamentally altering how software development is approached.
- AI Consulting in Practice – NLW, Superintelligent, @AIDailyBrief
This talk explores the current state of enterprise AI adoption, moving beyond the narrative of an AI bubble to focus on tangible value and return on investment (ROI). It presents findings from a study collecting self-reported ROI data across various AI use cases, highlighting that organizations are indeed finding value, with a significant portion reporting modest to high ROI. The discussion emphasizes the growing adoption of AI, particularly in software engineering, and the increasing deployment of agents within enterprises, despite challenges in scaling beyond pilot phases.
- From Vibe Coding To Vibe Engineering – Kitze, Sizzy
This talk explores the evolution of front-end development, contrasting traditional coding practices with emerging AI-assisted workflows. It introduces the concept of "vibe engineering" as a more sophisticated approach to using AI agents for coding, emphasizing the need for skilled engineers to guide and refine AI-generated code. The core thesis is that while AI can accelerate development, human expertise remains crucial for quality, abstraction, and complex problem-solving.
- VoiceVision RAG - Integrating Visual Document Intelligence with Voice Response — Suman Debnath, AWS
This talk introduces VoiceVision RAG, a system that integrates visual document intelligence with voice response capabilities. It explores multimodal RAG techniques, focusing on vision-based retrieval using the Pali model. The presentation demonstrates how to build agentic applications with the Strands framework, enabling voice-enabled retrieval and response generation from visual documents.
- Government Agents: AI Agents Meet Tough Regulations — Mark Myshatyn, Los Alamos National Lab
This talk explores the application of AI agents within a national laboratory setting, emphasizing their potential to accelerate scientific discovery and address complex national security challenges. It highlights the evolution from early computational methods to modern agentic systems, stressing the need for these tools to operate within stringent regulatory and security frameworks. The core thesis is that AI agents can significantly enhance scientific and operational capabilities, provided they are developed with explainability, isolation, governance, and speed in mind, fostering crucial partnerships between government entities and the broader tech community.
- Rishabh Garg, Tesla Optimus — Challenges in High Performance Robotics Systems
This talk addresses the complexities of high-performance robotics systems, focusing on the critical interplay between control policies and the underlying software and hardware infrastructure. It highlights how issues that appear to stem from the control policy often originate in the communication protocols, threading, synchronization, logging, and priority management within the system. The presentation uses a simplified toy robot architecture to illustrate common pitfalls and debugging strategies.
- Building an Agentic Platform — Ben Kus, CTO Box
This talk explores Box's journey in building an agentic platform, focusing on how agentic approaches can solve complex data extraction challenges that traditional methods and basic LLM calls struggle with. The core thesis is that an agentic abstraction layer provides a clean, flexible, and evolvable architecture for tackling sophisticated AI tasks, moving beyond simple chatbot interactions to orchestrate complex workflows.
- Five hard earned lessons about Evals — Ankur Goyal, Braintrust
This talk emphasizes the critical role of robust evaluation (evals) in the development and deployment of AI systems. It argues that evals should be actively engineered, not passively accepted, and that they are essential for both playing offense by identifying new use cases and defense by ensuring quality. The core thesis is that by treating evals as a continuous engineering process, organizations can better adapt to rapid model advancements and ship higher-quality AI products.
- Perceptual Evaluations: Evals for Aesthetics — Diego Rodriguez, Krea.ai
This talk explores the limitations of current AI evaluation metrics, particularly in generative media, by highlighting how they often fail to account for human perception and aesthetic judgment. The speaker argues that traditional metrics, like FID scores, can be misled by factors such as compression artifacts, leading to inaccurate assessments of AI model quality. The core thesis is that AI evaluation needs to evolve to incorporate human perceptual nuances and subjective qualities, moving beyond easily quantifiable but potentially superficial measures.
- How BlackRock Builds Custom Knowledge Apps at Scale — Vaibhav Page & Infant Vasanth, BlackRock
BlackRock has developed a framework to accelerate the creation of custom AI and knowledge applications at scale within their asset management operations. This system addresses the complexities of data extraction, LLM strategy selection, and deployment challenges, aiming to reduce app development time from months to days by empowering domain experts.
- Form factors for your new AI coworkers — Craig Wattrus, Flatfile
This talk explores new ways to interact with AI, moving beyond traditional interfaces to create more intuitive and collaborative "AI coworkers." It emphasizes understanding the capabilities and limitations of AI models as if they were physical materials, allowing for the development of novel form factors that enable AI to perform tasks more effectively and align with user goals. The core idea is to shift from simply automating tasks to fostering emergent behaviors and deeper collaboration between humans and AI.
- Fuzzing in the GenAI Era — Leonard Tang, Haize Labs
This talk introduces hazing, a form of fuzz testing adapted for Generative AI systems. It addresses the critical challenge of validating and verifying AI outputs, which are inherently subjective and unstructured. Traditional evaluation methods relying on static datasets are insufficient due to the brittle and non-deterministic nature of GenAI, where minor input variations can lead to drastically different outputs. Hazing aims to pressure-test AI systems through large-scale simulation and optimization before deployment to ensure robustness and reliability.
- Multi Agent AI and Network Knowledge Graphs for Change — Ola Mabadeje, Cisco
This talk introduces a system designed to reduce failures in network change management by leveraging multi-agent AI and network knowledge graphs. The core thesis is that by creating a digital twin of the production network, represented by a knowledge graph, and enabling specialized AI agents to interact with it, complex network operations can be made more robust and efficient. The system aims to provide a natural language interface for network operations teams and integrate with existing IT service management tools.
- Wisdom-Driven Knowledge Augmented Generation at Scale - Chin Keong Lam, Patho AI
This talk introduces Knowledge Augmented Generation (KAG) as an enhancement over traditional Retrieval Augmented Generation (RAG). KAG integrates structured knowledge graphs with large language models to enable more accurate, insightful, and decision-driving responses. The core thesis is that by representing wisdom as interconnected knowledge, AI systems can move beyond simple data retrieval to understanding and strategic advisory roles, particularly in complex domains.
- The Next Unicorns: 7 Top AI startups from the HF0 Residency
This talk highlights seven AI startups from the HF0 Residency program, showcasing innovative approaches to AI product development and deployment. The startups address diverse challenges, from creating personalized user experiences and improving data attribution to developing novel AI architectures and enabling widespread AI adoption. The core thesis emphasizes the shift from simply scaling AI models to focusing on reliability, developer ecosystems, and practical applications that solve real-world problems.
- #define AI Engineer - Greg Brockman, OpenAI (ft. Jensen Huang)
This talk features Greg Brockman discussing the evolution of AI engineering, his personal journey into coding and AI, and the future of building AI systems. He emphasizes the importance of practical building, the synergy between research and engineering, and the potential for AI to fundamentally transform various industries by lowering barriers to entry and enabling new forms of economic activity.
- The Future of Evals - Ankur Goyal, Braintrust
This talk discusses the evolution of AI model evaluation (evals), highlighting the shift from manual processes to automated optimization. It introduces Loop, an agent designed to enhance prompts, datasets, and scorers, thereby revolutionizing the eval process. The core thesis is that frontier models, particularly recent advancements, are now capable of significantly improving AI development workflows.
- Designing AI-Intensive Applications - swyx
This talk explores the evolving landscape of AI engineering, emphasizing the need for new standard models to guide the development of AI-intensive applications. The speaker posits that the field is moving beyond simple wrappers and demos towards robust production systems, drawing parallels to foundational periods in physics and other engineering disciplines. The core thesis is that identifying and adopting these new standard models will be crucial for building valuable and intelligent AI products.
- How to look at your data — Jeff Huber (Chroma) + Jason Liu (567)
This talk emphasizes the critical importance of measuring and analyzing both the inputs and outputs of AI systems to drive systematic improvement. The core thesis is that effective measurement, akin to Peter Drucker's adage, is essential for making informed decisions and enabling continuous enhancement. By looking at data, practitioners can move beyond guesswork and build more robust and user-centric AI products.
- On Engineering AI Systems that Endure The Bitter Lesson - Omar Khattab, DSPy & Databricks
This talk addresses the challenge of engineering AI systems in a rapidly evolving landscape, drawing parallels to Rich Sutton's "bitter lesson" in AI research. The core thesis is that while scaling and general methods are crucial for intelligence, AI engineering must focus on building reliable, robust, and understandable systems by abstracting away from rapidly changing low-level components. This involves investing in system design, clear specifications, and modularity rather than premature optimization or tightly coupled, model-specific implementations.
- Evals Are Not Unit Tests — Ido Pesok, Vercel v0
This talk introduces the concept of evals at the application layer, distinguishing them from traditional unit tests. It emphasizes that Large Language Models (LLMs) can be unreliable, leading to unexpected failures in AI applications even when basic functionality appears to work. The core thesis is that robust evals are crucial for building reliable AI products by systematically testing and measuring performance across a spectrum of user-driven scenarios.
- 2025 is the Year of Evals! Just like 2024, and 2023, and … — John Dickerson, CEO Mozilla AI
The core thesis of this talk is that 2025 will finally be the year for robust AI evaluation, driven by the convergence of several key factors. The increasing understanding of AI by business leaders, coupled with budget shifts towards generative AI projects, has set the stage. Crucially, the rise of agentic systems, which make decisions and take actions, introduces significant complexity and risk, making rigorous evaluation a necessity rather than an option.
- Vibe Coding with Confidence — Itamar Friedman, Qodo
This talk explores the evolution of AI in software development, moving beyond simple autocompletion and chat interfaces to more sophisticated, agent-driven workflows. The core thesis is that the Command Line Interface (CLI) will become a central hub for "vibe coding with confidence," enabling developers to manage complex, end-to-end tasks with greater reliability and control. This shift aims to address the limitations of current AI tools, particularly in enterprise settings, by integrating AI across the entire Software Development Life Cycle (SDLC).
- Full Workshop: Realtime Voice AI — Mark Backman, Daily
This workshop introduces Pipcat, an open-source Python framework for building real-time voice and multimodal AI agents. It emphasizes the challenges of creating natural, fast, and conversational voice AI, highlighting the advancements in speech-to-speech models that simplify pipeline complexity. The session provides a hands-on approach to building a voice bot, demonstrating Pipcat's modularity and orchestration capabilities.
- Vision AI in 2025 — Peter Robicheaux, Roboflow
This talk addresses the current state of AI vision, arguing that computer vision models lag significantly behind language models in terms of intelligence and pre-training leverage. The core thesis is that vision models are not yet "smart" due to limitations in evaluation metrics, a lack of effective large-scale pre-training utilization, and challenges in aligning visual and linguistic features. The presentation introduces new benchmarks and models aimed at improving vision AI's capabilities.
- Practical tactics to build reliable AI apps — Dmitry Kuchin, Multinear
This talk emphasizes a practical, user-centric approach to building reliable AI applications, moving beyond generic data science metrics. The core thesis is that AI app reliability stems from defining and testing against real-world scenarios and desired business outcomes, rather than abstract measures like factuality or groundness. This method allows for continuous improvement and confidence in the application's performance.
- How to Improve your Vibe Coding — Ian Butler
This talk addresses the current limitations of AI coding agents in identifying and fixing bugs, highlighting their tendency to generate false positives and struggle with complex codebases. It proposes practical strategies for developers to improve the effectiveness of these agents, focusing on structured rules, better context management, and leveraging advanced "thinking" models. The core thesis is that while agents show promise, their current implementation requires careful guidance and configuration to be truly useful for developers.
- Vibes won't cut it — Chris Kelly, Augment Code
This talk argues that the current hype around AI code generation overlooks the complexities of production software engineering. While AI can generate code, it doesn't inherently understand the nuanced decision-making, maintenance, and safety requirements of large-scale, production-ready systems. The core thesis is that professional software engineers remain essential, and the focus should shift from the quantity of AI-generated code to how AI can augment the existing, rigorous practices of software development.
- Real World Development with GitHub Copilot and VS Code — Harald Kirschner, Christopher Harrison
- Building Agents at Cloud Scale — Antje Barth, AWS
This talk explores building AI agents at cloud scale, emphasizing the reinvention of customer experiences and the development of new AI-powered applications. It highlights the significant integration of agentic capabilities and large language models (LLMs) in reimagining existing services like Alexa, which now operates on over 600 million devices. The presentation introduces a model-driven approach and an open-source Python SDK called Strand Agents, designed to accelerate the development and deployment of production-ready AI agents.
- State of Startups and AI 2025 - Sarah Guo, Conviction
The AI landscape is rapidly evolving, with significant advancements in reasoning capabilities and the emergence of sophisticated AI agents. While the potential for AI is vast, building impactful AI products is more challenging than anticipated, yet the value creation is immense, enabling companies to scale faster than ever before. The market for model capabilities is becoming increasingly competitive, with open-source contributions playing a vital role.
- Useful General Intelligence — Danielle Perszyk, Amazon AGI
This talk proposes a shift in the goal of AI development from creating "thinking machines" to augmenting human intelligence. It argues that general intelligence is not solely contained within an individual machine but emerges from social interaction and co-evolution with technology. The core thesis is that by building AI agents that can reliably interact with digital environments and align with human representations, we can enhance human capabilities and create useful general intelligence.
- The 2025 AI Engineering Report — Barr Yaron, Amplify
The 2025 AI Engineering Report survey reveals key trends in the rapidly evolving field. Despite varied job titles, the community is broad and technical, with a significant influx of experienced software engineers new to AI. The report highlights widespread LLM adoption for diverse use cases, with OpenAI models leading for external products. Customization methods like RAG and fine-tuning are prevalent, and frequent model and prompt updates are common.
- Agents vs Workflows: Why Not Both? — Sam Bhagwat, Mastra.ai
This talk explores the synergy between AI agents and workflows, arguing that they are not mutually exclusive but rather complementary tools for building robust AI applications. The core thesis is that combining agentic capabilities with structured workflows offers a more powerful and flexible approach than relying on either in isolation. The speaker critiques the tendency for some to advocate for one approach exclusively, emphasizing the practical benefits of integrating both.
- Why We Don’t Need More Data Centers - Dr. Jasper Zhang, Hyperbolic
The talk argues that the escalating demand for AI compute, particularly GPUs, does not necessitate solely building more data centers. Instead, it proposes that a GPU marketplace and intelligent resource allocation can significantly address this demand more efficiently and sustainably. The core idea is to leverage underutilized existing GPU capacity rather than exclusively expanding physical infrastructure.
- Infrastructure for the Singularity — Jesse Han, Morph
This talk introduces Morph's Infinibranch technology, a virtualization, storage, and networking solution designed for advanced AI agents. It posits that as AI develops intelligence and personhood, it requires infrastructure that can match its speed and complexity. Infinibranch aims to provide a substrate for a cloud environment where AI agents can operate with zero latency, enabling reversible actions, parallel exploration of possibilities, and a more dynamic interaction with the digital world.
- Hacking the Inference Pareto Frontier - Kyle Kranen, NVIDIA
This talk explores techniques for optimizing AI model inference to break the Pareto frontier, focusing on balancing quality, latency, and cost. The core thesis is that a well-designed system, tailored to specific application constraints, is crucial for successful deployment and application performance. By understanding and manipulating factors like scale, structure, and dynamism, engineers can achieve better service level agreements or reduce costs for existing ones.
- Pipecat Cloud: Enterprise Voice Agents Built On Open Source - Kwindla Hultman Kramer, Daily
This talk introduces Pipecat, an open-source, vendor-neutral framework for building reliable and performant voice AI agents. It emphasizes the challenges in voice AI development, such as achieving fast response times (targeting under 800 milliseconds for voice-to-voice) and accurate turn detection. The presentation also highlights Pipecat Cloud, a new offering designed to simplify the deployment and scaling of these agents by abstracting away complexities like Kubernetes.
- [Full Workshop] Building Conversational AI Agents - Thor Schaeff, ElevenLabs
This workshop focuses on building multilingual conversational AI agents, detailing the pipeline from speech-to-text to text-to-speech. It highlights the integration of large language models as the agent's "brain" and showcases ElevenLabs' tools for creating dynamic, responsive AI interactions across numerous languages. The session emphasizes practical application and developer experience, offering insights into configuring and deploying these agents.
- From Self-driving to Autonomous Voice Agents — Brooke Hopkins, Coval
This talk draws parallels between the development of self-driving car technology and the creation of robust evaluation strategies for autonomous voice agents. The core thesis is that the challenges and solutions in achieving reliability and scalability in self-driving can inform how we build trustworthy and effective voice AI systems. By adopting principles like large-scale simulation and continuous evaluation loops, developers can move beyond the limitations of deterministic approaches and unpredictable autonomous agents.
- Your realtime AI is ngmi — Sean DuBois (OpenAI), Kwindla Kramer (Daily)
This talk emphasizes the critical role of low latency in building effective voice AI experiences, arguing that current real-time AI is often inadequate. The core thesis is that achieving natural, human-like voice interactions requires a fundamental shift in how audio and video are handled over networks, moving beyond traditional methods like WebSockets to leverage technologies like WebRTC. The speakers suggest that voice is poised to become a primary interface for the next generation of AI applications.
- Why ChatGPT Keeps Interrupting You — Dr. Tom Shapland, LiveKit
This talk addresses the persistent issue of voice AI agents, like ChatGPT's advanced voice mode, interrupting users. The core problem lies in current voice AI's simplistic approach to turn-taking, which contrasts sharply with the complex, predictive, and parallel processing humans use in conversation. The presentation explores lessons from human conversation and introduces emerging techniques to improve voice AI's ability to manage conversational flow.
- Serving Voice AI at $1/hr: Open-source, LoRAs, Latency, Load Balancing - Neil Dwyer, Gabber
This talk details the experience of hosting an open-source voice AI model, Orpheus, for real-time consumer applications. The core challenge addressed is achieving low latency and high fidelity voice generation at a cost viable for widespread consumer use, which often requires near-free operation. The presentation highlights the technical hurdles in serving these models efficiently and presents solutions involving model fine-tuning, optimized infrastructure, and load balancing.
- How to defend your sites from AI bots — David Mytton, Arcjet
This talk addresses the increasing problem of automated traffic on websites, exacerbated by the rise of AI. While bots have always existed, AI crawlers and agents are intensifying issues like increased costs, bandwidth consumption, and denial-of-service attacks. The presentation outlines various defense mechanisms, ranging from voluntary standards to sophisticated detection and mitigation techniques, to help site owners manage and control bot traffic.
- The Unofficial Guide to Apple’s Private Cloud Compute - Jmo, CONFSEC
This talk provides an unofficial guide to Apple's Private Cloud Compute (PCC) system, explaining how it enables remote AI computation while maintaining user privacy. The core thesis is that Apple addresses the inherent privacy risks of sending data to remote servers by implementing a system with five key requirements: stateless computation, enforceable guarantees, non-targetability, no privileged runtime access, and verifiable transparency. The presentation outlines Apple's conceptual architecture and technical components designed to meet these requirements, offering insights into how similar privacy-preserving techniques can be adopted by developers outside the Apple ecosystem.
- How to Secure Agents using OAuth — Jared Hanson (Keycard, Passport.js)
This talk addresses the critical need for robust security in AI agents by advocating for the adoption of OAuth 2.0. It highlights the current security risks associated with broadly scoped, long-lived API keys used by agents and proposes OAuth as a solution to transition from static secrets to dynamic, delegated access. The discussion emphasizes that while OAuth is complex, its core principles are straightforward and essential for securing increasingly connected and useful AI agents.
- How we hacked YC Spring 2025 batch’s AI agents — Rene Brandel, Casco
This talk explores common security vulnerabilities found in AI agents by detailing a process of "hacking" Y Combinator batch AI agents. The core thesis is that agent security extends beyond traditional LLM prompt injection to encompass broader system-level risks, emphasizing the need to treat agents with the same security considerations as human users and to avoid building custom code execution environments.
- OpenAI on Securing Code-Executing AI Agents — Fouad Matin (Codex, Agent Robustness)
This talk addresses the critical security and safety considerations for AI agents capable of executing code. As AI models become increasingly proficient at writing and running code, the focus shifts from mere capability to responsible deployment and robust guardrails. The presentation emphasizes that code execution is becoming a standard feature for AI agents, moving beyond traditional software engineering tasks to achieve objectives more efficiently across various applications.
- Evaluating AI Search: A Practical Framework for Augmented AI Systems — Quotient AI + Tavily
This talk introduces a practical framework for evaluating AI search systems, particularly those operating in dynamic, real-world environments. It highlights the limitations of traditional static benchmarks and proposes dynamic datasets and reference-free metrics as crucial components for assessing the performance and reliability of augmented AI systems. The core thesis is that continuous improvement in AI systems requires robust evaluation methods that adapt to the evolving nature of information and user interactions.
- Scaling Enterprise-Grade RAG: Lessons from Legal Frontier - Calvin Qi (Harvey), Chang She (Lance)
This talk addresses the complexities of building enterprise-grade Retrieval Augmented Generation (RAG) systems, particularly within specialized domains like legal documents. It highlights the challenges of handling massive, complex datasets, sophisticated user queries, and stringent security requirements. The discussion emphasizes the critical role of robust evaluation strategies and modern data infrastructure to support these demanding applications.
- Building Alice’s Brain: an AI Sales Rep that Learns Like a Human - Sherwood & Satwik, 11x
This talk details the development of Alice's Brain, an AI sales representative's knowledge base designed to learn and operate more like a human. The system shifts from a manual context-pushing model to an automated one where the AI proactively pulls and utilizes relevant seller information. This approach aims to improve personalization and efficiency in sales outreach.
- Layering every technique in RAG, one query at a time - David Karam, Pi Labs (fmr. Google Search)
This talk provides a practical framework for improving Retrieval Augmented Generation (RAG) systems by systematically layering techniques based on their complexity and impact. It emphasizes a quality engineering approach, starting with defined outcomes and product problems, then analyzing failures to select appropriate RAG techniques. The core thesis is that understanding where and why a system fails is crucial for making informed decisions about which RAG enhancements to implement, rather than getting lost in hype or theoretical debates.
- Building a Smarter AI Agent with Neural RAG - Will Bryk, Exa.ai
This talk introduces Exa, a search engine designed for AI agents, moving beyond traditional keyword-based search. It argues that current search engines, optimized for human users, are insufficient for the complex, data-intensive needs of AI. Exa aims to provide a more powerful and flexible API that allows AI to query and retrieve information from the web with greater precision and comprehensiveness.
- [Full Workshop] Building Metrics that actually work — David Karam, Pi Labs (fmr Google Search)
This workshop focuses on building effective evaluation metrics for AI systems, moving beyond basic testing to create robust scoring systems. It emphasizes that evaluations are not just for testing but are the primary place where domain knowledge resides, enabling significant improvements in AI development. The session introduces a methodology for creating nuanced, calibrated metrics that correlate with desired outcomes, ultimately simplifying the AI development stack.
- Make your LLM app a Domain Expert: How to Build an Expert System — Christopher Lovejoy, Anterior
This talk outlines a playbook for building domain-native LLM applications, emphasizing that the system surrounding the model is more critical than the model's sophistication itself. The core challenge, termed the "last mile problem," involves providing LLMs with specific contextual understanding of industry workflows. By developing an adaptive domain intelligence engine, companies can leverage domain experts to translate industry insights into performance improvements, enabling continuous iteration and refinement of AI applications.
- Shipping Products When You Don't Know What they Can Do — Ben Stein, Teammates
This talk addresses the challenges of building and shipping products in the rapidly evolving AI agent space, where the capabilities of the underlying LLMs are not fully understood. It proposes a shift in product development from defining specific requirements to focusing on affordances and emergent behaviors, emphasizing the need for new tools and practices to navigate this uncertainty. The core thesis is that product management must transform to embrace discovery and adaptation in a probabilistic world.
- Shipping something to someone always wins — Kenneth Auchenberg (ex. Stripe, VSCode)
The core thesis is that successful product development, especially in the age of AI, hinges on rapid, iterative shipping and continuous user feedback, rather than large, infrequent releases. The speaker advocates for building "skateboard" equivalents—minimally viable products that deliver immediate value and allow for quick iteration based on real user input, which is far more effective than a traditional, linear development approach.
- Why your product needs an AI product manager, and why it should be you — James Lowe, i.AI
This talk argues for the critical role of an AI product manager, emphasizing that AI expertise is essential for this position. It builds on the idea that as AI coding agents and product features reduce the cost of software development, the demand for individuals who can effectively decide *what* to build will increase. The core thesis is that a dedicated AI product manager, or at least an AI product management mindset within a team, is crucial for navigating the complexities of AI product development.
- Everything is ugly, so go build something that isn't — Raiza Martin, Huxe (ex NotebookLM)
This talk argues that the current era of AI development is characterized by "ugly" or clunky products, which are merely the precursors to a more refined future. The core thesis is that building truly great AI products requires deep personal clarity, a singular focus on purpose, and a commitment to building user trust, which ultimately enables delight. The speaker emphasizes that restraint and judgment are key innovation multipliers in a landscape often dominated by overwhelming capabilities.
- Building the platform for agent coordination — Tom Moor, Linear
This talk explores Linear's journey in integrating AI into its product development tool, moving from early, pragmatic applications like natural language filters to more sophisticated agent-based features. The core thesis is that AI should be seamlessly integrated to provide practical value, acting as scalable, cloud-based teammates that enhance productivity and quality without being overly intrusive. The platform is evolving to support these agents as first-class citizens, enabling complex workflows and interactions.
- What Is a Humanoid Foundation Model? An Introduction to GR00T N1 - Annika & Aastha
This talk introduces GR00T N1, a humanoid robotics foundation model developed by Nvidia. It addresses the challenge of bridging the gap between the intelligence of large language models and the physical world by enabling robots to operate in human environments. The model's development is framed within a physical AI lifecycle involving data generation, model training, and deployment on edge devices.
- Real-time Experiments with an AI Co-Scientist - Stefania Druga, fmr. Google Deepmind
This talk introduces the concept of an AI co-scientist, analogous to a pair programmer but for real-world scientific experiments. The system integrates various sensors and cameras to collect empirical data in real-time, which an AI then analyzes to provide insights, generate hypotheses, and potentially accelerate scientific discovery. The presented system is built using open-source hardware and software, demonstrating a low-cost, accessible approach to AI-assisted experimentation.
- Scaling AI Agents Without Breaking Reliability — Preeti Somal, Temporal
This talk addresses the challenges of building reliable and scalable AI agent applications. It posits that these agents are complex distributed systems requiring robust orchestration, state management, and error handling, especially given the inherent unreliability of LLMs and external tools. The core thesis is that platforms like Temporal can abstract away the complexities of reliability and scalability, allowing developers to focus on core business logic and agent functionality.
- Ship Agents that Ship: A Hands-On Workshop - Kyle Penfound, Jeremy Adams, Dagger
This talk demonstrates how to build and deploy AI agents that can actively contribute to software development workflows using Dagger. It emphasizes creating agents with specific tools and environments, enabling them to perform tasks like code editing and testing within a sandboxed, containerized system. The core idea is to integrate AI capabilities into existing software engineering pipelines, providing guardrails and structure to manage the output of generative AI.
- The AI Engineer’s Guide to Raising VC — Dani Grant (Jam), Chelcie Taylor (Notable)
This talk, The AI Engineer’s Guide to Raising VC, features Dani Grant and Chelcie Taylor discussing the process and nuances of securing venture capital funding, particularly for early-stage companies. The core thesis is that founders, especially those from technical backgrounds, often overemphasize product and technology details when pitching VCs. Instead, investors primarily bet on the founder's vision, unique market insights, and ability to build a strong team. The discussion highlights that revenue and even a fully developed product are not prerequisites for raising capital, and emphasizes the importance of compelling storytelling and understanding investor psychology.
- Strategies for LLM Evals (GuideLLM, lm-eval-harness, OpenAI Evals Workshop) — Taylor Jordan Smith
This talk addresses the critical need for robust evaluation and benchmarking strategies when deploying large language models (LLMs) into production. It highlights the inherent complexities and potential pitfalls of generative AI, emphasizing that scalability, reliability, and safety are paramount. The presentation introduces practical tools and methods to assess LLM performance, ensuring that models meet enterprise-level requirements before and during deployment.
- Why you should care about AI interpretability - Mark Bissell, Goodfire AI
This talk explores mechanistic interpretability, a field focused on reverse-engineering neural networks to understand their internal workings. It argues that interpretability is moving from research labs into practical applications, offering AI engineers new tools for debugging, enhancing user experiences, and advancing scientific discovery. The core thesis is that understanding how AI models function internally is becoming crucial for building more reliable, controllable, and insightful AI systems.
- Information Retrieval from the Ground Up - Philipp Krenn, Elastic
This talk explores the fundamentals of information retrieval, focusing on the "R" in Retrieval Augmented Generation (RAG). It delves into both traditional keyword search and modern vector search, explaining their underlying mechanisms, strengths, and limitations. The presentation emphasizes that effective retrieval is crucial for accurate and relevant results in AI applications.
- Introduction to LLM serving with SGLang - Philip Kiely and Yineng Zhang, Baseten
This talk introduces SGLang, an open-source framework designed for high-performance serving of large language models (LLMs) and large vision models (LVMs). It emphasizes SGLang's production readiness, speed, and strong community support, highlighting its ability to provide day-zero support for new model releases and allow users to contribute to its development. The framework is presented as a valuable tool for optimizing LLM inference.
- Waymo's EMMA: Teaching Cars to Think - Jyh Jing Hwang, Waymo
This talk explores the evolution of autonomous driving systems, highlighting the transition from early, less capable models to sophisticated L4 systems like Waymo's. It introduces EMMA, an experimental system leveraging multimodal large language models (LLMs) like Gemini to enhance driving capabilities, particularly in handling rare and complex scenarios. The core thesis is that LLMs can significantly improve the generalizability and safety of autonomous driving by understanding and reacting to diverse real-world situations.
- Robotics: why now? - Quan Vuong and Jost Tobias Springberg, Physical Intelligence
This talk explores the advancements in robotics driven by the emergence of vision-language action (VLA) models. It highlights the transition from robots operating in highly constrained environments to performing complex tasks in semistructured real-world settings. The core thesis is that significant progress in robotics is now enabled by general AI developments, particularly VLA models, and that the primary bottleneck is shifting from hardware to software and model intelligence.
- A2A & MCP Workshop: Automating Business Processes with LLMs — Damien Murphy, Bench
This talk explores the automation of business processes using Large Language Models (LLMs) through two key protocols: A2A (Agent-to-Agent) and MCP (Model Context Protocol). A2A facilitates communication between remote agents, enabling specialization and parallel processing, while MCP acts as a standardized interface for agents to access context and tools, akin to a USB-C for AI. The workshop demonstrates how to integrate these protocols to build multi-agent systems triggered by webhooks, automating tasks like bug reporting and information dissemination.
- Ship Production Software in Minutes, Not Months — Eno Reyes, Factory
This talk argues that the future of software development is agent-driven, moving beyond human-driven approaches. It posits that AI agents, when properly integrated and provided with sufficient context, can handle a majority of tasks across the software lifecycle, enabling production software to be built in minutes rather than months. The core thesis is that organizations must transition to an agent-native development paradigm to unlock the true power of AI.
- Beyond the Prototype: Using AI to Write High-Quality Code - Josh Albrecht, Imbue
This talk addresses the gap between AI-generated code prototypes and production-ready software. It introduces Sculptor, an experimental coding agent environment designed to build trust in AI-generated code by focusing on identifying and preventing defects. The core thesis is that AI should be leveraged not just for code generation but also for ensuring its quality and reliability.
- Software Development Agents: What Works and What Doesn't - Robert Brennan, OpenHands
This talk explores the practical application and effectiveness of software development agents, emphasizing the shift from manual coding to higher-level problem-solving. It argues that while AI excels at the iterative process of writing and running code, human engineers remain crucial for critical thinking, user empathy, and architectural decisions. The presentation details the core components of these agents, their underlying mechanisms, and best practices for their integration into the development workflow.
- Devin 2.0 and the Future of SWE - Scott Wu, Cognition
The talk discusses the rapid advancement of AI agents in software engineering, highlighting a trend of capability doubling approximately every 70 days for coding tasks. This exponential growth has transformed AI's role from simple tab completion to handling complex, multi-file development tasks, with the interface and most critical capabilities evolving every few months.
- Your Coding Agent Just Got Cloned And Your Brain Isn't Ready - Rustin Banks, Google Jules
This talk introduces Jules, an asynchronous AI coding agent designed to run in the background and handle tasks in parallel, freeing developers to focus on creative coding. The core idea is to shift from serial task execution to a parallel workflow, leveraging AI for both the beginning and end of the software development lifecycle, from task creation to code merging and testing.
- Latent Space Paper Club: AIEWF Special Edition (Test of Time, DeepSeek R1/V3) — VIbhu Sapra
This talk reviews recent advancements in AI research, focusing on the DeepSeek models and the evolution of training methodologies. It introduces a new "Test of Time" paper club initiative aimed at systematically covering foundational AI concepts. The discussion highlights how reinforcement learning and extended inference time are enabling models to develop advanced reasoning capabilities, leading to significant performance improvements.
- Human seeded Evals — Samuel Colvin, Pydantic
This talk focuses on building AI applications more quickly and safely, emphasizing the importance of type safety in refactoring and avoiding bugs. It introduces techniques for developing AI applications, particularly highlighting the use of Pydantic AI for extracting structured data and managing agentic loops. The discussion also touches upon the challenges of determining when an agent loop should terminate and the benefits of using validation errors to guide model retries.
- Building AI Products That Actually Work — Ben Hylak (Raindrop), Sid Bendre (Oleve)
This talk focuses on the practical aspects of building AI products that are reliable and effective, moving beyond theoretical discussions of evaluations. It emphasizes that while AI technology is advancing, the core challenge lies in effective communication and managing the inherent complexity and undefined behaviors of AI systems. The speakers advocate for an iterative approach to product development, where continuous refinement based on real-world user data and signals is crucial for success.
- Rise of the AI Architect — Clay Bavor, Cofounder, Sierra w/ Alessio Fanelli
This talk introduces the concept of the AI Architect, a role emerging in the AI era analogous to the webmaster of the internet's early days. AI Architects are responsible for understanding the technology, shaping the agent's brand and user experience, and driving business outcomes. The discussion highlights the complexity of building AI agents, emphasizing that successful strategies involve a spirit of exploration, a focus on solving real problems, and a willingness to rearchitect existing teams and processes.
- AI That Pays: Lessons from Revenue Cycle — Nathan Wan, Ensemble Health
This talk explores the application of AI in healthcare's revenue cycle management (RCM), an often overlooked but critical financial process. The core thesis is that significant inefficiencies and lost revenue in healthcare stem from manual, complex, and error-prone administrative tasks, rather than clinical costs. AI offers a powerful opportunity to disrupt this area by reducing friction, preventing errors, and shifting resources towards more productive activities like patient care.
- Structuring a modern AI team — Denys Linkov, Wisedocs
This talk emphasizes that building a successful AI team hinges on understanding specific organizational needs and problems rather than solely chasing the latest AI research or hiring specialized roles prematurely. The core thesis is that technology adoption is often slow, and the effectiveness of AI solutions depends more on how they are integrated and utilized within a business context than on the cutting-edge nature of the technology itself. The speaker advocates for a pragmatic approach to team structure, prioritizing domain knowledge and business acumen alongside technical skills.
- The Rise of Open Models in the Enterprise — Amir Haghighat, Baseten
Enterprises are increasingly adopting AI, moving beyond simply purchasing verticalized solutions to building their own AI capabilities. While initial adoption often leverages closed-source models from providers like OpenAI and Anthropic, several factors are driving a shift towards open-source models. These include the need for specialized quality, reduced latency, better unit economics for agentic use cases, and a desire for competitive differentiation.
- Mentoring the Machine — Eric Hou, Augment Code
This talk explores how AI agents can be integrated into software development workflows, shifting the engineer's role from direct implementation to mentorship and orchestration. The core thesis is that by treating AI like junior engineers—providing context, guidance, and structured environments—teams can overcome the limitations of AI and unlock significant productivity gains. This approach transforms chaotic, interruption-filled days into more focused and efficient work, ultimately making software creation more of a science.
- Building Applications with AI Agents — Michael Albada, Microsoft
This talk explores the development of AI agents, highlighting both their promise and the significant obstacles encountered in production. It defines agents as entities capable of reasoning, acting, communicating, and adapting. The presentation emphasizes that agency exists on a spectrum and should be a tool to enhance effectiveness, not a goal in itself, cautioning against systems with low efficacy despite high agency.
- AX is the only Experience that Matters - Ivan Burazin, Daytona
The core thesis is that the future of software development and tooling lies in building for agents, not humans. As agents become the primary users of digital environments, the concept of "agent experience" (AX) will supersede traditional user and developer experiences. Tools that require human intervention will become obsolete, necessitating a shift towards building systems that agents can operate within autonomously.
- How to build Enterprise Aware Agents - Chau Tran, Glean
This talk explores the distinction between AI workflows and agents, highlighting their respective strengths and weaknesses. It proposes that agents can be viewed as tools that generate workflows, suggesting a powerful synergy between the two. The discussion also delves into methods for building enterprise-aware agents, emphasizing the importance of incorporating business-specific knowledge and processes.
- Monetizing AI — Alvaro Morales, Orb
This talk addresses the complexities of monetizing AI products, emphasizing the need for strategic pricing beyond intuition. It highlights the rapid evolution of AI technology, its impact on margins, and customer demand for clear ROI, all of which contribute to unique pricing challenges. The speaker proposes moving away from pricing based on vague notions towards data-driven frameworks and tools for effective monetization.
- Does AI Actually Boost Developer Productivity? (100k Devs Study) - Yegor Denisov-Blanch, Stanford
- How agents will unlock the $500B promise of AI - Donald Hruska, Retool
The talk argues that AI agents are poised to unlock significant economic value, moving beyond simple chatbots and code generation to integrate with real-world production systems. While building basic agents is becoming increasingly accessible, deploying them effectively within enterprises requires careful consideration of security, cost, and integration challenges. The future likely involves a mix of custom-built agents for core business logic and platform-hosted agents for broader workflows.
- How Intuit uses LLMs to explain taxes to millions of taxpayers - Jaspreet Singh, Intuit
Intuit leverages Large Language Models (LLMs) within TurboTax to enhance taxpayer understanding of their tax situations and potential deductions. The company has developed a proprietary generative OS called Geno, designed for large-scale, secure, and reliable use in the highly regulated tax industry. This platform integrates various components, including UI elements and an orchestrator, to manage different LLM solutions and deliver a cohesive user experience.
- 3 ingredients for building reliable enterprise agents - Harrison Chase, LangChain/LangGraph
Building reliable enterprise agents hinges on three core ingredients: maximizing value when the agent is correct, minimizing cost when it is wrong, and increasing the probability of success. This approach provides a foundational framework for developing agents that are more likely to be adopted and effective within an enterprise setting. The future vision involves numerous agents operating autonomously, coordinating tasks, and requiring careful management.
- From Hype to Habit: How We’re Building an AI-First SaaS Company—While Still Shipping the Roadmap
This talk explores the complexities of transitioning a SaaS company to an AI-first model, emphasizing that it's an evolutionary journey rather than a binary switch. It proposes a framework focusing on strategy, ways of working, and people to navigate this transformation. The core thesis is that becoming AI-first requires reimagining product development, embracing ambiguity, and fostering a cultural shift across the entire organization, not just within AI teams.
- Machines of Buying and Selling Grace - Adam Behrens, New Generation
This talk explores the evolution of the store concept in the age of AI, moving from physical locations to digitized online presences, and now towards an agentic future. The core thesis is that AI will digitize not just merchandise and distribution, but also the participants and their interactions, leading to dynamic, real-time, and generative interfaces for both human and agentic consumers. This shift aims to enhance transactions by moving beyond static websites to sophisticated agent-to-agent or agent-to-API interactions.
- How to Build Planning Agents without losing control - Yogendra Miraje, Factset
This talk explores how to build planning agents for AI applications, focusing on agentic workflows that combine the flexibility of agents with the reliability of traditional workflows. The core thesis is that by moving beyond simple reactive agents to proactive ones, and by employing a "planning by sub-goal division" design pattern, developers can create more controlled, scalable, and enterprise-ready AI systems. The approach emphasizes leveraging existing enterprise microservices and robust evaluation frameworks.
- Building Agents (the hard parts!) - Rita Kozlov, Cloudflare
This talk focuses on the practical challenges and components involved in building AI agents. It emphasizes that moving beyond simple AI augmentation to true automation requires a structured approach. The core idea is that agents need a client interface, an AI reasoning engine, workflow management for execution, and access to tools to perform actions.
- POC to PROD: Hard Lessons from 200+ Enterprise GenAI Deployments - Randall Hunt, Caylent
This talk shares hard-earned lessons from over 200 enterprise Generative AI deployments, emphasizing that AI is not a universal solution. It highlights the importance of understanding customer needs, robust evaluation, and efficient architecture over simply adopting the latest AI trends. The core thesis is that successful AI implementation requires a deep understanding of inputs, outputs, and user behavior, rather than relying solely on advanced models or prompt engineering.
- Build Dynamic Products, and Stop the AI Sideshow — Eliza Cabrera (Workday) + Jeremy Silva (Freeplay)
This talk argues for building dynamic, deeply integrated AI products rather than AI "sideshows" that are bolted onto existing systems. The core thesis is that companies should move beyond using AI primarily to demonstrate technological capability and instead focus on solving customer problems by integrating AI strategically into their core product development. This approach requires aligning AI and product strategies, teams, and roadmaps, and embracing a crawl, walk, run methodology for iterative development.
- The Billable Hour is Dead; Long Live the Billable Hour — Kevin Madura + Mo Bhasin, Alix Partners
This talk explores how AI, particularly generative AI, is reshaping knowledge work and professional services. The core thesis is that AI can significantly compress the initial data ingestion and analysis phases of engagements, freeing up experienced professionals for higher-value strategic tasks. This shift moves beyond traditional leverage models to a future where AI replicates and scales the expertise of senior individuals.
- From Copilot to Colleague: Trustworthy Agents for High-Stakes - Joel Hron, CTO Thomson Reuters
This talk explores the evolution of AI assistants from being merely helpful to becoming productive agents capable of making judgments and decisions. It emphasizes that agency in AI is not a binary state but a spectrum that can be tuned based on use case, risk tolerance, and user expectations. The presentation highlights the challenges and lessons learned in building trustworthy AI systems for high-stakes professional environments, particularly within legal, tax, and global trade sectors.
- How to Hire AI Engineers when EVERYONE is cheating with AI — Beth Glenfield, DevDay
The current technical hiring process is fundamentally broken due to the widespread use of AI, making traditional methods like LeetCode puzzles ineffective. Companies are struggling to identify genuine talent as AI assistants significantly boost candidate performance on these outdated assessments. This shift necessitates a reimagining of how to evaluate candidates, focusing on skills relevant to modern AI development and collaboration.
- Stateful environments for vertical agents — Josh Purtell, Synth Labs
This talk introduces the concept of stateful environments for AI agents, particularly for vertical applications. The core idea is to externalize and containerize the logic of a task or application, creating a distinct workspace that the agent can interact with. This separation allows for more robust agent development, easier updates, and advanced capabilities like multi-agent collaboration and state rollback.
- Books reimagined: AI to create new experiences for things you know — Lukasz Gandecki, TheBrain.pro
This talk explores reimagining books through AI to create novel user experiences. The core idea is to augment familiar content with dynamic, AI-generated elements like contextual summaries, character visualizations, and scene-appropriate music, transforming passive reading into an interactive, immersive journey. The approach emphasizes integrating AI subtly to enhance, rather than replace, the human element in content creation.
- AI powered entomology: Lessons from millions of AI code reviews — Tomas Reimers, Graphite
This talk explores the challenges and opportunities of using AI, specifically Large Language Models (LLMs), for code review. It highlights that while LLMs can identify bugs, their effectiveness is limited by the type of feedback they provide and the user's receptiveness to it. The core thesis is that by understanding the different categories of bugs and user preferences, AI can be trained to provide more valuable and actionable code review comments.
- Critical AI Inference your CIO can Trust — Sahil Yadav, Hariharan Ganesan, Telemetrak
This talk addresses the critical need for trustworthy AI, especially in mission-critical applications where AI inferences directly impact business decisions and financial outcomes. It highlights a significant gap between AI adoption and AI governance, leading to potential silent failures with substantial financial and operational consequences. The presentation introduces a framework for building and scaling AI systems that instill confidence through explainability, traceability, and robust guardrails.
- How to run Evals at Scale: Thinking beyond Accuracy or Similarity — Muktesh Mishra, Adobe
This talk emphasizes the critical role of evaluations (evals) in developing robust AI applications, moving beyond simple accuracy or similarity metrics. It highlights that effective evals are fundamental for aligning applications with system goals, ensuring continuous improvement, and building user trust. The core thesis is that a strategic, data-driven approach to evals, tailored to specific use cases, is essential for scaling AI development.
- Continuous Profiling for GPUs — Matthias Loibl, Polar Signals
This talk introduces continuous profiling for GPUs, a method to monitor and analyze GPU performance. It highlights the importance of profiling for improving application performance and reducing operational costs by enabling more efficient resource utilization. The approach leverages Linux eBPF for low-overhead, always-on profiling in production environments without requiring application instrumentation.
- Top Ten Challenges to Reach AGI — Stephen Chin, Andreas Kollegger
This talk explores ten potential challenges on the path to Artificial General Intelligence (AGI), drawing parallels with science fiction concepts. The speakers suggest that by examining these fictional scenarios, we can better understand and prepare for the real-world implications and ethical considerations of developing advanced AI. The core thesis is that a proactive, science-fiction-informed approach is crucial for responsible AGI development.
- Practical GraphRAG: Making LLMs smarter with Knowledge Graphs — Michael, Jesus, and Stephen, Neo4j
This talk introduces Graph RAG, a method for enhancing Large Language Models (LLMs) by integrating knowledge graphs. It addresses the limitations of standard LLMs, such as a lack of domain-specific knowledge, tendency to hallucinate, and difficulty in explaining answers. Graph RAG aims to provide more accurate, contextual, and explainable responses by leveraging structured data within knowledge graphs, moving beyond the limitations of purely vector-based retrieval.
- Knowledge Graphs in Litigation Agents — Tom Smoker, WhyHow
This talk explores the application of knowledge graphs and multi-agent systems within the legal industry, specifically for identifying and supporting class-action lawsuits. It details how unstructured data from web scraping and legal discovery can be transformed into structured graphs to aid legal professionals in case research and analysis. The core thesis is that by leveraging these technologies, complex legal data can be made more accessible, accurate, and actionable for lawyers.
- When Vectors Break Down: Graph-Based RAG for Dense Enterprise Knowledge - Sam Julien, Writer
This talk explores the limitations of traditional vector-based Retrieval Augmented Generation (RAG) for complex enterprise knowledge and proposes a graph-based RAG approach. It highlights how preserving relationships within data through knowledge graphs, combined with advanced techniques like fusion-decoder, significantly improves accuracy and reduces hallucinations in AI applications, especially for dense, specialized datasets.
- Stop Using RAG as Memory — Daniel Chalef, Zep
This talk argues against using Retrieval-Augmented Generation (RAG) as a generic memory solution for AI agents. Instead, it proposes modeling memory after specific business domains to create more cogent and capable memory systems. The core thesis is that semantic similarity alone is insufficient for accurate memory recall, and domain-aware memory structures are necessary.
- HybridRAG: A Fusion of Graph and Vector Retrieval - Mitesh Patel, NVIDIA
This talk introduces HybridRAG, a system that combines graph and vector retrieval for enhanced information retrieval. It highlights the advantages of knowledge graphs in capturing detailed relationships between entities, offering a more comprehensive view than purely semantic approaches. The system is broken down into data processing, graph and vector database creation, and inferencing, with a focus on optimizing retrieval strategies and evaluating performance.
- tldraw.computer - Steve Ruiz, tldraw
This talk introduces tldraw.computer, a platform that leverages AI for creative and functional applications. It showcases how a hackable canvas, built with standard web technologies, can serve as a foundation for AI-driven tools. The core thesis is that by treating AI models as programmable components within a visual interface, users can build complex applications, iterate rapidly, and achieve novel results, blurring the lines between design, coding, and AI execution.
- Excalidraw: AI and Human Whiteboarding Partnership - Christopher Chedeau
This talk explores the evolution of whiteboarding tools, drawing parallels between the transition from physical to digital whiteboards and the current integration of AI. It emphasizes that successful AI integrations should enhance, not merely replicate, existing functionalities, focusing on AI-native interactions rather than simply adding AI features. The core thesis is that AI and human collaboration in digital whiteboarding is entering a new phase, moving beyond basic AI augmentation to a more symbiotic partnership.
- The Bitter Layout or: How I Learned to Love the Model Picker — Maximillian Piras, Yutori
This talk explores the evolution of AI user interfaces, arguing that the current dominant chat-based layout, while functional, is a "bitter layout" that prioritizes adaptability to new models over optimal user experience. It suggests that as models become less commoditized, the interface itself becomes the commodity, necessitating a shift in design philosophy towards goal-oriented and constraint-based approaches, akin to gardening rather than construction.
- UX Design Principles for Semi Autonomous Multi Agent Systems — Victor Dibia, Microsoft
This talk explores the design principles for creating effective user experiences in semi-autonomous multi-agent systems. It emphasizes that while multi-agent systems offer powerful capabilities, their complexity requires careful consideration of user interaction, observability, and control. The presentation highlights a practical approach to building such systems, starting with defining the goal and tools before agent development, and advocates for an evaluation-driven design process.
- CIAM for AI: Authn/Authz for Agents — Michael Grinich, CEO of WorkOS
This talk addresses the critical need for robust identity, authentication, and authorization solutions for AI agents, a rapidly growing segment in the B2B SaaS landscape. As agents become more integrated into workflows, they require first-class identity support, distinct from traditional bots or human users. The presentation emphasizes the urgency of developing new standards for agent identity to ensure user safety and enable scalable adoption of AI technologies.
- Good design hasn’t changed with AI — John Pham, SF Compute
This talk argues that fundamental principles of good design remain unchanged despite advancements in AI tools. Design is defined not by aesthetics but by the entirety of a user's experience across all touchpoints, influencing how they feel about a product. In the current AI landscape, where feature parity is no longer a differentiator, design becomes the key element for creating unique and memorable user experiences.
- Building Effective Voice Agents — Toki Sherbakov + Anoop Kotha, OpenAI
This talk explores the advancements and practical applications of building effective voice agents, moving beyond traditional text-based AI. It highlights the emergence of speech-to-speech technology as a key component of the multimodal AI era, emphasizing that current models are now fast, expressive, and accurate enough for scalable production use. The discussion covers architectural patterns, key trade-offs in development, and best practices for creating robust and engaging audio experiences.
- What every AI engineer needs to know about GPUs — Charles Frye, Modal
This talk explains why AI engineers need to understand GPUs, shifting focus from API-based development to leveraging hardware capabilities. It draws an analogy to database usage, where developers don't build databases but must understand how to query them effectively. Similarly, AI engineers will increasingly need to understand GPU architecture, particularly tensor cores, to optimize performance for tasks like language model inference.
- Robots as professional Chefs - Nikhil Abraham, CloudChef
CloudChef has developed culinary intelligent robots capable of acting, sensing, and reasoning like human chefs. These robots utilize a combination of foundation models, teleoperation for edge cases, and specialized thermal and visual embeddings for food understanding. The system is designed to adapt to new kitchens, learn recipes from single demonstrations, and handle variations in ingredients and appliances, aiming to make high-quality food affordable through automation.
- [Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han
This talk delves into advanced AI training techniques, focusing on reinforcement learning (RL), quantization, and agent development. It explores the evolution of large language models, from early open-source efforts spurred by leaks to current sophisticated training methodologies. The discussion highlights the critical role of fine-tuning stages, including supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF), in transforming base models into capable conversational agents.
- A Taxonomy for Next-gen Reasoning — Nathan Lambert, Allen Institute (AI2) & Interconnects.ai
This talk explores the evolution of AI reasoning capabilities, moving beyond benchmark scores to focus on practical applications and future development. It posits that while current models excel at specific skills like math and code, the next frontier lies in planning, strategy, and abstraction. Achieving this requires a shift in how models are trained, emphasizing human effort in developing new algorithmic methods and data acquisition strategies.
- How to Train Your Agent: Building Reliable Agents with RL — Kyle Corbitt, OpenPipe
This talk details the process of building a reliable AI agent using reinforcement learning (RL), focusing on practical lessons learned from the ART E project, an email assistant. The core thesis is that while RL can significantly improve agent performance beyond prompted models, it's crucial to start with a strong prompted baseline, carefully design the training environment and reward functions, and be vigilant against reward hacking. The project demonstrates how RL can lead to substantial gains in accuracy, cost reduction, and latency improvements for specialized tasks.
- OpenThoughts: Data Recipes for Reasoning Models — Ryan Marten, Bespoke Labs
This talk introduces OpenThoughts 3, a project focused on creating high-quality, open-source reasoning datasets. The core thesis is that while models have shown significant performance gains on reasoning benchmarks, the underlying data recipes for achieving this are often undisclosed. OpenThoughts aims to fill this gap by providing a systematic approach and publicly available datasets to train more capable reasoning models, emphasizing that supervised fine-tuning (SFT) on well-crafted data can be highly effective.
- Google Photos Magic Editor: GenAI Under the Hood of a Billion-User App - Kelvin Ma, Google Photos
This talk details the engineering journey behind Google Photos' Magic Editor, highlighting the transition from on-device ML models to server-side generative AI. It emphasizes the challenges and learnings in integrating cutting-edge AI into a widely used application, focusing on user experience, model efficiency, and the inherent unpredictability of AI systems. The core thesis is that while AI offers powerful new capabilities, successful product integration requires rigorous engineering to manage its randomness and deliver reliable, user-friendly features.
- Dream Machine: Scaling to 1m users in 4 days — Keegan McCallum, Luma AI
This talk details Luma AI's experience scaling their Dream Machine video generation model to one million users in four days. It covers the immense infrastructure challenges faced during the launch, including the need to rapidly scale GPU resources from an initial 500 H100s to 5,000. The presentation also discusses the evolution of their serving stack from a brittle, coupled system to a more robust, decoupled architecture built on PyTorch, and the strategies implemented for fair job scheduling and efficient model management.
- ComfyUI Full Workshop — first workshop from ComfyAnonymous himself!
ComfyUI is an open-source, node-based design canvas for generative AI, supporting multimodal creative applications including image, video, audio, 3D, and text. It aims to provide maximal control by allowing users to interact with models at a granular level, going beyond simple prompts to manipulate elements like depth maps and masks. The platform emphasizes extensibility through community-developed custom nodes and a unique workflow sharing mechanism where generated images contain embedded metadata of the original workflow.
- Design like Karpathy is watching — Zeke Sikelianos, Replicate
This talk explores how to design AI products and services with language models as a primary audience, drawing lessons from Andrej Karpathy's experience building the Menuguen app. It emphasizes the shift towards LLMs consuming structured data like markdown and API schemas, rather than just human-readable web pages. The core thesis is that embracing these formats and improving developer experience is crucial for building successful AI-powered applications.
- On Curiosity — Sharif Shameem, Lexica
This talk argues that curiosity is the primary driving force for innovation, enabling the translation of future ideas into present realities. The speaker emphasizes that building and sharing demos is the most effective way to explore the potential of AI models, as these models are not fully understood even by their creators. Demos serve as a crucial tool for AI engineers, akin to excavators uncovering hidden capabilities, with curiosity acting as the guide.
- Real world MCPs in GitHub Copilot Agent Mode — Jon Peck, Microsoft
This talk introduces GitHub Copilot's Agent Mode, an advanced feature designed for completing moderately complex tasks autonomously. It moves beyond simple code completion and chat interactions to enable deep, iterative engagement with an AI agent. The presentation highlights how Agent Mode can build entire applications from a readme file or perform significant refactoring, with the developer providing permission for actions like terminal interactions.
- The rise of the agentic economy on the shoulders of MCP — Jan Curn, Apify
This talk explores how general intelligence in computing systems may emerge from the interaction of multiple agents, similar to how intelligence arises in biological systems or markets. The speaker posits that the Multi-agent Conversation Protocol (MCP) is a crucial development enabling agents to communicate and form an agentic mesh, facilitating the rise of an agentic economy where agents can discover and utilize services.
- Full Spec MCP: Hidden Capabilities of the MCP spec — Harald Kirschner, Microsoft/VSCode
This talk explores the underutilized capabilities of the Model Contract Protocol (MCP) beyond basic tool calling. It highlights how embracing the full MCP specification can unlock richer, stateful interactions for AI agents. The presentation emphasizes that while many are shipping products quickly using only tools, the spec offers advanced features like dynamic discovery, resources, and sampling that enable more sophisticated agent behavior and development experiences.
- Shipping an Enterprise Voice AI Agent in 100 Days - Peter Bar, Intercom Fin
This talk details the process of developing and shipping Finn Voice, an AI-powered voice agent for customer support, within a 100-day timeframe. The core thesis is that voice AI represents a significant frontier in customer service, offering substantial benefits in cost savings, availability, and user experience compared to traditional phone support. The presentation emphasizes that successful deployment requires more than just advanced AI models; it necessitates a product-centric approach that considers use cases, conversation design, integration with existing workflows, and building user trust.
- The State of Generative Media - Gorkem Yurtseven, FAL
The generative media landscape is rapidly evolving, with significant advancements in video and image generation models. The marginal cost of creation is approaching zero, leading to transformative impacts across various industries like advertising, e-commerce, and entertainment. While image generation has seen substantial progress, video generation is poised for exponential growth, potentially dwarfing the image market in scale and utility.
- Teaching Gemini to Speak YouTube: Adapting LLMs for Video Recommendations to 2B+DAU - Devansh Tandon
This talk details the adaptation of large language models (LLMs), specifically Gemini, for YouTube's recommendation system, aiming to improve user engagement for over 2 billion daily active users. The core thesis is that LLMs can revolutionize recommendation systems, potentially surpassing search applications in consumer impact. The approach involves creating a domain-specific language for videos and training a bilingual model capable of understanding both natural language and this new video token language.
- Transforming search and discovery using LLMs — Tejaswi & Vinesh, Instacart
This talk details how Instacart is leveraging Large Language Models (LLMs) to significantly enhance its search and discovery functionalities. The core thesis is that LLMs can overcome limitations of traditional search models, particularly for tail queries and new item discovery, by better understanding user intent and context, ultimately leading to improved user experience and business metrics.
- Netflix's Big Bet: One model to rule recommendations: Yesu Feng, Netflix
Netflix is betting on a single foundation model, based on transformer architecture, to handle all recommendation use cases. This approach aims to improve personalization by scaling semi-supervised learning and create high leverage by integrating the model across all systems, simultaneously enhancing downstream applications. The model learns user representations and incorporates rich event data, addressing challenges like cold-start problems and enabling faster innovation.
- 360Brew: LLM-based Personalized Ranking and Recommendation - Hamed and Maziar, LinkedIn AI
This talk details LinkedIn's journey in developing and deploying 360Brew, a large language model-based system for personalized ranking and recommendations. The core thesis is that a single, holistic LLM can replace disjointed, task-specific models, offering zero-shot capabilities, in-context learning, and instruction following to understand user journeys and serve diverse personalization needs on the LinkedIn platform.
- What We Learned from Using LLMs in Pinterest — Mukuntha Narayanan, Han Wang, Pinterest
This talk details Pinterest's experience integrating Large Language Models (LLMs) into their search relevance system. The core thesis is that LLMs significantly improve relevance prediction, and that techniques like knowledge distillation are crucial for productionizing these models at scale. The presentation also highlights the value of synthetic captions and user engagement data as content annotations.
- Measuring AGI: Interactive Reasoning Benchmarks for ARC-AGI-3 — Greg Kamradt, ARC Prize Foundation
This talk introduces ARC-AGI-3, a new benchmark designed to measure artificial general intelligence (AGI) by focusing on interactive reasoning and skill acquisition efficiency. The benchmark aims to create problems that are solvable by humans but challenging for current AI, thereby guiding AI research and development towards human-level intelligence. It moves beyond single-turn, static benchmarks to simulate more realistic, open-world exploration and learning scenarios.
- RL for Autonomous Coding — Aakanksha Chowdhery, Reflection.ai
This talk explores the progression of large language models (LLMs) from early scaling laws to the current frontier of autonomous coding. It highlights how techniques like chain-of-thought prompting and reinforcement learning with human feedback have improved LLM capabilities. The core thesis is that reinforcement learning, particularly in domains with automated verification like coding, represents the next significant step in scaling LLM performance and building more intelligent systems.
- Recsys Keynote: Improving Recommendation Systems & Search in the Age of LLMs - Eugene Yan, Amazon
This talk explores the integration of Large Language Models (LLMs) with recommendation systems and search functionalities. It proposes three key areas for advancement: semantic IDs to better represent item content and address cold-start problems, data augmentation using LLMs to improve data quality and scale for training, and unified models to streamline complex recommendation infrastructures.
- Benchmarks Are Memes: How What We Measure Shapes AI—and Us - Alex Duffy, Every.to
This talk argues that AI benchmarks function as memes, spreading ideas and shaping the development of AI models. As benchmarks become popular, models are trained and tested on them until saturation occurs. This cycle presents an opportunity for individuals to create new benchmarks that can influence the future direction of AI development, emphasizing the importance of human-defined goals and values in the process.
- Small AI Teams with Huge Impact — Vik Paruchuri, Datalab
This talk challenges the conventional Silicon Valley belief that increasing headcount directly correlates with greater productivity. The speaker argues that smaller, highly capable teams of generalists can achieve significantly more by focusing on core competencies, leveraging AI for low-leverage tasks, and maintaining a high degree of trust and customer focus. This approach prioritizes efficient collaboration and rapid feedback loops over bureaucratic processes often found in larger organizations.
- Rethinking Team Building: how a 30-person Startup serves 50 Million Users — Grant Lee, Gamma
This talk challenges traditional startup team-building and scaling strategies, advocating for a leaner, more agile approach. The speaker, Grant Lee of Gamma, proposes that instead of hierarchical growth, companies can achieve significant user reach with a small, highly capable team by embracing generalists, player-coaches, and a strong, shared culture. This model aims to foster innovation and productivity by prioritizing content and user needs over design complexities.
- Building a 10 person unicorn - Max Brodeur-Urbas, Gumloop
This talk focuses on how to scale a startup team to a "unicorn" status with fewer than ten people, emphasizing culture, hiring, and operational efficiency. The core thesis is that by being extremely selective in hiring, fostering a product-led approach, and automating extensively, small teams can achieve rapid growth and significant impact without ballooning in size. The speaker shares insights from their experience founding Gumloop, a workflow automation tool.
- Using OSS models to build AI apps with millions of users — Hassan El Mghari
This talk explores how to leverage open-source AI models to build applications capable of reaching millions of users. The speaker emphasizes that it's an opportune time for builders due to lowered barriers in development tools and the rapid release of groundbreaking AI models. The core thesis is that by focusing on simplicity, iterating quickly, and embracing open-source principles, developers can successfully create and scale AI-powered applications.
- Bolt.new: How we scaled $0-20m ARR in 60 days, with 15 people — Eric Simons, Bolt
This talk details Bolt.new's rapid growth from $0 to $20 million in Annual Recurring Revenue (ARR) within 60 days with a small team. The core thesis emphasizes the power of a lean, highly aligned team with deep context, enabling rapid iteration and scaling, even with an initial MVP product. It highlights the importance of independent decision-making and fostering user love through direct engagement and community building.
- Prompt Engineering and AI Red Teaming — Sander Schulhoff, HackAPrompt/LearnPrompting
This talk explores the evolving landscape of prompt engineering and AI red teaming, emphasizing that prompt engineering remains a critical skill despite claims of its demise. The speaker highlights the inherent difficulty in securing generative AI systems, drawing parallels and distinctions with classical cybersecurity. The presentation delves into various prompting techniques, the challenges of AI security, and the practical implications of prompt injection and jailbreaking.
- Survive the AI Knife Fight: Building Products That Win — Brian Balfour, Reforge
In a rapidly evolving AI landscape, building successful products requires a strategic approach beyond simply integrating AI features. The core thesis is that competitive advantage stems not from the AI technology itself, but from the unique combination of a product's proprietary data, its specific functionality, and a deep understanding of unmet customer needs. This talk emphasizes treating AI as modular components, or "Lego blocks," that can be assembled to create differentiated offerings.
- Automating Escrow with USDC and AI - Corey Cooper, Circle
This talk explores the integration of AI with programmable money, specifically Circle's USDC, to automate complex financial workflows like escrow. The core idea is to leverage AI for verifying conditions and USDC for instant, borderless settlement, creating a more efficient and programmable financial system. The presentation introduces Circle's developer tooling and demonstrates an open-sourced escrow agent application.
- How LLMs work for Web Devs: GPT in 600 lines of Vanilla JS - Ishan Anand
This talk breaks down how Large Language Models (LLMs) like GPT-2 work, demystifying them for web developers by implementing a core GPT-2 model in approximately 600 lines of vanilla JavaScript. The presentation emphasizes understanding the underlying mechanics rather than requiring deep ML expertise, using analogies and code walkthroughs to make complex concepts accessible. The goal is to transform the perception of LLMs from magic to understandable machinery.
- [Workshop] AI Pipelines and Agents in Pure TypeScript with Mastra.ai — Nick Nisi, Zack Proser
This workshop introduces Mastra.ai, a framework for building AI applications in TypeScript. It demonstrates how to create agentic workflows, focusing on composable pipelines, tools, and agents. The session highlights using TypeScript for production-ready AI applications, emphasizing patterns applicable to deploying AI solutions.
- AI Engineering with the Google Gemini 2.5 Model Family - Philipp Schmid, Google DeepMind
This talk introduces the Google Gemini 2.5 model family, focusing on practical applications for AI engineers and builders. It highlights the multimodal capabilities of Gemini 2.5 Pro and Flash, demonstrating how to leverage these models for text generation, image and audio understanding, function calling, and integrating with external tools via MCP servers. The session emphasizes hands-on learning through a workshop format, encouraging attendees to experiment with the models and SDK.
- The New Code — Sean Grove, OpenAI
This talk argues that structured communication, embodied in specifications, is more valuable than code itself for AI development. Specifications serve as the primary artifact for aligning human intent and values, enabling clearer communication, better testing, and more robust AI systems. As AI models advance, the ability to write effective specifications will become the most critical skill for programmers.
- Production software keeps breaking and it will only get worse — Anish Agarwal, Traversal.ai
The talk argues that as AI tools increasingly automate code development, the complexity of system design and production troubleshooting will grow, potentially leading engineers to spend more time on on-call duties. The core thesis is that current approaches to AI-assisted troubleshooting are insufficient and a new, more integrated method is needed to handle the escalating complexity of production incidents.
- Thinking Deeper in Gemini — Jack Rae, Google DeepMind
This talk explores the concept of "thinking" within Gemini, a new paradigm for AI models that allows for iterative computation and deeper reasoning before generating a final response. The core thesis is that by enabling models to spend more test-time compute on complex problems, we can unlock significant advancements in AI capabilities, moving beyond immediate response generation towards more profound problem-solving. This approach addresses bottlenecks in current AI by allowing for dynamic allocation of computational resources based on task difficulty.
- A year of Gemini progress + what comes next — Logan Kilpatrick, Google DeepMind
Logan Kilpatrick of Google DeepMind discussed a year of progress with Gemini models and outlined future developments. The talk highlighted the release of a new Gemini 2.5 Pro model, emphasizing its improved performance across benchmarks and its role as a turning point for the Gemini family. Kilpatrick also touched upon the organizational shifts within Google that have integrated research and product teams, enabling faster delivery of AI capabilities to both consumers and developers.
- 2025 in LLMs so far, illustrated by Pelicans on Bicycles — Simon Willison
The AI landscape has accelerated dramatically over the last six months, with numerous significant model releases making it challenging to assess their quality. Traditional benchmarks and leaderboards are losing credibility, prompting a shift towards more practical, self-devised evaluation methods. The speaker uses a unique benchmark involving generating SVGs of pelicans riding bicycles to assess model capabilities, highlighting the rapid progress in model performance and the increasing accessibility of powerful AI on consumer hardware.
- Trends Across the AI Frontier — George Cameron, ArtificialAnalysis.ai
This talk explores multiple frontiers in AI beyond just raw intelligence, emphasizing the trade-offs involved in accessing advanced AI capabilities. It highlights that the most intelligent models are not always the most suitable due to implications for cost, latency, and verbosity. The presentation uses benchmarking data to illustrate these trade-offs across reasoning capabilities, open-weight models, cost-effectiveness, and speed.
- Training Agentic Reasoners — Will Brown, Prime Intellect
This talk argues that reasoning and agents are fundamentally the same concept, with reinforcement learning (RL) being the key to developing powerful agentic systems. The speaker posits that traditional approaches to building agents often involve manual, iterative tuning that mirrors RL processes. By framing agent development through the lens of RL, developers can leverage established algorithms and techniques to create more robust and capable agents, especially for complex, multi-turn tasks.
- New York Times' Connections: A Case Study on NLP in Word Games — Shafik Quoraishee, NYT Games
This talk explores the New York Times Connections word game as a case study for evaluating AI's abstract reasoning capabilities. The presenter, a game developer at the NYT Games team, details independent research into how AI can approach solving the game, which involves grouping 16 words into four sets of four related terms. The research investigates the game's potential as a benchmarking tool, highlighting how its intentional decoys and need for explainable reasoning challenge current AI models.
- Claude Code & the evolution of agentic coding — Boris Cherny, Anthropic
The talk by Boris Cherny from Anthropic discusses the rapid evolution of AI models in coding and the challenges in developing user experiences (UX) that keep pace. It highlights how programming itself has evolved through layers of abstraction and interface changes, from punch cards to modern IDEs. The core thesis is that while models are advancing exponentially, the product development for these AI coding tools is still in its early stages, with a focus on building unopinionated, minimal-viable products to learn and adapt to the unknown optimal UX.
- 12-Factor Agents: Patterns of reliable LLM applications — Dex Horthy, HumanLayer
This talk proposes the "12-Factor Agents" pattern, drawing parallels to the 12-Factor App methodology for building reliable software. The core thesis is that building robust LLM applications requires applying established software engineering principles, focusing on modularity, control flow, and state management rather than solely relying on agentic capabilities. The approach emphasizes treating agents as software components that can be engineered for reliability and maintainability.
- MCP Is Not Good Yet — David Cramer, Sentry
David Cramer of Sentry discusses the current state of MCP (Model Communication Protocol), framing it as a pluggable architecture for agents. He emphasizes that while the concept is powerful, current implementations are often rough and require significant design effort. Cramer suggests that MCP's true value lies in its ability to integrate services into agent workflows, particularly for B2B SaaS companies, by providing context that LLMs can reason about.
- Your Personal Open-Source Humanoid Robot for $8,999 — JX Mo, K-Scale Labs
K-Scale Labs is developing open-source humanoid robots, the Kbot and Zbot, aimed at making advanced robotics accessible to developers and researchers. Their core mission is to democratize robotics by open-sourcing the entire hardware and software stack, enabling widespread adoption and innovation. The Kbot is a full-sized humanoid robot designed for affordability and modularity, while the Zbot is a smaller, more accessible option derived from a hackathon project.
- The Build-Operate Divide: Bridging Product Vision and AI Operational Reality
This talk addresses the common challenge of AI product concepts failing to reach their full potential due to operational difficulties. It emphasizes that bridging the gap between product vision and AI operational reality requires a deep understanding of how to deliver quality through rigorous evaluation, human review, and strategic team building. The core thesis is that successful AI product development hinges on a robust operational foundation that supports continuous iteration and quality assurance.
- The New Lean Startup — Sid Bendre, Oleve
This talk introduces a new model for building companies, termed the new lean startup, driven by the advent of AI tooling. It emphasizes the shift towards smaller, more profitable companies that achieve significant ARR with minimal teams. The core thesis is that AI tools enable tiny teams to build and scale successful products rapidly, challenging the traditional startup growth model.
- Conquering Agent Chaos — Rick Blalock, Agentuity
This talk addresses the significant challenges in deploying and running AI agents, often referred to as agent chaos. The speaker highlights common issues like timeouts, state management, and the need for agents to run for extended periods, pause, and resume. The core thesis is that successful agent deployment requires specific infrastructure and capabilities beyond typical stateless web applications, leading to the development of a platform designed to manage these complexities.
- Optimizing inference for voice models in production - Philip Kiely, Baseten
This talk focuses on optimizing inference for voice models in production, emphasizing runtime performance and infrastructure considerations. It highlights how the architectural similarity between Text-to-Speech (TTS) models and Large Language Models (LLMs) allows for the application of LLM optimization techniques. The core thesis is that while runtime optimizations are crucial, non-runtime factors like infrastructure and client code implementation can significantly impact overall latency and cost-efficiency.
- [Evals Workshop] Mastering AI Evaluation: From Playground to Production
This talk focuses on mastering AI evaluation, moving from initial development to production. It emphasizes that even the best large language models (LLMs) require a robust testing framework due to issues like hallucinations and performance degradation with changes. Effective evaluation helps answer critical questions about model selection, cost-effectiveness, brand consistency, and ongoing improvement, ultimately reducing development time, costs, and enabling faster iteration.
- Intro to GraphRAG — Zach Blumenfeld
This talk introduces GraphRAG, an architecture that integrates knowledge graphs with AI agents to enhance data retrieval and reasoning. It demonstrates how to build a knowledge graph using Neo4j and leverage it with tools like LangChain and LangGraph for applications such as talent management and skill analysis. The session covers data modeling, querying with Cypher, incorporating vector search for semantic similarity, and building agents that can interact with the graph.
- Securing Agents with Open Standards — Bobby Tiernay and Kam Sween, Auth0
This talk addresses the critical security challenges arising from increasingly capable AI agents that perform actions in the real world. It emphasizes the need for robust identity and access control mechanisms to prevent issues like secrets in prompts, overly broad scopes, and difficult troubleshooting. The core thesis is that by adopting open standards and implementing proper identity management, developers can build more secure and trustworthy AI agent systems.
- The emerging skillset of wielding coding agents — Beyang Liu, Sourcegraph / Amp
This talk explores the evolving skillset required to effectively utilize coding agents, moving beyond early chatbot paradigms. It argues that the current era demands a new approach to agent interaction, focusing on empowering agents to perform tasks autonomously rather than micromanaging them. The core thesis is that mastering coding agents is a high-ceiling skill that will significantly enhance developer productivity, akin to learning a new programming language or editor.
- Turning Fails into Features: Zapier’s Hard-Won Eval Lessons — Rafal Willinski, Vitor Balocco, Zapier
Building effective AI agents and the platforms that enable non-technical users to create them is a complex challenge due to the inherent nondeterminism of AI and unpredictable user behavior. The process requires a shift from traditional software development to building a data flywheel that continuously collects feedback, understands usage patterns, and identifies failures to improve the product iteratively. This involves robust instrumentation, strategic feedback collection, and a tiered approach to evaluation.
- Building voice agents with OpenAI — Dominik Kundel, OpenAI
This talk introduces OpenAI's Agents SDK for TypeScript, focusing on building voice agents. It defines agents as systems that accomplish tasks independently using a model, instructions, and tools, all managed by a runtime. The SDK offers features like handoffs, guardrails, streaming, tool calling, and native voice support, aiming to make technology more accessible and information exchange richer through voice.
- Containing Agent Chaos — Solomon Hykes, Dagger
This talk addresses the chaos and complexity arising from the use of coding agents, particularly from the perspective of platform engineers who enable developers. It proposes a new approach to managing agents by leveraging containerization and Git-like versioning to create isolated, customizable, and collaborative development environments. The core idea is to provide agents with dedicated, manageable environments that allow for background work, clear constraints, seamless human intervention, and flexibility in choosing underlying tools and infrastructure.
- Evals 101 — Doug Guthrie, Braintrust
This talk introduces the concept and practical application of evals for AI systems, emphasizing their role in ensuring quality, reliability, and correctness. Evals provide a structured testing framework to move beyond the non-deterministic nature of LLMs, enabling developers to rigorously build and iterate on AI applications. The presentation highlights how evals facilitate a feedback loop between offline development and online production monitoring, ultimately leading to improved AI products.
- Why should anyone care about Evals? — Manu Goyal, Braintrust
This talk argues that evals are crucial for the success and iteration of AI products, moving beyond simple unit tests for AI. Evals provide a simulated laboratory environment, allowing developers to test and refine models extensively before deploying to production. This process significantly speeds up development cycles and increases confidence in shipping AI features.
- Engineering Better Evals: Scalable LLM Evaluation Pipelines That Work — Dat Ngo, Aman Khan, Arize
This talk focuses on building scalable LLM evaluation pipelines. It emphasizes that effective evaluation is crucial for developing high-quality AI products and goes beyond simple "LLM as a judge" approaches. The core thesis is that a comprehensive evaluation strategy involves multiple methods, continuous tuning, and integration into the development lifecycle to accelerate iteration and improve AI system performance.
- To the moon! Navigating deep context in legacy code with Augment Agent — Forrest Brazeal, Matt Ball
This talk demonstrates how Augment Agent can be used to navigate and modernize complex, legacy codebases. By leveraging AI to understand historical code, developers can accelerate comprehension, identify potential issues, and even automate modernization tasks. The presentation uses the Apollo 11 guidance computer code as a case study to illustrate these capabilities.
- Serving Voice AI at Scale — Arjun Desai (Cartesia) & Rohit Talluri (AWS)
This talk addresses the challenges and advancements in serving voice AI at scale, focusing on the critical need for low latency and high quality in real-time interactive applications. It introduces state space models (SSMs) as a more efficient alternative to transformers for handling long sequences, enabling faster and more natural voice interactions across various devices. The discussion highlights how these advancements are crucial for enterprise voice AI, customer support, gaming, and content creation.
- Ship it! Building Production Ready Agents — Mike Chambers, AWS
This talk focuses on building production-ready AI agents, moving beyond simple prototypes to scalable, cloud-hosted solutions. It outlines the essential components of an agent, including the model, prompt, loop, history, and tools, and demonstrates how to leverage AWS services like Amazon Bedrock Agents to deploy these agents at cloud scale. The presentation emphasizes practical steps for integrating custom tools via Lambda functions and preparing agents for a production environment.
- Introducing Strands Agents, an Open Source AI Agents SDK — Suman Debnath, AWS
Strands Agents is an open-source SDK designed to simplify the creation of AI agents by minimizing scaffolding. The core philosophy is to leverage the increasing intelligence of AI models, allowing them to handle the reasoning process with minimal explicit prompting. Strands focuses on integrating models and tools, enabling developers to build agentic applications with greater ease.
- Data is Your Differentiator: Building Secure and Tailored AI Systems — Mani Khanuja, AWS
This talk emphasizes that data is the critical differentiator for building secure and tailored AI systems. Generative AI applications require special data treatment beyond standard transformation and loading, focusing on how data interacts with technology and people, and avoiding data silos. The approach to data must be tailored to specific use cases, such as travel agents, employee productivity chatbots, or marketing tools, each with unique data needs and responsibilities.
- How to build world-class AI products — Sarah Sachs (AI lead @ Notion) & Carlos Esteban (Braintrust)
This talk, presented by Sarah Sachs of Notion AI and Carlos Esteban of Braintrust, focuses on the critical role of observability and rigorous evaluation in building world-class AI products. The core thesis is that the quality and scalability of AI products stem directly from robust evaluation processes, which allow teams to iterate effectively and ensure consistent performance beyond simple demos.
- From Mixture of Experts to Mixture of Agents with Super Fast Inference - Daniel Kim & Daria Soboleva
This talk explores scaling large language models (LLMs) beyond traditional Mixture of Experts (MoE) architectures by introducing the concept of Mixture of Agents (MoA). MoA leverages multiple specialized LLMs, akin to agents, to collectively solve complex problems, aiming for higher intelligence and efficiency compared to monolithic models. The presentation also highlights Cerebras' hardware, designed for extremely fast inference, which enables practical implementation of MoA systems.
- Forget RAG Pipelines—Build Production Ready Agents in 15 Mins: Nina Lopatina, Rajiv Shah, Contextual
This talk introduces Contextual AI's platform for building production-ready Retrieval Augmented Generation (RAG) agents quickly, aiming to simplify the RAG pipeline. The core thesis is that RAG can be treated as a managed service, abstracting away the complexities of building and maintaining individual components like vector databases and LLM training. The platform offers an end-to-end solution for ingesting data, retrieving relevant information, and generating grounded responses, with a focus on enterprise-grade accuracy and modularity.
- Milliseconds to Magic: Real‑Time Workflows using the Gemini Live API and Pipecat
This talk explores the development of real-time voice AI workflows, emphasizing voice as a natural and universal interface for the next generation of AI applications. It highlights the complexities involved in creating seamless voice interactions, from foundational LLMs and real-time APIs like Gemini Live API to orchestration frameworks such as Pipecat and application code. The presentation suggests that while significant progress has been made, many aspects of voice AI are still in early stages of development, with capabilities progressively moving down the technology stack.
- Realtime Conversational Video with Pipecat and Tavus — Chad Bailey and Brian Johnson, Daily & Tavus
This talk introduces Pipecat, an open-source framework for building real-time conversational AI, and Tavus, a platform for creating conversational video replicas. It highlights the three key components needed for real-time AI: models, orchestration, and deployment. The presenters explain how Pipecat acts as an orchestration layer, managing the flow of data (frames) through various processors to enable low-latency audio and video communication.
- Vector Search Benchmark[eting] - Philipp Krenn, Elastic
This talk addresses the common pitfalls and deceptive practices in vector search benchmarking, often referred to as "benchmarketing." The core thesis is that many published benchmarks are misleading due to biased scenario selection, outdated competitor versions, and a lack of focus on crucial factors like data freshness and query relevance. The speaker emphasizes that truly meaningful benchmarks require rigorous, reproducible, and self-conducted evaluations tailored to specific use cases.
- Taming Rogue AI Agents with Observability-Driven Evaluation — Jim Bennett, Galileo
This talk addresses the challenge of ensuring AI agents function reliably by introducing observability-driven evaluation. It highlights that AI's non-deterministic nature makes traditional testing methods insufficient. The core thesis is that by using AI itself to evaluate AI outputs, developers can gain crucial insights into agent performance, identify failures at granular levels, and implement targeted improvements.
- Building agent fleet architectures your CISO doesn't hate — Lou Bichard, Gitpod
This talk discusses the evolution of Gitpod's architecture to support secure development environments, particularly for regulated industries, and how this foundation enables a privacy-first AI agent offering. The core thesis is that simplifying infrastructure, moving away from complex systems like Kubernetes towards native cloud services, leads to better security, lower operational overhead for customers, and a more robust platform for deploying advanced tools like AI agents.
- Don’t get one-shotted: Use AI to test, review, merge, and deploy code — Tomas Reimers, Graphite
The increasing adoption of AI tools in software development, particularly in the inner loop of coding, is creating a bottleneck in the outer loop of review, testing, merging, and deployment. This talk proposes that AI can also be leveraged to streamline and automate these outer loop processes, transforming the entire developer workflow rather than just the IDE.
- Effective agent design patterns in production — Laurie Voss, LlamaIndex
This talk introduces LlamaIndex as a framework for building generative AI applications, with a particular focus on agents. It emphasizes the necessity of Retrieval Augmented Generation (RAG) for agents to effectively process and utilize large, unstructured datasets. The presentation outlines several high-level design patterns that enhance agent performance in production environments.
- Foundry Local: Cutting-Edge AI experiences on device with ONNX Runtime/Olive — Emma Ning, Microsoft
Foundry Local is a Microsoft solution designed to enable developers to build cross-platform AI applications that run directly on user devices. It addresses the need for local AI by providing reasons such as offline access, enhanced privacy and security for sensitive data, cost efficiency for high-volume inference, and real-time latency requirements. The platform leverages ONNX Runtime for accelerated on-device inference across various hardware, integrating with Azure AI Foundry for model management and on-demand downloads.
- [Full Workshop] Vibe Coding at Scale: Customizing AI Assistants for Enterprise Environments
This talk explores "Vibe Coding," a methodology for leveraging AI assistants to accelerate software development by focusing on output rather than the intricacies of code. It outlines a progression from "YOLO vibes" for rapid prototyping to "structured vibes" for maintainability and "spectrum vibes" for scale and reliability, emphasizing trust-building and guardrails for enterprise environments.
- Unlocking AI Powered DevOps Within Your Organization — Jon Peck, GitHub
This talk explores how organizations can effectively integrate AI into their DevOps workflows to enhance efficiency and developer productivity. It emphasizes moving beyond basic AI code generation to leverage AI for planning, security, testing, deployment, and documentation. The presentation also touches on the evolution towards more autonomous AI agents and the importance of governance and safety.
- Vibe Coding at Scale: Customizing AI Assistants for Enterprise Environments - Harald Kirshner,
This talk explores different approaches to AI-assisted coding, moving from rapid, experimental "YOLO vibe coding" to more structured and spec-driven methods suitable for enterprise environments. The core thesis is that by understanding and applying these different "vibes," developers can leverage AI assistants more effectively, leading to increased productivity and maintainability in software development.
- The Agent Awakens: Collaborative Development with Copilot - Christopher Harrison, GitHub
- AI Red Teaming Agent: Azure AI Foundry — Nagkumar Arkalgud & Keiji Kanazawa, Microsoft
This talk introduces the AI Red Teaming Agent, a tool developed within Azure AI Foundry to help AI engineers proactively identify and mitigate risks in their AI systems. It emphasizes that building trustworthy AI is a collaborative effort, akin to building bridges and dams, and that AI engineering requires a systematic approach to testing and iteration. The agent provides a practical way for developers to simulate attacks and evaluate their models' vulnerabilities.
- Collaborating with Agents in your Software Dev Workflow - Jon Peck & Christopher Harrison, Microsoft
This talk explores how developers can effectively collaborate with AI agents, specifically focusing on GitHub Copilot's capabilities within the software development workflow. The core thesis is that understanding and providing proper context is crucial for maximizing the AI pair programmer's utility, moving beyond simple prompt engineering to encompass code readability, project structure, and clear intent. The discussion highlights various Copilot features, from basic code completion to advanced agent modes, and emphasizes best practices for integrating these tools into daily development tasks.
- Agentic Excellence: Mastering AI Agent Evals w/ Azure AI Evaluation SDK — Cedric Vidal, Microsoft
This talk focuses on the critical process of evaluating AI agents, emphasizing a shift from ad-hoc testing to a more methodical approach. It highlights that effective evaluation should begin at the earliest stages of AI development, not as an afterthought. The presentation introduces tools and frameworks, particularly the Azure AI Evaluation SDK, to ensure AI agents behave correctly and safely as they gain more independence.
- Building Code First AI Agents with Azure AI Agent Service — Cedric Vidal, Microsoft
This talk explores building AI agents using Azure AI Agent Service, focusing on a practical approach for developers. It defines an agent as a semi-autonomous software that pursues a goal by reasoning, integrating with data, and acting on the world. The presentation demonstrates how to leverage Azure AI Agent Service to simplify agent development by managing state, context, and tool integration in the cloud, enabling the creation of applications that can dynamically generate queries, user interfaces, and visualizations.
- How fast are LLM inference engines anyway? — Charles Frye, Modal
This talk explores the performance of open-source LLM inference engines, highlighting how recent advancements in model quality and inference software have made self-hosting viable. It presents benchmarking data to help engineers understand and optimize LLM performance for various use cases, emphasizing the trade-offs between different configurations and workloads.
- RAG in 2025: State of the Art and the Road Forward — Tengyu Ma, MongoDB (acq. Voyage AI)
This talk explores Retrieval Augmented Generation (RAG) as a key technology for enabling large language models (LLMs) to access and utilize proprietary enterprise data. The speaker argues that RAG offers a more efficient, cost-effective, and manageable approach compared to fine-tuning or relying solely on long context windows. The presentation also touches upon the evolution of RAG, advancements in embedding models, and future directions for improving retrieval accuracy and simplifying workflows.
- The State of AI Powered Search and Retrieval — Frank Liu, MongoDB (prev Voyage AI)
This talk explores the evolution and future of AI-powered search and retrieval, moving beyond traditional keyword matching to understand user intent and conceptual relationships. It highlights how AI search systems can provide more grounded and relevant responses, particularly in applications like Retrieval-Augmented Generation (RAG), by leveraging embeddings and advanced techniques. The discussion also touches upon the increasing importance of multimodality and agentic capabilities in shaping the next generation of search technologies.
- Architecting Agent Memory: Principles, Patterns, and Best Practices — Richmond Alake, MongoDB
This talk explores the critical role of memory in developing advanced AI agents, moving beyond stateless applications to create more believable, capable, and reliable systems. It posits that memory is fundamental to mimicking human intelligence and enabling agents to be reflective, interactive, proactive, and autonomous. The discussion emphasizes memory management as a core task for AI engineers, focusing on principles and practical patterns for integrating memory into agentic systems.
- Building Multimodal AI Agents From Scratch — Apoorva Joshi, MongoDB
This talk introduces the concept of building multimodal AI agents from scratch, focusing on integrating text and image processing capabilities. It explains the evolution from simple prompting and RAG to AI agents, highlighting agents' suitability for complex, multi-step tasks requiring reasoning and action. The session details the core components of an agent—perception, planning/reasoning, tools, and memory—and demonstrates how to construct a multimodal agent capable of answering questions about documents containing both text and images, and analyzing charts or diagrams.
- Why Your Agent’s Brain Needs a Playbook: Practical Wins from Using Ontologies - Jesús Barrasa, Neo4j
This talk explores the integration of knowledge graphs and Large Language Models (LLMs) for building robust AI applications, specifically focusing on the Graph RAG architecture. It highlights how ontologies, as formal, implementation-agnostic schemas, can significantly enhance knowledge graph creation and retrieval strategies, leading to more grounded and accurate AI responses. The core thesis is that a model-driven approach using ontologies provides practical wins by improving data quality and enabling dynamic retriever behavior.
- Memory Masterclass: Make Your AI Agents Remember What They Do! — Mark Bain, AIUS
This talk explores the critical role of memory in AI agents, positing that true AI memory encompasses all data, code, algorithms, and hardware, along with any causal changes affecting them. It draws parallels between the principles governing Large Language Models (LLMs), neuroscience, and mathematics, suggesting that asymmetries are necessary for existence and that preserving causal links through relationships is key to solving issues like hallucinations and optimizing hypothesis generation. The presentation highlights the potential of graph databases and agentic systems for building more robust and context-aware AI.
- Graph Intelligence: Enhance Reasoning and Retrieval Using Graph Analytics - Alison & Andreas, Neo4j
This talk explores how graph data science can enhance reasoning and retrieval in AI applications, particularly within the context of Retrieval Augmented Generation (RAG). It emphasizes that graphs provide a way to understand data relationships beyond simple vector similarity, enabling more comprehensive and context-aware responses. The session introduces graph concepts and demonstrates how to leverage graph analytics to manage and improve data quality for AI systems.
- GraphRAG methods to create optimized LLM context windows for Retrieval — Jonathan Larson, Microsoft
This talk introduces GraphRAG, a method for optimizing Large Language Model (LLM) context windows using graph structures for retrieval. The core thesis is that LLM memory with structure is a key enabler for building effective AI applications, and when paired with agents, this combination offers even greater power. The presentation demonstrates GraphRAG's application in code understanding and feature development, alongside the announcement of Benchmark QED, an open-source tool for evaluating LLM systems.
- Agentic GraphRAG: Simplifying Retrieval Across Structured & Unstructured Data — Zach Blumenfeld
This talk introduces Agentic GraphRAG, a method for simplifying data retrieval across both structured and unstructured sources. It proposes using a knowledge graph as an intermediary layer to enhance agentic workflows. By modeling data within a graph, agents can more accurately decompose complex questions, pull relevant information, and perform analytical tasks that go beyond simple semantic search.
- Revenue Engineering: How to Price (and Reprice) Your AI Product — Kshitij Grover, Orb
This talk explores the complexities of pricing AI products, emphasizing that pricing is a form of friction that must be carefully managed to align with product value and audience needs. It moves beyond traditional pricing models to discuss AI-native considerations like predictability, speed of value demonstration, and rapidly changing cost structures. The core thesis is that effective AI product pricing requires a deep understanding of the target audience, value delivery mechanisms, and flexible margin structures, allowing for continuous experimentation and adaptation.
- \"Data readiness\" is a Myth: Reliable AI with an Agentic Semantic Layer — Anushrut Gupta, PromptQL
The talk argues that the concept of "data readiness" for AI is a myth, as data is rarely perfect and constantly changing. Instead of striving for pristine data, the focus should be on building AI systems that can reliably work with messy, evolving data. This is achieved by creating an agentic semantic layer that learns and adapts to the specific business domain and its nuances over time, much like an experienced human analyst.
- Building Agentic Applications w/ Heroku Managed Inference and Agents — Julián Duque & Anush Dsouza
This talk introduces Heroku's Managed Inference and Agents service, designed to simplify the development and deployment of agentic AI applications. The service allows developers to integrate AI models directly within their application's infrastructure, enhancing security and control. It provides primitives for inference, model context protocol (MCP) integration, and a managed PostgreSQL database with PGVector for embeddings, enabling the creation of sophisticated AI-powered features.
- Events are the Wrong Abstraction for Your AI Agents - Mason Egger, Temporal.io
This talk argues that event-driven architecture (EDA) is the wrong abstraction for AI agents, leading to complex, tightly coupled systems. Instead, the speaker proposes durable execution as a superior model. Durable execution, exemplified by Temporal.io, offers crash-proof processing by automatically preserving application state, virtualizing execution across machines, and being time and hardware agnostic. This shift allows developers to focus on core business logic rather than managing the intricacies of event handling and system failures.
- Prompt Engineering is Dead — Nir Gazit, Traceloop
The talk argues that traditional prompt engineering is ineffective and proposes an automated approach to improving AI model responses. The speaker shares a personal experience of enhancing a Retrieval Augmented Generation (RAG) chatbot by building an agent that iteratively refines prompts based on an evaluation system, rather than manual prompt tuning. This method aims to achieve significant improvements without direct, manual prompt manipulation.
- The Eyes Are The (Context) Window to The Soul: How Windsurf Gets to Know You — Sam Fertig, Windsurf
This talk explores how AI coding tools can better understand and assist developers by deeply understanding context. The core thesis is that generating code is no longer the primary challenge; instead, the difficulty lies in creating code that integrates seamlessly into existing, complex codebases, adheres to specific standards, and aligns with individual developer preferences. This is achieved by focusing on the "what" and "how much" of context, rather than simply increasing context window sizes.
- Mastering Engineering Flow with Windsurf - Eashan Sinha, Windsurf
This talk introduces Windsurf's approach to enhancing the developer experience through agentic IDEs, focusing on the concept of "AI flows." The core thesis is that by treating AI coding assistants as collaborative teammates rather than separate tools, developers can achieve a more seamless and productive engineering flow. Windsurf's Cascade agent aims to achieve this by deeply understanding user intent and codebase context, moving beyond simple autocomplete or autonomous agents.
- (possible dupe but better sound) What does Enterprise Ready MCP mean? — Tobin South, WorkOS
This talk explores what it means for Model Communication Protocol (MCP) to be enterprise-ready, moving beyond basic tool-use capabilities to address the complexities of production environments. It highlights the evolution from simple chatbot interactions to sophisticated AI agents that require robust security, management, and scalability, particularly when integrating with internal enterprise systems. The core thesis is that while MCP offers a standardized way for AI to interact with external resources, achieving enterprise readiness involves overcoming significant challenges in authentication, authorization, and operational management.
- CI in the Era of AI: From Unit Tests to Stochastic Evals — Nathan Sobo, Zed
This talk explores the challenges and strategies for testing AI-enabled features, particularly agentic editing in code editors. It emphasizes the shift from deterministic testing to embracing stochastic evaluations when LLMs are involved. The core thesis is that rigorous, empirical software development practices, adapted for the non-deterministic nature of AI, are crucial for building reliable AI-powered products.
- Fun stories from building OpenRouter and where all this is going - Alex Atallah, OpenRouter
This talk explores the founding story of OpenRouter, initially conceived as an experiment to address the burgeoning AI inference market. It delves into the investigation of whether this market would be dominated by a single player or foster a diverse ecosystem. The discussion highlights the evolution from a model exploration platform to a comprehensive marketplace, emphasizing the challenges and innovations in aggregating and standardizing diverse AI models.
- Building AI Agents that actually automate Knowledge Work - Jerry Liu, LlamaIndex
This talk explores how AI agents can automate knowledge work, moving beyond simple chatbots to handle complex tasks involving unstructured data. It outlines a framework for building these agents, emphasizing the need for robust toolkits and well-defined agent architectures to process and act upon diverse data formats like documents and spreadsheets. The core thesis is that by combining advanced document understanding with flexible agent design patterns, AI can significantly enhance efficiency in knowledge-intensive roles.
- RFT, DPO, SFT: Fine-tuning with OpenAI — Ilan Bigio, OpenAI
This talk explores various fine-tuning techniques for OpenAI models, including Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Reinforcement Fine-Tuning (RFT). The presenter, Ilan Bigio from OpenAI's developer experience team, emphasizes that fine-tuning is a specialized tool for optimizing models beyond what prompt engineering can achieve, particularly for specific domains or behaviors. The discussion covers the data requirements, use cases, and limitations of each method, offering practical examples and best practices.
- Windsurf everywhere, doing everything, all at once - Kevin Hou, Windsurf
This talk introduces Windsurf, an AI-powered editor designed to revolutionize software creation by extending beyond traditional code completion. The core thesis is that AI should operate on a shared timeline with human developers, ingesting context from all sources a developer uses and taking actions across various platforms, not just within the IDE. This approach aims to enable AI to perform nearly all tasks a human software engineer would, moving towards a future where AI handles 99% of the workflow with human oversight only for final approval.
- Case Study + Deep Dive: Telemedicine Support Agents with LangGraph/MCP - Dan Mason
This talk presents a case study and deep dive into building telemedicine support agents using LangGraph and MCP. The core thesis is that LLM-powered agents can significantly enhance flexibility and capability in healthcare support systems, moving beyond traditional software limitations. The system aims to provide scalable, automated patient support while maintaining human oversight for complex or sensitive situations.
- Veo 3 for Developers — Paige Bailey, Google DeepMind
This talk introduces Google DeepMind's latest generative media models, focusing on V3 for video and audio generation, Imagine 4 for static images, and LIA 2 for music. The presentation highlights how these tools can revolutionize content creation, advertising, and user experiences by enabling the generation of novel content and offering enhanced creative control. The models are designed to improve stylistic and contextual consistency, making them powerful tools for developers and creators.
- Building Agents with Amazon Nova Act and MCP - Du'An Lightfoot, Amazon (Full Workshop)
This workshop introduces building intelligent autonomous AI systems using Amazon Nova Act and the Modern Communication Protocol (MCP). It focuses on a do-it-yourself approach, enabling developers to create agentic systems that can plan, act, and reason to achieve objectives. The session highlights how these agents can leverage tools, knowledge bases, and LLMs to tackle complex tasks, with a particular emphasis on browser automation and integrating various AI components.
- The Web Browser Is All You Need - Paul Klein IV, Browserbase
This talk argues that the web browser is the essential bridge for AI agents to interact with the vast majority of the internet, which lacks modern APIs. It posits that browsers, specifically headless browser MCP servers, are the key to unlocking AI's potential for interacting with legacy websites and services. The presentation explores different approaches to web agents and browser tools, emphasizing the browser's role as a universal integration point.
- Building Protected MCP Servers — Den Delimarsky and Julia Kasper, MCP Steering Committee & Microsoft
This talk addresses the critical need for security in Multi-Call Protocol (MCP) servers, particularly for remote instances. It introduces a new draft specification that simplifies authorization by separating the MCP server's role from that of an authorization server. This approach aims to reduce the burden on developers by allowing them to leverage existing OAuth 2.0 libraries and identity providers, rather than implementing complex authorization logic themselves.
- The State of MCP observability: Observable.tools — Alex Volkov and Benjamin Eckel, W&B and Dylibso
The increasing prevalence of Machine Communication Protocol (MCP) agents is creating an observability blind spot for developers. As agents utilize more tools via MCP, understanding their end-to-end execution becomes challenging. This talk introduces observable.tools as a manifesto to drive a community conversation around standardized, vendor-neutral MCP observability, advocating for the adoption of OpenTelemetry (otel) to address this growing issue.
- The Geopolitics of AI Infrastructure - Dylan Patel, SemiAnalysis
This talk examines the geopolitical landscape of AI infrastructure, focusing on the capabilities of China and the Middle East in contrast to US limitations. It highlights how China, despite sanctions, is advancing its AI chip manufacturing and deployment through innovative engineering and supply chain circumvention. Simultaneously, the Middle East is emerging as a significant hub for AI infrastructure development, attracting substantial investment and GPU resources, partly due to power constraints in the US.
- Remote MCPs: What we learned from shipping — John Welsh, Anthropic
This talk discusses the challenges and solutions for implementing and scaling Message Communication Protocol (MCP) clients within a large organization. It highlights how the rapid evolution of AI models capable of tool calling led to integration chaos, with duplicated functionality and inconsistent interfaces across services. The solution proposed is to standardize on MCP for providing model context, treating it as a plumbing layer for integrations rather than a competitive differentiator.
- MCP: Origins and Requests For Startups — Theodora Chu, Model Context Protocol PM, Anthropic
The Model Context Protocol (MCP) is an open-source, standardized protocol designed to give AI models agency by allowing them to interact with the outside world. Originating from the need to copy context from external sources into LLM context windows, MCP aims to enable models to reach a new level of usefulness and intelligence by facilitating tool calling and broader interaction capabilities. The protocol prioritizes server simplicity and encourages community contributions to evolve its standards and utility.
- How to Build Trustworthy AI — Allie Howe
This talk addresses the critical need for building trustworthy AI systems, emphasizing that responsibility ultimately lies with the user and developer. It defines trustworthy AI as a combination of AI security (protecting AI from external harm) and AI safety (preventing AI from harming the world). The presentation advocates for a shift from traditional DevSecOps to MLSecOps, highlighting the importance of integrating security and safety practices throughout the AI development lifecycle, from build to runtime.
- Exposing Agents as MCP servers with mcp-agent: Sarmad Qadri
This talk introduces the Model Context Protocol (MCP) as a standardized interface for connecting Large Language Models (LLMs) to external tools and resources, aiming to simplify and enhance agent development. It posits that 2025 will be the year agents hit mass production, facilitated by advancements in LLM reasoning capabilities, the widespread adoption of MCP, and simpler agent architectures. The presentation also explores modeling agents as asynchronous workflows and exposing them as MCP servers for greater composability and scalability.
- Supercharging developer workflow with Amazon Q Developer - Vikash Agrawal
This talk demonstrates how Amazon Q Developer, an AI coding assistant, can be integrated throughout the software development lifecycle (SDLC). It highlights Q Developer's capabilities in IDEs, CLIs, and GitHub to assist with planning, coding, testing, documentation, deployment, and even operational debugging, aiming to supercharge developer workflows.
- Just do it. (let your tools think for themselves) - Robert Chandler
This talk addresses the common challenges of unreliability, slowness, and cost associated with current AI agents that interact with external tools. The core thesis is that instead of creating simple, low-level wrappers around APIs, tools should be imbued with more agency, becoming specialized agents themselves. This approach blurs the line between tools and agents, enabling agents to offload complex tasks to more capable, specialized entities for improved performance and reliability.
- Break It 'Til You Make It: Building the Self-Improving Stack for AI Agents - Aparna Dhinakaran
This talk addresses the challenges of building and iterating on AI agents, particularly in evaluating their performance and identifying bottlenecks. It emphasizes the need for systematic evaluation beyond manual inspection of a few examples. The core thesis is that a self-improving stack for AI agents requires not only improving the agent's prompts and models but also continuously refining the evaluation methods themselves.
- MCPs are Boring (or: Why we are losing the Sparkle of LLMs) - Manuel Odendahl
This talk argues that the current focus on Multi-Call Protocol (MCP) and tool calling for LLMs is limiting their potential. The speaker contends that LLMs are powerful code generators and language producers, capable of much more than simply calling predefined functions with fixed schemas. The core thesis is that by treating LLMs as sophisticated code-generation engines and embracing recursive creation, developers can unlock greater flexibility and power, moving beyond the "boring" limitations of current tool-calling paradigms.
- The Many Ends of Programming - Ray Myers
This talk explores the evolving landscape of programming in the age of AI, moving beyond hype to discuss practical scenarios and their implications for software development. It argues that the future of programming isn't a single predetermined outcome but a spectrum of possibilities, emphasizing the need for empathy, careful consideration, and active participation in shaping this future. The core thesis is that while AI tools are rapidly advancing, human engineers play a crucial role in guiding their integration and ensuring positive outcomes.
- Why Bolt.new Won and Most DevTools AI Pivots Failed - Victoria Melnikova
This talk argues that most developer tool AI pivots fail because they incorrectly implement AI features. Instead of simply adding AI to existing workflows or trying to compete with large AI providers, successful products leverage their unique competitive advantages. By identifying what makes a product distinct and then exploring how AI can amplify that specific advantage, companies can create new categories and reinvent user experiences, leading to significant growth.
- Beyond Conversation: Why Documents Transform Natural Language into Code - Filip Kozera
This talk argues that document-based interfaces are superior to chat-based interfaces for complex tasks involving AI. Chat interfaces suffer from context pollution, lack structured iteration, and offer poor version control, leading to degraded model performance. Documents, conversely, provide a mechanism for forced clarity and structured communication, enabling more robust specification of complex systems and paving the way for background agents that can perform tasks autonomously.
- The 4 Patterns of AI Native Development — Patrick Debois
This talk outlines four key patterns emerging in AI-native development, shifting the developer's role beyond traditional coding. As AI tools evolve from simple code generation to complex agent teams, developers are transitioning into roles like managers, specifiers, discoverers, and knowledge curators. This evolution signifies a move towards more strategic and less purely executional work, fundamentally changing the software development lifecycle.
- Breaking the Chain: Agent Continuations for Resumable AI Workflows - Greg Benson
This talk introduces agent continuations, a new mechanism designed to address key challenges in deploying AI agents in production. It enables agents to pause their execution, save their complete state, and resume later, facilitating human approval workflows and robust error handling for long-running processes. This approach aims to make AI agents more reliable and manageable in complex, distributed environments.
- Are MCPs Overhyped? A Rant about MCPs — Henry Mao, Smithery
This talk critically examines the Model Context Protocol (MCP) ecosystem, arguing that despite initial excitement and the promise of standardizing AI agent interactions with services, significant challenges remain. The speaker contends that the current state of MCPs is fragmented and difficult to use, hindering the practical application of autonomous agents. The core thesis is that while the potential for AI agents is vast, the MCP infrastructure is not yet mature enough to fully realize this potential, necessitating further development and standardization.
- Why the Best AI Agents Are Built Without Frameworks (Primitives over Frameworks) — Ahmad Awais, CHAI
This talk argues that building AI agents on primitives rather than frameworks leads to more production-ready and scalable solutions. Frameworks are often bloated, slow, and introduce unnecessary abstractions. Instead, developers should leverage composable, low-level primitives, similar to how cloud services like Amazon S3 function, to create flexible and efficient AI agents.
- GPU-less, Trust-less, Limit-less: Reimagining the Confidential AI Cloud - Mike Bursell
This talk introduces confidential AI as a solution to trust issues in AI development and deployment. It explains how confidential computing, utilizing Trusted Execution Environments (TEs), protects data and models during processing, even from system administrators or hardware providers. This technology enables secure collaboration, monetization, and the use of sensitive data for AI tasks across various industries.
- Grounded Reasoning Systems for Cloud Architecture - Iman Makaremi
This talk explores the development of grounded reasoning systems for cloud architecture, emphasizing the need for AI systems that can understand, debate, and justify architectural decisions rather than simply automate tasks. It highlights the complexity of cloud architecture, the cognitive aspects of architectural decision-making, and the challenges of integrating textual requirements with graph-based architectural data. The core thesis is that multi-agent orchestration is key to building AI copilots capable of complex reasoning in this domain.
- The Agent Native Company — Rick Blalock, Agentuity
This talk introduces the concept of an "agent-native" or "AI-native" company, distinguishing it from merely AI-enhanced businesses. An agent-native company is built from the ground up with AI agents as a core component, fundamentally altering how work is done, teams are structured, and products are developed. This paradigm shift moves beyond AI as an add-on to AI as the central engine driving operations, culture, and product innovation.
- The Voice-First AI Overlay: Designing Conversational Co-Pilots - Gregory Bruss
This talk explores the concept of a voice-first AI overlay designed to enhance human-to-human conversations by providing real-time assistance without directly participating as a third speaker. The core idea is to leverage advancements in AI agent capabilities and voice technology to keep humans informed and on track during interactions, using voice as the most natural interface.
- Arrakis: How To Build An AI Sandbox From Scratch - Abhishek Bhardwaj, OpenAI
Arrakis is an open-source, self-hosted service for creating and managing AI sandboxes, designed for secure code execution and computer use by AI agents. It leverages microVMs to provide isolated environments, enabling AI models to safely utilize tools like code execution and search. The system emphasizes security, speed, and the ability for agents to backtrack and replan using snapshots, facilitating more complex task completion.
- 7 Habits of Highly Effective Generative AI Evaluations - Justin Muller
This talk emphasizes that robust evaluations are the most critical, yet often overlooked, component for successfully scaling generative AI workloads. The speaker argues that evaluations are not merely for measuring quality but are primarily a tool for discovering problems within AI systems. Implementing a strong evaluation framework is presented as the key differentiator between a project that remains a science experiment and one that achieves production-level success and scalability.
- Agents reported thousands of bugs, how many were real? - Ian Butler and Nick Gregory
This talk investigates the effectiveness of AI agents in identifying and fixing bugs within software development, moving beyond traditional feature development benchmarks. The presenters introduce a new benchmark designed to evaluate agents on maintenance tasks, highlighting that while agents can often patch simple bugs, their ability to comprehensively detect them, especially complex ones, is still in its early stages. The research suggests current agents struggle with holistic code evaluation and deep reasoning, leading to missed bugs and high false positive rates.
- The Coherence Trap: Why LLMs Feel Smart (But Aren’t Thinking) - Travis Frisinger
This talk argues that Large Language Models (LLMs) excel due to coherence, not intelligence. While LLMs can produce outputs that feel remarkably insightful and intelligent, they lack genuine understanding, intent, or desire. The speaker proposes that coherence is a system property, not a cognitive one, and explains how LLMs construct meaning on demand through pattern alignment within a high-dimensional latent space, rather than retrieving stored knowledge.
- The Knowledge Graph Mullet: Trimming GraphRAG Complexity - William Lyon
This talk introduces the "Knowledge Graph Mullet," a hybrid approach combining property graph and RDF triple concepts for more versatile knowledge graph management. It advocates for using a property graph model for data interaction and querying, while leveraging the scalability of RDF triples for the underlying storage. The presentation demonstrates this approach using Dgraph, an open-source graph database, and explores its application in building sophisticated Graph RAG (Retrieval Augmented Generation) workflows and AI agents.
- Building Reliable Support Agents Using the Effect Typescript Library - Michael Fester
This talk introduces the Effect TypeScript library as a robust solution for building reliable AI-powered customer support platforms. The core thesis is that while TypeScript provides a good foundation, Effect offers essential tools for managing the complexities of unreliable APIs, non-deterministic LLM outputs, and long-running workflows, thereby enhancing system stability, testability, and maintainability at scale.
- The Robots are coming for your job, and that's okay - Elmer Thomas and Maria Bermudez
This talk explores how AI agents can be used to enhance, rather than replace, human workflows, particularly within a small documentation team. The core thesis is that by automating repetitive, rule-based tasks, AI can free up human engineers to focus on higher-level judgment, clarity, and creativity, ultimately boosting productivity and reducing burnout.
- Blender MCP and The Future Of Creative Tools - Siddharth Ahuja
This talk explores the Blender MCP, an open-source project that enables Large Language Models (LLMs) to interact with and control Blender, a complex 3D creation tool. The core thesis is that by leveraging the MCP protocol and Blender's scripting capabilities, LLMs can significantly lower the barrier to entry for 3D content creation, allowing users to generate intricate scenes and assets through simple text prompts. This approach democratizes creative tools and points towards a future where LLMs act as orchestrators across various creative software.
- Agentic Enterprise - What your CEO must know about AI - Hubert Misztela
This talk explores the concept of agentic enterprise, positing that AI agents could fundamentally reshape organizations within three years. It defines AI agents as autonomous, LLM-based applications capable of planning, tool usage, and dynamic adaptation. The presentation emphasizes that digital assets should evolve to serve as tools or agents themselves, enabling capabilities like multi-step reasoning, adaptability, and computer vision for automation.
- The End of Awkward AI Transcriptions - Travis Bartley and Myungjong Kim
This talk details Nvidia's approach to developing enterprise-level speech AI models, focusing on robustness, coverage, personalization, and deployment efficiency. The core thesis is that a variety of specialized models, rather than a single monolithic solution, best meets diverse customer needs for conversational AI, emphasizing low latency and high efficiency for embedded devices.
- How agents broke app-level infrastructure - Evan Boyle
This talk addresses the challenges of building reliable AI applications, particularly focusing on the compute layer. Traditional web infrastructure is ill-suited for the long-running, non-deterministic nature of AI workflows, which often involve extensive data ingestion and complex agentic processes. The core thesis is that current infrastructure limitations lead to unreliable user experiences and hinder rapid experimentation, necessitating new architectural approaches.
- RAG Evaluation Is Broken! Here's Why (And How to Fix It) - Yuval Belfer and Niv Granot
This talk addresses the shortcomings of current Retrieval Augmented Generation (RAG) evaluation methods, arguing that they are fundamentally broken. The presenters contend that most benchmarks rely on simple, local questions with easily identifiable answers within specific text chunks, which does not reflect real-world data complexity. This leads to a cycle of optimizing for flawed benchmarks, resulting in RAG systems that perform poorly when deployed with actual user data.
- The Demo I Wish I'd Had: OpenAI's Agents SDK... serverless! - Brook Riggio
This talk presents a preferred architecture for building full-stack AI applications in a serverless environment. The core thesis is that by combining specific tools like Next.js, OpenAI's Agents SDK, and Ingest, developers can create resilient, agent-powered applications without managing complex infrastructure. The approach emphasizes integrating AI workflows directly into client applications for a seamless user experience.
- Will Agent evaluation via MCP Stabilize Agent Networks? - Ari Heljakka
This talk explores how the Model Contest Protocol (MCP) can be used to stabilize AI agents and agent networks. The core idea is that by systematically evaluating agent behavior and providing feedback, agents can learn to improve their performance and consistency, especially when tackling complex problems. This approach aims to create more controllable, transparent, and self-correcting agent systems.
- MCP Agent Fine tuning Workshop - Ronan McGovern
This workshop details the process of fine-tuning an AI agent that utilizes the Model Context Protocol (MCP) for tool access. It covers generating high-quality reasoning traces from agent interactions with tools, saving these multi-turn conversations, and using them to fine-tune a language model, specifically demonstrating with a Quen model. The goal is to improve the agent's performance by training it on its own successful interactions.
- From PM at Stripe to Building an AI startup, a recent founder's journey - Mounir Mouawad
This talk chronicles a founder's transition from a product role at a large tech company to establishing an AI startup. It highlights the unique challenges of identifying user problems in the nascent AI space, where problems are emergent rather than clearly defined. The presentation emphasizes the need for rapid iteration, hypothesis-driven development, and the creation of user narratives to navigate this evolving landscape.
- My AI Thinks I'm Eating My Feelings (and Other Nutritional Insights) - Rami Alhamad
This talk introduces Alma, an AI-powered nutrition companion designed to simplify healthy eating. The core thesis is that current nutrition tracking methods are overly complex, providing little value for the effort invested. Alma aims to solve this by making tracking natural and easy, building rich user context, and proactively connecting users with suitable food options.
- Rust is the language of the AGI - Michael Yuan
This talk argues that Rust is the ideal language for Artificial General Intelligence (AGI) development, positioning AI as the primary coder and humans as assistants. The core thesis is that Rust's rigorous compiler and strong type system, while challenging for humans, provide an excellent feedback loop for AI code generation, leading to more correct and efficient code. The Rust Coder project aims to facilitate this by making Rust easier for AI to generate and for humans to work with.
- Real AI Agents Need Planning, Not Just Prompting - Yuval Belfer
This talk argues that current large language models (LLMs), despite advancements, still struggle with complex instruction following. The core thesis is that AI agents require robust planning capabilities beyond simple prompting to effectively tackle intricate tasks, emphasizing the need for a lookahead strategy rather than step-by-step execution.
- The Current State of Browser Agents - Jerry Wu and Wyatt Marshall
This talk explores the current capabilities and limitations of browser agents, which are AI systems designed to control web browsers and perform tasks on behalf of users. The discussion covers the fundamental loop of observation, reasoning, and action that powers these agents, their common use cases like web scraping and form filling, and an evaluation of their performance based on a newly released benchmark dataset. The presentation highlights significant performance gaps between read tasks (information retrieval) and write tasks (state modification), and discusses the underlying reasons for these discrepancies.
- Text-to-Speech Data Preparation and Fine-tuning Workshop - Ronan McGovern
This workshop details the process of preparing data and fine-tuning a text-to-speech (TTS) model, specifically Sesame's CSM 1B model, to produce speech that mimics a target voice. It covers extracting audio from sources like YouTube, transcribing it, and formatting it into a dataset suitable for training. The process leverages the Unsloth library for efficient fine-tuning and demonstrates how to evaluate the model's performance before and after the fine-tuning process.
- Invisible Users, Invisible Interfaces: Accelerating Design Iteration with AI Simulation - Alex Liss
This talk proposes using AI simulation to accelerate design iteration by creating "intelligent twins" that act as invisible users. This approach aims to overcome the current AI trust gap, which stems from poorly implemented AI features, by enabling designers to identify and address user pain points more effectively. The methodology draws parallels to pilot training through simulation, suggesting a shift from traditional data collection to active AI-driven feedback loops within the design process.
- The RAG Stack We Landed On After 37 Fails - Jonathan Fernandes
This talk details the iterative process of building a Retrieval Augmented Generation (RAG) system, highlighting lessons learned from 37 failed attempts. It emphasizes practical choices for different components of a RAG stack, distinguishing between prototyping and production environments. The core thesis is that careful selection and integration of components like orchestration, embedding models, vector databases, and language models are crucial for effective RAG implementation.
- Luminal - Search-Based Deep Learning Compilers - Joe Fioti
Luminal is a deep learning library that aims for radical simplification through search-based compilation. Instead of the complex, multi-million line codebases of traditional libraries like PyTorch, Luminal represents models as simple directed acyclic graphs of a minimal set of core operations. This simplified representation allows for a compiler that uses search to discover highly optimized kernels, including complex ones like FlashAttention, which were previously difficult to implement.
- The Benchmarks Game: Why It's Rigged and How You Can (Really) Win - Darius Emrani
This talk argues that current AI benchmarks are fundamentally flawed and often manipulated, leading to misleading performance claims. The speaker contends that the immense financial and market value tied to benchmark scores incentivizes companies to game the system through various deceptive practices. Ultimately, the presentation advocates for building custom, use-case-specific evaluations rather than relying on public benchmarks.
- Stop Ordering AI Takeout A Cookbook for Winning When You Build In House - Jan Siml
This talk argues against the common practice of adopting complex, cutting-edge AI solutions for internal business needs, likening it to ordering expensive takeout when a simpler, in-house meal would suffice. The core thesis is that building AI solutions internally, when focused on specific, high-value workflows and leveraging existing data, can deliver significant revenue and operational improvements more effectively than off-the-shelf, overly complex systems.
- Buy Now, Maybe Pay Later: Dealing with Prompt-Tax While Staying at the Frontier - Andrew Thomspson
This talk addresses the challenges of building and shipping AI agentic products at the cutting edge of model development. It introduces the concept of "prompt tax," which refers to the unintended consequences and risks associated with integrating rapidly advancing AI models into applications. The core thesis is that to maximize opportunities from new AI capabilities, developers must ship products quickly, embracing the inherent uncertainties and managing them through iterative feedback and progressive rollout strategies.
- Cognitive Shield Real Time Real Smart - Rachna Srivastava
This talk introduces Cognitive Shield, a platform designed to combat the rising tide of AI-driven fraud. It highlights how advanced AI and machine learning can be leveraged not only to detect but also to prevent sophisticated scams, synthetic identities, and deepfake threats that bypass traditional security measures. The core thesis is that the same AI technologies used to perpetrate fraud can be repurposed to build robust defenses, thereby rebuilding and strengthening trust in the digital age.
- Unlocking Africa's Potential with AI — Thabang Ledwaba
This talk explores the transformative potential of Artificial Intelligence for Africa, arguing that the continent possesses immense creativity and a unique perspective that, when combined with AI, can lead to groundbreaking innovations. It challenges the perception of Africa as merely a consumer of technology and advocates for a shift towards becoming a producer and active player on the global stage. The core thesis is that by reimagining how Africa is perceived and by leveraging AI, the continent can overcome its challenges and unlock its full potential.
- Analyzing 10,000 Sales Calls With AI In 2 Weeks — Charlie Guo
This talk details how an AI engineer analyzed 10,000 sales calls in two weeks, a task previously requiring extensive manual effort or teams over months. The core thesis is that modern large language models, when integrated with sound engineering practices, can transform massive amounts of unstructured data into actionable insights, turning potential liabilities into valuable assets. The project highlights the feasibility of single engineers tackling complex data analysis challenges that were once insurmountable.
- Letting AI Interface with your App with MCP — Kent C Dodds
This talk introduces Model Context Protocol (MCP) as a standardized way for AI assistants to interface with various tools and services, aiming to bridge the gap between current AI capabilities and the desire for a universal assistant like Jarvis. It posits that the primary barrier to advanced AI assistants has been the difficulty of building numerous integrations, and MCP offers a solution by establishing a common protocol that any AI assistant can use to interact with any service provider.
- ChatGPT is poorly designed. So I fixed it
This talk argues that ChatGPT's user interface is poorly designed, leading to a confusing experience despite its rapid growth. The presenter demonstrates issues with its voice and text interaction, highlighting how separate functionalities feel disconnected. The core thesis is that by integrating multimodal capabilities and intelligently routing requests to appropriate models, the user experience can be significantly improved using off-the-shelf tools.
- Effective AI Agents Need Data Flywheels, Not The Next Biggest LLM – Sylendran Arunagiri, NVIDIA
This talk argues that building effective AI agents relies on data flywheels rather than simply using the largest available language models. A data flywheel is a continuous cycle of data processing, model customization, evaluation, and safety guardrailing. This process allows agents to refine their performance over time by learning from user feedback and production data, ultimately enabling the use of smaller, more cost-effective models without sacrificing accuracy.
- open-rag-eval: RAG Evaluation without \"golden\" answers — Ofer Mendelevitch, Vectara
This talk introduces open-rag-eval, an open-source project designed for scalable Retrieval-Augmented Generation (RAG) evaluation. It addresses the common challenge of RAG evaluation requiring "golden answers" or "golden chunks," which is often impractical and non-scalable. The project, developed in collaboration with the University of Waterloo, offers a research-backed approach to evaluate RAG pipelines without relying on pre-defined correct answers.
- Designing AI To Scale Human Thought — Jun Yu Tan, Tusk
This talk proposes a paradigm shift in AI interface design, moving from pure automation to augmentation that enhances human capabilities. Instead of AI systems attempting to automate complex tasks suboptimally, the focus should be on using AI to help humans produce higher quality work. This approach emphasizes interaction patterns that help users identify blind spots, foster creativity, and amplify thoughtful decision-making, ultimately aiming for trustworthy human-AI partnerships that grow with the user.
- The Future of Qwen: A Generalist Agent Model — Junyang Lin, Alibaba Qwen
This talk introduces the Qwen series of large language and multimodal models from Alibaba, focusing on their development towards generalist agent models. It highlights recent advancements in Qwen 3, including its hybrid thinking mode, extensive language support, and enhanced capabilities for agents and coding. The presentation also touches upon the future direction of AI development, emphasizing training agents through reinforcement learning and scaling multimodal capabilities.