World's Fair 2025
Does AI Actually Boost Developer Productivity? (100k Devs Study) - Yegor Denisov-Blanch, Stanford
Yegor Denisov-Blanch , Stanford
Overview
Summary pending.
Continue
Related
- (possible dupe but better sound) What does Enterprise Ready MCP mean? — Tobin South, WorkOS
This talk explores what it means for Model Communication Protocol (MCP) to be enterprise-ready, moving beyond basic tool-use capabilities to address the complexities of production environments. It highlights the evolution from simple chatbot interactions to sophisticated AI agents that require robust security, management, and scalability, particularly when integrating with internal enterprise systems. The core thesis is that while MCP offers a standardized way for AI to interact with external resources, achieving enterprise readiness involves overcoming significant challenges in authentication, authorization, and operational management.
- [Evals Workshop] Mastering AI Evaluation: From Playground to Production
This talk focuses on mastering AI evaluation, moving from initial development to production. It emphasizes that even the best large language models (LLMs) require a robust testing framework due to issues like hallucinations and performance degradation with changes. Effective evaluation helps answer critical questions about model selection, cost-effectiveness, brand consistency, and ongoing improvement, ultimately reducing development time, costs, and enabling faster iteration.
- [Full Workshop] Building Conversational AI Agents - Thor Schaeff, ElevenLabs
This workshop focuses on building multilingual conversational AI agents, detailing the pipeline from speech-to-text to text-to-speech. It highlights the integration of large language models as the agent's "brain" and showcases ElevenLabs' tools for creating dynamic, responsive AI interactions across numerous languages. The session emphasizes practical application and developer experience, offering insights into configuring and deploying these agents.
- [Full Workshop] Building Metrics that actually work — David Karam, Pi Labs (fmr Google Search)
This workshop focuses on building effective evaluation metrics for AI systems, moving beyond basic testing to create robust scoring systems. It emphasizes that evaluations are not just for testing but are the primary place where domain knowledge resides, enabling significant improvements in AI development. The session introduces a methodology for creating nuanced, calibrated metrics that correlate with desired outcomes, ultimately simplifying the AI development stack.
- [Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han
This talk delves into advanced AI training techniques, focusing on reinforcement learning (RL), quantization, and agent development. It explores the evolution of large language models, from early open-source efforts spurred by leaks to current sophisticated training methodologies. The discussion highlights the critical role of fine-tuning stages, including supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF), in transforming base models into capable conversational agents.
Unofficial community note. Prefer the recording for nuance.