World's Fair 2024
The ROI of AI: Why you need Eval Framework - Beyang Liu
Overview
This talk addresses the challenge of measuring the return on investment (ROI) of AI tools, particularly for developers. The core thesis is that while AI can significantly enhance developer productivity, demonstrating its business impact requires moving beyond anecdotal evidence or simple engagement metrics to more structured evaluation frameworks. The speaker emphasizes that the difficulty in measuring AI ROI is akin to the NP-hard problem of measuring developer productivity itself, suggesting that a precise, universally applicable solution is unlikely, but practical frameworks can be developed.
Who should watch
- Engineering leaders (Heads of Engineering, VPs)
- Product Managers (PMs)
- Individual Contributors (ICs) interested in evaluating AI tools
- Anyone struggling to quantify the business value of AI in software development
- Those facing tension between engineering and finance departments regarding AI tool costs and benefits
Key takeaways
- AI tools help developers bridge gaps in understanding and stay in flow, enabling them to focus on building features rather than getting sidetracked by dependencies or framework learning.
- Measuring AI ROI is complex because it often boils down to measuring developer productivity, which is itself an intractable problem.
- The "roles eliminated" framework, while theoretically sound for labor-saving tools, is rarely applied in software engineering as organizations prioritize building better user experiences over reducing headcount.
- AB testing velocity involves comparing a group using the AI tool against a control group to measure acceleration in task completion, though it requires significant effort and careful consideration of confounding factors.
- Time saved as a function of engagement tracks product metrics for time-saving actions, assigning an estimated time value to each, but this often provides a lower bound of the true value.
- Focusing on key performance indicators (KPIs) correlated with engineering quality and business impact, rather than generic metrics like lines of code, is crucial for effective AI ROI measurement.
- Impact on key initiatives, such as accelerating major code migrations or achieving specific OKRs, offers a high-level view of AI's value by quantifying the benefit of bringing forward critical projects.
- Surveys of developer satisfaction and preference, conducted within allocated budgets for productivity tools, provide valuable qualitative data, though they are less rigorous than quantitative methods.
Notable quotes
*Measuring AI ROI reduces to measuring developer productivity or the productivity of whatever class of knowledge worker that you're managing.*
*The most important thing I think is Define clear success criteria.*
*We think that the the next phase of evolution for us is going from kind of like these inline code completion scenarios to more what we call online agents that live in your editor but can still react to human feedback and guidance.*
Unofficial community note. Prefer the recording for nuance.