← Browse

World's Fair 2025

The End of Awkward AI Transcriptions - Travis Bartley and Myungjong Kim

Travis Bartley , Myungjong Kim

Overview

This talk details Nvidia's approach to developing enterprise-level speech AI models, focusing on robustness, coverage, personalization, and deployment efficiency. The core thesis is that a variety of specialized models, rather than a single monolithic solution, best meets diverse customer needs for conversational AI, emphasizing low latency and high efficiency for embedded devices.

Who should watch

Key takeaways

Notable quotes

*Our focus is generally on low latency, highly efficient models that can be used on embedded devices.*
*We have the model to meet the need as opposed to an idea that one model fits all.*
*The majority of the top five models do come from Nvidia, and all of it does is come down to this approach on a focus on customization and variety.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.