World's Fair 2025
Perceptual Evaluations: Evals for Aesthetics — Diego Rodriguez, Krea.ai
Overview
This talk explores the limitations of current AI evaluation metrics, particularly in generative media, by highlighting how they often fail to account for human perception and aesthetic judgment. The speaker argues that traditional metrics, like FID scores, can be misled by factors such as compression artifacts, leading to inaccurate assessments of AI model quality. The core thesis is that AI evaluation needs to evolve to incorporate human perceptual nuances and subjective qualities, moving beyond easily quantifiable but potentially superficial measures.
Who should watch
- AI engineers and researchers working on generative models for images, video, or audio.
- Product Managers and builders evaluating the quality and user experience of AI-generated content.
- Anyone interested in the challenges of measuring subjective qualities like aesthetics and perception in AI.
- Developers seeking to improve the robustness and perceptual accuracy of AI evaluation frameworks.
Key takeaways
- Current AI evaluation metrics often overlook human perception and aesthetic judgment, leading to flawed assessments.
- Standard metrics like FID scores can be sensitive to compression artifacts (e.g., JPEG), misrepresenting the perceptual quality of images.
- AI models are trained on human data, which can embed human biases and perceptual limitations into the models themselves.
- Exploiting human perceptual limitations, as seen in JPEG and MP3 compression, is a principle that should inform AI evaluation.
- There's a need for perceptually aware metrics that can better align with human judgment, especially for subjective qualities.
- AI excels at learning from subjective data, suggesting that training classifiers on human opinions can lead to better evaluation models.
- The ability of AI to handle translation can break down communication barriers, enabling new forms of collaboration and expression.
- Future evaluations should consider the nature of training data, including common artifacts and the subjective preferences of users.
Notable quotes
*We are now in we just entered the age where you can have models essentially they solve translation right or they solved it to a very high degree.*
*Hey man if you think about it like predicting the car back when everything was horses it's not that hard.*
*What are the traffics that we're missing now?*
Unofficial community note. Prefer the recording for nuance.