World's Fair 2025
AI Red Teaming Agent: Azure AI Foundry — Nagkumar Arkalgud & Keiji Kanazawa, Microsoft
Nagkumar Arkalgud , Keiji Kanazawa
Overview
This talk introduces the AI Red Teaming Agent, a tool developed within Azure AI Foundry to help AI engineers proactively identify and mitigate risks in their AI systems. It emphasizes that building trustworthy AI is a collaborative effort, akin to building bridges and dams, and that AI engineering requires a systematic approach to testing and iteration. The agent provides a practical way for developers to simulate attacks and evaluate their models' vulnerabilities.
Who should watch
- AI Engineers
- Product Managers
- Builders of AI applications
- Those concerned with AI safety and security
- Developers looking to implement robust testing for AI models
Key takeaways
- AI systems, especially agents, can be vulnerable to manipulation and unintended outputs, requiring rigorous testing.
- Trustworthy AI development is a team sport, necessitating collaboration between engineers and AI risk experts.
- The AI Red Teaming Agent, built on the pyrit Python package, offers a hosted solution with an SDK and dashboard for evaluating AI model security.
- The tool supports various attack strategies and risk categories, allowing for comprehensive testing of applications and direct model scans.
- Guardrails and prompt shields can be implemented within Azure AI Foundry to mitigate identified risks, acting as external filters to the raw model.
- Testing with different models, like GPT-4o versus GPT-3.5, demonstrates the impact of built-in security features and guardrails on attack success rates.
- The process involves mapping anticipated risks, implementing defenses, and then conducting evaluations, with red teaming being a key part of this cycle.
Notable quotes
*Trust is a team sport.*
*AI engineering is early. So we've got a lot of work to do to get to the point where people trust AI as much as they trust bridges.*
Unofficial community note. Prefer the recording for nuance.