AI Red Teaming
Agentic AI systems are powerful but remain vulnerable to adversarial manipulation and jailbreaking. We advance their robustness by continuously probing and debugging them through red teaming. This raises the following question.
How do we train red agents that continuously debug target AI systems to harden them?
Keywords
- Learning for Jailbreaking
- Learning for Indirect Prompt Injection
3 Related Publications
International Conference on Machine Learning (ICML), 2026
ICML