AI Red Teaming

Agentic AI systems are powerful but remain vulnerable to adversarial manipulation and jailbreaking. We advance their robustness by continuously probing and debugging them through red teaming. This raises the following question.

How do we train red agents that continuously debug target AI systems to harden them?

Keywords

  • Learning for Jailbreaking
  • Learning for Indirect Prompt Injection

3 Related Publications

Junyoung Park, Namgyu Park, Sechan Lee, Yoon-Chan Jhi, Jihoon Cho, Sangdon Park
2026
Jeongyeon Hwang, Sangdon Park, Jungseul Ok
International Conference on Machine Learning (ICML), 2026
ICML