AI Red Teaming
Agentic AI systems are powerful but remain vulnerable to adversarial manipulation and jailbreaking. We advance their robustness by continuously probing and debugging them through red teaming. This raises the following question.
How do we train red agents that continuously debug target AI systems to harden them?
Problems
- Learning for LLM Jailbreaking
- Learning for Agentic AI Jailbreaking