Agentic AI Alignment
Agentic AI systems act autonomously over long horizons, making alignment increasingly critical and challenging. We develop principled methods to ensure these agents remain truthful, safe, and secure — aligned with multiple objectives throughout deployment.
How do we train AI agents that remain aligned with multiple objectives?
Problems
- Constrained Learning
- Multi-objective Learning
- Learning for Planning
- Learning for SW Bug Discovery and Patching
10 Related Publications
The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP Findings), 2026
The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP Findings), 2026
🏆 Best Paper Award from CKAIA
The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP Findings), 2026
ICML26' RLxF Workshop
The Illusion of Rust Safety: Detecting Modular Unsafe Functions with LLMs
The ACM Conference on Computer and Communications Security (CCS), 2026
CCS
International Conference on Dependable Systems and Networks (DSN), 2026
Findings of the Association for Computational Linguistics: EMNLP (EMNLP Findings), 2025
International Conference on Computer Vision (ICCV), 2025
ICCV
Empirical Methods in Natural Language Processing (EMNLP), 2025
EMNLP
2025
🏆 DARPA AIxCC Winner - $(4+2)M Award
2025
🏆 Best Paper Finalist from CKAIA
ICLR'26 Reliable Autonomy Workshop