Agentic AI Alignment

Agentic AI systems act autonomously over long horizons, making alignment increasingly critical and challenging. We develop principled methods to ensure these agents remain truthful, safe, and secure — aligned with multiple objectives throughout deployment.

How do we train AI agents that remain aligned with multiple objectives?

Problems

  • Constrained Learning
  • Multi-objective Learning
  • Learning for Planning
  • Learning for SW Bug Discovery and Patching

10 Related Publications

Sechan Lee, Hyounghun Kim, Sangdon Park
The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP Findings), 2026
Jaewan Choi*, Junyoung Yang*, Sangdon Park
The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP Findings), 2026
🏆 Best Paper Award from CKAIA
Kyungmin Kim*, Youngbin Choi*, Seoyeon Lee, Suhyeon Jun, Dongwoo Kim^, Sangdon Park^
The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP Findings), 2026
ICML26' RLxF Workshop
The Illusion of Rust Safety: Detecting Modular Unsafe Functions with LLMs
Xiang Cheng, Fan Sang, Yibin Yang, Hang Zhang, Sangdon Park, Xiaokuan Zhang, Taesoo Kim
The ACM Conference on Computer and Communications Security (CCS), 2026
CCS
Xiang Cheng*, Sangdon Park*, HyungSeok Han, Xiaokuan Zhang, Taesoo Kim
International Conference on Dependable Systems and Networks (DSN), 2026
Kyungmin Kim, Youngbin Choi, Hyounghun Kim, Dongwoo Kim, Sangdon Park
Findings of the Association for Computational Linguistics: EMNLP (EMNLP Findings), 2025
Saemi Moon*, Minjong Lee*, Sangdon Park, Dongwoo Kim
International Conference on Computer Vision (ICCV), 2025
ICCV
Jeongyeon Hwang, Junyoung Park, Hyejin Park, Dongwoo Kim, Sangdon Park, Jungseul Ok
Empirical Methods in Natural Language Processing (EMNLP), 2025
EMNLP
Taesoo Kim, HyungSeok Han, Soyeon Park, Dae R. Jeong, Dohyeok Kim, Dongkwan Kim, Eunsoo Kim, Jiho Kim, Joshua Wang, Kangsu Kim, Sangwoo Ji, Woosun Song, Hanqing Zhao, Andrew Chin, Gyejin Lee, Kevin Stevens, Mansour Alharthi, Yizhuo Zhai, Cen Zhang, Joonun Jang, Yeongjin Jang, Ammar Askar, Dongju Kim, Fabian Fleischer, Jeongin Cho, Junsik Kim, Kyungjoon Ko, Insu Yun, Sangdon Park, Dowoo Baik, Haein Lee, Hyeon Heo, Minjae Gwon, Minjae Lee, Minwoo Baek, Seunggi Min, Wonyoung Kim, Yonghwi Jin, Younggi Park, Yunjae Choi, Jinho Jung, Gwanhyun Lee, Junyoung Jang, Kyuheon Kim, Yeonghyeon Cha, Youngjoon Kim
2025
🏆 DARPA AIxCC Winner - $(4+2)M Award
Minjae Lee*, Yoonjae Jung*, Sangdon Park
2025
🏆 Best Paper Finalist from CKAIA ICLR'26 Reliable Autonomy Workshop