Annual Meeting of the Association for Computational Linguistics (ACL) Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
International Conference on Machine Learning (ICML) Provably Cost-Sensitive Adversarial Defense via Randomized Smoothing
European Conference on Artificial Intelligence (ECAI) Inside the Black Box: Detecting Data Leakage in Pre-trained Language Encoders
ICML-Workshop (ICMLW) Provably Robust Cost-Sensitive Learning via Randomized Smoothing