Annual Meeting of the Association for Computational Linguistics (ACL) PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality
Annual Meeting of the Association for Computational Linguistics (ACL) Peering Behind the Shield: Guardrail Identification in Large Language Models
IEEE Transactions on Information Forensics and Security BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
IEEE International Conference on Computer Vision (ICCV) Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
Annual Meeting of the Association for Computational Linguistics (ACL) JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
Usenix Security Symposium (USENIX-Security) SecurityNet: Assessing Machine Learning Vulnerabilities on Public Models
International Conference on Learning Representations (ICLR) Data Poisoning Attacks Against Multimodal Encoders