Send email Copy Email Address

Email

Address

Im Oberen Werk 1
66386 St. Ingbert (Germany)

Awards (selection)

2025: AI 2000 Most Influential Scholar Award Honorable Mention 2025

2025: Best Machine Learning and Security Paper in Cybersecurity Award 2025

2022: Busy Beaver Award für "Privacy of Machine Learning"

2021: Busy Beaver teaching award for seminar “Privacy of Machine Learning” at Saarland University (2021 Winter)

2019: Best paper award at NDSS 

Short Bio

Dr. Yang Zhang is tenured faculty member at CISPA. His research concentrates on trustworthy machine learning (privacy, safety, and security). Moreover, he works on measuring and understanding misinformation and unsafe content like hateful memes on the Internet. Over the years, he has published multiple papers at top venues in computer science, including CCS, NDSS, Oakland, and USENIX Security. His work has received the NDSS 2019 distinguished paper award and the CCS 2022 best paper award runner-up.

CV: Last stations

Since 2020
Faculty at CISPA Helmholtz Center for Information Security
2019 - 2020
Research Group Leader at CISPA Helmholtz Center for Information Security
2017 - 2018
Postdoctoral Researcher - Host: Michael Backes - CISPA, Saarland University
2012 - 2016
Ph.D. in Computer Science at University of Luxembourg, highest honor

Publications by Yang Zhang

Year 2026

Article

IEEE Transactions on Information Forensics and Security BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning

Conference / Medium

Annual Meeting of the Association for Computational Linguistics (ACL)

Conference / Medium

Annual Meeting of the Association for Computational Linguistics (ACL) The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training

Conference / Medium

European Association for Computational Linguistics (EACL) Defeating Cerberus: Privacy-Leakage Mitigation in Vision Language Models

Article

IEEE Transactions on Dependable and Secure Computing Backdoor Complications: A Comprehensive Analysis and Mitigation of the Unforeseen Consequences of Backdoor Attacks

Conference / Medium

National Conference of the American Association for Artificial Intelligence (AAAI) SL-CBM: Enhancing Concept Bottleneck Models with Semantic Locality for Better Interpretability

Conference / Medium

Annual Meeting of the Association for Computational Linguistics (ACL) Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs

Conference / Medium

International Conference on Machine Learning (ICML) Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs

Conference / Medium

Annual Meeting of the Association for Computational Linguistics (ACL) InferPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents

Conference / Medium

Annual Meeting of the Association for Computational Linguistics (ACL) Rethinking Assessments of Prompt Injection Attacks