Send email Copy Email Address

Email

Address

Im Oberen Werk 1
66386 St. Ingbert (Germany)

Awards (selection)

2025: AI 2000 Most Influential Scholar Award Honorable Mention 2025

2025: Best Machine Learning and Security Paper in Cybersecurity Award 2025

2022: Busy Beaver Award für "Privacy of Machine Learning"

2021: Busy Beaver teaching award for seminar “Privacy of Machine Learning” at Saarland University (2021 Winter)

2019: Best paper award at NDSS 

Short Bio

Dr. Yang Zhang is tenured faculty member at CISPA. His research concentrates on trustworthy machine learning (privacy, safety, and security). Moreover, he works on measuring and understanding misinformation and unsafe content like hateful memes on the Internet. Over the years, he has published multiple papers at top venues in computer science, including CCS, NDSS, Oakland, and USENIX Security. His work has received the NDSS 2019 distinguished paper award and the CCS 2022 best paper award runner-up.

CV: Last stations

Since 2020
Faculty at CISPA Helmholtz Center for Information Security
2019 - 2020
Research Group Leader at CISPA Helmholtz Center for Information Security
2017 - 2018
Postdoctoral Researcher - Host: Michael Backes - CISPA, Saarland University
2012 - 2016
Ph.D. in Computer Science at University of Luxembourg, highest honor

Publications by Yang Zhang

Year 2025

Conference / Medium

Conference on Neural Information Processing Systems (NeurIPS) Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency

Conference / Medium

Conference on Neural Information Processing Systems (NeurIPS) Finding and Reactivating Post-Trained LLMs’ Hidden Safety Mechanisms

Conference / Medium

Conference on Empirical Methods in Natural Language Processing (EMNLP) Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification

Conference / Medium

IEEE International Conference on Computer Vision (ICCV) Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions

Conference / Medium

ACM Conference on Computer and Communications Security (CCS) UnsafeBench: Benchmarking Image Safety Classifiers onReal-World and AI-Generated Images

Article

IEEE Transactions on Dependable and Secure Computing Revealing the Risk of Hyper-parameter Leakage in Deep Reinforcement Learning Models

Conference / Medium

Usenix Security Symposium (USENIX-Security) Data Duplication: A Novel Multi-Purpose Attack Paradigm in Machine Unlearning

Conference / Medium

Usenix Security Symposium (USENIX-Security) Bridging the Gap in Vision Language Models in IdentifyingUnsafe Concepts Across Modalities

Conference / Medium

Usenix Security Symposium (USENIX-Security) On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts

Conference / Medium

Usenix Security Symposium (USENIX-Security) Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data