EDBT 2026 Demo / reviewers in the wild / expert
Chongyu Fan
dblp:359/3239
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Trustworthy machine learning · 67% Generative modeling · 24% Language models and text generation · 6% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
machine unlearning |
4.0 | 5 | 2025 | Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond · ICML 2025 Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills · EMNLP 2025 Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models · NeurIPS 2024 |
Machine learning › Generative modeling
diffusion model |
2.3 | 3 | 2024 | UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion Models · NeurIPS 2024 Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models · NeurIPS 2024 SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation · ICLR 2024 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.9 | 1 | 2025 | Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature attribution |
0.9 | 1 | 2025 | The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias Optimization · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias Optimization · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › adversarial machine learning › adversarial defense
jailbreak defense |
0.9 | 1 | 2025 | Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond · ICML 2025 |
Machine learning › Trustworthy machine learning
language model interpretability |
0.9 | 1 | 2025 | The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias Optimization · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model
reasoning model |
0.9 | 1 | 2025 | Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills · EMNLP 2025 |
Machine learning › Trustworthy machine learning › interpretability › visual explanation
saliency |
0.9 | 1 | 2025 | The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias Optimization · NeurIPS 2025 |
Machine learning › Generative modeling
concept erasure |
0.8 | 1 | 2024 | SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation · ICLR 2024 |
Machine learning › Generative modeling › concept erasure
concept unlearning in diffusion models |
0.8 | 1 | 2024 | Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › machine unlearning
unlearning evaluation |
0.8 | 1 | 2024 | UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
privacy |
0.3 | 1 | 2025 | Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills · EMNLP 2025 |
Machine learning › Optimization for machine learning › gradient-based optimization
sharpness-aware minimization |
0.3 | 1 | 2025 | Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond · ICML 2025 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2024 | SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation · ICLR 2024 |
Visual content generation and editing › image generation
stylized image generation |
0.2 | 1 | 2024 | UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion Models · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
machine unlearning · 2.3unlearning · 0.9sharpness-aware minimization · 0.9robust optimization · 0.9reasoning skill preservation · 0.9needle-in-a-haystack evaluation · 0.9attention bias optimization · 0.9model retraining · 0.8gradient-based weight saliency · 0.8benchmarking · 0.8adversarial training · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning SkillsabstractChangsheng Wang, Chongyu Fan, Yihua Zhang, Jinghan Jia, Dennis Wei, Parikshit Ram, Nathalie Baracaldo, Sijia Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Changsheng Wang, Chongyu Fan, Jinghan Jia, Dennis Wei, Parikshit Ram, Nathalie Baracaldo, Sijia Liu 0001 |
EMNLP | 2 |
| 2025 | Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and BeyondabstractThe LLM unlearning technique has recently been introduced to comply with data regulations and address the safety and ethical concerns of LLMs by removing the undesired data-model influence.
However, state-of-the-art unlearning methods face a critical vulnerability: they are susceptible to ``relearning'' the removed information from a small number of forget data points, known as relearning attacks. In this paper, we systematically investigate how to make unlearned models robust against such attacks. For the first time, we establish a connection between robust unlearning and sharpness-aware minimization (SAM) through a unified robust optimization framework, in an analogy to adversarial training designed to defend against adversarial attacks. Our analysis for SAM reveals that smoothness optimization plays a pivotal role in mitigating relearning attacks. Thus, we further explore diverse smoothing strategies to enhance unlearning robustness. Extensive experiments on benchmark datasets, including WMDP and MUSE, demonstrate that SAM and other smoothness optimization approaches consistently improve the resistance of LLM unlearning to relearning attacks. Notably, smoothness-enhanced unlearning also helps defend against (input-level) jailbreaking attacks, broadening our proposal's impact in robustifying LLM unlearning. Codes are available at https://github.com/OPTML-Group/Unlearn-Smooth. Chongyu Fan, Jinghan Jia, Anil Ramakrishna, Mingyi Hong 0001, Sijia Liu 0001 |
ICML | 1 |
| 2025 | Simplicity Prevails: Rethinking Negative Preference Optimization for LLM UnlearningabstractThis work studies the problem of large language model (LLM) unlearning, aiming to remove unwanted data influences (e.g., copyrighted or harmful content) while preserving model utility. Despite the increasing demand for unlearning, a technically-grounded optimization framework is lacking. Gradient ascent (GA)-type methods, though widely used, are suboptimal as they reverse the learning process without controlling optimization divergence (i.e., deviation from the pre-trained state), leading to risks of model collapse. Negative preference optimization (NPO) has been proposed to address this issue and is considered one of the state-of-the-art LLM unlearning approaches. In this work, we revisit NPO and identify another critical issue: reference model bias. This bias arises from using the reference model (i.e., the model prior to unlearning) to assess unlearning success, which can lead to a misleading impression of the true data-wise unlearning effectiveness. Specifically, it could cause (a) uneven allocation of optimization power across forget data with varying difficulty levels, and (b) ineffective gradient weight smoothing during the early stages of unlearning optimization. To overcome these challenges, we propose a simple yet effective unlearning optimization framework, called SimNPO, showing that simplicity—removing the reliance on a reference model (through the lens of simple preference optimization)—benefits unlearning. We provide deeper insights into SimNPO's advantages, including an analysis based on mixtures of Markov chains. Extensive experiments further validate its efficacy on benchmarks like TOFU, MUSE, and WMDP. Chongyu Fan, Jiancheng Liu, Licong Lin, Jinghan Jia, Song Mei, Sijia Liu 0001 |
NeurIPS | 1 |
| 2025 | The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias OptimizationabstractInput saliency aims to quantify the influence of input tokens on the output of large language models (LLMs), which has been widely used for prompt engineering, model interpretability, and behavior attribution. Despite the proliferation of saliency techniques, the field lacks a standardized and rigorous evaluation protocol. In this work, we introduce a stress-testing framework inspired by the needle-in-a-haystack (NIAH) setting to systematically assess the reliability of seven popular input saliency methods. Our evaluation reveals a surprising and critical flaw: existing methods consistently assign non-trivial importance to irrelevant context, and this attribution error worsens as input length increases. To address this issue, we propose a novel saliency method based on Attention Bias Optimization (ours), which explicitly optimizes the attention bias associated with each input token to quantify its causal impact on target token generation. ABO robustly outperforms existing methods by 10\sim30% in saliency accuracy across diverse NIAH tasks, maintains effectiveness up to 10K-token prompts, and enables practical applications including zero-shot detoxification, sentiment steering, and reasoning-error correction. Our findings highlight the limitations of prevalent attribution methods and establish ABO as a principled alternative for accurate token attribution. Changsheng Wang, Chongyu Fan, Jinghan Jia, Sijia Liu 0001 |
NeurIPS | 4 |
| 2024 | Challenging Forgets: Unveiling the Worst-Case Forget Sets in Machine Unlearning
Chongyu Fan, Jiancheng Liu, Alfred O. Hero III, Sijia Liu 0001 |
ECCV (21) | 1 |
| 2024 | SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and GenerationabstractWith evolving data regulations, machine unlearning (MU) has become an important tool for fostering trust and safety in today's AI models. However, existing MU methods focusing on data and/or weight perspectives often suffer limitations in unlearning accuracy, stability, and cross-domain applicability. To address these challenges, we introduce the concept of 'weight saliency' for MU, drawing parallels with input saliency in model explanation. This innovation directs MU's attention toward specific model weights rather than the entire model, improving effectiveness and efficiency. The resultant method that we call saliency unlearning (SalUn) narrows the performance gap with 'exact' unlearning (model retraining from scratch after removing the forgetting data points). To the best of our knowledge, SalUn is the first principled MU approach that can effectively erase the influence of forgetting data, classes, or concepts in both image classification and generation tasks. As highlighted below, For example, SalUn yields a stability advantage in high-variance random data forgetting, e.g., with a 0.2% gap compared to exact unlearning on the CIFAR-10 dataset. Moreover, in preventing conditional diffusion models from generating harmful images, SalUn achieves nearly 100% unlearning accuracy, outperforming current state-of-the-art baselines like Erased Stable Diffusion and Forget-Me-Not. Codes are available at https://github.com/OPTML-Group/Unlearn-Saliency.
**WARNING**: This paper contains model outputs that may be offensive in nature. Chongyu Fan, Jiancheng Liu, Eric Wong 0001, Dennis Wei, Sijia Liu 0001 |
ICLR | 1 |
| 2024 | Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsabstractDiffusion models (DMs) have achieved remarkable success in text-to-image generation, but they also pose safety risks, such as the potential generation of harmful content and copyright violations. The techniques of machine unlearning, also known as concept erasing, have been developed to address these risks. However, these techniques remain vulnerable to adversarial prompt attacks, which can prompt DMs post-unlearning to regenerate undesired images containing concepts (such as nudity) meant to be erased. This work aims to enhance the robustness of concept erasing by integrating the principle of adversarial training (AT) into machine unlearning, resulting in the robust unlearning framework referred to as AdvUnlearn. However, achieving this effectively and efficiently is highly nontrivial. First, we find that a straightforward implementation of AT compromises DMs’ image generation quality post-unlearning. To address this, we develop a utility-retaining regularization on an additional retain set, optimizing the trade-off between concept erasure robustness and model utility in AdvUnlearn. Moreover, we identify the text encoder as a more suitable module for robustification compared to UNet, ensuring unlearning effectiveness. And the acquired text encoder can serve as a plug-and-play robust unlearner for various DM types. Empirically, we perform extensive experiments to demonstrate the robustness advantage of AdvUnlearn across various DM unlearning scenarios, including the erasure of nudity, objects, and style concepts. In addition to robustness, AdvUnlearn also achieves a balanced tradeoff with model utility. To our knowledge, this is the first work to systematically explore robust DM unlearning through AT, setting it apart from existing methods that overlook robustness in concept erasing. Codes are available at https://github.com/OPTML-Group/AdvUnlearn.
Warning: This paper contains model outputs that may be offensive in nature. Xin Chen 0071, Jinghan Jia, Chongyu Fan, Jiancheng Liu, Mingyi Hong 0001, Sijia Liu 0001 |
NeurIPS | 5 |
| 2024 | UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion ModelsabstractThe technological advancements in diffusion models (DMs) have demonstrated unprecedented capabilities in text-to-image generation and are widely used in diverse applications. However, they have also raised significant societal concerns, such as the generation of harmful content and copyright disputes. Machine unlearning (MU) has emerged as a promising solution, capable of removing undesired generative capabilities from DMs. However, existing MU evaluation systems present several key challenges that can result in incomplete and inaccurate assessments. To address these issues, we propose UnlearnCanvas, a comprehensive high-resolution stylized image dataset that facilitates the evaluation of the unlearning of artistic styles and associated objects. This dataset enables the establishment of a standardized, automated evaluation framework with 7 quantitative metrics assessing various aspects of the unlearning performance for DMs. Through extensive experiments, we benchmark 9 state-of-the-art MU methods for DMs, revealing novel insights into their strengths, weaknesses, and underlying mechanisms. Additionally, we explore challenging unlearning scenarios for DMs to evaluate worst-case performance against adversarial prompts, the unlearning of finer-scale concepts, and sequential unlearning. We hope that this study can pave the way for developing more effective, accurate, and robust DM unlearning methods, ensuring safer and more ethical applications of DMs in the future. The dataset, benchmark, and codes are publicly available at this link. Chongyu Fan, Yuguang Yao, Jinghan Jia, Jiancheng Liu, Gaoyuan Zhang, Gaowen Liu, Ramana Rao Kompella, Xiaoming Liu 0002, Sijia Liu 0001 |
NeurIPS | 2 |