Zhengyue Zhao

dblp:314/0023 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 47% Trustworthy machine learning · 21% Generative modeling · 16%
Network and information security
2 papers
Security and privacy of machine learning · 59% Digital forensics and information hiding · 22% Privacy and data protection · 19%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › decoding
contrastive decoding
1.012026
Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions · AAAI 2026
Natural language and speech › Language models and text generation
large language model safety
1.012026
Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions · AAAI 2026
Machine learning › Trustworthy machine learning › AI safety
safety alignment
1.012026
Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions · AAAI 2026
Digital forensics and information hiding › watermarking › text watermarking
LLM watermarking
0.912025
Can Watermarks be Used to Detect LLM IP Infringement For Free? · ICLR 2025
Security and privacy of machine learning
model stealing
0.912025
Can Watermarks be Used to Detect LLM IP Infringement For Free? · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.812024
Can Protective Perturbation Safeguard Personal Data from Being Exploited by Stable Diffusion? · CVPR 2024
Security and privacy of machine learning › adversarial attack
adversarial perturbation
0.812024
Can Protective Perturbation Safeguard Personal Data from Being Exploited by Stable Diffusion? · CVPR 2024
Privacy and data protection
image privacy
0.812024
Can Protective Perturbation Safeguard Personal Data from Being Exploited by Stable Diffusion? · CVPR 2024
Security and privacy of machine learning › adversarial attack › adversarial perturbation
protective perturbation
0.812024
Can Protective Perturbation Safeguard Personal Data from Being Exploited by Stable Diffusion? · CVPR 2024
Data mining
causal inference
0.612022
A Counterfactual Modeling Framework for Churn Prediction · WSDM 2022
Data mining › predictive analytics
churn prediction
0.612022
A Counterfactual Modeling Framework for Churn Prediction · WSDM 2022
Natural language and speech › Language models and text generation
large language model
0.312025
Can Watermarks be Used to Detect LLM IP Infringement For Free? · ICLR 2025

Methods — techniques the papers use, named apart from their topics

watermark detection · 1.7anchor model query selection · 1.7adaptive thresholding · 1.7purification · 1.5fine-tuning · 1.5adversarial perturbation · 1.5prompt tuning · 1.0contrastive decoding · 1.0graph neural network · 0.6counterfactual data augmentation · 0.6
YearPublicationVenuePosition
2026 Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions
abstract
With the widespread application of Large Language Models (LLMs), it has become a significant concern to ensure their safety and prevent harmful responses. While current safe-alignment methods based on instruction fine-tuning and Reinforcement Learning from Human Feedback (RLHF) can effectively reduce harmful responses from LLMs, they often require high-quality datasets and heavy computational overhead during model training. Another way to align language models is to modify the logit of tokens in model outputs without heavy training. Recent studies have shown that contrastive decoding can enhance the performance of language models by reducing the likelihood of confused tokens. However, these methods require the manual selection of contrastive models or instruction templates, limiting the degree of contrast. To this end, we propose Adversarial Contrastive Decoding (ACD), an optimization-based framework to generate two opposite soft system prompts, the Safeguarding Prompt (SP) and the Adversarial Prompt (AP), for prompt-based contrastive decoding. The SP aims to promote safer outputs while the AP aims to exploit the harmful parts of the model, providing a strong contrast to align the model with safety. ACD only needs to apply a lightweight prompt tuning on a rather small anchor dataset without training the target model. Experiments conducted on extensive models and benchmarks demonstrate that the proposed method achieves much better safety performance than previous model training-free decoding methods without sacrificing its original generation ability.
Zhengyue Zhao, Kaidi Xu, Xing Hu 0001
AAAI2
2025 Can Watermarks be Used to Detect LLM IP Infringement For Free?
abstract
The powerful capabilities of LLMs stem from their rich training data and high-quality labeled datasets, making the training of strong LLMs a resource-intensive process, which elevates the importance of IP protection for such LLMs. Compared to gathering high-quality labeled data, directly sampling outputs from these fully trained LLMs as training data presents a more cost-effective approach. This practice—where a suspect model is fine-tuned using high-quality data derived from these LLMs, thereby gaining capabilities similar to the target model—can be seen as a form of IP infringement against the original LLM. In recent years, LLM watermarks have been proposed and used to detect whether a text is AI-generated. Intuitively, if data sampled from a watermarked LLM is used for training, the resulting model would also be influenced by this watermark. This raises the question: can we directly use such watermarks to detect IP infringement of LLMs? In this paper, we explore the potential of LLM watermarks for detecting model infringement. We find that there are two issues with direct detection: (1) The queries used to sample output from the suspect LLM have a significant impact on detectability. (2) The watermark that is easily learned by LLMs exhibits instability regarding the watermark's hash key during detection. To address these issues, we propose LIDet, a detection method that leverages available anchor LLMs to select suitable queries for sampling from the suspect LLM. Additionally, it adapts the detection threshold to mitigate detection failures caused by different hash keys. To demonstrate the effectiveness of this approach, we construct a challenging model set containing multiple suspect LLMs on which direct detection methods struggle to yield effective results. Our method achieves over 90\% accuracy in distinguishing between infringing and clean models, demonstrating the feasibility of using LLM watermarks to detect LLM IP infringement.
Zhengyue Zhao, Xiaogeng Liu, Somesh Jha, Patrick McDaniel, Bo Li 0026, Chaowei Xiao
ICLR1
2024 Can Protective Perturbation Safeguard Personal Data from Being Exploited by Stable Diffusion?
abstract
Stable Diffusion has established itself as a foundation model in generative AI artistic applications, receiving widespread research and application. Some recent fine-tuning methods have made it feasible for individuals to implant personalized concepts onto the basic Stable Diffusion model with minimal computational costs on small datasets. However, these innovations have also given rise to issues like facial privacy forgery and artistic copyright infringement. In recent studies, researchers have explored the addition of imperceptible adversarial perturbations to images to prevent potential unauthorized exploitation and infringements when personal data is used for fine-tuning Stable Dif-fusion. Although these studies have demonstrated the ability to protect images, it is essential to consider that these methods may not be entirely applicable in real-world scenarios. In this paper, we systematically evaluate the use of perturbations to protect images within a practical threat model. The results suggest that these approaches may not be sufficient to safeguard image privacy and copyright effectively. Furthermore, we introduce a purification method capable of removing protected perturbations while preserving the original image structure to the greatest extent possible. Experiments reveal that Stable Diffusion can effectively learn from purified images over all protective methods1.
Zhengyue Zhao, Jinhao Duan, Kaidi Xu, Chenan Wang, Rui Zhang 0040, Zidong Du, Qi Guo 0001, Xing Hu 0001
CVPR1
2024 Automated CPU Design by Learning from Input-Output Examples
Shuyao Cheng, Pengwei Jin, Qi Guo 0001, Zidong Du, Rui Zhang 0040, Xing Hu 0001, Yongwei Zhao 0001, Yifan Hao 0001, Xiangtao Guan, Husheng Han, Zhengyue Zhao, Xishan Zhang, Yuejie Chu, Weilong Mao, Tianshi Chen 0002, Yunji Chen
IJCAI11
2022 A Counterfactual Modeling Framework for Churn Prediction
abstract
Accurate churn prediction for retaining users is keenly important for online services because it determines their survival and prosperity. Recent research has specified social influence to be one of the most important reasons for user churn, and thereby many works start to model its effects on user churn to improve the prediction performance. However, existing works only use the data's correlational information while neglecting the problem's causal nature. Specifically, the fact that a user's churn is correlated with some social factors does not mean he/she is actually influenced by his/her friends, which results in inaccurate and unexplainable predictions of the existing methods. To bridge this gap, we develop a counterfactual modeling framework for churn prediction, which can effectively capture the causal information of social influence for accurate and explainable churn predictions. Specifically, we first propose a backbone framework that uses two separate embeddings to model users' endogenous churn intentions and the exogenous social influence. Then, we propose a counterfactual data augmentation module to introduce the causal information to the model by providing partially labeled counterfactual data. Finally, we design a three-headed counterfactual prediction framework to guide the model to learn causal information to facilitate churn prediction. Extensive experiments on two large-scale datasets with different types of social relations show our model's superior prediction performance compared with the state-of-the-art baselines. We further conduct an in-depth analysis of the prediction results demonstrating our proposed method's ability to capture causal information of social influence and give explainable churn predictions, which provide insights into designing better user retention strategies.
Guozhen Zhang 0001, Jinwei Zeng, Zhengyue Zhao, Depeng Jin, Yong Li 0008
WSDM3