Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Renjie Gu

dblp:245/3316 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0003-2895-444XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 60% Trustworthy machine learning · 21% Transfer learning and domain adaptation · 12%
Databases, data mining, and information retrieval
1 paper
Web and social media mining · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › distributed training › edge training
device-cloud collaborative learning
1.122022
Walle: An End-to-End, General-Purpose, and Large-Scale Production System for Device-Cloud Collaborative Machine Learning · OSDI 2022
On-Device Learning for Model Personalization with Large-Scale Cloud-Coordinated Domain Adaption · KDD 2022
Machine learning › Trustworthy machine learning
machine unlearning
1.012026
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning · ACL (1) 2026
Machine learning › Efficient and distributed learning
distributed training
0.612022
Walle: An End-to-End, General-Purpose, and Large-Scale Production System for Device-Cloud Collaborative Machine Learning · OSDI 2022
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.612022
On-Device Learning for Model Personalization with Large-Scale Cloud-Coordinated Domain Adaption · KDD 2022
Machine learning › Efficient and distributed learning
federated learning
0.612022
On-Device Learning for Model Personalization with Large-Scale Cloud-Coordinated Domain Adaption · KDD 2022
Machine learning › Efficient and distributed learning › edge computing › on-device machine learning
on-device learning
0.612022
On-Device Learning for Model Personalization with Large-Scale Cloud-Coordinated Domain Adaption · KDD 2022
Cloud and datacenter computing
cloud infrastructure
0.612022
Walle: An End-to-End, General-Purpose, and Large-Scale Production System for Device-Cloud Collaborative Machine Learning · OSDI 2022
Web and social media mining › misinformation detection
fake news detection
0.412019
Unsupervised Fake News Detection on Social Media: A Generative Approach · AAAI 2019
Web and social media mining
misinformation detection
0.412019
Unsupervised Fake News Detection on Social Media: A Generative Approach · AAAI 2019
Web and social media mining
social media analysis
0.412019
Unsupervised Fake News Detection on Social Media: A Generative Approach · AAAI 2019
Natural language and speech › Language models and text generation
large language model safety
0.312026
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

representation alignment · 1.0fine-tuning · 1.0feature randomization · 1.0retrieval of similar data · 0.6domain adaptation · 0.6latent variable inference · 0.4collapsed gibbs sampling · 0.4bayesian network · 0.4
YearPublicationVenuePosition
2026 Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
abstract
Unlearning in large language models (LLMs) aims to remove harmful training data while preserving overall utility.However, we find that existing methods often hallucinate, generate abnormal token sequences, or behave inconsistently, raising safety and trust concerns.According to prior literature on LLM honesty, such behaviors are often associated with dishonesty.This motivates us to investigate the notion of honesty in the context of model unlearning.We propose a formal definition of unlearning honesty, which includes: (1) preserving both utility and honesty on retained knowledge, and (2) ensuring effective forgetting while encouraging the model to acknowledge its limitations and respond consistently to questions related to forgotten knowledge.To systematically evaluate the honesty of unlearning, we introduce a suite of metrics that cover utility, honesty on the retained set, effectiveness of forgetting, rejection rate and refusal stability in Q&A and MCQ settings.Evaluating 9 methods across 3 mainstream families shows that all current methods fail to meet these standards.After experimental and theoretical analyses, we present ReVa, a representation-alignment procedure that fine-tunes feature-randomized unlearned models to better acknowledge forgotten knowledge.On Q&A tasks from the forget set, ReVa achieves the highest rejection rate after two rounds of interaction, nearly doubling the performance of the second-best method.Remarkably, It also improves honesty on the retained set.We release our data and code at https://github.com/renjiegu.
Renjie Gu, Jiazhen Du, Sijia Liu 0001
ACL (1)1
2025 Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning
abstract
Text-to-image (T2I) diffusion models have achieved impressive image generation quality and are increasingly fine-tuned for personalized applications. However, these models often inherit unsafe behaviors from toxic pretraining data, raising growing safety concerns. While recent safety-driven unlearning methods have made promising progress in suppressing model toxicity, they are found to be fragile to downstream fine-tuning, as we reveal that state-of-the-art methods largely fail to retain their effectiveness even when fine-tuned on entirely benign datasets. To mitigate this problem, in this paper, we propose ResAlign, a safety-driven unlearning framework with enhanced resilience against downstream fine-tuning. By modeling downstream fine-tuning as an implicit optimization problem with a Moreau envelope-based reformulation, ResAlign enables efficient gradient estimation to minimize the recovery of harmful behaviors. Additionally, a meta-learning strategy is proposed to simulate a diverse distribution of fine-tuning scenarios to improve generalization. Extensive experiments across a wide range of datasets, fine-tuning methods, and configurations demonstrate that ResAlign consistently outperforms prior unlearning approaches in retaining safety, while effectively preserving benign generation capability. Our code and pretrained models are publicly available at https://github.com/AntigoneRandy/ResAlign.
Boheng Li, Renjie Gu, Junjie Wang 0007, Leyi Qi, Yiming Li 0004, Run Wang 0001, Zhan Qin, Tianwei Zhang 0004
NeurIPS2
2022 On-Device Learning for Model Personalization with Large-Scale Cloud-Coordinated Domain Adaption
abstract
Cloud-based learning is currently the mainstream in both academia and industry. However, the global data distribution, as a mixture of all the users' data distributions, for training a global model may deviate from each user's local distribution for inference, making the global model non-optimal for each individual user. To mitigate distribution discrepancy, on-device training over local data for model personalization is a potential solution, but suffers from serious overfitting. In this work, we propose a new device-cloud collaborative learning framework under the paradigm of domain adaption, called MPDA, to break the dilemmas of purely cloud-based learning and on-device training. From the perspective of a certain user, the general idea of MPDA is to retrieve some similar data from the cloud's global pool, which functions as large-scale source domains, to augment the user's local data as the target domain. The key principle of choosing which outside data depends on whether the model trained over these data can generalize well over the local data. We theoretically analyze that MPDA can reduce distribution discrepancy and overfitting risk. We also extensively evaluate over the public MovieLens 20M and Amazon Electronics datasets, as well as an industrial dataset collected from Mobile Taobao over a period of 30 days. We finally build a device-tunnel-cloud system pipeline, deploy MPDA in the icon area of Mobile Taobao for click-through rate prediction, and conduct online A/B testing. Both offline and online results demonstrate that MPDA outperforms the baselines of cloud-based learning and on-device training only over local data, from multiple offline and online metrics.
Yikai Yan, Chaoyue Niu, Renjie Gu, Fan Wu 0006, Shaojie Tang 0001, Lifeng Hua, Chengfei Lyu, Guihai Chen
KDD3
2022 Walle: An End-to-End, General-Purpose, and Large-Scale Production System for Device-Cloud Collaborative Machine Learning
Chengfei Lv, Chaoyue Niu, Renjie Gu, Xiaotang Jiang, Zhaode Wang, Ziqi Wu, Qiulin Yao, Congyu Huang, Panos Huang, Hui Shu, Jinde Song, Peng Lan, Guohuan Xu, Fei Wu 0001, Shaojie Tang 0001, Fan Wu 0006, Guihai Chen
OSDI3
2019 Unsupervised Fake News Detection on Social Media: A Generative Approach
abstract
Social media has become one of the main channels for people to access and consume news, due to the rapidness and low cost of news dissemination on it. However, such properties of social media also make it a hotbed of fake news dissemination, bringing negative impacts on both individuals and society. Therefore, detecting fake news has become a crucial problem attracting tremendous research effort. Most existing methods of fake news detection are supervised, which require an extensive amount of time and labor to build a reliably annotated dataset. In search of an alternative, in this paper, we investigate if we could detect fake news in an unsupervised manner. We treat truths of news and users’ credibility as latent random variables, and exploit users’ engagements on social media to identify their opinions towards the authenticity of news. We leverage a Bayesian network model to capture the conditional dependencies among the truths of news, the users’ opinions, and the users’ credibility. To solve the inference problem, we propose an efficient collapsed Gibbs sampling approach to infer the truths of news and the users’ credibility without any labelled data. Experiment results on two datasets show that the proposed method significantly outperforms the compared unsupervised methods.
Shuo Yang 0001, Kai Shu, Suhang Wang, Renjie Gu, Fan Wu 0006, Huan Liu 0001
AAAI4