Rongzhe Wei

dblp:259/6894 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
9since 2021 · last 2025
0009-0008-9047-1303ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3
YearPublicationVenuePosition
2025 Generalization Principles for Inference over Text-Attributed Graphs with Large Language Models
abstract
Large language models (LLMs) have recently been introduced to graph learning, aiming to extend their zero-shot generalization success to tasks where labeled graph data is scarce. Among these applications, inference over text-attributed graphs (TAGs) presents unique challenges: existing methods struggle with LLMs' limited context length for processing large node neighborhoods and the misalignment between node embeddings and the LLM token space. To address these issues, we establish two key principles for ensuring generalization and derive the framework LLM-BP accordingly: (1) **Unifying the attribute space with task-adaptive embeddings**, where we leverage LLM-based encoders and task-aware prompting to enhance generalization of the text attribute embeddings; (2) **Developing a generalizable graph information aggregation mechanism**, for which we adopt belief propagation with LLM-estimated parameters that adapt across graphs. Evaluations on 11 real-world TAG benchmarks demonstrate that LLM-BP significantly outperforms existing approaches, achieving 8.10\% improvement with task-conditional embeddings and an additional 1.71\% gain from adaptive aggregation. The code and task-adaptive embeddings are publicly available.
Haoyu Peter Wang, Shikun Liu, Rongzhe Wei, Pan Li 0005
ICML3
2025 Underestimated Privacy Risks for Minority Populations in Large Language Model Unlearning
abstract
Large Language Models (LLMs) embed sensitive, human-generated data, prompting the need for unlearning methods. Although certified unlearning offers strong privacy guarantees, its restrictive assumptions make it unsuitable for LLMs, giving rise to various heuristic approaches typically assessed through empirical evaluations. These standard evaluations randomly select data for removal, apply unlearning techniques, and use membership inference attacks (MIAs) to compare unlearned models against models retrained without the removed data. However, to ensure robust privacy protections for every data point, it is essential to account for scenarios in which certain data subsets face elevated risks. Prior research suggests that outliers, particularly including data tied to minority groups, often exhibit higher memorization propensity which indicates they may be more difficult to unlearn. Building on these insights, we introduce a complementary, minority-aware evaluation framework to highlight blind spots in existing frameworks. We substantiate our findings with carefully designed experiments, using canaries with personally identifiable information (PII) to represent these minority subsets and demonstrate that they suffer at least 20\% higher privacy leakage across various unlearning methods, MIAs, datasets, and LLM scales. Our proposed minority-aware evaluation framework marks an essential step toward more equitable and comprehensive assessments of LLM unlearning efficacy.
Rongzhe Wei, Mufei Li, Mohsen Ghassemi, Eleonora Kreacic, Xiang Yue, Bo Li 0026, Vamsi K. Potluru, Pan Li 0005, Eli Chien
ICML1
2025 Differentially Private Relational Learning with Entity-level Privacy Guarantees
abstract
Learning with relational and network-structured data is increasingly vital in sensitive domains where protecting the privacy of individual entities is paramount. Differential Privacy (DP) offers a principled approach for quantifying privacy risks, with DP-SGD emerging as a standard mechanism for private model training. However, directly applying DP-SGD to relational learning is challenging due to two key factors: (i) entities often participate in multiple relations, resulting in high and difficult-to-control sensitivity; and (ii) relational learning typically involves multi-stage, potentially coupled (interdependent) sampling procedures that make standard privacy amplification analyses inapplicable. This work presents a principled framework for relational learning with formal entity-level DP guarantees. We provide a rigorous sensitivity analysis and introduce an adaptive gradient clipping scheme that modulates clipping thresholds based on entity occurrence frequency. We also extend the privacy amplification results to a tractable subclass of coupled sampling, where the dependence arises only through sample sizes. These contributions lead to a tailored DP-SGD variant for relational data with provable privacy guarantees. Experiments on fine-tuning text encoders over text-attributed network-structured relational data demonstrate the strong utility-privacy trade-offs of our approach.
Yinan Huang, Haoteng Yin, Eli Chien, Rongzhe Wei, Pan Li 0005
NeurIPS4
2025 Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
abstract
Machine unlearning techniques aim to mitigate unintended memorization in large language models (LLMs). However, existing approaches predominantly focus on the explicit removal of isolated facts, often overlooking latent inferential dependencies and the non-deterministic nature of knowledge within LLMs. Consequently, facts presumed forgotten may persist implicitly through correlated information. To address these challenges, we propose a knowledge unlearning evaluation framework that more accurately captures the implicit structure of real-world knowledge by representing relevant factual contexts as knowledge graphs with associated confidence scores. We further develop an inference-based evaluation protocol leveraging powerful LLMs as judges; these judges reason over the extracted knowledge subgraph to determine unlearning success. Our LLM judges utilize carefully designed prompts and are calibrated against human evaluations to ensure their trustworthiness and stability. Extensive experiments on our newly constructed benchmark demonstrate that our framework provides a more realistic and rigorous assessment of unlearning performance. Moreover, our findings reveal that current evaluation strategies tend to overestimate unlearning effectiveness.
Rongzhe Wei, Peizhi Niu, Hans Hao-Hsun Hsu, Ruihan Wu, Haoteng Yin, Mohsen Ghassemi, Vamsi K. Potluru, Eli Chien, Kamalika Chaudhuri, Olgica Milenkovic, Pan Li 0005
NeurIPS1
2024 Differentially Private Graph Diffusion with Applications in Personalized PageRanks
abstract
Graph diffusion, which iteratively propagates real-valued substances among the graph, is used in numerous graph/network-involved applications. However, releasing diffusion vectors may reveal sensitive linking information in the data such as transaction information in financial network data. However, protecting the privacy of graph data is challenging due to its interconnected nature. This work proposes a novel graph diffusion framework with edge-level different privacy guarantees by using noisy diffusion iterates. The algorithm injects Laplace noise per diffusion iteration and adopts a degree-based thresholding function to mitigate the high sensitivity induced by low-degree nodes. Our privacy loss analysis is based on Privacy Amplification by Iteration (PABI), which to our best knowledge, is the first effort that analyzes PABI with Laplace noise and provides relevant applications. We also introduce a novel $\infty$-Wasserstein distance tracking method, which tightens the analysis of privacy leakage and makes PABI more applicable in practice. We evaluate this framework by applying it to Personalized Pagerank computation for ranking tasks. Experiments on real-world network data demonstrate the superiority of our method under stringent privacy conditions.
Rongzhe Wei, Eli Chien, Pan Li 0005
NeurIPS1
2024 Learning Scalable Structural Representations for Link Prediction with Bloom Signatures
abstract
Graph neural networks (GNNs) have shown great potential in learning on graphs, but they are known to perform sub-optimally on link prediction tasks. Existing GNNs are primarily designed to learn node-wise representations and usually fail to capture pairwise relations between target nodes, which proves to be crucial for link prediction. Recent works resort to learning more expressive edge-wise representations by enhancing vanilla GNNs with structural features such as labeling tricks and link prediction heuristics, but they suffer from high computational overhead and limited scalability. To tackle this issue, we propose to learn structural link representations by augmenting the message-passing framework of GNNs with Bloom signatures. Bloom signatures are hashing-based compact encodings of node neighborhoods, which can be efficiently merged to recover various types of edge-wise structural features. We further show that any type of neighborhood overlap-based heuristic can be estimated by a neural network that takes Bloom signatures as input. GNNs with Bloom signatures are provably more expressive than vanilla GNNs and also more scalable than existing edge-wise models. Experimental results on five standard link prediction benchmarks show that our proposed model achieves comparable or better performance than existing edge-wise GNN models while being 3-200x faster and more memory-efficient for online inference. Source code is available at https://github.com/tonyzhang617/BloomSigLP.
Tianyi Zhang 0011, Haoteng Yin, Rongzhe Wei, Pan Li 0005, Anshumali Shrivastava
WWW3
2024 SLA$^{{\text{2}}}$2P: Self-Supervised Anomaly Detection With Adversarial Perturbation
abstract
Anomaly detection is a foundational yet difficult problem in machine learning. In this work, we propose a new and effective framework, dubbed as SLA2P, for unsupervised anomaly detection. Following the extraction of delegate embeddings from raw data, we implement random projections on the features and consider features transformed by disparate projections as being associated with separate pseudo-classes. We then train a neural network for classification on these transformed features to conduct self-supervised learning. Subsequently, we introduce adversarial disturbances to the modified attributes, and we develop anomaly scores built on the classifier's predictive uncertainties concerning these disrupted features. Our approach is motivated by the fact that as anomalies are relatively rare and decentralized, 1) the training of the pseudo-label classifier concentrates more on acquiring the semantic knowledge of regular data instead of anomalous data; 2) the altered attributes of the normal data exhibit greater resilience to disturbances compared to those of the anomalous data. Therefore, the disrupted modified attributes of anomalies can not be well classified and correspondingly tend to attain lesser anomaly scores. The results of experiments on various benchmark datasets for images, text, and inherently tabular data demonstrate that SLA2P achieves state-of-the-art performance consistently.
Yizhou Wang 0006, Can Qin, Rongzhe Wei, Yi Xu 0005, Yun Fu 0001
IEEE Trans. Knowl. Data Eng.3
2022 Self-supervision Meets Adversarial Perturbation: A Novel Framework for Anomaly Detection
abstract
Anomaly detection is a fundamental yet challenging problem in machine learning due to the lack of label information. In this work, we propose a novel and powerful framework, dubbed as SLA2P, for unsupervised anomaly detection. After extracting representative embeddings from raw data, we apply random projections to the features and regard features transformed by different projections as belonging to distinct pseudo-classes. We then train a classifier network on these transformed features to perform self-supervised learning. Next, we add adversarial perturbation to the transformed features to decrease their softmax scores of the predicted labels and design anomaly scores based on the predictive uncertainties of the classifier on these perturbed features. Our motivation is that because of the relatively small number and the decentralized modes of anomalies, 1) the pseudo label classifier's training concentrates more on learning the semantic information of normal data rather than anomalous data; 2) the transformed features of the normal data are more robust to the perturbations than those of the anomalies. Consequently, the perturbed transformed features of anomalies fail to be classified well and accordingly have lower anomaly scores than those of the normal samples. Extensive experiments on image, text, and inherently tabular benchmark datasets back up our findings and indicate that SLA2 achieves state-of-the-art anomaly detection performance consistently. Our code is made publicly available at https://github.com/wyzjack/SLA2P
Yizhou Wang 0006, Can Qin, Rongzhe Wei, Yi Xu 0005, Yun Fu 0001
CIKM3
2022 Understanding Non-linearity in Graph Neural Networks from the Bayesian-Inference Perspective
abstract
Graph neural networks (GNNs) have shown superiority in many prediction tasks over graphs due to their impressive capability of capturing nonlinear relations in graph-structured data. However, for node classification tasks, often, only marginal improvement of GNNs has been observed in practice over their linear counterparts. Previous works provide very few understandings of this phenomenon. In this work, we resort to Bayesian learning to give an in-depth investigation of the functions of non-linearity in GNNs for node classification tasks. Given a graph generated from the statistical model CSBM, we observe that the max-a-posterior estimation of a node label given its own and neighbors' attributes consists of two types of non-linearity, the transformation of node attributes and a ReLU-activated feature aggregation from neighbors. The latter surprisingly matches the type of non-linearity used in many GNN models. By further imposing Gaussian assumption on node attributes, we prove that the superiority of those ReLU activations is only significant when the node attributes are far more informative than the graph structure, which nicely explains previous empirical observations. A similar argument is derived when there is a distribution shift of node attributes between the training and testing datasets. Finally, we verify our theory on both synthetic and real-world networks. Our code is available at https://github.com/Graph-COM/Bayesian_inference_based_GNN.git.
Rongzhe Wei, Haoteng Yin, Junteng Jia, Austin R. Benson, Pan Li 0005
NeurIPS1
2020 Dual Adversarial Networks for Land-Cover Classification
abstract
River basin scene classification as an important application in the field of land-cover recognition has been arousing extensive concern. Traditional land-cover classification methods with multi-feature extractions on specific scene perform well on a single river basin, however, poorly address inter-basin classification owing to the varied texture shown in satellite images cross river basins (e.g., topography and climates). Current transfer learning approaches with domain adaptation, which can shorten the discrepancies between two river basins, pay less attention to diversity of multi-feature extractions given by remote sensing images, which may lead to negative transfer. To better address the above challenges, this paper proposes a model known as Dual Adversarial Networks for Land-cover Classification (DANLC). Our DANLC architecture consists of two domain adversarial networks in a paralleled structure, namely RGB and texture networks, for multi-feature extractions, which are able to capture the underlying representation of satellite images from different perspectives and get invariable transfer component. Results demonstrate the outstanding performance of our model in both the classification effect and robustness compared with traditional methods and state-of-the-art transfer learning approaches.
Jingyi An, Rongzhe Wei, Zhichao Zha
COMPSAC2
2020 NEUD-TRI: Network Embedding Based on Upstream and Downstream for Transaction Risk Identification
abstract
Invoices serve as records of financial transactions of taxpayers and significant basis to controlling tax source and collection of tax, via analyzing which, we can discern diversified tasks of tax risk, such as industry identification, hidden transaction detection, and illegal behavior mining. Among all the existing studies related to the identification of tax risk, there are some weaknesses through the machine learning model and network analysis because of the dependence on tax knowledge. Different from the manual selection of indicators and the manual definitions mode with the guidance of tax knowledge in the past, in this paper, we propose a novel method, namely, network embedding based on upstream and downstream for tax risk identification (NEUD-TRI), which considers the taxpayers serving as both seller and purchaser. The method designs optimization functions respectively to capture local and global static network structures and dynamic network structure. In view of the significant discrepancy of weights in the transaction network, this paper normalizes the weight within the range of the upstream and downstream of the vertex. Negative sampling and edge sampling are adopted to deal with the large-scale trait of the transaction network. Empirical results on tax data-sets of Shanxi province substantiate the effectiveness of our models.
Jingyi An, Rongzhe Wei, Bo Dong 0001, Xuanya Li
COMPSAC3
2020 A Novel Tax Evasion Detection Framework via Fused Transaction Network Representation
abstract
Tax evasion usually refers to the false declaration of taxpayers to reduce their tax obligations; this type of behavior leads to the loss of taxes and damage to the fair principle of taxation. Tax evasion detection plays a crucial role in reducing tax revenue loss. Currently, efficient auditing methods mainly include traditional data-mining-oriented methods, which cannot be well adapted to the increasingly complicated transaction relationships between taxpayers. Driven by this requirement, recent studies have been conducted by establishing a transaction network and applying the graphical pattern matching algorithm for tax evasion identification. However, such methods rely on expert experience to extract the tax evasion chart pattern, which is time-consuming and labor-intensive. More importantly, taxpayers' basic attributes are not considered and the dual identity of the taxpayer in the transaction network is not well retained. To address this issue, we have proposed a novel tax evasion detection framework via fused transaction network representation (TED-TNR), to detecting tax evasion based on fused transaction network representation, which jointly embeds transaction network topological information and basic taxpayer attributes into low-dimensional vector space, and considers the dual identity of the taxpayer in the transaction network. Finally, we conducted experimental tests on real-world tax data, revealing the superiority of our method, compared with state-of-the-art models.
Yingchao Wu, Bo Dong 0001, Rongzhe Wei, Xuanya Li
COMPSAC4
2019 Unsupervised Conditional Adversarial Networks for Tax Evasion Detection
abstract
The identification of tax evasion plays an important role in ensuring tax order, promoting the level of tax collection and management, and reducing tax losses. With the advancements in data mining technology, many machine learning techniques have yielded results in identifying tax evasion. However, to realize satisfactory performance, these models require large amounts of human annotated data. In the tax field, unlabeled tax data are abundant, data annotation in a single region is expensive, and the distributions of characteristics differ among regions; these factors pose substantial difficulties in the development of an identification model. Existing tax evasion detection methods are either trained for single-region tasks, in which case they perform poorly on inter-region tax evasion identification due to the discrepancies in feature distributions, or utilize labeled data from both the target-task field and different but related auxiliary fields to reuse and transfer knowledge of the target domain data, in which case they cannot deal with scenarios in which there are no labeled data in target audit tasks. Although current unsupervised transfer learning techniques can train models in labeled regions for unlabeled regions, large intra-class distribution discrepancies cannot be perfectly minimized in tax evasion detection scenarios. To better address the above challenges, this paper proposes a general architecture, namely, the unsupervised conditional adversarial networks (UCAN) for tax evasion detection, which is the first approach to solve audit tasks in unlabeled target domains via inter-region transfer. Our architecture establishes an adversarial neural network adding label information in the distribution adapter, which can granularly adapt the joint probability distribution (JPD) of the data. We introduce a constraint that is based on the conditional maximum mean discrepancy (CMMD) of the extracted features to align the conditional probability distribution (CPD) of the deep representation. Our model is formed by combining the distribution adapter and the label predictor to realize end-to-end learning of unsupervised feature transfer. The experimental results demonstrate the outstanding performance of our model in all migration tasks compared with state-of-the-art approaches.
Rongzhe Wei, Bo Dong 0001, Xulyu Zhu, Jianfei Ruan
IEEE BigData1
2019 ABR-HIC: Attention Based Bidirectional RNN for Hierarchical Industry Classification
abstract
Accurate industry classification of national economic activities as an important component in the construction of economic structure and as the basis of the formulation of economic policies and management of national economic activities has been gaining increasing attention. However, owing to the rapid growth in the number of industries, it is become increasingly difficult for tax bureaus to classify the registered taxpayers' industries. Conventional industrial classification methods only focus on the text features, which can not be analyzed and judged comprehensively according to the registration information, and can only carry on single-label classification since they neglect the primary and secondary relationships between the main and subsidiary industries, which can not meet application requirements. To better address these challenges, this paper proposes a model known as attention based bidirectional RNN for hierarchical industry classification (ABR-HIC), which is the first approach, to the best of our knowledge, to simultaneously address comprehensive registration information utilization and multi-label classification for the main and subsidiary industries. Our architecture establishes a bidirectional RNN using a word-attention mechanism, which is able to capture and fully utilize the text and non-text registration information for feature representation. By separating the taxpayer's primary and secondary multi-label classification problem corresponding to the main and subsidiary industries, respectively, into two subtasks and through multi-task learning, our model can provide comprehensive primary and secondary multi-industrial labels. Experiments were conducted on real tax data-sets of the Shaanxi Province, China and the results demonstrate the outstanding performance of our architecture in terms of both the classification effect and training time compared with those of state-of-the-art approaches.
Rongzhe Wei, Bo Dong 0001, Kuanzheng Yang, Jianfei Ruan
IEEE BigData1
2019 TEDM-PU: A Tax Evasion Detection Method Based on Positive and Unlabeled Learning
abstract
Tax evasion detection plays a crucial role in reducing tax revenue loss and many efforts have been made to develop detection models based on machine learning techniques. To train an effective model to detect tax evaders, a large amount of data is required, especially sufficient labeled data. However, the expensive and time-consuming annotation process results in small amount of labeled data being available, which makes the development of detection models difficult. To address this issue, we propose a tax evasion detection method based on positive and unlabeled learning (TEDM-PU), to identify tax evasion by utilizing limited annotated tax evasion taxpayers and a large amount of unlabeled data. The TEDM-PU framework consists of three stages: a preprocessing stage extracting taxpayer features based on random forest, a pseudo labeling stage assigning pseudo labels to unlabeled samples based on PUAdapter, and a model training stage based on LightGBM method. To evaluate the effectiveness of our proposed TEDM-PU, we conduct experimental tests on real-world tax data. The results demonstrate that TEDM-PU method can detect tax evaders with higher accuracy and better interpretability than state-of-the-art methods.
Yingchao Wu, Yuda Gao, Bo Dong 0001, Rongzhe Wei, Fa Zhang 0002
IEEE BigData5