VLDB 2026 Research / reviewers in the wild / expert
Zelei Cheng
dblp:258/0335
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0001-7478-933XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Security and privacy · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GPO: Learning from Critical Steps to Improve LLM ReasoningabstractLarge language models (LLMs) are increasingly used in various domains, showing impressive potential on various tasks.
Recently, reasoning LLMs have been proposed to improve the \textit{reasoning} or \textit{thinking} capabilities of LLMs to solve complex problems.
Despite the promising results of reasoning LLMs, enhancing the multi-step reasoning capabilities of LLMs still remains a significant challenge.
While existing optimization methods have advanced the LLM reasoning capabilities, they often treat reasoning trajectories as a whole, without considering the underlying critical steps within the trajectory. In this paper, we introduce \textbf{G}uided \textbf{P}ivotal \textbf{O}ptimization (GPO), a novel fine-tuning strategy that dives into the reasoning process to enable more effective improvements.
GPO first identifies the `critical step' within a reasoning trajectory - a point that the model must carefully proceed so as to succeed at the problem. We locate the critical step by estimating the advantage function.
GPO then resets the policy to the critical step and samples the new rollout and prioritizes learning process on those rollouts.
This focus allows the model to learn more effectively from pivotal moments within the reasoning process to improve the reasoning performance.
We demonstrate that GPO is not a standalone method, but rather a general strategy that can be integrated with various optimization methods to improve reasoning performance.
Besides theoretical analysis, our experiments across challenging reasoning benchmarks show that GPO can consistently and significantly enhances the performance of existing optimization methods, showcasing its effectiveness and generalizability in improving LLM reasoning by concentrating on pivotal moments within the generation process. Jiahao Yu 0001, Zelei Cheng, Xian Wu 0007, Xinyu Xing 0001 |
NeurIPS | 2 |
| 2024 | RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with ExplanationabstractDeep reinforcement learning (DRL) is playing an increasingly important role in real-world applications. However, obtaining an optimally performing DRL agent for complex tasks, especially with sparse rewards, remains a significant challenge. The training of a DRL agent can be often trapped in a bottleneck without further progress. In this paper, we propose RICE, an innovative refining scheme for reinforcement learning that incorporates explanation methods to break through the training bottlenecks. The high-level idea of RICE is to construct a new initial state distribution that combines both the default initial states and critical states identified through explanation methods, thereby encouraging the agent to explore from the mixed initial states. Through careful design, we can theoretically guarantee that our refining scheme has a tighter sub-optimality bound. We evaluate RICE in various popular RL environments and real-world applications. The results demonstrate that RICE significantly outperforms existing refining schemes in enhancing agent performance. Zelei Cheng, Xian Wu 0007, Jiahao Yu 0001, Sabrina Yang, Gang Wang 0011, Xinyu Xing 0001 |
ICML | 1 |
| 2024 | Soft-Label Integration for Robust Toxicity ClassificationabstractToxicity classification in textual content remains a significant problem. Data with labels from a single annotator fall short of capturing the diversity of human perspectives. Therefore, there is a growing need to incorporate crowdsourced annotations for training an effective toxicity classifier. Additionally, the standard approach to training a classifier using empirical risk minimization (ERM) may fail to address the potential shifts between the training set and testing set due to exploiting spurious correlations. This work introduces a novel bi-level optimization framework that integrates crowdsourced annotations with the soft-labeling technique and optimizes the soft-label weights by Group Distributionally Robust Optimization (GroupDRO) to enhance the robustness against out-of-distribution (OOD) risk. We theoretically prove the convergence of our bi-level optimization algorithm. Experimental results demonstrate that our approach outperforms existing baseline methods in terms of both average and worst-group accuracy, confirming its effectiveness in leveraging crowdsourced annotations to achieve more effective and robust toxicity classification. Zelei Cheng, Xian Wu 0007, Jiahao Yu 0001, Xin-Qiang Cai, Xinyu Xing 0001 |
NeurIPS | 1 |
| 2023 | StateMask: Explaining Deep Reinforcement Learning through State MaskabstractDespite the promising performance of deep reinforcement learning (DRL) agents in many challenging scenarios, the black-box nature of these agents greatly limits their applications in critical domains. Prior research has proposed several explanation techniques to understand the deep learning-based policies in RL. Most existing methods explain why an agent takes individual actions rather than pinpointing the critical steps to its final reward. To fill this gap, we propose StateMask, a novel method to identify the states most critical to the agent's final reward. The high-level idea of StateMask is to learn a mask net that blinds a target agent and forces it to take random actions at some steps without compromising the agent's performance. Through careful design, we can theoretically ensure that the masked agent performs similarly to the original agent. We evaluate StateMask in various popular RL environments and show its superiority over existing explainers in explanation fidelity. We also show that StateMask has better utilities, such as launching adversarial attacks and patching policy errors. Zelei Cheng, Xian Wu 0007, Jiahao Yu 0001, Wenhai Sun, Wenbo Guo 0002, Xinyu Xing 0001 |
NeurIPS | 1 |
| 2023 | Protecting Regression Models With Personalized Local Differential PrivacyabstractThe equation-solving model extraction attack is an intuitively simple but devastating attack to steal confidential information of regression models through a sufficient number of queries. Complete mitigation is difficult. Thus, the development of countermeasures is focused on degrading the attack effectiveness as much as possible without losing the model utilities. We investigate a novel personalized local differential privacy mechanism to defend against the attack. We obfuscate the model by adding high-dimensional Gaussian noise on model coefficients. Our solution can adaptively produce the noise to protect the model on the fly. We thoroughly evaluate the performance of our mechanisms using real-world datasets. The experiment shows that the proposed scheme outperforms the existing differential-privacy-enabled solution, i.e., 4 times more queries are required to achieve the same attack result. We also plan to publish the relevant codes to the community for further research. Haonan Yan, Zelei Cheng, Wenhai Sun, Hui Li 0006 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2022 | A face recognition algorithm based on feature fusionabstractSummary In the process of building a smart city, face recognition can be applied to the transformation of enterprises, communities, and parks. The combination of building security system and face recognition technology can improve the security experience of enterprises and citizens through the solution of hardware and software integration. Face recognition is still facing the challenges of illumination, occlusion, and attitude change in the actual application process. In addition, the end‐to‐end convolutional neural networks (CNN) seldom make use of the hierarchical feature of the network. So, we propose a hierarchy feature fusion method for face recognition, which uses supervisory information to learn shallow and deep facial features. The features are fused to enhance the recognition accuracy of face recognition against illumination and occlusion. The method is applied to the transformation of the visual geometry group network and Lightened CNN. The face recognition experiments are carried out using the hierarchy network. Our method has achieved good recognition results in the labeled faces in the wild (LFW) and AR face databases. Jiwei Zhang 0007, Xiaodan Yan, Zelei Cheng, Xueqi Shen |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Poisoning Attack for Inter-agent Transfer Learning
Zelei Cheng, Zuotian Li |
SecureComm (2) | 1 |
| 2020 | Hypergraph Attention NetworksabstractRecently, graph neural networks have achieved great success on the representation learning of the graph-structured data. However, these networks just consider the pairwise connection between nodes which cannot model the complicated connections of data in the real world. Thus, researchers began to pay attention to the hypergraph modeling. In recent years, some hypergraph neural networks have been proposed to aggregate the information of the hypergraph for representation learning. In this paper, we present hypergraph attention networks (HGATs) to encode the high-order data relation in the hypergraph. Specifically, our proposed HGATs consist of two modules: attentive vertex aggregation module and attentive hyperedge aggregation module. These two modules can implicitly assign different aggregation weights to different connected hyperedge/vertex to characterize the complex relations among data. We stack these modules to pass the messages between the hyperedges and vertices to refine the vertex/hyperedge features. Experimental results on the ModelNet40 and NTU2012 datasets show that our proposed HGATs can achieve superior performance for the visual object recognition tasks. Furthermore, we employ our HGAT for multi-view representation learning and better object classification results are achieved. Zelei Cheng, Zuotian Li, Manyi Wang |
TrustCom | 2 |
| 2020 | Deep Learning for Password Guessing and Password Strength Evaluation, A SurveyabstractText passwords are the most widely used authentication methods and will also be used in the future. Text passwords can be regarded as meaningful strings, and deep learning methods have an advantage of text processing. LSTM, RNN, GAN and other deep learning models have been using in password guessing and password strength measurements. In the paper, we make a survey on state-of-the-art deep learning methods for password guessing and password strength evaluation, including password pattern extraction, candidate password generation and password strength measurement. Compared with traditional methods, neural networks based methods can achieve better results and performance. Zelei Cheng |
TrustCom | 2 |