EDBT 2026 Demo / reviewers in the wild / expert
Zheyuan Liu 0010
dblp:191/0249-10
· DBLP profile ↗
16ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0001-7809-4586ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive and Context-rich Generative Self-supervised Learning on GraphsabstractGenerative self-supervised learning on graphs has emerged as a popular learning paradigm and demonstrated its efficacy in handling non-Euclidean data. However, several remaining issues limit the capability of existing methods: 1) the disregard of uneven node significance in masking, 2) the underutilization of holistic graph information, 3) the ignorance of semantic knowledge in the representation space due to the exclusive use of reconstruction loss in the output space, and 4) the unstable reconstructions caused by the large volume of masked contents. In light of this, we propose ACE-GSL, an adaptive and context-rich graph self-supervised learning framework to address these issues from the perspectives of adaptivity, integrity, complementarity, and consistency. Specifically, we first develop an adaptive feature mask generator to account for the unique significance of nodes and sample informative masks (adaptivity). We then design a ranking-based structure reconstruction objective joint with feature reconstruction to capture holistic graph information and emphasize the topological proximity between neighbors (integrity). After that, we present a bootstrapping-based similarity module to encode the high-level semantic knowledge in the representation space, complementary to the low-level reconstruction in the output space (complementarity). Finally, we build a consistency assurance module to provide reconstruction objectives with extra stabilized consistency targets (consistency). Extensive experiments demonstrate that ACE-GSL achieves state-of-the-art performance over 28 methods on 20 datasets across 3 tasks. Yijun Tian 0001, Chuxu Zhang, Ziyi Kou, Zheyuan Liu 0010, Xiangliang Zhang 0001, Nitesh V. Chawla |
AAAI | 4 |
| 2026 | Instant Personalized Large Language Model Adaptation via HypernetworkabstractZhaoxuan Tan, Zixuan Zhang, Haoyang Wen, Zheng Li, Rongzhi Zhang, Pei Chen, Fengran Mo, Zheyuan Liu, Qingkai Zeng, Qingyu Yin, Meng Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhaoxuan Tan, Haoyang Wen, Zheng Li 0018, Rongzhi Zhang, Fengran Mo, Zheyuan Liu 0010, Qingkai Zeng 0001, Qingyu Yin, Meng Jiang 0001 |
ACL (1) | 8 |
| 2026 | Behavior Knowledge Merge in Reinforced Agentic ModelsabstractXiangchi Yuan, Dachuan Shi, Chunhui Zhang, Zheyuan Liu, Shenglong Yao, Soroush Vosoughi, Wenke Lee. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiangchi Yuan, Dachuan Shi, Zheyuan Liu 0010, Shenglong Yao, Soroush Vosoughi, Wenke Lee |
ACL (1) | 4 |
| 2026 | OpenDecoder: Open Large Language Model Decoding to Incorporate Document Quality in RAGabstractThe development of large language models (LLMs) has achieved superior performance in a range of downstream tasks, including LLM-based retrieval-augmented generation (RAG). The quality of generated content heavily relies on the usefulness of the retrieved information and the capacity of LLMs' internal information processing mechanism to incorporate it in answer generation. It is generally assumed that the retrieved information is relevant to the question. However, the retrieved information may have a variable degree of relevance and usefulness, depending on the question and the document collection. It is important to take into account the relevance of the retrieved information in answer generation. In this paper, we propose OpenDecoder, a new approach that leverages explicit evaluation of the retrieved information as quality indicator features for generation. We aim to build a RAG model that is more robust to varying levels of noisy context. Three types of explicit evaluation information are considered: relevance score, ranking score, and QPP (query performance prediction) score. The experimental results on five benchmark datasets demonstrate the effectiveness and better robustness of OpenDecoder by outperforming various baseline methods. Importantly, this paradigm is flexible to be integrated with the post-training of LLMs for any purposes and incorporated with any type of external indicators. Fengran Mo, Zhan Su 0002, Yuchen Hui, Jinghan Zhang 0002, Jia Ao Sun, Zheyuan Liu 0010, Chao Zhang 0014, Tetsuya Sakai, Jian-Yun Nie |
WWW | 6 |
| 2025 | Modality-Aware Neuron Pruning for Unlearning in Multimodal Large Language ModelsabstractZheyuan Liu, Guangyao Dou, Xiangchi Yuan, Chunhui Zhang, Zhaoxuan Tan, Meng Jiang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zheyuan Liu 0010, Guangyao Dou, Xiangchi Yuan, Zhaoxuan Tan, Meng Jiang 0001 |
ACL (1) | 1 |
| 2025 | Disentangling Biased Knowledge from Reasoning in Large Language Models via Machine UnlearningabstractThe rapid development of Large Language Models (LLMs) has led to their widespread adoption across various domains, leveraging vast pre-training knowledge and impressive generalization capabilities. However, these models often inherit biased knowledge, resulting in unfair decisions in sensitive applications. It is challenging to remove this biased knowledge without compromising reasoning abilities due to the entangled nature of the learned knowledge within LLMs. To solve this problem, existing approaches have attempted to mitigate the bias using techniques such as fine-tuning with unbiased datasets, model merging, and gradient ascent. While these methods have experimentally proven effective, they can still be sub-optimum in fully disentangling biases from reasoning. To address this gap, we propose Selective Disentanglement Unlearning (SDU), a novel unlearning framework that selectively removes biased knowledge while preserving reasoning capabilities. SDU operates in three stages: identifying biased parameters using a shadow LLM, fine-tuning with unbiased data, and performing selective parameter updates based on weight saliency. Experimental results across multiple LLMs show that SDU improves fairness accuracy by 14.7% and enhances reasoning performance by 62.6% compared to existing baselines. Zheyuan Liu 0010, Suraj Maharjan, Fanyou Wu, Rahil Parikh, Belhassen Bayar, Srinivasan H. Sengamedu, Meng Jiang 0001 |
ACL (1) | 1 |
| 2025 | Superficial Self-Improved Reasoners Benefit from Model MergingabstractAs scaled language models (LMs) approach human-level reasoning capabilities, selfimprovement emerges as a solution to synthesizing high-quality data corpus.While previous research has identified model collapse as a risk in self-improvement, where model outputs become increasingly deterministic, we discover a more fundamental challenge: the superficial self-improved reasoners phenomenon.In particular, our analysis reveals that even when LMs show improved in-domain (ID) reasoning accuracy, they actually compromise their generalized reasoning capabilities on out-of-domain (OOD) tasks due to memorization rather than genuine learning.Through a systematic investigation of LM architecture, we discover that during self-improvement, LM weight updates are concentrated in less reasoning-critical layers, leading to superficial learning.To address this, we propose Iterative Model Merging (IMM), a method that strategically combines weights from original and self-improved models to preserve generalization while incorporating genuine reasoning improvements.Our approach effectively mitigates both LM collapse and superficial learning, moving towards more stable self-improving systems.Code is available 1 . Xiangchi Yuan, Zheyuan Liu 0010, Dachuan Shi, Leyan Pan, Soroush Vosoughi, Wenke Lee |
EMNLP | 3 |
| 2025 | Incorporating Rather Than Eliminating: Achieving Fairness for Skin Disease Diagnosis Through Group-Specific Experts
Gelei Xu, Yuying Duan, Zheyuan Liu 0010, Meng Jiang 0001, Michael Lemmon 0001, Wei Jin 0009, Yiyu Shi 0001 |
MICCAI (14) | 3 |
| 2025 | Protecting Privacy in Multimodal Large Language Models with MLLMU-BenchabstractZheyuan Liu, Guangyao Dou, Mengzhao Jia, Zhaoxuan Tan, Qingkai Zeng, Yongle Yuan, Meng Jiang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zheyuan Liu 0010, Guangyao Dou, Mengzhao Jia, Zhaoxuan Tan, Qingkai Zeng 0001, Yongle Yuan, Meng Jiang 0001 |
NAACL (Long Papers) | 1 |
| 2024 | Invited: Graph Learning for Parameter Prediction of Quantum Approximate Optimization AlgorithmabstractIn recent years, quantum computing has emerged as a transformative force in the field of combinatorial optimization, offering novel approaches to tackling complex problems that have long challenged classical computational methods. Among these, the Quantum Approximate Optimization Algorithm (QAOA) stands out for its potential to efficiently solve the Max-Cut problem, a quintessential example of combinatorial optimization. However, practical application faces challenges due to current limitations on quantum computational resource. Our work optimizes QAOA initialization, using Graph Neural Networks (GNN) as a warm-start technique. This sacrifices affordable computational resource on classical computer to reduce quantum computational resource overhead, enhancing QAOA's effectiveness. Experiments with various GNN architectures demonstrate the adaptability and stability of our framework, highlighting the synergy between quantum algorithms and machine learning. Our findings show GNN's potential in improving QAOA performance, opening new avenues for hybrid quantum-classical approaches in quantum computing and contributing to practical applications. Zhiding Liang, Gang Liu 0025, Zheyuan Liu 0010, Jinglei Cheng, Tianyi Hao 0003, Zhixin Song, Ji Liu 0007, Fanny Ye, Yiyu Shi 0001 |
DAC | 3 |
| 2024 | Democratizing Large Language Models via Personalized Parameter-Efficient Fine-tuningabstractPersonalization in large language models (LLMs) is increasingly important, aiming to align the LLMs' interactions, content, and recommendations with individual user preferences.Recent advances have highlighted effective prompt design by enriching user queries with non-parametric knowledge through behavior history retrieval and textual profiles.However, these methods faced limitations due to a lack of model ownership, resulting in constrained customization and privacy issues, and often failed to capture complex, dynamic user behavior patterns.To address these shortcomings, we introduce One PEFT Per User (OPPU) 1 , employing personalized parameter-efficient finetuning (PEFT) modules to store user-specific behavior patterns and preferences.By plugging in personal PEFT parameters, users can own and use their LLMs individually.OPPU integrates parametric user knowledge in the personal PEFT parameters with non-parametric knowledge from retrieval and profiles, adapting LLMs to user behavior shifts.Experimental results demonstrate that OPPU significantly outperforms existing prompt-based methods across seven diverse tasks in the LaMP benchmark.Further studies reveal OPPU's enhanced capabilities in handling user behavior shifts, modeling users at different activity levels, maintaining robustness across various user history formats, and displaying versatility with different PEFT methods. Zhaoxuan Tan, Qingkai Zeng 0001, Yijun Tian 0001, Zheyuan Liu 0010, Meng Jiang 0001 |
EMNLP | 4 |
| 2024 | Personalized Pieces: Efficient Personalized Large Language Models through Collaborative EffortsabstractPersonalized large language models (LLMs) aim to tailor interactions, content, and recommendations to individual user preferences.While parameter-efficient fine-tuning (PEFT) methods excel in performance and generalization, they are costly and limit communal benefits when used individually.To this end, we introduce PERSONALIZED PIECES (PER-PCS) 1 , a framework that allows users to safely share and assemble personalized PEFT efficiently with collaborative efforts.PER-PCS involves selecting sharers, breaking their PEFT into pieces, and training gates for each piece.These pieces are added to a pool, from which target users can select and assemble personalized PEFT using their history data.This approach preserves privacy and enables fine-grained user modeling without excessive storage and computation demands.Experimental results show PER-PCS outperforms non-personalized and PEFT retrieval baselines, offering performance comparable to OPPU with significantly lower resource use across six tasks.Further analysis highlights PER-PCS's robustness concerning sharer count and selection strategy, pieces sharing ratio, and scalability in computation time and storage space.PER-PCS's modularity promotes safe sharing, making LLM personalization more efficient, effective, and widely accessible through collaborative efforts. Zhaoxuan Tan, Zheyuan Liu 0010, Meng Jiang 0001 |
EMNLP | 2 |
| 2024 | Breaking the Trilemma of Privacy, Utility, and Efficiency via Controllable Machine UnlearningabstractMachine Unlearning (MU) algorithms have become increasingly critical due to the imperative adherence to data privacy regulations.The primary objective of MU is to erase the influence of specific data samples on a given model without the need to retrain it from scratch.Accordingly, existing methods focus on maximizing user privacy protection.However, there are different degrees of privacy regulations for each real-world web-based application.Exploring the full spectrum of trade-offs between privacy, model utility, and runtime efficiency is critical for practical unlearning scenarios.Furthermore, designing the MU algorithm with simple control of the aforementioned trade-off is desirable but challenging due to the inherent complex interaction.To address the challenges, we present Controllable Machine Unlearning (ConMU), a novel framework designed to facilitate the calibration of MU.The ConMU framework contains three integral modules: an important data selection module that reconciles the runtime efficiency and model generalization, a progressive Gaussian mechanism module that balances privacy and model generalization, and an unlearning proxy that controls the trade-offs between privacy and runtime efficiency.Comprehensive experiments on various benchmark datasets have demonstrated the robust adaptability of our control mechanism and its superiority over established unlearning methods.ConMU explores the full spectrum of the Privacy-Utility-Efficiency trade-off and allows practitioners to account for different real-world regulations. Zheyuan Liu 0010, Guangyao Dou, Eli Chien, Yijun Tian 0001, Ziwei Zhu 0001 |
WWW | 1 |
| 2023 | Chasing All-Round Graph Representation Robustness: Model, Training, and Optimization
Yijun Tian 0001, Mingxuan Ju, Zheyuan Liu 0010, Yanfang Ye 0001, Nitesh V. Chawla, Chuxu Zhang |
ICLR | 4 |
| 2023 | Fair Graph Representation Learning via Diverse Mixture-of-ExpertsabstractGraph Neural Networks (GNNs) have demonstrated a great representation learning capability on graph data and have been utilized in various downstream applications. However, real-world data in web-based applications (e.g., recommendation and advertising) always contains bias, preventing GNNs from learning fair representations. Although many works were proposed to address the fairness issue, they suffer from the significant problem of insufficient learnable knowledge with limited attributes after debiasing. To address this problem, we develop Graph-Fairness Mixture of Experts (G-Fame), a novel plug-and-play method to assist any GNNs to learn distinguishable representations with unbiased attributes. Furthermore, based on G-Fame, we propose G-Fame++, which introduces three novel strategies to improve the representation fairness from node representations, model layer, and parameter redundancy perspectives. In particular, we first present the embedding diversified method to learn distinguishable node representations. Second, we design the layer diversified strategy to maximize the output difference of distinct model layers. Third, we introduce the expert diversified method to minimize expert parameter similarities to learn diverse and complementary representations. Extensive experiments demonstrate the superiority of G-Fame and G-Fame++ in both accuracy and fairness, compared to state-of-the-art methods across multiple graph datasets. Zheyuan Liu 0010, Yijun Tian 0001, Erchi Zhang, Chao Huang 0001, Yanfang Ye 0001, Chuxu Zhang |
WWW | 1 |
| 2022 | GraphBERT: Bridging Graph and Text for Malicious Behavior Detection on Social MediaabstractThe development of social media (e.g., Twitter) allows users to make speeches with low cost and broad influence. Thus, social media has become a perfect place for users’ malicious behaviors like committing hate crimes, spreading toxic information, abetting crimes, etc. Malicious behaviors are covert and widespread, with potential relevance regarding topic, person, place, and so on. Therefore, it is necessary to develop novel techniques to detect and disrupt malicious behavior on social media effectively. Previous research has shown promising results in extracting semantic text (speech) representation using natural language processing methods. Yet the latent relation between speeches and the connection between users behind speeches is rarely explored. In light of this, we propose a holistic model named Graph adaption BERT (GraphBERT) to detect malicious behaviors on Twitter with both semantic and relational information. Specifically, we first present a novel and a large-scale corpus of tweet data to benefit both graph-based and language-based malicious behavior detection research. Then, we design a novel model GraphBERT to learn comprehensive tweet and user representation with the integration of both semantic information encoded by transformers (i.e., BERT) and relational information encoded by graph neural network. GraphBERT further leverages a weight adaption BERT module implemented between transformer layers to refine tweet embedding using relational information for malicious tweet classification. Finally, the adapted tweet embedding is used with the initial tweet representation to generate user embedding for malicious user detection. The extensive experiments on the collected Twitter data show that our model outperforms the state-of-the-art baseline methods for both tasks (i.e., malicious tweet classification and malicious user detection). Jiele Wu, Zheyuan Liu 0010, Erchi Zhang, Steven Lloyd Wilson, Chuxu Zhang |
ICDM | 3 |