EDBT 2026 Demo / reviewers in the wild / expert
Yi Liu 0071
dblp:97/4626-71
· DBLP profile ↗
15ranked-venue papers
5as first author
14since 2021 · last 2026
0009-0001-0354-1940ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse AutoencodersabstractSparse Autoencoder (SAE) has emerged as a powerful tool for mechanistic interpretability of large language models. Recent works apply SAE to protein language models (PLMs), aiming to extract and analyze biologically meaningful features from their latent spaces. However, SAE suffers from semantic entanglement, where individual neurons often mix multiple nonlinear concepts, making it difficult to reliably interpret or manipulate model behaviors. In this paper, we propose a semantically-guided SAE, called ProtSAE. Unlike existing SAE which requires annotation datasets to filter and interpret activations, we guide semantic disentanglement during training using both annotation datasets and domain knowledge to mitigate the effects of entangled attributes. We design interpretability experiments showing that ProtSAE learns more biologically relevant and interpretable hidden features compared to previous methods. Performance analyses further demonstrate that ProtSAE maintains high reconstruction fidelity while achieving better results in interpretable probing. We also show the potential of ProtSAE in steering PLMs for downstream generation tasks. Xiangyu Liu 0001, Haodi Lei, Yi Liu 0071, Wei Hu 0007 |
AAAI | 3 |
| 2026 | Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning ModelsabstractLarge reasoning models (LRMs) have shown remarkable progress on complex reasoning tasks. However, some questions posed to LRMs are inherently unanswerable, such as math problems lacking sufficient conditions. We find that LRMs continually fail to provide appropriate abstentions when confronted with these unanswerable questions. In this paper, we systematically analyze, investigate, and resolve this issue for trustworthy AI. We first conduct a detailed analysis of the distinct response behaviors of LRMs when facing unanswerable questions. Then, we show that LRMs possess sufficient cognitive capabilities to recognize the flaws in these questions. However, they fail to exhibit appropriate abstention behavior, revealing a misalignment between their internal cognition and external response. Finally, to resolve this issue, we propose a lightweight, two-stage method that combines cognitive monitoring with inference-time intervention. Experimental results demonstrate that our method significantly improves the abstention rate while maintaining the reasoning performance. Yi Liu 0071, Xiangyu Liu 0001, Zequn Sun 0001, Wei Hu 0007 |
AAAI | 1 |
| 2025 | Controllable Protein Sequence Generation with LLM Preference OptimizationabstractDesigning proteins with specific attributes offers an important solution to address biomedical challenges. Pre-trained protein large language models (LLMs) have shown promising results on protein sequence generation. However, to control sequence generation for specific attributes, existing work still exhibits poor functionality and structural stability. In this paper, we propose a novel controllable protein design method called CtrlProt. We finetune a protein LLM with a new multi-listwise preference optimization strategy to improve generation quality and support multi-attribute controllable generation. Experiments demonstrate that CtrlProt can meet functionality and structural stability requirements effectively, achieving state-of-the-art performance in both single-attribute and multi-attribute protein sequence generation. Xiangyu Liu 0001, Yi Liu 0071, Silei Chen, Wei Hu 0007 |
AAAI | 2 |
| 2025 | Unveiling the Impact of Multi-modal Content in Multi-modal Recommender SystemsabstractMulti-modal recommender systems (MRSs) have emerged as critical multi-modal technologies, but Do we truly leverage multi-modal content properly? Through an empirical study of four diverse, realworld datasets spanning various recommendation scenarios, we observe that MRSs exhibit a stronger tendency to recommend items with high similarity to users' past interactions in terms of multimodal content than conventional RSs. While this tendency improves the recommendation accuracy, it introduces a previously unexplored bias that significantly impacts user experience. We define this bias as User-side Content Bias: users who prefer items similar to their historical choices receive higher quality recommendations than those seeking diverse options. We show that User-side Content Bias is unrelated to the activity of users, indicating a fundamental limitation in current MRSs. We propose ISOLATOR: utIlizing uSer-side cOntent simiLarity via a model-AgnosTic framewORk to leverage multi-modal content more properly. ISOLATOR estimates the impact of User-side Content Similarity and proposes two intervention strategies to meet the needs for more accurate and unbiased recommendations. Extensive evaluations on several widely used datasets demonstrate that ISOLATOR consistently improves various state-of-the-art MRSs and effectively addresses the User-side Content Bias. Guipeng Xv, Yi Liu 0071, Chen Lin 0001, Xiaoli Wang 0002 |
ACM Multimedia | 3 |
| 2025 | Knowledge Graph-Guided Retrieval Augmented GenerationabstractXiangrong Zhu, Yuexiang Xie, Yi Liu, Yaliang Li, Wei Hu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xiangrong Zhu 0001, Yuexiang Xie, Yi Liu 0071, Yaliang Li, Wei Hu 0007 |
NAACL (Long Papers) | 3 |
| 2025 | Knowledge Enhancement and Temporal Aware for Multi-Behavior Contrastive RecommendationabstractA well-designed recommender system can accurately learn the embeddings of users and items, reflecting the unique preferences of users. Traditional recommendation techniques usually focus on modeling the singular type of behaviors between users and items. However, in many practical recommendation scenarios (e.g., social media, e-commerce), there exist multi-typed interactive behaviors in user–item relationships, such as click, tag-as-favorite, and purchase in online shopping platforms. Thus, how to make full use of multi-behavior information for recommendation is of great importance to the existing system, which presents challenges in two aspects that need to be explored: (1) Utilizing users’ personalized preferences to capture multi-behavioral dependencies; (2) Dealing with the insufficient recommendation caused by sparse supervision signal for target behavior. In this work, we propose the Knowledge Enhancement Multi-Behavior Contrastive Learning (KMCL) framework , including two Contrastive Learning tasks and three functional modules to tackle the above challenges, respectively. In particular, we design the multi-behavior learning module to extract users’ personalized behavior information for user-embedding enhancement and utilize knowledge graph in the knowledge enhancement module to derive more robust knowledge-aware representations for items. In addition, in the optimization stage, we also model the coarse-grained commonalities and the fine-grained differences between multi-behavior of users to further improve the recommendation effect and propose a joint training paradigm to enhance the learning effect of KMCLR in the joint learning module. Besides, we also considered how to make full use of temporal signals to enhance the effectiveness of multi-behavior recommendations in scenarios with time information and designed a novel encoder to address this issue. Extensive experiments and ablation tests on the three real-world datasets indicate that our KMCLR outperforms various state-of-the-art recommendation methods and verify the effectiveness of our method. Hongrui Xuan, Bohan Li 0001, Yi Liu 0071, Hongzhi Yin |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2024 | Multi-Aspect Controllable Text Generation with Disentangled Counterfactual AugmentationabstractMulti-aspect controllable text generation aims to control the generated texts in attributes from multiple aspects (e.g., "positive" from sentiment and "sport" from topic).For ease of obtaining training samples, existing works neglect attribute correlations formed by the intertwining of different attributes.Particularly, the stereotype formed by imbalanced attribute correlations significantly affects multi-aspect control.In this paper, we propose MAGIC, a new multi-aspect controllable text generation method with disentangled counterfactual augmentation.We alleviate the issue of imbalanced attribute correlations during training using counterfactual feature vectors in the attribute latent space by disentanglement.During inference, we enhance attribute correlations by target-guided counterfactual augmentation to further improve multi-aspect control.Experiments show that MAGIC outperforms stateof-the-art baselines in both imbalanced and balanced attribute correlation scenarios. Yi Liu 0071, Xiangyu Liu 0001, Xiangrong Zhu 0001, Wei Hu 0007 |
ACL (1) | 1 |
| 2024 | Task Oriented In-Domain Data AugmentationabstractLarge Language Models (LLMs) have shown superior performance in various applications and fields.To achieve better performance on specialized domains such as law and advertisement, LLMs are often continue pre-trained on in-domain data.However, existing approaches suffer from two major issues.First, in-domain data are scarce compared with general domainagnostic data.Second, data used for continual pre-training are not task-aware, such that they may not be helpful to downstream applications.We propose TRAIT, a task-oriented in-domain data augmentation framework.Our framework is divided into two parts: in-domain data selection and task-oriented synthetic passage generation.The data selection strategy identifies and selects a large amount of in-domain data from general corpora, and thus significantly enriches domain knowledge in the continual pre-training data.The synthetic passages contain guidance on how to use domain knowledge to answer questions about downstream tasks.By training on such passages, the model aligns with the need of downstream applications.We adapt LLMs to two domains: advertisement and math.On average, TRAIT improves LLM performance by 8% in the advertisement domain and 7.5% in the math domain. Simiao Zuo, Yeyun Gong, Qiang Lou, Yi Liu 0071, Shao-Lun Huang, Jian Jiao 0007 |
EMNLP | 6 |
| 2023 | Graph Convolution Synthetic Transformer for Chronic Kidney Disease Onset Prediction
Yi Liu 0071, Weitong Chen 0001, Yanda Wang, Yefan Huang, Xiaoli Wang 0002, Ken Cai, Bohan Li 0001 |
ADMA (3) | 2 |
| 2023 | Self-Supervised Dynamic Hypergraph Recommendation based on Hyper-Relational Knowledge GraphabstractKnowledge graphs (KGs) are commonly used as side information to enhance collaborative signals and improve recommendation quality. In the context of knowledge-aware recommendation (KGR), graph neural networks (GNNs) have emerged as promising solutions for modeling factual and semantic information in KGs. However, the long-tail distribution of entities leads to sparsity in supervision signals, which weakens the quality of item representation when utilizing KG enhancement. Additionally, the binary relation representation of KGs simplifies hyper-relational facts, making it challenging to model complex real-world information. Furthermore, the over-smoothing phenomenon results in indistinguishable representations and information loss. Yi Liu 0071, Hongrui Xuan, Bohan Li 0001, Meng Wang 0009, Tong Chen 0005, Hongzhi Yin |
CIKM | 1 |
| 2023 | Knowledge Enhancement for Contrastive Multi-Behavior RecommendationabstractA well-designed recommender system can accurately capture the attributes of users and items, reflecting the unique preferences of individuals. Traditional recommendation techniques usually focus on modeling the singular type of behaviors between users and items. However, in many practical recommendation scenarios (e.g., social media, e-commerce), there exist multi-typed interactive behaviors in user-item relationships, such as click, tag-as-favorite, and purchase in online shopping platforms. Thus, how to make full use of multi-behavior information for recommendation is of great importance to the existing system, which presents challenges in two aspects that need to be explored: (1) Utilizing users' personalized preferences to capture multi-behavioral dependencies; (2) Dealing with the insufficient recommendation caused by sparse supervision signal for target behavior. In this work, we propose a Knowledge Enhancement Multi-Behavior Contrastive Learning Recommendation (KMCLR) framework, including two Contrastive Learning tasks and three functional modules to tackle the above challenges, respectively. In particular, we design the multi-behavior learning module to extract users' personalized behavior information for user-embedding enhancement, and utilize knowledge graph in the knowledge enhancement module to derive more robust knowledge-aware representations for items. In addition, in the optimization stage, we model the coarse-grained commonalities and the fine-grained differences between multi-behavior of users to further improve the recommendation effect. Extensive experiments and ablation tests on the three real-world datasets indicate our KMCLR outperforms various state-of-the-art recommendation methods and verify the effectiveness of our method. Hongrui Xuan, Yi Liu 0071, Bohan Li 0001, Hongzhi Yin |
WSDM | 2 |
| 2023 | Bi-knowledge views recommendation based on user-oriented contrastive learning
Yi Liu 0071, Hongrui Xuan, Bohan Li 0001 |
J. Intell. Inf. Syst. | 1 |
| 2022 | GISDCN: A Graph-Based Interpolation Sequential Recommender with Deformable Convolutional Network
Yalei Zang, Yi Liu 0071, Weitong Chen 0001, Bohan Li 0001, Aoran Li, Lin Yue, Weihua Ma |
DASFAA (2) | 2 |
| 2021 | A Knowledge-Aware Recommender with Attention-Enhanced Dynamic Convolutional NetworkabstractSequential recommendation systems seek to learn users' preferences to predict their next actions based on the items engaged recently. Static behavior of users requires a long time to form, but short-term interactions with items usually meet some actual needs in reality and are more variable. RNN-based models are always constrained by the strong order assumption and are hard to model the complex and changeable data flexibly. Most of the CNN-based models are limited to the fixed convolutional kernel. All these methods are suboptimal when modeling the dynamics of item-to-item transitions. It is difficult to describe the items with complex relations and extract the fine-grained user preferences from the interaction sequence. To address these issues, we propose a knowledge-aware sequential recommender with the attention-enhanced dynamic convolutional network (KAeDCN). Our model combines the dynamic convolutional network with attention mechanisms to capture changing dependencies in the sequence. Meanwhile, we enhance the representations of items with Knowledge Graph (KG) information through an information fusion module to capture the fine-grained user preferences. The experiments on four public datasets demonstrate that KAeDCN outperforms most of the state-of-the-art sequential recommenders. Furthermore, experimental results also prove that KAeDCN can enhance the representations of items effectively and improve the extractability of sequential dependencies. Yi Liu 0071, Bohan Li 0001, Yalei Zang, Aoran Li, Hongzhi Yin |
CIKM | 1 |
| 2020 | A Survey on Blocking Technology of Entity Resolution
Bohan Li 0001, Yi Liu 0071, Anman Zhang, Wenhuan Wang, Shuo Wan |
J. Comput. Sci. Technol. | 2 |