Liyi Chen 0001

dblp:228/1335-1 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0003-2166-4386ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models
abstract
Wei Wu, Liyi Chen, Congxi Xiao, Tianfu Wang, Qimeng Wang, Chengqiang Lu, Yan Gao, Yiwu, Yao Hu, Hui Xiong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wei Wu 0045, Liyi Chen 0001, Congxi Xiao, Tianfu Wang 0002, Qimeng Wang, Chengqiang Lu, Yan Gao 0017, Yao Hu 0002, Hui Xiong 0001
ACL (1)2
2025 TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
abstract
Rapid advances in Large Language Models (LLMs) have spurred demand for processing extended context sequences in contemporary applications.However, this progress faces two challenges: performance degradation due to sequence lengths out-of-distribution, and excessively long inference times caused by the quadratic computational complexity of attention.These issues limit LLMs in long-context scenarios.In this paper, we propose Dynamic Token-Level KV Cache Selection (TokenSelect), a training-free method for efficient and accurate long-context inference.TokenSelect builds upon the observation of non-contiguous attention sparsity, using QK dot products to measure per-head KV Cache criticality at tokenlevel.By per-head soft voting mechanism, To-kenSelect selectively involves a few critical KV cache tokens in attention calculation without sacrificing accuracy.To further accelerate To-kenSelect, we design the Selection Cache based on observations of consecutive Query similarity and implemented the efficient Paged Dot Product Kernel, significantly reducing the selection overhead.A comprehensive evaluation of To-kenSelect demonstrates up to 23.84× speedup in attention computation and up to 2.28× acceleration in end-to-end latency, while providing superior performance compared to state-of-theart long-context inference methods.
Wei Wu 0045, Zhuoshi Pan, Kun Fu 0002, Chao Wang 0086, Liyi Chen 0001, Yunchu Bai, Tianfu Wang 0002, Zheng Wang 0027, Hui Xiong 0001
EMNLP5
2025 Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs
abstract
Proteins, as essential biomolecules, play a central role in biological processes, including metabolic reactions and DNA replication. Accurate prediction of their properties and functions is crucial in biological applications. Recent development of protein language models (pLMs) with supervised fine tuning provides a promising solution to this problem. However, the fine-tuned model is tailored for particular downstream prediction task, and achieving general-purpose protein understanding remains a challenge. In this paper, we introduce Structure-Enhanced Protein Instruction Tuning (SEPIT) framework to bridge this gap. Our approach incorporates a novel structure-aware module into pLMs to enrich their structural knowledge, and subsequently integrates these enhanced pLMs with large language models (LLMs) to advance protein understanding. In this framework, we propose a novel instruction tuning pipeline. First, we warm up the enhanced pLMs using contrastive learning and structure denoising. Then, caption-based instructions are used to establish a basic understanding of proteins. Finally, we refine this understanding by employing a mixture of experts (MoEs) to capture more complex properties and functional information with the same number of activated parameters. Moreover, we construct the largest and most comprehensive protein instruction dataset to date, which allows us to train and evaluate the general-purpose protein understanding model. Extensive experiments on both open-ended generation and closed-set answer tasks demonstrate the superior performance of SEPIT over both closed-source general LLMs and open-source LLMs trained with protein knowledge.
Wei Wu 0045, Chao Wang 0086, Liyi Chen 0001, Mingze Yin, Yiheng Zhu 0002, Kun Fu 0002, Jieping Ye, Hui Xiong 0001, Zheng Wang 0027
KDD (2)3
2025 Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential Recommendation
abstract
Multi-modal sequential recommendation (SR) leverages multi-modal data to learn more comprehensive item features and user preferences than traditional SR methods, which has become a critical topic in both academia and industry. Existing methods typically focus on enhancing multi-modal information utility through adaptive modality fusion to capture the evolving of user preference from user-item interaction sequences. However, most of them overlook the interference caused by redundant interest-irrelevant information contained in rich multi-modal data. Additionally, they primarily rely on implicit temporal information based solely on chronological ordering, neglecting explicit temporal signals that could more effectively represent dynamic user interest over time. To address these limitations, we propose a Hierarchical time-aware Mixture of experts for multi-modal Sequential Recommendation (HM4SR) with a two-level Mixture of Experts (MoE) and a multi-task learning strategy. Specifically, the first MoE, named Interactive MoE, extracts essential user interest-related information from the multi-modal data of each item. Then, the second MoE, termed Temporal MoE, captures user dynamic interests by introducing explicit temporal embeddings from timestamps in modality encoding. To further address data sparsity, we propose three auxiliary supervision tasks: sequence-level category prediction (CP) for item feature understanding, contrastive learning on ID (IDCL) to align sequence context with user interests, and placeholder contrastive learning (PCL) to integrate temporal information with modalities for dynamic interest modeling. Extensive experiments on four public datasets verify the effectiveness of HM4SR compared to several state-of-the-art approaches.
Shengzhe Zhang, Liyi Chen 0001, Dazhong Shen, Chao Wang 0086, Hui Xiong 0001
WWW2
2024 Temporal Graph Contrastive Learning for Sequential Recommendation
abstract
Sequential recommendation is a crucial task in understanding users' evolving interests and predicting their future behaviors. While existing approaches on sequence or graph modeling to learn interaction sequences of users have shown promising performance, how to effectively exploit temporal information and deal with the uncertainty noise in evolving user behaviors is still quite challenging. To this end, in this paper, we propose a Temporal Graph Contrastive Learning method for Sequential Recommendation (TGCL4SR) which leverages not only local interaction sequences but also global temporal graphs to comprehend item correlations and analyze user behaviors from a temporal perspective. Specifically, we first devise a Temporal Item Transition Graph (TITG) to fully leverage global interactions to understand item correlations, and augment this graph by dual transformations based on neighbor sampling and time disturbance. Accordingly, we design a Temporal item Transition graph Convolutional network (TiTConv) to capture temporal item transition patterns in TITG. Then, a novel Temporal Graph Contrastive Learning (TGCL) mechanism is designed to enhance the uniformity of representations between augmented graphs from identical sequences. For local interaction sequences, we design a temporal sequence encoder to incorporate time interval embeddings into the architecture of Transformer. At the training stage, we take maximum mean discrepancy and TGCL losses as auxiliary objectives. Extensive experiments on several real-world datasets show the effectiveness of TGCL4SR against state-of-the-art baselines of sequential recommendation.
Shengzhe Zhang, Liyi Chen 0001, Chao Wang 0086, Shuangli Li, Hui Xiong 0001
AAAI2
2024 Tackling Uncertain Correspondences for Multi-Modal Entity Alignment
abstract
Recently, multi-modal entity alignment has emerged as a pivotal endeavor for the integration of Multi-Modal Knowledge Graphs (MMKGs) originating from diverse data sources. Existing works primarily focus on fully depicting entity features by designing various modality encoders or fusion approaches. However, uncertain correspondences between inter-modal or intra-modal cues, such as weak inter-modal associations, description diversity, and modality absence, still severely hinder the effective exploration of aligned entity similarities. To this end, in this paper, we propose a novel Tackling uncertain correspondences method for Multi-modal Entity Alignment (TMEA). Specifically, to handle diverse attribute knowledge descriptions, we design alignment-augmented abstract representation that incorporates the large language model and in-context learning into attribute alignment and filtering for generating and embedding the attribute abstract. In order to mitigate the influence of the modality absence, we propose to unify all modality features into a shared latent subspace and generate pseudo features via variational autoencoders according to existing modal features. Then, we develop an inter-modal commonality enhancement mechanism based on cross-attention with orthogonal constraints, to address weak semantic associations between modalities. Extensive experiments on two real-world datasets validate the effectiveness of TMEA with a clear improvement over competitive baselines.
Liyi Chen 0001, Ying Sun 0006, Shengzhe Zhang, Yuyang Ye 0002, Wei Wu 0045, Hui Xiong 0001
NeurIPS1
2024 Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs
abstract
Large Language Models (LLMs) have shown remarkable reasoning capabilities on complex tasks, but they still suffer from out-of-date knowledge, hallucinations, and opaque decision-making. In contrast, Knowledge Graphs (KGs) can provide explicit and editable knowledge for LLMs to alleviate these issues. Existing paradigm of KG-augmented LLM manually predefines the breadth of exploration space and requires flawless navigation in KGs. However, this paradigm cannot adaptively explore reasoning paths in KGs based on the question semantics and self-correct erroneous reasoning paths, resulting in a bottleneck in efficiency and effect. To address these limitations, we propose a novel self-correcting adaptive planning paradigm for KG-augmented LLM named Plan-on-Graph (PoG), which first decomposes the question into several sub-objectives and then repeats the process of adaptively exploring reasoning paths, updating memory, and reflecting on the need to self-correct erroneous reasoning paths until arriving at the answer. Specifically, three important mechanisms of Guidance, Memory, and Reflection are designed to work together, to guarantee the adaptive breadth of self-correcting planning for graph reasoning. Finally, extensive experiments on three real-world datasets demonstrate the effectiveness and efficiency of PoG.
Liyi Chen 0001, Panrong Tong, Zhongming Jin 0001, Ying Sun 0006, Jieping Ye, Hui Xiong 0001
NeurIPS1
2024 AFDGCF: Adaptive Feature De-correlation Graph Collaborative Filtering for Recommendations
abstract
Collaborative filtering methods based on graph neural networks (GNNs) have witnessed significant success in recommender systems (RS), capitalizing on their ability to capture collaborative signals within intricate user-item relationships via message-passing mechanisms. However, these GNN-based RS inadvertently introduce excess linear correlation between user and item embeddings, contradicting the goal of providing personalized recommendations. While existing research predominantly ascribes this flaw to the over-smoothing problem, this paper underscores the critical, often overlooked role of the over-correlation issue in diminishing the effectiveness of GNN representations and subsequent recommendation performance. Up to now, the over-correlation issue remains unexplored in RS. Meanwhile, how to mitigate the impact of over-correlation while preserving collaborative filtering signals is a significant challenge. To this end, this paper aims to address the aforementioned gap by undertaking a comprehensive study of the over-correlation issue in graph collaborative filtering models. Firstly, we present empirical evidence to demonstrate the widespread prevalence of over-correlation in these models. Subsequently, we dive into a theoretical analysis which establishes a pivotal connection between the over-correlation and over-smoothing issues. Leveraging these insights, we introduce the Adaptive Feature De-correlation Graph Collaborative Filtering (AFDGCF) framework, which dynamically applies correlation penalties to the feature dimensions of the representation matrix, effectively alleviating both over-correlation and over-smoothing issues. The efficacy of the proposed framework is corroborated through extensive experiments conducted with four representative graph collaborative filtering models across four publicly available datasets. Our results show the superiority of AFDGCF in enhancing the performance landscape of graph collaborative filtering models.
Wei Wu 0045, Chao Wang 0086, Dazhong Shen, Chuan Qin 0002, Liyi Chen 0001, Hui Xiong 0001
SIGIR5
2024 COMI: COrrect and MItigate Shortcut Learning Behavior in Deep Neural Networks
abstract
Deep Neural Networks (DNNs), despite their notable progress across information retrieval tasks, encounter the issues of shortcut learning and struggle with poor generalization due to their reliance on spurious correlations between features and labels. Current research mainly mitigates shortcut learning behavior using augmentation and distillation techniques, but these methods could be laborious and introduce unwarranted biases. To tackle these, in this paper, we propose COMI, a novel method to COrrect and MItigate shortcut learning behavior. Inspired by the ways students solve shortcuts in educational scenarios, we aim to reduce model's reliance on shortcuts and enhance its ability to extract underlying information integrated with standard Empirical Risk Minimization (ERM). Specifically, we first design Correct Habit (CoHa) strategy to retrieve the top m challenging samples for priority training, which encourages model to rely less on shortcuts in the early training. Then, to extract more meaningful underlying information, the information derived from ERM is separated into task-relevant and task-irrelevant information, the former serves as the primary basis for model predictions, while the latter is considered non-essential. However, within task-relevant information, certain potential shortcuts contribute to overconfident predictions. To mitigate this, we design Deep Mitigation (DeMi) network with shortcut margin loss to adaptively control the feature weights of shortcuts and eliminate their influence. Besides, to counteract unknown shortcut tokens issue in NLP, we adopt locally interpretable module-LIME to help recognize shortcut tokens. Finally, extensive experiments conducted on NLP and CV tasks demonstrate the effectiveness of COMI, which can perform well on both IID and OOD samples.
Lili Zhao 0002, Qi Liu 0003, Linan Yue, Wei Chen 0156, Liyi Chen 0001, Ruijun Sun
SIGIR5
2024 Collaboration-Aware Hybrid Learning for Knowledge Development Prediction
abstract
In recent years, the rise of online Knowledge Management Systems (KMSs) has significantly improved work efficiency in enterprises. Knowledge development prediction, as a critical application within these online platforms, enables organizations to proactively address knowledge gaps and align their learning initiatives with evolving job requirements. However, it still confronts challenges in exploring the influence of collaborative networks on knowledge development and adapting to ecological situations in working environment. To this end, in this paper, we propose a Collaboration-Aware Hybrid Learning approach (CAHL) for predicting the future knowledge acquisition of employees and quantifying the impact of various knowledge learning patterns. Specifically, to fully harness the inherent rules of knowledge development, we first learn the knowledge co-occurrence and prerequisite relationships with an association prompt attention mechanism to generate effective knowledge representations through a specially-designed Job Knowledge Embedding module. Then, we aggregate the features of mastering knowledge and work collaborators for employee representations in another Employee Embedding module. Moreover, we propose to model the process of employee knowledge development via a Hybrid Learning Simulation module that integrates both collaborative learning and self learning to predict future-acquired job knowledge of employees. Finally, extensive experiments conducted on a real-world dataset clearly validate the effectiveness of CAHL.
Liyi Chen 0001, Chuan Qin 0002, Ying Sun 0006, Tong Xu 0001, Hengshu Zhu, Hui Xiong 0001
WWW1
2023 Entity Summarization via Exploiting Description Complementarity and Salience
abstract
Entity summarization is a novel and efficient way to understand real-world facts and solve the increasing information overload problem in large-scale knowledge graphs (KG). Existing studies mainly rely on ranking independent entity descriptions as a list under a certain scoring standard such as importance. However, they often ignore the relatedness and even semantic overlap between individual descriptions. This may seriously interfere with the contribution judgment of descriptions for entity summarization. Actually, the entity summary is a whole to comprehensively integrate the main aspects of entity descriptions, which could be naturally treated as a set. Unfortunately, the exploration of these set characteristics for entity summarization is still an open issue with great challenges. To that end, we draw inspiration from a set completion perspective and propose an entity summarization method with complementarity and salience (ESCS) to deeply exploit description complementarity and salience in order to form a summary set for the target entity. Specifically, we first generate entity description representations with textual features in the description embedding module. For the purpose of learning complementary relationships within the entire summary set, we devise a bi-directional long short-term memory structure to capture global complementarity for each summary in the summary complementarity learning module. Meanwhile, in order to estimate the salience of individual descriptions, we calculate similarities between semantic embeddings of the target entity and its property-value pairs in the description salience learning module. Next, with a joint learning stage, we can optimize ESCS from a set completion perspective. Finally, a summary generation strategy is designed to infer the entire summary set step-by-step for the target entity. Extensive experiments on a public benchmark have clearly demonstrated the effectiveness of ESCS and revealed the potential of set completion in entity summarization task.
Liyi Chen 0001, Zhi Li 0057, Weidong He, Gong Cheng 0001, Tong Xu 0001, Nicholas Jing Yuan, Enhong Chen
IEEE Trans. Neural Networks Learn. Syst.1
2022 Text Classification via Learning Semantic Dependency and Association
abstract
Text classification is a fundamental and classical problem in natural language processing. Existing methods in this area attach more attention to structure modeling of texts, while largely ignoring the cognitive principles of human reading. Actually, as an important aspect of exploring the characteristics of language comprehension, neuroscience research in recent years has demonstrated the human instinct for abstract thinking, where semantic processing and summarizing play essential roles. To this end, we propose a novel text classification method with Semantic Dependency and Association (SDA). To be specific, SDA comprises two coupled spaces, namely Dependency Space and Association Space, to imitate the real process when readers are inferring text semantics. Firstly, in Dependency Space, we devise a graph modeling structure to thoroughly capture the dependency relations between words. Then in Association Space, we associate the words with prior concepts and adopt attention mechanisms to match words with the most related concepts based on their dependency relations. Later, we synthesize the above two aspects to enhance the semantic representation of words. Afterwards, we design a Semantic Summarizing module to integrate the semantic information learned above. Finally, extensive experiments are conducted on four public English and Chinese text classification datasets, where experimental results not only demonstrate the effectiveness of our model but also reveal the superiority of the perspective from the human cognition for text comprehension.
Guanqi Zhu, Hanqing Tao, Han Wu 0002, Liyi Chen 0001, Ye Liu 0011, Qi Liu 0003, Enhong Chen
IJCNN4
2022 Multi-modal Siamese Network for Entity Alignment
abstract
The booming of multi-modal knowledge graphs (MMKGs) has raised the imperative demand for multi-modal entity alignment techniques, which facilitate the integration of multiple MMKGs from separate data sources. Unfortunately, prior arts harness multi-modal knowledge only via the heuristic merging of uni-modal feature embeddings. Therefore, inter-modal cues concealed in multi-modal knowledge could be largely ignored. To deal with that problem, in this paper, we propose a novel Multi-modal Siamese Network for Entity Alignment (MSNEA) to align entities in different MMKGs, in which multi-modal knowledge could be comprehensively leveraged by the exploitation of inter-modal effect. Specifically, we first devise a multi-modal knowledge embedding module to extract visual, relational, and attribute features of entities to generate holistic entity representations for distinct MMKGs. During this procedure, we employ inter-modal enhancement mechanisms to integrate visual features to guide relational feature learning and adaptively assign attention weights to capture valuable attributes for alignment. Afterwards, we design a multi-modal contrastive learning module to achieve inter-modal enhancement fusion with avoiding the overwhelming impact of weak modalities. Experimental results on two public datasets demonstrate that our proposed MSNEA provides state-of-the-art performance with a large margin compared with competitive baselines.
Liyi Chen 0001, Zhi Li 0057, Tong Xu 0001, Han Wu 0002, Zhefeng Wang 0001, Nicholas Jing Yuan, Enhong Chen
KDD1
2021 Linking the Characters: Video-oriented Social Graph Generation via Hierarchical-cumulative GCN
abstract
Recent years have witnessed the booming of online video platforms. Along this line, a graph to illustrate social relation among characters has been long expected to not only benefit the audiences for better understanding the story, but also support the fine-grained video analysis task in a semantic way. Unfortunately, though we humans could easily infer the social relations among characters, it is still an extremely challenging task for intelligent systems to automatically capture the social relation by absorbing multi-modal cues. Besides, they fail to describe the relations among multiple characters in a graph-generation perspective. To that end, inspired by the human inference ability on social relationship, we propose a novel Hierarchical- Cumulative Graph Convolutional Network (HC-GCN) to generate the social relation graph for multiple characters in the video. Specifically, we first integrate the short-term multi-modal cues, including visual, textual and audio information, to generate the frame-level graphs for part of characters via multimodal graph convolution technique. While dealing with the video-level aggregation task, we design an end-to-end framework to aggregate all frame-level subgraphs along the temporal trajectory, which results in a global video-level social graph with various social relationships among multiple characters. Extensive validations on two real-world large-scale datasets demonstrate the effectiveness of our proposed method compared with SOTA baselines.
Joya Chen, Tong Xu 0001, Liyi Chen 0001, Lingfei Wu 0001, Yao Hu 0002, Enhong Chen
ACM Multimedia4
2020 MMEA: Entity Alignment for Multi-modal Knowledge Graph
Liyi Chen 0001, Zhi Li 0057, Yijun Wang 0002, Tong Xu 0001, Zhefeng Wang 0001, Enhong Chen
KSEM (1)1