Kang Liu 0024

dblp:228/1231-24 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0001-6789-1811ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Bidirectional Counterfactual Distillation for Review-Based Recommendation
abstract
Review-based recommendation methods typically integrate multiple behaviors, including interactions, reviews, and ratings, to model user preferences. To effectively extract preference signals from diverse behaviors, some studies train multiple student models to capture distinct behavioral patterns, and leverage online distillation to facilitate collaborative learning among them. However, we argue that these techniques suffer from bias contamination from rating distributions and feature homogenization during cross-behavior knowledge transfer: (1) Rating distribution bias, arising from non-uniform historical ratings, propagates across behaviors through distillation, contaminating the true preference representations of other behaviors. (2) Static distillation strategies often lead to homogenized behavioral features, hindering the learning of behavior-specific preferences. To address these issues, we propose a novel Bidirectional Counterfactual Distillation (BiCoD) framework for review-based recommendation. In BiCoD, we first design an adversarial counterfactual distillation module to suppress the impact of non-uniform rating distributions on distillation, thereby preventing it from contaminating the user's true preference representations across behaviors. Subsequently, we introduce a stage-aware bidirectional distillation strategy to enhance the distinctiveness of behavioral features, facilitating the effective learning of behavior-specific preferences. Extensive experiments on five real-world datasets validate the effectiveness and superiority of the proposed framework.
Sheng Sang, Shujie Li 0002, Shuaiyang Li 0001, Kang Liu 0024, Wei Jia 0001, Dan Guo 0001, Feng Xue 0002
AAAI4
2026 Cross-modal feature disentangling via bidirectional distillation for multimodal recommendation
Shuaiyang Li 0001, Kang Liu 0024, Shujie Li 0002, Dan Guo 0001, Feng Xue 0002
Expert Syst. Appl.2
2026 Denoising precedes enrichment: aspect-sentiment-aware graph contrastive distillation network for review-based recommendation
Tongtong Yu, Sheng Sang, Kang Liu 0024
Inf. Sci.3
2026 ACD: Adversarial Counterfactual Distillation for Rating Prediction in Recommendation
abstract
Rating prediction is a classic task in recommendation systems, aiming to accurately estimate user ratings for various items. Historical ratings typically exhibit a non-uniform distribution, leading recommendation models to favor predicting high-frequency ratings. We refer to the inconsistency between predicted ratings and users' true preferences caused by non-uniform rating distributions asrating bias. To mitigate this bias, existing studies capture various interaction behavior patterns and employ knowledge distillation techniques to improve the network's ability to model user preferences. However, due to the model being trained on datasets with non-uniform rating distributions, the rating bias may propagate through the knowledge distillation process across different behaviors, thereby contaminating the modeling of users' true preferences. To this end, we propose a novelAdversarial Counterfactual Distillation(ACD) framework for the rating prediction task, aimed at eliminating rating bias. Specifically, we design aCounterfactual Distillation Modulefrom a causal reasoning perspective to facilitate knowledge transfer across various interaction behaviors while concurrently mitigating bias contamination. Furthermore, we introduce anAdversarial Debiasing Moduleto dynamically adjust the debiasing strength, ensuring that the model maintains an optimal balance between effective knowledge transfer and bias mitigation. Extensive experiments demonstrate the superior performance of our proposed ACD framework. The complete code is publicly available athttps://github.com/hfutmars/ACD.
Sheng Sang, Feng Xue 0002, Shuaiyang Li 0001, Kang Liu 0024, Richang Hong
IEEE Trans. Knowl. Data Eng.4
2025 BiKD2: Bidirectional Knowledge Distillation-enhanced explicit graph Disentangling network for review-based recommendation
Kang Liu 0024, Tongtong Yu, Yueli Song, Yuling Li 0001
Eng. Appl. Artif. Intell.1
2025 Hierarchical feature-guided prototypical network for few-shot knowledge graph completion
Yuling Li 0001, Kui Yu, Chunfeng Shen, Ji Chang, Kang Liu 0024
Neural Networks7
2025 Corrigendum to "Hierarchical Feature-guided Prototypical Network for Few-shot Knowledge Graph Completion" [Neural Networks Volume 191, November 2025, 107702/NN_107702]
Yuling Li 0001, Kui Yu, Chunfeng Shen, Ji Chang, Kang Liu 0024
Neural Networks7
2025 REDGCN: Rating-Oriented Explicit Disentangling Graph Convolution Network for Review-Aware Recommendation
abstract
Rating prediction is a challenging task in review-aware recommendation. Although current methods effectively combine collaborative signals with review data, they fail to differentiate user preferences across various ratings and overlook the independence between these ratings. In this article, we emphasize the importance of independence modeling among representations for different rating levels. To this end, we propose a rating-oriented explicit disentangling graph convolution network for review-aware recommendation, short for REDGCN. Specifically, we introduce a rating-oriented disentangled representation learning that segments representations and rating graph based on ratings. It also employs an explicit graph learning approach to ensure the independence of disentangled representations during information propagation, which mitigates noise from review features. Furthermore, we define and model one kind of cross-rating correlation, based on the characteristics of user rating behavior. By leveraging this approach, we introduce a cross-rating constraint as an additional task to further enhance the independence among disentangled representations and improve the stability of model training. We conduct extensive experiments on six public datasets to prove the effectiveness of REDGCN. The complete data and codes of REDGCN are available athttps://github.com/hfutmars/REDGCN.
Sheng Sang, Feng Xue 0002, Kang Liu 0024, Shuaiyang Li 0001, Richang Hong
IEEE Trans. Comput. Soc. Syst.4
2024 Multimodal Hierarchical Graph Collaborative Filtering for Multimedia-Based Recommendation
abstract
Multimedia-based recommendation (MMRec) is a challenging task, which goes beyond the collaborative filtering (CF) schema that only captures collaborative signals from interactions and explores multimodal user preference cues hidden in complex multimedia content. Despite the significant progress of current solutions for MMRec, we argue that they are limited by multimodal noise contamination. Specifically, a considerable amount of preference-irrelevant multimodal noise (e.g., the background, layout, and brightness in the image of the product) is incorporated into the representation learning of items, which contaminates the modeling of multimodal user preferences. Moreover, most of the latest researches are based on graph convolution networks (GCNs), which means that multimodal noise contamination is further amplified because noisy information is continuously propagated over the user–item interaction graph as recursive neighbor aggregations are performed. To address this problem, instead of the common MMRec paradigm which learns user preferences in an integrated manner, we propose a hierarchical framework to separately learn collaborative signals and multimodal preferences cues, thus preventing multimodal noise from flowing into collaborative signals. Then, to alleviate the noise contamination for multimodal user preference modeling, we propose to extract semantic entities from multimodal content that are more relevant to user interests, which can model semantic-level multimodal preferences and thus remove a large fraction of noise. Furthermore, we use the full multimodal features to model content-level multimodal preferences like the existing MMRec solutions, which ensures the sufficient utilization of multimodal information. Overall, we develop a novel model, multimodal hierarchical graph CF (MHGCF), which consists of three types of GCN modules tailored to capture collaborative signals, semantic-level preferences, and content-level preferences, respectively. We conduct extensive experiments to demonstrate the effectiveness of MHGCF and its components. The complete data and codes of MHGCF are available athttps://github.com/hfutmars/MHGCF.
Kang Liu 0024, Feng Xue 0002, Shuaiyang Li 0001, Sheng Sang, Richang Hong
IEEE Trans. Comput. Soc. Syst.1
2024 Multimodal Graph Causal Embedding for Multimedia-Based Recommendation
abstract
Multimedia-based recommendation (MMRec) models typically rely on observed user-item interactions and the multimodal content of items, such as visual images and textual descriptions, to predict user preferences. Among these, the user's preference for the displayed multimodal content of items is crucial for interacting with a particular item. We argue that users' preference behaviors (i.e., user-item interactions) for the modality content of items, beyond stemming from their real interest in the modality content, may also be influenced by their conformity to the popularity of items' modality-specific content (e.g., a user might be motivated to interact with a lipstick due to enthusiastic discussions among other users regarding textual reviews of the product). In essence, user-item interactions are jointly triggered by real interest and conformity. However, most existing MMRec models primarily concentrate on modeling users' interest preferences when capturing multimodal user preferences, neglecting the modeling of their conformity preferences, which results in sub-optimal recommendation performance. In this work, we resort to causal theory to propose a novel MMRec model, termed Multimodal Graph Causal Embedding (MGCE), revealing insights into the crucial causal relations of users' modality-specific interest and conformity in interaction behaviors within MMRec scenarios. Inspired by the colliding effect in causal inference and integrating the characteristics of real interest and conformity, we devise multimodal causal embedding learning networks to facilitate the learning of high-quality causal embeddings (multimodal interest and multimodal conformity embeddings) from both the structure-level and feature-level, yielding state-of-the-art performance. Extensive experimental results on three datasets demonstrate the effectiveness of MGCE.
Shuaiyang Li 0001, Feng Xue 0002, Kang Liu 0024, Dan Guo 0001, Richang Hong
IEEE Trans. Knowl. Data Eng.3
2023 Multimodal Counterfactual Learning Network for Multimedia-based Recommendation
abstract
Multimedia-based recommendation (MMRec) utilizes multimodal content (images, textual descriptions, etc.) as auxiliary information on historical interactions to determine user preferences. Most MMRec approaches predict user interests by exploiting a large amount of multimodal contents of user-interacted items, ignoring the potential effect of multimodal content of user-uninteracted items. As a matter of fact, there is a small portion of user preference-irrelevant features in the multimodal content of user-interacted items, which may be a kind of spurious correlation with user preferences, thereby degrading the recommendation performance. In this work, we argue that the multimodal content of user-uninteracted items can be further exploited to identify and eliminate the user preference-irrelevant portion inside user-interacted multimodal content, for example by counterfactual inference of causal theory. Going beyond multimodal user preference modeling only using interacted items, we propose a novel model called Multimodal Counterfactual Learning Network (MCLN), in which user-uninteracted items' multimodal content is additionally exploited to further purify the representation of user preference-relevant multimodal content that better matches the user's interests, yielding state-of-the-art performance. Extensive experiments are conducted to validate the effectiveness and rationality of MCLN. We release the complete codes of MCLN at https://github.com/hfutmars/MCLN.
Shuaiyang Li 0001, Dan Guo 0001, Kang Liu 0024, Richang Hong, Feng Xue 0002
SIGIR3
2023 Joint Multi-Grained Popularity-Aware Graph Convolution Collaborative Filtering for Recommendation
abstract
Graph convolution networks (GCNs), with their efficient ability to capture high-order connectivity in graphs, have been widely applied in recommender systems. Stacking multiple neighbor aggregation is the major operation in GCNs. It implicitly captures popularity features because the number of neighbor nodes reflects the popularity of a node. However, existing GCN-based methods ignore a universal problem: users’ sensitivity to item popularity is differentiated, but the neighbor aggregations in GCNs actually fix this sensitivity through graph Laplacian normalization, leading to suboptimal personalization. In this work, we propose to model multigrained popularity features and jointly learn them together with high-order connectivity to match the differentiation of user preferences exhibited in popularity features. Specifically, we develop a Joint Multigrained Popularity-aware Graph Convolution Collaborative Filtering model, short for JMP-GCF, which uses a popularity-aware embedding generation to construct multigrained popularity features and uses the idea of joint learning to capture the signals within and between different granularities of popularity features that are relevant for modeling user preferences. In addition, we propose a multistage stacked training strategy to speed up model convergence. We conduct extensive experiments on three public datasets to show the state-of-the-art performance of JMP-GCF. The complete codes of JMP-GCF are released athttps://github.com/hfutmars/JMP-GCF.
Kang Liu 0024, Feng Xue 0002, Xiangnan He 0001, Dan Guo 0001, Richang Hong
IEEE Trans. Comput. Soc. Syst.1
2023 Multimodal Graph Contrastive Learning for Multimedia-Based Recommendation
abstract
Multimedia-based recommendation is a challenging task that requires not only learning collaborative signals from user-item interaction, but also capturing modality-specific user interest clues from complex multimedia content. Though significant progress on this challenge has been made, we argue that current solutions remain limited by multimodal noise contamination. Specifically, a considerable proportion of multimedia content is irrelevant to the user preference, such as the background, overall layout, and brightness of images; the word order and semantic-free words in titles;etc. We take this irrelevant information as noise contamination to discover user preferences. Moreover, most recent research has been conducted by graph learning. This means that noise is diffused into the user and item representations with the message propagation; the contamination influence is further amplified. To tackle this problem, we develop a novel framework named Multimodal Graph Contrastive Learning (MGCL), which captures collaborative signals from interactions and uses visual and textual modalities to respectively extract modality-specific user preference clues. The key idea of MGCL involves two aspects: First, to alleviate noise contamination during graph learning, we construct three parallel graph convolution networks to independently generate three types of user and item representations, containing collaborative signals, visual preference clues, and textual preference clues. Second, to eliminate as much preference-independent noisy information as possible from the generated representations, we incorporate sufficient self-supervised signals into the model optimization with the help of contrastive learning, thus enhancing the expressiveness of the user and item representations. Note that MGCL is not limited to graph learning schema, but also can be applied to most matrix factorization methods. We conduct extensive experiments on three public datasets to validate the effectiveness and scalability of MGCL11We release the codes of MGCL athttps://github.com/hfutmars/MGCL..
Kang Liu 0024, Feng Xue 0002, Dan Guo 0001, Peijie Sun, Shengsheng Qian, Richang Hong
IEEE Trans. Multim.1
2023 MEGCF: Multimodal Entity Graph Collaborative Filtering for Personalized Recommendation
abstract
In most E-commerce platforms, whether the displayed items trigger the user’s interest largely depends on their most eye-catching multimodal content. Consequently, increasing efforts focus on modeling multimodal user preference, and the pressing paradigm is to incorporate complete multimodal deep features of the items into the recommendation module. However, the existing studies ignore the mismatch problem between multimodal feature extraction (MFE) and user interest modeling (UIM) . That is, MFE and UIM have different emphases. Specifically, MFE is migrated from and adapted to upstream tasks such as image classification. In addition, it is mainly a content-oriented and non-personalized process, while UIM, with its greater focus on understanding user interaction, is essentially a user-oriented and personalized process. Therefore, the direct incorporation of MFE into UIM for purely user-oriented tasks, tends to introduce a large number of preference-independent multimodal noise and contaminate the embedding representations in UIM. This paper aims at solving the mismatch problem between MFE and UIM, so as to generate high-quality embedding representations and better model multimodal user preferences. Towards this end, we develop a novel model, m ultimodal e ntity g raph c ollaborative f iltering, short for MEGCF. The UIM of the proposed model captures the semantic correlation between interactions and the features obtained from MFE, thus making a better match between MFE and UIM. More precisely, semantic-rich entities are first extracted from the multimodal data, since they are more relevant to user preferences than other multimodal information. These entities are then integrated into the user-item interaction graph. Afterwards, a symmetric linear Graph Convolution Network (GCN) module is constructed to perform message propagation over the graph, in order to capture both high-order semantic correlation and collaborative filtering signals. Finally, the sentiment information from the review data are used to fine-grainedly weight neighbor aggregation in the GCN, as it reflects the overall quality of the items, and therefore it is an important modality information related to user preferences. Extensive experiments demonstrate the effectiveness and rationality of MEGCF. 1
Kang Liu 0024, Feng Xue 0002, Dan Guo 0001, Le Wu 0001, Shujie Li 0002, Richang Hong
ACM Trans. Inf. Syst.1
2023 LCSNet: End-to-end Lipreading with Channel-aware Feature Selection
abstract
Lipreading is a task of decoding the movement of the speaker’s lip region into text. In recent years, lipreading methods based on deep neural network have attracted widespread attention, and the accuracy has far surpassed that of experienced human lipreaders. The visual differences in some phonemes are extremely subtle and pose a great challenge to lipreading. Most of the lipreading existing methods do not process the extracted visual features, which mainly suffer from two problems. First, the extracted features contain lot of useless information such as noise caused by differences in speech speed and lip shape, for example. In addition, the extracted features are not abstract enough to distinguish phonemes with similar pronunciation. These problems have a bad effect on the performance of lipreading. To extract features from the lip regions that are more distinguishable and more relevant to the speech content, this article proposes an end-to-end deep neural network-based lipreading model (LCSNet). The proposed model extracts the short-term spatio-temporal features and the motion trajectory features from the lip region in the video clips. The extracted features are filtered by the channel attention module to eliminate the useless features and then used as input to the proposed Selective Feature Fusion Module (SFFM) to extract the high-level abstract features. Afterwards, these features are used as input to the bidirectional GRU network in time order for temporal modeling to obtain the long-term spatio-temporal features. Finally, a Connectionist Temporal Classification (CTC) decoder is used to generate the output text. The experimental results show that the proposed model achieves a 1.0% CER and 2.3% WER on the GRID corpus database, which, respectively, represents an improvement of 52% and 47% compared to LipNet.
Feng Xue 0002, Kang Liu 0024, Zikun Hong, Mingwei Cao, Dan Guo 0001, Richang Hong
ACM Trans. Multim. Comput. Commun. Appl.3
2022 RGCF: Refined graph convolution collaborative filtering with concise and expressive embedding
abstract
Graph Convolution Networks (GCNs) have attracted significant attention and have become the most popular method for learning graph representations. In recent years, many efforts have focused on integrating GCNs into recommender tasks and have made remarkable progress. At its core is to explicitly capture the high-order connectivities between nodes in the user-item bipartite graph. However, we found some potential drawbacks existed in the traditional GCN-based recommendation models are that the excessive information redundancy yield by the nonlinear graph convolution operation reduces the expressiveness of the resultant embeddings, and the important popularity features that are effective in sparse recommendation scenarios are not encoded in the embedding generation process. In this work, we develop a novel GCN-based recommendation model, named Refined Graph convolution Collaborative Filtering (RGCF), where a refined graph convolution structure is designed to match non-semantic ID inputs. In addition, a new fine-tuned symmetric normalization is proposed to mine node popularity characteristics and further incorporate the popularity features into the embedding learning process. Extensive experiments were conducted on three public million-size datasets, and the RGCF improved by an average of 13.45% over the state-of-the-art baseline. Further comparative experiments validated the effectiveness and rationality of each part of our proposed RGCF. We released our code at https://github.com/hfutmars/RGCF.
Kang Liu 0024, Feng Xue 0002, Richang Hong
Intell. Data Anal.1