Feng Xue 0002

dblp:03/517-2 · DBLP profile ↗
← Back
38ranked-venue papers
14as first author
24since 2021 · last 2026
0000-0003-4962-9734ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Event-Guided Scene Text Image Super-Resolution
abstract
Scene text image super-resolution aims to enhance text legibility by recovering high-resolution text images from low-resolution inputs. However, maintaining fine details such as text strokes, edges, and textual accuracy remains challenging, particularly in low-light environments and high-speed motion scenarios, where degradation is more severe. Event cameras, with their high temporal resolution and ability to capture intensity changes, offer a promising solution for restoring lost fine details and mitigating degradation in these challenging conditions. In this paper, we propose EvTSR, the first framework that integrates Event data for scene Text image Super-Resolution. The core of EvTSR is the dual-stream frequency boost (DSFB) mechanism, which separates image features into high- and low-frequency components. High-frequency details like edges and strokes are enhanced using event data via the event-guided high-frequency (EGH) mechanism, while low-frequency components, responsible for global structure, are refined using the Text-Guided Low-frequency (TGL) mechanism with a pre-trained text recognizer, ensuring textual coherence. To further improve cross-modal integration, we introduce the cross-modal fusion (CMF) mechanism, which effectively aligns event and image features, enabling robust information fusion. Extensive experiments demonstrate that EvTSR achieves superior performance over existing methods.
Zihan Qi, Zeyu Xiao 0002, Haoyi Zhao, Yang Zhao 0002, Feng Xue 0002, Wei Jia 0001
AAAI5
2026 Bidirectional Counterfactual Distillation for Review-Based Recommendation
abstract
Review-based recommendation methods typically integrate multiple behaviors, including interactions, reviews, and ratings, to model user preferences. To effectively extract preference signals from diverse behaviors, some studies train multiple student models to capture distinct behavioral patterns, and leverage online distillation to facilitate collaborative learning among them. However, we argue that these techniques suffer from bias contamination from rating distributions and feature homogenization during cross-behavior knowledge transfer: (1) Rating distribution bias, arising from non-uniform historical ratings, propagates across behaviors through distillation, contaminating the true preference representations of other behaviors. (2) Static distillation strategies often lead to homogenized behavioral features, hindering the learning of behavior-specific preferences. To address these issues, we propose a novel Bidirectional Counterfactual Distillation (BiCoD) framework for review-based recommendation. In BiCoD, we first design an adversarial counterfactual distillation module to suppress the impact of non-uniform rating distributions on distillation, thereby preventing it from contaminating the user's true preference representations across behaviors. Subsequently, we introduce a stage-aware bidirectional distillation strategy to enhance the distinctiveness of behavioral features, facilitating the effective learning of behavior-specific preferences. Extensive experiments on five real-world datasets validate the effectiveness and superiority of the proposed framework.
Sheng Sang, Shujie Li 0002, Shuaiyang Li 0001, Kang Liu 0024, Wei Jia 0001, Dan Guo 0001, Feng Xue 0002
AAAI8
2026 LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition
abstract
Visual Speech Recognition (VSR), commonly known as lipreading, enables the recognition of spoken text by analyzing lip visual features. Due to the subtlety of lip movements, its recognition is much harder than other motion recognition tasks. Existing VSR models face the challenge of viseme ambiguity when processing phonemes with similar pronunciations—multiple phonemes share similar viseme features, leading to a notable drop in lipreading accuracy. To address this issue, this study proposes a Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition(LinProVSR) framework. First, an ambiguous sample set is constructed based on linguistic knowledge to provide supervisory signals for the model's training. Then, a Progressive Contrastive Disambiguation Network (PCDN) is designed, which progressively enhances the model's ability to capture the subtle viseme differences corresponding to similar phonemes through viseme-phoneme contrastive disambiguation in the encoding stage and text contrastive disambiguation in the decoding stage. Furthermore, we pioneer the Ambiguous Word Error Rate (AWER) metric specifically for evaluating recognition of phonetically ambiguous text, and verify the effectiveness of the proposed method on multiple public datasets, achieving a significant breakthrough especially in distinguishing visually similar phonemes.
Feng Xue 0002, Baochao Zhu, Wei Jia 0001, Shujie Li 0002, Yu Li 0053, Shengeng Tang, Dan Guo 0001
AAAI1
2026 Cross-modal feature disentangling via bidirectional distillation for multimodal recommendation
Shuaiyang Li 0001, Kang Liu 0024, Shujie Li 0002, Dan Guo 0001, Feng Xue 0002
Expert Syst. Appl.5
2026 D2M2Lip: dual-domain feature fusion and motion magnification for lip reading
Baochao Zhu, Shujie Li 0002, Feng Xue 0002
Multim. Syst.4
2026 CFLip: Generalizing Lipreading to Unseen Speakers by Learning Common Features
abstract
Lipreading refers to translating the lip movements observed in a video of a speaker into corresponding textual outputs, providing a visual alternative to auditory communication for individuals who are deaf or hard of hearing. Existing lipreading methods typically independently learn the lip movements of each speaker. This results in the model being highly sensitive to the individual visual features (lip color/shape) of the speakers in the training set, hindering the generalization of lipreading models. Despite the obvious visual variations in the lips of different speakers, we claim that there are still inherent common features when they pronounce the same phoneme. We attempt to learn the common pronunciation features across different speakers, so as to achieve better generalization of lipreading model to unseen speakers. In this article, we propose a sentence-level lipreading framework based on Learning Common Features (CFLip), designed to extract common pronunciation features fromvideo pairs. Specifically, we first employ data augmentation strategy to generate pseudo videos that share labels but with different speakers by replacing frame segments in real videos. With thesevideo pairs, we designed a dual-stream network to learn commonality feature by minimized the distance between the features of different speakers pronouncing the same words via Generalization Loss. Extensive experiments on benchmark datasets demonstrate that the proposed CFLip can effectively generalize to unseen speakers.
Yu Li 0053, Feng Xue 0002, Dan Guo 0001, Shengeng Tang, Shujie Li 0002, Richang Hong
IEEE Trans. Comput. Soc. Syst.2
2026 Wiener-Deconvolution-Driven Event-Based Deblurring for Low-Light Imaging
abstract
We address event-based deblurring for low-light imaging, where conventional frames suffer severe blur, noise and saturation, while events capture sharp high-frequency contrast changes with microsecond latency that can guide the recovery of lost structures. Existing event-based reconstruction methods neither explicitly model low-light noise and saturation nor enforce precise alignment between events and frames, which limits cross-modal fusion and deblurring quality. We propose the Wiener-Deconvolution-Driven Event-Based Deblurring Network (WiED-Net), which embeds the Wiener deconvolution into a deep architecture so that the physical imaging model and noise statistics are encoded in the frequency domain and high-frequency recovery is stabilized on noise dominated night data. WiED-Net adopts a two stage design. The first stage applies Wiener deconvolution in both image and feature spaces to suppress noise, recover saturated regions and reduce ringing, assisted by an eventguided cross-modal feature fusion (ECFF) module for accurate alignment. The second stage uses a multi-scale fusion module to integrate the complementary event and image branches. Training is constrained by a set of losses, including a tailored blur kernel loss that provides closed-loop regularization from physical priors. Together, these designs enable WiED-Net to recover fine details while robustly suppressing artifacts and noise, and to achieve superior quantitative and qualitative performance, achieving superior quantitative and qualitative performance with a notable improvement of 1.97 dB in PSNR and 5% in SSIM over the previous state-of-the-art methods in low-light deblurring. Code will be available at https://github.com/zhuzifeng38/WiED-Net.
Zeyu Xiao 0002, Jianlong Jin, Feng Xue 0002, Yu Liu 0023, Zhao Zhang 0001, Wei Jia 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 ACD: Adversarial Counterfactual Distillation for Rating Prediction in Recommendation
abstract
Rating prediction is a classic task in recommendation systems, aiming to accurately estimate user ratings for various items. Historical ratings typically exhibit a non-uniform distribution, leading recommendation models to favor predicting high-frequency ratings. We refer to the inconsistency between predicted ratings and users' true preferences caused by non-uniform rating distributions asrating bias. To mitigate this bias, existing studies capture various interaction behavior patterns and employ knowledge distillation techniques to improve the network's ability to model user preferences. However, due to the model being trained on datasets with non-uniform rating distributions, the rating bias may propagate through the knowledge distillation process across different behaviors, thereby contaminating the modeling of users' true preferences. To this end, we propose a novelAdversarial Counterfactual Distillation(ACD) framework for the rating prediction task, aimed at eliminating rating bias. Specifically, we design aCounterfactual Distillation Modulefrom a causal reasoning perspective to facilitate knowledge transfer across various interaction behaviors while concurrently mitigating bias contamination. Furthermore, we introduce anAdversarial Debiasing Moduleto dynamically adjust the debiasing strength, ensuring that the model maintains an optimal balance between effective knowledge transfer and bias mitigation. Extensive experiments demonstrate the superior performance of our proposed ACD framework. The complete code is publicly available athttps://github.com/hfutmars/ACD.
Sheng Sang, Feng Xue 0002, Shuaiyang Li 0001, Kang Liu 0024, Richang Hong
IEEE Trans. Knowl. Data Eng.2
2026 APMVS: Learning Multi-View Stereo Based on Adjacent Stage and Pair-Wise Stage Uncertainty Estimation
abstract
Many multi-view stereo (MVS) networks with a cascaded structure can effectively estimate depth while saving memory. However, the accuracy of the depth map in the fine stage depends on the depth map estimated in the coarse stage. Additionally, the multi-stage depth maps generated by the cascaded structure are used to compute losses but are not reused, resulting in a loss of inter-stage differentiation information. To address these issues, we propose a dual-uncertainty estimation MVS method that learns an MVS network based on adjacent stage and pair-wise stage uncertainty estimation, named APMVS. The core of the proposed APMVS is to employ dual-uncertainty estimation to mitigate the adverse effects of the cascaded structure. Specifically, it involves two estimation modules: adjacent stage uncertainty (ASU) and pair-wise stage uncertainty (PSU). The ASU estimation module dynamically adjusts the depth-hypothesis range by leveraging uncertainty from the previous stage, thereby improving the accuracy of depth-map prediction in the current stage. The PSU estimation module estimates the uncertainty between each pair of stages. Thus, regions with high uncertainty have minimal impact. We evaluate the proposed APMVS on the DTU, Tanks and Temples, and BlendedMVS datasets. Experimental results show that our method achieves superior reconstruction quality compared with other state-of-the-art methods.
Mingwei Cao, Siqi Nian, Haifeng Zhao 0001, Feng Xue 0002, Zhihan Lyu
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Eye See What You See: Query-Oriented Gaze Following
abstract
Gaze following is a fundamental challenge in computer vision that enables systems to infer where humans are looking in complex scenes, with applications spanning human-computer interaction, social robotics, and behavioral analysis. Traditional approaches to this task often rely on multi-stream architectures or transformer-based methods with inherent limitations. Current transformer-based methods face three major challenges: randomly initialized queries lack prior information about human positions; the conventional query-object paradigm necessitates numerous queries, resulting in expensive bipartite matching; and traditional global attention mechanism demands substantial resources, forcing trade-offs that sacrifice crucial facial detail features. We address these challenges by introducing a more efficient framework with three corresponding innovations: a position-aware query initialization that encodes head coordinates through sinusoidal positional encoding, incorporating prior knowledge of human head positions; a task-specific query assignment strategy that redefines the query-object relationship, remarkably reducing the number of required queries; and an enhanced sparse attention mechanism that is introduced to accelerate model convergence, enabling processing of higher resolution images without significantly increasing demands. Our approach achieves state-of-the-art performance among single-modality methods, reducing the two distance metrics by 20.4% and 28.1%, respectively, while cutting the number of learnable parameters by 71.9%. Additionally, our model outperforms most multi-modality approaches in terms of AUC metrics and distance metrics with faster convergence.
Feng Xue 0002, Dan Guo 0001
IJCNN5
2025 WPELip: enhance lip reading with word-prior information
Feng Xue 0002, Yu Li 0053, Shujie Li 0002
Multim. Syst.1
2025 REDGCN: Rating-Oriented Explicit Disentangling Graph Convolution Network for Review-Aware Recommendation
abstract
Rating prediction is a challenging task in review-aware recommendation. Although current methods effectively combine collaborative signals with review data, they fail to differentiate user preferences across various ratings and overlook the independence between these ratings. In this article, we emphasize the importance of independence modeling among representations for different rating levels. To this end, we propose a rating-oriented explicit disentangling graph convolution network for review-aware recommendation, short for REDGCN. Specifically, we introduce a rating-oriented disentangled representation learning that segments representations and rating graph based on ratings. It also employs an explicit graph learning approach to ensure the independence of disentangled representations during information propagation, which mitigates noise from review features. Furthermore, we define and model one kind of cross-rating correlation, based on the characteristics of user rating behavior. By leveraging this approach, we introduce a cross-rating constraint as an additional task to further enhance the independence among disentangled representations and improve the stability of model training. We conduct extensive experiments on six public datasets to prove the effectiveness of REDGCN. The complete data and codes of REDGCN are available athttps://github.com/hfutmars/REDGCN.
Sheng Sang, Feng Xue 0002, Kang Liu 0024, Shuaiyang Li 0001, Richang Hong
IEEE Trans. Comput. Soc. Syst.2
2025 Gloss-driven Conditional Diffusion Models for Sign Language Production
abstract
Sign Language Production (SLP) aims to convert text or audio sentences into sign language videos corresponding to their semantics, which is challenging due to the diversity and complexity of sign languages, and cross-modal semantic mapping issues. In this work, we propose a Gloss-driven Conditional Diffusion Model (GCDM) for SLP. The core of the GCDM is a diffusion model architecture, in which the sign gloss sequence is encoded by a Transformer-based encoder and input into the diffusion model as a semantic prior condition. In the process of sign pose generation, the textual semantic priors carried in the encoded gloss features are integrated into the embedded Gaussian noise via cross-attention. Subsequently, the model converts the fused features into sign language pose sequences through T-round denoising steps. During the training process, the model uses the ground-truth labels of sign poses as the starting point, generates Gaussian noise through T rounds of noise, and then performs T rounds of denoising to approximate the real sign language gestures. The entire process is constrained by the MAE loss function to ensure that the generated sign language gestures are as close as possible to the real labels. In the inference phase, the model directly randomly samples a set of Gaussian noise, generates multiple sign language gesture sequence hypotheses under the guidance of the gloss sequence, and outputs a high-confidence sign language gesture video by averaging multiple hypotheses. Experimental results on the Phoenix2014T dataset show that the proposed GCDM method achieves competitiveness in both quantitative performance and qualitative visualization.
Shengeng Tang, Feng Xue 0002, Jingjing Wu 0001, Shuo Wang 0008, Richang Hong
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Generalizing sentence-level lipreading to unseen speakers: a two-stream end-to-end approach
Yu Li 0053, Feng Xue 0002, Lin Wu 0001, Yincen Xie, Shujie Li 0002
Multim. Syst.2
2024 Multimodal Hierarchical Graph Collaborative Filtering for Multimedia-Based Recommendation
abstract
Multimedia-based recommendation (MMRec) is a challenging task, which goes beyond the collaborative filtering (CF) schema that only captures collaborative signals from interactions and explores multimodal user preference cues hidden in complex multimedia content. Despite the significant progress of current solutions for MMRec, we argue that they are limited by multimodal noise contamination. Specifically, a considerable amount of preference-irrelevant multimodal noise (e.g., the background, layout, and brightness in the image of the product) is incorporated into the representation learning of items, which contaminates the modeling of multimodal user preferences. Moreover, most of the latest researches are based on graph convolution networks (GCNs), which means that multimodal noise contamination is further amplified because noisy information is continuously propagated over the user–item interaction graph as recursive neighbor aggregations are performed. To address this problem, instead of the common MMRec paradigm which learns user preferences in an integrated manner, we propose a hierarchical framework to separately learn collaborative signals and multimodal preferences cues, thus preventing multimodal noise from flowing into collaborative signals. Then, to alleviate the noise contamination for multimodal user preference modeling, we propose to extract semantic entities from multimodal content that are more relevant to user interests, which can model semantic-level multimodal preferences and thus remove a large fraction of noise. Furthermore, we use the full multimodal features to model content-level multimodal preferences like the existing MMRec solutions, which ensures the sufficient utilization of multimodal information. Overall, we develop a novel model, multimodal hierarchical graph CF (MHGCF), which consists of three types of GCN modules tailored to capture collaborative signals, semantic-level preferences, and content-level preferences, respectively. We conduct extensive experiments to demonstrate the effectiveness of MHGCF and its components. The complete data and codes of MHGCF are available athttps://github.com/hfutmars/MHGCF.
Kang Liu 0024, Feng Xue 0002, Shuaiyang Li 0001, Sheng Sang, Richang Hong
IEEE Trans. Comput. Soc. Syst.2
2024 Multimodal Graph Causal Embedding for Multimedia-Based Recommendation
abstract
Multimedia-based recommendation (MMRec) models typically rely on observed user-item interactions and the multimodal content of items, such as visual images and textual descriptions, to predict user preferences. Among these, the user's preference for the displayed multimodal content of items is crucial for interacting with a particular item. We argue that users' preference behaviors (i.e., user-item interactions) for the modality content of items, beyond stemming from their real interest in the modality content, may also be influenced by their conformity to the popularity of items' modality-specific content (e.g., a user might be motivated to interact with a lipstick due to enthusiastic discussions among other users regarding textual reviews of the product). In essence, user-item interactions are jointly triggered by real interest and conformity. However, most existing MMRec models primarily concentrate on modeling users' interest preferences when capturing multimodal user preferences, neglecting the modeling of their conformity preferences, which results in sub-optimal recommendation performance. In this work, we resort to causal theory to propose a novel MMRec model, termed Multimodal Graph Causal Embedding (MGCE), revealing insights into the crucial causal relations of users' modality-specific interest and conformity in interaction behaviors within MMRec scenarios. Inspired by the colliding effect in causal inference and integrating the characteristics of real interest and conformity, we devise multimodal causal embedding learning networks to facilitate the learning of high-quality causal embeddings (multimodal interest and multimodal conformity embeddings) from both the structure-level and feature-level, yielding state-of-the-art performance. Extensive experimental results on three datasets demonstrate the effectiveness of MGCE.
Shuaiyang Li 0001, Feng Xue 0002, Kang Liu 0024, Dan Guo 0001, Richang Hong
IEEE Trans. Knowl. Data Eng.2
2023 Multimodal Counterfactual Learning Network for Multimedia-based Recommendation
abstract
Multimedia-based recommendation (MMRec) utilizes multimodal content (images, textual descriptions, etc.) as auxiliary information on historical interactions to determine user preferences. Most MMRec approaches predict user interests by exploiting a large amount of multimodal contents of user-interacted items, ignoring the potential effect of multimodal content of user-uninteracted items. As a matter of fact, there is a small portion of user preference-irrelevant features in the multimodal content of user-interacted items, which may be a kind of spurious correlation with user preferences, thereby degrading the recommendation performance. In this work, we argue that the multimodal content of user-uninteracted items can be further exploited to identify and eliminate the user preference-irrelevant portion inside user-interacted multimodal content, for example by counterfactual inference of causal theory. Going beyond multimodal user preference modeling only using interacted items, we propose a novel model called Multimodal Counterfactual Learning Network (MCLN), in which user-uninteracted items' multimodal content is additionally exploited to further purify the representation of user preference-relevant multimodal content that better matches the user's interests, yielding state-of-the-art performance. Extensive experiments are conducted to validate the effectiveness and rationality of MCLN. We release the complete codes of MCLN at https://github.com/hfutmars/MCLN.
Shuaiyang Li 0001, Dan Guo 0001, Kang Liu 0024, Richang Hong, Feng Xue 0002
SIGIR5
2023 Joint Multi-Grained Popularity-Aware Graph Convolution Collaborative Filtering for Recommendation
abstract
Graph convolution networks (GCNs), with their efficient ability to capture high-order connectivity in graphs, have been widely applied in recommender systems. Stacking multiple neighbor aggregation is the major operation in GCNs. It implicitly captures popularity features because the number of neighbor nodes reflects the popularity of a node. However, existing GCN-based methods ignore a universal problem: users’ sensitivity to item popularity is differentiated, but the neighbor aggregations in GCNs actually fix this sensitivity through graph Laplacian normalization, leading to suboptimal personalization. In this work, we propose to model multigrained popularity features and jointly learn them together with high-order connectivity to match the differentiation of user preferences exhibited in popularity features. Specifically, we develop a Joint Multigrained Popularity-aware Graph Convolution Collaborative Filtering model, short for JMP-GCF, which uses a popularity-aware embedding generation to construct multigrained popularity features and uses the idea of joint learning to capture the signals within and between different granularities of popularity features that are relevant for modeling user preferences. In addition, we propose a multistage stacked training strategy to speed up model convergence. We conduct extensive experiments on three public datasets to show the state-of-the-art performance of JMP-GCF. The complete codes of JMP-GCF are released athttps://github.com/hfutmars/JMP-GCF.
Kang Liu 0024, Feng Xue 0002, Xiangnan He 0001, Dan Guo 0001, Richang Hong
IEEE Trans. Comput. Soc. Syst.2
2023 LipFormer: Learning to Lipread Unseen Speakers Based on Visual-Landmark Transformers
abstract
Lipreading refers to understanding and further translating the speech of a video speaker into textual outputs. State-of-the-art lipreading methods excel in interpreting overlap speakers, i.e., speakers appear in both training and inference. However, generalizing those methods to unseen speakers incurs catastrophic performance degradation due to the limited number of speakers in training bank as well as the dominant visual variations caused by the shape/color of lips presented by different speakers. Therefore, merely depending on the visible changes of lips tends to overfit the model. To improve to generalise, in this paper we propose to use multi-modal features, i.e., visual and landmark, to describe the lip motion while being irrespective to speaker characteristics. The proposed sentence-level framework, dubbed LipFormer, is based on visual-landmark transformer architecture wherein a lip motion stream, a facial landmark stream, and a cross-modal fusion are interconnected. More specifically, the two-stream embeddings produced by self-attention are prompted into a cross-attention module to achieve the alignment across visual and landmark variations. The resulting fused features are decoded into linguistic texts by a cascaded sequence-to-sequence translation. Extensive experiments demonstrate that our method can generalise well to unseen speakers in multiple datasets.
Feng Xue 0002, Yu Li 0053, Deyin Liu, Yincen Xie, Lin Wu 0001, Richang Hong
IEEE Trans. Circuits Syst. Video Technol.1
2023 Multimodal Graph Contrastive Learning for Multimedia-Based Recommendation
abstract
Multimedia-based recommendation is a challenging task that requires not only learning collaborative signals from user-item interaction, but also capturing modality-specific user interest clues from complex multimedia content. Though significant progress on this challenge has been made, we argue that current solutions remain limited by multimodal noise contamination. Specifically, a considerable proportion of multimedia content is irrelevant to the user preference, such as the background, overall layout, and brightness of images; the word order and semantic-free words in titles;etc. We take this irrelevant information as noise contamination to discover user preferences. Moreover, most recent research has been conducted by graph learning. This means that noise is diffused into the user and item representations with the message propagation; the contamination influence is further amplified. To tackle this problem, we develop a novel framework named Multimodal Graph Contrastive Learning (MGCL), which captures collaborative signals from interactions and uses visual and textual modalities to respectively extract modality-specific user preference clues. The key idea of MGCL involves two aspects: First, to alleviate noise contamination during graph learning, we construct three parallel graph convolution networks to independently generate three types of user and item representations, containing collaborative signals, visual preference clues, and textual preference clues. Second, to eliminate as much preference-independent noisy information as possible from the generated representations, we incorporate sufficient self-supervised signals into the model optimization with the help of contrastive learning, thus enhancing the expressiveness of the user and item representations. Note that MGCL is not limited to graph learning schema, but also can be applied to most matrix factorization methods. We conduct extensive experiments on three public datasets to validate the effectiveness and scalability of MGCL11We release the codes of MGCL athttps://github.com/hfutmars/MGCL..
Kang Liu 0024, Feng Xue 0002, Dan Guo 0001, Peijie Sun, Shengsheng Qian, Richang Hong
IEEE Trans. Multim.2
2023 MEGCF: Multimodal Entity Graph Collaborative Filtering for Personalized Recommendation
abstract
In most E-commerce platforms, whether the displayed items trigger the user’s interest largely depends on their most eye-catching multimodal content. Consequently, increasing efforts focus on modeling multimodal user preference, and the pressing paradigm is to incorporate complete multimodal deep features of the items into the recommendation module. However, the existing studies ignore the mismatch problem between multimodal feature extraction (MFE) and user interest modeling (UIM) . That is, MFE and UIM have different emphases. Specifically, MFE is migrated from and adapted to upstream tasks such as image classification. In addition, it is mainly a content-oriented and non-personalized process, while UIM, with its greater focus on understanding user interaction, is essentially a user-oriented and personalized process. Therefore, the direct incorporation of MFE into UIM for purely user-oriented tasks, tends to introduce a large number of preference-independent multimodal noise and contaminate the embedding representations in UIM. This paper aims at solving the mismatch problem between MFE and UIM, so as to generate high-quality embedding representations and better model multimodal user preferences. Towards this end, we develop a novel model, m ultimodal e ntity g raph c ollaborative f iltering, short for MEGCF. The UIM of the proposed model captures the semantic correlation between interactions and the features obtained from MFE, thus making a better match between MFE and UIM. More precisely, semantic-rich entities are first extracted from the multimodal data, since they are more relevant to user preferences than other multimodal information. These entities are then integrated into the user-item interaction graph. Afterwards, a symmetric linear Graph Convolution Network (GCN) module is constructed to perform message propagation over the graph, in order to capture both high-order semantic correlation and collaborative filtering signals. Finally, the sentiment information from the review data are used to fine-grainedly weight neighbor aggregation in the GCN, as it reflects the overall quality of the items, and therefore it is an important modality information related to user preferences. Extensive experiments demonstrate the effectiveness and rationality of MEGCF. 1
Kang Liu 0024, Feng Xue 0002, Dan Guo 0001, Le Wu 0001, Shujie Li 0002, Richang Hong
ACM Trans. Inf. Syst.2
2023 LCSNet: End-to-end Lipreading with Channel-aware Feature Selection
abstract
Lipreading is a task of decoding the movement of the speaker’s lip region into text. In recent years, lipreading methods based on deep neural network have attracted widespread attention, and the accuracy has far surpassed that of experienced human lipreaders. The visual differences in some phonemes are extremely subtle and pose a great challenge to lipreading. Most of the lipreading existing methods do not process the extracted visual features, which mainly suffer from two problems. First, the extracted features contain lot of useless information such as noise caused by differences in speech speed and lip shape, for example. In addition, the extracted features are not abstract enough to distinguish phonemes with similar pronunciation. These problems have a bad effect on the performance of lipreading. To extract features from the lip regions that are more distinguishable and more relevant to the speech content, this article proposes an end-to-end deep neural network-based lipreading model (LCSNet). The proposed model extracts the short-term spatio-temporal features and the motion trajectory features from the lip region in the video clips. The extracted features are filtered by the channel attention module to eliminate the useless features and then used as input to the proposed Selective Feature Fusion Module (SFFM) to extract the high-level abstract features. Afterwards, these features are used as input to the bidirectional GRU network in time order for temporal modeling to obtain the long-term spatio-temporal features. Finally, a Connectionist Temporal Classification (CTC) decoder is used to generate the output text. The experimental results show that the proposed model achieves a 1.0% CER and 2.3% WER on the GRID corpus database, which, respectively, represents an improvement of 52% and 47% compared to LipNet.
Feng Xue 0002, Kang Liu 0024, Zikun Hong, Mingwei Cao, Dan Guo 0001, Richang Hong
ACM Trans. Multim. Comput. Commun. Appl.1
2022 RGCF: Refined graph convolution collaborative filtering with concise and expressive embedding
abstract
Graph Convolution Networks (GCNs) have attracted significant attention and have become the most popular method for learning graph representations. In recent years, many efforts have focused on integrating GCNs into recommender tasks and have made remarkable progress. At its core is to explicitly capture the high-order connectivities between nodes in the user-item bipartite graph. However, we found some potential drawbacks existed in the traditional GCN-based recommendation models are that the excessive information redundancy yield by the nonlinear graph convolution operation reduces the expressiveness of the resultant embeddings, and the important popularity features that are effective in sparse recommendation scenarios are not encoded in the embedding generation process. In this work, we develop a novel GCN-based recommendation model, named Refined Graph convolution Collaborative Filtering (RGCF), where a refined graph convolution structure is designed to match non-semantic ID inputs. In addition, a new fine-tuned symmetric normalization is proposed to mine node popularity characteristics and further incorporate the popularity features into the embedding learning process. Extensive experiments were conducted on three public million-size datasets, and the RGCF improved by an average of 13.45% over the state-of-the-art baseline. Further comparative experiments validated the effectiveness and rationality of each part of our proposed RGCF. We released our code at https://github.com/hfutmars/RGCF.
Kang Liu 0024, Feng Xue 0002, Richang Hong
Intell. Data Anal.2
2021 Graph Attention-Based Deep Neural Network for 3D Point Cloud Processing
abstract
Due to the increasing popularity of 3D sensors, it has become easier and easier to obtain point cloud data. In many fields such as autonomous driving, how to fully extract the features of point cloud to better understand and perceive 3D scenes requires further research. Therefore, this paper proposes a novel end-to-end deep learning network for the features of disorder and irregularity of 3D point cloud data. Our network uses an encoder-decoder network architecture, and a three-layer structure is used in both stages. Each encoder layer consists of graph attention convolution and graph attention pooling. Graph attention convolution reflects the spatial distribution relationship in the neighborhood area. And graph attention pooling merges the spatial distribution information in the neighborhood into the feature of the sampling point. The experimental results show that our method achieves the best results in shape classification and has competitive performance in other tasks.
Feng Xue 0002, Xiaohui Yuan 0001, Qiang Lu 0002
ICME2
2020 Cross-modal retrieval via label category supervised matrix factorization hashing
Feng Xue 0002
Pattern Recognit. Lett.1
2020 Knowledge-Based Topic Model for Multi-Modal Social Event Analysis
abstract
With the accumulation of data on the Internet and progress in representation learning techniques, knowledge priors learned from a large-scale knowledge base has been increasingly used in probabilistic topic models. However, it is challenging to learn interpretable topics and a discriminative event representation based on multi-modal information. To address these issues, we propose a knowledge priors- and max-margin-based topic model for multi-modal social event analysis, called the KGE-MMSLDA, in which feature representation and knowledge priors are jointly learned. Our model has three main advantages over current methods: (1) It integrates additional knowledge from external knowledge base into a unified topic model in which the max-margin classifier, and multi-modal information are exploited to increase the number of event descriptions obtained. (2) We mined knowledge priors from over 74,000 web documents. Multi-modal data with these knowledge priors are then incorporated into the topic model to increase the number of coherent topics learned. (3) A large-scale multi-modal dataset (containing 10 events, where each event contained approximately 7,000 Flickr pages) was collected and has been released publicly for event topic mining and classification research. In comparative experiments, the proposed method outperformed state-of-the-art models on topic coherence, and obtained a classification accuracy of 85.1%.
Feng Xue 0002, Richang Hong, Xiangnan He 0001, Shengsheng Qian, Changsheng Xu
IEEE Trans. Multim.1
2019 Social multi-modal event analysis via knowledge-based weighted topic model
Feng Xue 0002, Xueliang Liu, Tianpeng Liu, Qiang Lu 0002
J. Vis. Commun. Image Represent.1
2019 Multi-modal max-margin supervised topic model for social event analysis
Feng Xue 0002, Shengsheng Qian, Tianzhu Zhang 0001, Xueliang Liu, Changsheng Xu
Multim. Tools Appl.1
2019 Deep Item-based Collaborative Filtering for Top-N Recommendation
abstract
Item-based Collaborative Filtering (ICF) has been widely adopted in recommender systems in industry, owing to its strength in user interest modeling and ease in online personalization. By constructing a user’s profile with the items that the user has consumed, ICF recommends items that are similar to the user’s profile. With the prevalence of machine learning in recent years, significant processes have been made for ICF by learning item similarity (or representation) from data. Nevertheless, we argue that most existing works have only considered linear and shallow relationships between items, which are insufficient to capture the complicated decision-making process of users. In this article, we propose a more expressive ICF solution by accounting for the nonlinear and higher-order relationships among items. Going beyond modeling only the second-order interaction (e.g., similarity) between two items, we additionally consider the interaction among all interacted item pairs by using nonlinear neural networks. By doing this, we can effectively model the higher-order relationship among items, capturing more complicated effects in user decision-making. For example, it can differentiate which historical itemsets in a user’s profile are more important in affecting the user to make a purchase decision on an item. We treat this solution as a deep variant of ICF, thus term it as DeepICF. To justify our proposal, we perform empirical studies on two public datasets from MovieLens and Pinterest. Extensive experiments verify the highly positive effect of higher-order item interaction modeling with nonlinear neural networks. Moreover, we demonstrate that by more fine-grained second-order interaction modeling with attention network, the performance of our DeepICF method can be further improved.
Feng Xue 0002, Xiangnan He 0001, Xiang Wang 0010, Jiandong Xu, Richang Hong
ACM Trans. Inf. Syst.1
2017 Erratum to "Local line directional pattern for palmprint recognition" [Pattern Recognit. 50(2016) 26-44]
Yue-Tong Luo, Lan-Ying Zhao, Bob Zhang 0001, Wei Jia 0001, Feng Xue 0002, Yihai Zhu, Bing-Qing Xu
Pattern Recognit.5
2016 Design of digital camouflage by recursive overlapping of pattern templates
Feng Xue 0002, Yue-Tong Luo, Wei Jia 0001
Neurocomputing1
2016 Camouflage performance analysis and evaluation framework based on features fusion
Feng Xue 0002, Chengxi Yong, Yue-Tong Luo, Wei Jia 0001
Multim. Tools Appl.1
2016 Local line directional pattern for palmprint recognition
Yue-Tong Luo, Lan-Ying Zhao, Bob Zhang 0001, Wei Jia 0001, Feng Xue 0002, Yihai Zhu, Bing-Qing Xu
Pattern Recognit.5
2015 Towards efficient support relation extraction from RGBD images
Feng Xue 0002, Meng Wang 0001, Richang Hong
Inf. Sci.1
2015 Camouflage texture evaluation using a saliency map
Feng Xue 0002, Cui Guoying, Richang Hong
Multim. Syst.1
2015 An Intensity-Texture model based level set method for image segmentation
Hai Min, Wei Jia 0001, Yang Zhao 0002, Rong-Xiang Hu, Yue-Tong Luo, Feng Xue 0002
Pattern Recognit.7
2014 Image quality assessment based on matching pursuit
Richang Hong, Jianxin Pan, Shijie Hao, Meng Wang 0001, Feng Xue 0002, Xindong Wu 0001
Inf. Sci.5
2007 Real-Time Texture Synthesis Using s-Tile Set
Feng Xue 0002, You-Sheng Zhang, Julang Jiang, Xindong Wu 0001, Ronggui Wang
J. Comput. Sci. Technol.1