Shujie Li 0002

dblp:78/6393-2 · also Shu-Jie Li 0002 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-6525-2706ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Bidirectional Counterfactual Distillation for Review-Based Recommendation
abstract
Review-based recommendation methods typically integrate multiple behaviors, including interactions, reviews, and ratings, to model user preferences. To effectively extract preference signals from diverse behaviors, some studies train multiple student models to capture distinct behavioral patterns, and leverage online distillation to facilitate collaborative learning among them. However, we argue that these techniques suffer from bias contamination from rating distributions and feature homogenization during cross-behavior knowledge transfer: (1) Rating distribution bias, arising from non-uniform historical ratings, propagates across behaviors through distillation, contaminating the true preference representations of other behaviors. (2) Static distillation strategies often lead to homogenized behavioral features, hindering the learning of behavior-specific preferences. To address these issues, we propose a novel Bidirectional Counterfactual Distillation (BiCoD) framework for review-based recommendation. In BiCoD, we first design an adversarial counterfactual distillation module to suppress the impact of non-uniform rating distributions on distillation, thereby preventing it from contaminating the user's true preference representations across behaviors. Subsequently, we introduce a stage-aware bidirectional distillation strategy to enhance the distinctiveness of behavioral features, facilitating the effective learning of behavior-specific preferences. Extensive experiments on five real-world datasets validate the effectiveness and superiority of the proposed framework.
Sheng Sang, Shujie Li 0002, Shuaiyang Li 0001, Kang Liu 0024, Wei Jia 0001, Dan Guo 0001, Feng Xue 0002
AAAI2
2026 LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition
abstract
Visual Speech Recognition (VSR), commonly known as lipreading, enables the recognition of spoken text by analyzing lip visual features. Due to the subtlety of lip movements, its recognition is much harder than other motion recognition tasks. Existing VSR models face the challenge of viseme ambiguity when processing phonemes with similar pronunciations—multiple phonemes share similar viseme features, leading to a notable drop in lipreading accuracy. To address this issue, this study proposes a Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition(LinProVSR) framework. First, an ambiguous sample set is constructed based on linguistic knowledge to provide supervisory signals for the model's training. Then, a Progressive Contrastive Disambiguation Network (PCDN) is designed, which progressively enhances the model's ability to capture the subtle viseme differences corresponding to similar phonemes through viseme-phoneme contrastive disambiguation in the encoding stage and text contrastive disambiguation in the decoding stage. Furthermore, we pioneer the Ambiguous Word Error Rate (AWER) metric specifically for evaluating recognition of phonetically ambiguous text, and verify the effectiveness of the proposed method on multiple public datasets, achieving a significant breakthrough especially in distinguishing visually similar phonemes.
Feng Xue 0002, Baochao Zhu, Wei Jia 0001, Shujie Li 0002, Yu Li 0053, Shengeng Tang, Dan Guo 0001
AAAI4
2026 Cross-modal feature disentangling via bidirectional distillation for multimodal recommendation
Shuaiyang Li 0001, Kang Liu 0024, Shujie Li 0002, Dan Guo 0001, Feng Xue 0002
Expert Syst. Appl.3
2026 D2M2Lip: dual-domain feature fusion and motion magnification for lip reading
Baochao Zhu, Shujie Li 0002, Feng Xue 0002
Multim. Syst.2
2026 CFLip: Generalizing Lipreading to Unseen Speakers by Learning Common Features
abstract
Lipreading refers to translating the lip movements observed in a video of a speaker into corresponding textual outputs, providing a visual alternative to auditory communication for individuals who are deaf or hard of hearing. Existing lipreading methods typically independently learn the lip movements of each speaker. This results in the model being highly sensitive to the individual visual features (lip color/shape) of the speakers in the training set, hindering the generalization of lipreading models. Despite the obvious visual variations in the lips of different speakers, we claim that there are still inherent common features when they pronounce the same phoneme. We attempt to learn the common pronunciation features across different speakers, so as to achieve better generalization of lipreading model to unseen speakers. In this article, we propose a sentence-level lipreading framework based on Learning Common Features (CFLip), designed to extract common pronunciation features fromvideo pairs. Specifically, we first employ data augmentation strategy to generate pseudo videos that share labels but with different speakers by replacing frame segments in real videos. With thesevideo pairs, we designed a dual-stream network to learn commonality feature by minimized the distance between the features of different speakers pronouncing the same words via Generalization Loss. Extensive experiments on benchmark datasets demonstrate that the proposed CFLip can effectively generalize to unseen speakers.
Yu Li 0053, Feng Xue 0002, Dan Guo 0001, Shengeng Tang, Shujie Li 0002, Richang Hong
IEEE Trans. Comput. Soc. Syst.6
2025 WPELip: enhance lip reading with word-prior information
Feng Xue 0002, Yu Li 0053, Shujie Li 0002
Multim. Syst.4
2024 Generalizing sentence-level lipreading to unseen speakers: a two-stream end-to-end approach
Yu Li 0053, Feng Xue 0002, Lin Wu 0001, Yincen Xie, Shujie Li 0002
Multim. Syst.5
2023 MEGCF: Multimodal Entity Graph Collaborative Filtering for Personalized Recommendation
abstract
In most E-commerce platforms, whether the displayed items trigger the user’s interest largely depends on their most eye-catching multimodal content. Consequently, increasing efforts focus on modeling multimodal user preference, and the pressing paradigm is to incorporate complete multimodal deep features of the items into the recommendation module. However, the existing studies ignore the mismatch problem between multimodal feature extraction (MFE) and user interest modeling (UIM) . That is, MFE and UIM have different emphases. Specifically, MFE is migrated from and adapted to upstream tasks such as image classification. In addition, it is mainly a content-oriented and non-personalized process, while UIM, with its greater focus on understanding user interaction, is essentially a user-oriented and personalized process. Therefore, the direct incorporation of MFE into UIM for purely user-oriented tasks, tends to introduce a large number of preference-independent multimodal noise and contaminate the embedding representations in UIM. This paper aims at solving the mismatch problem between MFE and UIM, so as to generate high-quality embedding representations and better model multimodal user preferences. Towards this end, we develop a novel model, m ultimodal e ntity g raph c ollaborative f iltering, short for MEGCF. The UIM of the proposed model captures the semantic correlation between interactions and the features obtained from MFE, thus making a better match between MFE and UIM. More precisely, semantic-rich entities are first extracted from the multimodal data, since they are more relevant to user preferences than other multimodal information. These entities are then integrated into the user-item interaction graph. Afterwards, a symmetric linear Graph Convolution Network (GCN) module is constructed to perform message propagation over the graph, in order to capture both high-order semantic correlation and collaborative filtering signals. Finally, the sentiment information from the review data are used to fine-grainedly weight neighbor aggregation in the GCN, as it reflects the overall quality of the items, and therefore it is an important modality information related to user preferences. Extensive experiments demonstrate the effectiveness and rationality of MEGCF. 1
Kang Liu 0024, Feng Xue 0002, Dan Guo 0001, Le Wu 0001, Shujie Li 0002, Richang Hong
ACM Trans. Inf. Syst.5
2022 An iterative solution for improving the generalization ability of unsupervised skeleton motion retargeting
Shujie Li 0002, Wei Jia 0001, Yang Zhao 0002, Liping Zheng
Comput. Graph.1
2022 EEPNet: An efficient and effective convolutional neural network for palmprint recognition
Wei Jia 0001, Yang Zhao 0002, Shujie Li 0002, Hai Min
Pattern Recognit. Lett.4
2021 Real-time automatic helmet detection of motorcyclists in urban traffic using improved YOLOv5 detector
abstract
Abstract In traffic accidents, motorcycle accidents are the main cause of casualties, especially in developing countries. The main cause of fatal injuries in motorcycle accidents is that motorcycle riders or passengers do not wear helmets. In this paper, an automatic helmet detection of motorcyclists method based on deep learning is presented. The method consists of two steps. The first step uses the improved YOLOv5 detector to detect motorcycles (including motorcyclists) from video surveillance. The second step takes the motorcycles detected in the previous step as input and continues to use the improved YOLOv5 detector to detect whether the motorcyclists wear helmets. The improvement of the YOLOv5 detector includes the fusion of triplet attention and the use of soft‐NMS instead of NMS. A new motorcycle helmet dataset (HFUT‐MH) is being proposed, which is larger and more comprehensive than the existing dataset derived from multiple traffic monitoring in Chinese cities. Finally, the proposed method is verified by experiments and compared with other state‐of‐the‐art methods. Our method achieves mAP of 97.7%, F1‐score of 92.7% and frames per second (FPS) of 63, which outperforms other state‐of‐the‐art detection methods.
Wei Jia 0001, Shiquan Xu, Yang Zhao 0002, Hai Min, Shujie Li 0002
IET Image Process.6
2020 Blind Quality Assessment for Cartoon Images
abstract
Current blind image quality assessment (BIQA) algorithms are mainly designed for natural images. Unfortunately, cartoon and cartoon-like images are quite different from natural images. Hence, recent BIQA methods are not very robust to cartoon images. In this paper, we propose a specific BIQA algorithm designed for cartoon images, which consists of the following terms. First, a cartoon image is divided into edge areas and nonedge areas via a Tchebichef moment (TM)-based process. Second, a multiorder sharpness statistic term is used to measure the quality of the edges, and a sharpness statistic prior model of high-quality (HQ) cartoon images is built. Finally, a local encoding statistic term is adopted to describe the textural complexity in the nonedge areas, and a texture statistic prior model is also established. The experimental results on the cartoon image datasets demonstrate that the proposed method can accurately evaluate the visual quality of cartoon images and is more suitable for cartoon scenarios than some traditional BIQA algorithms.
Yuan Chen 0012, Yang Zhao 0002, Shujie Li 0002, Wangmeng Zuo, Wei Jia 0001, Xiaoping Liu 0003
IEEE Trans. Circuits Syst. Video Technol.3
2019 Bidirectional recurrent autoencoder for 3D skeleton motion data refinement
Shujie Li 0002, Haisheng Zhu, Wenjun Xie, Yang Zhao 0002, Xiaoping Liu 0003
Comput. Graph.1