Xingye Li

dblp:286/5971 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards robust sentiment analysis with multimodal interaction graph and hybrid contrastive learning
Peizhu Gong, Jin Liu 0009, Xiliang Zhang, Xingye Li, Lai Wei 0001, Huihua He
Pattern Recognit.4
2026 Implicit Knowledge and Emotional Cues-Enhanced Multimodal Sarcasm Detection Model
abstract
Multimodal sarcasm detection aims to determine whether sarcasm exists in the interplay between text and accompanying images or audio. Effectively exploiting commonsense knowledge is crucial in multimodal sarcasm detection. However, previous methods only utilize the inherent attributes within modalities to construct superficial inconsistencies. They fail to fully leverage commonsense knowledge and implicit sarcasm cues and cannot uncover the deep contradictions between modalities. Accordingly, we propose a novel model named the Knowledge and Emotion Enhanced (KEE) model, which enhances sarcasm detection by integrating commonsense reasoning and cross-modal emotional contrast. Specifically, the KEE model leverages a generative language model to infer commonsense knowledge from the visual modality. It then selects highly relevant knowledge phrases through an attention mechanism and integrates them with text embeddings via a gated weighting mechanism to reveal sarcastic clues behind literal semantics. Additionally, the model captures emotional polarity inconsistencies across modalities and optimizes detection through a joint distribution framework. Experimental results demonstrate that the KEE model outperforms state-of-the-art methods on benchmark datasets.
Jin Liu 0009, Xingye Li, Huihua He
IEEE Trans. Affect. Comput.3
2026 Text-Centric Sparse Interaction Fusion Network With a Modality Calibrating Module for Multimodal Sentiment Analysis
abstract
Despite significant progress in Multimodal Sentiment Analysis (MSA), particularly with methods based on Transformers and Attentions, three major issues remain unresolved. First, existing unimodal representation learning methods overlook the noise and redundancy present within unimodal sequences. Second, current methods ignore the uneven distribution of emotional information across different modalities and fail to consider the highest contribution of text. Third, they neglect the issue of inter-modal noise redundancy in modal interactions. To address the aforementioned issues, this paper proposes a Text-Centric Sparse Interaction Fusion Network with a Modality Calibrating Module (TC-SIFN-MCM). A Modality Calibrating Module (MCM) is introduced during the unimodal representation learning phase to suppress noise within modalities. A Text-Centric Sparse Interaction Fusion Network (TC-SIFN) is designed to guide the learning of resource-poor non-text modalities with the resource-rich text modality, fully leveraging the highest contribution of the text modality. A Top-K Sparse Attention Transformer is utilized as a means of modal inter action, suppressing inter-modal noise generated during modal interactions. We conducted extensive experiments on the CMU MOSI, CMU-MOSEI, and SIMS datasets, and the results show that our TC-SIFN-MCM model outperforms current methods, with a notably impressive Acc-2 score of 86.07% on the MOSEI dataset.
Hongkun Zhou, Jin Liu 0009, Xingye Li, Huihua He
IEEE Trans. Affect. Comput.3
2025 Vessel re-identification by a hierarchical perceptual aggregation network with inclination-aware attention
abstract
Abstract Vessel re-identification (re-ID) is a crucial task in maritime supervision, enhancing maritime safety and improving the maritime situational awareness system. However, distinct from land-based scenarios involving vehicles or pedestrians, vessels, as enormous rigid bodies situated in the dynamic marine environment, face unique challenges such as significant variations in the scale of discriminative features and unpredictable sway. Furthermore, there is a limited number of publicly available datasets for vessel re-ID in complex backgrounds. In this paper, to overcome these challenges, a novel Hierarchical Perceptual Aggregation Network with Inclination-Aware Attention (HPAN-IAA) is proposed. HPAN-IAA comprises two main modules: the Hierarchical Perceptual Aggregation Block (HPAB) and the Inclination-Aware Attention Block (IAAB). Specifically, in HPAB, a hierarchical perceptual function is introduced to decompose visual information of vessels into discriminative features at multiple levels. These feature maps with different levels of detail from diverse network layers are then fused together by concatenation, resulting in a comprehensive feature representation that effectively integrates information across various scales. Conversely, to address the irregular variations and random omissions in discriminative feature distribution caused by unpredictable vessel sway, in IAAB, the Channel Collaborative Attention Module and the Pyramidal Spatial Attention Module are designed to adaptively extract potential discriminative features within each channel and spatial dimension, enhancing model’s ability in effectively extracting and utilizing irregularly changing discriminative features. Moreover, we propose a novel vessel re-ID dataset—VesselReID-2258. Extensive experiments conducted on VesselReID-2258 and the publicly available dataset VesselReID demonstrate that HPAN-IAA outperforms the current state-of-the-art methods,achieving superior performance with mean Average Precision scores of 0.861 and 0.823.
Yuetian Cao, Jin Liu 0009, Zijun Yu, Xingye Li, Lai Wei 0001, Zhongdai Wu
Comput. J.4
2025 DNMCN: Dual-Stage Normalization Based Modality-Collaborative Fusion Network for Multimodal Sentiment Analysis
abstract
Due to the high-quality semantic information provided by the text modality, text-driven models have become the dominant approach for Multimodal Sentiment Analysis (MSA) in recent years. Despite notable progress in previous studies, two primary limitations remain: (i) aligning multimodal features often relies on simple matching of sequence length or feature dimension, which overlooks cross-modal heterogeneity. (ii) existing fusion techniques tend to over-rely on text, potentially diminishing the emotional data contributed by other modalities. To address these issues, in this paper, we propose a Dual-stage Normalization based Modality-Collaborative Fusion Network (DNMCN). Initially, to reduce modality discrepancies, we introduce a dual-stage normalization strategy, where features from different modalities were mapped into a common dimensional space in the first stage to facilitate effective cross-modal comparisons; sequence length inconsistencies caused by cropping and multiscale dimension reduction were addressed in the second stage. Additionally, to achieve high-quality cross-modal mapping without losing non-textual modality information, we propose an Adaptive modality-Collaborative Fusion Transformer (ACF-T) block. Specifically, in ACF-T block, textual semantics are first integrated into any non-text modality via multi-head attention. Next, a novel adaptive weighting strategy is introduced to balance the contribution of fused features and other non-textual modality features, thereby enhancing crossmodal interaction. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches on the public benchmark datasets CH-SIMS, CMU-MOSI and CMU-MOSEI.
Jin Liu 0009, Xingye Li, Jiajia Jiao, Huihua He
IEEE Trans. Affect. Comput.3
2025 AGT-Net: It Takes Two to Tango in Long-Term Person Reidentification
abstract
The key to address long-term person reidentification (Re-ID) in videos is to extract invariant spatio-temporal features (ISTF), which can be broadly categorized into two forms: 1) clothes-irrelevant appearance such as facial characteristics and body shape; 2) identity-distinctive motion such as posture and gait. However, existing studies mainly focus on mining either appearance- or motion-based features in the sequences without sufficient utilization of the ISTF. In this article, we propose an Appearance and gait-based tango network (AGT-Net) to comprehensively mine ISTF from appearance details and gait motions. Specifically, on the appearance detail branch, we introduce an invariance-aware video Swin transformer (IA-VST) to extract clothes-irrelevant appearance from the original RGB sequences. On the gait motion branch, we propose a motion-sensitive gait feature extractor (MS-GFE) to learn identity-distinctive motion from the pre-processed gait sequences. In addition, a score-level fusion strategy is introduced to integrate information from the two streams for prediction. Besides, since there is a lack of publicly available datasets, we propose a style-transferring synthetic long-term Video Re-ID (STYLE-VID) dataset, particularly for long-term Re-ID. Extensive experiments demonstrate that AGT-Net not only outperforms the state-of-the-art methods by up to 1.5% mAP on STYLE-VID, but also achieves comparable performance to other models on the traditional short-term Re-ID dataset MARS.
Zijun Yu, Jin Liu 0009, Peizhu Gong, Xingye Li, Lai Wei 0001, Huihua He, Zhongdai Wu
IEEE Trans. Comput. Soc. Syst.4
2025 MST-ARGCN: modality-squeeze transformer with attentional recurrent graph capsule network for multimodal sentiment analysis
Xingye Li, Meijing Li, Huihua He
J. Supercomput.3
2025 Dual-constraint capsule network for visible-infrared cross-modality pedestrian re-identification
Zhengjie Xi, Xingye Li
J. Supercomput.3
2024 IDPonzi: An interpretable detection model for identifying smart Ponzi schemes
Xia Feng, Qichen Shi, Xingye Li, Liangmin Wang 0001
Eng. Appl. Artif. Intell.3
2024 MAGDRA: A Multi-modal Attention Graph Network with Dynamic Routing-By-Agreement for multi-label emotion recognition
Xingye Li, Jin Liu 0009, Yurong Xie, Peizhu Gong, Xiliang Zhang, Huihua He
Knowl. Based Syst.1
2024 MEDMCN: a novel multi-modal EfficientDet with multi-scale CapsNet for object detection
Xingye Li, Jin Liu 0009, Zhengyu Tang, Bing Han 0009, Zhongdai Wu
J. Supercomput.1
2023 A Multi-Stage Hierarchical Relational Graph Neural Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis targets at accurately perceiving the emotional states by incorporating related information from multiple sources. However, existing methods mostly neglect the unbalanced contributions and inherent relational interactions across distinct modalities. In this paper, we propose a multi-stage hierarchical relational graph neural network (MHRG), catering to intra- and inter-modal dynamics learning with modality calibration. In the first stage, modality-specific graph convolution modules are introduced to learn the intra-modal sequential semantics. In the second, we design a modality-adaptive modification module to determine the contribution of each modality based on the prediction confidence. Finally, diverse inter-modal dynamics are considered respectively by a novel hierarchical relational graph fusion method for further aggregation according to the type of interactions. Extensive experiments on benchmark datasets demonstrate that MHRG outperforms the existing methods and achieves the state-of-the-art performance.
Peizhu Gong, Xiliang Zhang, Xingye Li
ICASSP4
2022 Circulant-interactive Transformer with Dimension-aware Fusion for Multimodal Sentiment Analysis
Peizhu Gong, Jin Liu 0009, Xiliang Zhang, Xingye Li, Zijun Yu
ACML4
2022 SnapshotNet: Self-supervised feature learning for point cloud data segmentation using minimal labeled data
Xingye Li, Zhigang Zhu 0001
Comput. Vis. Image Underst.1
2021 Bioinformatic analysis of SMN1-ACE/ACE2 interactions hinted at a potential protective effect of spinal muscular atrophy against COVID-19-induced lung injury
abstract
Patients with spinal muscular atrophy (SMA) are susceptible to the respiratory infections and might be at a heightened risk of poor clinical outcomes upon contracting coronavirus disease 2019 (COVID-19). In the face of the COVID-19 pandemic, the potential associations of SMA with the susceptibility to and prognostication of COVID-19 need to be clarified. We documented an SMA case who contracted COVID-19 but only developed mild-to-moderate clinical and radiological manifestations of pneumonia, which were relieved by a combined antiviral and supportive treatment. We then reviewed a cohort of patients with SMA who had been living in the Hubei province since November 2019, among which the only 1 out of 56 was diagnosed with COVID-19 (1.79%, 1/56). Bioinformatic analysis was carried out to delineate the potential genetic crosstalk between SMN1 (mutation of which leads to SMA) and COVID-19/lung injury-associated pathways. Protein-protein interaction analysis by STRING suggested that loss-of-function of SMN1 might modulate COVID-19 pathogenesis through CFTR, CXCL8, TNF and ACE. Expression quantitative trait loci analysis also revealed a link between SMN1 and ACE2, despite low-confidence protein-protein interactions as suggested by STRING. This bioinformatic analysis could give hint on why SMA might not necessarily lead to poor outcomes in patients with COVID-19.
Xingye Li, Jianxiong Shen, Haining Tan, Tianhua Rong, Youxi Lin, Erwei Feng, Zhengguang Chen, Lin Zhang 0015, Matthew Tak Vai Chan, William Ka Kei Wu
Briefings Bioinform.2