Jiang Li 0004

dblp:41/3068-4 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0002-0116-5662ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 HAFUNet: A Hierarchical Attention Fusion Network for Monocular Depth Estimation Integrating Event and Frame Data
abstract
In robotics and autonomous driving, accurate depth estimation is vital yet challenging under dynamic scenes and extreme lighting. Conventional frame-based cameras offer rich context but suffer from motion blur and limited dynamic range, while event cameras provide high temporal resolution and dynamic range but lack global scene structure. Therefore, recent studies explore frame-event fusion depth estimation methods to leverage these two complementary modalities to achieve robust performance. However, due to the mismatch in temporal and spatial resolution, there is an inherent contradiction between high spatial resolution frames captured at sparse temporal intervals and event streams characterized by spatial sparsity but high temporal resolution, rendering cross-modal feature fusion ineffective. Moreover, the limited availability of frame-event depth datasets further undermines the model's generalization capability across different scenes. To address the above challenges, we propose HAFUNet, a Hierarchical Attention Fusion Network for depth estimation via frame-event fusion. Our method contains: (1) a pre-trained Dual-Stream Encoder (DSEer) to extract complementary features from frame and event inputs; (2) a Cross-modal Feature Interaction Module (CFIM) that aligns and fuses spatial-channel features across modalities; and (3) a Hierarchical Attention Decoder (HADer) that progressively refines depth predictions via attention-guided convolution. Experiments on synthetic and real-world datasets show that HAFUNet surpasses existing methods in depth accuracy and robustness. These results demonstrate the strength of our fusion strategy in diverse environments. Code is available at https://github.com/SiYZhangwh/HAFUNet.
Xiaoping Wang 0001, Jiang Li 0004, Weibin Feng, Xin Zhan, Hongzhi Huang
ACM Multimedia3
2025 Tracing Intricate Cues in Dialogue: Joint Graph Structure and Sentiment Dynamics for Multimodal Emotion Recognition
abstract
Multimodal emotion recognition in conversation (MERC) has garnered substantial research attention recently. Existing MERC methods face several challenges: (1) they fail to fully harness direct inter-modal cues, possibly leading to less-than-thorough cross-modal modeling; (2) they concurrently extract information from the same and different modalities at each network layer, potentially triggering conflicts from the fusion of multi-source data; (3) they lack the agility required to detect dynamic sentimental changes, perhaps resulting in inaccurate classification of utterances with abrupt sentiment shifts. To address these issues, a novel approach named GraphSmile is proposed for tracking intricate emotional cues in multimodal dialogues. GraphSmile comprises two key components, i.e., GSF and SDP modules. GSF ingeniously leverages graph structures to alternately assimilate inter-modal and intra-modal emotional dependencies layer by layer, adequately capturing cross-modal cues while effectively circumventing fusion conflicts. SDP is an auxiliary task to explicitly delineate the sentiment dynamics between utterances, promoting the model's ability to distinguish sentimental discrepancies. GraphSmile is effortlessly applied to multimodal sentiment analysis in conversation (MSAC), thus enabling simultaneous execution of MERC and MSAC tasks. Empirical results on multiple benchmarks demonstrate that GraphSmile can handle complex emotional and sentimental patterns, significantly outperforming baseline models.
Jiang Li 0004, Xiaoping Wang 0001, Zhigang Zeng
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Producing Considerate Responses: Progressive Staged Training for Emotional Support Conversation
abstract
Emotional support conversation (ESC) aims to alleviate the negative emotions of help-seekers by providing psychological assistance. Existing approaches typically overlook the abundant annotations contained in the ESC dataset, such as the situation descriptions and feedback scores of seekers, which limits their performance. In an effort to utilize the annotation information to enhance the emotional support ability of the backbone, we propose a three-stage training method called BlenderBot-ThTra for ESC systems. The proposed BlenderBot-ThTra involves the following three training processes: fine-tuning with supplemental feedback utterance, fine-tuning with auxiliary situation restoration, and calibration with the helpfulness estimation. The first stage aims to intensify the backbone's perception of conversational context, the second stage propels the backbone into excavating the causes of the emotional distress faced by the seeker. In the third stage, we leverage a Bayesian method based on the seeker's feedback scores to train a helpfulness evaluation model, then exploit a contrastive learning method to calibrate the ESC backbone. We conduct experiments on the standard multiturn ESC dataset, and the results demonstrate that BlenderBot-ThTra has a significant advantage in generating more supportive and adaptive responses.
Guoqing Lv, Jiang Li 0004, Xiaoping Wang 0001, Xin Zhan, Zhigang Zeng
IEEE Trans. Comput. Soc. Syst.2
2024 EmotionIC: emotional inertia and contagion-driven dependency modeling for emotion recognition in conversation
Jiang Li 0004, Zhigang Zeng
Sci. China Inf. Sci.2
2024 A dual-stream recurrence-attention network with global-local awareness for emotion recognition in textual dialog
Jiang Li 0004, Xiaoping Wang 0001, Zhigang Zeng
Eng. Appl. Artif. Intell.1
2024 ERNetCL: A novel emotion recognition network in textual conversation based on curriculum learning strategy
Jiang Li 0004, Xiaoping Wang 0001, Zhigang Zeng
Knowl. Based Syst.1
2024 GA2MIF: Graph and Attention Based Two-Stage Multi-Source Information Fusion for Conversational Emotion Detection
abstract
Multimodal Emotion Recognition in Conversation (ERC) plays an influential role in the field of human-computer interaction and conversational robotics since it can motivate machines to provide empathetic services. Multimodal data modeling is an up-and-coming research area in recent years, which is inspired by human capability to integrate multiple senses. Several graph-based approaches claim to capture interactive information between modalities, but the heterogeneity of multimodal data makes these methods prohibit optimal solutions. In this work, we introduce a multimodal fusion approach named Graph and Attention based Two-stage Multi-source Information Fusion (GA2MIF) for emotion detection in conversation. Our proposed method circumvents the problem of taking heterogeneous graph as input to the model while eliminating complex redundant connections in the construction of graph. GA2MIF focuses on contextual modeling and cross-modal modeling through leveraging Multi-head Directed Graph ATtention networks (MDGATs) and Multi-head Pairwise Cross-modal ATtention networks (MPCATs), respectively. Extensive experiments on two public datasets (i.e., IEMOCAP and MELD) demonstrate that the proposed GA2MIF has the capacity to validly capture intra-modal long-range contextual information and inter-modal complementary information, as well as outperforms the prevalent State-Of-The-Art (SOTA) models by a remarkable margin.
Jiang Li 0004, Xiaoping Wang 0001, Guoqing Lv, Zhigang Zeng
IEEE Trans. Affect. Comput.1
2024 CFN-ESA: A Cross-Modal Fusion Network With Emotion-Shift Awareness for Dialogue Emotion Recognition
abstract
Multimodal emotion recognition in conversation (ERC) has garnered growing attention from research communities in various fields. In this paper, we propose a Crossmodal Fusion Network with Emotion-Shift Awareness (CFNESA) for ERC. Extant approaches employ each modality equally without distinguishing the amount of emotional information in these modalities, rendering it hard to adequately extract complementary information from multimodal data. To cope with this problem, in CFN-ESA, we treat textual modality as the primary source of emotional information, while visual and acoustic modalities are taken as the secondary sources. Besides, most multimodal ERC models ignore emotion-shift information and overfocus on contextual information, leading to the failure of emotion recognition under emotion-shift scenario. We elaborate an emotion-shift module to address this challenge. CFNESA mainly consists of unimodal encoder (RUME), cross-modal encoder (ACME), and emotion-shift module (LESM). RUME is applied to extract conversation-level contextual emotional cues while pulling together data distributions between modalities; ACME is utilized to perform multimodal interaction centered on textual modality; LESM is used to model emotion shift and capture emotion-shift information, thereby guiding the learning of the main task. Experimental results demonstrate that CFN-ESA can effectively promote performance for ERC and remarkably outperform state-of-the-art models.
Jiang Li 0004, Xiaoping Wang 0001, Zhigang Zeng
IEEE Trans. Affect. Comput.1
2024 GraphCFC: A Directed Graph Based Cross-Modal Feature Complementation Approach for Multimodal Conversational Emotion Recognition
abstract
Emotion Recognition in Conversation (ERC) plays a significant part in Human-Computer Interaction (HCI) systems since it can provide empathetic services. Multimodal ERC can mitigate the drawbacks of uni-modal approaches. Recently, Graph Neural Networks (GNNs) have been widely used in a variety of fields due to their superior performance in relation modeling. In multimodal ERC, GNNs are capable of extracting both long-distance contextual information and inter-modal interactive information. Unfortunately, since existing methods such as MMGCN directly fuse multiple modalities, redundant information may be generated and diverse information may be lost. In this work, we present a directed Graph based Cross-modal Feature Complementation (GraphCFC) module that can efficiently model contextual and interactive information. GraphCFC alleviates the problem of heterogeneity gap in multimodal fusion by utilizing multiple subspace extractors and Pair-wise Cross-modal Complementary (PairCC) strategy. We extract various types of edges from the constructed graph for encoding, thus enabling GNNs to extract crucial contextual and interactive information more accurately when performing message passing. Furthermore, we design a GNN structure called GAT-MLP, which can provide a new unified network framework for multimodal learning. The experimental results on two benchmark datasets show that our GraphCFC outperforms the state-of-the-art (SOTA) approaches.
Jiang Li 0004, Xiaoping Wang 0001, Guoqing Lv, Zhigang Zeng
IEEE Trans. Multim.1
2023 InferEM: Inferring the Speaker's Intention for Empathetic Dialogue Generation
Guoqing Lv, Jiang Li 0004, Xiaoping Wang 0001, Zhigang Zeng
CogSci2
2023 Watch the Speakers: A Hybrid Continuous Attribution Network for Emotion Recognition in Conversation With Emotion Disentanglement
abstract
Emotion Recognition in Conversation (ERC) has attracted widespread attention in the natural language processing field due to its enormous potential for practical applications. Existing ERC methods face challenges in achieving generalization to diverse scenarios due to insufficient modeling of context, ambiguous capture of dialogue relationships and overfitting in speaker modeling. In this work, we present a Hybrid Continuous Attributive Network (HCAN) to address these issues in the perspective of emotional continuation and emotional attribution. Specifically, HCAN adopts a hybrid recurrent and attention-based module to model global emotion continuity. Then a novel Emotional Attribution Encoding (EAE) is proposed to model intra- and inter-emotional attribution for each utterance. Moreover, aiming to enhance the robustness of the model in speaker modeling and improve its performance in different scenarios, A comprehensive loss function emotional cognitive loss $\mathcal{L}_{EC}$ is proposed to alleviate emotional drift and overcome the overfitting of the model to speaker modeling. Our model achieves state-of-the-art performance on three datasets, demonstrating the superiority of our work. Another extensive comparative experiments and ablation studies on three benchmarks are conducted to provided evidence to support the efficacy of each module. Further exploration of generalization ability experiments shows the plug-and-play nature of the EAE module in our method.
Shanglin Lei, Xiaoping Wang 0001, Guanting Dong 0001, Jiang Li 0004
ICTAI4
2023 GraphMFT: A graph network based multimodal fusion technique for emotion recognition in conversation
Jiang Li 0004, Xiaoping Wang 0001, Guoqing Lv, Zhigang Zeng
Neurocomputing1
2021 A Novel Text Classification Approach based on Meta-path Similarities and Graph Neural Networks
abstract
With the rise of neural networks, studies on text classification have transitioned from traditional methods to deep learning, especially to graph neural networks on text graphs constructed from corpora.In this paper, we model the complex instances and rich interactions in text classification as a heterogeneous graph.Nevertheless, due to the overlook of indirect relations between documents, graph neural networks have not been fully exploited for the heterogeneous text graph with different types of nodes and links.Consequently, we propose a Meta-Path-based Text Graph Neural Network (MPTGNN) for text classification.Specifically, we first construct a heterogeneous text graph from corpora; we then transform the text graph into several homogeneous weighted graphs via some pre-defined metapaths; we also propose a Two-stage Multi-graph Information Fusion method (TMIF) for document representation.Empirical results on multiple benchmark datasets have proved that our proposed method outperforms state-of-the-art graph-based methods like Text GCN.
Jiang Li 0004, Qing Zhou 0002
SEKE2
2020 Discovering the Lonely Among the Students with Weighted Graph Neural Networks
abstract
Nowadays, college students are prone to feel lonely, thus loneliness identification is an essential task. Existing methods like questionnaire are mainly used to identify loners through loneliness scales. These methods are subjective and often ineffective because loners may avoid reporting their real conditions. In this paper, we propose a new method based on pair-wise course collaborative relationships for identifying lonely students in an end-to-end fashion. Due to the overlook of relations among instances, traditional machine learning methods are insufficient to distill collaborative information from course records. In order to make full use of the interaction among students in course, we use weighted Graph Neural Networks (GNNs) to model our problem. The number of times that two students have collaborated is extracted as edge weight to build a weighted graph. Meanwhile, we propose a Binary tree-based Graph Oversampling Algorithm (BGOA) to tackle class-imbalanced problem. The experimental results have proved that the proposed approach can achieve above 88% of performance.
Qing Zhou 0002, Jiang Li 0004, Yinchun Tang
ICTAI2
2020 Identifying Loners from Their Project Collaboration Records - A Graph-Based Approach
Qing Zhou 0002, Jiang Li 0004, Yinchun Tang
KSEM (1)2