Yuntao Shou

dblp:325/3357 · DBLP profile ↗
← Back
30ranked-venue papers
14as first author
30since 2021 · last 2026
0000-0003-3270-6238ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 10 first-author · 21 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning
abstract
Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs. However, real-world data often suffer from missing modalities due to transmission errors or user behavior, severely degrading model performance. Existing methods enhance robustness via cross-modal consistency learning but largely ignore modality complementarity, leading to biased reconstructions. To address this limitation, we propose C²MOE, a novel Consistency and Complementarity-guided Mixture of Experts framework for incomplete multimodal emotion learning. Our approach unifies representation learning and missing modality imputation within a principled information-theoretic framework. Specifically, multimodal knowledge is factorized into consistency and complementarity components via interaction-aware experts. Consistency is captured by maximizing cross-modal predictability, while complementarity is preserved by maximizing conditional entropy between modalities. Building upon this decomposition, C²MOE introduces a dual-branch prediction mechanism for robust imputation under missing modalities. The consistency branch aligns imputed features with the joint distribution by minimizing uncertainty, and the complementarity branch exploits modality-unique cues via entropy maximization. Finally, C²MOE employs a learnable reweighting module that dynamically assigns importance scores to each expert’s output, yielding a robust and adaptive fusion for imputation. Extensive experiments on multiple MERC benchmarks demonstrate that C²MOE consistently surpasses state-of-the-art methods across various missing-modality settings, validating its robustness and generalization.
Yuntao Shou, Wei Ai 0001, Keqin Li 0001
ICMR1
2026 AMB-DSGDN: Adaptive modality-balanced dynamic semantic graph differential network for multimodal emotion recognition
abstract
Multimodal dialogue emotion recognition captures emotional cues by fusing text, visual, and audio modalities. However, existing approaches still suffer from notable limitations in modeling emotional dependencies and learning multimodal representations. On the one hand, they are unable to effectively filter out redundant or noisy signals within multimodal features, which hinders the accurate capture of the dynamic evolution of emotional states across and within speakers. On the other hand, during multimodal feature learning, dominant modalities (e.g., textual cues) tend to overwhelm the fusion process, thereby suppressing the complementary contributions of non-dominant modalities such as speech and vision, ultimately constraining the overall recognition performance. To address these challenges, we propose an Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network (AMB-DSGDN). Concretely, we first construct modality-specific subgraphs for text, speech, and vision, where each modality contains intra-speaker and inter-speaker graphs to capture both self-continuity and cross-speaker emotional dependencies. On top of these subgraphs, we introduce a differential graph attention mechanism, which computes the discrepancy between two sets of attention maps. By explicitly contrasting these attention distributions, the mechanism cancels out shared noise patterns while retaining modality-specific and context-relevant signals, thereby yielding purer and more discriminative emotional representations. In addition, we design an adaptive modality balancing mechanism, which estimates a dropout probability for each modality according to its relative contribution in emotion modeling. This mechanism randomly discards a portion of features from dominant modalities to suppress their overwhelming influence, while proportionally rescaling the preserved features based on the dropout probability to maintain overall information balance. Extensive experiments on IEMOCAP and MELD datasets validate that AMB-DSGDN significantly outperforms state-of-the-art baselines, demonstrating its effectiveness and robustness in multimodal conversational emotion recognition.
Yuntao Shou, Yilong Tan, Wei Ai 0001, Keqin Li 0001
Expert Syst. Appl.2
2026 Relational graph-driven differential denoising and diffusion attention fusion for multimodal conversational emotion recognition
Yuntao Shou, Wei Ai 0001, Keqin Li 0001
Neurocomputing2
2026 Efficient long-distance latent relation-aware graph neural network for multi-modal federated emotion recognition
Yuntao Shou, Wei Ai 0001, Jiayi Du
J. Parallel Distributed Comput.1
2026 Dynamic fusion-aware graph convolutional neural network for multimodal emotion recognition in conversations
Weilun Tang, Yuntao Shou, Yilong Tan, Wei Ai 0001, Keqin Li 0001
Knowl. Based Syst.3
2026 CILF-CIAE: CLIP-driven image-language fusion for correcting inverse age estimation
Yuntao Shou, Wei Ai 0001, Keqin Li 0001
Neural Networks1
2026 A Comprehensive Survey on Multi-modal Conversational Emotion Recognition with Deep Learning
abstract
Multi-modal Conversation Emotion Recognition (MCER) aims to recognize and track the speaker’s emotional state using text, speech, and visual information. Compared with traditional single-utterance multi-modal emotion recognition or single-modal conversation emotion recognition, MCER is more challenging. It requires modeling complex emotional interactions and learning consistent and complementary semantics across multiple modalities. Although many deep learning-based approaches have been proposed for MCER, there is still a lack of systematic reviews summarizing existing modeling methods. Therefore, a timely and comprehensive overview of MCER’s recent advances in deep learning is of great significance. In this survey, we provide a comprehensive overview of MCER modeling methods and roughly divide MCER methods into four categories, i.e., context-free modeling, sequential context modeling, speaker-differentiated modeling, and speaker-relationship modeling. Unlike conventional taxonomies based on modality combinations or task-stage decomposition, our framework focuses on how models structurally capture conversational dynamics, speaker roles, and emotional dependencies. In addition, we further discuss MCER’s publicly available popular datasets, multi-modal feature extraction methods, application areas, existing challenges, and future development directions. We hope this review provides valuable insights into the current state of MCER research and inspires the development of more effective models.
Yuntao Shou, Wei Ai 0001, Fangze Fu, Keqin Li 0001
ACM Trans. Inf. Syst.1
2025 Revisiting Multimodal Emotion Recognition in Conversation from the Perspective of Graph Spectrum
abstract
Efficiently capturing consistent and complementary semantic features in context is crucial for Multimodal Emotion Recognition in Conversations (MERC). However, limited by the over-smoothing or low-pass filtering characteristics of spatial graph neural networks, are insufficient to accurately capture the long-distance consistency low-frequency information and complementarity high-frequency information of the utterances. To this end, this paper revisits the task of MERC from the perspective of the graph spectrum and proposes a Graph-Spectrum-based Multimodal Consistency and Complementary collaborative learning framework GS-MCC. First, GS-MCC uses a sliding window to construct a multimodal interaction graph to model conversational relationships and designs efficient Fourier graph operators (FGO) to extract long-distance high-frequency and low-frequency information, respectively. FGO can be stacked in multiple layers, which can effectively alleviate the over-smoothing problem. Then, GS-MCC uses contrastive learning to construct self-supervised signals that reflect complementarity and consistent semantic collaboration with high and low-frequency signals, thereby improving the ability of high and low-frequency information to reflect genuine emotions. Finally, GS-MCC inputs the coordinated high and low-frequency information into the MLP network and softmax function for emotion prediction. Extensive experiments have proven the superiority of the GS-MCC architecture proposed in this paper on two benchmark data sets.
Wei Ai 0001, Fuchen Zhang, Yuntao Shou, Keqin Li 0001
AAAI3
2025 A Flow Model with Low-Rank Transformers for Incomplete Multimodal Survival Analysis
abstract
In recent years, multimodal medical data-based survival analysis has attracted much attention. However, realworld datasets often suffer from the problem of incomplete modality, where some patient modality information is missing due to acquisition limitations or system failures. Existing methods typically infer missing modalities directly from observed ones using deep neural networks, but they often ignore the distributional discrepancy across modalities, resulting in inconsistent and unreliable modality reconstruction. To address these challenges, we propose a novel framework that combines a low-rank Transformer with a flow-based generative model for robust and flexible multimodal survival prediction. Specifically, we first formulate the concerned problem as incomplete multimodal survival analysis using the multi-instance representation of whole slide images (WSIs) and genomic profiles. To realize incomplete multimodal survival analysis, we propose a classspecific flow for cross-modal distribution alignment. Under the condition of class labels, we model and transform the cross-modal distribution. By virtue of the reversible structure and accurate density modeling capabilities of the normalizing flow model, the model can effectively construct a distribution-consistent latent space of the missing modality, thereby improving the consistency between the reconstructed data and the true distribution. Finally, we design a lightweight Transformer architecture to model intra-modal dependencies while alleviating the overfitting problem in high-dimensional modality fusion by virtue of the low-rank Transformer. Extensive experiments have demonstrated that our method not only achieves state-of-the-art performance under complete modality settings, but also maintains robust and superior accuracy under the incomplete modalities scenario.
Yuntao Shou, Zao Dai, Wei Ai 0001, Keqin Li 0001
BIBM2
2025 Dynamic Graph Neural ODE Network for Multi-modal Emotion Recognition in Conversation
abstract
Multimodal emotion recognition in conversation (MERC) refers to identifying and classifying human emotional states by combining data from multiple different modalities (e.g., audio, images, text, video, etc.). Specifically, human emotional expressions are often complex and diverse, and these complex emotional expressions can be captured and understood more comprehensively through the fusion of multimodal information. Most existing graph-based multimodal emotion recognition methods can only use shallow GCNs to extract emotion features and fail to capture the temporal dependencies caused by dynamic changes in emotions. To address the above problems, we propose a Dynamic Graph Neural Ordinary Differential Equation Network (DGODE) for multimodal emotion recognition in conversation, which combines the dynamic changes of emotions to capture the temporal dependency of speakers’ emotions. Technically, the key idea of DGODE is to use the graph ODE evolution network to characterize the continuous dynamics of node representations over time and capture temporal dependencies. Extensive experiments on two publicly available multimodal emotion recognition datasets demonstrate that the proposed DGODE model has superior performance compared to various baselines. Furthermore, the proposed DGODE can also alleviate the over-smoothing problem, thereby enabling the construction of a deep GCN network.
Yuntao Shou, Wei Ai 0001, Keqin Li 0001
COLING1
2025 Graph Domain Adaptation With Dual-Branch Encoder and Two-Level Alignment for Whole Slide Image-Based Survival Prediction
Yuntao Shou, Xiangyong Cao, Peiqiang Yan, Qiaohui, Qian Zhao 0002, Deyu Meng
ICCV1
2025 GSDNet: Revisiting Incomplete Multimodality-Diffusion Emotion Recognition from the Perspective of Graph Spectrum
abstract
Multimodal Emotion Recognition (MER) combines technologies from multiple fields (e.g., computer vision, natural language processing, and audio signal processing), aiming to infer an individual's emotional state by analyzing information from different sources (i.e., video, audio, and text). Compared with single modality, by fusing complementary semantic information from different modalities, the model can obtain more robust knowledge representation. However, the modality missing problem limits the performance of MERC in practical scenarios. Recent work has achieved impressive performance on modality completion using graph neural networks and diffusion models, respectively. This inspires us to combine these two dimensions in the completion network to obtain more powerful representation capabilities. However, we argue that directly running a full-rank score-based diffusion model on the entire graph adjacency matrix space may adversely affect the learning process of the diffusion model. This is because the model assumes a direct relationship between each pair of nodes and ignores local structural features and sparse connections between nodes, thereby significantly reducing the quality of the generated data. Based on the above ideas, we propose a novel Graph Spectral Diffusion Network (GSDNet), which utilizes a low-rank score-based diffusion model to map Gaussian noise to the graph spectral distribution space of missing modalities and recover the missing data according to its original distribution. Extensive experiments have demonstrated that GSDNet achieves state-of-the-art emotion recognition performance in various modality loss scenarios.
Yuntao Shou, Wei Ai 0001, Cen Chen 0002, Keqin Li 0001
IJCAI1
2025 Revisiting Multi-modal Emotion Learning with Broad State Space Models and Probability-Guidance Fusion
Yuntao Shou, Wei Ai 0001, Keqin Li 0001
ECML/PKDD (4)1
2025 Enhancing domain adaptation for plant diseases detection through Masked Image Consistency in Multi-Granularity Alignment
Guinan Guo, Songning Lai, Qingyang Wu, Yuntao Shou, Wenxu Shi
Expert Syst. Appl.4
2025 LRA-GNN: Latent Relation-Aware Graph Neural Network with initial and Dynamic Residual for facial age estimation
Yuntao Shou, Wei Ai 0001, Keqin Li 0001
Expert Syst. Appl.2
2025 SDR-GNN: Spectral Domain Reconstruction Graph Neural Network for incomplete multimodal learning in conversational emotion recognition
Fangze Fu, Wei Ai 0001, Fan Yang 0044, Yuntao Shou, Keqin Li 0001
Knowl. Based Syst.4
2025 Contrastive Graph Representation Learning with Adversarial Cross-View Reconstruction and Information Bottleneck
Yuntao Shou, Haozhi Lan, Xiangyong Cao
Neural Networks1
2025 Masked contrastive graph representation learning for age estimation
Yuntao Shou, Xiangyong Cao, Huan Liu 0012, Deyu Meng
Pattern Recognit.1
2025 A Low-Rank Matching Attention Based Cross-Modal Feature Fusion Method for Conversational Emotion Recognition
abstract
Conversational emotion recognition (CER) is an important research topic in human-computer interactions. Although recent advancements in transformer-based cross-modal fusion methods have shown promise in CER tasks, they tend to overlook the crucial intra-modal and inter-modal emotional interaction or suffer from high computational complexity. To address this, we introduce a novel and lightweight cross-modal feature fusion method called Low-Rank Matching Attention Method (LMAM). LMAM effectively captures contextual emotional semantic information in conversations while mitigating the quadratic complexity issue caused by the self-attention mechanism. Specifically, by setting a matching weight and calculating inter-modal features attention scores row by row, LMAM requires only one-third of the parameters of self-attention methods. We also employ the low-rank decomposition method on the weights to further reduce the number of parameters in LMAM. As a result, LMAM offers a lightweight model while avoiding overfitting problems caused by a large number of parameters. Moreover, LMAM is able to fully exploit the intra-modal emotional contextual information within each modality and integrates complementary emotional semantic information across modalities by computing and fusing similarities of intra-modal and inter-modal features simultaneously. Experimental results verify the superiority of LMAM compared with other popular cross-modal fusion methods on the premise of being more lightweight. Also, LMAM can be embedded into any existing state-of-the-art CER methods in a plug-and-play manner, and can be applied to other multi-modal recognition tasks, e.g., session recommendation and humour detection, demonstrating its remarkable generalization ability.
Yuntao Shou, Huan Liu 0012, Xiangyong Cao, Deyu Meng, Bo Dong 0001
IEEE Trans. Affect. Comput.1
2025 GroupFace: Imbalanced Age Estimation Based on Multi-Hop Attention Graph Convolutional Network and Group-Aware Margin Optimization
abstract
With the recent advances in computer vision, age estimation has significantly improved in overall accuracy. However, owing to the most common methods do not take into account the class imbalance problem in age estimation datasets, they suffer from a large bias in recognizing long-tailed groups. To achieve high-quality imbalanced learning in long-tailed groups, the dominant solution lies in that the feature extractor learns the discriminative features of different groups and the classifier is able to provide appropriate and unbiased margins for different groups by the discriminative features. Therefore, in this novel, we propose an innovative collaborative learning framework (GroupFace) that integrates a multi-hop attention graph convolutional network and a dynamic group-aware margin strategy based on reinforcement learning. Specifically, to extract the discriminative features of different groups, we design an enhanced multi-hop attention graph convolutional network. This network is capable of capturing the interactions of neighboring nodes at different distances, fusing local and global information to model facial deep aging, and exploring diverse representations of different groups. In addition, to further address the class imbalance problem, we design a dynamic group-aware margin strategy based on reinforcement learning to provide appropriate and unbiased margins for different groups. The strategy divides the sample into four age groups and considers identifying the optimum margins for various age groups by employing a Markov decision process. Under the guidance of the agent, the feature representation bias and the classification margin deviation between different groups can be reduced simultaneously, balancing inter-class separability and intra-class proximity. After joint optimization, our architecture achieves excellent performance on several age estimation benchmark datasets. It not only achieves large improvements in overall estimation accuracy but also gains balanced performance in long-tailed group estimation.
Yuntao Shou, Wei Ai 0001, Keqin Li 0001
IEEE Trans. Inf. Forensics Secur.2
2025 SE-GNN: Seed Expanded-Aware Graph Neural Network With Iterative Optimization for Semi-Supervised Entity Alignment
abstract
Entity alignment aims to use pre-aligned seed pairs to find other equivalent entities from different knowledge graphs and is widely used in graph fusion-related fields. However, as the scale of knowledge graphs increases, manually annotating pre-aligned seed pairs becomes difficult. Existing research utilizes entity embeddings obtained by aggregating single structural information to identify potential seed pairs, thus reducing the reliance on pre-aligned seed pairs. However, due to the structural heterogeneity of KG, the quality of potential seed pairs obtained using only a single structural information is not ideal. In addition, although existing research improves the quality of potential seed pairs through semi-supervised iteration, they underestimate the impact of embedding distortion produced by noisy seed pairs on the alignment effect. In order to solve the above problems, we propose a seed expanded-aware graph neural network with iterative optimization for semi-supervised entity alignment, named SE-GNN. First, we utilize the semantic attributes and structural features of entities, combined with a conditional filtering mechanism, to obtain high-quality initial potential seed pairs. Next, we designed a local and global awareness mechanism. It introduces initial potential seed pairs and combines local and global information to obtain a more comprehensive entity embedding representation, which alleviates the impact of KG structural heterogeneity and lays the foundation for the optimization of initial potential seed pairs. Then, we designed the threshold nearest neighbor embedding correction strategy. It combines the similarity threshold and the bidirectional nearest neighbor method as a filtering mechanism to select iterative potential seed pairs and also uses an embedding correction strategy to eliminate the embedding distortion. Finally, we will reach the optimized potential seeds after iterative rounds to input local and global sensing mechanisms, obtain the final entity embedding, and perform entity alignment. Experimental results on public datasets demonstrate the excellent performance of our SE-GNN, showcasing the effectiveness of the model. Our code is publicly available athttps://github.com/ShuoShan1/SE-GNN.
Hongen Shao, Yuntao Shou, Wei Ai 0001, Keqin Li 0001
IEEE Trans. Knowl. Data Eng.4
2025 DER-GCN: Dialog and Event Relation-Aware Graph Convolutional Neural Network for Multimodal Dialog Emotion Recognition
abstract
With the continuous development of deep learning (DL), the task of multimodal dialog emotion recognition (MDER) has recently received extensive research attention, which is also an essential branch of DL. The MDER aims to identify the emotional information contained in different modalities, e.g., text, video, and audio, and in different dialog scenes. However, the existing research has focused on modeling contextual semantic information and dialog relations between speakers while ignoring the impact of event relations on emotion. To tackle the above issues, we propose a novel dialog and event relation-aware graph convolutional neural network (DER-GCN) for multimodal emotion recognition method. It models dialog relations between speakers and captures latent event relations information. Specifically, we construct a weighted multirelationship graph to simultaneously capture the dependencies between speakers and event relations in a dialog. Moreover, we also introduce a self-supervised masked graph autoencoder (SMGAE) to improve the fusion representation ability of features and structures. Next, we design a new multiple information Transformer (MIT) to capture the correlation between different relations, which can provide a better fuse of the multivariate information between relations. Finally, we propose a loss optimization strategy based on contrastive learning to enhance the representation learning ability of minority class features. We conduct extensive experiments on the benchmark datasets, Interactive Emotional Dyadic Motion Capture (IEMOCAP) and Multimodal EmotionLines Dataset (MELD), which verify the effectiveness of the DER-GCN model. The results demonstrate that our model significantly improves both the average accuracy and the value of emotion recognition. Our code is publicly available at https://github.com/yuntaoshou/DER-GCN.
Wei Ai 0001, Yuntao Shou, Keqin Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 SpeGCL: Self-Supervised Graph Spectrum Contrastive Learning Without Positive Samples
abstract
Graph contrastive learning (GCL) has emerged as a powerful method for dealing with noise and fluctuations in graph-structured data, and can be applied to social networks and knowledge graphs. Although various graph augmentation strategies have emerged in the field of GCL, traditional graph convolutional network (GCN) mainly tends to preserve smooth features and has difficulty capturing fine-grained changes between different views. To address the above issue, we first construct Fourier graph neural network (FourierGNN) from the perspective of graph spectrum learning, which captures different frequency components by stacking multiple Fourier graph operations (FGO) layers in Fourier space. Then, we find that the difference between the high-frequency information of two augmented graphs should be larger than the difference between the low-frequency information. Next, we theoretically prove that focusing only on pushing negative pairs farther away can more effectively achieve performance advantages. By leveraging these discoveries, we propose a novel self-supervised graph spectrum contrastive learning framework, i.e., SpeGCL, and design an effective contrastive strategy to optimize this goal. We also provide a theoretical justification for the efficacy of using only negative samples in SpeGCL. Extensive experiments have been conducted on unsupervised, transfer, and semi-supervised learning tasks to show that SpeGCL outperforms existing state-of-the-art (SOTA) GCL methods.
Yuntao Shou, Xiangyong Cao, Deyu Meng
IEEE Trans. Neural Networks Learn. Syst.1
2024 Edge-enhanced minimum-margin graph attention network for short text classification
Wei Ai 0001, Hongen Shao, Yuntao Shou, Keqin Li 0001
Expert Syst. Appl.4
2024 A multi-message passing framework based on heterogeneous graphs in conversational emotion recognition
Yuntao Shou, Wei Ai 0001, Jiayi Du, Keqin Li 0001
Neurocomputing2
2024 A multi-view mask contrastive learning graph convolutional neural network for age estimation
Yuntao Shou, Wei Ai 0001, Keqin Li 0001
Knowl. Inf. Syst.2
2024 Masked Graph Learning With Recurrent Alignment for Multimodal Emotion Recognition in Conversation
abstract
Since Multimodal Emotion Recognition in Conversation (MERC) can be applied to public opinion monitoring, intelligent dialogue robots, and other fields, it has received extensive research attention in recent years. Unlike traditional unimodal emotion recognition, MERC can fuse complementary semantic information between multiple modalities (e.g., text, audio, and vision) to improve emotion recognition. However, previous work ignored the inter-modal alignment process and the intra-modal noise information before multimodal fusion but directly fuses multimodal features, which will hinder the model for representation learning. In this study, we have developed a novel approach called Masked Graph Learning with Recursive Alignment (MGLRA) to tackle this problem, which uses a recurrent iterative module with memory to align multimodal features, and then uses the masked GCN for multimodal feature fusion. First, we employ LSTM to capture contextual information and use a graph attention-filtering mechanism to eliminate noise effectively within the modality. Second, we build a recurrent iteration module with a memory function, which can use communication between different modalities to eliminate the gap between modalities and achieve the preliminary alignment of features between modalities. Then, a cross-modal multi-head attention mechanism is introduced to achieve feature alignment between modalities and construct a masked GCN for multimodal feature fusion, which can perform random mask reconstruction on the nodes in the graph to obtain better node feature representation. Finally, we utilize a multilayer perceptron (MLP) for emotion recognition. Extensive experiments on two benchmark datasets (i.e., IEMOCAP and MELD) demonstrate that MGLRA outperforms state-of-the-art methods.
Fuchen Zhang, Yuntao Shou, Hongen Shao, Wei Ai 0001, Keqin Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 A Two-Stage Multimodal Emotion Recognition Model Based on Graph Contrastive Learning
abstract
In terms of human-computer interaction, it is becoming more and more important to correctly understand the user’s emotional state in a conversation, so the task of multimodal emotion recognition (MER) started to receive more attention. However, existing emotion classification methods usually perform classification only once. Sentences are likely to be misclassified in a single round of classification. Previous work usually ignores the similarities and differences between different morphological features in the fusion process. To address the above issues, we propose a two-stage emotion recognition model based on graph contrastive learning (TS-GCL). First, we encode the original dataset with different preprocessing modalities. Second, a graph contrastive learning (GCL) strategy is introduced for these three modal data with other structures to learn similarities and differences within and between modalities. Finally, we use MLP twice to achieve the final emotion classification. This staged classification method can help the model to better focus on different levels of emotional information, thereby improving the performance of the model. Extensive experiments show that TS-GCL has superior performance on IEMOCAP and MELD datasets compared with previous methods.
Wei Ai 0001, Fuchen Zhang, Yuntao Shou, Hongen Shao, Keqin Li 0001
ICPADS4
2023 GraphUnet: Graph Make Strong Encoders for Remote Sensing Segmentation
abstract
Remote sensing segmentation are widely applied in environmental protection, and urban change detection, etc. Despite the success of deep learning-based remote sensing segmentation methods (e.g., CNN and Transformer), they are not flexible enough to model irregular objects. To tackle the above problems, this paper treats images as graph structures and introduces a simple vision GNN (GraphUNet) architecture for remote sensing segmentation. Specifically, we construct a graph view and utilize GNN to extract the visual feature. To extract the global contextual location information in the image, we introduce a multi-head attention mechanism for global information extraction. Furthermore, we introduce a feed-forward network for each node to perform feature transformation on node features to encourage information diversity. Extensive experiments on three real datasets demonstrate that our model GraphUNet outperforms SOTA remote sensing image segmentation methods.
Yuntao Shou, Wei Ai 0001, Fuchen Zhang, Keqin Li 0001
ICPADS1
2022 Conversational emotion recognition studies based on graph convolutional neural networks and a dependent syntactic analysis
Yuntao Shou, Wei Ai 0001, Keqin Li 0001
Neurocomputing1