EDBT 2026 Demo / reviewers in the wild / expert
Wei Ai 0001
dblp:40/9555-1
· DBLP profile ↗
38ranked-venue papers
11as first author
36since 2021 · last 2026
0000-0001-9047-7977ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 7 first-author · 20 since 2021Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion LearningabstractRecent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs. However, real-world data often suffer from missing modalities due to transmission errors or user behavior, severely degrading model performance. Existing methods enhance robustness via cross-modal consistency learning but largely ignore modality complementarity, leading to biased reconstructions. To address this limitation, we propose C²MOE, a novel Consistency and Complementarity-guided Mixture of Experts framework for incomplete multimodal emotion learning. Our approach unifies representation learning and missing modality imputation within a principled information-theoretic framework. Specifically, multimodal knowledge is factorized into consistency and complementarity components via interaction-aware experts. Consistency is captured by maximizing cross-modal predictability, while complementarity is preserved by maximizing conditional entropy between modalities. Building upon this decomposition, C²MOE introduces a dual-branch prediction mechanism for robust imputation under missing modalities. The consistency branch aligns imputed features with the joint distribution by minimizing uncertainty, and the complementarity branch exploits modality-unique cues via entropy maximization. Finally, C²MOE employs a learnable reweighting module that dynamically assigns importance scores to each expert’s output, yielding a robust and adaptive fusion for imputation. Extensive experiments on multiple MERC benchmarks demonstrate that C²MOE consistently surpasses state-of-the-art methods across various missing-modality settings, validating its robustness and generalization. Yuntao Shou, Wei Ai 0001, Keqin Li 0001 |
ICMR | 3 |
| 2026 | DRL-based privacy-preserving video streaming task offloading under energy constraints in MEC
Wei Ai 0001, Zhixiong He, Keqin Li 0001 |
Comput. Networks | 5 |
| 2026 | AMB-DSGDN: Adaptive modality-balanced dynamic semantic graph differential network for multimodal emotion recognitionabstractMultimodal dialogue emotion recognition captures emotional cues by fusing text, visual, and audio modalities. However, existing approaches still suffer from notable limitations in modeling emotional dependencies and learning multimodal representations. On the one hand, they are unable to effectively filter out redundant or noisy signals within multimodal features, which hinders the accurate capture of the dynamic evolution of emotional states across and within speakers. On the other hand, during multimodal feature learning, dominant modalities (e.g., textual cues) tend to overwhelm the fusion process, thereby suppressing the complementary contributions of non-dominant modalities such as speech and vision, ultimately constraining the overall recognition performance. To address these challenges, we propose an Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network (AMB-DSGDN). Concretely, we first construct modality-specific subgraphs for text, speech, and vision, where each modality contains intra-speaker and inter-speaker graphs to capture both self-continuity and cross-speaker emotional dependencies. On top of these subgraphs, we introduce a differential graph attention mechanism, which computes the discrepancy between two sets of attention maps. By explicitly contrasting these attention distributions, the mechanism cancels out shared noise patterns while retaining modality-specific and context-relevant signals, thereby yielding purer and more discriminative emotional representations. In addition, we design an adaptive modality balancing mechanism, which estimates a dropout probability for each modality according to its relative contribution in emotion modeling. This mechanism randomly discards a portion of features from dominant modalities to suppress their overwhelming influence, while proportionally rescaling the preserved features based on the dropout probability to maintain overall information balance. Extensive experiments on IEMOCAP and MELD datasets validate that AMB-DSGDN significantly outperforms state-of-the-art baselines, demonstrating its effectiveness and robustness in multimodal conversational emotion recognition. Yuntao Shou, Yilong Tan, Wei Ai 0001, Keqin Li 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Relational graph-driven differential denoising and diffusion attention fusion for multimodal conversational emotion recognition
Yuntao Shou, Wei Ai 0001, Keqin Li 0001 |
Neurocomputing | 3 |
| 2026 | Efficient long-distance latent relation-aware graph neural network for multi-modal federated emotion recognition
Yuntao Shou, Wei Ai 0001, Jiayi Du |
J. Parallel Distributed Comput. | 2 |
| 2026 | Dynamic fusion-aware graph convolutional neural network for multimodal emotion recognition in conversations
Weilun Tang, Yuntao Shou, Yilong Tan, Wei Ai 0001, Keqin Li 0001 |
Knowl. Based Syst. | 6 |
| 2026 | CILF-CIAE: CLIP-driven image-language fusion for correcting inverse age estimation
Yuntao Shou, Wei Ai 0001, Keqin Li 0001 |
Neural Networks | 3 |
| 2026 | A two-stage entity event deduplication method based on graph node selection and node optimization strategy
Wei Ai 0001, Hongen Shao, Keqin Li 0001 |
Soft Comput. | 1 |
| 2026 | A Comprehensive Survey on Multi-modal Conversational Emotion Recognition with Deep LearningabstractMulti-modal Conversation Emotion Recognition (MCER) aims to recognize and track the speaker’s emotional state using text, speech, and visual information. Compared with traditional single-utterance multi-modal emotion recognition or single-modal conversation emotion recognition, MCER is more challenging. It requires modeling complex emotional interactions and learning consistent and complementary semantics across multiple modalities. Although many deep learning-based approaches have been proposed for MCER, there is still a lack of systematic reviews summarizing existing modeling methods. Therefore, a timely and comprehensive overview of MCER’s recent advances in deep learning is of great significance. In this survey, we provide a comprehensive overview of MCER modeling methods and roughly divide MCER methods into four categories, i.e., context-free modeling, sequential context modeling, speaker-differentiated modeling, and speaker-relationship modeling. Unlike conventional taxonomies based on modality combinations or task-stage decomposition, our framework focuses on how models structurally capture conversational dynamics, speaker roles, and emotional dependencies. In addition, we further discuss MCER’s publicly available popular datasets, multi-modal feature extraction methods, application areas, existing challenges, and future development directions. We hope this review provides valuable insights into the current state of MCER research and inspires the development of more effective models. Yuntao Shou, Wei Ai 0001, Fangze Fu, Keqin Li 0001 |
ACM Trans. Inf. Syst. | 3 |
| 2026 | PAHPA: Revolutionizing Kubernetes Autoscaling With Integrated Predictive Analytics and Real-Time MonitoringabstractKubernetes provides powerful container orchestration features, like auto-scaling, which dynamically adjusts the scale of containerized applications. However, the default auto scaling mechanism in Kubernetes typically responds to workload changes only after they have occurred, which can result in resource mismatches. While there has been significant research into proactive auto-scaling approaches, most existing solutions struggle to adapt to the constantly evolving characteristics of prediction targets. This paper presents PAHPA, an intelligent, prediction-assisted approach aimed at enhancing Kubernetes' autoscaling capabilities, which integrates an prediction model, SLMD-LightGBM proposed in this paper, featuring a self updating mechanism for continuous prediction optimization, queueing theory analysis to improve resource allocation decisions, and a correction mechanism that refines prediction metrics using real-time data. A central insight of this paper is that while predictions are crucial, they should not be the sole basis for decisions and must be complemented by real-time monitoring data to ensure robust and adaptive autoscaling. Experimental results demonstrate that PAHPA significantly enhances system stability and reduces service latency, achieving a 9.5% lower Violation Rate (16.3%) than HPA (25.8%) and a 63% reduction in 99th percentile latency (880ms vs. 2400ms). This research highlights how combining predictive analytics with real-time monitoring can lead to more effective autoscaling strategies in cloud-native environments. Junwei Xiao, Fan Yang 0044, Wei Ai 0001, Guoqi Xie |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | Revisiting Multimodal Emotion Recognition in Conversation from the Perspective of Graph SpectrumabstractEfficiently capturing consistent and complementary semantic features in context is crucial for Multimodal Emotion Recognition in Conversations (MERC). However, limited by the over-smoothing or low-pass filtering characteristics of spatial graph neural networks, are insufficient to accurately capture the long-distance consistency low-frequency information and complementarity high-frequency information of the utterances. To this end, this paper revisits the task of MERC from the perspective of the graph spectrum and proposes a Graph-Spectrum-based Multimodal Consistency and Complementary collaborative learning framework GS-MCC. First, GS-MCC uses a sliding window to construct a multimodal interaction graph to model conversational relationships and designs efficient Fourier graph operators (FGO) to extract long-distance high-frequency and low-frequency information, respectively. FGO can be stacked in multiple layers, which can effectively alleviate the over-smoothing problem. Then, GS-MCC uses contrastive learning to construct self-supervised signals that reflect complementarity and consistent semantic collaboration with high and low-frequency signals, thereby improving the ability of high and low-frequency information to reflect genuine emotions. Finally, GS-MCC inputs the coordinated high and low-frequency information into the MLP network and softmax function for emotion prediction. Extensive experiments have proven the superiority of the GS-MCC architecture proposed in this paper on two benchmark data sets. Wei Ai 0001, Fuchen Zhang, Yuntao Shou, Keqin Li 0001 |
AAAI | 1 |
| 2025 | A Flow Model with Low-Rank Transformers for Incomplete Multimodal Survival AnalysisabstractIn recent years, multimodal medical data-based survival analysis has attracted much attention. However, realworld datasets often suffer from the problem of incomplete modality, where some patient modality information is missing due to acquisition limitations or system failures. Existing methods typically infer missing modalities directly from observed ones using deep neural networks, but they often ignore the distributional discrepancy across modalities, resulting in inconsistent and unreliable modality reconstruction. To address these challenges, we propose a novel framework that combines a low-rank Transformer with a flow-based generative model for robust and flexible multimodal survival prediction. Specifically, we first formulate the concerned problem as incomplete multimodal survival analysis using the multi-instance representation of whole slide images (WSIs) and genomic profiles. To realize incomplete multimodal survival analysis, we propose a classspecific flow for cross-modal distribution alignment. Under the condition of class labels, we model and transform the cross-modal distribution. By virtue of the reversible structure and accurate density modeling capabilities of the normalizing flow model, the model can effectively construct a distribution-consistent latent space of the missing modality, thereby improving the consistency between the reconstructed data and the true distribution. Finally, we design a lightweight Transformer architecture to model intra-modal dependencies while alleviating the overfitting problem in high-dimensional modality fusion by virtue of the low-rank Transformer. Extensive experiments have demonstrated that our method not only achieves state-of-the-art performance under complete modality settings, but also maintains robust and superior accuracy under the incomplete modalities scenario. Yuntao Shou, Zao Dai, Wei Ai 0001, Keqin Li 0001 |
BIBM | 6 |
| 2025 | Dynamic Graph Neural ODE Network for Multi-modal Emotion Recognition in ConversationabstractMultimodal emotion recognition in conversation (MERC) refers to identifying and classifying human emotional states by combining data from multiple different modalities (e.g., audio, images, text, video, etc.). Specifically, human emotional expressions are often complex and diverse, and these complex emotional expressions can be captured and understood more comprehensively through the fusion of multimodal information. Most existing graph-based multimodal emotion recognition methods can only use shallow GCNs to extract emotion features and fail to capture the temporal dependencies caused by dynamic changes in emotions. To address the above problems, we propose a Dynamic Graph Neural Ordinary Differential Equation Network (DGODE) for multimodal emotion recognition in conversation, which combines the dynamic changes of emotions to capture the temporal dependency of speakers’ emotions. Technically, the key idea of DGODE is to use the graph ODE evolution network to characterize the continuous dynamics of node representations over time and capture temporal dependencies. Extensive experiments on two publicly available multimodal emotion recognition datasets demonstrate that the proposed DGODE model has superior performance compared to various baselines. Furthermore, the proposed DGODE can also alleviate the over-smoothing problem, thereby enabling the construction of a deep GCN network. Yuntao Shou, Wei Ai 0001, Keqin Li 0001 |
COLING | 3 |
| 2025 | GSDNet: Revisiting Incomplete Multimodality-Diffusion Emotion Recognition from the Perspective of Graph SpectrumabstractMultimodal Emotion Recognition (MER) combines technologies from multiple fields (e.g., computer vision, natural language processing, and audio signal processing), aiming to infer an individual's emotional state by analyzing information from different sources (i.e., video, audio, and text). Compared with single modality, by fusing complementary semantic information from different modalities, the model can obtain more robust knowledge representation. However, the modality missing problem limits the performance of MERC in practical scenarios. Recent work has achieved impressive performance on modality completion using graph neural networks and diffusion models, respectively. This inspires us to combine these two dimensions in the completion network to obtain more powerful representation capabilities. However, we argue that directly running a full-rank score-based diffusion model on the entire graph adjacency matrix space may adversely affect the learning process of the diffusion model. This is because the model assumes a direct relationship between each pair of nodes and ignores local structural features and sparse connections between nodes, thereby significantly reducing the quality of the generated data. Based on the above ideas, we propose a novel Graph Spectral Diffusion Network (GSDNet), which utilizes a low-rank score-based diffusion model to map Gaussian noise to the graph spectral distribution space of missing modalities and recover the missing data according to its original distribution. Extensive experiments have demonstrated that GSDNet achieves state-of-the-art emotion recognition performance in various modality loss scenarios. Yuntao Shou, Wei Ai 0001, Cen Chen 0002, Keqin Li 0001 |
IJCAI | 4 |
| 2025 | Revisiting Multi-modal Emotion Learning with Broad State Space Models and Probability-Guidance Fusion
Yuntao Shou, Wei Ai 0001, Keqin Li 0001 |
ECML/PKDD (4) | 3 |
| 2025 | Contrastive multi-graph learning with neighbor hierarchical sifting for semi-supervised text classification
Wei Ai 0001, Keqin Li 0001 |
Expert Syst. Appl. | 1 |
| 2025 | LRA-GNN: Latent Relation-Aware Graph Neural Network with initial and Dynamic Residual for facial age estimation
Yuntao Shou, Wei Ai 0001, Keqin Li 0001 |
Expert Syst. Appl. | 3 |
| 2025 | MFLM-GCN: Multi-relation Fusion and Latent-relation Mining Graph Convolutional Network for entity alignment
Wei Ai 0001, Hongen Shao, Zhixiong He, Keqin Li 0001 |
Knowl. Based Syst. | 1 |
| 2025 | SDR-GNN: Spectral Domain Reconstruction Graph Neural Network for incomplete multimodal learning in conversational emotion recognition
Fangze Fu, Wei Ai 0001, Fan Yang 0044, Yuntao Shou, Keqin Li 0001 |
Knowl. Based Syst. | 2 |
| 2025 | SE-GCL: an event-based simple and effective graph contrastive learning for text representation
Wei Ai 0001, Keqin Li 0001 |
Neural Comput. Appl. | 2 |
| 2025 | GroupFace: Imbalanced Age Estimation Based on Multi-Hop Attention Graph Convolutional Network and Group-Aware Margin OptimizationabstractWith the recent advances in computer vision, age estimation has significantly improved in overall accuracy. However, owing to the most common methods do not take into account the class imbalance problem in age estimation datasets, they suffer from a large bias in recognizing long-tailed groups. To achieve high-quality imbalanced learning in long-tailed groups, the dominant solution lies in that the feature extractor learns the discriminative features of different groups and the classifier is able to provide appropriate and unbiased margins for different groups by the discriminative features. Therefore, in this novel, we propose an innovative collaborative learning framework (GroupFace) that integrates a multi-hop attention graph convolutional network and a dynamic group-aware margin strategy based on reinforcement learning. Specifically, to extract the discriminative features of different groups, we design an enhanced multi-hop attention graph convolutional network. This network is capable of capturing the interactions of neighboring nodes at different distances, fusing local and global information to model facial deep aging, and exploring diverse representations of different groups. In addition, to further address the class imbalance problem, we design a dynamic group-aware margin strategy based on reinforcement learning to provide appropriate and unbiased margins for different groups. The strategy divides the sample into four age groups and considers identifying the optimum margins for various age groups by employing a Markov decision process. Under the guidance of the agent, the feature representation bias and the classification margin deviation between different groups can be reduced simultaneously, balancing inter-class separability and intra-class proximity. After joint optimization, our architecture achieves excellent performance on several age estimation benchmark datasets. It not only achieves large improvements in overall estimation accuracy but also gains balanced performance in long-tailed group estimation. Yuntao Shou, Wei Ai 0001, Keqin Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | SE-GNN: Seed Expanded-Aware Graph Neural Network With Iterative Optimization for Semi-Supervised Entity AlignmentabstractEntity alignment aims to use pre-aligned seed pairs to find other equivalent entities from different knowledge graphs and is widely used in graph fusion-related fields. However, as the scale of knowledge graphs increases, manually annotating pre-aligned seed pairs becomes difficult. Existing research utilizes entity embeddings obtained by aggregating single structural information to identify potential seed pairs, thus reducing the reliance on pre-aligned seed pairs. However, due to the structural heterogeneity of KG, the quality of potential seed pairs obtained using only a single structural information is not ideal. In addition, although existing research improves the quality of potential seed pairs through semi-supervised iteration, they underestimate the impact of embedding distortion produced by noisy seed pairs on the alignment effect. In order to solve the above problems, we propose a seed expanded-aware graph neural network with iterative optimization for semi-supervised entity alignment, named SE-GNN. First, we utilize the semantic attributes and structural features of entities, combined with a conditional filtering mechanism, to obtain high-quality initial potential seed pairs. Next, we designed a local and global awareness mechanism. It introduces initial potential seed pairs and combines local and global information to obtain a more comprehensive entity embedding representation, which alleviates the impact of KG structural heterogeneity and lays the foundation for the optimization of initial potential seed pairs. Then, we designed the threshold nearest neighbor embedding correction strategy. It combines the similarity threshold and the bidirectional nearest neighbor method as a filtering mechanism to select iterative potential seed pairs and also uses an embedding correction strategy to eliminate the embedding distortion. Finally, we will reach the optimized potential seeds after iterative rounds to input local and global sensing mechanisms, obtain the final entity embedding, and perform entity alignment. Experimental results on public datasets demonstrate the excellent performance of our SE-GNN, showcasing the effectiveness of the model. Our code is publicly available athttps://github.com/ShuoShan1/SE-GNN. Hongen Shao, Yuntao Shou, Wei Ai 0001, Keqin Li 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | DER-GCN: Dialog and Event Relation-Aware Graph Convolutional Neural Network for Multimodal Dialog Emotion RecognitionabstractWith the continuous development of deep learning (DL), the task of multimodal dialog emotion recognition (MDER) has recently received extensive research attention, which is also an essential branch of DL. The MDER aims to identify the emotional information contained in different modalities, e.g., text, video, and audio, and in different dialog scenes. However, the existing research has focused on modeling contextual semantic information and dialog relations between speakers while ignoring the impact of event relations on emotion. To tackle the above issues, we propose a novel dialog and event relation-aware graph convolutional neural network (DER-GCN) for multimodal emotion recognition method. It models dialog relations between speakers and captures latent event relations information. Specifically, we construct a weighted multirelationship graph to simultaneously capture the dependencies between speakers and event relations in a dialog. Moreover, we also introduce a self-supervised masked graph autoencoder (SMGAE) to improve the fusion representation ability of features and structures. Next, we design a new multiple information Transformer (MIT) to capture the correlation between different relations, which can provide a better fuse of the multivariate information between relations. Finally, we propose a loss optimization strategy based on contrastive learning to enhance the representation learning ability of minority class features. We conduct extensive experiments on the benchmark datasets, Interactive Emotional Dyadic Motion Capture (IEMOCAP) and Multimodal EmotionLines Dataset (MELD), which verify the effectiveness of the DER-GCN model. The results demonstrate that our model significantly improves both the average accuracy and the value of emotion recognition. Our code is publicly available at https://github.com/yuntaoshou/DER-GCN. Wei Ai 0001, Yuntao Shou, Keqin Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | SSD-GNN: Fraud Detection based on Spectral and Spatial Dual Graph Neural NetworksabstractGraph Neural Networks (GNNs) have attracted much attention due to their outstanding performance in fraud detection tasks, which reveal fraudsters by aggregating neighborhood features. However, most GNNs assume that neighborhood features are homogeneous, which does not match the fundamentally heterogeneous fraud graph, while heterogeneous information is equally important for fraud detection. In addition, the label-imbalanced neighborhood also deteriorates fraud detection accuracy. To address the above issues, we propose an innovative spectral and spatial dual graph neural network for fraud detection, namely SSD-GNN. In SSD-GNN, we design a dual-pathway framework, including spectral and spatial processing paths, and aggregate neighborhood features according to the homogeneity and heterogeneity assumptions, respectively. In the spectral processing path, we design high-pass filtering and low-pass filtering to form a mixed-pass filter to extract low-frequency information of homogeneity and high-frequency information of heterogeneity. In the spatial processing path, we first enhance the connection between the target node and the global node by clustering information to alleviate the label imbalance problem. Then we design similarity aggregation and disparity aggregation to capture homogeneity and heterogeneity information, respectively. We run experiments on two real-world fraud detection datasets, using AUC and F1-macro as test metrics. The comprehensive experimental results show that the AUC and F1-macro on the average source exceed at least 5 percent. demonstrate the superiority of the proposed SSD-GNN. Boyi He, Jianzhe Zhao, Wei Ai 0001 |
HPCC | 4 |
| 2024 | Edge-enhanced minimum-margin graph attention network for short text classification
Wei Ai 0001, Hongen Shao, Yuntao Shou, Keqin Li 0001 |
Expert Syst. Appl. | 1 |
| 2024 | A multi-message passing framework based on heterogeneous graphs in conversational emotion recognition
Yuntao Shou, Wei Ai 0001, Jiayi Du, Keqin Li 0001 |
Neurocomputing | 3 |
| 2024 | Stackelberg Game-Based Bandwidth Allocation and Resource Pricing for Multiuser in MEC SystemabstractWith the rapid development of artificial intelligence, a substantial number of computing-intensive applications have emerged in Internet of Things (IoT) devices. The mobile edge computing (MEC) architecture enables the provision of abundant computing and storage resources in close proximity to end users (EUs), thereby effectively enhancing their quality of experience (QoE). Nonetheless, both the MEC server and EUs are self-interests, it is crucial to establish suitable incentive mechanism to promote active engagement from both parties in the offloading process. Therefore, we employ the Stackelberg game to describe the interaction process between EUs and the MEC server, and an optimal relationship between bandwidth and offloading task size is established to simplify the decision problem for EUs. Then, the optimal strategies for the MEC server and EUs are solved using reverse induction. Given the limited resources of the MEC server, we propose a dynamic programming-based resource allocation (DPRA) algorithm to maximize the revenue of the MEC server while ensuring the cost of each EU. The simulation results demonstrate that the DPRA algorithm can reduce latency and energy consumption costs, significantly outperforming other comparative strategies in terms of performance at both EUs and the MEC server. Zhao Tong 0001, Yuanyang Zhang, Jing Mei, Wei Ai 0001, Kenli Li 0001, Keqin Li 0001 |
IEEE Internet Things J. | 4 |
| 2024 | A multi-view mask contrastive learning graph convolutional neural network for age estimation
Yuntao Shou, Wei Ai 0001, Keqin Li 0001 |
Knowl. Inf. Syst. | 4 |
| 2024 | Masked Graph Learning With Recurrent Alignment for Multimodal Emotion Recognition in ConversationabstractSince Multimodal Emotion Recognition in Conversation (MERC) can be applied to public opinion monitoring, intelligent dialogue robots, and other fields, it has received extensive research attention in recent years. Unlike traditional unimodal emotion recognition, MERC can fuse complementary semantic information between multiple modalities (e.g., text, audio, and vision) to improve emotion recognition. However, previous work ignored the inter-modal alignment process and the intra-modal noise information before multimodal fusion but directly fuses multimodal features, which will hinder the model for representation learning. In this study, we have developed a novel approach called Masked Graph Learning with Recursive Alignment (MGLRA) to tackle this problem, which uses a recurrent iterative module with memory to align multimodal features, and then uses the masked GCN for multimodal feature fusion. First, we employ LSTM to capture contextual information and use a graph attention-filtering mechanism to eliminate noise effectively within the modality. Second, we build a recurrent iteration module with a memory function, which can use communication between different modalities to eliminate the gap between modalities and achieve the preliminary alignment of features between modalities. Then, a cross-modal multi-head attention mechanism is introduced to achieve feature alignment between modalities and construct a masked GCN for multimodal feature fusion, which can perform random mask reconstruction on the nodes in the graph to obtain better node feature representation. Finally, we utilize a multilayer perceptron (MLP) for emotion recognition. Extensive experiments on two benchmark datasets (i.e., IEMOCAP and MELD) demonstrate that MGLRA outperforms state-of-the-art methods. Fuchen Zhang, Yuntao Shou, Hongen Shao, Wei Ai 0001, Keqin Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | A D-Truss-Equivalence Based Index for Community Search Over Large Directed GraphsabstractCommunity Search (CS) aims to enable online and personalized discovery of communities. Recently, attention to the CS problem in directed graphs (di-graph) needs to be improved despite the extensive study conducted on undirected graphs. Nevertheless, the existing studies are plagued by several shortcomings, e.g., Achieving high-performance CS while ensuring the retrieved community is cohesive is challenging. This paper uses the D-truss model to address the limitations of investigating the CS problem in large di-graphs. We aim to implement millisecond-level D-truss CS in di-graphs by building a summarized graph index. To capture the interconnectedness of edges within D-truss communities, we propose an innovative equivalence relation known as D-truss-equivalence, which allows us to divide the edges in a di-graph into a sequence of super nodes (s-nodes). These s-nodes form the D-truss-equivalencebased index, DEBI, an index structure that preserves the truss properties and ensures efficient space utilization. Using DEBI, CS can be performed without time-consuming access to the original graph. The experiments indicate that our method can achieve millisecond-level D-truss community query while ensuring high community quality. In addition, dynamic maintenance of indexes can also be achieved at a lower cost. Our code is available athttps://github.com/XieCanhao04/DEBI. Wei Ai 0001, CanHao Xie, Jiayi Du, Keqin Li 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | An Effective Index for Truss-based Community Search on Large Directed GraphsabstractCommunity search is a derivative of community detection that enables online and personalized discovery of communities and has found extensive applications in massive real-world networks. Recently, there needs to be more focus on the community search issue within directed graphs, even though substantial research has been carried out on undirected graphs. The recently proposed D-truss model has achieved good results in the quality of retrieved communities. However, existing D-truss-based work cannot perform efficient community searches on large graphs because it consumes too many computing resources to retrieve the maximal D-truss. To overcome this issue, we introduce an innovative merge relation known as D-truss-connected to capture the inherent density and cohesiveness of edges within D-truss. This relation allows us to partition all the edges in the original graph into a series of D-truss-connected classes. Then, we construct a concise and compact index, ConDTruss, based on D-truss-connected. Using ConDTruss, the efficiency of maximum D-truss retrieval will be greatly improved, making it a theoretically optimal approach. Experimental evaluations conducted on large directed graphs certificate the effectiveness of our proposed method. Wei Ai 0001, CanHao Xie, Yinghao Wu, Keqin Li 0001 |
ICPADS | 1 |
| 2023 | A Two-Stage Multimodal Emotion Recognition Model Based on Graph Contrastive LearningabstractIn terms of human-computer interaction, it is becoming more and more important to correctly understand the user’s emotional state in a conversation, so the task of multimodal emotion recognition (MER) started to receive more attention. However, existing emotion classification methods usually perform classification only once. Sentences are likely to be misclassified in a single round of classification. Previous work usually ignores the similarities and differences between different morphological features in the fusion process. To address the above issues, we propose a two-stage emotion recognition model based on graph contrastive learning (TS-GCL). First, we encode the original dataset with different preprocessing modalities. Second, a graph contrastive learning (GCL) strategy is introduced for these three modal data with other structures to learn similarities and differences within and between modalities. Finally, we use MLP twice to achieve the final emotion classification. This staged classification method can help the model to better focus on different levels of emotional information, thereby improving the performance of the model. Extensive experiments show that TS-GCL has superior performance on IEMOCAP and MELD datasets compared with previous methods. Wei Ai 0001, Fuchen Zhang, Yuntao Shou, Hongen Shao, Keqin Li 0001 |
ICPADS | 1 |
| 2023 | Fast Butterfly-Core Community Search For Large Labeled GraphsabstractCommunity Search (CS) aims to identify densely interconnected subgraphs corresponding to query vertices within a graph. However, existing heterogeneous graph-based community search methods need help identifying cross-group communities and suffer from efficiency issues, making them unsuitable for large graphs. This paper presents a fast community search model based on the Butterfly-Core Community (BCC) structure for heterogeneous graphs. The Random Walk with Restart (RWR) algorithm and butterfly degree comprehensively evaluate the importance of vertices within communities, allowing leader vertices to be rapidly updated to maintain cross-group cohesion. Moreover, we devised a more efficient method for updating vertex distances, which minimizes vertex visits and enhances operational efficiency. Extensive experiments on several real-world temporal graphs demonstrate the effectiveness and efficiency of this solution. Jiayi Du, Yinghao Wu, Wei Ai 0001, CanHao Xie, Keqin Li 0001 |
ICPADS | 3 |
| 2023 | GraphUnet: Graph Make Strong Encoders for Remote Sensing SegmentationabstractRemote sensing segmentation are widely applied in environmental protection, and urban change detection, etc. Despite the success of deep learning-based remote sensing segmentation methods (e.g., CNN and Transformer), they are not flexible enough to model irregular objects. To tackle the above problems, this paper treats images as graph structures and introduces a simple vision GNN (GraphUNet) architecture for remote sensing segmentation. Specifically, we construct a graph view and utilize GNN to extract the visual feature. To extract the global contextual location information in the image, we introduce a multi-head attention mechanism for global information extraction. Furthermore, we introduce a feed-forward network for each node to perform feature transformation on node features to encourage information diversity. Extensive experiments on three real datasets demonstrate that our model GraphUNet outperforms SOTA remote sensing image segmentation methods. Yuntao Shou, Wei Ai 0001, Fuchen Zhang, Keqin Li 0001 |
ICPADS | 2 |
| 2023 | A multi-semantic passing framework for semi-supervised long text classification
Wei Ai 0001, Hongen Shao, Keqin Li 0001 |
Appl. Intell. | 1 |
| 2022 | Conversational emotion recognition studies based on graph convolutional neural networks and a dependent syntactic analysis
Yuntao Shou, Wei Ai 0001, Keqin Li 0001 |
Neurocomputing | 3 |
| 2017 | DHCRF: A Distributed Conditional Random Field Algorithm on a Heterogeneous CPU-GPU Cluster for Big DataabstractAs one of the most recognized models in machine learning, the conditional random fields (CRF) has been widely used in many applications. As the parameter estimation of CRF is highly time-consuming, how to improve the performance of CRF has received significant attention, in particular in the big data environment. To deal with large-scale data, CPU-based or GPU-based parallelization solutions have been proposed to improve performance. However, the problem is an ongoing one. In this paper, we focus on the big data environment and propose a distributed CRF on a heterogeneous CPU-GPU cluster called DHCRF. Our approach differs from previous work. Specifically, it leverages a three-stage heterogeneous Map and Reduce operation to improve the performance, making full use of CPU-GPU collaborative computing capabilities in a big data environment. Furthermore, by combining elastic data partition and intermediate results multiplexing method, the distributed CRF is optimized. Elastic data partition is performed to keep the load balanced, and the intermediate results multiplexing method is adopted to reduce data communication. Experimental results show that the DHCRF outperforms the baseline CRF algorithm and the CPU-based parallel CRF algorithm with notable performance improvement while maintaining competitive correctness at the same time. Wei Ai 0001, Kenli Li 0001, Cen Chen 0002, Jiwu Peng, Keqin Li 0001 |
ICDCS | 1 |
| 2015 | Hadoop Recognition of Biomedical Named Entity Using Conditional Random FieldsabstractProcessing large volumes of data has presented a challenging issue, particularly in data-redundant systems. As one of the most recognized models, the conditional random fields (CRF) model has been widely applied in biomedical named entity recognition (Bio-NER). Due to the internally sequential feature, performance improvement of the CRF model is nontrivial, which requires new parallelized solutions. By combining and parallelizing the limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) and Viterbi algorithms, we propose a parallel CRF algorithm called MapReduce CRF (MRCRF) in this paper, which contains two parallel sub-algorithms to handle two time-consuming steps of the CRF model. The MapReduce L-BFGS (MRLB) algorithm leverages the MapReduce framework to enhance the capability of estimating parameters. Furthermore, the MapReduce Viterbi (MRVtb) algorithm infers the most likely state sequence by extending the Viterbi algorithm with another MapReduce job. Experimental results show that the MRCRF algorithm outperforms other competing methods by exhibiting significant performance improvement in terms of time efficiency as well as preserving a guaranteed level of correctness. Kenli Li 0001, Wei Ai 0001, Zhuo Tang, Fan Zhang 0003, Lingang Jiang, Keqin Li 0001, Kai Hwang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |