VLDB 2026 Research / reviewers in the wild / expert
Lili Guo 0001
dblp:49/6485-1
· DBLP profile ↗
61ranked-venue papers
10as first author
51since 2021 · last 2026
0000-0002-5526-0980ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 5 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 12 · 12 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-view Hierarchical Graph Contrastive Learning based on Asynchronous Asymmetric StructureabstractContrastive learning has strong generalization ability and the capability to learn automatically without labeled information. However, it still faces challenges such as insufficient feature diversity, a lack of multi-level semantics, and the balance between tolerance and consistency. To address these challenges, This study propose a Multi-view Hierarchical Graph Contrastive Learning method. First, a new view is generated through a diffusion matrix to provide multi-view data for contrastive learning. Then, these multi-view data are fed into an asynchronous asymmetric network structure, specifically using graph network models to learn diversified features. Next, we adopt a self-designed hierarchical contrastive learning framework, constructing a three-level contrastive loss for joint optimization of nodes, subgraphs, and global graphs. Meanwhile, we introduce alignment and consistency and appropriately adjust the loss function through a temperature coefficient. Ultimately, the model achieves excellent classification performance on multiple datasets through node classification and graph classification tasks. Chuangui Cao, Shifei Ding, Jian Zhang 0019, Lili Guo 0001, Xuan Li 0004 |
WWW | 4 |
| 2026 | Contextual Structure-Enhanced Selective Graph Convolutional NetworkabstractGraph Neural Networks fundamentally rely on homophily assumptions where connected nodes are expected to share similar labels, consequently suffering severe performance degradation in heterophilic graphs due to the indiscriminate neighbor aggregation mechanism. Although recent solutions have attempted to incorporate higher-order neighborhoods or reweighting schemes, they often inadvertently amplify structural noise by introducing a larger proportion of dissimilar nodes than similar ones, while simultaneously failing to capture nuanced contextual patterns due to their inability to discern subtle local structural variations across subgraphs. To holistically address these intractable and co-existing challenges, we propose the Contextual Structure Enhanced Selective Graph Convolutional Network (CSS-GCN), a novel architecture that organically synergizes contextual structure modeling with adaptive neighbor selection. Specifically, our approach employs ego-network partitioning and group fairness constraints to effectively quantify domain-invariant structural patterns, thereby countering the contextual blindness often observed in conventional GNNs. Complementarily, we design a selective propagation mechanism unifying adaptive neighborhood distribution-based similarity computation with the gated fusion of three distinct information pathways: potential homophilic neighbors identified through attribute-topology synergy, first-hop connections, and ego-representations. This dual-component framework enables nodes to dynamically filter out irrelevant signals while preserving structural consistency across diverse homophily-heterophily landscapes. Extensive validation on 10 real-world graphs demonstrates the effectiveness and superiority of our proposed approach. Shifei Ding, Fangchen Li, Lili Guo 0001, Jian Zhang 0019 |
WWW | 3 |
| 2026 | EmoKEG: Knowledge-enhanced heterogeneous graph for emotion recognition in conversation
Lili Guo 0001, Yanan Cui, Yikang Song, Haiwei Hou, Shifei Ding |
Pattern Recognit. | 1 |
| 2026 | A comprehensive survey of image clustering based on deep learning
Haiwei Hou, Shifei Ding, Chuangui Cao, Xiao Xu 0006, Lili Guo 0001, Xuan Li 0004 |
Pattern Recognit. | 5 |
| 2025 | A Medical Image Classification Network Based on Multi-View Consistent Momentum Contrastive LearningabstractDue to variations in imaging conditions, images often exhibit discrepancies in color reproduction. Furthermore, motion-induced blur can lead to edge degradation, making color sensitivity and edge blurriness two prevalent and challenging issues in both natural image processing and medical image analysis. To address these challenges, we propose a model termed the Three-View Consistency Mo-mentum Contrastive with Sobel Operator (SVCMC). Specifically, we first design a three-view momen-tum-update architecture that employs a So-bel-augmented ResNet as the backbone. We then introduce a novel contrastive loss, referred to as the Three-View Consistency Momentum Contrastive Loss. Next, to mitigate the oscillations and slow convergence commonly observed in contrastive learning, we construct a dynamic contrastive loss function that adapts in real time over the training process. Finally, we validated the superiority of our model on two medical image datasets and one natural image dataset, where its classification ac-curacy and convergence speed significantly out-performed existing state-of-the-art contrastive models. Chuangui Cao, Shifei Ding, Lili Guo 0001 |
IJCAI | 3 |
| 2025 | Global Information Compensation Network for Image DenoisingabstractIn image denoising research, discriminative models have achieved impressive results which mainly owes to the powerful ability of convolutional networks in local feature extraction. However, there is still room for improvement due to insufficient utilization of global information. Although using fully connected layers or increasing network depth can supplement global information, this results in a significant increase in parameters and computational cost. To address these issues, we propose a global information compensation network (GICN) for image denoising in this paper. Firstly, at the shallow network part, we propose a global feature mining block that enhances the network's ability to extract global information by combining non-local blocks and the Fourier transform while improving the interpretability of the model. Secondly, between the encoder and decoder, we propose a cross-scale feature aggregation block to fuse information at different scales. Finally, we employ attention blocks to improve skip connections to better capture long-distance dependencies. Extensive experimental results show that our proposed GICN effectively compensates for global information, achieves a balance between denoising efficiency and effect, and surpasses mainstream methods in multiple benchmark tests. Shifei Ding, Qidong Wang, Lili Guo 0001 |
IJCAI | 3 |
| 2025 | Multi-modal Anchor Gated Transformer with Knowledge Distillation for Emotion Recognition in ConversationabstractEmotion Recognition in Conversation (ERC) aims to detect the emotions of individual utterances within a conversation. Generating efficient and modality-specific representations for each utterance remains a significant challenge. Previous studies have proposed various models to integrate features extracted using different modality-specific encoders. However, they neglect the varying contributions of modalities to this task and introduce high complexity by aligning modalities at the frame level. To address these challenges, we propose the Multi-modal Anchor Gated Transformer with Knowledge Distillation (MAGTKD) for the ERC task. Specifically, prompt learning is employed to enhance textual modality representations, while knowledge distillation is utilized to strengthen representations of weaker modalities. Furthermore, we introduce a multi-modal anchor gated transformer to effectively integrate utterance-level representations across modalities. Extensive experiments on the IEMOCAP and MELD datasets demonstrate the effectiveness of knowledge distillation in enhancing modality representations and achieve state-of-the-art performance in emotion recognition. Our code is available at: https://github.com/JieLi-dd/MAGTKD. Jie Li 0069, Shifei Ding, Lili Guo 0001, Xuan Li 0004 |
IJCAI | 3 |
| 2025 | L2DGCN: Learnable Enhancement and Label Selection Dynamic Graph Convolutional Networks for Mitigating Degree BiasabstractGraph Neural Networks (GNNs) are powerful models for node classification, but their performance is heavily reliant on manually labeled data, which is often costly and results in insufficient labeling. Recent studies have shown that message-passing neural networks struggle to propagate information in low-degree nodes, negatively affecting overall performance. To address the information bias caused by degree imbalance, we propose a Learnable Enhancement and Label Selection Dynamic Graph Convolutional Network (L2DGCN). L2DGCN consists of a teacher model and a student model. The teacher model employs an improved label propagation mechanism that enables remote label information dissemination among all nodes. The student model introduces a dynamically learnable graph enhancement strategy, perturbing edges to facilitate information exchange among low-degree nodes. This approach maintains the global graph structure while learning graph representations. Additionally, we have designed a label selector to mitigate the impact of unreliable pseudo-labels on model learning. To validate the effectiveness of our proposed model with limited labeled data, we conducted comprehensive evaluations of semi-supervised node classification across various scenarios with a limited number of annotated nodes. Experimental results demonstrate that our data enhancement model significantly contributes to node classification tasks under sparse labeling conditions. Jingxiao Zhang, Shifei Ding, Lili Guo 0001, Xuan Li 0004 |
NeurIPS | 4 |
| 2025 | Enhancing histopathological image classification through multi-view momentum encoding learning
Chuangui Cao, Shifei Ding, Lili Guo 0001, Xiaohao Xie |
Appl. Intell. | 3 |
| 2025 | A novel robust semi-supervised stochastic configuration network for regression tasks with noise
Shifei Ding, Zi Zhang, Chenglong Zhang 0001, Lili Guo 0001, Xuan Li 0004 |
Inf. Sci. | 5 |
| 2025 | Semi-supervised classification model with stochastic configuration networks
Shifei Ding, Zi Zhang, Chenglong Zhang 0001, Lili Guo 0001, Xuan Li 0004 |
Knowl. Inf. Syst. | 5 |
| 2025 | Triple-view graph clustering network based on high-confidence contrastive learning strategy
Shifei Ding, Zhe Li 0071, Xiao Xu 0006, Lili Guo 0001, Ling Ding 0001 |
Knowl. Based Syst. | 4 |
| 2025 | Robust Density Peaks Clustering for Manifold Data With Multiple PeaksabstractDensity peaks clustering (DPC) is an excellent clustering algorithm that does not need any prior knowledge. However, DPC still has the following shortcomings: (1) The Euclidean distance used by it is not applicable to manifold data with multiple peaks. (2) The local density calculation for DPC is too simple, and the final results may fluctuate due to the cutoff-distancedc. (3) Manually selected centers by decision-graph may lead to a wrong number of clusters and poor performance. To address these shortcomings and improve the performance, a robust density peaks clustering algorithm for manifold data with multiple peaks (RDPCM) is proposed to reduce the sensitivity of clustering results to parameters. Motivated by DPC-GD, RDPCM replaces the Euclidean distance with geodesic distance, which is optimized by the improved mutual K-nearest neighbors. It better considers the local manifold structure of the datasets and obtains excellent results. In addition, the Davies-Bouldin Index based on Minimum Spanning Tree (MDBI) is proposed to select the ideal number of classes adaptively. Numerous experiments have established that RDPCM is more effective and superior than other advanced clustering algorithms. Ling Ding 0001, Chao Li 0102, Shifei Ding, Xiao Xu 0006, Lili Guo 0001, Xindong Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Multi-Agent policy gradients with dynamic weighted value decomposition
Shifei Ding, Xiaomin Dong, Jian Zhang 0019, Lili Guo 0001, Wei Du 0010, Chenglong Zhang 0001 |
Pattern Recognit. | 4 |
| 2025 | Multi-channel set polynomial based label regularized graph neural networks against extreme data scarcity
Jingxiao Zhang, Shifei Ding, Jian Zhang 0019, Lili Guo 0001, Ling Ding 0001 |
Pattern Recognit. | 4 |
| 2025 | APIN: Amplitude- and phase-aware interaction network for speech emotion recognition
Lili Guo 0001, Jie Li 0069, Shifei Ding, Jianwu Dang 0001 |
Speech Commun. | 1 |
| 2025 | Fast Density Peaks Clustering Algorithm Based on Approximate k-Nearest NeighborsabstractDensity peaks clustering (DPC) is one of the density-based clustering algorithms and has been widely studied and applied in recent years because of its unique parameter, non-iteration and good robustness. However, it cannot effectively identify the cluster centers, and time and space complexities are too high. To this end, this paper proposes a fast density peaks clustering algorithm based on approximatek-nearest neighbors (FDPAN). Firstly, it uses Balanced K-means based Hierarchical K-means (BKHK) method to partition the data and quickly find the approximatek-nearest neighbors (AKNN), improving the algorithm’s efficiency on large-scale high-dimensional data. Meanwhile, three-way clustering is used to improve the neighbor search of the boundary points of the partition. Then, the local density and relative distance of DPC are recalculated by AKNN. Finally, according to the similar density chain, the connected high-density points are labeled while searching for the cluster center, and the remaining points are assigned to the clusters where their nearest higher-density points are located. Theoretical analysis and experiments on synthetic and real datasets show that FDPAN can obtain higher clustering results and shorten the operation time on large-scale high-dimensional data compared with DPC and its variants. Shifei Ding, Chao Li 0102, Xiao Xu 0006, Lili Guo 0001, Ling Ding 0001, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Vertical Federated Density Peaks Clustering Under Nonlinear MappingabstractAs the representative density-based clustering algorithm, density peaks clustering (DPC) has wide recognition, and many improved algorithms and applications have been extended from it. However, the DPC involving privacy protection has not been deeply studied. In addition, there is still room for improvement in the selection of centers and allocation methods of DPC. To address these issues, vertical federated density peaks clustering under nonlinear mapping (VFDPC) is proposed to address privacy protection issues in vertically partitioned data. Firstly, a hybrid encryption privacy protection mechanism is proposed to protect the merging process of distance matrices generated by client data. Secondly, according to the merged distance matrix, a more effective cluster merging under nonlinear mapping is proposed to ameliorate the process of DPC. Results on man-made, real, and multi-view data fully prove the improvement of VFDPC on clustering accuracy. Chao Li 0102, Shifei Ding, Xiao Xu 0006, Lili Guo 0001, Ling Ding 0001, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Parameter-Adaptive Border Peeling Clustering AlgorithmabstractMost clustering algorithms require setting one or more parameters, which rely on prior knowledge or are constantly adjusted based on external indicators. To address the issues of requiring external index guidance, blindness, and time-consuming parameter setting for clustering algorithms on complex data, we propose a novel Parameter-Adaptive Border Peeling clustering algorithm (PABP). The PABP algorithm initially employs the maximum number of neighbors identified through natural neighbor search to automatically ascertain the number of local neighborhoods. At the same time, the Gaussian kernel bandwidth can be adaptively obtained in density measurement, which can highlight high-density areas. Secondly, the number of peels is adaptively determined by the coefficient of variation of density during the iterative border peeling process. Lastly, labels are assigned to core points based on graph connections, while the clustering of border points is accomplished via label propagation. PABP does not require users to adjust parameters based on prior knowledge or external indicators throughout the entire process. In the experiment, PABP was compared with seven other advanced clustering algorithms on 13 synthetic datasets, 10 UCI datasets, and Olivetti Face and MNIST datasets. The results indicate that the clustering performance of PABP is superior to the compared algorithms. Hui Tu, Shifei Ding, Xiao Xu 0006, Lili Guo 0001, Ling Ding 0001, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Multiagent Reinforcement Learning With Graphical Mutual Information MaximizationabstractCommunication learning is an important research direction in the multiagent reinforcement learning (MARL) domain. Graph neural networks (GNNs) can aggregate the information of neighbor nodes for representation learning. In recent years, several MARL methods leverage GNN to model information interactions between agents to coordinate actions and complete cooperative tasks. However, simply aggregating the information of neighboring agents through GNNs may not extract enough useful information, and the topological relationship information is ignored. To tackle this difficulty, we investigate how to efficiently extract and utilize the rich information of neighbor agents as much as possible in the graph structure, so as to obtain high-quality expressive feature representation to complete the cooperation task. To this end, we present a novel GNN-based MARL method with graphical mutual information (MI) maximization to maximize the correlation between input feature information of neighbor agents and output high-level hidden feature representations. The proposed method extends the traditional idea of MI optimization from graph domain to multiagent system, in which the MI is measured from two aspects: agent features information and agent topological relationships. The proposed method is agnostic to specific MARL methods and can be flexibly integrated with various value function decomposition methods. Considerable experiments on various benchmarks demonstrate that the performance of our proposed method is superior to the existing MARL methods. Shifei Ding, Wei Du 0010, Ling Ding 0001, Jian Zhang 0019, Lili Guo 0001, Bo An 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Horizontal Federated Density Peaks ClusteringabstractDensity peaks clustering (DPC) is a popular clustering algorithm, which has been studied and favored by many scholars because of its simplicity, fewer parameters, and no iteration. However, in previous improvements of DPC, the issue of privacy data leakage was not considered, and the "Domino" effect caused by the misallocation of noncenters has not been effectively addressed. In view of the above shortcomings, a horizontal federated DPC (HFDPC) is proposed. First, HFDPC introduces the idea of horizontal federated learning and proposes a protection mechanism for client parameter transmission. Second, DPC is improved by using similar density chain (SDC) to alleviate the "Domino" effect caused by multiple local peaks in the flow pattern dataset. Finally, a novel data dimension reduction and image encryption are used to improve the effectiveness of data partitioning. The experimental results show that compared with DPC and some of its improvements, HFDPC has a certain degree of improvement in accuracy and speed. Shifei Ding, Chao Li 0102, Xiao Xu 0006, Lili Guo 0001, Ling Ding 0001, Xindong Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | DSTCNet: Deep Spectro-Temporal-Channel Attention Network for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) plays an important role in human-computer interaction, which can provide better interactivity to enhance user experiences. Existing approaches tend to directly apply deep learning networks to distinguish emotions. Among them, the convolutional neural network (CNN) is the most commonly used method to learn emotional representations from spectrograms. However, CNN does not explicitly model features' associations in the spectral-, temporal-, and channel-wise axes or their relative relevance, which will limit the representation learning. In this article, we propose a deep spectro-temporal-channel network (DSTCNet) to improve the representational ability for speech emotion. The proposed DSTCNet integrates several spectro-temporal-channel (STC) attention modules into a general CNN. Specifically, we propose the STC module that infers a 3-D attention map along the dimensions of time, frequency, and channel. The STC attention can focus more on the regions of crucial time frames, frequency ranges, and feature channels. Finally, experiments were conducted on the Berlin emotional database (EmoDB) and interactive emotional dyadic motion capture (IEMOCAP) databases. The results reveal that our DSTCNet can outperform the traditional CNN-based and several state-of-the-art methods. Lili Guo 0001, Shifei Ding, Longbiao Wang, Jianwu Dang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Learning Efficient and Robust Multi-Agent Communication via Graph Information BottleneckabstractEfficient communication learning among agents has been shown crucial for cooperative multi-agent reinforcement learning (MARL), as it can promote the action coordination of agents and ultimately improve performance. Graph neural network (GNN) provide a general paradigm for communication learning, which consider agents and communication channels as nodes and edges in a graph, with the action selection corresponding to node labeling. Under such paradigm, an agent aggregates information from neighbor agents, which can reduce uncertainty in local decision-making and induce implicit action coordination. However, this communication paradigm is vulnerable to adversarial attacks and noise, and how to learn robust and efficient communication under perturbations has largely not been studied. To this end, this paper introduces a novel Multi-Agent communication mechanism via Graph Information bottleneck (MAGI), which can optimally balance the robustness and expressiveness of the message representation learned by agents. This communication mechanism is aim at learning the minimal sufficient message representation for an agent by maximizing the mutual information (MI) between the message representation and the selected action, and simultaneously constraining the MI between the message representation and the agent feature. Empirical results demonstrate that MAGI is more robust and efficient than state-of-the-art GNN-based MARL methods. Shifei Ding, Wei Du 0010, Ling Ding 0001, Lili Guo 0001, Jian Zhang 0019 |
AAAI | 4 |
| 2024 | Expressive Multi-Agent Communication via Identity-Aware LearningabstractInformation sharing through communication is essential for tackling complex multi-agent reinforcement learning tasks. Many existing multi-agent communication protocols can be viewed as instances of message passing graph neural networks (GNNs). However, due to the significantly limited expressive ability of the standard GNN method, the agent feature representations remain similar and indistinguishable even though the agents have different neighborhood structures. This further results in the homogenization of agent behaviors and reduces the capability to solve tasks effectively. In this paper, we propose a multi-agent communication protocol via identity-aware learning (IDEAL), which explicitly enhances the distinguishability of agent feature representations to break the diversity bottleneck. Specifically, IDEAL extends existing multi-agent communication protocols by inductively considering the agents' identities during the message passing process. To obtain expressive feature representations for a given agent, IDEAL first extracts the ego network centered around that agent and then performs multiple rounds of heterogeneous message passing, where different parameter sets are applied to the central agent and the other surrounding agents within the ego network. IDEAL fosters expressive communication between agents and generates distinguishable feature representations, which promotes action diversity and individuality emergence. Experimental results on various benchmarks demonstrate IDEAL can be flexibly integrated into various multi-agent communication methods and enhances the corresponding performance. Wei Du 0010, Shifei Ding, Lili Guo 0001, Jian Zhang 0019, Ling Ding 0001 |
AAAI | 3 |
| 2024 | A fine-tuning framework for cross-domain RUL prediction based on relevant health indicatorsabstractBearings are crucial elements in mechanical equipment, and their deterioration can result in equipment malfunction or failure. Therefore, precise analysis of the degradation process and proactive prediction of the remaining useful life are essential measures to prevent significant accidents. Practically, the intricacy and fluctuation of operational circumstances result in cross-domain issues in the training and test datasets, meaning that the source and target domains have inconsistent distributions. The disparities from data collection to feature extraction stages might exacerbate the discrepancies between the two domains. Additionally, selecting metrics can restrict the effectiveness of cross-domain prediction. To solve the above problems, this paper proposes a cross-domain transfer prediction method based on generalized HI for fine-tuning. Initially, multidimensional characteristics are retrieved, followed by acquiring HI metrics using a secondary optimization indicator and correlation extraction. Ultimately, the original dataset from the source domain is expanded to enhance the accuracy of prediction through time-series data augmentation. The effectiveness and practicality of the suggested strategy are convincingly showcased through experiments conducted on the PHM2012 dataset. Junrong Du, Zimeng Fan 0003, Lei Song 0011, Xuanang Gui, Lili Guo 0001, Xuzhi Li |
CSCWD | 5 |
| 2024 | Robust Classification of Incomplete Time Series with Noisy LabelsabstractMissing data and noisy labeling are common problems in time series analysis. The traditional approach to deal with missing data is to separate interpolation and classification, which is not interactive and provides unsatisfactory performance. While advanced methods can learn features from missing information, feature representation is limited due to the accumulation of interpolation errors. For noisy label interference, a robust loss function is a simpler and more general solution for robust learning. This study proposes an end-to-end neural network that unifies data interpolation and feature learning within a single framework. The focus is placed on extracting useful information from incomplete time series data, and for the computation of classification loss, a robustness loss function is used which effectively reduces the impact of noisy labels. The model is evaluated on 20 univariate time series from the UCR archive after noise processing. The results show that the model outperforms state-of-the-art methods in classifying incomplete time series under noisy labels, especially at high missing rates with high noise rates. Pengshuai Yao, Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Lili Guo 0001 |
CSCWD | 6 |
| 2024 | A novel image denoising algorithm combining attention mechanism and residual UNet network
Shifei Ding, Qidong Wang, Lili Guo 0001, Jian Zhang 0019, Ling Ding 0001 |
Knowl. Inf. Syst. | 3 |
| 2024 | Speaker-aware cognitive network with cross-modal attention for multimodal emotion recognition in conversation
Lili Guo 0001, Yikang Song, Shifei Ding |
Knowl. Based Syst. | 1 |
| 2024 | Robust Multi-Agent Communication With Graph Information Bottleneck OptimizationabstractRecent research on multi-agent reinforcement learning (MARL) has shown that action coordination of multi-agents can be significantly enhanced by introducing communication learning mechanisms. Meanwhile, graph neural network (GNN) provides a promising paradigm for communication learning of MARL. Under this paradigm, agents and communication channels can be regarded as nodes and edges in the graph, and agents can aggregate information from neighboring agents through GNN. However, this GNN-based communication paradigm is susceptible to adversarial attacks and noise perturbations, and how to achieve robust communication learning under perturbations has been largely neglected. To this end, this paper explores this problem and introduces a robust communication learning mechanism with graph information bottleneck optimization, which can optimally realize the robustness and effectiveness of communication learning. We introduce two information-theoretic regularizers to learn the minimal sufficient message representation for multi-agent communication. The regularizers aim at maximizing the mutual information (MI) between the message representation and action selection while minimizing the MI between the agent feature and message representation. Besides, we present a MARL framework that can integrate the proposed communication mechanism with existing value decomposition methods. Experimental results demonstrate that the proposed method is more robust and efficient than state-of-the-art GNN-based MARL methods. Shifei Ding, Wei Du 0010, Ling Ding 0001, Jian Zhang 0019, Lili Guo 0001, Bo An 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Better value estimation in Q-learning-based multi-agent reinforcement learning
Ling Ding 0001, Wei Du 0010, Jian Zhang 0019, Lili Guo 0001, Chenglong Zhang 0001, Di Jin 0001, Shifei Ding |
Soft Comput. | 4 |
| 2024 | Wavelet and Adaptive Coordinate Attention Guided Fine-Grained Residual Network for Image DenoisingabstractConvolutional neural networks (CNN) have achieved remarkable performance in image denoising. However, most existing CNNs cannot accurately capture and remove tiny noises during the denoising process and lose edge detail information easily. In this paper, we propose a fine-grained residual network guided by wavelet and adaptive coordinate attention (WACAFRN) for image denoising. Firstly, we propose an adaptive coordinate attention mechanism and combine it with cascaded Res2Net residual blocks to form an encoder network for more accurate noise removal. Secondly, we propose a wavelet attention mechanism that combines global and local residual blocks to form a decoder network, aiming to address the problem of edge detail information loss. At last, we complement the noise information through a noise estimation block to further enhance the model’s ability to adapt to noise. Extensive experiment results demonstrate that our proposed method outperforms existing denoising methods in both qualitative and quantitative aspects. Notably, our method significantly improves real-world noise removal tasks on the CC dataset, with an average increase of 2.08 dB in PSNR and 0.0264 in SSIM over the state-of-the-art methods. Additionally, WACAFRN exhibits faster inference speeds, underscoring its efficiency in real-world applications. Shifei Ding, Qidong Wang, Lili Guo 0001, Xuan Li 0004, Ling Ding 0001, Xindong Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Intuitionistic Fuzzy Stochastic Configuration Networks for Solving Binary Classification ProblemsabstractStochastic configuration network (SCN) is an emerging type of random neural network that allocates node parameters via supervised mechanisms to ensure universal approximation capability. However, outliers and noise in realworld data can adversely affect SCN's classification performance. To enhance SCN for binary classification, we introduce intuitionistic fuzzy set concepts to propose intuitionistic fuzzy SCN (IFSCN). Unlike SCN, IFSCN allocates an intuitionistic fuzzy number to every sample by computing degrees of membership and non-membership, thereby crafting an optimal classifier through strategic weighting of samples to mitigate the influence of noise. Furthermore, stochastic configuration sparse autoencoders (SCSAE) effectively learn sparse features using L1 regularization. By integrating multiple SC-SAE models, we extract robust sparse feature representations. We then propose hierarchical IFSCN (IFHSCN) built on SC-SAE for enhanced accuracy. Comprehensive experiments on 8 benchmark datasets demonstrate IFSCN and IFHSCN achieve superior binary classification performance over state-of-the-art models like intuitionistic fuzzy twin SVM, kernel ridge regression, random vector functional link networks. Overall, this study successfully equips SCN for real-world noisy data via intuitionistic fuzzy sets and sparsity, providing an effective and scalable solution for robust classification. Lili Guo 0001, Jianglan Zhu, Chenglong Zhang 0001, Shifei Ding |
IEEE Trans. Fuzzy Syst. | 1 |
| 2024 | Towards Faster Deep Graph Clustering via Efficient Graph Auto-EncoderabstractDeep graph clustering (DGC) has been a promising method for clustering graph data in recent years. However, existing research primarily focuses on optimizing clustering outcomes by improving the quality of embedded representations, resulting in slow-speed complex models. Additionally, these methods do not consider changes in node similarity and corresponding adjustments in the original structure during the iterative optimization process after updating node embeddings, which easily falls into the representation collapse issue. We introduce an Efficient Graph Auto-Encoder (EGAE) and a dynamic graph weight updating strategy to address these issues, forming the basis for our proposed Fast DGC (FastDGC) network. Specifically, we significantly reduce feature dimensions using a linear transformation that preserves the original node similarity. We then employ a single-layer graph convolutional filtering approximation to replace multiple layers of graph convolutional neural network, reducing computational complexity and parameter count. During iteration, we calculate the similarity between nodes using the linearly transformed features and periodically update the original graph structure to reduce edges with low similarity, thereby enhancing the learning of discriminative and cohesive representations. Theoretical analysis confirms that EGAE has lower computational complexity. Extensive experiments on standard datasets demonstrate that our proposed method improves clustering performance and achieves a speedup of 2–3 orders of magnitude compared to state-of-the-art methods, showcasing outstanding performance. The code for our model is available at https://github.com/Marigoldwu/FastDGC . Furthermore, we have organized a portion of the DGC code into a unified framework, available at https://github.com/Marigoldwu/A-Unified-Framework-for-Deep-Attribute-Graph-Clustering . Shifei Ding, Benyu Wu, Ling Ding 0001, Xiao Xu 0006, Lili Guo 0001, Hongmei Liao, Xindong Wu 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2024 | Graph-Based Semi-Supervised Deep Image Clustering With Adaptive Adjacency MatrixabstractImage clustering is a research hotspot in machine learning and computer vision. Existing graph-based semi-supervised deep clustering methods suffer from three problems: 1) because clustering uses only high-level features, the detailed information contained in shallow-level features is ignored; 2) most feature extraction networks employ the step odd convolutional kernel, which results in an uneven distribution of receptive field intensity; and 3) because the adjacency matrix is precomputed and fixed, it cannot adapt to changes in the relationship between samples. To solve the above problems, we propose a novel graph-based semi-supervised deep clustering method for image clustering. First, the parity cross-convolutional feature extraction and fusion module is used to extract high-quality image features. Then, the clustering constraint layer is designed to improve the clustering efficiency. And, the output layer is customized to achieve unsupervised regularization training. Finally, the adjacency matrix is inferred by actual network prediction. A graph-based regularization method is adopted for unsupervised training networks. Experimental results show that our method significantly outperforms state-of-the-art methods on USPS, MNIST, street view house numbers (SVHN), and fashion MNIST (FMNIST) datasets in terms of ACC, normalized mutual information (NMI), and ARI. Shifei Ding, Haiwei Hou, Xiao Xu 0006, Jian Zhang 0019, Lili Guo 0001, Ling Ding 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | SFEMGN: Image Denoising with Shallow Feature Enhancement Network and Multi-Scale ConvGRUabstractImage denoising methods based on convolutional neural networks have been popular and achieved relatively excellent performance. However, most of the existing methods cannot fully obtain and use the shallow feature information when removing noise, and cannot better combine information between various network layers. In this paper, we propose an image denoising algorithm based on a feature enhancement network and multi-scale convGRU, named a shallow feature enhancement and multi-scale convGRU denoising network (SFEMGN), through an in-depth study of convolutional networks and GRU networks. We first propose a feature enhancement block to extract richer shallow features and enhance the protection of image details. Furthermore, the proposed SFEMGN integrates a multi-scale convolution GRU module, which can combine spatial features and temporal features at the same time. Comparative experiments and ablation studies demonstrate that our proposed model can achieve competitive performance in both gray and color image denoising tasks. Qidong Wang, Lili Guo 0001, Shifei Ding, Jian Zhang 0019, Xiao Xu 0006 |
ICASSP | 2 |
| 2023 | FEMRNet: Feature-enhanced multi-scale residual network for image denoising
Xiao Xu 0006, Qidong Wang, Lili Guo 0001, Jian Zhang 0019, Shifei Ding |
Appl. Intell. | 3 |
| 2023 | Multi-agent dueling Q-learning with mean field and value decomposition
Shifei Ding, Wei Du 0010, Ling Ding 0001, Lili Guo 0001, Jian Zhang 0019, Bo An 0001 |
Pattern Recognit. | 4 |
| 2023 | A Sampling-Based Density Peaks Clustering Algorithm for Large-Scale Data
Shifei Ding, Chao Li 0102, Xiao Xu 0006, Ling Ding 0001, Jian Zhang 0019, Lili Guo 0001, Tianhao Shi |
Pattern Recognit. | 6 |
| 2023 | Graph clustering network with structure embedding enhanced
Shifei Ding, Benyu Wu, Xiao Xu 0006, Lili Guo 0001, Ling Ding 0001 |
Pattern Recognit. | 4 |
| 2023 | A novel capsule network based on deep routing and residual learning
Jian Zhang 0019, Qinghai Xu, Lili Guo 0001, Ling Ding 0001, Shifei Ding |
Soft Comput. | 3 |
| 2022 | An optimized twin support vector regression algorithm enhanced by ensemble empirical mode decomposition and gated recurrent unit
Shifei Ding, Zichen Zhang 0002, Lili Guo 0001 |
Inf. Sci. | 3 |
| 2022 | Value function factorization with dynamic weighting for deep multi-agent reinforcement learning
Wei Du 0010, Shifei Ding, Lili Guo 0001, Jian Zhang 0019, Chenglong Zhang 0001, Ling Ding 0001 |
Inf. Sci. | 3 |
| 2022 | Hypergraph regularized semi-supervised support vector machine
Shifei Ding, Lili Guo 0001, Zichen Zhang 0002 |
Inf. Sci. | 3 |
| 2022 | Low-degree term first in ResNet, its variants and the whole neural network family
Tongfeng Sun, Shifei Ding, Lili Guo 0001 |
Neural Networks | 3 |
| 2022 | Broad learning system based ensemble deep model
Chenglong Zhang 0001, Shifei Ding, Lili Guo 0001, Jian Zhang 0019 |
Soft Comput. | 3 |
| 2022 | Learning affective representations based on magnitude and dynamic relative phase information for speech emotion recognition
Lili Guo 0001, Longbiao Wang, Jianwu Dang 0001, Chng Eng Siong, Seiichi Nakagawa |
Speech Commun. | 1 |
| 2021 | Representation Learning with Spectro-Temporal-Channel Attention for Speech Emotion RecognitionabstractConvolutional neural network (CNN) is found to be effective in learning representation for speech emotion recognition. CNNs do not explicitly model the associations or relative importance of features in the spectral/temporal/channel-wise axes. In this paper, we propose an attention module, named spectro-temporal-channel (STC) attention module that is integrated with CNN to improve representation learning ability. Our module infers an attention map along the three dimensions, namely time, frequency, and CNN channel. Experiments are conducted on the IEMOCAP database to evaluate the effectiveness of the proposed representation learning method. The results demonstrate that the proposed method outperforms the traditional CNN method by an absolute increase of 3.13% in terms of F1 score. Lili Guo 0001, Longbiao Wang, Chenglin Xu, Jianwu Dang 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 1 |
| 2021 | Multimodal Emotion Recognition with Capsule Graph Convolutional Based Representation FusionabstractDue to the more robust characteristics compared to unimodal, audio-video multimodal emotion recognition (MER) has attracted a lot of attention. The efficiency of representation fusion algorithm often determines the performance of MER. Although there are many fusion algorithms, information redundancy and information complementarity are usually ignored. In this paper, we propose a novel representation fusion method, Capsule Graph Convolutional Network (CapsGCN). Firstly, after unimodal representation learning, the extracted audio and video representations are distilled by capsule network and encapsulated into multimodal capsules respectively. Multimodal capsules can effectively reduce data redundancy by the dynamic routing algorithm. Secondly, the multimodal capsules with their inter-relations and intra-relations are treated as a graph structure. The graph structure is learned by Graph Convolutional Network (GCN) to get hidden representation which is a good supplement for information complementarity. Finally, the multimodal capsules and hidden relational representation learned by CapsGCN are fed to multihead self-attention to balance the contributions of source representation and relational representation. To verify the performance, visualization of representation, the results of commonly used fusion methods, and ablation studies of the proposed CapsGCN are provided. Our proposed fusion method achieves 80.83% accuracy and 80.23% F1 score on eNTERFACE05’. Jiaxing Liu 0001, Longbiao Wang, Zhilei Liu, Yahui Fu 0001, Lili Guo 0001, Jianwu Dang 0001 |
ICASSP | 6 |
| 2021 | CONSK-GCN: Conversational Semantic- and Knowledge-Oriented Graph Convolutional Network for Multimodal Emotion RecognitionabstractEmotion recognition in conversations (ERC) has received significant attention in recent years due to its widespread applications in diverse areas, such as social media, health care, and artificial intelligence interactions. However, different from nonconversational text, it is particularly challenging to model the effective context-aware dependence for the task of ERC. To address this problem, we propose a new Conversational Semantic- and Knowledge-oriented Graph Convolutional Network (ConSK-GCN) approach that leverages both semantic dependence and commonsense knowledge. First, we construct the contextual inter-interaction and intradependence of the interlocutors via a conversational graph-based convolutional network based on multimodal representations. Second, we incorporate commonsense knowledge to guide ConSK-GCN to model the semantic-sensitive and knowledge-sensitive contextual dependence. The results of extensive experiments show that the proposed method outperforms the current state of the art on the IEMOCAP dataset. Yahui Fu 0001, Shogo Okada, Longbiao Wang, Lili Guo 0001, Yaodong Song, Jiaxing Liu 0001, Jianwu Dang 0001 |
ICME | 4 |
| 2021 | A Sentiment Similarity-Oriented Attention Model with Multi-task Learning for Text-Based Emotion Recognition
Yahui Fu 0001, Lili Guo 0001, Longbiao Wang, Zhilei Liu, Jiaxing Liu 0001, Jianwu Dang 0001 |
MMM (1) | 2 |
| 2021 | Robust unsupervised anomaly detection via multi-time scale DCGANs with forgetting mechanism for industrial multivariate time series
Lei Song 0011, Jianxing Wang, Lili Guo 0001, Xuzhi Li |
Neurocomputing | 4 |
| 2020 | Speech Emotion Recognition with Local-Global Aware Deep Representation LearningabstractConvolutional neural network (CNN) based deep representation learning methods for speech emotion recognition (SER) have demonstrated great success. The basic design of CNN restricts the ability to model only local information well. Capsule network (CapsNet) can overcome the shortages of CNNs to capture the shallow global features from the spectrogram, although CapsNet cannot learn the local and deep global information. In this paper, we propose a local-global aware deep representation learning system that mainly includes two modules. One module contains a multi-scale CNN, time- frequency CNN (TFCNN) to learn the local representation. In the other module, we introduce a structure with dense connections of multiple blocks to learn shallow and deep global information. Every block in this structure is a complete CapsNet improved by a new routing algorithm. The local and global representations are fed to the classifier and achieve an absolute increase of at least 4.25% than benchmarks on IEMOCAP. Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Lili Guo 0001, Jianwu Dang 0001 |
ICASSP | 4 |
| 2020 | Temporal Attention Convolutional Network for Speech Emotion Recognition with Latent Representation
Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Yuan Gao 0040, Lili Guo 0001, Jianwu Dang 0001 |
INTERSPEECH | 5 |
| 2020 | Speaker-Aware Speech Emotion Recognition by Fusing Amplitude and Phase Information
Lili Guo 0001, Longbiao Wang, Jianwu Dang 0001, Zhilei Liu, Haotian Guan |
MMM (1) | 1 |
| 2020 | An improvement growing neural gas method for online anomaly detection of aerospace payloads
Lei Song 0011, Taisheng Zheng, Jianxing Wang, Lili Guo 0001 |
Soft Comput. | 4 |
| 2019 | Time-Frequency Deep Representation Learning for Speech Emotion Recognition Integrating Self-attention
Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Lili Guo 0001, Jianwu Dang 0001 |
ICONIP (4) | 4 |
| 2018 | Gender-Aware CNN-BLSTM for Speech Emotion Recognition
Linjuan Zhang, Longbiao Wang, Jianwu Dang 0001, Lili Guo 0001, Qiang Yu 0005 |
ICANN (1) | 4 |
| 2018 | A Feature Fusion Method Based on Extreme Learning Machine for Speech Emotion RecognitionabstractSpeech emotion recognition is important to understand users' intention in human-computer interaction. However, it is a challenging task partly because we cannot clearly know which feature and model are effective to distinguish emotions. Previous studies utilize convolutional neural network (CNN) directly on spectrograms to extract features, and bidirectional long short term memory (BLSTM) is the state-of-the-art model. However, there are two problems of CNN-BLSTM. Firstly, it doesn't utilize heuristic features based on priori knowledge. Secondly, BLSTM has a complex structure and high complexity in training. To address the first problem, we propose a feature fusion method that combines CNN-based features and heuristic-based discriminative features which are extracted from heuristic features using deep neural network (DNN). In addition, we utilize extreme learning machine (ELM) instead of BLSTM to solve the second problem. The experiments conducted on EmoDB and our method leads to 40% relative error reduction in Fl-score compared to CNN-BLSTM. Lili Guo 0001, Longbiao Wang, Jianwu Dang 0001, Linjuan Zhang, Haotian Guan |
ICASSP | 1 |
| 2018 | Convolutional Neural Network with Spectrogram and Perceptual Features for Speech Emotion Recognition
Linjuan Zhang, Longbiao Wang, Jianwu Dang 0001, Lili Guo 0001, Haotian Guan |
ICONIP (4) | 4 |
| 2018 | Speech Emotion Recognition by Combining Amplitude and Phase Information Using Convolutional Neural Network
Lili Guo 0001, Longbiao Wang, Jianwu Dang 0001, Linjuan Zhang, Haotian Guan, Xiangang Li |
INTERSPEECH | 1 |
| 2017 | Extreme learning machine with kernel model based on deep learning
Shifei Ding, Lili Guo 0001, Yanlu Hou |
Neural Comput. Appl. | 2 |