VLDB 2026 Research / reviewers in the wild / expert
Ming Gao 0026
dblp:71/4173-26
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-0723-1197ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Source Patch Feature Fusion With Neighborhood Flash Attention Transformer for Pixel-Level Vehicle and Road Recognition in Hyperspectral ImageabstractHyperspectral imaging can capture the spectrum of each pixel in an image across various wavelengths, providing unparalleled opportunities for precise detection, classification, and analysis of transportation infrastructure. However, traditional methods often struggle with the curse of dimensionality, inter-class variability, and the spectral-spatial trade-off inherent in hyperspectral data. To address these challenges, we introduce a novel Multi-Source Patch Feature fusion based Neighborhood Flash Attention Transformer (MSPF-NFAT) for pixel-level vehicle and road recognition in hyperspectral images (HSIs). Our methodology hinges on the insight that the integration of complementary features from multiple sources and scales can significantly enhance classification performance. Specifically, the MSPF is designed to aggregate and harmonize features extracted from both spectral and spatial dimensions, as well as from different contextual scales within the image. This fusion process ensures a richer representation of the data, capturing both the fine-grained details and the broader contextual information essential for accurate classification. Building upon this enriched feature set, we employ the NFAT, a state-of-the-art attention mechanism that focuses on capturing local spatial relationships while efficiently scaling to accommodate the high-resolution characteristics of hyperspectral data. In addition, extensive experimental results on four widely used HSIs datasets show that our newly proposed method provides superior performance compared to other state-of-the-art methods. Weiwei Cai 0001, Pengjiang Qian, Chuang Wang 0011, Jian Yao 0005, Ming Gao 0026, E. Y. K. Ng |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | MMFNet: A Multi-modal and Multi-attention Fusion Network for Multi-task Skin Disease ClassificationabstractSkin diseases rank among the most prevalent ailments in humans, which underscores the critical importance of early detection and diagnosis. Considering clinical practice, the integration of information from various modalities holds significant potential to enhance diagnostic accuracy and precision. In this study, we propose a multi-modal and multi-attention fusion network for multi-task skin disease classification. We initially devise the Shifted Window Cross-attention Fusion (SWCF) module based on the shifted window self-attention mechanism, leveraging the advantages of the attention mechanism to learn the correlations between different modalities and fully integrate multi-modal features. Subsequently, taking into account the heterogeneity of the data, we employ the Heterogeneous Data Cross-attention Fusion (HDCF) module using the idea of information decoupling and separation processing to integrate image and text data. Additionally, we suggest the Dynamic Loss Weight Allocation (DLWA) method for multi-task learning to refine the training procedure. We confirm the superiority of the proposed method on the publicly available multi-modal skin lesion dataset, Derm7pt. The average accuracy of multi-modal skin disease classification is 79.14%, surpassing the current state-of-the-art methods. Xinlei Zhu, Pengjiang Qian, Chuang Wang 0011, Ming Gao 0026, Weiwei Cai 0001, Eddie Yin-Kwee Ng |
BIBM | 5 |
| 2023 | Exponential linear units-guided Depthwise separable convolution network with cross attention mechanism for hyperspectral image classificationabstractHyperspectral images (HSI) are more informative than other remote sensing techniques and hence frequently utilized in numerous domains. Convolutional neural network algorithm provides outstanding performance in image processing and currently has established as the main approach in the field of HSI classification . HSI contain numerous channels with various spatial and spectral feature information, and simultaneously contain massive redundant information, leading to dimensional explosion or gradient disappearance in the process of classification. A multilayer network model based on Exponential Linear Units-guided Depthwise Separable Convolution is proposed in this paper to extract both spectral and spatial features from HSI at scales ranging, and the feature maps are then fed into a cross-attention mechanism for weight allocation to improve local feature information and optimize computational resource allocation. Extensive experiments on three well-known hyperspectral datasets are conducted, and the results show that the proposed network model can accurately complete the required HSI classification operation and outperforms other well-established techniques in terms of computing efficiency. Ming Gao 0026, Pengjiang Qian |
Signal Process. | 1 |
| 2023 | Hierarchical Domain Adaptation Projective Dictionary Pair Learning Model for EEG Classification in IoMT SystemsabstractEpilepsy recognition based on electroencephalogram (EEG) and artificial intelligence technology is the main tool of health analysis and diagnosis in Internet of medical things (IoMT). As a distributed learning framework, federated learning can train a shared model from multiple independent edge nodes using local data, which has greatly promoted the development of IoMT. One of the main challenges of EEG-based epilepsy recognition in IoMT is that EEG records show varying distributions in different devices, different times, and different people. This nonstationary characteristic of EEG reduces the accuracy of the recognition model. To improve the classification performance in IoMT, a hierarchical domain adaptation projective dictionary pair learning (HDA-PDPL) model is developed in the study. HDA-PDPL integrates EEG signals from different domains (person, edge nodes, devices, etc.) into a set of hierarchical subspace and simultaneously learns synthesis and analysis dictionary pairs in each layer. Specifically, a nonlinear transform function is introduced to seek hierarchical feature projection. The domain adaptation term on sparse coding builds a connection between different domains. Thus, the shared synthesis and analysis dictionaries can encode domain-invariant representation and discrimination knowledge from different domains. Besides, the local preserved term of projective codes is introduced to capture the potential discriminative local structures of samples. The experimental results on two EEG epilepsy classifications verified that the HDA-PDPL model can outperform other comparisons by utilizing more shared knowledge of different domains. Weiwei Cai 0001, Ming Gao 0026, Yizhang Jiang, Xiaoqing Gu, Xin Ning 0001, Pengjiang Qian, Tongguang Ni |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | Stereo Attention Cross-Decoupling Fusion-Guided Federated Neural Learning for Hyperspectral Image ClassificationabstractFederated learning is a promising solution in several industries for co-training models among distributed clients via centralized servers without leaving private user data on the devices. Thus, federated learning can be seen as a stimulus for the edge computing paradigm as it supports collaborative learning and model optimization. In view of the strict requirements for data security and system reliability of hyperspectral classification techniques for surveillance, aerospace, and military missions, this paper proposes a novel stereo attention cross-decoupling fusion-guided federated neural learning algorithm for hyperspectral image classification, which first trains client devices using a scalable federated learning approach consisting of master server, secure aggregator and edge client devices of a certain size.The distributed devices train local models of the neural network for classifying hyperspectral images and send them to the secure aggregator, which aggregates the local models using a weighted averaging strategy and sends them to the master server for iteration. In addition, the stereo attention cross-decoupling fusion module is used to mine the multidimensional spatial details of the hyperspectral images, specifically by first extracting the most discriminative features from different directions (horizontal, vertical, and spatial) using the attention mechanism, and then using the decoupling fusion strategy to classify the original feature map into three levels: significant, minor, and redundant, and use them to model the multidimensional spatial relationships, thus strengthening the capability to represent features. Extensive experiments on several public datasets have shown that the proposed method provides competitive performance and, more importantly, is effective in enhancing privacy and reliability for hyperspectral image classification. Weiwei Cai 0001, Ming Gao 0026, Yao Ding 0010, Xin Ning 0001, Xiao Bai 0001, Pengjiang Qian |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Multi-Source Domain Transfer Discriminative Dictionary Learning Modeling for Electroencephalogram-Based Emotion RecognitionabstractCognitive computing is dedicated to researching a computing principle and method that can simulate the intelligence ability of human brain. Human emotion is the basic component of human cognitive activities. Electroencephalogram (EEG) computer signals obtained from a brain computer interface are difficult to conceal, and using machine learning methods to analyze EEG emotion is a hot topic in artificial intelligence. However, the EEG signal is non-stationary, making it difficult to select sufficient data from the same person to train a classifier for a subject. To promote the performance of emotion recognition methods, a multi-source domain transfer discriminative dictionary learning modeling (MDTDDL) is proposed in this study. The method integrates transfer learning and dictionary learning in a learning model, including the concepts of subspace learning, manifold smoothness, margin-based discriminant embedding, and large margin. The domain-specific transformation matrix projects EEG signals from various domains into the transfer subspace. The domain-invariant dictionary can find potential connections between multiple source domains and target domain. The manifold smoothness and margin-based discriminant embedding term further improve the model’s learning ability. The alternating optimization technique is used in model solving to efficiently compute model parameters. Experiments on the SEED and DEAP datasets demonstrate the effectiveness of MDTDDL. Xiaoqing Gu, Weiwei Cai 0001, Ming Gao 0026, Yizhang Jiang, Xin Ning 0001, Pengjiang Qian |
IEEE Trans. Comput. Soc. Syst. | 3 |