VLDB 2026 Research / reviewers in the wild / expert
Weiwei Cai 0001
dblp:73/7383-1 · also Wei-Wei Cai 0001
· DBLP profile ↗
24ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0001-8992-9999ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Knowledge Calibration Fusion and Label Space Graph Regularization-Based Multicenter Fuzzy SystemsabstractTraditional single-center learning algorithms often face significant limitations in handling heterogeneous data integration, including insufficient generalization ability, weak privacy protection, and difficulties adapting to multi-center scenarios. To address these challenges, multi-center learning has emerged as a critical technological framework. Although our previously proposed MKTC-R0T algorithm partially addressed the integration and modeling of multi-center data through the knowledge transfer calibration strategy, it still exhibits notable shortcomings in terms of knowledge fusion stability, model interpretability and generalization, as well as the utilization of complementary information across centers. To overcome these limitations, we propose a Knowledge Calibration Fusion and Label Space Graph Regularization-based Multi-center TSK Fuzzy System (KCF-LSG-MTSK). Specifically, we introduce an enhanced knowledge calibration and fusion strategy to effectively integrate heterogeneous information between the base center (BC) and auxiliary center (AC). We also propose a novel label space graph regularization scheme that constructs both intracenter and intercenter graph structures, leveraging data consistency and complementarity to enhance the quality of knowledge sharing. Furthermore, building upon firstorder TSK fuzzy system optimization, our approach incorporates a projected maximum mean discrepancy (PMMD) transfer term to effectively reduce data distribution discrepancies between the BC and AC. Experimental results on thirteen benchmark datasets demonstrate that KCFLSGMTSK achieves an average accuracy of 88.7%, significantly outperforming stateoftheart singlecenter and multicenter methods, thereby validating the superiority of our approach in heterogeneous data integration, knowledge transfer, and interpretable classification. Chuang Wang 0011, Pengjiang Qian, Weiwei Cai 0001, Jian Yao 0005, Yizhang Jiang, E. Y. K. Ng, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2025 | FefDM-Transformer: Dual-channel multi-stage Transformer-based encoding and fusion mode for infrared-visible images
Junwu Li, Yaomin Wang, Xin Ning 0001, Wenguang He, Weiwei Cai 0001 |
Expert Syst. Appl. | 5 |
| 2025 | Sorted Texture-Aware Glance and Gaze Network for Hyperspectral Image Classification With Low Training SamplesabstractHyperspectral images (HSI) provide a wealth of information surpassing human visual capabilities, enabling precise identification of remote sensing targets. However, it faces significant challenges, including insufficient long-range dependency modeling, difficulties in data collection, and the tendency of models to get trapped in local optima during training. To overcome these obstacles, we present the sorted texture-aware glance and gaze network (ST-GGNet) tailored for HSI classification. First, we propose the glance and gaze attention (GGA) mechanism, which employs feature interaction-based long-term modeling to minimize information loss across spectral bands and focus on critical land cover features within HSI. Subsequently, the sorted texture-aware module (STM) is introduced to deeply mine and efficiently utilizes detailed texture and spectral information, thereby enhancing accuracy even with limited training data. Additionally, we propose the budding growth optimization algorithm (BGO), which integrates a budding growth mechanism to help the model discover better solutions, boosting optimization and classification performance. Experimental evaluations conducted on four public HSI datasets—Pavia University, Salinas, Houston, and WHU-Longkou—demonstrate the superior performance of ST-GGNet compared to nine state-of-the-art (SOTA) classification methods. Specifically, under limited training samples, ST-GGNet achieves overall accuracies (OA) of 99.42%, 96.88%, 96.86%, and 97.74%; average accuracies (AA) of 98.90%, 98.01%, 97.07%, and 92.48%; and Kappa coefficients of 99.24%, 96.53%, 96.59%, and 97.03% respectively. The findings reveal that ST-GGNet not only maintains strong robustness and generalization but also effectively suppresses noise and excels at distinguishing spatially similar adjacent land covers, especially in low-samples scenarios, consistently outperforming existing SOTA methods. We have released our code and models at https://github.com/Pluviophile-sy/ST-GGNet. Taiyong Li, Jialei Zhan, Jialang Liu, Xuan Xiong, Weiwei Cai 0001, Exian Liu, Yingmei Wei, Yaowen Hu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Multi-Source Patch Feature Fusion With Neighborhood Flash Attention Transformer for Pixel-Level Vehicle and Road Recognition in Hyperspectral ImageabstractHyperspectral imaging can capture the spectrum of each pixel in an image across various wavelengths, providing unparalleled opportunities for precise detection, classification, and analysis of transportation infrastructure. However, traditional methods often struggle with the curse of dimensionality, inter-class variability, and the spectral-spatial trade-off inherent in hyperspectral data. To address these challenges, we introduce a novel Multi-Source Patch Feature fusion based Neighborhood Flash Attention Transformer (MSPF-NFAT) for pixel-level vehicle and road recognition in hyperspectral images (HSIs). Our methodology hinges on the insight that the integration of complementary features from multiple sources and scales can significantly enhance classification performance. Specifically, the MSPF is designed to aggregate and harmonize features extracted from both spectral and spatial dimensions, as well as from different contextual scales within the image. This fusion process ensures a richer representation of the data, capturing both the fine-grained details and the broader contextual information essential for accurate classification. Building upon this enriched feature set, we employ the NFAT, a state-of-the-art attention mechanism that focuses on capturing local spatial relationships while efficiently scaling to accommodate the high-resolution characteristics of hyperspectral data. In addition, extensive experimental results on four widely used HSIs datasets show that our newly proposed method provides superior performance compared to other state-of-the-art methods. Weiwei Cai 0001, Pengjiang Qian, Chuang Wang 0011, Jian Yao 0005, Ming Gao 0026, E. Y. K. Ng |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | MMFNet: A Multi-modal and Multi-attention Fusion Network for Multi-task Skin Disease ClassificationabstractSkin diseases rank among the most prevalent ailments in humans, which underscores the critical importance of early detection and diagnosis. Considering clinical practice, the integration of information from various modalities holds significant potential to enhance diagnostic accuracy and precision. In this study, we propose a multi-modal and multi-attention fusion network for multi-task skin disease classification. We initially devise the Shifted Window Cross-attention Fusion (SWCF) module based on the shifted window self-attention mechanism, leveraging the advantages of the attention mechanism to learn the correlations between different modalities and fully integrate multi-modal features. Subsequently, taking into account the heterogeneity of the data, we employ the Heterogeneous Data Cross-attention Fusion (HDCF) module using the idea of information decoupling and separation processing to integrate image and text data. Additionally, we suggest the Dynamic Loss Weight Allocation (DLWA) method for multi-task learning to refine the training procedure. We confirm the superiority of the proposed method on the publicly available multi-modal skin lesion dataset, Derm7pt. The average accuracy of multi-modal skin disease classification is 79.14%, surpassing the current state-of-the-art methods. Xinlei Zhu, Pengjiang Qian, Chuang Wang 0011, Ming Gao 0026, Weiwei Cai 0001, Eddie Yin-Kwee Ng |
BIBM | 6 |
| 2024 | Practical and secure multifactor authentication protocol for autonomous vehicles in 5GabstractAbstract Autonomous vehicles (AV) can not only improve traffic safety and congestion, but also have strategic significance for the development of the transportation industry. With the continuous updating of core technologies such as artificial intelligence, sensor detection, synchronous positioning, and high‐precision mapping, the development of AV has been promoted. When 5G network is combined with Internet of Vehicles, the problems of AV can be solved by taking advantage of 5G ultra‐large bandwidth, low latency and high reliability. However, when the user controls the vehicle remotely, a real‐time and reliable authentication process is needed, while minimizing the overhead of security protocols. Therefore, this article proposes a practical and secure multifactor user authentication protocol for AV in 5G network. By introducing non‐interactive zero‐knowledge proof technology and physical uncloning function, the protocol completes mutual authentication and key agreement without revealing any sensitive information. The article proves the security of the protocol through BAN logic and the simulation of Scyther. And it can resist malicious attacks and provide more security features. The informal security analysis shows that the protocol can meet the proposed security requirements. Finally, we evaluate the efficiency of the protocol, and the results show that the protocol can provide better performance. Junfeng Miao, Zhaoshun Wang, Xin Ning 0001, Weiwei Cai 0001, Ruimin Liu |
Softw. Pract. Exp. | 5 |
| 2024 | The Configurational Paths in BoP Rural E-Commerce Entrepreneurial Opportunity: A Fuzzy-Set Qualitative Comparative AnalysisabstractBottom of the pyramid (BoP) e-commerce entrepreneurship has emerged as a new phenomenon in China’s rural regions, but its origins have not been thoroughly explored. This study analyzed the shaping of entrepreneurial opportunity (EO) based on the entrepreneurial ecosystem theory as the study framework. A questionnaire survey was conducted among 213 rural e-commerce practitioners in Guangdong and Wuling Mountains, China, using the fuzzy-set qualitative comparative analysis approach to study the various factors and causal mechanisms affecting EOs. Our findings highlight the following: 1) there is no single necessary condition for high BoP rural e-commerce EO formation. However, low human capital leads to nonhigh BoP rural e-commerce EO formation; 2) the driving mechanisms for high BoP rural e-commerce EO are of three types (five paths): government-led, dual-interaction between government and potential subjects, and market-led; and 3) there is an asymmetry between the driving mechanisms of high and nonhigh entrepreneurship opportunity formation in rural e-commerce. The findings of this article contribute to the expansion of the entrepreneurial ecosystem theory’s application in rural e-commerce entrepreneurship and provide practical implications for the BoP population in effectively obtaining EOs in rural e-commerce. Lijuan Huang, Guojie Xie 0001, Weiwei Cai 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2024 | Multicenter Knowledge Transfer Calibration With Rapid Zeroth-Order TSK Fuzzy System for Small Sample Epileptic EEG SignalsabstractThe diagnosis and treatment of epilepsy necessitate the precise identification and classification of electroencephalogram (EEG) signals. However, EEG samples from different medical institutions often exhibit variability due to factors such as institutional characteristics, geographic locations, and the professional levels of physicians. This variability limits the widespread application of existing methods in small or single medical institutions, as they typically rely on large-scale and high-quality datasets. To rapidly assist small or single medical institutions in constructing models for the diagnosis and classification of epileptic EEG signals that are both highly generalizable and interpretable while ensuring patient privacy, this article proposes an innovative learning framework named Multicenter Knowledge Transfer Calibration with rapid zeroth-order TSK fuzzy system (MKTC-R0T). This method employs the zeroth-order TSK fuzzy system as the baseline model for each center and integrates a knowledge transfer calibration strategy within a multicenter learning framework, aiming to enhance the model's generalizability and classification accuracy in the face of inconsistent sample quality and sample heterogeneity. Specifically, MKTC-R0T first establishes a base center model in a large medical institution, and then, by imitating the forgetting mechanism of the human brain, a portion of the knowledge at the base center is randomly forgotten, while the remaining knowledge is utilized to assist the auxiliary centers in rapidly deploying models. Ultimately, through a knowledge integration strategy, all centers collectively guide the target center in building an efficient linear system for the diagnosis and classification of epileptic EEG signals. Extensive experiments conducted on 12 epilepsy EEG signal datasets have validated that MKTC-R0T outperforms other typical algorithms in terms of running time, deployment speed, rule complexity, and the model's generalization and robustness, which demonstrates the substantial potential of MKTC-R0T in the field of epilepsy EEG signal diagnosis and classification. Chuang Wang 0011, Pengjiang Qian, Zhihuang Wang, Weiwei Cai 0001, Jian Yao 0005, Yizhang Jiang, Xiangyu Yan |
IEEE Trans. Fuzzy Syst. | 4 |
| 2024 | S²GFormer: A Transformer and Graph Convolution Combining Framework for Hyperspectral Image ClassificationabstractTransformer-based methods have a great ability to model nonlocal interactions between spectral and spatial information, while the local features are easily ignored. Graph convolutional neural networks (GCNs) tend to do well in exploiting neighborhood vertex interactions based on their unique aggregation mechanism, while the ability to extract global information is limited. In this article, we study to comprehensively utilize the advantages of transformer and graph convolution by combining the two structures into a unified Transformer (Graphormer) to construct both local and global interactions for hyperspectral image (HSI) classification, and spatial–spectral features enhanced Graphormer framework (S2GFormer) is proposed. Specifically, a follow patch mechanism is first proposed to transform the pixel in HSI to patches while preserving the local spatial features and reducing the computational cost. Moreover, a patchwise spectral embedding block is designed to extract the spectral features of the patch, in which a neighborhood convolution is inserted for comprehensive spectral information extraction. Finally, a multilayer Graphormer Encoder module is proposed to extract the representative spatial–spectral features from the patch for HSI classification. In our network, we jointly integrate the three aforementioned parts into a unified network, and each component benefits the other. The experimental results demonstrate its suitability for HSI classification when compared with other state-of-the-art (SOTA) classifiers, particularly in scenarios with very limited labeled samples. The code of S2GFormer will be made publicly available at:https://github.com/DY-HYX. Yao Ding 0010, Aitao Yang, Shujun Yang, Yaoming Cai, Weiwei Cai 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | A review of research on co-trainingabstractSummary Co‐training algorithm is one of the main methods of semi‐supervised learning in machine learning, which explores the effective information in unlabeled data by multi‐learner collaboration. Based on the development of co‐training algorithm, the research work in recent years was further summarized in this article. In particular, three main steps of relevant co‐training algorithms are introduced: view acquisition, learners' differentiation, and label confidence estimation. Finally, we summarized the problems existing in the current co‐training methods, gave some suggestions for improvement, and looked forward to the future development direction of the co‐training algorithm. Xin Ning 0001, Shaohui Xu, Weiwei Cai 0001, Liping Zhang 0014, Wenfa Li |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | Identification of grape leaf diseases based on VN-BWT and Siamese DWOAM-DRNet
Chuang Cai, Weiwei Cai 0001, Yahui Hu, Liujun Li, Guoxiong Zhou |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | MDMASNet: A dual-task interactive semi-supervised remote sensing image segmentation methodabstractRemote sensing image (RSIs) segmentation is widely used in urban planning, natural disaster detection and many other fields. Compared with natural scene images, RSIs have higher resolution, complex imaging, and diverse object shapes and sizes, while semantic segmentation methods based on deep learning often require many data labels. In this paper, we propose a semi-supervised RSIs segmentation network with multi-scale deformable threshold feature extraction module and mixed attention (MDMANet). First, a pyramid ensemble structure is used, which incorporates deformable convolution and bole convolution, to extract features of objects with different shapes and sizes and reduce the influence of redundant features. Meanwhile, a mixed attention (MA) is proposed to aggregate long-range contextual relationships and fuse low-level features with high-level features. Second, an FCN-based full convolution discriminator task network is designed to help evaluate the feasibility of unlabeled image prediction results. We performed experimental validation on three datasets, and the results show that MDMANet segmentation provides more significant improvement in accuracy and better generalization than existing segmentation networks. Liangji Zhang, Zaichun Yang, Guoxiong Zhou, Aibin Chen, Yao Ding 0010, Liujun Li, Weiwei Cai 0001 |
Signal Process. | 9 |
| 2023 | Multi-Modality Fusion & Inductive Knowledge Transfer Underlying Non-Sparse Multi-Kernel Learning and Distribution AdaptionabstractWith the development of sensors, more and more multimodal data are accumulated, especially in biomedical and bioinformatics fields. Therefore, multimodal data analysis becomes very important and urgent. In this study, we combine multi-kernel learning and transfer learning, and propose a feature-level multi-modality fusion model with insufficient training samples. To be specific, we firstly extend kernel Ridge regression to its multi-kernel version under the lp-norm constraint to explore complementary patterns contained in multimodal data. Then we use marginal probability distribution adaption to minimize the distribution differences between the source domain and the target domain to solve the problem of insufficient training samples. Based on epilepsy EEG data provided by the University of Bonn, we construct 12 multi-modality & transfer scenarios to evaluate our model. Experimental results show that compared with baselines, our model performs better on most scenarios. Yuanpeng Zhang 0001, Kaijian Xia, Yizhang Jiang, Pengjiang Qian, Weiwei Cai 0001, Chengyu Qiu, Khin Wee Lai, Dongrui Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | Hierarchical Domain Adaptation Projective Dictionary Pair Learning Model for EEG Classification in IoMT SystemsabstractEpilepsy recognition based on electroencephalogram (EEG) and artificial intelligence technology is the main tool of health analysis and diagnosis in Internet of medical things (IoMT). As a distributed learning framework, federated learning can train a shared model from multiple independent edge nodes using local data, which has greatly promoted the development of IoMT. One of the main challenges of EEG-based epilepsy recognition in IoMT is that EEG records show varying distributions in different devices, different times, and different people. This nonstationary characteristic of EEG reduces the accuracy of the recognition model. To improve the classification performance in IoMT, a hierarchical domain adaptation projective dictionary pair learning (HDA-PDPL) model is developed in the study. HDA-PDPL integrates EEG signals from different domains (person, edge nodes, devices, etc.) into a set of hierarchical subspace and simultaneously learns synthesis and analysis dictionary pairs in each layer. Specifically, a nonlinear transform function is introduced to seek hierarchical feature projection. The domain adaptation term on sparse coding builds a connection between different domains. Thus, the shared synthesis and analysis dictionaries can encode domain-invariant representation and discrimination knowledge from different domains. Besides, the local preserved term of projective codes is introduced to capture the potential discriminative local structures of samples. The experimental results on two EEG epilepsy classifications verified that the HDA-PDPL model can outperform other comparisons by utilizing more shared knowledge of different domains. Weiwei Cai 0001, Ming Gao 0026, Yizhang Jiang, Xiaoqing Gu, Xin Ning 0001, Pengjiang Qian, Tongguang Ni |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Stereo Attention Cross-Decoupling Fusion-Guided Federated Neural Learning for Hyperspectral Image ClassificationabstractFederated learning is a promising solution in several industries for co-training models among distributed clients via centralized servers without leaving private user data on the devices. Thus, federated learning can be seen as a stimulus for the edge computing paradigm as it supports collaborative learning and model optimization. In view of the strict requirements for data security and system reliability of hyperspectral classification techniques for surveillance, aerospace, and military missions, this paper proposes a novel stereo attention cross-decoupling fusion-guided federated neural learning algorithm for hyperspectral image classification, which first trains client devices using a scalable federated learning approach consisting of master server, secure aggregator and edge client devices of a certain size.The distributed devices train local models of the neural network for classifying hyperspectral images and send them to the secure aggregator, which aggregates the local models using a weighted averaging strategy and sends them to the master server for iteration. In addition, the stereo attention cross-decoupling fusion module is used to mine the multidimensional spatial details of the hyperspectral images, specifically by first extracting the most discriminative features from different directions (horizontal, vertical, and spatial) using the attention mechanism, and then using the decoupling fusion strategy to classify the original feature map into three levels: significant, minor, and redundant, and use them to model the multidimensional spatial relationships, thus strengthening the capability to represent features. Extensive experiments on several public datasets have shown that the proposed method provides competitive performance and, more importantly, is effective in enhancing privacy and reliability for hyperspectral image classification. Weiwei Cai 0001, Ming Gao 0026, Yao Ding 0010, Xin Ning 0001, Xiao Bai 0001, Pengjiang Qian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | A Novel Hyperspectral Image Classification Model Using Bole Convolution With Three-Direction Attention Mechanism: Small Sample and Unbalanced LearningabstractCurrently, the use of rich spectral and spatial information of hyperspectral images (HSIs) to classify ground objects is a research hotspot. However, the classification ability of existing models is significantly affected by its high data dimensionality and massive information redundancy. Therefore, we focus on the elimination of redundant information and the mining of promising features and propose a novel Bole convolution (BC) neural network with a tandem three-direction attention (TDA) mechanism (BTA-Net) for the classification of HSI. A new BC is proposed for the first time in this algorithm, whose core idea is to enhance effective features and eliminate redundant features through feature punishment and reward strategies. Considering that traditional attention mechanisms often assign weights in a one-direction manner, leading to a loss of the relationship between the spectra, a novel three-direction (horizontal, vertical, and spatial directions) attention mechanism is proposed, and an addition strategy and a maximization strategy are used to jointly assign weights to improve the context sensitivity of spatial–spectral features. In addition, we also designed a tandem TDA mechanism module and combined it with a multiscale BC output to improve classification accuracy and stability even when training samples are small and unbalanced. We conducted scene classification experiments on four commonly used hyperspectral datasets to demonstrate the superiority of the proposed model. The proposed algorithm achieves competitive performance on small samples and unbalanced data, according to the results of comparison and ablation experiments. The source code for BTA-Net can be found athttps://github.com/vivitsai/BTA-Net. Weiwei Cai 0001, Xin Ning 0001, Guoxiong Zhou, Xiao Bai 0001, Yizhang Jiang, Wei Li 0032, Pengjiang Qian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Graph-Structured Convolution-Guided Continuous Context Threshold-Aware Networks for Hyperspectral Image ClassificationabstractAlthough convolutional neural networks (CNNs) have shown superior performance to traditional machine learning algorithms for hyperspectral image classification tasks, the ability of traditional CNNs to model remote dependencies in the spatial orientation of HSIs is still limited, and they always extract similar low-level features, leading to feature redundancy. To cope with this limitation, this paper proposes a novel multi-order statistical representation-guided graph convolution and continuous context threshold-aware network for the classification of hyperspectral images with limited training samples. Initially, the spectral spatial information is separately modeled using first-order features and second-order pooling operators. Secondly, we propose graph-structuring the patch’s features. By employing a random walk transition probability matrix, graph-structured convolution can mine more discriminative direction features. In addition, we design a continuous context threshold-aware network to model multidimensional spatial relationships, thereby enhancing the representation of graph features. Specifically, the cross-attention mechanism is used to calculate the attention weights in the vertical and horizontal directions, and the features are divided into two levels—important and secondary—by solving the cosine distance between feature vectors, and the former is retained and the latter is punished. Extensive experiments on multiple HSIs datasets demonstrated that the proposed method delivers competitive performance. The code will be available at: https://github.com/vivitsai/GSC-CCTA. Weiwei Cai 0001, Pengjiang Qian, Yao Ding 0010, Meiqiao Bi, Xin Ning 0001, Danfeng Hong, Xiao Bai 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | DGPF-RENet: A Low Data Dependence Network With Low Training Iterations for Hyperspectral Image ClassificationabstractThe classification of ground objects from hyperspectral images (HSIs) is of great importance for human perception of information about the terrain and landscape. HSIs have numerous dimensions, and obtaining the data is difficult. The issue of slow convergence of neural network training is brought on by high dimensional data, and the neural network’s performance is impacted by the challenging data acquisition process. In order to achieve the effects of low data dependence and rapid convergence, we propose a redundancy elimination network architecture with decoupled-gaze attention mechanism and phantom fractal modules (DGPF-RENet) for HSIs classification. First, we propose the decoupled-gaze attention mechanism (DGA) to make full use of correlation between adjacent bands and the continuity of neighboring pixels in HSIs. Then, a redundancy elimination module (REM) is proposed to reduce the number of feature points and eliminate redundant information while preserving the contextual information and relationships between pixels. Finally, the phantom fractal module (PFM) is proposed, which improves the scale of feature learning by fractalising convolutions at multiple scales. Four publicly available HSIs datasets, including Indian Pines, Salinas, DFC2018, and WHUHi-HongHu, were used in our experiments. According to experimental findings, when compared to other state-of-the-art methods, our method performs best with a small number of training samples and few iterations. We have released our code and models at https://github.com/yuhua666/DGPF-RENet. Jialei Zhan, Yaowen Hu, Guoxiong Zhou, Weiwei Cai 0001, Aibin Chen, Liu Xie, Maopeng Li, Liujun Li |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Multi-feature fusion: Graph neural network and CNN combining for hyperspectral image classification
Yao Ding 0010, Danfeng Hong, Chengguo Yu, Nengjun Yang, Weiwei Cai 0001 |
Neurocomputing | 8 |
| 2022 | Hybrid Dilated Convolution Guided Feature Filtering and Enhancement Strategy for Hyperspectral Image ClassificationabstractWith the increasing maturity of optics and photonics, hyperspectral technology has also greatly advanced. Hyperspectral images composed of hundreds of adjacent bands and containing useful information can be easily obtained. However, unlike ordinary remote sensing images, each sample in hyperspectral remote sensing images has high-dimensional features and contains rich spatial and spectral information, which greatly increases the difficulty of feature selection and mining, increases the computational complexity, and limits the recognition accuracy of the model. Therefore, in this letter, a novel hybrid dilated-convolution-guided feature filtering and enhancement strategy (HDCFE-Net) model is proposed to classify hyperspectral images. Dilated convolution can reduce the spatial feature loss without reducing the receptive field and can obtain distant features. It can also be combined with the traditional convolution without losing its original information. We propose a feature filtering and enhancement strategy that eliminates redundant features and reduces computational complexity. The core concept is to set a threshold feature value, like the rounding method, to filter and enhance features. Experiments on three well-known hyperspectral datasets—Indian Pines (IPs), Pavia University (PU), and Salinas—show that in less than 1% (IPs: 5%) of the training samples, the overall accuracy (OA) of our method is 77%, 89%, and 91%, respectively, which is superior to several well-known methods. The experiments demonstrated the effectiveness and superiority of HDCFE-Net. Runmin Liu, Weiwei Cai 0001, Guangjun Li, Xin Ning 0001, Yizhang Jiang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Multi-Source Domain Transfer Discriminative Dictionary Learning Modeling for Electroencephalogram-Based Emotion RecognitionabstractCognitive computing is dedicated to researching a computing principle and method that can simulate the intelligence ability of human brain. Human emotion is the basic component of human cognitive activities. Electroencephalogram (EEG) computer signals obtained from a brain computer interface are difficult to conceal, and using machine learning methods to analyze EEG emotion is a hot topic in artificial intelligence. However, the EEG signal is non-stationary, making it difficult to select sufficient data from the same person to train a classifier for a subject. To promote the performance of emotion recognition methods, a multi-source domain transfer discriminative dictionary learning modeling (MDTDDL) is proposed in this study. The method integrates transfer learning and dictionary learning in a learning model, including the concepts of subspace learning, manifold smoothness, margin-based discriminant embedding, and large margin. The domain-specific transformation matrix projects EEG signals from various domains into the transfer subspace. The domain-invariant dictionary can find potential connections between multiple source domains and target domain. The manifold smoothness and margin-based discriminant embedding term further improve the model’s learning ability. The alternating optimization technique is used in model solving to efficiently compute model parameters. Experiments on the SEED and DEAP datasets demonstrate the effectiveness of MDTDDL. Xiaoqing Gu, Weiwei Cai 0001, Ming Gao 0026, Yizhang Jiang, Xin Ning 0001, Pengjiang Qian |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2022 | Self-Supervised Locality Preserving Low-Pass Graph Convolutional Embedding for Large-Scale Hyperspectral Image ClusteringabstractDue to prior knowledge deficiency, large spectral variability, and high dimension of hyperspectral image (HSI), HSI clustering is extremally a fundamental but challenging task. Deep clustering methods have achieved remarkable success and have attracted increasing attention in unsupervised HSI classification (HSIC). However, the poor robustness, adaptability, and feature presentation limit their practical applications to complex large-scale HSI datasets. Thus, this article introduces a novel self-supervised locality preserving low-pass graph convolutional embedding method (L2GCC) for large-scale hyperspectral image clustering. Specifically, a spectral–spatial transformation HSI preprocessing mechanism is introduced to learn superpixel-level spectral–spatial features from HSI and reduce the number of graph nodes for subsequent network processing. In addition, locality preserving low-pass graph convolutional embedding autoencoder is proposed, in which the low-pass graph convolution and layerwise graph attention are designed to extract the smoother features and preserve layerwise locality features, respectively. Finally, we develop a self-training strategy, in which a self-training clustering objective employs soft labels to supervise the clustering process and obtain appropriate hidden representations for node clustering. L2GCC is an end-to-end training network, which is jointly optimized by graph reconstruction loss and self-training clustering loss. On Indian Pines, Salinas, and University of Houston 2013 datasets, the clustering accuracy overall accuracies (OAs) of the proposed L2GCC are 73.51%, 83.15%, and 64.12%, respectively. Yao Ding 0010, Yaoming Cai, Siye Li, Biao Deng, Weiwei Cai 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Unsupervised Self-Correlated Learning Smoothy Enhanced Locality Preserving Graph Convolution Embedding Clustering for Hyperspectral ImagesabstractHyperspectral image (HSI) clustering is an extremely fundamental but challenging task with no labeled samples. Deep clustering methods have attracted increasing attention and have achieved remarkable success in HSI classification. However, most existing clustering methods are ineffective for large-scale HSI, due to their poor robustness, adaptability, and feature presentation. In this paper, to address these issues, we introduce unsupervised self-correlated learning smoothy enhanced locality preserving graph convolution embedding clustering (S2LGCC) for large-scale HSI. Specifically, the spectral-spatial transformation is introduced to transform the original HSI into a graph while preserving the local spectral features and spatial structures. After that, a locality preserving graph convolutional embedding encoder is designed to learn the hidden representation from the graph, in which the deep layer-wise graph convolutional network (LGAT) is proposed to preserve the adaptive layer-wise locality features. In addition, the self-correlated learning smoothy module is developed to learn the smoothy information and the non-local relationship in the hidden representation space for clustering. Finally, a self-training strategy is proposed to cluster the graph node, in which a self-training clustering objective employs soft labels to supervise the clustering process. The proposed S2LGCC is jointly optimized by the fusion graph reconstruction loss and self-training clustering loss, and the two benefit each other. On IP, Salinas, and UH2013 datasets, the OAs of our S2LGCC are 71.76%, 82.61%, and 63.82%, respectively. Yao Ding 0010, Nengjun Yang, Haojie Hu, Xianxiang Huang, Weiwei Cai 0001 |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2021 | MC-Net: Multiple max-pooling integration module and cross multi-scale deconvolution network
Hongfeng You, Long Yu 0001, Shengwei Tian, Yan Xing 0004, Xin Ning 0001, Weiwei Cai 0001 |
Knowl. Based Syst. | 7 |