VLDB 2026 Research / reviewers in the wild / expert
Shuiping Gou
dblp:41/676
· DBLP profile ↗
74ranked-venue papers
17as first author
41since 2021 · last 2027
0000-0002-2619-6481ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 35 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 26 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Needle tracking for free-hand ultrasound-guided percutaneous liver tumor ablations
Ningtao Liu, Shuwei Xing, Derek W. Cool, Jing Yuan 0001, Luguang Huang, Shuiping Gou, Aaron Fenster |
Expert Syst. Appl. | 7 |
| 2026 | TIM++: Transductive Information Maximization for Few-Shot CLIPabstractTransductive Information Maximization (TIM) is a leading transductive few-shot learning method that maximizes the mutual information between query features and their predicted labels, while incorporating supervision from the support set. However, its potential remains underexplored, primarily due to the limited utilization of textual knowledge provided by vision-language models (VLMs) such as CLIP. To address this, we propose TIM++, an enhanced framework that incorporates both visual and textual information for few-shot CLIP adaptation. Specifically, TIM++ introduces a Kullback-Leibler (KL) divergence-based regularization term that encourages the model’s posterior predictions to align with CLIP’s zero-shot output distribution, especially focusing on the most confident predictions. Additionally, we develop an improved prototype initialization strategy that leverages both support and query features enriched with CLIP-guided semantics. Extensive experiments on 11 public datasets demonstrate that TIM++ consistently outperforms the standard TIM, achieving average accuracy gains of 19.25% and 10.88% in 1-shot and 2-shot settings, respectively. TIM++ also surpasses other existing state-of-the-art methods, establishing a new benchmark for few-shot learning with VLMs. Yingping Li, Yutong Zou, Yunshi Huang, Changzhe Jiao, Shen Peng, Zhang Guo 0001, Shuiping Gou |
AAAI | 8 |
| 2026 | Deep Hierarchical Knowledge Loss for Fault Intensity DiagnosisabstractFault intensity diagnosis (FID) plays a pivotal role in intelligent manufacturing while neglecting dependencies among target classes hinders its practical deployment. This paper introduces a novel and general framework with deep hierarchical knowledge loss (DHK) to achieve hierarchical consistent representation and prediction. We develop a novel hierarchical tree loss to enable a holistic mapping of same-attribute classes, leveraging tree-based positive and negative hierarchical knowledge constraints. We further design a focal hierarchical tree loss to enhance its extensibility and devise two adaptive weighting schemes based on tree height. In addition, we propose a group tree triplet loss with hierarchical dynamic margin by incorporating hierarchical group concepts and tree distance to model boundary structural knowledge across classes. The joint two losses significantly improve the recognition of subtle faults. Extensive experiments are performed on four real-world datasets from various industrial domains (three cavitation datasets from SAMSON AG and one publicly available dataset) for FID, all showing superior results and outperforming recent state-of-the-art FID methods. Yu Sha, Shuiping Gou, Bo Liu 0009, Ningtao Liu, Horst Stöcker, Domagoj Vnucec, Nadine Wetzstein, Andreas Widl, Kai Zhou 0017 |
KDD (1) | 2 |
| 2026 | Wavelet-based high-frequency fusion for multi-class segmentation of fecal pathological components in microscopic images
Nuo Tong, Shuiping Gou, Bianping Liang, Shaobin Deng, Jisheng Li, Mingxue Wang |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Mutual risk prompt learning with multi-objective optimization for collaborative tumor and peritumor segmentation
Nuo Tong, Qingyang Meng, Chunsheng Xu, Shuiping Gou, Mei Shi, Mengbin Li |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Interpretable refinement of medical foundation-model segmentation via visual thinking states and uncertainty gating
Bassam M. Kanber, Shuiping Gou, Bo Liu 0009, Naglaa F. Noaman, Magd Mukred, Ahmad Al Smadi |
Neurocomputing | 2 |
| 2026 | Uncertainty-guided vertex-parameter bidirectional refinement for hand pose and shape estimation
Shuiping Gou, Yalong Jiang, Yu Sha, Yingping Li |
Neurocomputing | 3 |
| 2025 | Cross-Rejective Open-Set SAR Image RegistrationabstractSynthetic Aperture Radar (SAR) image registration is an essential upstream task in geoscience applications, in which pre-detected keypoints from two images are employed as observed objects to seek matched-point pairs. In general, the registration is regarded as a typical closed-set classification, which forces each keypoint to be classified into the given classes, but ignoring an essential issue that numerous redundant keypoints are beyond the given classes, which unavoidably results in capturing incorrect matched-point pairs. Based on this, we propose a Cross-Rejective Open-set SAR Image Registration (CroR-OSIR) method. In this work, these redundant keypoints are regarded as out-of-distribution (OOD) samples, and we formulate the registration as a special open-set task with two modules: supervised contrastive feature-tuning and cross-rejective open-set recognition (CroR-OSR). Unlike traditional open-set recognition, all samples, including OOD samples, are available in the CroR-OSR module. CroR-OSR conducts the closed-set classifications in individual open-set domains from two images, meanwhile employing the cross-domain rejection during training, to exclude these OOD samples based on confidence and consistency. Moreover, a new supervised contrastive tuning strategy is incorporated for feature-tuning. Especially, the cross-domain estimation labels obtained by CroR-OSR are fed back to the feature-tuning module for feature-tuning, to enhance feature discriminability. The experimental results illustrate that the proposed method achieves more precise registration than the state-of-the-art methods. The code is released at https://github.com/XDyaoshi/CroR-OSIR-main. Shasha Mao, Shiming Lu, Zhaolong Du, Licheng Jiao, Shuiping Gou, Luntian Mou, Xuequan Lu |
CVPR | 5 |
| 2025 | TopicGeo: An Efficient Unified Framework for Geolocation
Shuiping Gou |
ICCV | 3 |
| 2025 | Adaptation and learning of spatio-temporal thresholds in spiking neural networks
Shuiping Gou, Peizhao Wang, Licheng Jiao, Zhang Guo 0001, Jisheng Li |
Neurocomputing | 2 |
| 2025 | GRU-TV: Time- and Velocity-aware Gated Recurrent Unit for patient representation
Ningtao Liu, Shuiping Gou, Ruoxi Gao, Binxiao Su, Claire Keun Sun Park, Shuwei Xing, Jing Yuan 0001, Aaron Fenster |
J. Biomed. Informatics | 2 |
| 2025 | Collaborative Knowledge Injection for Concealed Object Detection in MMW Human InspectionabstractMillimeter-wave body screening technology has gained significant attention at inspection sites due to its noncontact and safety. However, existing concealed object detection methods still face challenges in real-world security scenarios, especially the missed detection of dim-small objects posing security risks. The challenge stems primarily from insufficient discrimination of features between small concealed objects and backgrounds. To this end, we propose a collaborative knowledge injection detection network (CKID-Net). It injects the object semantic knowledge learned from an external object database into the concealed object detection model, which forces the model to push background representations apart from the object prior knowledge, and pull together concealed object representations and the prior knowledge, thereby improving the model's discrimination. Our method collaboratively learns representations of prior objects and objects to be detected via excavating their semantic relation. Experiments on active millimeter-wave (AMMW) and terahertz (THz) human datasets show that the CKID-Net outperforms state-of-the-art methods, especially on detection rate. Nuo Tong, Shuiping Gou, Shasha Mao |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Self-Supervised, Non-Contact Heartbeat Detection Based on Ballistocardiograms Utilizing Physiological Information GuidanceabstractBallistocardiograms (BCG) is a passive, non-contact heart rate detection technology that requires no action on the part of the individual. However, during the BCG signal acquisition process, the surface pressure generated by cardiac contraction is easily disturbed by external factors, and as people's health deteriorates, the j-peak (the main peak of the BCG signal) is no longer prominent. Our aim is to establish a non-contact, self-supervised heart rate detection method based on physiological information, to improve the accuracy and robustness of BCG heart rate detection under wider and more adverse conditions. The algorithm is guided by the heart rate estimation based on BCG itself, thereby reconstructing a signal with physiological significance. We also propose a heartbeat mapping algorithm based on Bidirectional Long Short-Term Memory Network (BiLSTM) for extracting global deep features, achieving real-time heartbeat prediction, and eliminating local deviations brought about by reconstruction. To verify the effectiveness of the proposed method, this paper evaluated 40 young subjects and 4 elderly subjects. Compared with the existing state-of-the-art methods, beat-to-beat heart rate estimation and heartbeat detection both performed excellently, surpassing most methods using precise labels. The experimental results show that the proposed method achieves effective heartbeat detection, demonstrating robustness and effectiveness in the face of unavoidable noise and variations. Changzhe Jiao, Aoyu Yang, Hantao Zhao, Ruhan Yi, Shuiping Gou, Yu Sha, Wanshun Wen, Licheng Jiao, Marjorie Skubic |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Multi-Task Hybrid Conv-Transformer With Emotional Localized Ambiguity Exploration for Facial Pain AssessmentabstractRecently, there has been significant progress in automatic pain assessment based on facial expression analysis. However, the performance of pain assessment remains unsatisfactory, due to a lack of analysis on local pain-related action units and emotional ambiguity. In particular, ambiguous pain expressions complicate the estimation of pain. It is argued that certain facial local regions related to pain should receive more attention while estimating pain intensities. Based on this, we propose a multi-task hybrid Conv-Transformer method for facial pain assessment, which utilizes the self-attention mechanism to explore facial local features related to pain intensities and constructs a multi-task joint optimizing module to mitigate facial emotional ambiguity. In particular, the proposed method modifies the network structure of the vision transformer model to better estimate continuous pain intensities. Meanwhile, a multi-task module is constructed to jointly optimize the classification and the regression tasks of pain assessment, which effectively regularizes the extracted features and facilitates a better fit of the regressed prediction to the given label. Finally, experimental results on the UNBC Pain dataset illustrate that the proposed method performs better with pain assessment compared with state-of-the-art methods. Shasha Mao, Angze Li, Yanjia Luo, Shuiping Gou, Mengnan Qi, Tianhuan Li, Xinyi Wei, Binxiao Su, Nan Gu |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | DASCE: Long-Tailed Data Augmentation Based Sparse Class-Correlation ExploitationabstractThe long-tailed data distribution frequently occurs in the real-world scenarios, whereas deep learning is not effective enough for such distribution. In order to improve the effectiveness for the long-tailed data, data augmentation is widely used to balance the distribution of classes by generating new samples. However, most existing studies are designed from the perspective of the class-independence assumption by default, ignoring the effect of interrelation among classes for data augmentation, which causes that some generated samples may be unrepresentative and useless for balancing the class-distribution. Inspired by this, we propose a new data augmentation method based the sparse class-correlation exploitation in this paper, which can generate more representative samples by utilizing the class-correlation, to effectively balance the class-distribution for the long-tailed data. In the proposed method, a sparse class-correlation exploration module is first proposed to explore the potential correlations among multiple classes for boosting the classification performance. Based on the class-correlations, the pivotal seed-samples are generated by maximizing the sparse representation of challenging samples. Meanwhile, an ambiguity-filtered translation module is designed to generate more representative new samples for the target classes based the obtained seed-samples by enhancing the class-consistency and suppressing the deviation from the target classes. In addition, we introduce the self-supervised feature and fuse it with the discriminative feature to explore more accurate class-correlations. Experimental results illustrate that the proposed method obtains better performance only with a small number of generated samples than the state-of-the-art methods. Mengnan Qi, Shasha Mao, Shuiping Gou, Licheng Jiao |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | PEPC-Net: Progressive Edge Perception and Completion Network for Precise Identification of Safe Resection Margins in Maxillofacial CystsabstractMaxillofacial cysts pose significant surgical risks due to their proximity to critical anatomical structures, such as blood vessels and nerves. Precise identification of the safe resection margins is essential for complete lesion removal while minimizing damage to surrounding at-risk tissues, which highly relies on accurate segmentation in CT images. However, due to the limited space and complex anatomical structures in the maxillofacial region, along with heterogeneous compositions of bone and soft tissues, accurate segmentation is extremely challenging. Thus, a Progressive Edge Perception and Completion Network (PEPC-Net) is presented in this study, which integrates three novel components: 1) Progressive Edge Perception Branch, which progressively fuses semantic features from multiple resolution levels in a dual-stream manner, enabling the model to handle the varying forms of maxillofacial cysts at different stages. 2) Edge Information Completion Module, which captures subtle, differentiated edge features from adjacent layers within the encoding blocks, providing more comprehensive edge information for identifying heterogeneous boundaries. 3) Edge-Aware Skip Connection to adaptively fuse multi-scale edge features, preserving detailed edge information, to facilitate precise identification of the cyst boundaries. Extensive experiments on clinically collected maxillofacial lesion datasets validate the effectiveness of the proposed PEPC-Net, achieving a DSC of 88.71% and an ASD of 0.489mm. It's generalizability is further assessed using an external validation set, which includes more diverse range of maxillofacial cyst cases and images of varying qualities. These experiments highlight the superior performance of PEPC-Net in delineating the polymorphic edges of heterogeneous lesions, which is critical for safe resection margins decision. Nuo Tong, Yuanlin Liu, Yueheng Ding, Lingnan Hou, Mei Shi, Shuiping Gou |
IEEE Trans. Medical Imaging | 8 |
| 2024 | Hierarchical Knowledge Guided Fault Intensity Diagnosis of Complex Industrial SystemsabstractFault intensity diagnosis (FID) plays a pivotal role in monitoring and maintaining mechanical devices within complex industrial systems.As current FID methods are based on chain of thought without considering dependencies among target classes.To capture and explore dependencies, we propose a hierarchical knowledge guided fault intensity diagnosis framework (HKG) inspired by the tree of thought, which is amenable to any representation learning methods.The HKG uses graph convolutional networks to map the hierarchical topological graph of class representations into a set of interdependent global hierarchical classifiers, where each node is denoted by word embeddings of a class.These global hierarchical classifiers are applied to learned deep features extracted by representation learning, allowing the entire model to be end-toend learnable.In addition, we develop a re-weighted hierarchical knowledge correlation matrix (Re-HKCM) scheme by embedding inter-class hierarchical knowledge into a data-driven statistical correlation matrix (SCM) which effectively guides the information sharing of nodes in graphical convolutional neural networks and avoids over-smoothing issues.The Re-HKCM is derived from the Yu Sha, Shuiping Gou, Bo Liu 0009, Johannes Faber, Ningtao Liu, Stefan Schramm, Horst Stöcker, Thomas Steckenreiter, Domagoj Vnucec, Nadine Wetzstein, Andreas Widl, Kai Zhou 0017 |
KDD | 2 |
| 2024 | Interpretable Matching of Optical-SAR Image via Dynamically Conditioned Diffusion ModelsabstractDriven by the complementary information fusion of optical and synthetic aperture radar (SAR) images, the optical-SAR image matching has drawn much attention. However, the significant radiometric differences between them imposes great challenges on accurate matching. Most existing approaches convert SAR and optical images into a shared feature space to perform the matching, but these methods often fail to achieve the robust matching since the feature spaces are unknown and uninterpretable. Motivated by the interpretable latent space of diffusion models, this paper formulates an optical-SAR image translation and matching framework via a dynamically conditioned diffusion model (DCDM) to achieve the interpretable and robust optical-SAR cross-modal image matching. Specifically, in the denoising process, to filter out outlier matching regions, a gated dynamic sparse cross-attention module is proposed to facilitate efficient and effective long-range interactions of multi-grained features between the cross-modal data. In addition, a spatial position consistency constraint is designed to promote the cross-attention features to perceive the spatial corresponding relation in different modalities, improving the matching precision. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods in terms of both the matching accuracy and the interpretability. Shuiping Gou, Yunzhi Chen |
ACM Multimedia | 1 |
| 2024 | Physically Informed Prior and Cross-Correlation Constraint for Fine-Grained Road Crack Segmentation
Shuiping Gou, Yunzhi Chen |
PRCV (3) | 2 |
| 2024 | Hierarchical cavitation intensity recognition using Sub-Master Transition Network-based acoustic signals in pipeline systems
Shuiping Gou, Yu Sha, Bo Liu 0009, Ningtao Liu, Johannes Faber, Stefan Schramm, Horst Stöcker, Thomas Steckenreiter, Domagoj Vnucec, Nadine Wetzstein, Andreas Widl, Kai Zhou 0017 |
Expert Syst. Appl. | 1 |
| 2024 | Joint-Wise Temporal Self-Similarity Periodic Selection Network for Repetitive Fitness Action CountingabstractAccurate repetitive action counting has crucial applications in the era of AI-assisted universal fitness. Existing methods are prone to large errors in spatially fine-grained action counting scenarios. In this study, we propose a joint-wise temporal self-similarity periodic selection network (JTSPS-Net) with a human skeleton as its input. Periodic knowledge is embedded in skeleton joint units and selected in a coarse-to-fine manner to focus on the temporal repetition that occurs in the local space. The proposed JTSPS-Net adopts a temporal multiscale fusion strategy to better handle videos with various lengths. To maintain the interpretability of the model, we design an impulse map regression module that uses one random frame per action unit as its labels. Furthermore, to fill the action counting gap in real physical fitness scenarios and to scale up the current repetition count dataset, we construct a high-quality dataset named FitnessRep, which consists of 2,110 fitness videos collected in realistic scenarios. Experiments demonstrate that the proposed JTSPS-Net outperforms the state-of-the-art approach on our dataset and two other public datasets, especially on fine-grained action samples. In addition, it has a good ability to generalize to repetitive actions belonging to unseen categories. Shuiping Gou, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Fully Automatic Fine-Grained Grading of Lumbar Intervertebral Disc Degeneration Using Regional Feature Recalibration NetworkabstractAccurate fine-grained grading of lumbar intervertebral disc (LIVD) degeneration is essential for the diagnosis and treatment design of high-incidence low back pain. However, the grading accuracy is still challenged by lacking the fine-grained degenerative details, which is mainly due to the existing grading methods are easily dominated by the salient nucleus pulposus regions in LIVD, overlooking the inconspicuous degeneration changes of the surrounding structures. In this study, a novel regional feature recalibration network (RFRecNet) is proposed to achieve accurate and reliable LIVD degeneration grading. Detection transformer (DETR) is first utilized to detect all LIVDs and then input to the proposed RFRecNet for the fine-grained grading. To obtain sufficient features from both the salient nucleus pulposus and the surrounding regions, a regional cube-based feature boosting and suppression (RC-FBS) module is designed to adaptively recalibrate the feature extraction and utilization from the various regions in LIVD, and a feature diversification (FD) module is proposed to capture the complementary semantic information from the multi-scale features for the comprehensive fine-grained degeneration grading. Extensive experiments were conducted on a clinically collected dataset, which consists of 500 MR scans with a total of 10225 LIVDs. An average grading accuracy of 90.5%, specificity of 97.5%, sensitivity of 90.8%, and Cohen's kappa correlation coefficient of 0.876 are obtained, which indicate that the proposed framework is promising to provide doctors with reliable and consistent fine-grained quantitative evaluation results of the LIVD degeneration conditions for the optimal surgical plan design. Nuo Tong, Shuiping Gou, Bo Liu 0009, Yufeng Bai, Jingzhong Liu, Tan Ding |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Multiscale Cross-Modal Homogeneity Enhancement and Confidence-Aware Fusion for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection has shown many advantages in a variety of environments, particularly poor illumination conditions, by leveraging visible-thermal modalities. However, in-depth insight into distinguishing the complementary content of multimodal data and exploring the extent of multimodal feature fusion is still lacking. In this paper, we propose a novel multispectral pedestrian detector with multiscale cross-modal homogeneity enhancement and confidence-aware feature fusion. RGB and thermal streams are constructed to extract features and generate candidate proposals. During feature extraction, multiscale cross-modal homogeneity enhancement is proposed to enhance single-modal features using the separated homogeneous features via modal interactions. Homogeneity features encode the semantic information of the scene and are extracted from the RGB-thermal pairs by employing a channel attention mechanism. Proposals from two modalities are united to obtain multimodal proposals. Then, confidence measurement fusion is proposed to achieve multispectral feature fusion in each proposal by measuring the internal confidence of each modality and the interaction confidence between modalities. In addition, a confidence transfer loss function is designed to focus more on hard-to-detect samples during training. Experimental results on two challenging datasets demonstrate that the proposed method achieves better performance compared to existing methods. Jiajun Xiang, Feixiang Sun, Longwu Yuan, Shuiping Gou |
IEEE Trans. Multim. | 6 |
| 2024 | SPMHand: Segmentation-Guided Progressive Multi-Path 3D Hand Pose and Shape EstimationabstractHand pose and shape estimation plays an important role in numerous applications. A cost-effective and practical-friendly approach is to perform accurate hand estimation from a single RGB image, but this task is challenging due to ubiquitous hand self-occlusion and hand-object interaction occlusions. In this paper, we propose a novel SPMHand network to alleviate the effect of occlusions, inspired by the process that humans infer the whole hand when the hand is occluded. The proposed SPMHand consists of two main modules to generate hand segmentations as guidance and conduct hand regressions in a progressive multi-path manner. The segmentation-guided deocclusion module enables the network to “see” the occluded hand by inferring the whole hand segmentation. Specifically, the visible hand segmentation is first obtained and then a hand morphology attention block is introduced to infer the whole hand segmentation by fusing the visible information with the learned hand features. The progressive multi-path regression module is designed to gradually regress the fine hand with intermediate supervisions. Features from deep to shallow are utilized for the hand regressions from coarse to decent. Subsequently, the structure feature, joint heatmaps and segmentations that provide guidance for deocclusion are embedded and fused for the final fine hand regression. Experiments on four challenging datasets illustrate that the proposed SPMHand outperforms the state-of-the-arts in both real-world and synthetic scenes, especially in the present of severe hand-object occlusions. Shuiping Gou |
IEEE Trans. Multim. | 2 |
| 2023 | A CAM-Enhancing Generative Person Re-ID Method Based Global and Local FeaturesabstractFor GAN-based Person Re-identification (Re-ID), the key is to generate pedestrian images with higher identity consistency and meanwhile larger intra-class diversity. Generally, the main discriminative parts focus on some local regions from the foreground of each pedestrian image for Re-ID, and they should be irrelevant to the background. Whereas, most existing methods generate pedestrian images only based on global features, which difficultly achieves emphasizing crucial local regions and weakening the background. Based on this, we propose a CAM-enhancing generative Re-ID method in which the global and local features are jointly used. In the proposed method, an adaptive CAM-enhancing local encoder is designed to explore the significance of local appearances and enhance the effect of crucial local features in generations, where the foreground is divided into multiple local parts and separated from the background by pedestrian segmentation. Moreover, a new generation loss is proposed to supervise the identity consistency by reducing the inconsistency of crucial regions in foregrounds and meanwhile enrich the intra-class diversity by generating variant backgrounds. Experimental results indicate that the proposed method obtains better generation images and Re-ID performance than other methods. Angze Li, Shasha Mao, Mengnan Qi, Shuiping Gou, Licheng Jiao |
ICIP | 5 |
| 2023 | Feature Adversarial Network for Multimodal Template MatchingabstractRecently generative adversarial network (GAN) has been explored to multimodal template matching. Existing GAN-based multimodal template matching methods exploit the image generation to transform the multimodal template matching task as the unimodal one. However, image synthesis-based multimodal template matching methods relay on the generated image quality, which is unstable. Towards this end, this paper proposes a feature adversarial network, which maps different modal images into a common subspace and learns the correlation in the subspace. Specifically, the feature mapper is designed to map the multimodal features into intermediate features, and the modality discriminator is proposed to optimize the multimodal intermediate features until they are indistinguishable. Thus, an effective common feature subspace is generated for correlation learning. The experimental results on a public dataset demonstrate the superiority of the proposed method. Xiushe Zhang, Chunlei Han, Jinming Mu, Shuiping Gou |
IGARSS | 5 |
| 2023 | RGMIL: Guide Your Multiple-Instance Learning Model with RegressorabstractIn video analysis, an important challenge is insufficient annotated data due to the rare occurrence of the critical patterns, and we need to provide discriminative frame-level representation with limited annotation in some applications. Multiple Instance Learning (MIL) is suitable for this scenario. However, many MIL models paid attention to analyzing the relationships between instance representations and aggregating them, but neglecting the critical information from the MIL problem itself, which causes difficultly achieving ideal instance-level performance compared with the supervised model.
To address this issue, we propose the $\textbf{\textit{Regressor-Guided MIL network} (RGMIL)}$, which effectively produces discriminative instance-level representations in a general multi-classification scenario. In the proposed method, we make full use of the $\textit{regressor}$ through our newly introduced $\textit{aggregator}$, $\textbf{\textit{Regressor-Guided Pooling} (RGP)}$. RGP focuses on simulating the correct inference process of humans while facing similar problems without introducing new parameters, and the MIL problem can be accurately described through the critical information from the $\textit{regressor}$ in our method.
In experiments, RGP shows dominance on more than 20 MIL benchmark datasets, with the average bag-level classification accuracy close to 1.
We also perform a series of comprehensive experiments on the MMNIST dataset. Experimental results illustrate that our $\textit{aggregator}$ outperforms existing methods under different challenging circumstances. Instance-level predictions are even possible under the guidance of RGP information table in a long sequence. RGMIL also presents comparable instance-level performance with S-O-T-A supervised models in complicated applications. Statistical results demonstrate the assumption that a MIL model can compete with a supervised model at the instance level, as long as a structure that accurately describes the MIL problem is provided. The codes are available on $\url{https://github.com/LMBDA-design/RGMIL}$. Zhaolong Du, Shasha Mao, Shuiping Gou, Licheng Jiao |
NeurIPS | 4 |
| 2023 | Weakly-Supervised Semantic Feature Refinement Network for MMW Concealed Object DetectionabstractThe concealed object detection in millimeter-wave human body images is a challenging task due to the noise and dim-small objects. Exploiting the spatial dependencies to mine the difference between the object and the noise is vital for the discrimination of objects. However, most approaches ignore the context around the object. In this paper, a concealed object detection framework based on structural context is proposed to suppress noise interference and refine localizable semantic features. The framework consists of two subnetworks, structural region-based multi-scale weakly supervised feature refinement and local context-based concealed object detection. The multi-scale weakly supervised feature refinement is constructed to learn position-aware semantics of objects of various sizes while suppressing background noises in structural regions. Specifically, a multi-scale pooling method is proposed to better localize objects of different sizes, and an object-activated region enhancement module is designed to strengthen object semantic representations and suppress the background interference. Moreover, an adaptive local context aggregation module is designed to integrate the local context around the bounding box in the concealed object detection, which improves the discrimination of the model for the dim-small objects. Experimental results on the AMMW and the PMMW datasets demonstrate that the proposed approach improves detection performance with lower false alarm rates. Shuiping Gou, Shasha Mao, Licheng Jiao, Yinghai Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | ₁ Sparsity-Regularized Attention Multiple-Instance Network for Hyperspectral Target DetectionabstractAttention-based deep multiple-instance learning (MIL) has been applied to many machine-learning tasks with imprecise training labels. It is also appealing in hyperspectral target detection, which only requires the label of an area containing some targets, relaxing the effort of labeling the individual pixel in the scene. This article proposes an L1 sparsity-regularized attention multiple-instance neural network (L1-attention MINN) for hyperspectral target detection with imprecise labels that enforces the discrimination of false-positive instances from positively labeled bags. The sparsity constraint applied to the attention estimated for the positive training bags strictly complies with the definition of MIL and maintains better discriminative ability. The proposed algorithm has been evaluated on both simulated and real-field hyperspectral (subpixel) target detection tasks, where advanced performance has been achieved over the state-of-the-art comparisons, showing the effectiveness of the proposed method for target detection from imprecisely labeled hyperspectral data. Changzhe Jiao, Chao Chen 0040, Shuiping Gou, Xiuxiu Wang, Bo Yang 0047, Licheng Jiao |
IEEE Trans. Cybern. | 3 |
| 2023 | Adaptive Self-Supervised SAR Image Registration With Modifications of Alignment TransformationabstractConsidering that deep learning achieves the prominent performance, it has been applied to synthetic aperture radar (SAR) image registration to improve the registration accuracy. In most methods, a deep registration model is constructed to classify matched points and unmatched points, in which SAR image registration is regarded as a supervised two-classification problem. However, it is difficult to annotate massive matched points manually in practice, which limits the performance of deep networks. Besides, inevitable differences among SAR images easily cause that some training and testing samples are inconsistent, which probably brings negative effects for training a robust registration model. To address these problems, we propose an adaptive self-supervised SAR image registration method, where SAR image registration is regarded as a self-supervised task rather than the supervised two-classification task. Inspired by self-supervised learning, we consider each point on SAR images as a category-independent instance, which mitigates the requirement of manual annotations. Based on key points from images, a self-supervised model is constructed to explore the latent feature of each key point, and then, pairs of match points are sought via evaluating similarities among key points and used to calculate the alignment transformation matrix. Meanwhile, to enhance the consistency of samples, we design a new strategy that constructs multiscale samples by transforming key points from one image into another, which avoids inevitable diversities between two images effectively. In particular, the constructed samples feeding to the self-supervised model are adaptively updated with the modification of the transformation matrix in iterations. Moreover, the similarity of maximal public areas (MPAS) indicator is proposed to assist in estimating the transformation. Finally, experimental results illustrate that the proposed method achieves more accurate registrations than other compared methods. Shasha Mao, Jinyuan Yang, Shuiping Gou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Siamese Network with Feature Alignment Method for Non-Homologous SAR-ATRabstractMost of the existing synthetic aperture radar(SAR) automatic target recognition(ATR) algorithm are based on data-driven. However, there is not enough data for some specific target recognition to be trained. In this paper, a siamese network with parameter sharing is created, and then the simulated and real SAR images of the available targets are used as sample pair inputs, and labeled as positive and negative sample pairs based on whether the input sample pairs are of the same category, so that the network can be trained to extract domain invariant features, and then use classifiers to achieve the recognition task of non-homologous targets. The proposed method is validated on the moving and stationary target acquisition and recognition (MSTAR) dataset, The results are that the accuracy of on trained by ten pairs simulation and real SAR images are higher 93.92%, and the accuracy trained by on only one reaches 81.17%. Haiyang Ren, Yong Qiang, Baozhu Liu, Shuiping Gou |
IGARSS | 4 |
| 2022 | RGB-Thermal based Pedestrian Detection with Single-Modal Augmentation and ROI Pooling Multiscale FusionabstractRGB-Thermal based pedestrian detection has received more extensive attention due to the provided detailed information and thermal sensitivity of pedestrians. In this paper, a single-modal feature augmentation network (SMA-Net) is proposed. Firstly, two single-modal branches are trained separately to optimize the feature extraction of each branch in addition to the training of pedestrian detection based on fused features. Secondly, a lightweight ROI pooling multiscale fusion module (PMSF) is proposed to obtain more fine-grained and abundant features, in which pooling features of different scales are integrated by adaptively weighting. Finally, a generative constraint strategy is designed to constrain fusion by minimizing the loss function between the generated fusion image and RGB-Thermal pairs. Experimental result on the challenging dataset KAIST demonstrates that the proposed SMA-Net achieves great performance in terms of accuracy and computational efficiency. Jiajun Xiang, Shuiping Gou, Zhihui Zheng |
IGARSS | 2 |
| 2022 | Regional-Local Adversarially Learned One-Class Classifier Anomalous Sound Detection in Global Long-Term SpaceabstractAnomalous sound detection (ASD) is one of the most significant tasks of mechanical equipment monitoring and maintaining in complex industrial systems. In practice, it is vital to efficiently identify abnormal status of the working mechanical system, which can further facilitate the failure troubleshooting. In this paper, we propose a multi-pattern adversarial learning one-class classification framework, which allows us to use both the generator and the discriminator of an adversarial model for efficient ASD. The core idea is to learn reconstructing the normal patterns of acoustic data through two different patterns from auto-encoding generators, which succeeds in generalizing the fundamental role of a discriminator from identifying real and fake data to distinguishing between regional and local pattern reconstructions. Moreover, we design a novel balanceable detection strategy using both generators and a discriminator to achieve anomaly detection efficiently. Furthermore, we present a global filter layer for long-term interactions in the frequency domain space, which directly learns from the original data without introducing any human priors. Extensive experiments are performed on four real-world datasets from different industrial domains (three cavitation datasets from SAMSON AG, and one existing publicly) for anomaly detection, all showing superior results and outperform recent state-of-the-art ASD methods. Yu Sha, Shuiping Gou, Johannes Faber, Bo Liu 0009, Stefan Schramm, Horst Stöcker, Thomas Steckenreiter, Domagoj Vnucec, Nadine Wetzstein, Andreas Widl, Kai Zhou 0017 |
KDD | 2 |
| 2022 | A multi-task learning for cavitation detection and cavitation intensity recognition of valve acoustic signals
Yu Sha, Johannes Faber, Shuiping Gou, Bo Liu 0009, Stefan Schramm, Horst Stöcker, Thomas Steckenreiter, Domagoj Vnucec, Nadine Wetzstein, Andreas Widl, Kai Zhou 0017 |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | Cross-Connected Bidirectional Pyramid Network for Infrared Small-Dim Target DetectionabstractInfrared small-dim target detection is an important technology in the fields of infrared guidance, anti-missile, and tracking system. Due to the small size of targets, no obvious structure information, and low image signal-to-noise ratio (SNR), infrared small-dim target detection is still a challenging task. In this letter, a cross-connected bidirectional pyramid network (CBP-Net) is proposed for infrared small-dim target detection. The main body of the CBP-Net is to embed a bottom-up pyramid in the feature pyramid network (FPN), which is designed to provide more comprehensive target information by connecting with the original multi-scale features and the top-down pyramid. The bottom-up pyramid together with the top-down pyramid forms the proposed bidirectional pyramid structure. Then, an region of interest (ROI) feature augment module (RFA) composed of deformable ROI pooling and position attention is designed to fuse multi-scale ROI features and enhance the spatial information of the small-dim target. Besides, a regular constraint loss (RCL) is introduced to restrict multi-scale feature fusion to learn more precise target location information. Experimental results on two challenging datasets show that the performance of the proposed CBP-Net is superior to the state-of-the-art methods. Yuanning Bai, Shuiping Gou, Yaohong Chen, Zhihui Zheng |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Ship Segmentation via Encoder-Decoder Network With Global Attention in High-Resolution SAR ImagesabstractShip detection in the synthetic aperture radar (SAR) image is of great significance in the fields of military and coastal defense. Most ship detection methods are designed based on the object detection framework, which can only provide the vertices’ coordinates of the bounding box covering the ship targets but cannot provide more detailed contour information. Target segmentation can further explore the shape and edge information of the objects, which can be used as a blazing novel means for automatic object detection. In this letter, a 3-D atrous encoder–decoder neural network with global attention modules (GAM-EDNet) is proposed to achieve ship segmentation in SAR images. The encoder–decoder structure with atrous convolution is developed as the network body to fully exploit the structural information of the ship targets with various sizes. To increase the structural information of the single-polarization SAR images, a 3-D image cube is designed as the input of the GAM-EDNet. A global attention module is proposed to further improve the segmentation performance by integrating the high-level semantic features with the low-level location features. Besides, an SAR ship segmentation dataset (SAR-HR4) is built to evaluate the segmentation performance, and the experimental results show that the proposed GAM-EDNet achieves better performance than other state-of-the-art methods. Jichao Li 0003, Shuiping Gou, Jiawei Chen 0001, Xiaolong Sun |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Fast and Low-GPU-memory abdomen CT organ segmentation: The FLARE challengeabstractAutomatic segmentation of abdominal organs in CT scans plays an important role in clinical practice. However, most existing benchmarks and datasets only focus on segmentation accuracy, while the model efficiency and its accuracy on the testing cases from different medical centers have not been evaluated. To comprehensively benchmark abdominal organ segmentation methods, we organized the first Fast and Low GPU memory Abdominal oRgan sEgmentation (FLARE) challenge, where the segmentation methods were encouraged to achieve high accuracy on the testing cases from different medical centers, fast inference speed, and low GPU memory consumption, simultaneously. The winning method surpassed the existing state-of-the-art method, achieving a 19× faster inference speed and reducing the GPU memory consumption by 60% with comparable accuracy. We provide a summary of the top methods, make their code and Docker containers publicly available, and give practical suggestions on building accurate and efficient abdominal organ segmentation models. The FLARE challenge remains open for future submissions through a live platform for benchmarking further methodology developments at https://flare.grand-challenge.org/. Jun Ma 0016, Yao Zhang 0010, Song Gu, Xingle An, Zhihe Wang, Cheng Ge, Yinan Xu 0004, Shuiping Gou, Franz Thaler, Christian Payer, Darko Stern, Edward G. A. Henderson, Dónal M. McSweeney, Andrew Green 0001, Price Jackson, Lachlan McIntosh, Quoc-Cuong Nguyen, Abdul Qayyum 0002, Pierre-Henri Conze, Ziyan Huang, Deng-Ping Fan, Huan Xiong, Guoqiang Dong, Qiongjie Zhu, Xiaoping Yang 0001 |
Medical Image Anal. | 11 |
| 2022 | Self-Paced Feature Attention Fusion Network for Concealed Object Detection in Millimeter-Wave ImageabstractThe active millimeter-wave (AMMW) scanner has been widely used for inspecting human security in public places in recent years owing to its ability to detect all kinds of objects under the clothes and be harmless to the body. However, it is really challenging to detect all concealed objects automatically and accurately due to inherent imaging noise, unknown object kind, and uncertain position. Recently, many existing methods, especially deep learning-based, have achieved good performances on concealed object detection. These methods work well for detecting a few kinds of large objects, but fail to perform on dim and incomplete hard objects. To address this task, a concealed object detection model with self-paced feature attention fusion network (SPFAFN) is proposed in this article. To be specific, the features with different scales are fused in a top-down manner to integrate details and global semantics to better detect small objects. During fusing multi-scale features, a hierarchical pyramid attention mechanism composed of channel and spatial attention is developed to perceive the object. Moreover, boosting self-paced learning is exploited to guide the model to learn hard samples that are difficultly detected. The proposed method is validated on two real-world datasets: an AMMW dataset and a publicly available passive millimeter-wave (PMMW) dataset. Experimental results demonstrate that the proposed approach is superior to the state-of-the-art methods, and achieves better performances on the two datasets with Average Precision (AP). Shuiping Gou, Jichao Li 0003, Yinghai Zhao, Changzhe Jiao, Shasha Mao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | A Stepwise Matching Method for Multi-modal Image based on Cascaded NetworkabstractTemplate matching of multi-modal image has been a challenge to image matching, and it is difficult to balance the speed and the accuracy, especially for images with large sizes. Based on this, we propose a stepwise image matching method to achieve a precise location from the coarse-to-fine image matching by utilizing cascaded networks. In the proposed method, a coarse-grained matching network is firstly constructed to locate a rough matching position based on cross-correlating features of optical and SAR images. Specially, to enhance the credible matching position, a suppression network is designed to evaluate for the obtained cross-correlation feature and added into the coarse-grained network as a feedback. Secondly, a fine-grained matching network is constructed based on the obtained rough matching result to gain a more precise matching. In this part, ternary groups are utilized to construct the training samples. Interestingly, we apply the region with a few pixels offset as the negative class, which effectively distinguishes similar neighbourhoods of the rough matching position. Moreover, a modified Siamese network is used to extract features of SAR and optical images, respectively. Finally, experimental results illustrate that the proposed method obtains more precise matching compared with the state-of-the-art methods. Jinming Mu, Shuiping Gou, Shasha Mao, Shankui Zheng |
ACM Multimedia | 2 |
| 2021 | End-to-End Ensemble Learning by Exploiting the Correlation Between Individuals and WeightsabstractEnsemble learning performs better than a single classifier in most tasks due to the diversity among multiple classifiers. However, the enhancement of the diversity is at the expense of reducing the accuracies of individual classifiers in general and, thus, how to balance the diversity and accuracies is crucial for improving the ensemble performance. In this paper, we propose a new ensemble method which exploits the correlation between individual classifiers and their corresponding weights by constructing a joint optimization model to achieve the tradeoff between the diversity and the accuracy. Specifically, the proposed framework can be modeled as a shallow network and efficiently trained by the end-to-end manner. In the proposed ensemble method, not only can a high total classification performance be achieved by the weighted classifiers but also the individual classifier can be updated based on the error of the optimized weighted classifiers ensemble. Furthermore, the sparsity constraint is imposed on the weight to enforce that partial individual classifiers are selected for final classification. Finally, the experimental results on the UCI datasets demonstrate that the proposed method effectively improves the performance of classification compared with relevant existing ensemble methods. Shasha Mao, Weisi Lin, Licheng Jiao, Shuiping Gou, Jiawei Chen 0001 |
IEEE Trans. Cybern. | 4 |
| 2021 | Non-Invasive Heart Rate Estimation From Ballistocardiograms Using Bidirectional LSTM RegressionabstractNon-invasive heart rate estimation is of great importance in daily monitoring of cardiovascular diseases. In this paper, a bidirectional long short term memory (bi-LSTM) regression network is developed for non-invasive heart rate estimation from the ballistocardiograms (BCG) signals. The proposed deep regression model provides an effective solution to the existing challenges in BCG heart rate estimation, such as the mismatch between the BCG signals and ground-truth reference, multi-sensor fusion and effective time series feature learning. Allowing label uncertainty in the estimation can reduce the manual cost of data annotation while further improving the heart rate estimation performance. Compared with the state-of-the-art BCG heart rate estimation methods, the strong fitting and generalization ability of the proposed deep regression model maintains better robustness to noise (e.g., sensor noise) and perturbations (e.g., body movements) in the BCG signals and provides a more reliable solution for long term heart rate monitoring. Changzhe Jiao, Chao Chen 0040, Shuiping Gou, Dong Hai 0001, Bo Yu Su, Marjorie Skubic, Licheng Jiao, Alina Zare, K. C. Ho 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Hyperspectral Target Detection via Multiple Instance LSTM Target Localization NetworkabstractModeling target detection problem given inaccurate annotations as a multiple instance learning (MIL) problem is an effective way for addressing the ground truth uncertainties of remotely sensed hyperspectral imagery. In this paper, we propose a hyperspectral target detection method based on 1D convolution neural network (1DCNN) feature extraction and long short term memory network (LSTM) under the MIL framework, where the LSTM features for each hyperspectral pixel is further refined by a scoring network as to discriminate the real target instance from the inaccurately labeled hyperspectral regions. The proposed method has achieved superior results on both simulated data and real hyperspectral data over the state-of-the-art methods, showing the prospects for further investigation. Xiuxiu Wang, Chubing Guo, Chao Chen 0040, Shuiping Gou, Changzhe Jiao |
IGARSS | 5 |
| 2020 | Ship Segmentation on High-Resolution Sar Image by a 3D Dilated Multiscale U-NetabstractTargets detection and segmentation in a synthetic aperture radar (SAR) image is a vital step for its interpretation. It is quite challenging for most conventional methods due to complex background and the speckle. Furthermore, the sizes of targets in a scene are variable. Inspired by the success of neural networks in computer vision, In this paper, we propose a 3D dilated multi-scale U-shape convolutional neural network (3DDM-UNet). In the proposed method, we first build a 3D image block via a multiscale stationary wavelet transform to exploit the structural information of targets with various sizes. Then, the built 3D image block is fed into a 3D dilated multiscale U-Net. To train the proposed network, we build a dataset from a scene of SAR image with various sizes and shapes of ship targets. Finally, the trained network is employed to the testing set to obtain the segmentation results. Experimental results on test images show that the proposed method achieved better performance than conventional methods. Jichao Li 0003, Chubing Guo, Shuiping Gou, Yuanbo Chen, Jiawei Chen 0010 |
IGARSS | 3 |
| 2019 | Bitemporal Fully Polarimetric Sar Images Change Detection Via Nearest Regularized Joint Sparse and Transfer Dictionary LearningabstractMost current synthetic aperture radar (SAR) images change detection methods are developed based on a difference image. In this paper, we propose a novel bitemporal polarmetric SAR (PolSAR) images change detection framework based on their land-covers classifications. First, a nearest regularized joint sparse representation (NRJSR) model is developed to exploit the correlations among various polarimetric information and spatial context. Next, a transfer dictionary learning method is proposed for bitemporal PolSAR images classifications. Finally, the changed map can be obtained by comparing these two classification results. The comparison experiment results show that the proposed algorithm obtains better performance. Yao Tan, Jichao Li 0003, Peiyang Zhang, Shuiping Gou, Yuanbo Chen, Jiawei Chen 0001, Changyan Sun |
IGARSS | 4 |
| 2019 | Hyperspectral Target Detection Via Deep Multiple Instance Self-Attention Neural NetworkabstractMultiple instance learning (MIL) can be used for solving the imprecisely labeled hyperspectral target detection problems, which only needs the label of an area containing some targets. Furthermore, existing methods decompose this task into a target signature learning task and a follow-on similarity measurement between the estimated signature and the test points. In this paper, we propose a deep multiple instance learning method based on self-attention mechanism, in which the max operation and 1D convolution neural network (1D CNN) are adopted to realize an end-to-end hyperspectral target detection structure without learning the target signature. In the proposed deep MIL target detection method, self-attention mechanism with max operation has advantage in estimating the labels of instances from the positive bag via calculating the contribution of each instance to the bag-level classification. The simulated and real hyperspectral target detection experiments are shown to illustrate the performance of the method. Xiuxiu Wang, Shuiping Gou, Chao Chen 0040, Yuanbo Chen, Xu Tang 0004, Changzhe Jiao |
IGARSS | 3 |
| 2018 | Classification of PolSAR Images Based on SVM with Self-Paced Learning OptimizationabstractA novel classification method for polarimetric synthetic aperture radar (PolSAR) images using support vector machine (SVM) with self-paced learning (SPL) optimization is proposed in the paper. In our method, Cloude-Pottier polarimetric decomposition components and the eigenvalues of coherency matrix are used as features. Classification is carried out using SVM, and SPL is used to improve the classifier and achieve a stronger generalization capacity. Under SPL paradigm, the classifier learns the easier samples first and gradually involves more difficult samples into the training process. The proposed method achieves the overall classification accuracies of 89.79% on the Flevoland dataset. Such results are comparable with the compared algorithms. Wenshuai Chen, Dong Hai 0001, Shuiping Gou, Licheng Jiao |
IGARSS | 3 |
| 2018 | Patch-Based Gaussian Mixture Model for Concealed Object Detection in Millimeter-Wave imagesabstractIn this paper, we proposed a novel method called patch based mixture of Gaussians-Low Rank Matrix Factorization (patch based MoG-LRMF) to detect concealed objects in a body image acquired by millimeter wave scanner at airport security procedures. Concealed objects vary significantly from size to type, which is comparatively random. MoG enables to model complex and uncertain information, which exactly match the characteristics of objects. However, related work is only able to model objects with pixel-wise information, which neglects the structure of the object, our patch based MoG utilizes structure and uncertainty of objects to detect concealed items. We demonstrate the effectiveness of our approach with enough experiment results. We find that many small and different material objects can be detected with our method, which performs well under relatively complicated data. Shuiping Gou, Xiuxiu Wang, Yinghai Zhao |
TENCON | 2 |
| 2018 | Classification of Heterogeneous Scenes in POL-SAR Image Based on Statistical AnalysisabstractIn this paper, a new method is presented based on statistical framework for the classification of heterogeneous scenes in Polarimetric SAR (POL-SAR) image. Firstly, a new measurement on the heterogeneous scenes is presented based on statistical analysis to describe the local space complexity and variability. Then, an adaptive threshold is set to classify the heterogeneous scenes of POL-SAR images. The performance of our method is tested by three real POL-SAR data sets. The experiment results show that our method works well for the heterogeneous scenes classification, and can effectively suppress noise. Shuiping Gou |
TENCON | 3 |
| 2017 | Full polarization SAR image classification using deep learning with shallow featureabstractThe classification of the POL-SAR image become more and more important with the development of the polarization of synthetic aperture radar system. Generally, the classification of POL-SAR images are based on polarization feature, such as support vector machine (SVM), Wishart clustering and other methods. Specifically, some ground objects usually have some weak scattering characteristics which cannot obtain good results by only using the traditional classification based on polarization features. So, the deep learning based on T matrix is used to mine the powerful feature of SAR data. In order to speed up computation and improve classification accuracy, a classification of full-polarization SAR images based on Deep Learning with Shallow features is proposed in this paper. The proposed method can get better classification for those weak scatter objects than those methods only using polarization features. Debo Li, Yu Gu 0015, Shuiping Gou, Licheng Jiao |
IGARSS | 3 |
| 2017 | Polarimetric SAR image change detection based on low rank and sparse representation with Freeman-Durden decompositionabstractPolarimetric synthetic aperthetic radar(POLAR) is more advanced imaging radar, which includes four channels and provides more information than a single-channel SAR image. However, existing change detection methods cannot make use of polarimetric information from POLSAR data to extract difference image. In this paper, a novel method of change detection based on low rank and sparse decomposition with Freeman-Durden decomposition(FDD) is proposed. The FDD is used to extract scattering features from POLSAR data. Low rank and sparse algorithm is applied to obtain difference information. The results demonstrate the better performance than classical methods for change detection of POLSAR image. Wei Liu 0127, Shuiping Gou, Licheng Jiao |
IGARSS | 3 |
| 2017 | A weighted joint sparse of three channels method for full POL-SAR data classificationabstractIn recent years, both passive and active (i.e., Synthetic Aperture Radar or SAR) satellite remote sensing has proven to be valuable tools for mapping land cover. Most of the classification algorithms are based on image-intensity and they do not perform well in different coastal zone types, because these terrains have similar optical or radar backscattering signals. In this paper we propose a weighted joint sparse on the three-channel to mine the polarimetric features and texture information. The proposed method can update weights automatically according to the difference of the three channels' contribution and the least residual error. The similarity and distinctiveness of three channels are used to deal with the complex object classification. Hybrid sparse coefficients are input to the support vector machine for fully polarimetric image classification. The proposed algorithm performed well in distinguishing some coastal land-use types. A comparison study is also conducted to show that proposed algorithm outperforms two commonly classification methods. Wenshuai Chen, Shuiping Gou, Xiangrong Zhang, Xiaofeng Li 0001, Licheng Jiao |
IGARSS | 3 |
| 2017 | Fast Classification for Large Polarimetric SAR Data Based on Refined Spatial-Anchor GraphabstractThe graph model-based semisupervised machine learning is well established. However, its computational complexity is still high in terms of the time consumption especially for large data. In this letter, we propose a fast semisupervised classification algorithm using the recently presented spatial-anchor graph for a large polarimetric synthetic aperture radar (Pol-SAR) data, named as Fast Spatial-Anchor Graph (FSAG) based algorithm. Based on an initial superpixel segmentation on the PolSAR image, the homogenous regions are obtained. The border pixels are reassigned to the most similar superpixel according to majority voting and distance measurement. Then, feature vectors are weighted within local homogenous regions. The refined spatial-anchor graph is constructed with these regions, and the semisupervised classification is conducted. Experimental results on synthesized and real PolSAR data indicate that the proposed FSAG greatly reduces time consumption and maintains the accuracy for terrain classifications compared with state-of-the-art graph-based approaches. Hongying Liu 0001, Shuyuan Yang 0001, Shuiping Gou, Puhua Chen, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Accelerating Learning to Rank via SVM with OpenCL and OpenMP on Heterogeneous PlatformsabstractSupport vector machine (SVM) is a popular algorithm for learning to rank, but the training speed of SVM is the bottleneck when dealing with large size data problems. Recently, heterogeneous computing platforms, such as graphics processing unit (GPU) and Many Integrated Core (MIC), have exhibited huge superiority in High Performance Computing domain. Open Computing Language (OpenCL) and Open Multi-Processing (OpenMP) are two popular parallel programming interface for different Heterogeneous Platforms. To resolve the speed problem of RSVM, comparison of the performance of different parallel programming models on different heterogeneous platforms is important. We designed OpenMPbased parallel learning to Rank SVM (PLRSVM) for multi-core CPU and MIC, and OpenCL-based PLRSVM for multi-core CPU, GPU and MIC. The experimental result shows the different performance between OpenMP based program and OpenCL based program. The OpenCL based program significantly speeds up training process of SVM and shows good portability on heterogeneous devices. The experiment also suggests that selection of suitable programming models according to the hardware platform and the structure of serial algorithm is an important step to acquire high performance of parallel algorithm. Huming Zhu, Yanfei Wu, Peng Zhang 0003, Shuiping Gou, Licheng Jiao |
ICPADS | 6 |
| 2016 | Classification of PolSAR image with non-negative tensor factorization approachabstractPolarimetric synthetic aperture radar (PolSAR) is of great importance in the remote sensing, which can be used widely in both civil and military fields. However, existing classification methods cannot effectively utilize the spatial structure information of the SAR data. In this study, a classification method for PolSAR images based on non-negative tensor factorization (NTF) is proposed. The proposed method uses tensor to represent the original data and extract the spacial structure feature by using NTF. Classification results are obtained by using support vector machines (SVM) and the Wishart clustering technology. The results show the validity and accuracy of the proposed method on PolSAR images classification. Shuiping Gou, Wenshuai Chen, Licheng Jiao |
IGARSS | 1 |
| 2016 | Coastal zone land-use classification of full-polarization SAR data based on joint sparseabstractToo many terrains of coastal zone require effective management to better understand coastal changes. With the development of full-polarization Synthetic Aperture Radar (POL-SAR) imaging, these terrains from coastal zone classification are available. But the classification of coastal land-use types based on full-polarization SAR data has not been investigated. In fact, the signal return coming from the sea can be frequently indistinct from one coming from the land. And coastal zone terrains modeled as a strong multiplicative noise as well as they are very similar with scattering mechanism, which makes the coastline zone types classification a very complicated issue. A joint sparse representation-based method for classification of coastal land-use types with full-polarization SAR data is proposed in this paper. Shuiping Gou, Xiaofeng Li 0001, Licheng Jiao |
IGARSS | 1 |
| 2016 | Image super-resolution based on the pairwise dictionary selected learning and improved bilateral regularisationabstractA pairwise dictionary selected learning (PDSL) model is proposed in this study, which is specially tailored to synthesise a low‐resolution image. This is accomplished with the use of an external dictionary of high‐resolution images, which is selectively learned based on the internal dictionary of the target reconstruction image. The PDSL can avoid interpolating the fictitious information to the reconstructed image. Optimisation is performed using a bilateral regularisation term in an edge reserving based on the aid of spatial and directional proximity. Using the proposed approach, the results achieved are superior to the existing dictionary learning‐based methods. The authors also present a quantitative evaluation of super‐resolution reconstruction using various statistics which demonstrated significant average peak signal‐to‐noise ratio improvements by their model. A comparison of the proposed method with five other state‐of‐the‐art methods is presented and the authors’ method achieves better visual effects in edge structures. Shuiping Gou, Shuzhen Liu, Yaosheng Wu, Licheng Jiao |
IET Image Process. | 1 |
| 2016 | Coastal Zone Classification With Fully Polarimetric SAR ImageryabstractClassifying different types of land cover in coastal zones using synthetic aperture radar (SAR) imagery is a challenge due to the fact that many types of coastal zone have similar backscattering characteristics. In this letter, we propose an unsupervised method based on a three-channel joint sparse representation (SR) classification with fully polarimetric SAR (PolSAR) data. The proposed method utilizes both texture and polarimetric feature information extracted from the HH, HV, and VV channels of a SAR image. The texture features are extracted by applying a wavelet transform to a SAR image, and then sparsely represented based on the correlation among the three channels. The polarimetric features, i.e., the scattering entropy and scattering angle from the H/α model, are also sparsely represented. A joint SR algorithm using both texture and polarimetric features is constructed to establish target dictionaries. An orthogonal matching pursuit algorithm is then used to calculate sparse coefficients. Hybrid coefficients are inputted to the kernel support vector machine for a fully PolSAR image classification. We applied the proposed algorithm to an Advanced Land Observing Satellite-2 L-band SAR image acquired in the Yellow River Delta, China. The classified land types are validated against the official survey map. The algorithm performs well in distinguishing six coastal land-use types. A comparison study is also conducted to show that proposed algorithm outperforms two commonly used classification methods. Shuiping Gou, Xiaofeng Li 0001, Xiaofeng Yang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | Weighted classifier ensemble based on quadratic form
Shasha Mao, Licheng Jiao, Shuiping Gou, Bo Chen 0001, Sai-Kit Yeung |
Pattern Recognit. | 4 |
| 2014 | Pol-SAR image classification using eigenvalue-based joint statistical frameworkabstractIn this paper, a novel method based on joint statistical framework is proposed for classification of polarimetric SAR image. The Gaussian model of the maximum eigenvalue and volume scattering power for coherency matrix is estimated to describe their statistical distribution. And Bayesian classifier is used to classify the polarimetric SAR image. In order to make full use of the local context structure of image, the local statistical model is used based on the maximum posterior probability (MAP) rule. The method is tested with the NASA/JPL AIRSAR data. Shuiping Gou, W. F. Wang, Licheng Jiao, Shuang Wang 0001, X. R. Zhang |
IGARSS | 1 |
| 2014 | Classification of imbalanced hyperspectral imagery data using support vector samplingabstractDue to the imbalance in obtaining labeled samples for different land-cover classes, hyperspectral image classification encounters the issue of imbalanced classification. In this paper, a novel and effective method is proposed to address the imbalanced learning problem in hyperspectral image classification, which combines support vector machine (SVM) and sampling strategy. The main novelty and contribution of our paper are that we propose to do sampling referring to the support vectors (SVs) rather than the training data to provide a balanced distribution during the model learning. Sampling among the training data may be time consuming, while sampling referring to the SVs is more efficient and representative with much lower complexity. Therefore, the proposed method is expected to be simple and effective for imbalanced learning problem. Experimental results on real hyperspectral image dataset show that our method can effectively improve the classification accuracy for the minority classes in the imbalanced dataset. Xiangrong Zhang, Yaoguo Zheng, Biao Hou, Shuiping Gou |
IGARSS | 5 |
| 2014 | Eigenvalue Analysis-Based Approach for POL-SAR Image ClassificationabstractA novel polarimetric synthetic aperture radar (POL-SAR) image classification approach is proposed in this paper by exploiting coherency matrix eigenvalues for polarimetric information representation and understanding. The approach consists of two parts. Initially, the statistical distributions of eigenvalue for homogeneous areas are analyzed by taking eigenvalues as the features of polarimetric information. The Bayesian classification method is applied to verify the feasibility of distinguishing different homogeneous areas. As a result, this method can work well those pixels with the similar scatter mechanism by using different polarimetric intensity information from eigenvalues. But this process cannot adequately distinguish those pixels with similar eigenvalues distribution. So, an eigenvalues-based local operator is defined to overcome the insufficient of the similar pixels by introducing a similar measure and eigenvalues-based texture information. After all pixels are classified by Bayesian classification, if the similarity of the pixel is larger than the given threshold, this pixel will be further classified by support vector machine using texture information. The proposed method is tested on three POL-SAR datasets, in which the average classification accuracy of eight categories for the Flevoland data from our method reaches nearly 90%. Shuiping Gou, Xiangrong Zhang, Weifang Wang, Fangfang Du |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | Multi-elitist immune clonal quantum clustering algorithm
Shuiping Gou, Xiong Zhuang, Yangyang Li 0001, Licheng Jiao |
Neurocomputing | 1 |
| 2012 | Graph based SAR images change detectionabstractA change detection method for SAR images based on graph is proposed in this paper. In order to avoid the information loss of constructing difference image, we operated directly on the original and changed images. Two adjacent graphs based on the original and changed images are constructed, and two adjacent graphs are connected together to obtain the difference graph. And then we cut the difference graph with spectral clustering. As a result, the global information and local structure information of each pixel from original and changed images are used by constructing three adjacent graphs graph. Compared with traditional difference image based change detection method, the proposed method has lower overall alarms. In addition, the patches are produced by using graph cut and computation cost is reduced greatly for our algorithm. Shuiping Gou, Tiantian Yu |
IGARSS | 1 |
| 2012 | SAR image change detection based on low rank matrix decompositionabstractIn this paper we propose an unsupervised approach for SAR image change detection task. A new method based on compressed sensing is applied. First using the PPB method for the speckle reduction, and then the logarithm ratio method is applied to generate a simple change map, and then the compressed sensing-based method is used to part the change map into a low rank part and a sparse part, where the sparse part is correspond to the changed area, finally k-means algorithm is applied to cluster the sparse part into two clusters. Experiment results show the effectiveness and feasibility of the proposed method. Xiangrong Zhang, Yaoguo Zheng, Jie Feng 0003, Shuiping Gou |
IGARSS | 4 |
| 2012 | Quantum Immune Fast Spectral Clustering for SAR Image SegmentationabstractSpectral clustering algorithm suffers from memory use and computational time bottleneck when handling large-scale image segmentation. By optimizing the selection of representative points before spectral embedding, a fast spectral clustering method with quantum immune optimization is proposed. The incorporation of quantum computing and immune clonal selection theory makes the selection of representative points more reasonable. The empirical study on the University of California Irvine standard data set clustering and synthetic aperture radar image segmentation demonstrates the efficiency of our algorithm and the capability to deal with large-scale data rapidly. Shuiping Gou, Xiong Zhuang, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2011 | An analysis of the eigenvalues method for polarimetric SAR image classificationabstractIn this paper, we focus on the distribution of eigenvalues, and based on Gaussian assumption, then we do an analysis of the eigenvalues potential for POL-SAR classification. Generally, we use Gaussian mixture model to describe the distribution of the eigenvalue and Bayesian classifier to achieve the POL-SAR pixel classification. The method is tested with the NASA/JPL AIRSAR data. Shuiping Gou, Licheng Jiao |
IGARSS | 1 |
| 2011 | Greedy optimization classifiers ensemble based on diversity
Shasha Mao, Licheng Jiao, Shuiping Gou |
Pattern Recognit. | 4 |
| 2010 | Algorithm of Partition based Network Boosting for imbalanced data classificationabstractNetwork Boosting (NB) is an ensemble learning method which combines weak learners together based on a network and can learn the target hypothesis asymptotically. NB has higher generalization ability compared to Bagging and AdaBoost. But, when datasets are class-imbalanced, the performance of NB will decrease quickly. In order to solve this problem, we present a Partition based Network Boosting method (PNB) to classify imbalanced data. For PNB method, every classifier node of the classifier network is provided with the same number of training data which are all of same weights. The classifier in the network is built by the balanced training set sampled from the training data according to the weights record of the training data it holds. And then, the weights of the instances of every node classifier are updated based on the classification results of self-node and its neighbor nodes. The classifier network is trained repeatedly in such a way. Weight factor of hypothesis in the training progress is introduced to improve the performance. The final classification is formed by all the hypotheses of the classifier network learned during the training progress so that the label of new instances can be decided by the weight voting. The experimental results on UCI data and imbalanced biomedical data show that the PNB algorithm has better AUC and recall performance compared with NB learning machine. Shuiping Gou, Licheng Jiao, Xiong Zhuang |
IJCNN | 1 |
| 2009 | Distributed Transfer Network Learning Based Intrusion DetectionabstractIn order to solve the problem that there exists unbalanced detection performance on different types of attacks in current large-scale network intrusion detection algorithms, Distributed Transfer Network Learning algorithm is proposed in this paper. The algorithm introduces transfer learning into Distributed Network Boosting algorithm for instructing the attacks learning with poor performance, in which the instances transfer learning is adopted for different domain adaptation. The experimental results on the Kdd Cup’99 Data Set show that the proposed algorithm has higher efficacy and better performance. Further, the detection accuracy of R2L attacks has been improved greatly while maintaining higher detection accuracy of other attack types. Shuiping Gou, Licheng Jiao |
ISPA | 1 |
| 2007 | Solving multidimensional knapsack problems by an immune-inspired algorithmabstractThis paper introduces a computational model simulating the dynamic process of human immune response to solve multidimensional knapsack problems. The new model is a quaternion (G, I, R, Al), where G denotes exterior stimulus or antigen, I denotes the set of valid antibodies, R denotes the set of reaction rules describing the interactions between antibodies, and Al denotes the dynamic algorithm describing how the reaction rules are applied to antibody population. The set of antibody-adjusting rules, the set of clonal selection rules, and a dynamic algorithm, named M P-PAISA, are designed for solving multidimensional knapsack problems. The efficiency of the proposed algorithm was validated by testing on 57 benchmark problems and comparing with three genetic algorithms. The results indicated that the proposed algorithm was suitable for solving multidimensional knapsack problems. Maoguo Gong, Licheng Jiao, Wenping Ma 0001, Shuiping Gou |
IEEE Congress on Evolutionary Computation | 4 |
| 2007 | Solving multiobjective clustering using an immune-inspired algorithmabstractIn this study, we introduced a novel multiobjective optimization algorithm, Nondominated Neighbor Immune Algorithm (NNIA), to solve the multiobjective clustering problems. NNIA solves multiobjective optimization problems by using a nondominated neighbor-based selection technique, an immune inspired operator, two heuristic search operators and elitism. The main novelty of NNIA is that the selection technique only selects minority isolated nondominated individuals in current population to clone proportionally to the crowding-distance values, recombine and mutate. As a result, NNIA pays more attention to the less-crowded regions in the current trade-off front. The experimental results on seven artificial data sets with different manifold structure and six real-world data sets show that the NNIA is an effective algorithm for solving multiobjective clustering problems, and the NNIA based multiobjective clustering technique is a cogent unsupervised learning method. Maoguo Gong, Lining Zhang, Licheng Jiao, Shuiping Gou |
IEEE Congress on Evolutionary Computation | 4 |
| 2007 | SVMs ensemble for radar target recognition based on evolutionary feature selectionabstractA novel radar target recognition method based on SVMs ensemble is presented, in which a set of suitable feature subsets are selected for component SVMs by Immune Clonal Algorithm, a new artificial immune system algorithm. With Immune Clonal Algorithm, high quality and high diversity of the components for SVMs ensemble are ensured. Experimental results on one-dimension radar high resolution range profiles demonstrate the validity and reliability of this new radar target recognition method. Xiangrong Zhang, Licheng Jiao, Shuiping Gou |
IEEE Congress on Evolutionary Computation | 3 |
| 2005 | Image Recognition Using Synergetic Neural Network
Shuiping Gou, Licheng Jiao |
ISNN (2) | 1 |
| 2004 | SAR Image Recognition Using Synergetic Neural Networks Based on Immune Clonal Programming
Shuiping Gou, Licheng Jiao |
ISNN (1) | 1 |