EDBT 2026 Demo / reviewers in the wild / expert
Hao Sun 0014
dblp:82/2248-14
· DBLP profile ↗
29ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0002-1314-4957ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Consensus Labeling: Prompt-Guided Clustering Refinement for Weakly Supervised Text-Based Person Re-IdentificationabstractWeakly supervised text-based person re-identification aims to retrieve specific pedestrians based on textual descriptions without identity labels available during training. This task remains challenging due to the inherent cross-modal heterogeneity and lack of identity annotations. There is a common issue of modality gap in vision language models, which in turn affects the performance of downstream tasks such as cross-modal retrieval and multimodal clustering. Specifically, in our research and experiments, we found that there is a problem of inter-modal misalignment between image and text modalities. However, existing methods rely on mutual enhancement strategies between image and text clustering, leading to the accumulation of clustering noise and affecting the final retrieval performance. To address this issue, we propose a Consensus Labelling: Prompt-guided Clustering refinement (CLPC) framework for weakly supervised text-based person re-identification. Specifically, we introduce a textual inversion network to learn a pseudo token that captures visual context, which is then integrated into natural language sentences as personalized textual prompt. To further improve clustering quality, we introduce a Nearest Neighbor-Guided Pseudo Label Mining (NGPM) method, which uses the clusters derived from personalized textual prompts to refine the clustering of image features. Additionally, we design a Dynamic Margin Triplet (DMT) loss, where the margin is adaptively adjusted using a sigmoid-based function to enhance the model’s ability to distinguish hard negative samples. We have also introduce a Normalized Distribution Matching (NDM) loss to minimize the KL divergence between the image-text matching scores and the normalized soft matching scores. The extensive experimental results on three public datasets have demonstrated the superiority of our method. Our code is available at https://github.com/LeviWeiZhi/CLPC. Chengji Wang, Weizhi Nie, Hongbo Zhang 0002, Hao Sun 0014, Mang Ye |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | Reconstruction-Contrast Coupling Learning for Open-Set Semi-Supervised Hyperspectral Image ClassificationabstractAlthough numerous semi-supervised learning methods have been elaborately designed for hyperspectral image (HSI) classification, most existing semi-supervised learning paradigms still rely on a closed-set assumption. These methods implicitly assume that the category spaces of labeled and unlabeled samples are completely aligned, that is, all unlabeled samples must belong to a pre-defined known category set. However, the closed-set assumption is particularly problematic in practical remote sensing scenarios because partial unlabeled data inevitably belong to unknown categories. To address this challenge, this paper proposes a reconstruction-contrast coupling learning (ReCo2L) method for open-set semi-supervised HSI classification, fully leveraging the complementarity between masked feature reconstruction learning and contrastive learning to enhance the encoder’s local detail sensitivity and global discriminative ability. Specifically, we first apply a masked feature reconstruction learning with an adaptive masking strategy to enhance the encoder’s ability to capture local details by high-quality spectral-spatial feature reconstruction. Then, we employ contrastive learning to strengthen the encoder’s capability to extract global characteristics by pulling semantically similar samples closer and pushing dissimilar ones farther apart in the feature space. Finally, a pixel-prototype deviation loss is proposed to further improve both inter-category distinguishability and intra-category compactness by reducing the distances between labeled sample features and their corresponding class anchors. Extensive experiments on three benchmark datasets demonstrate that our proposed ReCo2L achieves superior classification performance in both known and unknown categories and significantly surpasses 10 state-of-the-art HSI classification methods. The code will be available at https://github.com/repository-AI-chen/ReCo2L. Hao Sun 0014, Renyi Chen, Yong Chen 0024, Wenjing Chen 0003, Wei Xie 0008, Xiaoqiang Lu |
IEEE Trans. Image Process. | 1 |
| 2026 | P-CLIP: Progressive Discrepancy Learning for One-Shot Text-to-Image Person Re-IdentificationabstractOne-shot Text-to-Image Person Re-Identification (One-shot TIReID) aims to construct a TIReID model using only a single labeled image-text pair per identity, along with a large pool of unlabeled person images. While supervised learning in text-to-image person re-identification has demonstrated high effectiveness, the requirement for extensive annotated data, both in terms of identities and corresponding textual descriptions, makes it impractical for large-scale camera networks. One-shot TIReID presents a promising approach to reduce the annotation burden. The primary challenge in one-shot TIReID lies in establishing consistent visual-textual correspondences across diverse viewing conditions, particularly in the absence of cross-view paired data. To address this challenge, we propose a novel progressive discrepancy learning framework, termed P-CLIP, which aims to establish a shared embedding space that is robust to view-specific biases. To achieve this goal, we dynamically construct multi-view image-text pairs based on a single labeled pair and simultaneously project the multi-view data into a unified embedding space. Specifically, we propose a Progressive Multi-View Generation method (MVG) to generate multiple noisy views from a single labeled instance for training. To mitigate cross-view ambiguities, we introduce a Cross-View Discrepancy Learning module (CDL) that leverages the discrepancies among different views to guide the learning of cross-view visual-textual correspondences. This approach effectively integrates multimodal error correction into the person re-identification domain. Furthermore, to enhance the effectiveness of visual-textual correspondence learning, we propose a Compact Cross-Modal Matching Loss (CCM), which suppresses unmatched pairs while emphasizing matched ones. Extensive experiments were conducted on three benchmark datasets, and the experimental results demonstrate the effectiveness of our proposed method. The data and codes are available at https://github.com/Itachjw/P-CLIP/tree/main. Chengji Wang, Ming Dong 0004, Mang Ye, Hao Sun 0014, Xingpeng Jiang |
IEEE Trans. Image Process. | 4 |
| 2025 | SMILE: Semantic Multi-Scale Integration and LLM-Enhanced Influenza-Like Illness ForecastingabstractInfluenza-like illness (ILI) forecasting is crucial for effective public health intervention, but existing models often fail to capture the complex temporal and semantic patterns inherent in epidemic data. Traditional statistical techniques and even advanced deep learning methods predominantly leverage numerical time series data, thereby overlooking contextual medical and epidemiological insights that could enhance the performance. Recent progress in large language models (LLMs) has illustrated their exceptional effectiveness in integrating semantic understanding into natural language processing tasks within the medical context. Motivated by these developments, we propose SMILE (Semantic Multi-scale Integration and LLM-Enhanced network), a novel multi-modal forecasting framework designed to integrate LLM-derived semantic features with multi-scale temporal analysis. Built upon TimeMixer architecture, SMILE introduces an automatic semantic feature extraction system using LLMs, adaptive fusion mechanisms for integrating textual and temporal data, and demonstrates robust performance improvements. Extensive experiments on ILI and benchmark datasets confirm that SMILE significantly outperforms state-of-the-art forecasting methods, highlighting the value of incorporating semantic context into time series disease prediction. Ming Dong 0004, Qianxiao Fang, Hao Sun 0014, Weizhong Zhao, Tingting He 0003 |
BIBM | 3 |
| 2025 | PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric NetworkabstractWith the exponential growth of video content, aiming at localizing relevant video moments based on natural language queries, video moment retrieval (VMR) has gained significant attention. Existing weakly supervised VMR methods focus on designing various feature modeling and modal interaction modules to alleviate the reliance on precise temporal annotations. However, these methods have poor generalization capabilities on compositional queries with novel syntactic structures or vocabulary in real-world scenarios. To this end, we propose a new task: weakly supervised compositional moment retrieval (WSCMR). This task trains models using only video-query pairs without precise temporal annotations, while enabling generalization to complex compositional queries. Furthermore, a proposal-centric network (PC-Net) is proposed to tackle this challenging task. First, video and query features are extracted through frozen feature extractors, followed by modality interaction to obtain multimodal features. Second, to handle compositional queries with explicit temporal associations, a dual-granularity proposal generator decodes multimodal global and frame-level features to obtain query-relevant proposal boundaries with fine-grained temporal perception. Third, to improve the discrimination of proposal features, a proposal feature aggregator is constructed to conduct semantic alignment of frames and queries, and employ a learnable peak-aware Gaussian distributor to fit the frame weights within the proposals to derive proposal features from the video frame features. Finally, the proposal quality is assessed based on the results of reconstructing the masked query using the obtained proposal features. To further enhance the model's ability to capture semantic associations between proposals and queries, a quality margin regularizer is constructed to dynamically stratify proposals into high and low query-relevance subsets and enhance the association between queries and common elements within proposals, and suppress spurious correlations via inter-subset contrastive learning. Notably, PC-Net achieves superior performance with 54\% fewer parameters than prior works by parameter-efficient design. Experiments on Charades-CG and ActivityNet-CG demonstrate PC-Net’s ability to generalize across diverse compositional queries. Code is available at https://github.com/mingyao1120/PC-Net. Mingyao Zhou, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004, Chengji Wang, Mang Ye |
NeurIPS | 2 |
| 2025 | Correlation-based switching mean teacher for semi-supervised medical image segmentation
Guiyuhan Deng, Hao Sun 0014, Wei Xie 0008 |
Neurocomputing | 2 |
| 2025 | Bow Direction Detection Based on Angular Coding With Heading Intersection Over Union LossabstractAccurate bow direction detection is essential for ship trajectory prediction and port monitoring. Existing ship detection networks typically output angles within 180°, while extending to 360° introduces cyclic issues affecting rotation intersection over union (RIoU) accuracy. This study proposes a novel bow direction detection algorithm that extends network output to 360° and integrates a heading intersection over union (HIoU) loss to enhance detection accuracy and robustness. Additionally, an HIoU loss function is designed to improve bow direction identification and reduce quantization errors in hash codes. The algorithm is evaluated on three datasets: FGSD, OHD-SJTU-S, and OHD-SJTU-L. On FGSD, it achieves mean average precision (mAP) of 91.14%. On OHD-SJTU-S, it attains an$\text {mAP}_{50:95}$of 63.3% and a bow direction prediction accuracy of 90.7%. On OHD-SJTU-L, the$\text {mAP}_{50:95}$is 29.2%, with an accuracy of 80.2%. Yaxiong Chen, Qiangqiang Huang, Hao Sun 0014, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Learning Positive-Negative Prompts for Open-Set Remote Sensing Scene Classification
Hao Sun 0014, Hanlizi Chen, Wenjing Chen 0003, Chengji Wang, Wei Xie 0008, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Class-Aware Consistency Learning for Open-Set Semi-Supervised Hyperspectral Image ClassificationabstractSemi-supervised hyperspectral image (HSI) classification methods focus on exploring the spectral and spatial information of unlabeled samples. However, existing methods generally follow the closed-set setting, assuming that unlabeled samples do not contain novel classes, which is hard to hold in practical applications. This paper aims to study semi-supervised HSI classification in the open-set setting, i.e., unlabeled samples fall into novel classes, and proposes a class-aware consistency learning (CACL) method. First, to explore discriminative spectral-spatial features, a position-aware transformer is developed, which effectively models spatial position priors between the center pixel and its neighboring pixels via a symmetric position-aware encoding. Then, to reduce the interference from novel class samples on the model’s discrimination, a prototype-driven consistency learning is proposed, which accurately selects unlabeled samples belonging to known classes via a known class sampler, and efficiently utilizes their spectral-spatial information by modeling consistent predictions across different views. Finally, to further improve the distinguishability between known classes, a prototype contrastive optimization is proposed to decrease the distance between samples from the same class and increase the distances between those from different classes in the feature domain. Furthermore, an adaptive segmentation threshold is designed to accurately predict known classes and reject novel classes. Extensive experiments verify that our CACL outperforms the state-of-the-art methods, achieving the overall accuracy of 81.65%, 88.61%, and 92.88% on the Indian Pines, Salinas, and Pavia University datasets, with 10 labeled samples in each known class. The code is available at https://github.com/rock-in/CACL-main. Hao Sun 0014, Renyi Chen, Huaxiong Yao, Yaxiong Chen, Wei Xie 0008, Guirong Feng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight DetectionabstractVideo moment retrieval (MR) and highlight detection (HD) based on natural language queries are two highly related tasks, which aim to obtain relevant moments within videos and highlight scores of each video clip. Recently, several methods have been devoted to building DETR-based networks to solve both MR and HD jointly. These methods simply add two separate task heads after multi-modal feature extraction and feature interaction, achieving good performance. Nevertheless, these approaches underutilize the reciprocal relationship between two tasks. In this paper, we propose a task-reciprocal transformer based on DETR (TR-DETR) that focuses on exploring the inherent reciprocity between MR and HD. Specifically, a local-global multi-modal alignment module is first built to align features from diverse modalities into a shared latent space. Subsequently, a visual feature refinement is designed to eliminate query-irrelevant information from visual features for modal interaction. Finally, a task cooperation module is constructed to refine the retrieval pipeline and the highlight score prediction process by utilizing the reciprocity between MR and HD. Comprehensive experiments on QVHighlights, Charades-STA and TVSum datasets demonstrate that TR-DETR outperforms existing state-of-the-art methods. Codes are available at https://github.com/mingyao1120/TR-DETR. Hao Sun 0014, Mingyao Zhou, Wenjing Chen 0003, Wei Xie 0008 |
AAAI | 1 |
| 2024 | Segmentation Foundation Model-Aided Medical Image SegmentationabstractAccurate medical image segmentation is significant for reliable clinical diagnoses and pathology research. Deep learning methods rely on large high-quality dataset for supervised training to achieve satisfactory performance, but manually annotating large-scale datasets is time-consuming and costly. A recent breakthrough in the segmentation foundation model SAM shows promise for aiding annotation. However, when using SAM for annotation, errors are inevitably introduced. To utilize SAM for effective labeling, we adopt a dual-stream collaborative learning framework. Initially, a subset from the dataset is divided and then SAM is used to generate noisy labels. The first branch incorporates a dynamic weight fusion module, which adaptively fuses complementary features from model predictions and noisy labels, reconstructing more informative labels. The second branch aims to improve segmentation accuracy by training the model on an accurate subset and sharing parameters with the other branch. Our method outperforms state-of-the-art segmentation methods on the ISIC2018 and BUSI datasets. Shiqi Hua, Dunbo Ning, Wei Xie 0008, Hao Sun 0014 |
BIBM | 4 |
| 2024 | Spatial Formation-Guided Network for Group Activity RecognitionabstractEffectively modeling the interactions among actors is critical and challenging for Group Activity Recognition (GAR). Previous methods usually divide actors into subgroups based on the similarity of appearance features for modeling multilevel interactions among actors. However, the appearance feature-based grouping scheme does not fully consider the spatial relations of actors, which can provide a discriminative clue for GAR. In this paper, we propose a Spatial Formation-Guided Network (SFGN) to capture effective interactions under the guidance of spatial formations. We first design a spatial formation extractor to excavate latent spatial relations among actors for extracting spatial formation features. Then, a formation-guided interaction module is built to utilize the spatial formation features to guide the interactions among actors. Finally, a cross-formation interaction module is further designed to explore the complementarity among diverse spatial formations. Extensive experiments on the volleyball dataset and the collective activity dataset demonstrate that SFGN outperforms the state-of-the-art methods. Dunbo Ning, Wenjing Chen 0003, Wei Xie 0008, Hao Sun 0014 |
ICASSP | 4 |
| 2024 | Cross-Modal Multiscale Difference-Aware Network for Joint Moment Retrieval and Highlight DetectionabstractSince the goals of both Moment Retrieval (MR) and Highlight Detection (HD) are to quickly obtain the required content from the video according to user needs, several works have attempted to take advantage of the commonality between both tasks to design transformer-based networks for joint MR and HD. Although these methods achieve impressive performance, they still face some problems: a) Semantic gaps across different modalities. b) Various durations of different query-relevant moments and highlights. c) Smooth transitions among diverse events. To this end, we propose a Cross-modal Multiscale Difference-aware Network, named CMDNet. First, a clip-text alignment module is constructed to narrow semantic gaps between different modalities. Second, a multiscale difference perception module is utilized to mine the differential information between adjacent clips and perform multiscale modeling to obtain discriminative representations. Finally, these representations are fed into the MR and HD task heads to retrieve relevant moments and estimate highlight scores precisely. Extensive experiments on three popular datasets demonstrate that CMDNet achieves state-of-the-art performance. Mingyao Zhou, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008 |
ICASSP | 3 |
| 2024 | Spatial Dual Context Learning for Weakly-supervised Group Activity Recognition in Still-imagesabstractThis paper investigates a new task, Weakly- supervised Group Activity Recognition in Still-images (WGARS), which aims to extend the applicability of Group Activity Recognition (GAR) to broader scenarios, such as low-latency domains. To tackle this challenge, we propose a Spatial Dual Context Transformer (SDCT), comprising a Dual Context Encoder (DCE) and a Dual Context Decoder (DCD). The DCE module individually encodes holistic context with integral relations of overall actors, and encodes partial context with individual features in still images. Subsequently, the DCD module explores the complementarity between holistic and partial contexts, and alternatively updates these encoded contexts to enhance the interaction of actors. Additionally, auxiliary supervised contrastive learning is incorporated to mitigate activity confusion. The proposed SDCT attains state-of-the-art performance on Volleyball and NBA datasets in WGARS. Notably, SDCT even outperforms recent methods when extended to the weakly-supervised GAR in videos task on Volleyball dataset. Dunbo Ning, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004 |
ICME | 4 |
| 2024 | Image-Centered Pseudo Label Generation for Weakly Supervised Text-Based Person Re-Identification
Weizhi Nie, Chengji Wang, Hao Sun 0014, Wei Xie 0008 |
PRCV (12) | 3 |
| 2024 | Query-aware multi-scale proposal network for weakly supervised temporal sentence grounding in videos
Mingyao Zhou, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Ming Dong 0004, Xiaoqiang Lu |
Knowl. Based Syst. | 3 |
| 2024 | Prototype-Based Pseudo-Label Refinement for Semi-Supervised Hyperspectral Image ClassificationabstractPseudo-label learning-based methods usually regard class confidence above a certain threshold for unlabeled samples as pseudo-labels, which may result in pseudo-labels still containing wrong labels. In this letter, we propose a prototype-based pseudo-label refinement (PPLR) for semi-supervised hyperspectral image classification. The proposed PPLR filters wrong labels from pseudo-labels using class prototypes, which can improve the discrimination of the network. First, PPLR uses multi-head attentions to extract the spectral-spatial features, and designs an adaptive threshold that can be dynamically adjusted to generate high-confidence pseudo-labels. Then, PPLR constructs class prototypes for different categories using labeled sample features and unlabeled sample features with refined pseudo-labels to improve the quality of pseudo-labels by filtering wrong labels. Finally, PPLR further assigns reliable weights to these pseudo-labels in calculating their supervised loss, and introduces a center loss to improve the discrimination of features. When 10 labeled samples per category are utilized for training, PPLR achieves the overall accuracies of 82.11%, 86.70% and 92.50% on the Indian Pines, Houston2013 and Salinas datasets, respectively. Renyi Chen, Huaxiong Yao, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Xiaoqiang Lu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Orientational Clustering Learning for Open-Set Hyperspectral Image ClassificationabstractRecently, some literature has begun to pay attention to the open-set problem in remote sensing application scenarios and studied various open-set hyperspectral image classification (OSHIC) methods. These OSHIC methods are usually based on deep neural networks, using the nondirectional Euclidean distance losses to constrain latent sample representations of known classes to be compact. Nonetheless, the potential effect of the spatial distribution of sample representations is ignored, resulting in degraded classification performance in OSHIC. In this letter, we propose an orientational clustering learning (OCL) method for OSHIC. First, in the feature space generated by the convolutional neural network, a class anchor strategy is employed to bring features of the same class closer while keeping features of different classes distant. Then, we utilize the orientational learning to further tighten the intraclass feature space. OCL directionally optimizes the spatial distribution of hyperspectral sample representations to improve the ability to identify known classes and distinguish unknown classes. Experiments show that the OCL achieves overall accuracies of 94.43%, 92.27%, and 76.94% on the Pavia University, Salinas, and Indian Pines datasets, respectively. Wenjing Chen 0003, Hailong Ning, Hao Sun 0014, Wei Xie 0008 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Cross-Modal Feature Fusion-Based Knowledge Transfer for Text-Based Person SearchabstractText-based person search aims to retrieve corresponding images of person from a large gallery based on text descriptions. Existing methods strive to bridge the modality gap between images and texts and have made promising progress. However, these approaches disregard the knowledge imbalance between images and texts caused by the reporting bias. To resolve this issue, we present a cross-modal feature fusion-based knowledge transfer network to balance identity information between images and texts. First, we design an identity information emphasis module to enhance person-relevant information and suppress person-irrelevant information. Second, we design an intermediate modal-guided knowledge transfer module to balance the knowledge between images and texts. Experimental results on CUHK-PEDES, ICFG-PEDE, and RSTPReid datasets demonstrate that our method achieves state-of-the-art performance. Kaiyang You, Wenjing Chen 0003, Chengji Wang, Hao Sun 0014, Wei Xie 0008 |
IEEE Signal Process. Lett. | 4 |
| 2023 | Semi-Supervised Facial Expression Recognition by Exploring False Pseudo-LabelsabstractPseudo-labels are popular in semi-supervised facial expression recognition. Recent methods usually exploit the confidence as the criterion for pseudo-label generation, and utilize the high-confidence pseudo-labels as the ground-truth for training. However, high confidence cannot guarantee the correctness of pseudo-labels. False pseudo-labels can weaken the feature discrimination and degrade recognition performance. In this paper, we propose a Critical Feature Refinement Network (CFRN) to alleviate the interference of false pseudo-labels on the model performance. Specially, a feature dropout module and a feature emphasis module are proposed to improve the feature discrimination of CFRN. Then, a mean-absolute error loss is further exploited to improve the robustness against false pseudo-labels. Experimental results on three challenging datasets RAF-DB, SFEW and Affectnet demonstrate that the proposed CFRN outperforms the state-of-the-art methods. Hao Sun 0014, Chenchen Pi, Wei Xie 0008 |
ICME | 1 |
| 2023 | Deep Feature Reconstruction Learning for Open-Set Classification of Remote-Sensing ImageryabstractExisting remote sensing scene image (RSSI) classification methods usually rely on static closed-set assumption that testing samples do not belong to unknown classes. However, practical applications are usually the open-set classification problem, which means that RSSIs from unknown classes will appear in the testing set. Most existing methods are prone to forcibly misclassify RSSIs of unknown classes into known classes, resulting in poor practical performance. In this letter, a deep feature reconstruction learning (DFRL) framework is proposed for open-set classification of RSSIs. The proposed DFRL unifies discriminative feature learning and feature reconstruction into an end-to-end network. Firstly, a feature extraction module is utilized to project raw input data from the image space to the feature space to extract deep features. Then, the deep features are fed to a deep feature reconstruction module for distinguishing known and unknown classes based on feature-level reconstruction errors. The feature-level reconstruction can effectively suppress the interference of complex backgrounds. In addition, a sparse regularization is introduced to improve the discrimination of image representation. Experiments on three RSSI datasets demonstrate the effectiveness of DFRL for open-set classification of RSSIs. Hao Sun 0014, Jie Yu 0003, Dongbo Zhou, Wenjing Chen 0003, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Pseudolabel-Based Unreliable Sample Learning for Semi-Supervised Hyperspectral Image ClassificationabstractRecently, pseudo-label-based deep learning methods have shown excellent performance in semi-supervised hyperspectral image (HSI) classification. These methods usually select high-confidence unlabeled samples to help optimize backbone classification networks. However, a large number of remaining low-confidence unlabeled samples, which contain rich land-covers information, are underutilized. In this paper, we propose a pseudo-label-based unreliable sample learning (PUSL) method to fully exploit low-confidence unlabeled samples for semi-supervised HSI classification. Firstly, to avoid overfitting the spatial distribution of labeled samples, we build a position-free transformer (PFT) as the backbone classification network. Secondly, PFT is initially trained with labeled samples in a supervised learning manner to obtain an initial classifier, which is then used to split unlabeled samples into reliable and unreliable unlabeled samples based on the predicted confidence. Thirdly, reliable unlabeled samples participate in training along with labeled samples. Finally, unreliable unlabeled samples are treated as negative samples for corresponding categories to improve the discrimination of PFT in a contrastive learning paradigm. Extensive experiments on three HSI datasets demonstrate that PUSL outperforms compared methods. Huaxiong Yao, Renyi Chen, Wenjing Chen 0003, Hao Sun 0014, Wei Xie 0008, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Deformable Attention U-Shaped Network with Progressively Supervised Learning for Subarachnoid Hemorrhage Image SegmentationabstractSubarachnoid hemorrhage (SAH) is a common acute disease, which belongs to a subtype of intracranial hemorrhage. In this paper, a deformable attention u-shaped network (DAUN) is specially designed for SAH image segmentation. Firstly, a deformable attention module is embedded at the end of each encoding layer in Res-UNet to adaptively adjust the attention domain for alleviating the introduction of irrelevant information. Then, to improve the segmentation accuracy on irregular edges and small lesions, a region-boundary-aware loss is utilized to optimize the model. Finally, a progressively supervised learning strategy is proposed to train the proposed DAUN, which enables DAUN to find a balance between the focus on semantic information and position information of each pixel. A novel SAHCT dataset is constructed to demonstrate the performance of DAUN. In addition, the Monuseg dataset is utilized to evaluate the generalization ability of DAUN. Hao Sun 0014, Lianghao Jin, Wei Xie 0008 |
BIBM | 1 |
| 2022 | Cross-Attention Spectral-Spatial Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification aims to identify categories of hyperspectral pixels. Recently, many convolutional neural networks (CNNs) have been designed to explore the spectrums and spatial information of HSI for classification. In recent CNN-based methods, 2-D or 3-D convolutions are inevitably utilized as basic operations to extract the spatial or spectral–spatial features. However, 2-D and 3-D convolutions are sensitive to the image rotation, which may result in that recent CNN-based methods are not robust to the HSI rotation. In this article, a cross-attention spectral–spatial network (CASSN) is proposed to alleviate the problem of HSI rotation. First, a cross-spectral attention component is proposed to exploit the local and global spectrums of the pixel to generate band weight for suppressing redundant bands. Second, a spectral feature extraction component is utilized to capture spectral features. Then, a cross-spatial attention component is proposed to generate spectral–spatial features from the HSI patch under the guidance of the pixel to be classified. Finally, the spectral–spatial feature is fed to a softmax classifier to obtain the category. The effectiveness of CASSN is demonstrated on three public databases. Hao Sun 0014, Chunbo Zou, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Rotation-Invariant Attention Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification refers to identifying land-cover categories of pixels based on spectral signatures and spatial information of HSIs. In recent deep learning-based methods, to explore the spatial information of HSIs, the HSI patch is usually cropped from original HSI as the input. And 3 ×3 convolution is utilized as a key component to capture spatial features for HSI classification. However, the 3 ×3 convolution is sensitive to the spatial rotation of inputs, which results in that recent methods perform worse in rotated HSIs. To alleviate this problem, a rotation-invariant attention network (RIAN) is proposed for HSI classification. First, a center spectral attention (CSpeA) module is designed to avoid the influence of other categories of pixels to suppress redundant spectral bands. Then, a rectified spatial attention (RSpaA) module is proposed to replace 3 ×3 convolution for extracting rotation-invariant spectral-spatial features from HSI patches. The CSpeA module, the 1 ×1 convolution and the RSpaA module are utilized to build the proposed RIAN for HSI classification. Experimental results demonstrate that RIAN is invariant to the spatial rotation of HSIs and has superior performance, e.g., achieving an overall accuracy of 86.53% (1.04% improvement) on the Houston database. The codes of this work are available at https://github.com/spectralpublic/RIAN. Xiangtao Zheng, Hao Sun 0014, Xiaoqiang Lu, Wei Xie 0008 |
IEEE Trans. Image Process. | 2 |
| 2021 | A Supervised Segmentation Network for Hyperspectral Image ClassificationabstractRecently, deep learning has drawn broad attention in the hyperspectral image (HSI) classification task. Many works have focused on elaborately designing various spectral-spatial networks, where convolutional neural network (CNN) is one of the most popular structures. To explore the spatial information for HSI classification, pixels with its adjacent pixels are usually directly cropped from hyperspectral data to form HSI cubes in CNN-based methods. However, the spatial land-cover distributions of cropped HSI cubes are usually complicated. The land-cover label of a cropped HSI cube cannot simply be determined by its center pixel. In addition, the spatial land-cover distribution of a cropped HSI cube is fixed and has less diversity. For CNN-based methods, training with cropped HSI cubes will result in poor generalization to the changes of spatial land-cover distributions. In this paper, an end-to-end fully convolutional segmentation network (FCSN) is proposed to simultaneously identify land-cover labels of all pixels in a HSI cube. First, several experiments are conducted to demonstrate that recent CNN-based methods show the weak generalization capabilities. Second, a fine label style is proposed to label all pixels of HSI cubes to provide detailed spatial land-cover distributions of HSI cubes. Third, a HSI cube generation method is proposed to generate plentiful HSI cubes with fine labels to improve the diversity of spatial land-cover distributions. Finally, a FCSN is proposed to explore spectral-spatial features from finely labeled HSI cubes for HSI classification. Experimental results show that FCSN has the superior generalization capability to the changes of spatial land-cover distributions. Hao Sun 0014, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Image Process. | 1 |
| 2020 | Remote Sensing Scene Classification by Gated Bidirectional NetworkabstractRemote sensing (RS) scene classification is a challenging task due to various land covers contained in RS scenes. Recent RS classification methods demonstrate that aggregating the multilayer convolutional features, which are extracted from different hierarchical layers of a convolutional neural network, can effectively improve classification accuracy. However, these methods treat the multilayer convolutional features as equally important and ignore the hierarchical structure of multilayer convolutional features. Multilayer convolutional features not only provide complementary information for classification but also bring some interference information (e.g., redundancy and mutual exclusion). In this paper, a gated bidirectional network is proposed to integrate the hierarchical feature aggregation and the interference information elimination into an end-to-end network. First, the performance of each convolutional feature is quantitatively analyzed and a superior combination of convolutional features is selected. Then, a bidirectional connection is proposed to hierarchically aggregate multilayer convolutional features. Both the top–down direction and the bottom–up direction are considered to aggregate multilayer convolutional features into the semantic-assist feature and appearance-assist feature, respectively, and a gated function is utilized to eliminate interference information in the bidirectional connection. Finally, the semantic-assist feature and appearance-assist feature are merged for classification. The proposed method can compete with the state-of-the-art methods on four RS scene classification data sets (AID, UC-Merced, WHU-RS19, and OPTIMAL-31). Hao Sun 0014, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Spectral-Spatial Attention Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification aims to assign each hyperspectral pixel with a proper land-cover label. Recently, convolutional neural networks (CNNs) have shown superior performance. To identify the land-cover label, CNN-based methods exploit the adjacent pixels as an input HSI cube, which simultaneously contains spectral signatures and spatial information. However, at the edge of each land-cover area, an HSI cube often contains several pixels whose land-cover labels are different from that of the center pixel. These pixels, named interfering pixels, will weaken the discrimination of spectral-spatial features and reduce classification accuracy. In this article, a spectral-spatial attention network (SSAN) is proposed to capture discriminative spectral-spatial features from attention areas of HSI cubes. First, a simple spectral-spatial network (SSN) is built to extract spectral-spatial features from HSI cubes. The SSN is composed of a spectral module and a spatial module. Each module consists of only a few 3-D convolution and activation operations, which make the proposed method easy to converge with a small number of training samples. Second, an attention module is introduced to suppress the effects of interfering pixels. The attention module is embedded into the SSN to obtain the SSAN. The experiments on several public HSI databases demonstrate that the proposed SSAN outperforms several state-of-the-art methods. Hao Sun 0014, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | A Feature Aggregation Convolutional Neural Network for Remote Sensing Scene ClassificationabstractRemote sensing scene classification (RSSC) refers to inferring semantic labels based on the content of the remote sensing scenes. Recently, most works take the pretrained convolutional neural network (CNN) as the feature extractor to build a scene representation for RSSC. The activations in different layers of CNN (named intermediate features) contain different spatial and semantic information. Recent works demonstrate that aggregating intermediate features into a scene representation can significantly improve the classification accuracy for RSSC. However, the intermediate features are aggregated by some unsupervised feature encoding methods (e.g., Bag-of-Visual-Words). Little attention has been paid to explore the information of semantic labels for the feature aggregation. In this paper, in order to explore the semantic label information, an end-to-end feature aggregation CNN (FACNN) is proposed to learn a scene representation for RSSC. In FACNN, a supervised convolutional features' encoding module and a progressive aggregation strategy are proposed to leverage the semantic label information to aggregate the intermediate features. The FACNN integrates the feature learning, feature aggregation, and classifier into a unified end-to-end framework for joint training. In FACNN, the scene representation is learned by considering the information of semantic labels, which can result in better performance for RSSC. Extensive experiments on AID, UC-Merged, and WHU-RS19 databases demonstrate that FACNN performs better than several state-of-the-art methods. Xiaoqiang Lu, Hao Sun 0014, Xiangtao Zheng |
IEEE Trans. Geosci. Remote. Sens. | 2 |