Haixia Bi

dblp:200/3134 · DBLP profile ↗
← Back
31ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0002-3629-0332ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 10 first-author · 17 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly Detection
abstract
Video Anomaly Detection (VAD) aims to locate events that deviate from normal patterns in videos. Traditional approaches often rely on extensive labeled data and incur high computational costs. Recent tuning-free methods based on Multimodal Large Language Models (MLLMs) offer a promising alternative by leveraging their rich world knowledge. However, these methods typically rely on textual outputs, which introduces information loss, exhibits normalcy bias, and suffers from prompt sensitivity, making them insufficient for capturing subtle anomalous cues. To address these constraints, we propose HeadHunt-VAD, a novel tuning-free VAD paradigm that bypasses textual generation by directly hunting robust anomaly-sensitive internal attention heads within the frozen MLLM. Central to our method is a Robust Head Identification module that systematically evaluates all attention heads using a multi-criteria analysis of saliency and stability, identifying a sparse subset of heads that are consistently discriminative across diverse prompts. Features from these expert heads are then fed into a lightweight anomaly scorer and a temporal locator, enabling efficient and accurate anomaly detection with interpretable outputs. Extensive experiments show that HeadHunt-VAD achieves state-of-the-art performance among tuning-free methods on two major VAD benchmarks while maintaining high efficiency, validating head-level probing in MLLMs as a powerful and practical solution for real-world anomaly detection.
Zhaolin Cai, Fan Li 0003, Ziwei Zheng, Haixia Bi, Lijun He 0001
AAAI4
2026 Invisible Triggers, Visible Threats! Road-Style Adversarial Creation Attack for Visual 3D Detection in Autonomous Driving
abstract
Modern autonomous driving (AD) systems leverage 3D object detection to perceive foreground objects in 3D environments for subsequent prediction and planning. Visual 3D detection based on RGB cameras provides a cost-effective solution compared to the LiDAR paradigm. While achieving promising detection accuracy, current deep neural network-based models remain highly susceptible to adversarial examples. The underlying safety concerns motivate us to investigate realistic adversarial attacks in AD scenarios. Previous work has demonstrated the feasibility of placing adversarial posters on the road surface to induce hallucinations in the detector. However, the unnatural appearance of the posters makes them easily noticeable by humans, and their fixed content can be readily targeted and defended. To address these limitations, we propose the AdvRoad to generate diverse road-style adversarial posters. The adversaries have naturalistic appearances resembling the road surface while compromising the detector to perceive non-existent objects at the attack locations. We employ a two-stage approach, termed Road-Style Adversary Generation and Scenario-Associated Adaptation, to maximize the attack effectiveness on the input scene while ensuring the natural appearance of the poster, allowing the attack to be carried out stealthily without drawing human attention. Extensive experiments show that AdvRoad generalizes well to different detectors, scenes, and spoofing locations. Moreover, physical attacks further demonstrate the practical threats in real-world environments.
Jian Wang 0113, Lijun He 0001, Yixing Yong, Haixia Bi, Fan Li 0003
AAAI4
2026 RectMamba: Exploring state space models with entropy-divergence framework for noisy label rectification
Ningwei Wang, Weiqiang Jin, Haixia Bi, Guang Yang 0006
Neurocomputing3
2026 Polarimetric diffusion model with hybrid convolutional Transformer for PolSAR image classification
Zuzheng Kuang, Haixia Bi, Lijun He 0001, Fan Li 0003
Pattern Recognit.2
2026 IDEAL: Independent domain embedding augmentation learning
Zelin Yang, Lin Xu 0001, Shiyang Yan, Haixia Bi, Fan Li 0003
Pattern Recognit.4
2026 LSFMamba: Local-Enhanced Spiral Fusion Mamba for Multi-Modal Land Cover Classification
abstract
Multi-modal learning, which fuses complementary information from different modalities, has significantly improved the accuracy of land cover classification, especially under adverse conditions like cloudy or rainy weather. Recent advancements in multi-modal remote sensing land cover classification (MMRLC) have witnessed the efficacy of approaches based on CNN and Transformer. However, CNN exhibits limitations in capturing long-range dependencies, whereas Transformer suffers from high computational complexity. Recently, Mamba has garnered widespread attention due to its superior long-range modeling capabilities with linear complexity. Nevertheless, Mamba demonstrates notable limitations when directly applied to MMRLC, including limited local contextual modeling capacity, suboptimal multi-modal feature fusion and lack of a task-specific spatial continuity scanning strategy. Hence, to fully explore the potential of Mamba in multi-modal land cover classification, we propose LSFMamba, which comprises multiple hierarchically connected local-enhanced fusion Mamba (LFM) modules. Within each LFM module, a local-enhanced visual state space (LVSS) block is designed to extract features from different modalities, while a cross-modal interaction state space (CISS) block is created to fuse these multi-modal features. In the LVSS block, we integrate a multi-kernel CNN block into the gating branch in Mamba to enhance its local modeling capabilities. In the CISS block, features from different modalities are interleaved, facilitating cross-modal feature interaction through the state space model. Furthermore, we introduce a novel spiral scanning strategy to reassess the significance of central pixels, a design driven by the unique characteristics of pixel-wise classification task. Extensive experimental results on three multi-modal remote sensing datasets demonstrate that the proposed LSFMamba achieves state-of-the-art performance with lower complexity. The code will be released at https://github.com/hhchhang78/LSFMamba.
Honghao Chang, Haixia Bi, Fan Li 0003
IEEE Trans. Circuits Syst. Video Technol.2
2025 Improved Baselines with Synchronized Encoding for Universal Medical Image Segmentation
Jiadong Feng, Xuande Mi, Haixia Bi, Jian Sun 0009
MICCAI (2)4
2025 ECP-Mamba: An Efficient Multiscale Self-Supervised Contrastive Learning Method With State Space Model for PolSAR Image Classification
abstract
Recently, polarimetric synthetic aperture radar (PolSAR) image classification has been greatly promoted by deep neural networks. However, current deep learning-based PolSAR image classification methods are caught in the dilemma of obtaining high accuracy with sparse labels while maintaining high computational efficiency. To solve this issue, we present ECP-Mamba, an efficient framework integrating multi-scale self-supervised contrastive learning with a state space model backbone. Specifically, we design a cross-scale predictive pretext task, which learns representations via aligning local and global polarimetric features, effectively mitigating the annotation scarcity issue. To enhance computational efficiency, we introduce Mamba architecture to PolSAR image classification for the first time. A spiral scanning strategy tailored for pixel-wise classification task is proposed within this framework, prioritizing causally relevant features near the central pixel. Additionally, a lightweight cross Mamba module is proposed to facilitate complementary multi-scale feature interaction. Extensive experiments on four benchmark datasets demonstrate the effectiveness of ECP-Mamba in balancing high accuracy with computational efficiency. On the Flevoland 1989 dataset, ECP-Mamba achieves state-of-the-art performance with an overall accuracy of 99.70%, an average accuracy of 99.64% and a Kappa coefficient of 0.9962. Our code will be available at https://github.com/HaixiaBi1982/ECP_Mamba.
Zuzheng Kuang, Haixia Bi, Fan Li 0003
IEEE Trans. Geosci. Remote. Sens.2
2025 Unsupervised and Unpaired Fusion-Based Hyperspectral Image Super-Resolution Based on Reflectance Migration
abstract
Fusion-based hyperspectral image (HSI) super-resolution has garnered growing attention in recent years due to its strong capability in reconstructing high spatial resolution (HR) hyperspectral images. However, most existing methods either rely on a supervised training strategy or require pairwise images which are difficult to acquire. To address these issues, we propose an unsupervised and unpaired fusion-based super-resolution network, called U2SRnet, which achieves super-resolution without the aid of any supervision nor pairwise images. Instead, an arbitrary low spatial resolution (LR) HSI is used as a fixed hyperspectral input to provide reflectance information for the spectral reconstruction of RGB images. U2SRnet consists of two modules, which are spectral correction module (SCM) and band generation module (BGM) respectively. The former aims to enforce material reflectance consistency between unpaired HS-RGB images, which is implemented via a joint perception attention module (JPAM). Considering the spatial inconsistency of material positions between HS and RGB images, we propose an elaborately-designed material matching mechanism in BGM to obtain HS-RGB block pairs with high material-semantic similarity. Furthermore, a differential migration mechanism guided by physical reflectance priors, universally applicable to unpaired images, is introduced to migrate the material reflectance properties between these block pairs. Extensive experiments were performed on Cave, Harvard and TG1HRSSC datasets. Experimental results validated the effectiveness and superiority of the proposed method.
Haixia Bi, Xuehu Zhu, Bin Pan
IEEE Trans. Geosci. Remote. Sens.2
2025 EDADet: Encoder-Decoder Domain Augmented Alignment Detector for Tiny Objects in Remote Sensing Images
abstract
In recent years, deep learning has shown great potential in object detection applications, but it is still difficult to accurately detect tiny objects with an area proportion of less than 1% in remote sensing images. Most existing studies focus on designing complex networks to learn discriminative features of tiny objects, usually resulting in a heavy computational burden. In contrast, this article proposes an accurate and efficient single-stage detector called EDADet for tiny objects. First, domain conversion technology is used to realize cross-domain multimodal data fusion based on single-modal data input. Then, a tiny object-aware backbone is designed to extract features at different scales. Next, an encoder–decoder feature fusion (EDFF) structure is devised to achieve efficient cross-scale propagation of semantic information. Finally, a center-assist loss and an alignment self-supervised loss are adopted to alleviate the position sensitivity issue and drift of tiny objects. A series of experiments on the AI-TODv2 dataset demonstrate the effectiveness and practicality of our EDADet. It achieves state-of-the-art (SOTA) performance and surpasses the second-best method by 9.65% in AP50 and 4.86% in mAP.
Wenguang Tao, Tian Yan, Haixia Bi
IEEE Trans. Geosci. Remote. Sens.4
2024 Diffusion-Based Generative Self-Supervised Model for Few-Shot PolSAR Image Classification
abstract
With the emergence of deep neural networks, polarimetric synthetic aperture radar (PolSAR) image classification has seen significant advancements in recent years. However, most deep learning-based PolSAR image classification methods heavily rely on a large amount of labeled data which is difficult and expensive to collect for PolSAR. Generative self-supervised learning, which aims to learn representations from unlabeled data via generative auxiliary tasks, is a promising solution to address this issue. Specifically, diffusion models have shown favorable generative capabilities in recent years. Inspired by this, we propose a diffusion-based generative self-supervised model for PolSAR image classification in this work. Via simulating the noise-adding and denoising process, the proposed method is able to learn discriminative and robust polarimetric representations without any annotations, which greatly facilitates the downstream classification. Experiments on the benchmark Flevoland dataset demonstrate the effectiveness of our proposed model.
Zuzheng Kuang, Shirou Jing, Haixia Bi, Yaochen Li
IGARSS4
2024 Dual-Branch PolSAR Image Classification Based On Graphmae And Local Feature Extraction
abstract
The annotation of polarimetric synthetic aperture radar (PolSAR) images is a labor-intensive and time-consuming process. Therefore, classifying PolSAR images with limited labels is a challenging task in remote sensing domain. In recent years, self-supervised learning approaches have proven effective in PolSAR image classification with sparse labels. However, we observe a lack of research on generative self-supervised learning in the studied task. Motivated by this, we propose a dual-branch classification model based on generative self-supervised learning in this paper. The first branch is a superpixel-branch, which learns superpixel-level polarimetric representations using a generative self-supervised graph masked autoencoder. To acquire finer classification results, a convolutional neural networks-based pixel-branch is further incorporated to learn pixel-level features. Classification with fused dual-branch features is finally performed to obtain the predictions. Experimental results on the benchmark Flevoland dataset demonstrate that our approach yields promising classification results.
Haixia Bi, Danfeng Hong
IGARSS3
2024 Cross-Attention-Driven Adaptive Graph Relational Network for Multilabel Remote Sensing Scene Classification
abstract
Multilabel remote sensing scene classification (MLRSSC) has garnered growing attention in recent years, owing to its more comprehensive description of land covers compared to its single-label counterpart. However, challenges arise inevitably. First, the relations among multiple scene labels are sophisticated. How to excavate the interclass dependencies is, therefore, a key challenge for the MLRSSC task. Second, extracting discriminative semantic features is essential, yet challenging for scene prediction of remote sensing images. Another issue is that the multilabel dataset usually shows twofold sample imbalances, that is, class imbalance and positive-negative imbalance, which have not been explored in MLRSSC tasks so far. To overcome the above hurdles, we put forward a cross-attention-driven adaptive graph relational network for the MLRSSC task. Different from the chain-like long short-term memory (LSTM) or static label co-occurrence matrices, we propose to use image-specific relational graphs to dynamically model the interclass dependencies. We innovatively devise a cross-attention-driven representation learning approach, which uses learnable label embeddings to query the class-wise semantic features, explicitly establishing the feature-label connections. Moreover, we design a balanced focal loss (BFL) function, where the loss contributions of positive and negative samples are rebalanced based on the respective imbalance degrees of diverse classes. Extensive experiments were performed on UCM, AID, and DFC15 multilabel datasets. Experimental results demonstrated that our proposed method achieves state-of-the-art performance in the studied task.
Haixia Bi, Honghao Chang, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.1
2024 Deep Symmetric Fusion Transformer for Multimodal Remote Sensing Data Classification
abstract
In recent years, multimodal remote sensing data classification (MMRSC) has evoked growing attention due to its more comprehensive and accurate delineation of Earth’s surface compared to its single-modal counterpart. However, it remains challenging to capture and integrate local and global features from single-modal data. Moreover, how to fully excavate and exploit the interactions between different modalities is still an intricate issue. To this end, we propose a novel dual-branch transformer-based framework named deep symmetric fusion transformer (DSymFuser). Within the framework, each branch contains a stack of local-global mixture (LGM) blocks, to extract hierarchical and discriminative single-modal features. In each LGM block, a local-global feature mixer with learnable weights is specifically devised to adaptively aggregate the local and global features extracted with a convolutional neural network (CNN)–transformer network. Furthermore, we innovatively design a symmetric fusion transformer (SFT) that trails behind each LGM block. The elaborately designed SFT symmetrically facilitates cross-modal correlation excavation, comprehensively exploiting the complementary cues underlying heterogeneous modalities. The hierarchical construction of the LGM and SFT blocks enables feature extraction and fusion in a multilevel manner, further promoting the completeness and descriptiveness of the learned features. We conducted extensive ablation studies and comparative experiments on three benchmark datasets, and the experimental results validated the effectiveness and superiority of the proposed method. The source code of the proposed method will be available publicly athttps://github.com/HaixiaBi1982/DSymFuser.
Honghao Chang, Haixia Bi, Fan Li 0003, Jocelyn Chanussot, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.2
2024 Polarimetry-Inspired Contrastive Learning for Class-Imbalanced PolSAR Image Classification
abstract
In recent years, deep neural networks have significantly boosted the performance of polarimetric synthetic aperture radar (PolSAR) image classification. However, existing deep learning-based approaches still suffer from the following limitations. First, the performance of them is subject to the availability of massive annotations which are difficult to acquire for PolSAR images. Secondly, the class imbalance in PolSAR data greatly hinders the correct classification of minority yet equally pivotal classes. To overcome the above shortcomings, we propose a polarimetry-inspired contrastive learning PolSAR image classification approach, in the hope of elevating the classification accuracy by taking advantage of the polarimetric domain knowledge. Firstly, a complex-valued contrastive learning framework is designed, via which powerful polarimetric representations are learnt without any manual annotations. Specifically, we innovatively design two distribution-inspired positive sample generation strategies, i.e., WishartPSG and NoisePSG, to enable discriminative and domain-specific representation learning. A novel hybrid anti-imbalance scheme is further devised to tackle the class imbalance issue, which combines a contextual consistency-based pseudo-label generation and a weighted feature-level synthetic data over-sampling technique. It should be highlighted that the domain knowledge of PolSAR, including the data and noise distributions, complex-valued characteristics and the spatial consistency prior, is fully exploited throughout our model design. Extensive experiments on four benchmark datasets demonstrated the effectiveness of the proposed model. For the Flevoland 1989 dataset, our method improves the overall accuracy, average accuracy and Kappa metrics by 3.54%, 6.81% and 7.29% respectively, compared to existing state-of-the-art method. Our code will be available at https://github.com/HaixiaBi1982/PiCL.
Zuzheng Kuang, Haixia Bi, Fan Li 0003, Jian Sun 0009
IEEE Trans. Geosci. Remote. Sens.2
2023 Complex-Valued Self-Supervised PolSAR Image Classification Integrating Attention Mechanism
abstract
Polarimetric synthetic aperture radar (PolSAR) image classification is acknowledged as a critical task in remote sensing image processing, and the performance of this task has witnessed a substantial improvement owing to developing deep neural networks. However, existing approaches still suffer from at least one of the following limitations. First, their performance mainly depends on huge amounts of annotations. Secondly, the physical mechanism and characteristics of PolSAR data are not fully exploited in these methods. To tackle the label scarcity issue, we establish an attention mechanism-incorporated self-supervision framework for PolSAR image classification via designing a predictive auxiliary learning task. According to the properties of PolSAR data, we adapt the framework to complex-valued. Additionally, noise injection augmentation scheme considering the speckle noise distribution is designed to enhance the robustness of our model. Involving the complex-valued characteristics of the PolSAR data and noise into the architecture and loss function design makes it essentially distinctive from existing self-supervised methods.
Zuzheng Kuang, Haixia Bi, Fan Li 0003
IGARSS2
2023 Complex-Valued Fully Convolutional Network for PolSAR Image Classification with Noisy Labels
abstract
The process of annotating PolSAR data is highly intricate and demands a proficient understanding of the subject matter. Due to the complexity inherent in this task, it may result in imprecise or erroneous labels being incorporated. The presence of noisy labels inevitably impacts the performance of models in this context, making PolSAR classification a challenging task. This paper proposes a module for correcting noisy labels, which utilizes a CV-CNN with two convolutional layers as its backbone and presents two key contributions: (1) an effective label correction method that leverages the inherent similarities between training samples to repair imprecise or erroneous labels, and (2) a rebalancing loss function that adjusts the weights of different classes to enhance the accuracy of smaller classes. Experimental evaluations on the Flevoland dataset demonstrate the efficacy of our proposed approach.
Ningwei Wang, Haixia Bi
IGARSS2
2023 Tropical Cyclone Intensity Prediction by Spectral-Temporal Dislocation and Attention-Based Networks
abstract
Accurate prediction of Tropical cyclone (TC) intensity using multispectral images (MSIs) is critical to avoid economic loss and life casualty. Although existing methods have achieved good prediction results, they neglect changes in cloud patterns such as cyclone eyes and cloud spirals, which are closely related to TC intensity. How to leverage temporal-spatial-spectral features of MSIs to improve prediction accuracy is challenging task. In this paper, we propose a novel framework with Spectral-Temporal Dislocation and Attention-Based Networks (STD-AN) to predict MSW speed values near cyclone centers. The STD technique allows the framework to learn temporal-spatial-spectral features of TC. Meanwhile, the Self-Attention Modules (SAM) enable global attention feature extraction and Cross-Attention Modules (CAM) fuse different band features to improve prediction accuracy. Experimental results show that the proposed framework outperforms several state-of-the-art methods for TC intensity prediction.
Yahui Xiu, Xinyang Pu, Haixia Bi, Feng Xu 0001
IGARSS4
2023 Noise-Tolerant Unsupervised Classification for PolSAR Images Via Deep Clustering And Markov Random Field
abstract
Due to the difficulty of obtaining manual annotations for polarimetric synthetic aperture radar (PolSAR) images, the problem of analyzing these images without or with few labels has become a current challenge. Considering the scarcity of labels, this paper proposes a noise-tolerant deep clustering-based PolSAR image classification that mainly uses autoencoders to learn discriminative features. In addition, in order to improve the performance and noise resistance of this method, we adopt Markov Random Field (MRF) to enhance the smoothness of class labels. We conducted experiments on a real benchmark PolSAR image, and the results show that our method achieves state-of-the-art PolSAR image classification results without any manual annotations.
Haixia Bi, Danfeng Hong
IGARSS2
2022 PolSAR Image Classification Based on Robust Low-Rank Feature Extraction and Markov Random Field
abstract
Polarimetric synthetic aperture radar (PolSAR) image classification has been investigated vigorously in various remote sensing applications. However, it is still a challenging task nowadays. One significant barrier lies in the speckle effect embedded in the PolSAR imaging process, which greatly degrades the quality of the images and further complicates the classification. To this end, we present a novel PolSAR image classification method that removes speckle noise via low-rank (LR) feature extraction and enforces smoothness priors via the Markov random field (MRF). Especially, we employ the mixture of Gaussian-based robust LR matrix factorization to simultaneously extract discriminative features and remove complex noises. Then, a classification map is obtained by applying a convolutional neural network with data augmentation on the extracted features, where local consistency is implicitly involved, and the insufficient label issue is alleviated. Finally, we refine the classification map by MRF to enforce contextual smoothness. We conduct experiments on two benchmark PolSAR data sets. Experimental results indicate that the proposed method achieves promising classification performance and preferable spatial consistency.
Haixia Bi, Jing Yao 0002, Zhiqiang Wei 0004, Danfeng Hong, Jocelyn Chanussot
IEEE Geosci. Remote. Sens. Lett.1
2022 Gaussian Focal Loss: Learning Distribution Polarized Angle Prediction for Rotated Object Detection in Aerial Images
abstract
With the increasing availability of aerial data, object detection in aerial images has aroused more and more attention in remote sensing community. The difficulty lies in accurately predicting the angular information for each target when using the oriented bounding boxes to represent the arbitrary oriented objects, as the periodicity of the angle could cause inconsistency between target angle values. To resolve the problem, recent works propose to perform angular prediction from a regression problem to a classification task with circular smooth label. However, we find that current loss functions applying to binary soft labels need to approximate the soft label values at each position. When summed over all the negative angle categories, these relatively insignificant loss values can overwhelm the target angle category, thus preventing the network from predicting precise angle information. In this paper, we propose a novel loss function that acts as a more effective alternative to the classification-based rotated detectors. By constructing the classification loss with adaptive Gaussian attenuation on the negative locations, our training objective can not only avoid discontinuous angle boundaries but also enable the network to obtain more accurate angle predictions with higher response at peaks. Moreover, an aspect ratio-aware factor was proposed based on our loss function to enhance the robustness of the model for determining the orientation for square-like objects. Extensive experiments on aerial image datasets DOTA, HRSC2016, and UCAS-AOD demonstrated the effectiveness and superior performances of our approaches.
Jian Wang 0113, Fan Li 0003, Haixia Bi
IEEE Trans. Geosci. Remote. Sens.3
2021 An Enhanced 3-D Discrete Wavelet Transform for Hyperspectral Image Classification
abstract
In the classification of hyperspectral image (HSI), there exists a common issue that the collected HSI data set is always contaminated by various noise (e.g., Gaussian, stripe, and deadline), degrading the classification results. To tackle this issue, we modify the 3-dimensional discrete wavelet transform (3DDWT) method by considering the noise effect on feature quality and propose an enhanced 3DDWT (E-3DDWT) approach to extract the feature and meanwhile alleviate the noise. Specifically, the proposed E-3DDWT method first applies classical 3DDWT method to the HSI data cube and thus can generate eight subcubes in each level. Then, the stripe noise is concentrated into several subcubes due to its spatial vertical property. Finally, we abandon these subcubes and obtain the feature cube by stacking the remaining ones. After acquiring the feature, we then adopt the convolutional neural network (CNN) model with an active learning strategy for classification since CNN has been verified to be a state-of-the-art feature extraction method for HSI classification, and active learning strategy can alleviate the insufficient labeled sample issue to some extent. In addition, we apply the Markov random field to enhance the final categorized results. Experiments on two synthetically striped data sets show that our proposed approach achieves better categorized results than other advanced methods.
Xiangyong Cao, Jing Yao 0002, Xueyang Fu, Haixia Bi, Danfeng Hong
IEEE Geosci. Remote. Sens. Lett.4
2021 Human Activity Recognition Based on Dynamic Active Learning
abstract
Activity of daily living is an important indicator of the health status and functional capabilities of an individual. Activity recognition, which aims at understanding the behavioral patterns of people, has increasingly received attention in recent years. However, there are still a number of challenges confronting the task. First, labelling training data is expensive and time-consuming, leading to limited availability of annotations. Secondly, activities performed by individuals have considerable variability, which renders the generally used supervised learning with a fixed label set unsuitable. To address these issues, we propose a dynamic active learning-based activity recognition method in this work. Different from traditional active learning methods which select samples based on a fixed label set, the proposed method not only selects informative samples from known classes, but also dynamically identifies new activities which are not included in the predefined label set. Starting with a classifier that has access to a limited number of labelled samples, we iteratively extend the training set with informative labels by fully considering the uncertainty, diversity and representativeness of samples, based on which better-informed classifiers can be trained, further reducing the annotation cost. We evaluate the proposed method on two synthetic datasets and two existing benchmark datasets. Experimental results demonstrate that our method not only boosts the activity recognition performance with considerably reduced annotation cost, but also enables adaptive daily activity analysis allowing the presence and detection of novel activities and patterns.
Haixia Bi, Miquel Perelló-Nieto, Raúl Santos-Rodríguez, Peter A. Flach
IEEE J. Biomed. Health Informatics1
2020 Polsar Image Classification via Robust Low-Rank Feature Extraction and Markov Random Field
abstract
Polarimetric synthetic aperture radar (PolSAR) image classification has been investigated rigorously in various remote sensing applications. However, it is still a challenging task nowadays. One significant barrier lies in the speckle effect embedded in the PolSAR imaging process, which significantly degrades the quality of the images and further complicates the classification. To address this issue, we present a novel Pol-SAR image classification method which removes speckle noise via robust low-rank feature extraction and enforces smoothness priors through Markov Random Field (MRF). Specifically, we employ the mixture of Gaussian (MoG) based low-rank matrix factorization (LRMF) to simultaneously extract robust features and remove noise. Then, a classification map is obtained by applying Random Forest (RF) classifier on the extracted LRMF features. Finally, we refine the classification map by Markov random field (MRF) to enforce contextual smoothness. We conduct experiments on two real benchmark PolSAR data sets. Experimental results indicate that the proposed method achieves promising classification performance and preferable spatial consistency.
Haixia Bi, Raúl Santos-Rodríguez, Peter A. Flach
IGARSS1
2020 Polarimetric SAR Image Semantic Segmentation With 3D Discrete Wavelet Transform and Markov Random Field
abstract
Polarimetric synthetic aperture radar (PolSAR) image segmentation is currently of great importance in image processing for remote sensing applications. However, it is a challenging task due to two main reasons. Firstly, the label information is difficult to acquire due to high annotation costs. Secondly, the speckle effect embedded in the PolSAR imaging process remarkably degrades the segmentation performance. To address these two issues, we present a contextual PolSAR image semantic segmentation method in this paper. With a newly defined channel-wise consistent feature set as input, the three-dimensional discrete wavelet transform (3D-DWT) technique is employed to extract discriminative multi-scale features that are robust to speckle noise. Then Markov random field (MRF) is further applied to enforce label smoothness spatially during segmentation. By simultaneously utilizing 3D-DWT features and MRF priors for the first time, contextual information is fully integrated during the segmentation to ensure accurate and smooth segmentation. To demonstrate the effectiveness of the proposed method, we conduct extensive experiments on three real benchmark PolSAR image data sets. Experimental results indicate that the proposed method achieves promising segmentation accuracy and preferable spatial consistency using a minimal number of labeled pixels.
Haixia Bi, Lin Xu 0001, Xiangyong Cao, Yong Xue, Zongben Xu
IEEE Trans. Image Process.1
2019 Unsupervised PolSAR Image Factorization with Deep Convolutional Networks
abstract
This paper presents a novel unsupervised polarimetric synthetic aperture radar (PolSAR) image classification method, which incorporates polarimetric image factorization and deep convolutional networks into a principled framework. To implement this idea, we design a convolutional neural network (CNN) with a newly defined loss function which measures the probability distribution distance between the initial distribution maps and CNN predictions. In the proposed method, we firstly execute polarimetric image factorization to generate a dictionary of meaningful atom scatters and their corresponding distribution maps, where the strongest scatters are selected as training samples for CNN. Next, we train the CNN by iteratively optimizing the defined energy function, producing the final distribution maps and classification result. The proposed approach is applied on a real UAVSAR image. Experimental results justify that our approach can effectively classify the PolSAR image in an unsupervised way and produce favorable classification results.
Haixia Bi, Feng Xu 0001, Zhiqiang Wei 0004, Yibo Han, Yuanlong Cui, Yong Xue, Zongben Xu
IGARSS1
2019 An Active Deep Learning Approach for Minimally-Supervised Polsar Image Classification
abstract
Aiming at improving the classification performance with greatly reduced annotation cost, this paper presents an active deep learning approach for minimally-supervised PolSAR image classification, which integrates active learning and fine-tuning convolutional neural network (CNN) into a principled framework. Starting from a CNN trained using a very limited number of labeled pixels, we iteratively and actively select the most informative candidates for annotation, and incrementally fine-tune the CNN by incorporating the newly annotated pixels. Moreover, to boost the performance and robustness of the proposed method, we employ Markov random field to enforce label smoothness, and data augmentation technique to enlarge the training set. Extensive experiments demonstrated that our approach achieved state-of-the-art classification results with significantly reduced annotation cost.
Haixia Bi, Feng Xu 0001, Zhiqiang Wei 0004, Yibo Han, Yuanlong Cui, Yong Xue, Zongben Xu
IGARSS1
2019 A Graph-Based Semisupervised Deep Learning Model for PolSAR Image Classification
abstract
Aiming at improving the classification accuracy with limited numbers of labeled pixels in polarimetric synthetic aperture radar (PolSAR) image classification task, this paper presents a graph-based semisupervised deep learning model for PolSAR image classification. It models the PolSAR image as an undirected graph, where the nodes correspond to the labeled and unlabeled pixels, and the weighted edges represent similarities between the pixels. Upon the graph, we design an energy function incorporating a semisupervision term, a convolutional neural network (CNN) term, and a pairwise smoothness term. The employed CNN extracts abstract and data-driven polarimetric features and outputs class label predictions to the graph model. The semisupervision term enforces the category label constraints on the human-labeled pixels. The pairwise smoothness term encourages class label smoothness and the alignment of class label boundaries with the image edges. Starting from an initialized class label map generated based on K-Wishart distribution hypothesis or superpixel segmentation of PauliRGB images, we iteratively and alternately optimize the defined energy function until it converges. We conducted experiments on two real benchmark PolSAR images, and extensive experiments demonstrated that our approach achieved the state-of-the-art results for PolSAR image classification.
Haixia Bi, Jian Sun 0009, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.1
2019 An Active Deep Learning Approach for Minimally Supervised PolSAR Image Classification
abstract
Recently, deep neural networks have received intense interests in polarimetric synthetic aperture radar (PolSAR) image classification. However, its success is subject to the availability of large amounts of annotated data which require great efforts of experienced human annotators. Aiming at improving the classification performance with greatly reduced annotation cost, this paper presents an active deep learning approach for minimally supervised PolSAR image classification, which integrates active learning and fine-tuned convolutional neural network (CNN) into a principled framework. Starting from a CNN trained using a very limited number of labeled pixels, we iteratively and actively select the most informative candidates for annotation, and incrementally fine-tune the CNN by incorporating the newly annotated pixels. Moreover, to boost the performance and robustness of the proposed method, we employ Markov random field (MRF) to enforce class label smoothness, and data augmentation technique to enlarge the training set. We conducted extensive experiments on four real benchmark PolSAR images, and experiments demonstrated that our approach achieved state-of-the-art classification results with significantly reduced annotation cost.
Haixia Bi, Feng Xu 0001, Zhiqiang Wei 0004, Yong Xue, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.1
2017 Polsar image classification based on three-dimensional wavelet texture features and Markov random field
abstract
The speckle effect embedded in polarimetric synthetic aperture radar (PolSAR) data damages the performance of PolSAR image classification greatly. To alleviate this issue, a new supervised classification method, which introduces spatial consistency in both feature extraction and classification steps is proposed. Specifically, three-dimensional discrete wavelet transform (3D-DWT) is used to extract spectral-spatial texture features, which are proved to be more discriminative than original ones. Afterward, label smoothness prior is incorporated in the classification, which is implemented using a Markov random field (MRF). To demonstrate the validity of the proposed method, real PolSAR image is used in experiments. Compared with the other state-of-the-art methods, this method achieves higher classification accuracy and better visual spatial connectivity.
Haixia Bi, Lin Xu 0001, Xiangyong Cao, Zongben Xu
IGARSS1
2017 Unsupervised PolSAR Image Classification Using Discriminative Clustering
abstract
This paper presents a novel unsupervised image classification method for polarimetric synthetic aperture radar (PolSAR) data. The proposed method is based on a discriminative clustering framework that explicitly relies on a discriminative supervised classification technique to perform unsupervised clustering. To implement this idea, we design an energy function for unsupervised PolSAR image classification by combining a supervised softmax regression model with a Markov random field smoothness constraint. In this model, both the pixelwise class labels and classifiers are taken as unknown variables to be optimized. Starting from the initialized class labels generated by Cloude-Pottier decomposition and $K$ -Wishart distribution hypothesis, we iteratively optimize the classifiers and class labels by alternately minimizing the energy function with respect to them. Finally, the optimized class labels are taken as the classification result, and the classifiers for different classes are also derived as a side effect. We apply this approach to real PolSAR benchmark data. Extensive experiments justify that our approach can effectively classify the PolSAR image in an unsupervised way and produce higher accuracies than the compared state-of-the-art methods.
Haixia Bi, Jian Sun 0009, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.1