EDBT 2026 Demo / reviewers in the wild / expert
Zaidao Wen
dblp:142/6366
· DBLP profile ↗
26ranked-venue papers
12as first author
12since 2021 · last 2025
0000-0003-1258-7737ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CLIP-Based Modality Compensation for Visible-Infrared Image Re-IdentificationabstractVisible-infrared image re-identification (VIReID) aims to match objects with the same identity appearing across different modalities. Given the significant differences between visible and infrared images, VIReID poses a formidable challenge. Most existing methods focus on extracting modality-shared features while ignore modality-specific features, which often also contain crucial important discriminative information. In addition, high-level semantic information of the objects, such as shape and appearance, is also crucial for the VIReID task. To further enhance the retrieval performance, we propose a novel one-stage CLIP-based Modality Compensation (CLIP-MC) method for the VIReID task. Our method introduces a new prompt learning paradigm that leverages the semantic understanding capabilities of CLIP to recover missing modality information. CLIP-MC comprises three key modules: Instance Text Prompt Generation (ITPG), Modality Compensation (MC), and Modality Context Learner (MCL). Specifically, the ITPG module facilitates effective alignment and interaction between image tokens and text tokens, enhancing the text encoder's ability to capture detailed visual information from the images. This ensures that the text encoder generates fine-grained descriptions of the images. The MCL module captures the unique information of each modality and generates modality-specific context tokens, which are more flexible compared to fixed text descriptions. Guided by the modality-specific context, the text encoder discovers missing modality information from the images and produces compensated modality features. Finally, the MC module combines the original and compensated modality features to obtain complete modality features that contain more discriminative information. We conduct extensive experiments on three VIReID datasets and compare the performance of our method with other existing approaches to demonstrate its effectiveness and superiority. Gang Hu 0008, Yafei Lv, Zaidao Wen |
IEEE Trans. Multim. | 5 |
| 2024 | Rendering-Inspired Cross-Source Feature Disentanglement for Domain Adaptation- Based SAR Ship ClassificationabstractThis letter presents a novel generative model-based domain adaptation framework for synthetic aperture radar (SAR) ship classification. The framework introduces a unified image representation model tailored for multiple sensor observations, without distinguishing between the source and target domains. Two types of feature encoders are designed, leveraging domain-level and category-level supervised signals to extract ship-specific and source-specific factors. The paper elaborates on a novel conditional rendering-inspired imaging function that utilizes a conditional convolutional neural network and the adaptive instance normalization module to simulate the interaction effects between rays and the ship in the image rendering algorithm. By modulating both appearance and statistics, the imaging function captures textural distinctions and statistical variances between SAR and optical images in a physical plausible way. Furthermore, we develop an efficient intervention-based regularization strategy for feature disentanglement through a category source dual aware discriminators module. The effectiveness of the proposed model is demonstrated through experimental results on the FUSAR-Ship dataset, assisted with optical ship images from FGSCR42. Comparative analysis with other domain adaptation algorithms and several typical deep classification architectures demonstrates the superior classification accuracy achieved by our model. Yafei Lv, Zaidao Wen |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Global-Local Information Soft-Alignment for Cross-Modal Remote-Sensing Image-Text RetrievalabstractCross-modal remote-sensing image-text retrieval (CMRSITR) is a challenging task that aims to retrieve target remote-sensing (RS) images based on textual descriptions. However, the modal gap between texts and RS images poses a significant challenge. RS images comprise multiple targets and complex backgrounds, necessitating the mining of both global and local information for effective CMRSITR. Existing approaches primarily focus on local image features while disregarding the local features of the text and their correspondence. These methods typically fuse global and local image features and align them with global text features. However, they struggle to eliminate the influence of cluttered backgrounds and may overlook crucial targets. To address these limitations, we propose a novel framework for CMRSITR based on a transformer architecture, which leverages global-local information soft alignment (GLISA) to enhance retrieval performance. Our framework incorporates a global image extraction module, which captures the global semantic features of image-text pairs and effectively represents the relationships among multiple targets in RS images. Additionally, we introduce an adaptive local information extraction module that adaptively mines discriminative local clues from both RS images and texts, aligning the corresponding fine-grained information. To mitigate semantic ambiguities during the alignment of local features, we design a local information soft-alignment module. In comparative evaluations using two public CMRSITR datasets, our proposed method achieves state-of-the-art results, surpassing not only traditional cross-modal retrieval methods by a substantial margin, but also other CLIP-based methods. Gang Hu 0008, Zaidao Wen, Yafei Lv |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Multimodal Discriminative Feature Learning for SAR ATR: A Fusion Framework of Phase History, Scattering Topology, and ImageabstractDeep learning has emerged as the dominant paradigm for synthetic aperture radar (SAR) automatic target recognition (ATR), which induces discriminative visual features from target images. Due to the imaging mechanism, the SAR image is a reconstruction of the interaction between the electromagnetic wave and the target, which intrinsically entangles the radar characteristics, target scattering properties, and visual signatures. However, current learning algorithms focus primarily on the visual signatures without explicit specification or ineffective integration of the full range of SAR features into the learning process, leading to a weak generalization ability across different radar systems and imaging conditions. To address this issue, we propose a novel multimodal feature fusion learning framework, which captures a comprehensive set of target features from different domains for enhanced complementarity. First, we encode the target response and radar characteristics into the phase-history data. A cross-direction sequence learning module is designed to extract their range and azimuth dependence. Next, a hierarchical graph node-aggregating neural network is developed to learn the scattering topology features from scattering points according to the part-to-whole learning bias. Finally, the learned features from the phase-history and scattering domains are fused with the image features obtained from an off-the-shelf deep feature extractor for final target recognition. Experiments on the moving and stationary target acquisition and recognition (MSTAR) benchmark demonstrate its effectiveness. Compared with the other SAR ATR algorithms, our approach can achieve the state-of-the-art recognition accuracy without using any data augmentation trick, especially in cases of limited training samples. Zaidao Wen, Youlan Yu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Cross-Modality Vessel Re-Identification With Deep Alignment Decomposition NetworkabstractCross-modality vessel re-identification (ReID) presents a formidable challenge in the domain of maritime surveillance, necessitating the development of robust methodologies to accurately match vessels across disparate imaging modalities. This paper introduces a novel Cross-modality Alignment Decomposition Network (CAD-Net) to address the inherent complexities associated with this task. CAD-Net incorporates a geometric-semantic cross-modal alignment module for effectively mitigating geometric and modality variances within the global features. Additionally, it integrates an adaptive local decomposition module associated with a diversity regularization, enabling the capture of local vessel features, all while circumventing the reliance on predefined part separation criteria. To address the scarcity of cross-modal vessel datasets, which are predominantly biased towards visible light modality, and to evaluate the performance of the proposed framework, we have constructed a novel dataset named KongTong-boat (KT-boat). It comprises 2,826 high-resolution images, including 1,443 RGB images and 1,383 IR images, featuring 117 distinct vessels. This dataset can be served as a new fundamental benchmark for evaluating the efficacy of cross-modality vessel ReID algorithms, filling a critical gap in the field. The experimental results obtained on the KT-boat dataset unequivocally demonstrate the remarkable effectiveness of CAD-Net in the context of cross-modality ReID. Notably, when compared to state-of-the-art cross-modality ReID algorithms applied to general cross-modality pedestrian benchmarks on KT-boat and RegDB dataset, CAD-Net consistently outperforms them across key evaluation metrics, including the rank-1 index and mean Average Precision (mAP). Zaidao Wen, Yafei Lv |
IEEE Trans. Multim. | 1 |
| 2023 | View-Semantic Transformer With Enhancing Diversity for Sparse-View SAR Target RecognitionabstractWith the rapid development of supervised learning-based SAR target recognition technology, it is easy to find that the recognition performance is proportional to the amount of training samples. However, the biased data distribution and under-representation of the model caused by incomplete data within categories exacerbate the challenge of SAR interpretation. In this paper, we propose a new view-semantic transformer network (VSTNet) that generates synthesized samples to complete the statistical distribution of training data and improve the discriminative representation of the model. First, SAR images from different views are encoded into a disentangled latent space, which allows us to synthesize data with more diverse views by manipulating view-semantic features. Second, the synthesized data as a complement effectively expands the training set and alleviates the overfitting problem of limited data in sparse views. Third, the proposed method unifies SAR image synthesis and SAR target recognition into an end-to-end framework to boost their performance against each other. Experiments conducted on moving and stationary target acquisition and recognition (MSTAR) data demonstrate the robustness and effectiveness of the proposed method. Zhunga Liu, Feiyan Wu, Zaidao Wen, Zuowei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Contrastive Feature Disentangling for Partial Aspect Angles SAR Noncooperative Target RecognitionabstractDeep learning algorithms have achieved state-of-the-art progress in synthetic aperture radar (SAR) automatic target recognition (ATR) tasks. They theoretically assume that training and test samples are independent and identically distributed (i.i.d.) for generalization, but it is intractable for practical ATR scenarios. In this paper, we propose a novel contrastive feature disentangling framework termed ConFeDent to learn features with improved generalization performance under a condition of a weaker distribution consistency. More specifically, ConFeDent aims to describe the semantic interactions between two arbitrary SAR training samples instead of treating them independently. It can implicitly disentangle features encoding the pose and identity knowledge from the whole samples with a semi-parametric geometric transformation model and a second-order energy model. In particular, except for the identity label, we use deductive-based geometry knowledge as supervision to teach the model to learn the concept of aspect angle variation. A progressively amortized inference scheme is constructed for efficient feature learning and recognition in an end-to-end manner. Finally, we further release a strengthened version, called ConFeDent+, which can explicitly utilize and learn more information from cross-category samples. Experimental results on the moving and stationary target acquisition and recognition (MSTAR) benchmark demonstrate the effectiveness of our proposed models in the SAR ATR. In particular, we validate the algorithms in a more challenging scenario where the range of aspect angles for training and testing samples is permitted to be disparate. Our model can achieve much higher recognition accuracy than other SAR ATR algorithms. Zaidao Wen, Zhunga Liu, Sijian Li, Quan Pan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | A Statistical-Spatial Feature Learning Network for PolSAR Image ClassificationabstractPolarimetric synthetic aperture radar (PolSAR) image classification plays an important role in the development of remote sensing image interpretation for the rich polarization information. Generative methods learn the statistical distribution characteristic of the scattered echoes data with heavy noise. However, it is tedious to design and solve different generative likelihood functions. Discriminative methods learn image features and classifier in an end-to-end framework whose classification performance is also limited without polarization statistical characteristics. In order to make full use of the statistical characteristic of PolSAR echoes and spatial feature of PolSAR image, a novel real-value hybrid generative/discriminative (HGD) deep network is proposed to learn statistical-spatial feature for PolSAR image classification with the data expressing ability of the generative term and the end-to-end learning ability of the discriminative term. First, a derived replacement form is obtained from the typical Wishart distribution of PolSAR data with an eigenvalue generative term. In this way, the complex-value form of PolSAR data converts into real-value expression, which makes the PolSAR image classification easy to understand and implement in the deep network. Finally, these two terms are integrated from the variational Bayesian theory with an alternative optimization equation derived for PolSAR image classification. It provides a normal framework for PolSAR multifeature learning. Experiments are tested on different PolSAR datasets compared with several state-of-the-art individual generative or discriminative methods. All visual experimental results and accuracy values demonstrate the superiority of the proposed method. Zaidao Wen, Yanbo Luo |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Multilevel Scattering Center and Deep Feature Fusion Learning Framework for SAR Target RecognitionabstractIn synthetic aperture radar (SAR) automatic target recognition (ATR), there are mainly two types of methods: physics-driven model and data-driven network. The physics-driven model can exploit electromagnetic theory to obtain physical properties, while the data-driven network will extract deep discriminant feature of targets. These two types of features represent the target characteristics in scattering domain and image domain, respectively. However, the representation discrepancy caused by the different modalities between them hinders the further comprehensive utilization and fusion of both features. In order to take full advantage of physical knowledge and deep discriminant feature for SAR ATR, we propose a new feature fusion learning framework SDF-Net to combine scattering and deep image features. In this work, we treat the attributed scattering centers (ASC) as set-data instead of multiple individual points, which can well mine the topological interaction among scatterers. Then multi-region multi-scale sub-sets are constructed at both component and target levels. To be specific, the most significant scattering intensity and overall representation in these sub-sets are exploited successively to learn permutation-invariant scattering features according to a set-oriented deep network. The scattering representations can provide mid-level semantic and structural features that are subsequently fused with the complementary deep image features to yield an end-to-end high-level feature learning framework, which helps enhance the generalization ability of networks especially under complex observation conditions. Extensive experiments on Moving and Stationary Target Acquisition and Recognition database verify the effectiveness and robustness of the SDF-Net compared against both typical SAR ATR networks and ASC-based models. Zhunga Liu, Zaidao Wen, Kun Li 0002, Quan Pan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Evidential Combination of Classifiers for Imbalanced DataabstractIt remains an important research topic for the classification of imbalanced data. There exist some methods to solve this problem, such as hybrid-sampling, over-sampling, and under-sampling. Each method has its own advantage, and different methods generally provide some complementary knowledge. We want to combine these three methods at the decision level in an appropriate way for achieving as good as possible classification performance. Evidence theory is expert at representing and combining uncertain information. So a new method called an evidential combination of classifiers (ECC) is proposed for dealing with imbalanced data. The classification result generated by different strategies (i.e., hybrid-sampling, over-sampling, or under-sampling) may have different reliabilities for query patterns. A cautious reliability evaluation rule is developed for each classification result based on the close neighborhoods. After that, the classification result is revised with a new belief redistribution way according to the reliability evaluation, and the probability/belief of one class can be partially transferred to other classes as well as the total ignorance, which is defined by the whole frame of classes. By doing this, we can reduce the error risk of each classification method. Then, the revised classification results from different methods are combined by evidence theory to make the final class decision. The effectiveness of the ECC method has been demonstrated using several experiments, and it shows that ECC can effectively improve the classification performance comparing with other related methods. Jiawei Niu, Zhunga Liu, Yao Lu 0010, Zaidao Wen |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | Cost-Sensitive Latent Space Learning for Imbalanced PolSAR Image ClassificationabstractLand cover classification is an important application for polarimetric synthetic aperture radar (PolSAR) image interpretation. The classification performance of a promising parametric feature and classifier learning-based algorithm is limited when the amounts of pixels from different classes vary greatly. PolSAR data from minority classes is difficult to recognize correctly owing to a strong learning bias toward the majority classes, resulting in under-performing features for minority classes. To address this issue, a cost-sensitive latent space learning network based on the feature and classifier learning framework is proposed to reduce the learning bias for supporting the classification of imbalanced data in PolSAR images. First, a new cost-sensitive method is developed by adaptively computing the cost coefficient from predicted labels in the optimization process. Thus, the imbalanced distribution of PolSAR data can be obtained for both the labeled and unlabeled pixels rather than a predefined misclassifying matrix for labeled pixels. Second, latent space learning is used as an auxiliary task to assist the main task of classifier learning. By weighting the distance between the learned feature and the basis of the latent space with a different cost-sensitive coefficient, pixels in minority and majority classes are promoted to be more separable. Thus, the strong bias to majority classes is reduced from both the feature learning and classification process. Finally, the proposed method is studied through experiments on three different PolSAR images with several existing state-of-the-art methods. The experiments validate the effectiveness of the proposed method for balanced and imbalanced PolSAR land cover classification. Biao Hou, Zaidao Wen, Zhongle Ren, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Rotation Awareness Based Self-Supervised Learning for SAR Target Recognition With Limited Training SamplesabstractThe scattering signatures of a synthetic aperture radar (SAR) target image will be highly sensitive to different azimuth angles/poses, which aggravates the demand for training samples in learning-based SAR image automatic target recognition (ATR) algorithms, and makes SAR ATR a more challenging task. This paper develops a novel rotation awareness-based learning framework termed RotANet for SAR ATR under the condition of limited training samples. First, we propose an encoding scheme to characterize the rotational pattern of pose variations among intra-class targets. These targets will constitute several ordered sequences with different rotational patterns via permutations. By further exploiting the intrinsic relation constraints among these sequences as the supervision, we develop a novel self-supervised task which makes RotANet learn to predict the rotational pattern of a baseline sequence and then autonomously generalize this ability to the others without external supervision. Therefore, this task essentially contains a learning and self-validation process to achieve human-like rotation awareness, and it serves as a task-induced prior to regularize the learned feature domain of RotANet in conjunction with an individual target recognition task to improve the generalization ability of the features. Extensive experiments on moving and stationary target acquisition and recognition benchmark database demonstrate the effectiveness of our proposed framework. Compared with other state-of-the-art SAR ATR algorithms, RotANet will remarkably improve the recognition accuracy especially in the case of very limited training samples without performing any other data augmentation strategy. Zaidao Wen, Zhunga Liu, Quan Pan 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | A Distribution and Structure Match Generative Adversarial Network for SAR Image ClassificationabstractSynthetic aperture radar (SAR) image classification is a fundamental research in the interpretation of SAR images. The previous methods are unilaterally based on statistical features or spatial features, which cannot capture features with complete SAR image characteristics and unavoidably limits the performance for classification. In this article, novel sample weighting and class adversarial training strategies are proposed to fuse complementary SAR characteristics. Based on these, a distribution and structure match auxiliary classifier generative adversarial network (DSM-ACGAN) is constructed for high-quality discriminative feature learning. Particularly, the characteristics of statistical distribution and spatial structure are jointly considered in class adversarial training of DSM-ACGAN. On the one hand, DSM-ACGAN sets the true SAR image characteristics as goals for the generator to learn generative models of each category. On the other hand, and more importantly, it guides the discriminator to simultaneously capture the desired statistical and structural features. Through the class adversarial processing, the discriminative feature learning progressively improves and contributes to classification. Additionally, class-balanced and plausible samples can be generated. Experimental results on three broad SAR images from different satellites confirm the effectiveness of class adversarial training and the superiority of discriminative feature learning in DSM-ACGAN. Visual performance and quantitative metrics also show the state-of-the-art performance of the novel model. Zhongle Ren, Biao Hou, Zaidao Wen, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Rotation Awareness Based Self-Supervised Learning for SAR Target RecognitionabstractIn this paper, we newly suggest that more attention should be paid on learning rotation-equivariant and label-invariant features for each target instead of the conventional rotation-invariant ones. To achieve this goal, we present a novel rotation awareness based self-supervised learning (RR-SSL) deep model to recognize the behavior of target rotation, which is also benefit from the discriminative training scheme without manual labeling. Then this model is incorporated into another deep discriminative model of target recognition to form a dual-task learning framework, where their bottom layers are shared to capture the expected features. Sufficient experimental results on moving and stationary target acquisition and recognition (MSTAR) database demonstrate the effectiveness of our proposed model. The overall framework can achieve a better or comparative recognition accuracy compared with other state-of-the-art SAR-ATR algorithms. Zaidao Wen, Zhunga Liu, Quan Pan 0001 |
IGARSS | 2 |
| 2019 | Polar-Spatial Feature Fusion Learning With Variational Generative-Discriminative Network for PolSAR ClassificationabstractFeature learning-based polarimetric synthetic aperture radar (PolSAR) classification model will generally suffer from the challenge of deficient labeled pixels. In this paper, we propose a novel generative-discriminative network for PolSAR polar-spatial feature fusion learning and classification, which comprises of a deep generative network and a discriminative network with their bottom layers shared. With this architecture, it enables to make use of both labeled and unlabeled pixels in a PolSAR image for model learning in a semisupervised way. Moreover, the proposed network imposes a Gaussian random field prior and a conditional random field posterior on the learned fusion features and the output label configuration, respectively. Without the need of the complicated recurrent iterations, our network can still efficiently produce the structured fusion feature as well as a smoothed classification map by involving some auxiliary variables, and it is specifically optimized via variational inference within an alternating direction method of multipliers iteration scheme. Extensive experiments on different benchmark PolSAR imageries demonstrate the effectiveness and superiority of the proposed network. Compared with other state-of-the-art algorithms of PolSAR feature learning and classification, our model can achieve a much better performance in terms of the visual quality of the label map and overall classification accuracy, facilitating the much less labeling pixels. Zaidao Wen, Zhunga Liu, Quan Pan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Variational Learning of Mixture Wishart Model for PolSAR Image ClassificationabstractThe phase difference, amplitude product, and amplitude ratio between two polarizations are important discriminators for terrain classification, which derives a significant statistical-distribution-based polarimetric synthetic aperture radar (PolSAR) image classification. Traditionally, statistical-distribution-based PolSAR image classification models pay attention to two aspects: searching for a suitable distribution to model certain PolSAR image and a satisfactory solution for the corresponding distribution model with samples in every terrain. Usually, the described distribution form is too complicated to build. Besides, inaccurate parameter estimation may lead to poor classification performance for PolSAR image. In order to refrain from this phenomenon, a variational thought is adopted for the statistical-distribution-based PolSAR classification method in this paper. First, a mixture Wishart model is built to model the PolSAR image to replace the complicated distribution for the PolSAR image. Second, a learning-based method is suggested instead of inaccurate point estimation of parameters to determine the distribution for every class in the mixture Wishart model. Finally, the proposed learning-based mixture Wishart model will be built as a variational form to realize a parametric model for PolSAR image classification. In the experiments, it will be proved that the class centers are easier to distinguish among different terrains learned from the proposed variational model. In addition, a classification performance on the PolSAR image is superior to the original point estimation Wishart model on both visual classification result and accuracy. Biao Hou, Zaidao Wen, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Discriminative Feature Learning for Real-Time SAR Automatic Target Recognition With the Nonlinear Analysis Cosparse ModelabstractThis letter presents an efficient application of the nonlinear analysis cosparse model (NACM) to the task of real-time synthetic aperture radar automatic target recognition (ATR). In contrast to the conventional synthesis sparse representation model, NACM enables efficient sparse feature extraction and selection using a feed-forward mechanism. Furthermore, NACM does not require a sparsity-inducing regularizer. This model uses a task-driven learning framework, in which a naive Bayes or a discriminative classifier is adaptively learned along with the regularized features. Experimental results with the moving and stationary target acquisition and recognition benchmark demonstrate the effectiveness and efficiency of our proposed approach. Compared with traditional classification algorithms using sparse representation, our approach not only achieves higher or comparable recognition accuracy but also dramatically reduces the execution time for real-time ATR. Zaidao Wen, Biao Hou, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Discriminative Transformation Learning for Fuzzy Sparse Subspace ClusteringabstractThis paper develops a novel iterative framework for subspace clustering (SC) in a learned discriminative feature domain. This framework consists of two modules of fuzzy sparse SC and discriminative transformation learning. In the first module, fuzzy latent labels containing discriminative information and latent representations capturing the subspace structure will be simultaneously evaluated in a feature domain. Then the linear transforming operator with respect to the feature domain will be successively updated in the second module with the advantages of more discrimination, subspace structure preservation, and robustness to outliers. These two modules will be alternatively carried out and both theoretical analysis and empirical evaluations will demonstrate its effectiveness and superiorities. In particular, experimental results on three benchmark databases for SC clearly illustrate that the proposed framework can achieve significant improvements than other state-of-the-art approaches in terms of clustering accuracy. Zaidao Wen, Biao Hou, Licheng Jiao |
IEEE Trans. Cybern. | 1 |
| 2018 | Target-Oriented High-Resolution SAR Image Formation via Semantic Information Guided RegularizationsabstractSparsity-regularized synthetic aperture radar (SAR) imaging framework has shown its remarkable performance to generate a feature-enhanced high-resolution image, in which a sparsity-inducing regularizer is involved by exploiting the sparsity priors of some visual features in the underlying image. However, since the simple prior of low-level features is insufficient to describe different semantic contents in the image, this type of regularizer will be incapable of distinguishing between the target of interest and unconcerned background clutters. As a consequence, the features belonging to the target and clutters are simultaneously affected in the generated image without concerning their underlying semantic labels. To address this problem, we propose a novel semantic information guided generative framework for target-oriented SAR image formation, which aims at enhancing the interested target scatters while suppressing the background clutters. First, we develop a new semantics-specific regularizer for image formation by exploiting the statistical properties of different semantic categories in a target scene SAR image. In order to infer the semantic label for each pixel in an unsupervised way, we moreover induce a novel high-level prior-driven regularizer and some semantic causal rules from the prior knowledge. Finally, our regularized framework for image formation is further derived as a simple iteratively reweighted $\ell _{1}$ minimization problem that can be conveniently solved by many off-the-shelf solvers. Experimental results demonstrate the effectiveness and superiority of our framework for SAR image formation in terms of target enhancement and clutters suppression, compared with the state of the arts. Additionally, the proposed framework opens a new direction of devoting some machine learning strategies to image formation, which can benefit the subsequent decision-making tasks. Biao Hou, Zaidao Wen, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Joint Sparse Recovery With Semisupervised MUSICabstractDiscrete multiple signal classification (MUSIC) with its low computational cost and mild condition requirement becomes a significant noniterative algorithm for joint sparse recovery (JSR). However, it fails in rank defective problem caused by coherent or limited amount of multiple measurement vectors (MMVs). In this letter, we provide a novel sight to address this problem by interpreting JSR as a binary classification problem with respect to atoms. Meanwhile, MUSIC essentially constructs a supervised classifier based on the labeled MMVs so that its performance will heavily depend on the quality and quantity of these training samples. From this viewpoint, we develop a semisupervised MUSIC (SS-MUSIC) in the spirit of machine learning, which declares that the insufficient supervised information in the training samples can be compensated from those unlabeled atoms. Instead of constructing a classifier in a fully supervised manner, we iteratively refine a semisupervised classifier by exploiting the labeled MMVs and some reliable unlabeled atoms simultaneously. Through this way, the required conditions and iterations can be greatly relaxed and reduced. Numerical experimental results demonstrate that SS-MUSIC can achieve much better recovery performances than other MUSIC extended algorithms as well as some typical greedy algorithms for JSR in terms of iterations and recovery probability. Zaidao Wen, Biao Hou, Licheng Jiao |
IEEE Signal Process. Lett. | 1 |
| 2017 | Discriminative Dictionary Learning With Two-Level Low Rank and Group Sparse Decomposition for Image ClassificationabstractDiscriminative dictionary learning (DDL) framework has been widely used in image classification which aims to learn some class-specific feature vectors as well as a representative dictionary according to a set of labeled training samples. However, interclass similarities and intraclass variances among input samples and learned features will generally weaken the representability of dictionary and the discrimination of feature vectors so as to degrade the classification performance. Therefore, how to explicitly represent them becomes an important issue. In this paper, we present a novel DDL framework with two-level low rank and group sparse decomposition model. In the first level, we learn a class-shared and several class-specific dictionaries, where a low rank and a group sparse regularization are, respectively, imposed on the corresponding feature matrices. In the second level, the class-specific feature matrix will be further decomposed into a low rank and a sparse matrix so that intraclass variances can be separated to concentrate the corresponding feature vectors. Extensive experimental results demonstrate the effectiveness of our model. Compared with the other state-of-the-arts on several popular image databases, our model can achieve a competitive or better performance in terms of the classification accuracy. Zaidao Wen, Zaidao Hou, Licheng Jiao |
IEEE Trans. Cybern. | 1 |
| 2017 | Robust Semisupervised Classification for PolSAR Image With Noisy LabelsabstractThe robustness of the supervised polarimetric synthetic aperture radar (PolSAR) image classification is severely affected by two main aspects, namely, the quantity and quality of the labeled training pixels. Specifically, limited manually labeled pixels with respect to the large scale of PolSAR image have limited the performance of the automatic classification methods, while manually labeled training pixels shall be unfaithful with the speckle and impure cell for their low qualities. In order to address the above two fundamental problems, we propose a robust semisupervised probability graphic-based classification framework. First, a semisupervised learning scheme is implemented to simultaneously exploit both labeled and unlabeled pixels for information compensation. Moreover, structural relationship among neighboring pixels inducing from the prior information is further benefit to reduce the influence of limited labeled pixels. Second, a robust classification loss function is added in the process of training classifier to enhance the robustness to the noisy labeled pixels. Third, unfaithful limited labeled data can be settled with a hybrid generative/discriminative classification framework, where labeled and unlabeled pixels are simultaneously exploited for learning high-level feature for the low-quality pixels. The effectiveness of the proposed framework on the specific aspect is validated in experiments on real PolSAR data sets, which reveal the superiority in both visual performance and classification accuracy compared with the state-of-the-art methods. Totally speaking, our model has improved the classification accuracy by at least 20% on data set Flevoland, 10% on Oberpfaffenhofen, and 5% on Weihe River than the compared ones. Biao Hou, Zaidao Wen, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Discriminative Nonlinear Analysis Operator Learning: When Cosparse Model Meets Image ClassificationabstractA linear synthesis model-based dictionary learning framework has achieved remarkable performances in image classification in the last decade. Behaved as a generative feature model, it, however, suffers from some intrinsic deficiencies. In this paper, we propose a novel parametric nonlinear analysis cosparse model (NACM) with which a unique feature vector will be much more efficiently extracted. Additionally, we derive a deep insight to demonstrate that NACM is capable of simultaneously learning the task-adapted feature transformation and regularization to encode our preferences, domain prior knowledge, and task-oriented supervised information into the features. The proposed NACM is devoted to the classification task as a discriminative feature model and yield a novel discriminative nonlinear analysis operator learning framework (DNAOL). The theoretical analysis and experimental performances clearly demonstrate that DNAOL will not only achieve the better or at least competitive classification accuracies than the state-of-the-art algorithms, but it can also dramatically reduce the time complexities in both training and testing phases. Zaidao Wen, Biao Hou, Licheng Jiao |
IEEE Trans. Image Process. | 1 |
| 2016 | Unsupervised PolSAR image classification using boundary-preserving region division and region-based affinity propagation clusteringabstractThis paper presents a new method for polarimetric synthetic aperture radar (PolSAR) image classification. Firstly, to get a reasonable edge strength map, polarimetric information is used in edge strength calculation, and watershed algorithm is used to obtain the oversegmentation using the edge strength. Secondly, a searching table is used to determine the most suitable region to be merged. Finally, region-based affinity propagation clustering is employed to achieve an initial classification map, and the method provides an adjacent Wishart classifier with spatial relations to obtain the final classification result. Biao Hou, Yuheng Jiang, Bo Ren 0001, Zaidao Wen, Shuang Wang 0001, Licheng Jiao |
IGARSS | 4 |
| 2016 | Learning task-driven polarimetric target decomposition: A new perspectiveabstractPolarimetric target decomposition aims to decompose a polarimetric synthetic aperture (PolSAR) data on a base reflecting some scattering mechanisms. The corresponding coefficients will be further exploited as the feature vector for the subsequent interpretation task. Intuitively, its performance heavily depends on the choice of bases and many off-the-shelf ones have been constructed based on mathematical or physical model since last two decades. However, these fixed bases are generally insufficient to characterize all types of data in a PolSAR image so that the extracted features are not beneficial to the subsequent task. To address this issue, we propose a novel target decomposition framework to learn a set of task-desired bases as well as feature vectors from the input polarimetric data. Focusing on the classification task, involve a supervised regularizer is further involved in our framework to increase the discrimination of features. Experimental results demonstrate the effectiveness of proposed framework. Zaidao Wen, Biao Hou, Shuang Wang 0001, Licheng Jiao |
IGARSS | 1 |
| 2013 | High resolution SAR target reconstruction from compressive measurements with prior knowledgeabstractIn this paper, an effective prior knowledge based framework for target reconstruction from compressive measurements is proposed. In this framework, a traditional compressed imaging method is firstly introduced which indicates that for a range cell containing K strongest scattering points can be reconstructed based on the theory of compressive sensing. Secondly, a greedy iteration algorithm is modified which utilizes some prior knowledge of the target during the reconstruction step. The experiments are carried on the Moving and Stationary Target Acquisition and Recognition (MSTAR) database and the results show the effectiveness of our framework for target reconstruction. Zaidao Wen, Biao Hou, Shuang Wang 0001 |
IGARSS | 1 |