VLDB 2026 Research / reviewers in the wild / expert
Wei Wei 0008
dblp:24/4105-8
· DBLP profile ↗
116ranked-venue papers
14as first author
66since 2021 · last 2026
0000-0002-0655-056XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 4 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 4 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 45 · 7 first-author · 30 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust-R1: Degradation-Aware Reasoning for Robust Visual UnderstandingabstractMultimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that focuses solely on visual encoder generalization, suffering from limited interpretability and isolated optimization. To overcome these limitations, we propose Robust-R1, a novel framework that explicitly models visual degradations through structured reasoning chains. Our approach integrates: (i) supervised fine-tuning for degradation-aware reasoning foundations, (ii) reward-driven alignment for accurately perceiving degradation parameters, and (iii) dynamic reasoning depth scaling adapted to degradation intensity. To facilitate this approach, we introduce a specialized 11K dataset featuring realistic degradations synthesized across four critical real-world visual processing stages, each annotated with structured chains connecting degradation parameters, perceptual influence, pristine semantic reasoning chain, and conclusion. Comprehensive evaluations demonstrate state-of-theart robustness: Robust-R1 outperforms all general and robust baselines on the real-world degradation benchmark R-Bench, while maintaining superior anti-degradation performance under multi-intensity adversarial degradations on MMMB, MMStar, and RealWorldQA. Jiaqi Tang 0005, Jianmin Chen, Wei Wei 0008, Xiaogang Xu 0002, Runtao Liu, Qipeng Xie, Jiafei Wu, Lei Zhang 0001, Qifeng Chen 0001 |
AAAI | 3 |
| 2026 | JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation PromotionabstractGiven the inherently costly and time-intensive nature of pixel-level annotation, the generation of synthetic datasets comprising sufficiently diverse synthetic images paired with ground-truth pixel-level annotations has garnered increasing attention recently for training high-performance semantic segmentation models. However, existing methods necessitate to either predict pseudo annotations after image generation or generate images conditioned on manual annotation masks, which incurs image-annotation semantic inconsistency or scalability problem. To migrate both problems with one stone, we present a novel dataset generative diffusion framework for semantic segmentation, termed JoDiffusion. Firstly, given a standard latent diffusion model, JoDiffusion incorporates an independent annotation variational auto-encoder (VAE) network to map annotation masks into the latent space shared by images. Then, the diffusion model is tailored to capture the joint distribution of each image and its annotation mask conditioned on a text prompt. By doing these, JoDiffusion enables simultaneously generating paired images and semantically consistent annotation masks solely conditioned on text prompts, thereby demonstrating superior scalability. Additionally, a mask optimization strategy is developed to mitigate the annotation noise produced during generation. Experiments on Pascal VOC, COCO, and ADE20K datasets show that the annotated dataset generated by JoDiffusion yields substantial performance improvements in semantic segmentation compared to existing methods. Haoyu Wang 0016, Lei Zhang 0054, Dengyang Jiang, Wei Wei 0008, Chen Ding 0002 |
AAAI | 5 |
| 2026 | AMG-Net: A multitask network with adaptive mutual guidance for Semantic Change Detection
Yuduo Bian, Wei Wei 0008, Chen Ding 0002, Lei Zhang 0038, Jiangbin Zheng 0001, Yanning Zhang 0001 |
Pattern Recognit. | 2 |
| 2026 | Push the limit of scene text recognition using character and text length guided text super-resolution
Jiangtao Nie, Boxiong Wu, Wenyu Peng, Wei Wei 0008, Lei Zhang 0054, Chen Ding 0002, Yanning Zhang 0001 |
Pattern Recognit. | 4 |
| 2026 | Category text-guided RGBT tracking with shared-specific feature representation
Wei Wei 0008, Haolie Wang, Yuduo Bian, Haijiao Xing, Chen Ding 0002, Lei Zhang 0054, Tao Zhou 0009, Jiangbin Zheng 0001, Yanning Zhang 0001 |
Pattern Recognit. | 1 |
| 2026 | Do it yourself dynamic single image super resolution network via ODE
Xiao Zhang 0058, Zhen Zhang 0008, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2026 | Hyperspectral Image Compression With Spectral-Spatial Coupling and Group-Wise Context ModelingabstractThe rich spectral information within hyperspectral images (HSIs) results in large data volumes. Thus finding a compact representation for HSIs while maintaining reconstruction quality is a fundamental task for numerous applications. Though the existing learning-based compression methods and context models have shown strong rate-distortion (RD) performance, these methods only pay their attention on spatial redundancy without considering the spectral redundancy of HSIs, which thus impedes further improvement of their performance on HSI. Moreover, the strictly sequential autoregressive nature of context models leads to inefficiency, further limiting their practical applications. In this paper, leveraging the spectral priors unique to HSIs, we propose a hybrid Transformer-CNN architecture to find compact latent representations of HSIs. In specific, we construct Spectral-Spatial Coupling Transformer Group (SSCTG) to cooperatively extract spatial and spectral features of HSIs. Additionally, we propose Group-wise Context Model (GCM) to further enhance the parallel processing capability of autoregression within context models, significantly improving the coding efficiency. Extensive experiments demonstrate the effectiveness of the proposed method, achieving superior RD performance compared to state-of-the-art methods while maintaining high efficiency of codecs. Wei Wei 0008, Shuyi Zhao, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Meta-Exploiting Complementary Semantic Consistency for Cross-Domain Few-Shot Learning PromotionabstractMeta-learning has emerged as an effective solver for cross-domain few-shot learning (CD-FSL) tasks. Despite achieving obvious progress recently, the typical episodic learning paradigm often causes the feature embedding model collapsing into the simplicity bias pitfall, viz., the model tends to prioritize some shortcut patterns (e.g., color, style, background) that are only sufficient to distinguish categories in source domain, while fail to generalize across domains. To mitigate this problem, we present a novel meta-learning framework which emphasizes meta-exploiting inductive bias to alleviate simplicity bias for CD-FSL promotion, and mainly contributes in the following four aspects. 1) We establish a novel inductive bias for CD-FSL, termed complementary semantic consistency (CSC). The rationale behind lies in that forcing the semantic consistency between two complementary feature learning schemes is beneficial to distill cross-domain transferable features. 2) We establish a solid theoretical foundation, supported by rigorous mathematical proofs and key lemmas, which demonstrates that CSC establishes a tighter generalization bound and facilitates the learning of domain-invariant features. 3) Inspired by CSC, we propose a general meta-learning framework, which implements complementary feature embedding models using parallel networks with the same architecture but different input forms, and introduce proper knowledge distillation losses to encourage the semantic consistency between different branches during meta-training. This framework can be seamlessly integrated with any complementary feature learning schemes. 4) To clarify this point, we instantiate two effective meta-learners based on the proposed framework. The former establishes a two-branch network that simultaneously classifies both the query image and its random local crops. The latter decomposes the query image into high-frequency and low-frequency components, which are then integrated into a parallel feature embedding network for category prediction, analogous to the original query image. Subsequently, a KL divergence based knowledge distillation loss is separately leveraged to force the prediction consistency between the complementary branches (e.g., local-global, spatial-frequency) during meta-training. By doing these, both learners are able to distill cross-domain transferable features with better generalization performance. Empirical results on diverse benchmarks consistently affirm the proposed framework's advantages, while additional analysis provides compelling support for our key claims. Fei Zhou 0008, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Guosheng Lin, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Low-Biased General Annotated Dataset GenerationabstractPre-training backbone networks on a general annotated dataset (e.g., ImageNet) that comprises numerous manually collected images with category annotations has proven to be indispensable for enhancing the generalization capacity of downstream visual tasks. However, those manually collected images often exhibit bias, which is non-transferable across either categories or domains, thus causing the model’s generalization capacity degeneration. To mitigate this problem, we present a low-biased general annotated dataset generation framework (lbGen). Instead of expensive manual collection, we aim at directly generating low-biased images with category annotations. To achieve this goal, we propose to leverage the advantage of a multimodal foundation model (e.g., CLIP), in terms of aligning images in a low-biased semantic space defined by language. Specifically, we develop a bi-level semantic alignment loss, which not only forces all generated images to be consistent with the semantic distribution of all categories belonging to the target dataset in an adversarial learning manner, but also requires each generated image to match the semantic description of its category name. In addition, we further cast an existing image quality scoring model into a quality assurance loss to preserve the quality of the generated image. By leveraging these two loss functions, we can obtain a low-biased image generation model by simply fine-tuning a pre-trained diffusion model using only all category names in the target dataset as input. Experimental results confirm that, compared with the manually labeled dataset or other synthetic datasets, the utilization of our generated low-biased dataset leads to stable generalization capacity enhancement of different backbone networks across various tasks, especially in tasks where the manually labeled samples are scarce. Code is available at: https://github.com/vvvvvjdy/lbGen Dengyang Jiang, Haoyu Wang 0016, Lei Zhang 0054, Wei Wei 0008, Guang Dai, Yanning Zhang 0001 |
CVPR | 4 |
| 2025 | Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive Injection
Wei Wei 0008, Jiaqi Tang 0005, Jiangtao Nie, Yanyu Ye, Xiaogang Xu 0002, Ying-Cong Chen, Lei Zhang 0001 |
ICCV | 2 |
| 2025 | Rhythmguassian: Repurposing Generalizable Gaussian Model for Remote Physiological Measurement
Hao Lu 0009, Yuting Zhang 0008, Jiaqi Tang 0005, Wenhang Ge, Wei Wei 0008, Kaishun Wu, Ying-Cong Chen |
ICCV | 6 |
| 2025 | Towards Effective Foundation Model Adaptation for Extreme Cross-Domain Few-Shot Learning
Fei Zhou 0008, Lei Zhang 0038, Wei Wei 0008, Chen Ding 0002, Guosheng Lin, Yanning Zhang 0001 |
ICCV | 4 |
| 2025 | Prompt-Free Conditional Diffusion for Multi-object Image AugmentationabstractDiffusion model has underpinned much recent advances of dataset augmentation in various computer vision tasks. However, when involving generating multi-object images as real scenarios, most existing methods either rely entirely on text condition, resulting in a deviation between the generated objects and the original data, or rely too much on the original images, resulting in a lack of diversity in the generated images, which is of limited help to downstream tasks. To mitigate both problems with one stone, we propose a prompt-free conditional diffusion framework for multi-object image augmentation. Specifically, we introduce a local-global semantic fusion strategy to extract semantics from images to replace text, and inject knowledge into the diffusion model through LoRA to alleviate the category deviation between the original model and the target dataset. In addition, we design a reward model based counting loss to assist the traditional reconstruction loss for model training. By constraining the object counts of each category instead of pixel-by-pixel constraints, bridging the quantity deviation between the generated data and the original data while improving the diversity of the generated data. Experimental results demonstrate the superiority of the proposed method over several representative state-of-the-art baselines and showcase strong downstream task gain and out-of-domain generalization capabilities. Code is available at \href{https://github.com/00why00/PFCD}{here}. Haoyu Wang 0016, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
IJCAI | 3 |
| 2025 | Generalized pixel-aware deep function-mixture network for effective spectral super-resolution
Jiangtao Nie, Lei Zhang 0054, Chongxing Song, Zhiqiang Lang, Weixin Ren, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
Knowl. Based Syst. | 6 |
| 2025 | SAR remote sensing image segmentation based on feature enhancement
Wei Wei 0008, Yanyu Ye, Guochao Chen, Yanning Zhang 0001 |
Neural Networks | 1 |
| 2025 | Multi-scale feature extraction and fusion with attention interaction for RGB-T tracking
Haijiao Xing, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
Pattern Recognit. | 2 |
| 2025 | Domain consistency learning for continual test-time adaptation in image semantic segmentation
Yanyu Ye, Wei Wei 0008, Lei Zhang 0054, Chen Ding 0002, Yanning Zhang 0001 |
Pattern Recognit. | 2 |
| 2025 | A visual prompt learning network for hyperspectral object tracking
Haijiao Xing, Wei Wei 0008, Lei Zhang 0054, Chen Ding 0002 |
Pattern Recognit. Lett. | 2 |
| 2025 | Leapfrog Polymorphic Neural Ordinary Differential EquationabstractNeural Ordinary Differential Equations (NODEs) revolutionize the way we view residual networks as solvers for initial value problems (IVPs), with layer depth serving as the time step. In this study, we propose a more efficient extension of NODEs called Leap-Frog Polymorphic Neural ODEs (LF-NODEs). LF-NODEs introduce the leap-frog updating scheme, breaking free from the specific structure of time-evolving mixtures of multiple dynamical systems. In each time step (corresponding to each residual network layer), LF-NODEs integrate the features of multiple dynamical systems (residual network layers) by employing time-variant weights from different dynamical systems. Our LF-NODEs not only overcome the limitations in the representation power of individual residual networks but also provide a more flexible and effective structure for residual networks based on multiple dynamical systems. The efficiency of the leap-frog updating scheme is theoretically derived and demonstrated. Furthermore, we expand the model by incorporating a multi-scale approach (LF-MSNODEs) to achieve even more accurate model representations. Empirical results showcase the improvements provided by our proposed methods in 2D concentric spheres task, irregularly sampled time series prediction, and classification tasks. Additionally, the results provide empirical evidence that the learned feature space enhances system efficiency with fewer parameters and function evaluations compared to the baseline methods. Xiao Zhang 0058, Wei Wei 0008, Zhen Zhang 0008, Lei Zhang 0054 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Adapt Anything: Tailor Any Image Classifier Across Domains and Categories Using Text-to-Image Diffusion ModelsabstractWe study a novel problem in this paper, that is, if a modern text-to-image diffusion model can tailor any image classifier across domains and categories. Existing domain adaption works exploit both source and target data for domain alignment so as to transfer the knowledge from the labeled source data to the unlabeled target data. However, as the development of text-to-image diffusion models, we wonder if the high-fidelity synthetic data can serve as a surrogate of the source data in real world. In this way, we do not need to collect and annotate the source data for each image classification task in a one-for-one manner. Instead, we utilize only one off-the-shelf text-to-image model to synthesize images with labels derived from text prompts, and then leverage them as a bridge to dig out the knowledge from the task-agnostic text-to-image generator to the task-oriented image classifier via domain adaptation. Such a one-for-all adaptation paradigm allows us to adapt anything in the world using only one text-to-image generator as well as any unlabeled target data. Extensive experiments validate the feasibility of this idea, which even surprisingly surpasses the state-of-the-art domain adaptation works using the source data collected and annotated in real world. Weijie Chen 0006, Haoyu Wang 0016, Shicai Yang, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Luojun Lin, Di Xie, Yueting Zhuang |
IEEE Trans. Big Data | 5 |
| 2025 | Diffusion-Augmented Cross-Domain Prototypical Knowledge Distillation for Few-Shot Learning in Hyperspectral Image ClassificationabstractCross-domain few-shot learning (FSL) has demonstrated remarkable new classes recognition capabilities in hyperspectral image classification tasks. However, existing domain adaptation methods face two critical challenges in the cross-domain feature alignment process: first, the domain shift leads to misaligned feature transfer and diminished classification accuracy; second, the intra-class feature dispersion and inter-class boundary blurring in few-shot tasks result in degraded classification performance for novel classes. Moreover, the impact of redundant and noisy data on model discriminability is rarely considered in existing approaches. To solve these issues, this article proposes a cross-domain FSL hyperspectral image classification method based on diffusion-augmented prototype knowledge distillation (DAPKD-CFSL). Firstly, we introduce a diffusion-augmented unsupervised domain adaptation pre-training (DA-PT) framework to address the domain shift by performing a domain-adversarial denoising and reconstruction task using visible source data and masked target data. Second, our dual-branch spatial-spectral attention (DB-SSA) captures global and local spectral-spatial dependencies to enhance feature representation. Then, the proposed global-local prototype knowledge distillation (GL-PKD) performs global prototype alignment while conducting local contrastive learning, addressing feature dispersion and boundary ambiguity. Finally, a dynamic learning strategy prioritizes feature alignment early and gradually strengthens classification supervision through adaptive loss weights, and incorporates an SNR-enhanced loss to effectively mitigate noise interference. The experimental results on three HSI datasets demonstrate the superiority and effectiveness of the proposed DAPKD-CFSL. Chen Ding 0002, Sirui Zheng, Mengmeng Zheng, Yizhou Dong, Wenqiang Hua, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Constructing a Multi-Modal Based Underwater Acoustic Target Recognition Method With a Pre-Trained Language-Audio ModelabstractUnderwater acoustic target recognition (UATR) aims to accurately identify radiated acoustic signals from ships in complex maritime environments. The challenges of this task lay in how to explore discriminative representation from complex and limited acoustic samples. Recently, various deep learning-based UATR methods have been proposed. However, their performance on real sonar-collected signals remains restricted. On one hand, most methods currently adopt different representation extraction strategies to extract features from acoustic signals such as time-frequency (T-F) representation, wave representation, and joint representation. However, the limited feature representation capability and simple feature fusion strategies often limit the recognition performance improvement. On the other hand, they often overlook the knowledge gains brought by pre-trained models and the extraction of multifeature semantic correlation knowledge. This leads to unsatisfactory performance and even overfitting issues. To mitigate these issues, this article proposes a multifeature UATR (MF-UATR) method. It introduces a strongly generalized multi-modal pre-trained language-audio model and contrastive learning-based feature-level fusion strategy to semantically guide and fuse multiple features. This strategy facilitates the model in learning prior knowledge and the semantic correlations between features thereby improving recognition performance. In addition, we also considered the few-shot scenarios with extremely limited data, in which a multi-modal few-shot UATR (MMFS-UATR) scheme is proposed. It efficiently completes the few-shot UATR (FS-UATR) task by combining parameter-efficient fine-tuning (PEFT) techniques, semantic supervision strategy, and pre-trained MF-UATR. Extensive experiments on two public datasets, DeepShip and ShipsEar, demonstrate that the proposed frameworks achieve optimal target recognition performance under regular and few-shot settings. Jiangtao Nie, Wei Wei 0008, Lei Zhang 0054 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | IrregFusion: A Generalized Framework for Hyperspectral Image Fusion Across Diverse Spectral DataabstractFusing a low-resolution (LR) hyperspectral image (HSI) with a high-resolution (HR) multispectral image (MSI) has emerged as a promising strategy for reconstructing high-quality HSIs that combine rich spectral and fine spatial information. However, most existing HSI fusion methods operate under the restrictive assumption that the LR HSI and HR MSI are spatially aligned and fully consistent on the field-of-view (FoV), which significantly limits their applicability in real-world scenarios when such alignment is unavailable. To overcome these limitations, we propose IrregFusion, a generalized HSI fusion framework capable of handling both FoV-consistent and inconsistent fusion scenarios. Specifically, IrregFusion incorporates a Transformer-based reconstruction module that captures both intra- and inter-modal correlations between the diverse spectra data and the MSI, enhancing the model’s ability to perceive and reconstruct non-local spectral–spatial structures. To further address the challenges posed by FoV inconsistencies, we introduce a spectral propagation strategy that diffuses observed spectral information into adjacent spectral-blank regions, thereby easing the reconstruction of missing spectral content. Additionally, a self-supervised adaptation mechanism is integrated into the framework, enabling robust spectral–spatial representation learning and enhancing generalization across diverse and challenging conditions. Extensive experiments conducted on benchmark datasets demonstrate that IrregFusion effectively addresses the challenges of diverse spectral data fusion and consistently outperforms state-of-the-art methods in both reconstruction accuracy and visual fidelity. The source code will be released in https://github.com/JiangtaoNie/IrregFusion.git. Jiangtao Nie, Wei Wei 0008, Lei Zhang 0054, Chen Ding 0002, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | MCINet: Fusing Low-Light Visible-Infrared Image via Max-Merge Complementary InformationabstractFusing complementary information in the visible-infrared image offers a promising approach to enhance the performance of downstream computer vision tasks (e.g., object detection, segmentation etc) in complicated imaging conditions (e.g., low-illumination). However, due to the robust imaging capacity of the infrared sensor in complicated imaging conditions, most existing methods primarily rely on the salient object intensity information in the infrared modality for fusion, while the visible information (e.g., color, texture etc) is not adequately utilized, and thus limit their generalization capacity in downstream computer vision tasks. In this study, we present a novel image fusion framework, i.e.,MCInet, which attempts toMaximize and merge theComplementaryInformation across visible-infrared modalities for more informative image fusion. To this end, we first introduce the modality-specific processing module into the fusion framework to improve the information representation of each modality image. For visible images, a pre-trained low-light enhance module is adopted to enhance its color and texture information. In addition, for infrared images, a nonlinear mapping module is constructed to suppress the excessive salient object intensity information of infrared modality. Then we establish a reusable MCI block that embeds a cross-image mutual information minimization scheme into an input-aware fusion module. This empowers us to dynamically maximize and merge the complementary information between two input images according to their feature representation. In addition, we introduce a cycle reconstruction loss to self-supervised regularize the fusion results for further enhancement. Experiments on image fusion, object detection, and segmentation tasks demonstrate that the proposed framework can produce more informative fusion results and exhibit better performance in downstream computer vision tasks. Jiangtao Nie, Boxiong Wu, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | CR-SSRNet: Cross-Sensor Robust Spectral Super-Resolution Network Guided by Cognition FeaturesabstractSpectral Super-Resolution (SSR) aims at reconstructing a latent hyperspectral images (HSI) from a RGB image. Recent progress mainly focused on building a deep spectral super-resolution networkto directly map the input RGB image to the corresponding HSI. Their pleasing performance depends on the assumption that the spectral response function determined by the RGB sensor is consistent across training and test data. However, in practice, the training and test data are inevitably captured by different RGB sensors, thus resulting in obvious performance drop when using these networks. To mitigate this problem, we present a novel cognitive feature guided cross-sensor robust spectral super-resolution network. In a specific, a U-shape multi-scale network is first established to learn the deep mapping between input RGB image and the latent HSI. Then, a large-scale foundation cognitive model is introduced to extract multi-level cross-sensor invariant cognitive features from the input RGB. Moreover, these features are separately adapted and injected into different decoder blocks in the U-shape spectral super-resolution network. By doing these, the proposed network learns to appropriately guide the coarse-to-fine spectral reconstruction process using multilevel cognitive features, and thus shows better generalization performance in the cross-sensor SSR tasks. Experiments on two benchmark datasets demonstrate the superiority of the proposed method over several state-of-the-art baselines. Weixin Ren, Ruiling Liu, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
IGARSS | 4 |
| 2024 | Accurate SAR Aircraft Detection Algorithm Based on Feature EnhancementabstractAircraft detection in synthetic aperture radar (SAR) images is of much significance because of its all-weather, all-day, and strong penetrating characteristics. However, existing algorithms exhibit inadequate capacity for feature extraction due to imaging discontinuities and background interference of SAR images. To overcome these task-specific issues, we proposed a feature enhancement-based SAR aircraft detection algorithm. In detail, we employed Adaptive Contrast Enhancement (ACE) in the preprocessing stage to reduce noises, and then we embed Scatter Point Focused Module (SPFM) into network to enhance the feature extraction of aircraft scattering points. Furthermore, we devised Background Interference Suppression Module (BISM) to accentuate salient points and suppress non-essential pixels. Experimental results on the GaoFen-3 SAR aircraft dataset demonstrate the effectiveness of the proposed feature enhancement-based method. Yizun Wang, Lei Zhang 0054, Chen Ding 0002, Chunna Tian, Wei Wei 0008 |
IGARSS | 7 |
| 2024 | Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot LearningabstractMeta-learning offers a promising avenue for few-shot learning (FSL), enabling models to glean a generalizable feature embedding through episodic training on synthetic FSL tasks in a source domain. Yet, in practical scenarios where the target task diverges from that in the source domain, meta-learning based method is susceptible to over-fitting. To overcome this, we introduce a novel framework, Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot Learning, which is crafted to comprehensively exploit the cross-domain transferable image prior that each image can be decomposed into complementary low-frequency content details and high-frequency robust structural characteristics. Motivated by this insight, we propose to decompose each query image into its high-frequency and low-frequency components, and parallel incorporate them into the feature embedding network to enhance the final category prediction. More importantly, we introduce a feature reconstruction prior and a prediction consistency prior to separately encourage the consistency of the intermediate feature as well as the final category prediction between the original query image and its decomposed frequency components. This allows for collectively guiding the network's meta-learning process with the aim of learning generalizable image feature embeddings, while not introducing any extra computational cost in the inference phase. Our framework establishes new state-of-the-art results on multiple cross-domain few-shot learning benchmarks. Fei Zhou 0008, Peng Wang 0023, Lei Zhang 0054, Zhenghua Chen, Wei Wei 0008, Chen Ding 0002, Guosheng Lin, Yanning Zhang 0001 |
NeurIPS | 5 |
| 2024 | FDE-Net: A memory-efficiency densely connected network inspired from fractional-order differential equations for single image super-resolution
Xiao Zhang 0058, Lei Zhang 0054, Wei Wei 0008, Chunna Tian, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2024 | Cross-Domain Distribution Calibration of Hyperspectral Image ClassificationabstractDue to the huge number of trainable parameters, deep learning based hyperspectral image(HSI) classification method frequently struggle to achieve satisfactory accuracy when providing small amount of labeled training samples. This study proposes a novel few-shot transfer learning based HSI classification method, which can exploit samples from multiple other HSI datasets(termed as multi-source domain) to address the issues of limited labeled samples in target domain. For this purpose, we first construct a feature extractor utilizing both convolution neural network (CNN) and transformer. Specifically, CNN extracts features of HSI in spatial domain, while transformer is used to capture both global and local features within spectral domain. Since the constructed feature extractor is trained on multiple HSIs from source domain, it has a good generalization ability. Then, we propose to utilize the distribution calibration to decrease the difference between the features of the source domain and the target domain. By selecting samples with similar distribution with the target domain from the multi-source domain for distribution calibration, the generalization ability of the proposed method for the target domain classification HSI is further enhanced. Experimental results demonstrate the proposed method has better HSI classification results compared with other competing methods. Junyuan Ding, Wei Wei 0008, Lei Zhang 0054 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Unsupervised Test-Time Adaptation Learning for Effective Hyperspectral Image Super-Resolution With Unknown DegenerationabstractFusing a low-resolution hyperspectral image (HSI) with a high-resolution (HR) multi-spectral image has provided an effective way for HSI super-resolution (SR). The key lies on inferring the posteriori of the latent (i.e., HR) HSI using an appropriate image prior and the likelihood determined by the degeneration between the latent HSI and the observed images. However, in scenarios with complex imaging environments and various imaging scenes, the prior of HSIs can be prohibitively complicated and the degeneration is often unknown, which causes it difficult to accurately infer the posteriori of each latent HSI. To tackle this problem, we present an unsupervised test-time adaptation learning (UTAL) framework for HSI SR under unknown degeneration. Instead of directly modeling the complicated image prior, it first implicitly learns a content-agnostic prior shared across different images through supervisedly pre-training a mutual-guiding fusion module on extensive synthetic data. Then, it adapts the shared prior to those private characteristics in the latent HSI for posteriori inference through unsupervisedly learning a self-guiding adaptation module and a degeneration estimation network on two observed images in the test phase. Such a two-stage learning scheme models the complicated image prior in a divide-and-conquer manner, which eases the modeling difficulty and improves the prior accuracy. Moreover, the unknown degeneration can be estimated properly. Both of these two advantages empower us to accurately infer the posteriori of the latent HSI, thereby increasing the generalization performance in real applications. Additionally, in order to further mitigate the over-fitting in coping with more challenging cases (e.g., degenerations in both spectral and spatial domains are unknown) and speed up, we propose to meta-train UTAL on extensive synthetic SR tasks and solve it using an alternative optimization strategy such that UTAL learns to produce good generalization performance in real challenging cases with a small number of gradient descent steps. To verify the efficacy of UTAL, we evaluate it on HSI SR tasks with different unknown degenerations as well as some other HSI restoration tasks (e.g., compressive sensing), and report strong results superior to that of existing competitors. Lei Zhang 0054, Jiangtao Nie, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Meta-collaborative comparison for effective cross-domain few-shot learning
Fei Zhou 0008, Peng Wang 0023, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
Pattern Recognit. | 4 |
| 2024 | Adjustable Visible and Infrared Image FusionabstractThe visible and infrared image fusion (VIF) method aims to utilize the complementary information between these two modalities to synthesize a new image containing richer information. Although it has been extensively studied, the synthesized image that has the best visual results is difficult to reach consensus since users have different opinions. To address this problem, we propose an adjustable VIF framework termed AdjFusion, which introduces a global controlling coefficient into VIF to enforce it can interact with users. Within AdjFusion, a semantic-aware modulation module is proposed to transform the global controlling coefficient into a semantic-aware controlling coefficient, which provides pixel-wise guidance for AdjFusion considering both interactivity and semantic information within visible and infrared images. In addition, the introduced global controlling coefficient not only can be utilized as an external interface for interaction with users but also can be easily customized by the downstream tasks (e.g., VIF-based detection and segmentation), which can help to select the best fusion result for the downstream tasks. Taking advantage of this, we further propose a lightweight adaptation module for AdjFusion to learn the global controlling coefficient to be suitable for the downstream tasks better. Experimental results demonstrate the proposed AdjFusion can 1) provide ways to dynamically synthesize images to meet the diverse demands of users; and 2) outperform the previous state-of-the-art methods on both VIF-based detection and segmentation tasks, with the constructed lightweight adaptation method. Our code will be released after accepted athttps://github.com/BearTo2/AdjFusion. Boxiong Wu, Jiangtao Nie, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Arbitrary-Scale Hyperspectral Image Super-Resolution From a Fusion Perspective With Spatial PriorsabstractHigh-resolution hyperspectral image (HR HSI) plays a crucial role in remote sensing applications. The single HSI super-resolution (SR) method aims to obtain an HR HSI in the spatial domain from its low-resolution (LR) counterpart. Although it has been widely studied, the performance of the existing HSI SR method is still limited because the HSI data structure itself cannot provide sufficient spatial information for reconstruction, especially with a large SR factor. In this study, we cast single HSI SR as a task fusing LR HSI with its spectral response RGB image, from which the prevalent extra high-resolution RGB images can be introduced to provide sufficient and high-quality spatial prior information for HSI SR even with a large SR factor. Within this framework, we further propose an HSI arbitrary-scale SR method, which naturally incorporates such a spatial prior in both feature extraction and local implicit image function (LIIF). Extensive experiments on two benchmark remote sensing HSI datasets, showcasing the exceptional SR performance of our proposed method. The proposed SPG-ASSR method outperforms state-of-the-art (SOTA) approaches, demonstrating its effectiveness and practical applicability. Guochao Chen, Jiangtao Nie, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | GLGAT-CFSL: Global-Local Graph Attention Network-Based Cross-Domain Few-Shot Learning for Hyperspectral Image ClassificationabstractFew-shot learning (FSL) is an effective approach to address the issue of limited labeled data in hyperspectral image classification (HSIC). However, it overlooks the domain shift between the source domain (SD) and the target domain (TD) in cross-domain tasks. Most existing domain adaptation (DA) methods alleviate the domain shift problem to some extent, but DA methods based on traditional convolutional operators overlook the nonlocal spatial relationships in HSI, while methods based on graph neural networks (GNNs), although effective in leveraging nonlocal spatial information for domain alignment, overly emphasize global relationships, which is disadvantageous for pixel-level classification in HSI. To solve these issues, this article proposes a novel globalp-local graph attention network-based cross-domain FSL (GLGAT-CFSL), which comprehensively reduces domain shift through global-to-local domain alignment. It has the following advantages: 1) an innovative dynamic triplet graph attention network is devised to identify nonlocal spatial relationships in HSI for global graph alignment (GGA) while also addressing common overfitting and oversmoothing issues in GNNs; 2) an ingenious local similarity learning (LSL) strategy is designed after global domain alignment, utilizing intradomain connectivity structures and interdomain node similarities for local DA, promoting cross-domain information propagation and more comprehensive reduction of domain shift; and 3) we propose a novel triaxial dynamic convolutional neural network (TDCNN) as the feature extractor, promoting cross-dimensional interaction between spectral and spatial dimensions, establishing a more generalizable and rich feature representation between the SD and the TD. The experimental results on three HSI datasets demonstrate the superiority and effectiveness of the proposed GLGAT-CFSL. Chen Ding 0002, Zhicong Deng, Yaoyang Xu, Mengmeng Zheng, Lei Zhang 0054, Yu Cao 0016, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Integrating Prototype Learning With Graph Convolution Network for Effective Active Hyperspectral Image ClassificationabstractIn recent years, active learning (AL) methods have provided a feasible approach to alleviate the problem of limited labeled samples in deep learning projects. Existing AL algorithms generally tend to select sample without labeled, whose category is difficult to distinguish. However, the sample in the category center is difficult to determine in AL operations, resulting in inaccurate category measuring and inaccurate sample selection. In addition, hyperspectral images (HSIs) have rich spectral reflective bands with strong correlations, which leads to the phenomenon that the spatial distribution between different categories in HSIs characterizes staggered distribution, which undoubtedly influences the HSI classification effect. In this article, we propose a new AL method (called PLGCN) which combines prototype learning (PL) and graph convolution network (GCN) to solve few-shot HSI classification tasks, and this method can add into existing deep learning-based HSI classification models. It includes two advantages: 1) the prototype of each category is iteratively updated to ensure the optimality of prototype in each sampling stage and 2) the spatial distribution of unlabeled samples is extracted via graph convolution neural network in order to obtain the better features in new space for easier discriminating. Experimental results on three commonly used benchmark HSI datasets demonstrate the effectiveness of the PLGCN in HSI classification tasks with limited labeled samples. Chen Ding 0002, Mengmeng Zheng, Sirui Zheng, Yaoyang Xu, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Dual ODE: Spatial-Spectral Neural Ordinary Differential Equations for Hyperspectral Image Super-ResolutionabstractSignificant advancements have been made in hyperspectral image (HSI) super-resolution with the development of deep-learning techniques. However, the current application of deep neural network architectures to HSI super-resolution heavily relies on empirical design strategies, which can potentially impede the improvement of image reconstruction performance and introduce distortions in the results. To address this, we propose an innovative HSI super-resolution network called dual ordinary differential equations (Dual ODEs). Drawing inspiration from ordinary differential equations (ODEs), our approach offers reliable guidelines for the design of HSI super-resolution networks. The Dual ODE model leverages a spatial ODE block to extract spatial information and a spectral ODE block to capture internal spectral features. This is accomplished by redefining the conventional residual module using the multiple ODE functions method. To evaluate the performance of our model, we conducted extensive experiments on four benchmark HSI datasets. The results conclusively demonstrate the superiority of our Dual ODE approach over state-of-the-art models. Moreover, our approach incorporates a small number of parameters while maintaining an interpretable model design, thereby reducing model complexity. Xiao Zhang 0058, Chongxing Song, Tao You, Qicheng Bai, Wei Wei 0008, Lei Zhang 0054 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Mind the Gap: Polishing Pseudo Labels for Accurate Semi-supervised Object DetectionabstractExploiting pseudo labels (e.g., categories and bounding boxes) of unannotated objects produced by a teacher detector have underpinned much of recent progress in semi-supervised object detection (SSOD). However, due to the limited generalization capacity of the teacher detector caused by the scarce annotations, the produced pseudo labels often deviate from ground truth, especially those with relatively low classification confidences, thus limiting the generalization performance of SSOD. To mitigate this problem, we propose a dual pseudo-label polishing framework for SSOD. Instead of directly exploiting the pseudo labels produced by the teacher detector, we take the first attempt at reducing their deviation from ground truth using dual polishing learning, where two differently structured polishing networks are elaborately developed and trained using synthesized paired pseudo labels and the corresponding ground truth for categories and bounding boxes on the given annotated objects, respectively. By doing this, both polishing networks can infer more accurate pseudo labels for unannotated objects through sufficiently exploiting their context knowledge based on the initially produced pseudo labels, and thus improve the generalization performance of SSOD. Moreover, such a scheme can be seamlessly plugged into the existing SSOD framework for joint end-to-end learning. In addition, we propose to disentangle the polished pseudo categories and bounding boxes of unannotated objects for separate category classification and bounding box regression in SSOD, which enables introducing more unannotated objects during model training and thus further improves the performance. Experiments on both PASCAL VOC and MS-COCO benchmarks demonstrate the superiority of the proposed method over existing state-of-the-art baselines. The code can be found at https://github.com/snowdusky/DualPolishLearning. Lei Zhang 0054, Wei Wei 0008 |
AAAI | 3 |
| 2023 | Glocal Energy-based Learning for Few-Shot Open-Set RecognitionabstractFew-shot open-set recognition (FSOR) is a challenging task of great practical value. It aims to categorize a sample to one of the predefined, closed-set classes illustrated by few examples while being able to reject the sample from unknown classes. In this work, we approach the FSOR task by proposing a novel energy-based hybrid model. The model is composed of two branches, where a classification branch learns a metric to classify a sample to one of closed-set classes and the energy branch explicitly estimates the open-set probability. To achieve holistic detection of open-set samples, our model leverages both class-wise and pixel-wise features to learn a glocal energy-based score, in which a global energy score is learned using the class-wise features, while a local energy score is learned using the pixel-wise features. The model is enforced to assign large energy scores to samples that are deviated from the few-shot examples in either the class-wise features or the pixel-wise features, and to assign small energy scores otherwise. Experiments on three standard FSOR datasets show the superior performance of our model.11Code is available at https://github.com/00why00/Glocal Haoyu Wang 0016, Guansong Pang, Peng Wang 0023, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
CVPR | 5 |
| 2023 | Revisiting Prototypical Network for Cross Domain Few-Shot LearningabstractPrototypical Network is a popular few-shot solver that aims at establishing a feature metric generalizable to novel few-shot classification (FSC) tasks using deep neural networks. However, its performance drops dramatically when generalizing to the FSC tasks in new domains. In this study, we revisit this problem and argue that the devil lies in the simplicity bias pitfall in neural networks. In specific, the network tends to focus on some biased shortcut features (e.g., color, shape, etc.) that are exclusively sufficient to distinguish very few classes in the meta-training tasks within a pre-defined domain, but fail to generalize across domains as some desirable semantic features. To mitigate this problem, we propose a Local-global Distillation Prototypical Network (LDP-net). Different from the standard Prototypical Network, we establish a two-branch network to classify the query image and its random local crops, respectively. Then, knowledge distillation is conducted among these two branches to enforce their class affiliation consistency. The rationale behind is that since such global-local semantic relationship is expected to hold regardless of data domains, the local-global distillation is beneficial to exploit some cross-domain transferable semantic features for feature metric establishment. Moreover, such local-global semantic consistency is further enforced among different images of the same class to reduce the intra-class semantic variation of the resultant feature. In addition, we propose to update the local branch as Exponential Moving Average (EMA) over training episodes, which makes it possible to better distill cross-episode knowledge and further enhance the generalization performance. Experiments on eight cross-domain FSC benchmarks empirically clarify our argument and show the state-of-the-art results of LDP-net. Code is available in https://github.com/NWPUZhoufei/LDP-Net Fei Zhou 0008, Peng Wang 0023, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
CVPR | 4 |
| 2023 | U-Shape Spectral-Transformer for Robust Fusion Based Hyperspectral Super-ResolutionabstractFusing a high-spatial-resolution (HR) multi-spectral image (MSI) with a low-spatial resolution (LR) hyperspectral image (HSI) provides an effective way for HSI super-resolution (SR). Although recent deep neural network-based methods have shown pleasing fusion performance, most of them assume both input images for fusion to be clean without any noise corruption. When random noise exists in real applications, their performance drops greatly. To mitigate this problem, we present a U-shape spectral transformer for robust fusion-based HSI SR, which mainly contributes in the following three aspects. 1) A two-stage network is established to end-to-end denoise both input images and fuse them for SR. 2) A U-shape spectral transformer is constructed to simultaneously exploit the multi-scale spatial information and the long-range correlation in spectral domain, which enables sufficiently fusing the supplementary spatial-spectral information in both input images for accurate HSI SR. 3) A mutual information maximization based loss is composed with the conventional reconstruction loss to more accurately supervise the training process, thus further enhance the performance. Experimental results on two datasets demonstrate the efficacy of the proposed method in terms of HSI SR under different levels of noise corruption. Guochao Chen, Boxiong Wu, Haijiao Xing, Wei Wei 0008, Lei Zhang 0054 |
IGARSS | 5 |
| 2023 | Semi-Supervised Classification of Hyperspectral Images based on Contrastive Learning ConstraintabstractDespite significant advancements in deep learning-based algorithms for classifying hyperspectral image (HSI), this task remains challenging when only few labeled training examples are available. In this paper, we introduce a contrastive learning constraint and propose a semi-supervised HSI classification approach. We first build a multi-scale feature extraction module, which extracts fine-grained features from a small number of labeled samples together with a huge amount of unlabeled samples. Then, by modeling contrastive constraints on the unlabeled data, we construct a contrastive sub-network module, which can efficiently support the supervised HSI classification sub-network trained on the labeled dataset and hence enhance the generalization ability. Experimental results on two datasets demonstrate the effectiveness of the proposed semi-supervised HSI classification methods. Junyuan Ding, Yue Wen, Weixin Ren, Lei Zhang 0054, Wei Wei 0008 |
IGARSS | 5 |
| 2023 | Correlated NMS: Establishing Correlations Between Dense Predictions of Remote Sensing ImagesabstractObject detection is an important task for remote sensing image analysis, which aims to identify and locate objects within captured remote sensing images. Several object detection methods have been proposed, among which Non-maximum suppression (NMS) is an essential ingredient of these methods. Although simple and useful for object detection, the performance of detection methods with NMS will degeneradte when the objects within remote sensing images are dense. One of the main reasons is the presence of severe occlusion in some remote sensing images, which can easily mislead NMS to suppress the nearby candidate box of different object from its central box. To address this problem effectively, we propose to measure the correlation between the candidate box and the central box, from which if these two boxes come from the same object can be estimated. Building on this idea, we propose Correlated NMS, which can adaptively adjust the suppression threshold between the candidate box and its central box based on whether they tend to represent the same object. Experimental results demonstrate the effectiveness of the proposed method. Wei Li 0219, Guochao Chen, Lei Zhang 0054, Wei Wei 0008 |
IGARSS | 6 |
| 2023 | Efficient and Accurate Giraffe-Det for UAV Image Based Object DetectionabstractObject detection based on unmanned aerial vehicle (UAV) images has become an important area of research within remote sensing community. However, detecting objects on UAV image datasets, such as Visdrone [1] and UAVDT [2], encounters greater challenges compared with detecting objects on ordinary image datasets like COCO. It can be attributed to the fact that UAV image datasets frequently include a significant quantity of small objects, which are more difficult to detect due to the limited information available. In this study, we introduce a new object detection method for UAV images, termed as HRGiraffe-Det, which builds upon the small-object-friendly detection model(i.e., Giraffe-Det). To preserve more spatial information of small targets, we utilize upsampled image instead of the original image as input. Additionally, we construct a Multi-Proxy Head (MPHead) to deal with objects those have diverse appearance variations. Experimental results on UAV image dataset demonstrate the effectiveness of the proposed method for object detection. Qinglin Ran, Wei Wei 0008, Lei Zhang 0054 |
IGARSS | 3 |
| 2023 | Wavelet Transform Based Network for Spectral Super-ResolutionabstractSpectral super-resolution (SSR) aims at reconstructing a hyperspectral image (HSI) from an observed RGB image through interpolation in the spectral domain. Recent progress mainly focus on establishing various deep interpolation networks to directly exploit the spatial-spectral information of the RGB image for SSR. However, few of them pay attention on its frequency information, which proves to be orthogonal to the spatial-spectral information and also crucial for SSR, and thus their generalization performance can be further improved. To mitigate this problem, in this study we proposes a wavelet transform based network (WTNet) for SSR. Different from existing SSR networks in image-domain, the Haar wavelet transform is employed to decompose the input RGB image into four different frequency bands. Moreover, a multi-scale convolution and self-attention based feature extraction block and a cross-attention based band interaction block are constructed to separately exploit the statistics within each band as well as the inter-band frequency correlation. By doing these, the proposed WTNet is able to sufficiently exploit the frequency information of the input RGB image for accurate SSR. Experimental results on two datasets demonstrate the efficacy and superior SSR performance of the proposed WTNet. Weixin Ren, Qianyue Duan, Tiange Huang, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
IGARSS | 5 |
| 2023 | IVJDN: An End-to-End Network for Joint Infrared and Visible Image Fusion and DetectionabstractFusing infrared and visible images has been an active research topic within the remote sensing community since these two kinds of images can provide complementary information. Though different methods have been proposed, most of the existing infrared and visible image fusion methods only focus on obtaining visually pleasing results without considering if the fusion results fit well for the subsequential object detection task. To obtain better detection results from infrared and visible image fusion, we propose an end-to-end network that incorporates an image fusion module and object detection module into a unified framework. Within the network architecture constructed, attention mechanisms as well as intensity loss and gradient loss are utilized to effectively preserve the distinguishing characteristics of both infrared and visible modalities for object detection, yielding advantageous attributes for detection purposes. By jointly training the image fusion module and the object detection module, our proposed method achieves improved object detection performance. Experimental results corroborate the effectiveness of the proposed approach. Qinglin Ran, Wei Wei 0008, Chen Ding 0002, Lei Zhang 0054 |
IGARSS | 3 |
| 2023 | Meta-hallucinating prototype for few-shot learning promotion
Lei Zhang 0054, Fei Zhou 0008, Wei Wei 0008, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2023 | Milstein-driven neural stochastic differential equation model with uncertainty estimates
Xiao Zhang 0058, Wei Wei 0008, Zhen Zhang 0008, Lei Zhang 0054, Wei Li 0219 |
Pattern Recognit. Lett. | 2 |
| 2023 | Learning to Class-Adaptively Manipulate Embeddings for Few-Shot LearningabstractIn few-shot learning (FSL), meta-learning approach (MLA) mainly focuses on learning transferable knowledge from plenty of auxiliary FSL tasks to facilitate fast generalization to a new task. For a given FSL task, due to the inter-class distribution discrepancy, each class necessitates a specific embedding (i.e., a mapping function) to map samples into an ideal semantic space where samples from this class can be well separately from other classes. Moreover, these embeddings may vary with different tasks. Hence, one crucial knowledge for MLA is how to separately construct optimal embeddings for each class based on a few training samples given in a FSL task. However, most existing MLAs rarely consider this and thus show limited generalization capacity. To mitigate this problem, instead of directly construct class-adaptive embeddings, we present a new MLA that aims at learning to class-adaptively manipulate the features of samples for accurate classification in a new FSL task. In a specific, for a new FSL task, the proposed MLA first learns to generate some class-specific weights based on training samples via exploiting the inter-class distribution discrepancy between this class and the others. Then, the generated weights are utilized to compute the Hadamard product of features produced by a task-agnostic embedding module. By doing this, the proposed MLA can dynamically enhance or depress some specific semantic dimensions of sample features depending on the distribution of each class for accurate classification, and thus equals to constructing class-adaptive embeddings for each class but in a simpler way which can appropriately avoid over-fitting and is scalable to cases with extensive classes. To show its efficacy, we test the proposed MLA on four benchmark FSL datasets under various settings and report superior performance over existing state-of-the-arts. Fei Zhou 0008, Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Semi-Supervised Neural Architecture Search for Hyperspectral Imagery Classification Method With Dynamic Feature ClusteringabstractHyperspectral image(HSI) contains rich spatial and spectral information, which makes HSI classification task the research focus of HSI analysis within remote sensing community. Though deep learning based HSI classification methods obtain good performance in recent years, how to learn network structure better suitable for a given HSI instead of utilizing a manually designed one for HSI classification is still a challenging problem, especially providing only small amount of labeled samples. To address this problem, we propose the first semi-supervised HSI classification network constructed via the neural architecture search. Specifically, we propose a two-head semi-supervised HSI classification framework utilizing both labeled and unlabeled data, which consists of a shared feature extraction module, a classifier module for labeled samples together with a clustering module for unlabeled samples. To boost the performance of the constructed two-head network, we propose to utilize deep features instead of the original pixels for HSI clustering to generate pseudo labels for the unlabeled data. Within the conducted semi-supervised network, we specifically design a method to automatically search for the shared feature extraction module better suitable for the given HSI data, which leads to better HSI classification results. Experimental results on three HSI datasets demonstrate the effectiveness of the proposed method, providing only limited number of labeled training samples. Wei Wei 0008, Shuyi Zhao, Songzheng Xu, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Dynamic Super-Pixel Normalization for Robust Hyperspectral Image ClassificationabstractDeep neural networks (DNNs) have underpinned most of recent progress of hyperspectral image (HSI) classification. One premise of their success lies in the high image quality without noise corruption. However, due to the limitation of the imaging sensor and imaging conditions, HSIs captured in practice inevitably suffer from random noise, which will degrade the generalization performance and robustness of most existing DNN-based methods. In this study, we propose a dynamic super-pixel normalization (DSN) based DNN for HSI classification, which can adaptively relieve the negative effect caused by various types of noise corruption and improve the generalization performance. To achieve this goal, we propose a DSN module, for a given super-pixel which normalizes the inner pixel features using parameters dynamically generated based on themselves. By doing this, such a module enables adaptively restoring the similarity among pixels within the super-pixel corrupted by random noise through aligning their feature distribution, thus enhancing the generalization performance on noisy HSI. Moreover, it can be directly plugged into any other existing DNN architectures. To appropriately train the proposed DNN model, we further present a semi-supervised learning framework, which integrates the cross entropy loss and Kullback-Leibler (KL) divergence loss on labeled samples with the information entropy loss on the unlabeled samples for joint learning to well sidestep over-fitting. Experiments on three benchmark HSI classification datasets demonstrate the advantages of the proposed method over several state-of-the-art competitors in handling HSIs under different types of noise corruption. Cong Wang 0013, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Lightweighted Hyperspectral Image Classification Network by Progressive Bi-QuantizationabstractConvolutional neural network (CNN) has shown its powerful ability for hyperspectral image (HSI) classification, which however, is difficult to deploy on resource-limited or low-latency platforms due to its parameter and computation redundancy. Though binary neural network (BNN) has attracted attention for its extreme compressing and speeding up ability by binarizing both weights and activations, it has rarely been explored for HSI classification. In this study, we elaborately design a BNN with good performance for HSI classification task. Specifically, an adaptive gradient scale module is proposed to flexibly modify the gradient during training stage to better optimize the BNN and does not add any extra computation for inference. Furthermore, a curriculum learning-based progressive binarization strategy is utilized to improve the performance. Compared with the existing BNN works, our method can increase the HSI classification accuracy by a large margin while maintaining the compressing ratio. Abundant experiments on three datasets demonstrate the effectiveness of the proposed method. Wei Wei 0008, Chongxing Song, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | E2FIF: Push the Limit of Binarized Deep Imagery Super-Resolution Using End-to-End Full-Precision Information FlowabstractBinary neural network (BNN) provides a promising solution to deploy parameter-intensive deep single image super-resolution (SISR) models onto real devices with limited storage and computational resources. To achieve comparable performance with the full-precision counterpart, most existing BNNs for SISR mainly focus on compensating for the information loss incurred by binarizing weights and activations in the network through better approximations to the binarized convolution. In this study, we revisit the difference between BNNs and their full-precision counterparts and argue that the key to good generalization performance of BNNs lies on preserving a complete full-precision information flow along with an accurate gradient flow passing through each binarized convolution layer. Inspired by this, we propose to introduce a full-precision skip connection, or a variant thereof, over each binarized convolution layer across the entire network, which can increase the forward expressive capability and the accuracy of back-propagated gradient, thus enhancing the generalization performance. More importantly, such a scheme can be applied to any existing BNN backbones for SISR without introducing any additional computation cost. To validate the efficacy of the proposed approach, we evaluate it using four different backbones for SISR on four benchmark datasets and report obviously superior performance over existing BNNs and even some 4-bit competitors. Chongxing Song, Zhiqiang Lang, Wei Wei 0008, Lei Zhang 0054 |
IEEE Trans. Image Process. | 3 |
| 2022 | Non-Local Proposal Dynamic Enhancement Learning for Few-Shot Object Detection in Remote Sensing ImagesabstractDeep neural networks have underpinned much of recent progress in few-shot object detection (FSOD) in remote sensing images. The key lies in accurately inferring the object categories and bounding boxes depending on the feature of each proposal region. However, due to lack of sufficient labeled samples for training model well-fitting, the feature of each proposal fails to be discriminative and informative enough for accurate inference, thus limiting the generalization capacity. To mitigate this problem, we propose a non-local proposal dynamic enhancement learning (NPDEL) methods for FSOD in remote sensing images. In contrast to directly utilizing the proposal features extracted from the backbone, we propose to enhance them before inference using a non-local dynamic enhancement module which first carries out a non-local graph convolution on all proposal features and then dynamically fuses the convolved results with the original features for enhancement. By doing this, the enhanced proposal features can adaptively aggregate the related semantic information from the whole image, thus improving their discriminability as well as the generalization capacity in FSOD. Experiments results on different FSOD tasks demonstrate the efficacy of the proposed method. Haoyu Wang 0016, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
IGARSS | 3 |
| 2022 | Source-Free Domain Adaptation for Cross-Scene Hyperspectral Image ClassificationabstractDeep learning based cross-domain hyperspectral image (HSI) classification methods were proposed to train a classifier adapted to unlabeled target domain with the help of abundant labeled data in source domain. Although the existing methods show their potential for cross-domain HSI classification, the data in source domain may not be provided due to the data privacy, which limits the availability of these methods. In this case, how to utilize the model or knowledge trained from source domain becomes a more challenging problem. In this study, we emphasize on this problem, and propose source-free unsupervised domain adaptation method for HSI classification. Specifically, we firstly design a source domain HSI spectral feature generator, and then realize the class-wised alignment between the generated source domain HSI spectral features and the target domain features of HSI through contrastive learning. To solve the dilemma of without labels in the target domain, we also utilize a logits-weighted prototype classifier to iteratively obtain the data label of the target domain. Experiments on two cross-scene HSI datasets demonstrate the effectiveness of the proposed method when only providing the model trained from the source domain. Zun Xu, Wei Wei 0008, Lei Zhang 0054, Jiangtao Nie |
IGARSS | 2 |
| 2022 | Hyperspectral Classification with Gradient Based Active LearningabstractAlthough the hyperspectral image classification method based on convolutional neural network(CNN) has made great progress, the classification task with less training samples remains a challenging problem. Active learning technology is able to alleviate the problem above by improving the quality of labeled samples. In this paper, we propose to choose informative samples in gradient space, which considers both uncertainty and diversity in selection process. At each iteration, we performs clustering on gradient vector of network parameters corresponding to each sample and its predicted label, then the samples nearest to cluster centers are selected and labeled. Moreover, in order to alleviate the overfitting problem duing to less training samples in the early stage of active learning process, we utilize a two-branch network with shared feature extraction module to learn both supervised classification and unsupervised clustering knowledge. The proposed method is validated on widely used hyperspectral image data set, and achieves better performance. Songzheng Xu, Wei Wei 0008, Lei Zhang 0054, Xiao Zhang 0058 |
IGARSS | 2 |
| 2022 | Dynamic Long-Short Range Structure Learning for Low-Illumination Remote Sensing Imagery HDR ReconstructionabstractA promising way for low-illumination (LI) remote sensing images high-dynamic range (HDR) reconstruction is to model the mapping function from the input LI images to the corresponding high-quality counterpart using deep convolution neural networks. Due to various image contents, the key for achieving pleasing performance lies on comprehensively exploit the image-specific long-rang (e.g., non-local similarity, low-rank) and short-range (e.g., local similarity, texture etc.) structures in the LI images using appropriate network architecture. However, most existing methods can only exploit either short-range or long-range structures that are contentagnostic shared across all images, thus limiting their generalization capacity. To tackle this problem, we propose a dynamic long-short range structure learning framework for LR remote sensing images HDR reconstruction. In contrast to existing methods, we introduce a novel two-branch network architecture including a pixel-aware dynamic module that can adaptively exploit the pixel-aware short-range structure surrounding each pixel depending on its feature representation, and a long-range transformer module that dynamically exploit the long-range correlation between image patchesin the deep feature space. Then, the learned long-short range structures are integrated and cast into pixel-wise scaling factors of an illumination enhance module to restore the LI image. It empowers us to effectively exploit the image-specific long-short range structures of each input IL images for accurate HDR reconstruction. Experimental results on remote sensing images with different levels of IL demonstrate the effectiveness of the proposed method. Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
IGARSS | 3 |
| 2022 | Meta-Generating Deep Attentive Metric for Few-Shot ClassificationabstractLearning to generate a task-aware base learner proves a promising direction to deal with few-shot learning (FSL) problem. Existing methods mainly focus on generating an embedding model utilized with a fixed metric (e.g., cosine distance) for nearest neighbour classification or directly generating a linear classifier. However, due to the limited discriminative capacity of such a simple metric or classifier, these methods fail to generalize to challenging cases appropriately. To mitigate this problem, we present a novel deep metric meta-generation method that turns to an orthogonal direction, i.e., learning to adaptively generate a specific metric for a new FSL task based on the task description (e.g., a few labelled samples). In this study, we structure the metric using a three-layers deep attentive network that is flexible enough to produce a discriminative metric for each task. Moreover, different from existing methods that utilize an uni-modal weight distribution conditioned on labelled samples for network generation, the proposed meta-learner establishes a multi-modal weight distribution conditioned on cross-class sample pairs using a tailored variational autoencoder, which can separately capture the specific inter-class discrepancy statistics for each class and jointly embed the statistics for all classes into metric generation. By doing this, the generated metric can be appropriately adapted to a new FSL task with pleasing generalization performance. To demonstrate this, we test the proposed method on three benchmark FSL datasets and gain competitive results with state-of-the-art competitors. Fei Zhou 0008, Lei Zhang 0054, Wei Wei 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Contrastive Haze-Aware Learning for Dynamic Remote Sensing Image DehazingabstractImage dehazing methods aim to recover a clear image from its hazy counterpart. While various dehazing methods have been proposed, their performance on real-world remote sensing (RS) images remains unsatisfying. A key reason is that the complex weather and imaging conditions (e.g., large fields of view) cause the haze condition to dramatically change in different images, while most existing methods fail to flexibly adapt their dehazing model to the specific haze condition in each image. To mitigate this problem, we present a contrastive haze-aware learning based dynamic dehazing method which demonstrates two aspects of advantage. On one hand, a contrastive clustering scheme is utilized to learn the image-wise haze representation using a set of real-world hazy images in an unsupervised manner, which enables identifying and discriminating the specific haze condition in each given hazy image. On the other hand, with the learned haze representation, a parameter generator can produce haze-aware parameters to dynamically construct a dehazing model for the given hazy image, which empowers us to adaptively dehaze the image based on its specific haze condition and thus improves the generalization ability. In addition, a new contrastive loss defined based on the learned haze representation is further utilized for model training and leads to better performance. To demonstrate the effectiveness of the proposed method, we evaluate it on two benchmark RS image datasets including various real-world hazy images, and observe obviously superiority over other state-of-the-art competitors. Jiangtao Nie, Wei Wei 0008, Lei Zhang 0054, Jianlong Yuan, Zhibin Wang 0004, Hao Li 0030 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Toward Effective Hyperspectral Image Classification Using Dual-Level Deep Spatial Manifold RepresentationabstractHyperspectral image (HSI) contains an abundant spatial structure that can be embedded into feature extraction (FE) or classifier (CL) components for pixelwise classification enhancement. Although some existing works have exploited some simple spatial structures (e.g., local similarity) to enhance either the FE or CL component, few of them consider the latent manifold structure and how to simultaneously embed the manifold structure into both components seamlessly. Thus, their performance is still limited, especially in cases with limited or noisy training samples. To solve both problems with one stone, we present a novel dual-level deep spatial manifold representation (SMR) network for HSI classification, which consists of two kinds of blocks: an SMR-based FE block and an SMR-based CL block. In both blocks, graph convolution is utilized to adaptively model the latent manifold structure lying in each local spatial area. The difference is that the former block condenses the SMR in deep feature space to form the representation for each center pixel, while the later block leverages the SMR to propagate the label information of other pixels within the local area to the center one. To train the network well, we impose an unsupervised information loss on unlabeled samples and a supervised cross-entropy loss on the labeled samples for joint learning, which empowers the network to utilize sufficient samples for SMR learning. Extensive experiments on two benchmark HSI data set demonstrate the efficacy of the proposed method in terms of pixelwise classification, especially in the cases with limited or noisy training samples. Cong Wang 0013, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Unsupervised Recurrent Hyperspectral Imagery Super-Resolution Using Pixel-Aware RefinementabstractUnsupervised fusion-based hyperspectral imagery (HSI) super-resolution (SR) is an essential task of HSI processing, which aims to reconstruct a high-resolution (HR) HSI using only an observed low-resolution HSI and a conventional HR image. Although a large number of unsupervised HSI SR methods have been proposed, the heuristic handcrafted image priors adopted by the majority of these methods restrict their capacity to capture specific characteristics of the HSI, as well as their ability to generalize to noisy observation images. In this study, we investigate a fusion-based HSI SR framework with the deep image prior, in which the deep neural network (rather than a heuristic handcrafted image prior) is exploited to capture plenty of image statistics. Within this framework, we further propose an unsupervised recurrence-based HSI SR method using pixel-aware refinement, which utilizes the intermediate reconstruction results to self-supervise unsupervised learning. Due to containing the information of the image-specific characteristic, the proposed method achieves better performance, in terms of both accuracy and robustness to noise, compared with the existing methods. Extensive experiments on four HSI data sets demonstrate the effectiveness of the proposed method. Wei Wei 0008, Jiangtao Nie, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Boosting Hyperspectral Image Classification With Unsupervised Feature LearningabstractThe deep learning-based method has shown promising competence in image classification. Its success can be attributed to the ability to learn discriminative feature representation given plenty of labeled data. However, in real-hyperspectral image (HSI) classification applications, since pixel labeling is difficult and costly, the labels we can obtain within an HSI are always limited and noisy (i.e., inaccurate), which consequently causes overfitting of the deep learning-based method. To address this problem, we propose a novel unified deep learning network to employ both labeled and unlabeled data for training, with which the unsupervised structure knowledge, e.g., intracluster similarity and intercluster dissimilarity, inherently contained in those unlabeled data can be exploited to boost the conventional supervised classification. Specifically, we first explore the unsupervised structure knowledge in unlabeled data via a clustering method and formulate a supervised clustering task on those data with the obtained cluster labels. Then, we propose a multitask network to jointly address both the conventional classification task and the formulated supervised clustering task. With a shared feature extraction module and a high-level feature fusion module, the unsupervised structure knowledge contained in unlabeled data can be effectively introduced into the classification task, which is beneficial to learn a more discriminative feature representation and, thus, well mitigates the overfitting problem and improves the classification results. Experimental results on three data sets demonstrate the proposed method can effectively label the unlabeled data within an HSI, especially when the training labels are limited and noisy. Wei Wei 0008, Songzheng Xu, Lei Zhang 0054, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Neural Stochastic Differential Equation for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is an essental task of HSI analysis, which aims to assign each pixel a pre-defined class label. Though deep learning based methods dominate the HSI classification methods to date, the existing methods seldom consider how to directly model the uncertainty broadly exists in the HSI applications, which impedes their usage in real applications. To address this problem, we propose to directly model the uncertainty into the deep learning based HSI classification model and construct a specific network based on stochastic differential equation (SDE). The constructed network consists two subnets, in which one is utilized to well fit the HSI classification task and one is exploited to capture the uncertainty within the HSI classification. The constructed network can better depict the uncertainty, and thus result in better HSI classification performance. Experimental results demonstrate the effectiveness of the constructed model for HSI classification. Xiao Zhang 0058, Wei Wei 0008, Lei Zhang 0054, Chen Ding 0002 |
IGARSS | 2 |
| 2021 | Meta Transfer Learning for Few-Shot Hyperspectral Image ClassificationabstractWe propose a novel meta-learning approach for few-shot hyperspectral image (HSI) classification, which learns to distil transferable prior knowledge from a base dataset with sufficient labeled samples and generalize the knowledge to an unseen dataset with extremely limited labeled samples for performance improvement. Specifically, we first construct a backbone classification model using an embedding module and a linear classifier. Then, we sample extensive synthetic few-shot tasks from the base dataset, each of which consists of a support set with limited labeled samples and a query set with some unlabeled test samples. Given these tasks, we propose to optimize the embedding module using an episode learning scheme where for each task we train the linear classier based on an initialized embedding module using the support set and ultimately optimize the embedding module based on the test error on the query set until the test error on all tasks is minimized. By doing this, the resultant embedding module is able to appropriately generalize to an unseen few-shot classification task and lead to good performance with the linear classifier. Experiments on two standard classification benchmarks under different few-shot settings demonstrate the efficacy of the proposed method. Fei Zhou 0008, Lei Zhang 0054, Wei Wei 0008, Zongwen Bai, Yanning Zhang 0001 |
IGARSS | 3 |
| 2021 | GSDet: Object Detection in Aerial Images Based on Scale ReasoningabstractVariations in both object scale and style under different capture scenes (e.g., downtown, port) greatly enhance the difficulties associated with object detection in aerial images. Although ground sample distance (GSD) provides an apparent clue to address this issue, no existing object detection methods have considered utilizing this useful prior knowledge. In this paper, we propose the first object detection network to incorporate GSD into the object detection modeling process. More specifically, built on a two-stage detection framework, we adopt a GSD identification subnet converting the GSD regression into a probability estimation process, then combine the GSD information with the sizes of Regions of Interest (RoIs) to determine the physical size of objects. The estimated physical size can provide a powerful prior for detection by reweighting the weights from the classification layer of each category to produce RoI-wise enhanced features. Furthermore, to improve the discriminability among categories of similar size and make the inference process more adaptive, the scene information is also considered. The pipeline is flexible enough to be stacked on any two-stage modern detection framework. The improvement over the existing two-stage object detection methods on the DOTA dataset demonstrates the effectiveness of our method. Wei Li 0219, Wei Wei 0008, Lei Zhang 0054 |
IEEE Trans. Image Process. | 2 |
| 2021 | Embarrassingly Simple Binarization for Deep Single Imagery Super-Resolution NetworksabstractDeep convolutional neural networks (DCCNs) have shown pleasing performance in single image super-resolution (SISR). To deploy them onto real devices with limited storage and computational resources, a promising solution is to binarize the network, i.e., quantize each float-point weight and activation into 1 bit. However, existing works on binarizing DCNNs still suffer from severe performance degradation in SISR. To mitigate this problem, we argue that the performance degradation mainly comes from no appropriate constraint on the network weights, which causes it difficult to sensitively reverse the binarization results of these weights using the backpropagated gradient during training and thus limits the flexibility of network in respect of fitting extensive training samples. Inspired by this, we present an embarrassingly simple but effective binarization scheme for SISR, which can obviously relieve the performance degeneration resulted from network binarization and is applicable to different DCNN architectures. Specifically, we force each weight to follow a compact uniform prior, with which the weight will be given a very small absolute value close to zero and its binarization result can be straightforwardly reversed even by a small backpropagated gradient. By doing this, the flexibility and the generalization performance of the binarized network can be improved. Moreover, such a prior performs much better when introducing real identity shortcuts into the network. In addition, to avoid falling into bad local minima during training, we employ a pixel-wise curriculum learning strategy to learn the constrained weights in an easy-to-hard manner. Experiments on four SISR benchmark datasets demonstrate the effectiveness of the proposed binarization method in terms of binarizing different SISR network architectures, e.g., it even achieves performance comparable to the baseline with 5 quantization bits. Lei Zhang 0054, Zhiqiang Lang, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Deep Blind Hyperspectral Image Super-ResolutionabstractThe production of a high spatial resolution (HR) hyperspectral image (HSI) through the fusion of a low spatial resolution (LR) HSI with an HR multispectral image (MSI) has underpinned much of the recent progress in HSI super-resolution. The premise of these signs of progress is that both the degeneration from the HR HSI to LR HSI in the spatial domain and the degeneration from the HR HSI to HR MSI in the spectral domain are assumed to be known in advance. However, such a premise is difficult to achieve in practice. To address this problem, we propose to incorporate degeneration estimation into HSI super-resolution and present an unsupervised deep framework for "blind" HSIs super-resolution where the degenerations in both domains are unknown. In this framework, we model the latent HR HSI and the unknown degenerations with deep network structures to regularize them instead of using handcrafted (or shallow) priors. Specifically, we generate the latent HR HSI with an image-specific generator network and structure the degenerations in spatial and spectral domains through a convolution layer and a fully connected layer, respectively. By doing this, the proposed framework can be formulated as an end-to-end deep network learning problem, which is purely supervised by those two input images (i.e., LR HSI and HR MSI) and can be effectively solved by the backpropagation algorithm. Experiments on both natural scene and remote sensing HSI data sets show the superior performance of the proposed method in coping with unknown degeneration either in the spatial domain, spectral domain, or even both of them. Lei Zhang 0054, Jiangtao Nie, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Pixel-Aware Deep Function-Mixture Network for Spectral Super-ResolutionabstractSpectral super-resolution (SSR) aims at generating a hyperspectral image (HSI) from a given RGB image. Recently, a promising direction is to learn a complicated mapping function from the RGB image to the HSI counterpart using a deep convolutional neural network. This essentially involves mapping the RGB context within a size-specific receptive field centered at each pixel to its spectrum in the HSI. The focus thereon is to appropriately determine the receptive field size and establish the mapping function from RGB context to the corresponding spectrum. Due to their differences in category or spatial position, pixels in HSIs often require different-sized receptive fields and distinct mapping functions. However, few efforts have been invested to explicitly exploit this prior.To address this problem, we propose a pixel-aware deep function-mixture network for SSR, which is composed of a new class of modules, termed function-mixture (FM) blocks. Each FM block is equipped with some basis functions, i.e., parallel subnets of different-sized receptive fields. Besides, it incorporates an extra subnet as a mixing function to generate pixel-wise weights, and then linearly mixes the outputs of all basis functions with those generated weights. This enables us to pixel-wisely determine the receptive field size and the mapping function. Moreover, we stack several such FM blocks to further increase the flexibility of the network in learning the pixel-wise mapping. To encourage feature reuse, intermediate features generated by the FM blocks are fused in late stage, which proves to be effective for boosting the SSR performance. Experimental results on three benchmark HSI datasets demonstrate the superiority of the proposed method. Lei Zhang 0054, Zhiqiang Lang, Peng Wang 0023, Wei Wei 0008, Shengcai Liao, Ling Shao 0001, Yanning Zhang 0001 |
AAAI | 4 |
| 2020 | Unsupervised Adaptation Learning for Hyperspectral Imagery Super-ResolutionabstractThe key for fusion based hyperspectral image (HSI) super-resolution (SR) is to infer the posteriori of a latent HSI using appropriate image prior and likelihood that depends on degeneration. However, in practice the priors of high-dimensional HSIs can be extremely complicated and the degeneration is often unknown. Consequently most existing approaches that assume a shallow hand-crafted image prior and a pre-defined degeneration, fail to well generalize in real applications. To tackle this problem, we present an unsupervised adaptation learning (UAL) framework. Instead of directly modelling the complicated image prior, we propose to first implicitly learn a general image prior using deep networks and then adapt it to a specific HSI. Following this idea, we develop a two-stage SR network that leverages two consecutive modules: a fusion module and an adaptation module, to recover the latent HSI in a coarse-to-fine scheme. The fusion module is pretrained in a supervised manner on synthetic data to capture a spatial-spectral prior that is general across most HSIs. To adapt the learned general prior to the specific HSI under unknown degeneration, we introduce a simple degeneration network to assist learning both the adaptation module and the degeneration in an unsupervised way. In this way, the resultant image-specific prior and the estimated degeneration can benefit the inference of a more accurate posteriori, thereby increasing generalization capacity. To verify the efficacy of UAL, we extensively evaluate it on four benchmark datasets and report strong results that surpass existing approaches. Lei Zhang 0054, Jiangtao Nie, Wei Wei 0008, Yanning Zhang 0001, Shengcai Liao, Ling Shao 0001 |
CVPR | 3 |
| 2020 | Unsupervised Deep Hyperspectral Super-Resolution With Unregistered ImagesabstractFusion based hyperspectral image (HSI) super-resolution has long been the research focus of hyperspectral image processing since it can generate a high-resolution (HR) HSI in both spatial and spectral domains. However, the success of the existing fusion based HSI super-resolution methods depends on the premise that the images utilized for fusion (i.e. the input low-spatial-resolution HSI and the low-spectral-resolution multispectral image) are exactly registered. Although such a premise is too idealistic to comply with in real cases, few efforts have considered this problem. To fill this gap, we propose to incorporate image registration into HSI super-resolution for joint unsupervised learning in this study. Specifically, a spatial transformer network (STN) is introduced to learn the parameters of the affine transformation between the input two images. In order to avoid over-fitting, we constrain the STN with a novel constraint during learning. By doing this, both the STN and super-resolution network can be cast into a weighted joint learning model without any supervision from the latent HR HSI. Experimental results demonstrate the effectiveness of the proposed method in coping with unregistered input images. Jiangtao Nie, Lei Zhang 0054, Wei Wei 0008, Chen Ding 0002, Yanning Zhang 0001 |
ICME | 3 |
| 2020 | Deep Self-Supervised Learning for Few-Shot Hyperspectral Image ClassificationabstractDespite the success of deep learning based methods for hyperspectral imagery (HSI) classification, they demand amounts of labeled samples for training whereas the labeled samples in lots of applications are always insufficient due to the expensive manual annotation cost. To address this problem, we propose a two-branch deep learning based method for few-shot HSI classification, where two branches separately accomplish HSI classification in a cube-wise level and a cube-pair level. With a shared feature extractor sub-network, the self-supervised knowledge contained in the cube-pair branch provides an effective way to regularize the original few-shot HSI classification branch (i.e., cube-wise branch) with limited labeled samples, which thus improves the performance of HSI classification. The superiority of the proposed method on few-shot HSI classification is demonstrated experimentally on two HSI benchmark datasets. Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IGARSS | 3 |
| 2020 | Adaptive Importance Learning for Improving Lightweight Image Super-Resolution Network
Lei Zhang 0054, Peng Wang 0023, Chunhua Shen, Lingqiao Liu, Wei Wei 0008, Yanning Zhang 0001, Anton van den Hengel |
Int. J. Comput. Vis. | 5 |
| 2020 | Hyperspectral Image Classification With Data Augmentation and Classifier FusionabstractRecently, deep convolutional neural network (DCNN)-based methods have achieved much success in hyperspectral image (HSI) classification, when sufficient labeled samples are provided during training. However, due to the expensive cost of labeling in HSIs, only limited labeled samples can be given in practice, which often causes these methods to be overfitting. To address this problem, we present a new HSI classification method in this study, which is constructed in the following two steps. First, we establish a data mixture model to augment the labeled training set quadratically and train a DCNN-based classifier on it. Then, through randomly sampling the coefficient in the data mixture model, we obtain several independent classifiers and fuse them with a voting strategy to produce the final classification results. Since both data augmentation and classifier fusion are effective to deal with limited samples, the proposed method shows superior performance in the classification of HSIs, which can be demonstrated by the experimental results on two benchmark HSI data sets. Cong Wang 0013, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Towards Effective Deep Embedding for Zero-Shot LearningabstractZero-shot learning (ZSL) can be formulated as a cross-domain matching problem: after being projected into a joint embedding space, a visual sample will match against all candidate class-level semantic descriptions and be assigned to the nearest class. In this process, the embedding space underpins the success of such matching and is crucial for ZSL. In this paper, we conduct an in-depth study on the construction of embedding space for ZSL and posit that an ideal embedding space should satisfy two criteria: intra-class compactness and inter-class separability. While the former encourages the embeddings of visual samples of one class to distribute tightly close to the semantic description embedding of this class, the latter requires embeddings from different classes to be well separated from each other. Towards this goal, we present a simple but effective two-branch network to simultaneously map semantic descriptions and visual samples into a joint space, on which visual embeddings are forced to regress to their class-level semantic embeddings and the embeddings crossing classes are required to be distinguishable by a trainable classifier. Furthermore, we extend our method to a transductive setting to better handle the model bias problem in ZSL (i.e., samples from unseen classes tend to be categorized into seen classes) with minimal extra supervision. Specifically, we propose a pseudo labeling strategy to progressively incorporate the testing samples into the training process and thus balance the model between seen and unseen classes. Experimental results on five standard ZSL datasets show the superior performance of the proposed method and its transductive extension. Lei Zhang 0054, Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Wei Wei 0008, Yanning Zhang 0001, Anton van den Hengel |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Accurate Tensor Completion via Adaptive Low-Rank RepresentationabstractLow-rank representation-based approaches that assume low-rank tensors and exploit their low-rank structure with appropriate prior models have underpinned much of the recent progress in tensor completion. However, real tensor data only approximately comply with the low-rank requirement in most cases, viz., the tensor consists of low-rank (e.g., principle part) as well as non-low-rank (e.g., details) structures, which limit the completion accuracy of these approaches. To address this problem, we propose an adaptive low-rank representation model for tensor completion that represents low-rank and non-low-rank structures of a latent tensor separately in a Bayesian framework. Specifically, we reformulate the CANDECOMP/PARAFAC (CP) tensor rank and develop a sparsity-induced prior for the low-rank structure that can be used to determine tensor rank automatically. Then, the non-low-rank structure is modeled using a mixture of Gaussians prior that is shown to be sufficiently flexible and powerful to inform the completion process for a variety of real tensor data. With these two priors, we develop a Bayesian minimum mean-squared error estimate framework for inference. The developed framework can capture the important distinctions between low-rank and non-low-rank structures, thereby enabling more accurate model, and ultimately, completion. For various applications, compared with the state-of-the-art methods, the proposed model yields more accurate completion results. Lei Zhang 0054, Wei Wei 0008, Qinfeng Shi, Chunhua Shen, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Vehicle Re-Identification in Aerial Imagery: Dataset and ApproachabstractIn this work, we construct a large-scale dataset for vehicle re-identification (ReID), which contains 137k images of 13k vehicle instances captured by UAV-mounted cameras. To our knowledge, it is the largest UAV-based vehicle ReID dataset. To increase intra-class variation, each vehicle is captured by at least two UAVs at different locations, with diverse view-angles and flight-altitudes. We manually label a variety of vehicle attributes, including vehicle type, color, skylight, bumper, spare tire and luggage rack. Furthermore, for each vehicle image, the annotator is also required to mark the discriminative parts that helps them to distinguish this particular vehicle from others. Besides the dataset, we also design a specific vehicle ReID algorithm to make full use of the rich annotation information. It is capable of explicitly detecting discriminative parts for each specific vehicle and significantly outperforming the evaluated baselines and state-of-the-art vehicle ReID approaches. Peng Wang 0015, Bingliang Jiao, Lu Yang 0016, Shizhou Zhang, Wei Wei 0008, Yanning Zhang 0001 |
ICCV | 6 |
| 2019 | Deep Spectral Super-Resolution with Noisy InputabstractLearning based methods, e.g., sparse coding or deep convolutional neural networks (DCNNs) have underpinned much of recent progress in increasing the spectral resolution of an RGB image for hyperspectral image (HSI) super-resolution. However, these methods suffer severe performance loss, when the test RGB image distributed differently from the training set, e.g., being corrupted with random noise. To mitigate this problem, we propose an unsupervised deep spectral super-resolution method, which employs a DCNN to generate the latent HSI from an input RGB and encourages it to fit the input RGB image through down-sampling in spectral domain as well as a sparse gradient prior in spatial domain. Due to the powerful capacity of DCNN in capturing the low-level image statistics, the proposed method is able to automatically accommodate the noise corruption in the input RGB image. Experimental results shows the superior performance of the proposed method. Zhiqiang Lang, Lei Zhang 0054, Wei Wei 0008, Jiangtao Nie, Chunna Tian, Yanning Zhang 0001 |
IGARSS | 3 |
| 2019 | Unsupervised deep domain adaptation for hyperspectral image classificationabstractDeep neural networks have been proven to be a promising way for hyperspectral image (HSI) classification. Their success depends on a premise that source domain (i.e., training) and target domain (i.e., test) samples are identically distributed. However, due to various imaging environments, in practice obvious distribution discrepancy often exists between these two domains, which can dramatically reduce the capacity of the classifier trained in source domain generalizing to target domain. To mitigate this problem, we present a novel deep unsupervised domain adaptation framework for HSI classification, which can simultaneously align the distributions of two domains and learn a classifier in source domain. Firstly, we employ two auto-encoder networks to separately project the samples from two domains into two low-dimensional feature spaces. Then, a multi-level maximum mean discrepancy (MMD) loss is imposed on the feature space to reduce the distribution discrepancy between two domains. Given the resultant features, a classification subnet is further learned to classify the labeled samples in source domain. Since the classifier is trained based on the domain-invariant features, it can well generalize to the target domain. Experimental results on one benchmark cross-domain HSI datasets prove the superior performance of the proposed method. Wei Li 0219, Wei Wei 0008, Lei Zhang 0054, Cong Wang 0013, Yanning Zhang 0001 |
IGARSS | 2 |
| 2019 | Robust Deep Hyperspectral Imagery Super-ResolutionabstractFusing a low spatial resolution (LR) hyperspectral image (HSI) with a high spatial resolution (HR) multi-spectral image (MSI) is an effective way for HSI super-resolution. When the input LR HSI and the HR MSI are clean, most of existing fusion based methods can produce pleasing results. However, the input HSI and MSI are often corrupted with random noise in practice, which can greatly degrade the performance of these methods. To address this problem, we present a robust deep HSI super-resolution method in this study. In contrast to leveraging a heuristic shallow sparsity or low-rank prior in previous methods, we propose to employ a deep convolution neural network as the prior of the latent HR HSI. With such a prior, the fusion based HSI super-resolution can be formulated as an end-to-end deep learning problem, which can be effectively solved with the back-propagation algorithm. Due to the deep structure, the proposed image prior is able to capture more powerful statistics of the latent HR HSI, and thus can still produce pleasing results with noisy input images. Experimental results on two benchmark datasets demonstrate the effectiveness of the proposed method. Jiangtao Nie, Lei Zhang 0054, Cong Wang 0013, Wei Wei 0008, Yanning Zhang 0001 |
IGARSS | 4 |
| 2019 | Improving Hyperspectral Image Classification with Unsupervised Knowledge LearningabstractRecently, deep convolutional neural networks(DCNNs) based methods have shown pleasing performance in hyperspectral image(HSI) classification. However, due to extensive coefficients resulted by the deep structure, these methods are prone to be overfitting during training, especially when the labeled samples are limited. To address this problem, we propose to learn the unsupervised knowledge from both unlabeled and labeled samples to regularize the conventional supervised learning. Following this idea, we present a two-branch network, in which two branches are separately utilized to perform the clustering and classification based on a shared feature extraction module. Thanks to the shared structure, the crucial unsupervised information (e.g., intra-cluster similarity & inter-cluster dissimilarity, etc.) can be injected into the supervised learning procedure, and thus leads to improved generalization capacity. Experiments on two widely used HSI datasets show the superior performance of the proposed method. Wei Wei 0008, Lei Zhang 0054, Yanning Zhang 0001 |
IGARSS | 2 |
| 2019 | Jointing Cross-Modality Retrieval to Reweight Attributes for Image Caption Generation
Mengmeng Jiang, Donghu Deng, Wei Wei 0008, Chunna Tian |
PRCV (3) | 6 |
| 2019 | SS-GANs: Text-to-Image via Stage by Stage Generative Adversarial Networks
Ming Tian, Yuting Xue, Chunna Tian, Donghu Deng, Wei Wei 0008 |
PRCV (2) | 6 |
| 2019 | Accurate imagery recovery using a multi-observation patch model
Lei Zhang 0054, Wei Wei 0008, Qinfeng Shi, Chunhua Shen, Anton van den Hengel, Yanning Zhang 0001 |
Inf. Sci. | 2 |
| 2019 | Robust Hyperspectral Image Domain Adaptation With Noisy LabelsabstractIn hyperspectral image (HSI) classification, domain adaptation (DA) methods have been proved effective to address unsatisfactory classification results caused by the distribution difference between training (i.e., source domain) and testing (i.e., target domain) pixels. However, these methods rely on accurate labels in source domain, and seldom consider the performance drop resulted by noisy label, which often happens since labeling pixel in HSI is a challenging task. To improve the robustness of DA method to label noise, we propose a new unsupervised HSI DA method, which is constructed from both feature-level and classifier-level. First, a linear transformation function is learned in feature-level to align the source (domain) subspace with the target (domain) subspace. Then, a robust low-rank representation based classifier is developed to well cope with the features obtained from the aligned subspace. Since both subspace alignment and the classifier are immune to noisy labels, the proposed method obtains good classification results when confronting with noisy labels in source domain. Experimental results on two DA benchmarks demonstrate the effectiveness of the proposed method. Wei Wei 0008, Wei Li 0219, Lei Zhang 0054, Cong Wang 0013, Peng Zhang 0005, Yanning Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | Fast-Convergent Fully Connected Deep Learning Model Using Constrained Nodes Input
Chen Ding 0002, Ying Li 0017, Lei Zhang 0054, Lu Yang 0016, Wei Wei 0008, Yong Xia 0001, Yanning Zhang 0001 |
Neural Process. Lett. | 6 |
| 2019 | Unsupervised Domain Adaptation Using Robust Class-Wise MatchingabstractUnsupervised domain adaptation (DA) enables a classifier trained on data from one domain to be applied to data from another without labels. Given that the key to transferring a classifier across domains is to mitigate the data distribution mismatch for each class, most previous works completely or partially focus on global distribution matching across domains. The global data space, however, can be complicated, which makes modeling the global distribution difficult. To mitigate this problem, we present a novel unsupervised DA framework where the DA problem is addressed by proposing a robust class-wise matching strategy. Specifically, through minimizing a maximum mean discrepancy-based class-wise fisher discriminant across domains, this framework jointly optimizes two modules: a transferable feature learning module that reduces the distribution discrepancy between the same classes as well as increasing the distribution discrepancy between different classes across domains by a linear projection, and a robust classifier that exploits both the supervised information in source domain and the unsupervised low-rank property of target domain. In experiments on three DA benchmark data sets, the proposed framework shows the state-of-the-art performance. Lei Zhang 0054, Peng Wang 0023, Wei Wei 0008, Hao Lu 0003, Chunhua Shen, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Intracluster Structured Low-Rank Matrix Analysis Method for Hyperspectral DenoisingabstractHyperspectral images (HSIs) denoising aims at eliminating the noise generated during the acquisition and transmission of HSIs. Since denoising is an ill-posed problem, utilizing proper knowledge of HSIs as regularization is essential for a good denoiser. Many HSI denoising methods have been proposed to leverage various prior knowledge, e.g., total variation, sparsity, and so on. Among those knowledge, a low-rank property has been shown to be effective for HSI denoising since it has the ability to deal with the missing values. However, most existing low-rank methods seldom consider mining the useful structures inside the low-rank matrix for a better denoising result. In addition, the rank number needs to be assigned manually. To address these problems, we propose an intracluster structured low-rank matrix analysis method for HSI denoising. First, we divide the original HSI into some clusters by taking advantages of both local similarity and nonlocal similarity structures, with which the resulted clusters are simpler and show more obvious low-rank property. Second, with singular value decomposition on the low-rank matrix in each cluster, the structured sparsity is modeled among the singular values to capture the structure of the low-rank matrix. Finally, an efficient optimization method is proposed to learn the structured sparsity adaptively from the data, as well as to inversely estimate the latent clean HSI from the noisy counterpart. The proposed method can not only obtain better denoising results compared with the-state-of-the-art methods but also automatically determine the rank number. Extensive experimental results demonstrate the effectiveness of the proposed method. Wei Wei 0008, Lei Zhang 0054, Yining Jiao, Chunna Tian, Cong Wang 0013, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Learning Discriminative Compact Representation for Hyperspectral Imagery ClassificationabstractAbundant spectral information of hyperspectral images (HSIs) has shown an obvious advantage in improving the performance of classification in the remote sensing domain. However, this is paid by the expensive consumption on the computation, transmission, as well as storage of HSIs. To address this problem, we propose to learn the discriminative compact representation for HSIs classification, which not only greatly reduces the data redundancy in the image but also preserves the discriminative information required for pixelwise classification in HSIs. To this end, we present a multi-task deep learning framework, which integrates HSIs autoencoding and classification into a two-branch deep neural network for jointly learning. In the network, we employ an encoder block to learn the compact representation of the input HSI via compression in the spectral domain. Being fed with the compact representation, the autoencoding branch then employs a decoder block to reconstruct the input HSI, while the classification branch utilizes a classifier block to predict the label for each pixel. Through end-to-end joint learning, the compact representation is not only informative enough to accurately reconstruct the original HSI but also discriminative enough to appropriately label each pixel with the trained classier. Sufficient experimental results on four HSIs classification data sets demonstrate the effectiveness of the proposed framework. Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Urban Impervious Surface Extraction Based on the Integration of Remote Sensing Images and Social Media DataabstractThis paper presents an inspiring approach for accurate estimation of impervious surfaces, which exploits the strength of two kind of heterogeneous features, i.e., physical features derived from satellite images and social features derived from social media datasets, respectively. On the one hand, we use a morphological attribute profiles guided spectral mixture analysis model to achieve estimates of physical features. On the other hand, we mine the social features from textual information of social media datasets. Then, a multivariable linear regression model is conducted to obtain the impervious surfaces. Experiment results, conducted with multi-spectral images collected by LANDSAT-8 and social media datasets scraped from Sina Weibo of Guangzhou city, suggest that our approach could lead to reliable and good estimation of the imperviousness. Wei Wei 0008, Jun Li 0009, Yanning Zhang 0001 |
IGARSS | 2 |
| 2018 | Accurate Spectral Super-Resolution from Single RGB Image Using Multi-scale CNN
Yiqi Yan, Lei Zhang 0054, Jun Li 0009, Wei Wei 0008, Yanning Zhang 0001 |
PRCV (2) | 4 |
| 2018 | Cluster Sparsity Field: An Internal Hyperspectral Imagery Prior for Reconstruction
Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Chunhua Shen, Anton van den Hengel, Qinfeng Shi |
Int. J. Comput. Vis. | 2 |
| 2018 | Salient object detection in hyperspectral imagery using multi-scale spectral-spatial gradient
Lei Zhang 0054, Yanning Zhang 0001, Hangqi Yan, Wei Wei 0008 |
Neurocomputing | 5 |
| 2018 | Color pornographic image detection based on color-saliency preserved mixture deformable part model
Chunna Tian, Xiangnan Zhang, Wei Wei 0008, Xinbo Gao 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Exploiting Clustering Manifold Structure for Hyperspectral Imagery Super-ResolutionabstractFusing a low-resolution hyperspectral image (HSI) with a high-resolution (HR) conventional image into an HR HSI has become a prevalent HSIs super-resolution scheme. However, in most previous works, little attention has been paid on exploiting the underlying manifold structure in the spatial domain of the latent HR HSI. In this paper, we advance a provable prior knowledge that the clustering manifold structure of the latent HSI can be well preserved in the spatial domain of the input conventional image. Inspired by this, we first conduct clustering in the spatial domain of the input conventional image and adopt the intra-cluster self-expressiveness model to implicitly depict the clustering manifold structure, which enables learning the complicated manifold structure via solving a constrained ridge regression model without knowing the exact form of the manifold. Then, we incorporate the learned structure into a variational super-resolution framework to regularize the latent HSI. The resulted framework can be effectively optimized by a standard alternating direction method of multipliers. Since the learned structure can well depict the underlying spatial manifold of the latent HSI, the proposed method shows the state-of-the-art super-resolution performance on two benchmark data sets. Lei Zhang 0054, Wei Wei 0008, Chengcheng Bai, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Solving Constrained Combinatorial Optimisation Problems via MAP Inference without High-Order PenaltiesabstractSolving constrained combinatorial optimisation problems via MAP inference is often achieved by introducing extra potential functions for each constraint. This can result in very high order potentials, e.g. a 2nd-order objective with pairwise potentials and a quadratic constraint over all N variables would correspond to an unconstrained objective with an order-N potential. This limits the practicality of such an approach, since inference with high order potentials is tractable only for a few special classes of functions. We propose an approach which is able to solve constrained combinatorial problems using belief propagation without increasing the order. For example, in our scheme the 2nd-order problem above remains order 2 instead of order N. Experiments on applications ranging from foreground detection, image reconstruction, quadratic knapsack, and the M-best solutions problem demonstrate the effectiveness and efficiency of our method. Moreover, we show several situations in which our approach outperforms commercial solvers like CPLEX and others designed for specific constrained MAP inference problems. Zhen Zhang 0008, Qinfeng Shi, Julian J. McAuley, Wei Wei 0008, Yanning Zhang 0001, Rui Yao 0006, Anton van den Hengel |
AAAI | 4 |
| 2017 | When Unsupervised Domain Adaptation Meets Tensor RepresentationsabstractDomain adaption (DA) allows machine learning methods trained on data sampled from one distribution to be applied to data sampled from another. It is thus of great practical importance to the application of such methods. Despite the fact that tensor representations are widely used in Computer Vision to capture multi-linear relationships that affect the data, most existing DA methods are applicable to vectors only. This renders them incapable of reflecting and preserving important structure in many problems. We thus propose here a learning-based method to adapt the source and target tensor representations directly, without vectorization. In particular, a set of alignment matrices is introduced to align the tensor representations from both domains into the invariant tensor subspace. These alignment matrices and the tensor subspace are modeled as a joint optimization problem and can be learned adaptively from the data using the proposed alternative minimization scheme. Extensive experiments show that our approach is capable of preserving the discriminative power of the source domain, of resisting the effects of label noise, and works effectively for small sample sizes, and even one-shot DA. We show that our method outperforms the state-of-the-art on the task of cross-domain visual recognition in both efficacy and efficiency, and particularly that it outperforms all comparators when applied to DA of the convolutional activations of deep convolutional networks. Hao Lu 0003, Lei Zhang 0054, Zhiguo Cao 0001, Wei Wei 0008, Ke Xian, Chunhua Shen, Anton van den Hengel |
ICCV | 4 |
| 2017 | Hyperspectral image super-resolution extending: An effective fusion based method without knowing the spatial transformation matrixabstractHyperspectral image (HSI) super-resolution, a technique to obtain higher (often spatial) resolution image from the original image, has been extensively studied and applied to lots of fields such as computer vision, remote sensing, etc. Though fusion based method has achieved state-of-the-art result, it always assume the spatial transformation matrix is given in advance, whereas such a matrix is actually unknown in reality. An unsuitable given matrix will deteriorate the superresolution result greatly. To address this issue, we propose a novel fusion based HSI super-resolution method without knowing the spatial transformation matrix. Specifically, we incorporate super-resolution and spatial transformation matrix estimation into a unified framework. We alternately estimate the matrix and the higher spatial resolution HSI. We find that without given the spatial transformation matrix, the proposed method can obtain more accurate reconstruction result compared with other competing methods. Experimental results demonstrate the effectiveness of the proposed method. Lei Zhang 0054, Chunna Tian, Chen Ding 0002, Yanning Zhang 0001, Wei Wei 0008 |
ICME | 6 |
| 2017 | Dynamic Programming Bipartite Belief Propagation For Hyper Graph MatchingabstractHyper graph matching problems have drawn attention recently due to their ability to embed higher order relations between nodes. In this paper, we formulate hyper graph matching problems as constrained MAP inference problems in graphical models. Whereas previous discrete approaches introduce several global correspondence vectors, we introduce only one global correspondence vector, but several local correspondence vectors. This allows us to decompose the problem into a (linear) bipartite matching problem and several belief propagation sub-problems. Bipartite matching can be solved by traditional approaches, while the belief propagation sub-problem is further decomposed as two sub-problems with optimal substructure. Then a newly proposed dynamic programming procedure is used to solve the belief propagation sub-problem. Experiments show that the proposed methods outperform state-of-the-art techniques for hyper graph matching. Zhen Zhang 0008, Julian J. McAuley, Wei Wei 0008, Yanning Zhang 0001, Qinfeng Shi |
IJCAI | 4 |
| 2017 | Structured Sparse Coding-Based Hyperspectral Imagery Denoising With Intracluster FilteringabstractSparse coding can exploit the intrinsic sparsity of hyperspectral images (HSIs) by representing it as a group of sparse codes. This strategy has been shown to be effective for HSI denoising. However, how to effectively exploit the structural information within the sparse codes (structured sparsity) has not been widely studied. In this paper, we propose a new method for HSI denoising, which uses structured sparse coding and intracluster filtering. First, due to the high spectral correlation, the HSI is represented as a group of sparse codes by projecting each spectral signature onto a given dictionary. Then, we cast the structured sparse coding into a covariance matrix estimation problem. A latent variable-based Bayesian framework is adopted to learn the covariance matrix, the sparse codes, and the noise level simultaneously from noisy observations. Although the considered strategy is able to perform denoising through accurately reconstructing spectral signatures, an inconsistent recovery of sparse codes may corrupt the spectral similarity in each spatial homogeneous cluster within the scene. To address this issue, an intracluster filtering scheme is further employed to restore the spectral similarity in each spatial cluster, which results in better denoising results. Our experimental results, conducted using both simulated and real HSIs, demonstrate that the proposed method outperforms several state-of-the-art denoising methods. Wei Wei 0008, Lei Zhang 0054, Chunna Tian, Antonio Plaza, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Pairwise Matching through Max-Weight Bipartite Belief PropagationabstractFeature matching is a key problem in computer vision and pattern recognition. One way to encode the essential interdependence between potential feature matches is to cast the problem as inference in a graphical model, though recently alternatives such as spectral methods, or approaches based on the convex-concave procedure have achieved the state-of-the-art. Here we revisit the use of graphical models for feature matching, and propose a belief propagation scheme which exhibits the following advantages: (1) we explicitly enforce one-to-one matching constraints, (2) we offer a tighter relaxation of the original cost function than previous graphical-model-based approaches, and (3) our sub-problems decompose into max-weight bipartite matching, which can be solved efficiently, leading to orders-of-magnitude reductions in execution time. Experimental results show that the proposed algorithm produces results superior to those of the current state-of-the-art. Zhen Zhang 0008, Qinfeng Shi, Julian J. McAuley, Wei Wei 0008, Yanning Zhang 0001, Anton van den Hengel |
CVPR | 4 |
| 2016 | Cluster Sparsity Field for Hyperspectral Imagery Denoising
Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Chunhua Shen, Anton van den Hengel, Qinfeng Shi |
ECCV (5) | 2 |
| 2016 | Salient object detection in hyperspectral imagery using spectral gradient contrastabstractSalient object detection in hyperspectral imagery has drawn people's attention in recent years. Some detection methods which focus on extending Itti's visual saliency model into spectral domain have been proposed. However, these methods are sensitive to high-contrast edges and cannot preserve boundary of salient object well. To address these shortcomings, we propose a region-based spectral gradient contrast method for salient object detection in this paper. First, we calculate gradient along each spectral vector of the input hyperspectral imagery. Then over segmentation and clustering methods are applied on gradient data to get a group of image regions. Finally, center prior and local contrast are employed to compute the saliency score of each region, with which the salient object can be obtained. Experimental results on four datasets demonstrate that the proposed method outperforms several competing methods on detection accuracy. Hangqi Yan, Yanning Zhang 0001, Wei Wei 0008, Lei Zhang 0054 |
IGARSS | 3 |
| 2016 | Hyperspectral imagery denoising using covariance matrix estimation based structured sparse coding and intra-cluster filteringabstractSparse coding provides an excellent image prior for hyperspectral images (HSIs) denoising. However, on one hand, it is challenging to capture the structure within each sparse code for improving the reconstruction accuracy, on the other hand, the inconsistent recovery of the sparse codes corrupts the spectrum similarity in each homogeneous cluster of the HSI. To address these problems, we first propose a novel covariance matrix estimation based structured sparse coding method, where the sparse code matrix is modeled by a matrix normal distribution with a full covariance matrix. By estimating the covariance matrix with a latent variable based Bayesian framework, the data-dependent and noise-robust structure for each sparse code is learned from the noisy observation, with which the sparse codes are reconstructed accurately. Then, an intra-cluster filtering is employed to restore the spectrum similarity in each cluster. Experimental results demonstrate that the proposed method outperforms several state-of-the-art methods in HSIs denoising. Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Cong Wang 0013 |
IGARSS | 2 |
| 2016 | Facial expression transfer method based on frequency analysis
Wei Wei 0008, Chunna Tian, Stephen J. Maybank, Yanning Zhang 0001 |
Pattern Recognit. | 1 |
| 2016 | Dictionary Learning for Promoting Structured Sparsity in Hyperspectral Compressive SensingabstractThe ability to accurately represent a hyperspectral image (HSI) as a combination of a small number of elements from an appropriate dictionary underpins much of the recent progress in hyperspectral compressive sensing (HCS). Preserving structure in the sparse representation is critical to achieving an accurate reconstruction but has thus far only been partially exploited because existing methods assume a predefined dictionary. To address this problem, a structured sparsity-based hyperspectral blind compressive sensing method is presented in this study. For the reconstructed HSI, a data-adaptive dictionary is learned directly from its noisy measurements, which promotes the underlying structured sparsity and obviously improves reconstruction accuracy. Specifically, a fully structured dictionary prior is first proposed to jointly depict the structure in each dictionary atom as well as the correlation between atoms, where the magnitude of each atom is also regularized. Then, a reweighted Laplace prior is employed to model the structured sparsity in the representation of the HSI. Based on these two priors, a unified optimization framework is proposed to learn both the dictionary and sparse representation from the measurements by alternatively optimizing two separate latent variable Bayes models. With the learned dictionary, the structured sparsity of HSIs can be well described by the reweighted Laplace prior. In addition, both the learned dictionary and sparse representation are robust to noise corruption in the measurements. Extensive experiments on three hyperspectral data sets demonstrate that the proposed method outperforms several state-of-the-art HCS methods in terms of the reconstruction accuracy achieved. Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Chunhua Shen, Anton van den Hengel, Qinfeng Shi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Exploring Structured Sparsity by a Reweighted Laplace Prior for Hyperspectral Compressive SensingabstractHyperspectral compressive sensing (HCS) can greatly reduce the enormous cost of hyperspectral images (HSIs) on imaging, storage, and transmission by only collecting a few compressive measurements in the image acquisition. One of the most challenging problems for HCS is how to reconstruct the HSI accurately from such a few measurements. It has been proved that introducing structure information into sparsity prior can improve the reconstruction performance of standard compressive sensing models. However, the structured sparsity of HSIs is unknown in reality and easily affected by random noise, which makes it difficult to explore the structured sparsity in HCS. To address this problem, we propose a novel reweighted Laplace prior-based HCS method in this paper. First, a hierarchical reweighted Laplace prior is proposed to model the distribution of sparsity in an HSI, which relieves the undemocratic penalization of traditional Laplace prior on nonzero coefficients of a sparse signal. Then, a latent variable-based Bayesian model is employed to learn the optimal configuration of the reweighted Laplace prior from the measurements. This model unifies signal recovery, sparsity prior learning, and noise estimation into a variational framework, where these three tasks are alternatively optimized till convergence. The finally learned sparsity prior can well represent the underlying structure in the sparse signal and is adaptive to the unknown noise. These advantages together improve the reconstruction accuracy of HCS obviously. Moreover, the proposed method is extended to learn a matrix normal distribution-based prior with a full covariance matrix, which depicts the underlying structure in the sparse signal better. As a result, the reconstruction accuracy is further improved. Extensive experimental results on three hyperspectral data sets demonstrate that the proposed method outperforms several state-of-the-art HCS methods in terms of the reconstruction accuracy. Lei Zhang 0054, Wei Wei 0008, Chunna Tian, Fei Li 0011, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Reweighted laplace prior based hyperspectral compressive sensing for unknown sparsityabstractCompressive sensing(CS) has been exploited for hype-spectral image(HSI) compression in recent years. Though it can greatly reduce the costs of computation and storage, the reconstruction of HSI from a few linear measurements is challenging. The underlying sparsity of HSI is crucial to improve the reconstruction accuracy. However, the sparsity of HSI is unknown in reality and varied with different noise, which makes the sparsity estimation difficult. To address this problem, a novel reweighted Laplace prior based hyperspectral compressive sensing method is proposed in this study. First, the reweighted Laplace prior is proposed to model the distribution of sparsity in HSI. Second, the latent variable Bayes model is employed to learn the optimal configuration of the reweighted Laplace prior from the measurements. The model unifies signal recovery, prior learning and noise estimation into a variational framework to infer the parameters automatically. The learned sparsity prior can represent the underlying structure of the sparse signal very well and is adaptive to the unknown noise, which improves the reconstruction accuracy of HSI. The experimental results on three hyperspectral datasets demonstrate the proposed method outperforms several state-of-the-art hyperspectral CS methods on the reconstruction accuracy. Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Chunna Tian, Fei Li 0011 |
CVPR | 2 |
| 2015 | Hyperspectral Compressive Sensing Using Manifold-Structured Sparsity PriorabstractTo reconstruct hyperspectral image (HSI) accurately from a few noisy compressive measurements, we present a novel manifold-structured sparsity prior based hyperspectral compressive sensing (HCS) method in this study. A matrix based hierarchical prior is first proposed to represent the spectral structured sparsity and spatial unknown manifold structure of HSI simultaneously. Then, a latent variable Bayes model is introduced to learn the sparsity prior and estimate the noise jointly from measurements. The learned prior can fully represent the inherent 3D structure of HSI and regulate its shape based on the estimated noise level. Thus, with this learned prior, the proposed method improves the reconstruction accuracy significantly and shows strong robustness to unknown noise in HCS. Experiments on four real hyperspectral datasets show that the proposed method outperforms several state-of-the-art methods on the reconstruction accuracy of HSI. Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Fei Li 0011, Chunhua Shen, Qinfeng Shi |
ICCV | 2 |
| 2015 | Visual Tracking Based on the Adaptive Color Attention Tuned Sparse Generative Object ModelabstractThis paper presents a new visual tracking framework based on an adaptive color attention tuned local sparse model. The histograms of sparse coefficients of all patches in an object are pooled together according to their spatial distribution. A particle filter methodology is used as the location model to predict candidates for object verification during tracking. Since color is an important visual clue to distinguish objects from background, we calculate the color similarity between objects in the previous frames and the candidates in current frame, which is adopted as color attention to tune the local sparse representation-based appearance similarity measurement between the object template and candidates. The color similarity can be calculated efficiently with hash coded color names, which helps the tracker find more reliable objects during tracking. We use a flexible local sparse coding of the object to evaluate the degeneration degree of the appearance model, based on which we build a model updating mechanism to alleviate drifting caused by temporal varying multi-factors. Experiments on 76 challenging benchmark color sequences and the evaluation under the object tracking benchmark protocol demonstrate the superiority of the proposed tracker over the state-of-the-art methods in accuracy. Chunna Tian, Xinbo Gao 0001, Wei Wei 0008 |
IEEE Trans. Image Process. | 3 |
| 2014 | Non-sparse infinite-kernel learning for automated identification of Alzheimer's disease using PET imagingabstractMulti-kernel learning machine (MKLM) has recently been introduced to the research of computer-aided dementia identification and pathology progress tracking. Despite its good performance especially in case of using heterogeneous data, such learning schema and its variants usually utilize a L-l norm constraint that promotes sparse solutions, which may cause loss of potentially important information. In this paper, we propose the non-sparse infinite-kernel learning machine (NS-IKLM) for automated identification of Alzheimer cases from normal controls. In our approach, a modified constraint is utilized to promotes non-sparse solutions and kernel parameters are automatically tuned during the learning process. The proposed algorithm has been evaluated on a set of FDG-PET images selected from the Alzheimer's disease neuroimaing initiative (ADNI) cohort. Our results demonstrate that the proposed non-sparse NS-IKLM is able to achieve satisfying dementia identification at a relatively low computational cost. Yong Xia 0001, Shen Lu, Wei Wei 0008, David Dagan Feng, Yanning Zhang 0001 |
ICARCV | 3 |
| 2014 | Robust face pose classification method based on geometry-preserving visual phraseabstractConstructing the discriminative feature for face pose classification is challenging. Since the key facial points are co-occurring with different spatial layout in different poses, we propose a pose classification framework based on the local geometry-preserving visual phrase (GVP). The weighting strategy on GVP enhances the discriminability of the single word and the spatial layout in the high order phrase simultaneously. Thus, the co-occurring words and local geometric structure of phrase are discriminative to distinguish pose, flexible and robust to the multi-factor variations. The experimental results on Oriental Face database and PIE database show the superiority of our method compared with the PCA, LDA based methods and the tf-idf weighted BoW method. Wei Wei 0008, Chunna Tian, Yanning Zhang 0001 |
ICIP | 1 |
| 2014 | 3D total variation hyperspectral compressive sensing using unmixingabstractTo reduce the huge resource consumption in the hyperspectral imaging and transmission, this paper proposes a high-performance compression method. Specially, a novel 3D total variation prior is imposed on abundance fractions of end-members. In this method, compressed data is obtained by a random observation matrix in a compressive sensing way. Based on the hyperspectral linear mixed model and known endmembers, abundance fractions are estimated by an augmented Lagrangian method with the devised prior and then the original data is reconstructed. Extensive experimental results demonstrate the superiority of the proposed method to several state-of-art methods. Lei Zhang 0054, Yanning Zhang 0001, Wei Wei 0008, Fei Li 0011 |
IGARSS | 3 |
| 2014 | A jointly distributed semi-supervised topic model
Yanning Zhang 0001, Wei Wei 0008 |
Neurocomputing | 2 |
| 2013 | An associative saliency segmentation method for infrared targetsabstractAutomatic infrared (IR) target segmentation plays an important role in IR image analysis. Recent works have shown that exploiting visual attention model can improve target segmentation performance in visible images. However, when directly applied to IR images, those methods cannot guarantee the effectiveness due to the low contrast between targets and background, high noise, etc. To address above problem, a novel associative saliency-based visual attention model for IR images is proposed in this paper. First, an IR image is decomposed into assemble of homogeneous regions. With those regions, saliency based on region and edge contrast is constructed, respectively. Then associative saliency, generated from those two kinds of saliency, is used to extract IR target from background. The superiority of the proposed method is examined and demonstrated through a large number of the experiments using IR images. Lei Zhang 0054, Yanning Zhang 0001, Wei Wei 0008, Qingjie Meng |
ICIP | 3 |
| 2012 | A realistic dynamic facial expression transfer method
Yanning Zhang 0001, Wei Wei 0008 |
Neurocomputing | 2 |
| 2009 | Incremental Multi-view Face Tracking Based on General View Manifold
Wei Wei 0008, Yanning Zhang 0001 |
ACCV (2) | 1 |
| 2009 | Illumination Invariant Multi-pose Face TrackingabstractA novel face tracking algorithm is proposed which is less insensitive to large face appearance variation caused by lighting and viewpoints. The proposed dynamic face appearance model includes an illumination invariant pose manifold to represent the pose nonlinearity and an object-specific incremental multi-pose face model. The pose manifold is built with the patch alignment algorithm, where the tensor based illumination reduction is used to define the local patch. Particularly, the probability of the face belonging to the pose manifold can be used for choosing the object-specific face model. The experimental results show the face tracking model can successfully track faces under unseen poses and changing illuminations. Wei Wei 0008, Yanning Zhang 0001, Zenggang Lin |
ICIG | 1 |