EDBT 2026 Demo / reviewers in the wild / expert
Wenda Zhao 0003
dblp:166/7309-3
· DBLP profile ↗
43ranked-venue papers
27as first author
37since 2021 · last 2026
0000-0002-7463-6103ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 14 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 13 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 6 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | First-Order Cross-Domain Meta Learning for Few-Shot Remote Sensing Object ClassificationabstractRemote sensing images exhibit intrinsic domain complexity arising from multi-source sensor variances, which heterogeneity fundamentally challenges conventional cross-domain few-shot methods that assume simple distribution shifts. Addressing this, we propose a first-order Cross-Domain Meta Learning (CDML) for few-shot remote sensing object classification. CDML implements a dual-stage domain adaptation task as the fundamental meta-learning unit, and includes a cross-domain meta-train phase (CDMTrain) and a cross-domain meta-test phase (CDMTest). In CDMTrain, we propose an inner-loop multi-domain few-shot task sampling, which enables a teacher model encapsulate both cross-category discriminative features and authentic inter-domain distributional divergence. This alternating cyclic learning paradigm captures genuine domain shifts, with each update direction progressively guiding the model toward parameters that balance multi-domain performance. In CDMTest, we evaluate a domain diversity enhancement by transferring teacher parameters to the student model for cross-domain capability assessment on the reserved pseudo-unseen domain. The task-level design progressively improves domain generalization through iterative domain adaptive task learning. Meanwhile, to mitigate the conflicts and inadequacies caused by multi-domain scenarios, we propose a learnable affine transformation model. It adaptively learns affine transformation parameters through intermediate layer features to fine-tune the update direction. Extensive experiments on five remote sensing classification benchmarks demonstrate a superior performance of the proposed method compared with the state-of-the-art methods. Wenda Zhao 0003, Huchuan Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | CDTFusion: Crossing Domain and Task for Infrared and Visible Image FusionabstractInfrared and visible images present different domains that hinder the fusion process, thereby losing texture details. Besides, the low-level fusion and subsequent high-level segmentation appear cross-task feature gap that impedes their mutual promotion, causing blurred object edges. Addressing the above issues, this paper proposes a novel infrared and visible image fusion method that simultaneously crosses domain and task. First, a swap image translation strategy is built to transfer the features of visible and infrared images into an adaptive domain. Meanwhile, a global-local constraint is introduced to achieve overall domain space transfer, and shorten their feature distance. Second, a task interaction & query module is designed to explore the cross-task feature interactive relationship, which is then used as a bridge to realize the gradient backpropagation. Thus, a fine-grained mapping from the segmentation feature to fusion feature is obtained. Extensive experiments demonstrate that the proposed method exhibits superior fusion and segmentation performance than the state-of-the-art methods. Wenda Zhao 0003, You He 0002, Huchuan Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Context-Infused Trajectories: Enhancing Context and Frame Consistency in Reasoning Video Object SegmentationabstractReasoning video object segmentation (ReaVOS) aims to segment referred objects in video sequences based on implicit and complex linguistic queries. Existing methods typically compress limited video frames into pooled representations and prompt multimodal large language models (MLLMs) to generate a single global segmentation token. However, this strategy lacks explicit contextual guidance and causes substantial loss of spatial details, limiting capability and segmentation consistency. To overcome these limitations, we introduce Context-infused Consistent Video Segmentor (CiCVS), a novel framework leveraging contextual information to guide generation of temporally coherent and accurate mask trajectories. CiCVS incorporates a Hierarchical Frame Sampling (HFS) module, which globally samples support frames across the entire video to ensure broad temporal coverage, and then uniformly selects target frames within the support set. It also employs a Contextual Token Prompting (CTP) module, which utilizes contextual cues from support frames to guide the MLLM in generating specialized tokens for various target frames, enabling the model to capture intricate temporal patterns and ensure consistency across long-range sequences. At the core of CTP is the Multimodal Injection Compressor (MIC) block, which efficiently integrates support frame features and textual semantic information into a compact set of latent queries, enhancing temporal-level object perception. To further advance the ReaVOS field, we introduce the CoCoRVOS benchmark, which features more temporally intricate reasoning instructions and a diverse set of video scenarios. Extensive experiments demonstrate that CiCVS establishes a new state-of-the-art on multiple benchmarks, achieving significant improvements in $\mathcal {J}\& \mathcal {F}$ scores, including +2.7 on CoCoRVOS, +1.4 on ReVOS, and +7.0 on ReasonVOS, underscoring its superior contextual reasoning and segmentation capabilities. Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Qi Xu 0008, Wenda Zhao 0003, Jin Zhan, Huchuan Lu |
IEEE Trans. Image Process. | 5 |
| 2025 | UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation EnhancementabstractZero-shot domain adaptation (ZSDA) presents substantial challenges due to the lack of images in the target domain. Previous approaches leverage Vision-Language Models (VLMs) to tackle this challenge, exploiting their zero-shot learning capabilities. However, these methods primarily address domain distribution shifts and overlook the misalignment between the detection task and VLMs, which rely on manually crafted prompts. To overcome these limitations, we propose the unified prompt and representation enhancement (UPRE) framework, which jointly optimizes both textual prompts and visual representations. Specifically, our approach introduces a multi-view domain prompt that combines linguistic domain priors with detection-specific knowledge, and a visual representation enhancement module that produces domain style variations. Furthermore, we introduce multi-level enhancement strategies, including relative domain distance and positive-negative separation, which align multi-modal representations at the image level and capture diverse visual representations at the instance level, respectively. Extensive experiments conducted on nine benchmark datasets demonstrate the superior performance of our framework in ZSDA detection scenarios. Code is available at https://github.com/AMAP-ML/UPRE. Xiao Zhang 0050, Fei Wei, Wenda Zhao 0003, Feiyi Li, Xiangxiang Chu |
ICCV | 4 |
| 2025 | Self-calibrated region-level regression for crowd counting
Jiawen Zhu 0003, Wenda Zhao 0003, You He 0002, Huchuan Lu |
Sci. China Inf. Sci. | 2 |
| 2025 | FreeFusion: Infrared and Visible Image Fusion via Cross Reconstruction LearningabstractExisting fusion methods empirically design elaborate fusion losses to retain the specific features from source images. Since image fusion has no ground truth, the hand-crafted losses may not make the fused images cover all the vital features, and then affect the performance of the high-level tasks. Here, there are two main challenges: domain discrepancy among source images and semantic mismatch at different-level tasks. This paper proposes an infrared and visible image fusion via cross reconstruction learning, which doesn't using any hand-crafted fusion losses, but prompts the network to adaptively fuse complementary information of source images. Firstly, we design a cross reconstruction learning model that decouples the fusion features to reconstruct another-modality source image. Thus, the fusion network is forced to learn the domain-adaptive representations of two modal features, which enables their domain alignment in a latent space. Secondly, we propose a dynamic interactive fusion strategy that builds a correlation matrix between fusion features and object semantic features to overcome the semantic mismatch. Further, we enhance the strong correlation features and suppress the weak correlation features to improve the interactive ability. Extensive experiments on three datasets demonstrate the superior fusion performance compared to the state-of-the-art methods, concurrently facilitating the segmentation accuracy. Wenda Zhao 0003, Hengshuai Cui, You He 0002, Huchuan Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Hybrid Gaussian Deformation for Efficient Remote Sensing Object DetectionabstractLarge-scale high-resolution remote sensing images (LSHR) are increasingly adopted for object detection, since they capture finer details. However, LSHR imposes a substantial computational cost. Existing methods explore lightweight backbones and advanced oriented bounding box regression mechanisms. Nevertheless, they still rely on high-resolution inputs to maintain detection accuracy. We observe that LSHR comprise extensive background areas that can be compressed to reduce unnecessary computation, while object regions contain details that can be reserved to improve detection accuracy. Thus, we propose a hybrid Gaussian deformation module that dynamically adjusts the sampling density at each location based on its relevance to the detection task, i.e., high-density sampling preserves more object regions and better retains detailed features, while low-density sampling diminishes the background proportion. Further, we introduce a bilateral deform-uniform detection framework to exploit the potential of the deformed sampled low-resolution images and original high-resolution images. Specifically, a deformed deep backbone takes the deformed sampled images as inputs to produce high-level semantic information, and a uniform shallow backbone takes the original high-resolution images as inputs to generate precise spatial location information. Moreover, we incorporate a deformation-aware feature registration module that calibrates the spatial information of deformed features, preventing regression degenerate solutions while maintaining feature activation. Subsequently, we introduce a feature relationship interaction fusion module to balance the contributions of features from both deformed and uniform backbones. Comprehensive experiments on three challenging datasets show that our method achieves superior performance compared with the state-of-the-art methods. Wenda Zhao 0003, Xiao Zhang 0050, Huchuan Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Remote Sensing Image Generation via Object Text DecouplingabstractRemote sensing images usually reveal various objects with complex structures and different locations within vast ground area backgrounds. That leads to a major challenge for conventional generative models in handling remote sensing objects with correct shapes and clear textures. Integrating additional object-level controls can be a potential solution to improve generation quality, yet previous approaches inject the object-related conditions by specifying their locations, causing a limitation in object layout in generated results. To enable high object fidelity, high layout diversity and object customizable generation for remote sensing images, we propose a remote sensing image generation via object text decoupling, namely OTD-GAN. OTD-GAN takes advantage of the inherent text-to-image generation procedure and adaptively integrates the decoupled textual representations of visual objects into the global captions, thus achieving object-level controls without layout restrictions. Specifically, we design an object text decoupling module to predict a semantically consistent textual representation for each object. By decoupling the textual representation into a class invariant part and an object specific part, the converted representation is able to catch general semantic for similar objects as well as differentiated details for individual objects. After that, we use an object text semantic enhancement module to fuse the obtained object text representations with the global captions to enrich the object-related semantic within the textual modality. As a result, the generator will benefit from the object conditions and reinforce the generation quality while remaining flexibility to create diverse layouts. Extensive experiments on remote sensing image-caption datasets including NWPU-Captions and RSICD demonstrate that our method achieves leading performance compared to existing state-of-the-art approaches. Wenda Zhao 0003, Zhepu Zhang, Fan Zhao 0005, You He 0002, Huchuan Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | GCBF: Grouped Cross-Band Fusion Network for Multispectral Scene ClassificationabstractRemote sensing scene classification is a crucial task for remote sensing image interpretation. Existing multispectral scene classification methods have overlooked the interrelationships between different spectral bands, which limits the mining of complementary information within the images. Addressing this issue, we propose a grouped cross-band fusion (GCBF) network for remote sensing multispectral scene classification to take full advantage of complementary information between various spectral bands. Firstly, we separate the various bands of the given multispectral image into different groups to better capture the characteristics of each spectral band. Then, we use the existing UniFormer as a feature extractor to learn the representations of red, green, and blue (RGB) bands. For the spectral bands other than RGB, we propose a new network called multi-stage grouped spectral feature extraction (MGSFE) network to learn discriminative representations. We also draw inspiration from the band combination in the field of remote sensing and introduce a cross-band attention fusion (CBAF) module designed to adaptively merge features from both the RGB bands and other spectral bands. Extensive experiments on three widely used remote sensing multispectral scene classification datasets of BigEarthNet, SEN12MS, and EuroSAT demonstrate the superiority of our proposed method compared with several state-of-the-art (SOTA) methods. Jin Li 0069, Yu Liu 0005, Wenda Zhao 0003, Zhizhuo Jiang, Xueqian Wang 0002, Bolun Zheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Rotation-Invariant Knowledge Distillation for Remote Sensing Object DetectionabstractDetecting small-rotated objects in remote sensing remains a challenging task due to feature dilution and insufficient rotation invariance. Feature dilution arises when small object features are overwhelmed by background noise and progressively lost as network depth increases. Meanwhile, the lack of rotation invariance stems from the fixed nature of convolution, which struggles to handle arbitrary orientations. To address these challenges, we propose a rotation-invariant knowledge distillation, a visual-language models (VLMs) driven knowledge distillation framework tailored for optimizing small-rotated object detection in remote sensing. Our method introduces two novel components:Enhanced-Consistency Feature Distillation(ECFD) andRotation-Invariant Feature Distillation(RIFD). ECFD mitigates feature dilution by aligning consistent language representations from VLMs with cross-depth features, ensuring consistent small-rotated object representation across different depths. RIFD enhances rotation invariance by leveraging VLMs to distill robust rotational knowledge into detectors, aligning positive and negative language features with detector features to reduce sensitivity to orientation changes and mitigate class confusion. Without introducing additional computational overhead during inference, our method significantly improves the performance of remote sensing object detectors. Extensive experiments on public remote sensing datasets with complex scenes demonstrate the state-of-the-art results. Code is available at https://github.com/Shower-Lee9527/CRKD. Feiyi Li, Xiao Zhang 0050, Wenda Zhao 0003, You He 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Weakly Supervised Cross Mixer for Infrared and Visible Image FusionabstractRemote sensing infrared and visible image fusion aims to integrate information from multiple source images to enhance visual representation and support high-level visual tasks. However, due to the lack of ground truth supervision, most existing methods depend on predefined fusion-specific loss functions. Such manual definitions of cross-modal feature representations are often incomplete, potentially leading to information loss and reduced fusion performance. Complete information retention in fusion results can significantly benefit high-level tasks. Motivated by this, we propose a novel approach that utilizes weakly-supervised segmentation guidance for comprehensive information representation ability, eliminating the need for fusion-specific losses. Firstly, a self-supervised reconstruction model is proposed to obtain a decoder with robust cross-modal feature representation, which can adaptively reconstruct images based on cross-modal features. Then we design a weakly-supervised feature mining module to capture fused features with complete cross-modal information representation under the guidance of the segmentation task. The finial fusion result is adaptively reconstructed from the fused features by the cross-modal adaptive decoder without using any constraints. Extensive experiments demonstrate that the proposed method effectively mitigates information loss and achieves superior performance in both fusion and segmentation tasks compared to the state-of-the-art methods. Model and code are available at https://github.com/wangwenbo26/WSCM. Wenda Zhao 0003, You He 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Diverse Text-Prompt Generation for Remote Sensing Image ClassificationabstractInadequate remote sensing image training data usually makes remote sensing image classification models achieve low accuracy. Thus, we propose a diverse text-prompt generation learning (DPL) method. The context optimization (CoOp) model transfers the feature extraction capabilities of the CLIP model to downstream tasks with learnable text prompts. However, the small number of samples in remote sensing images can easily lead to noise. In contrast, DPL introduces a diverse text-prompt generation structure. Due to the diversity of multiple prompts, noise generated by inadequate samples is suppressed. Moreover, in order to keep the prompts diverse, we propose a prompt diversity loss. This loss pulls the prompts away from each other, which suppresses the noise generated by the limited training samples. Extensive experiments show that our method achieves superior performance than the existing methods on DOTA, HRRSD, and NWPU VHR-10 datasets. The model and code are available athttps://github.com/LvXiangzhu/DPL. Wenda Zhao 0003, Xiangzhu Lv, Ruikun He, Fan Zhao 0005, You He 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Cross-Domain Few-Shot Remote Sensing Object Classification via Triplet Relation-Aware MetricabstractIn real-world scenarios, peculiar remote sensing categories are difficult to collect on account of high cost and technical requirements. Moreover, there exists domain distribution gap among different datasets. Existing methods leverage inter-class and intra-class relations to enhance feature representation. Since remote images are shot from top to bottom, there is little difference between classes. Thus, such distance constraint only forms decision boundary between different classes. This paper proposes a triplet relation-aware metric for cross-domain few-shot remote sensing object classification, where the triplet relation-aware metric adjusts the distances among three kinds of inter-instance relations (i.e., same instance, same class and different class relations) to obtain a precise and effective feature representation. Especially, the distance of the same instance is regarded as a distance coordinate origin to guide distance metric learning. In this way, we constitute richer feature relations to promote representation learning in the source domain. Concretely, this procedure is optimized by the supervision of the designed relation-aware soft label based on the distance coordinate origin. Then, we align the triplet relation-aware metric between source domain and pseudo domain generated by the proposed episode style adversarial attack, thereby obtaining a domain-invariant feature representation. Extensive experiments on five widely-used remote sensing datasets demonstrate the superior performance of the proposed method compared with the state of the arts. Code is available at: https://github.com/jackhdpbl/TRAM. Ruikun He, Wenda Zhao 0003, You He 0002 |
IEEE Trans. Image Process. | 2 |
| 2024 | SRRT: Exploring Search Region Regulation for Visual Object TrackingabstractThe dominant trackers generate a fixed-size rectangular region based on the previous prediction or initial bounding box as the model input, i.e., search region. While this manner obtains promising tracking efficiency, a fixed-size search region lacks flexibility and is likely to fail in some cases, e.g., fast motion and distractor interference. Trackers tend to lose the target object due to the limited search region or experience interference from distractors due to the excessive search region. Drawing inspiration from the pattern humans track an object, we propose a novel tracking paradigm, called Search Region Regulation Tracking (SRRT) that applies a small eyereach when the target is captured and zooms out the search field when the target is about to be lost. SRRT applies a proposed search region regulator to estimate an optimal search region dynamically for each frame, by which the tracker can flexibly respond to transient changes in the location of object occurrences. To adapt the object’s appearance variation during online tracking, we further propose a locking-state determined updating strategy for reference frame updating. The proposed SRRT is concise without bells and whistles, yet achieves evident improvements and competitive results with other state-of-the-art trackers on eight benchmarks. On the large-scale LaSOT benchmark, SRRT improves SiamRPN++ and TransT with absolute gains of 4.6% and 3.1% in terms of AUC. The code and models will be released. Jiawen Zhu 0003, Xin Chen 0032, Xinying Wang 0005, Dong Wang 0004, Wenda Zhao 0003, Huchuan Lu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Text-Guided Diverse Image Synthesis for Long-Tailed Remote Sensing Object ClassificationabstractRemote sensing datasets pose long-tailed data distribution, and such unbalanced datasets will reduce the performance of existing remote sensing object classification models. Existing methods mainly rely on resampling datasets, modifying loss functions, data augmentation, and transfer learning to cope with such challenges. Unlike these, our study takes a novel perspective and focuses on mitigating the long-tailed distribution problem by generating a large number of tail-class images with consistency and diversity. Specifically, this article introduces a novel text-guided tail-class generation network (TGN). TGN comprises two main components: knowledge mutual distillation network (KMDN) and class-consistent diverse tail-class generation network (CDTG). KMDN resolves the isolation issue of the head and tail knowledge by facilitating mutual learning of feature representations between the head and tail data, thereby improving the feature extraction capability of the tail model. CDTG focuses on generating class-consistency diverse tail-class images that uses tail-class features extracted by KMDN. Especially, the class consistency is guaranteed by contrastive language-image pre-trainings (CLIP’s) powerful text-image alignment capability. These generated images are then added back into the original dataset to alleviate the long-tailed distribution, thereby improving the tail-class accuracy. Extensive experiments on the widely used DIOR, FGSC-23 and DOTA datasets demonstrate that the proposed method outperforms state-of-the-art methods. Dataset and code are publicly available athttps://github.com/XinR-Tang/TGN. Haojun Tang, Wenda Zhao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | High-Frequency Feature Transfer for Multispectral Image Super-ResolutionabstractLow-resolution characteristics of multispectral images restrict their usability. Various approaches (e.g., single-image super-resolution reconstruction (SISR) and pansharpening method) have been proposed to enrich the spatial details of low-resolution multispectral images (LRMSs) to obtain high-resolution ones. While the pansharpening method inevitably depends on panchromatic (PAN) images, which limits its application scenarios, SISR does not need auxiliary images, yet the blurred edge details within the reconstructed super-resolution multispectral (SRMS) images still remain a big challenge. In this work, we propose a novel high-frequency feature transfer (HFFT-PAN) method for multispectral image super-resolution to tackle the above drawbacks. Specifically, we first exploit the inherent low-frequency features among LRMS images to facilitate the extraction of high-frequency features from PAN images. After that, the high-frequency features from PAN images are transferred to the reconstruction procedure of multispectral images so that the SRMS images can not only benefit from the edge detail information from PAN images but also avoid the utilization of any PAN images during inference. Moreover, we employ an additional contrastive loss during training to ensure the fidelity of the generated SRMS images. Qualitative and quantitative evaluations exhibit that the proposed method performs favorably against state-of-the-art methods. The model and code are available athttps://github.com/wx0110wx/HFFT. Fan Zhao 0005, Wenda Zhao 0003, Zhepu Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Center-Wise Feature Consistency Learning for Long-Tailed Remote Sensing Object RecognitionabstractLong-tailed distribution of remote sensing data generally limits the object recognition performance of deep neural networks. We notice that too many samples from head class will induce the neural network to learn features of tail class samples being biased towards the head. To solve this, we propose a novel center-wise feature consistency learning (CFCL) mechanism for long-tailed remote sensing object recognition. Firstly, we implement a head-tail center feature generation procedure that builds two teacher models to extract the knowledge from the head class and tail class samples respectively, so as to avoid the extracted tail class features being affected by the head classes. Secondly, a center-wise feature consistency learning strategy is introduced, which distills the central feature of each class to a student model, thereby making the classification boundaries more prominent. Especially, the central feature is estimated by referring to the features which are correctly classified by the teacher models, thus the inaccurate knowledge is abandoned. Extensive experiments on widely-adopted remote sensing recognition datasets including FGSC-23, DIOR, xView and HRSC2016 demonstrate that our method achieves superior performance compared to the state-of-the-art approaches.Code and data are available at: https://github.com/wdzhao123/CWFC. Wenda Zhao 0003, Zhepu Zhang, Jiani Liu 0004, Yu Liu 0005, You He 0002, Huchuan Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Attacking Defocus Detection With Blur-Aware Transformation for Defocus DeblurringabstractPrevious fully-supervised defocus deblurring has made significant progress. However, training such deep models requires abundant paired ground truth, which is expensive and error-prone. This paper makes an attempt to train a defocus deblurring model without using paired ground truth and any other unpaired data. Related reblur-to-deblur schemes generally use physics-based reblur or GAN-based reblur, suffering from the robustness of blur kernel and hallucination generated by GAN. Besides, the domain gap between the realistic blurred image and reblurred image hinders deblurring performance. Addressing these challenges, we propose a weakly-supervised defocus deblurring framework via defocus detection attack. On one hand, we build a focused area detection attack (FADA) to enforce the focused area to reblur, thereby reversing its detection result by a pretrained defocus blur detection network. Moreover, we introduce a blur-aware transfer modulated from the defocused region to help FADA render a robust reblurred region. On the other hand, we implement a defocused region detection attack to guide the realistic blurred region to deblur in the process of training deblurring network with simulated-paired areas. Extensive experiments on three widely-used datasets verify the effectiveness of our framework. Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu |
IEEE Trans. Multim. | 1 |
| 2024 | Defocus Blur Detection Attack via Mutual-Referenced Feature TransferabstractBenefiting from deep learning, defocus blur detection (DBD) has made prominent progress. Existing DBD methods generally study multiscale and multilevel features to improve performance. In this article, from a different perspective, we explore to generate confrontational images to attack DBD network. Based on the observation that defocus area and focus region in an image can provide mutual feature reference to help improve the quality of the confrontational image, we propose a novel mutual-referenced attack framework. Firstly, we design a divide-and-conquer perturbation image generation model, where the focus region attack image and defocus area attack image are generated respectively. Then, we integrate mutual-referenced feature transfer (MRFT) models to improve attack performance. Comprehensive experiments are provided to verify the effectiveness of our method. Moreover, related applications of our study are presented, e.g., sample augmentation to improve DBD and paired sample generation to boost defocus deblurring. Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Interactive Feature Embedding for Infrared and Visible Image FusionabstractGeneral deep learning-based methods for infrared and visible image fusion rely on the unsupervised mechanism for vital information retention by utilizing elaborately designed loss functions. However, the unsupervised mechanism depends on a well-designed loss function, which cannot guarantee that all vital information of source images is sufficiently extracted. In this work, we propose a novel interactive feature embedding in a self-supervised learning framework for infrared and visible image fusion, attempting to overcome the issue of vital information degradation. With the help of a self-supervised learning framework, hierarchical representations of source images can be efficiently extracted. In particular, interactive feature embedding models are tactfully designed to build a bridge between self-supervised learning and infrared and visible image fusion learning, achieving vital information retention. Qualitative and quantitative evaluations exhibit that the proposed method performs favorably against state-of-the-art methods. Fan Zhao 0005, Wenda Zhao 0003, Huchuan Lu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Confusion Region Mining for Crowd CountingabstractExisting works mainly focus on crowd and ignore the confusion regions which contain extremely similar appearance to crowd in the background, while crowd counting needs to face these two sides at the same time. To address this issue, we propose a novel end-to-end trainable confusion region discriminating and erasing network called CDENet. Specifically, CDENet is composed of two modules of confusion region mining module (CRM) and guided erasing module (GEM). CRM consists of basic density estimation (BDE) network, confusion region aware bridge and confusion region discriminating network. The BDE network first generates a primary density map, and then the confusion region aware bridge excavates the confusion regions by comparing the primary prediction result with the ground-truth density map. Finally, the confusion region discriminating network learns the difference of feature representations in confusion regions and crowds. Furthermore, GEM gives the refined density map by erasing the confusion regions. We evaluate the proposed method on four crowd counting benchmarks, including ShanghaiTech Part_A, ShanghaiTech Part_B, UCF_CC_50, and UCF-QNRF, and our CDENet achieves superior performance compared with the state-of-the-arts. Jiawen Zhu 0003, Wenda Zhao 0003, Libo Yao, You He 0002, Maodi Hu, Huchuan Lu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Style-Content Metric Learning for Multidomain Remote Sensing Object RecognitionabstractPrevious remote sensing recognition approaches predominantly perform well on the training-testing dataset. However, due to large style discrepancies not only among multidomain datasets but also within a single domain, they suffer from obvious performance degradation when applied to unseen domains. In this paper, we propose a style-content metric learning framework to address the generalizable remote sensing object recognition issue. Specifically, we firstly design an inter-class dispersion metric to encourage the model to make decision based on content rather than the style, which is achieved by dispersing predictions generated from the contents of both positive sample and negative sample and the style of input image. Secondly, we propose an intra-class compactness metric to force the model to be less style-biased by compacting classifier's predictions from the content of input image and the styles of positive sample and negative sample. Lastly, we design an intra-class interaction metric to improve model's recognition accuracy by pulling in classifier's predictions obtained from the input image and positive sample. Extensive experiments on four datasets show that our style-content metric learning achieves superior generalization performance against the state-of-the-art competitors. Code and model are available at: https://github.com/wdzhao123/TSCM. Wenda Zhao 0003, Ruikai Yang, Yu Liu 0005, You He 0002 |
AAAI | 1 |
| 2023 | MetaFusion: Infrared and Visible Image Fusion via Meta-Feature Embedding from Object DetectionabstractFusing infrared and visible images can provide more texture details for subsequent object detection task. Conversely, detection task furnishes object semantic information to improve the infrared and visible image fusion. Thus, a joint fusion and detection learning to use their mutual promotion is attracting more attention. However, the feature gap between these two different-level tasks hinders the progress. Addressing this issue, this paper proposes an infrared and visible image fusion via meta-feature embedding from object detection. The core idea is that meta-feature embedding model is designed to generate object semantic features according to fusion network ability, and thus the semantic features are naturally compatible with fusion features. It is optimized by simulating a meta learning. Moreover, we further implement a mutual promotion learning between fusion and detection tasks to improve their performances. Comprehensive experiments on three public datasets demonstrate the effectiveness of our method. Code and model are available at: https://github.com/wdzhao123/MetaFusion. Wenda Zhao 0003, Shigeng Xie, Fan Zhao 0005, You He 0002, Huchuan Lu |
CVPR | 1 |
| 2023 | Frequency-Adaptive Learning for SAR Ship Detection in Clutter ScenesabstractConvolutional neural networks (CNNs) have been widely applied in the context of ship detection in synthetic aperture radar (SAR) images, but the detection performance is still not ideal in scenarios with clutter interference. Mining frequency-domain information to suppress the sea clutter in SAR ship detection has attracted wide attention. However, existing frequency-domain ship detection methods do not process frequency-domain information adaptively, which results in the degradation of ship detection performance. To overcome this problem, this article proposes a novel deep learning network called YOLO-FA. YOLO-FA contains the proposed frequency attention module (FAM), which can process frequency-domain information of SAR images adaptively. The proposed method can suppress the sea clutter in the SAR images with the help of frequency-domain information. We evaluate the proposed method YOLO-FA on two datasets, i.e., the high-resolution SAR images’ dataset (HRSID) and SAR ship detection dataset (SSDD). Compared with the baseline method YOLOv5 and the existing commonly used methods, YOLO-FA achieves state-of-the-art detection performance on both the datasets. Linping Zhang, Yu Liu 0005, Wenda Zhao 0003, Xueqian Wang 0002, Gang Li 0008, You He 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Weakly Correlated Distillation for Remote Sensing Object RecognitionabstractRemote sensing object labels require high specialization, resulting in a limited number of labeled samples. Without large labeled samples to support training, general remote sensing object recognition models have limited accuracy. Addressing this issue, this paper proposes a weakly correlated distillation learning framework for remote sensing object recognition with small number of samples. Benefitting from large-scale natural image datasets, many recognition models achieve superior feature extraction capabilities. Thus, we use them as backbones to build teacher models, and then fine-tune the teacher models with a small-scale remote sensing dataset. However, due to the limited number of remote sensing samples, the teacher models may produce noisy features that reduce the performance of the student model. Therefore, we propose a weakly correlated distillation method that selects the weakly correlated features from teacher models to distill the student. Since the weakly correlated features contain different noise distributions which can be mutually suppressed, thereby improving the performance of the student. Extensive experiments on three widely-used datasets of DOTA, HRRSD and NWPU VHR-10 demonstrate the superior performance of our method compared with the state of the arts. Code is available at: https://github.com/wdzhao123/WCD. Wenda Zhao 0003, Xiangzhu Lv, Yu Liu 0005, You He 0002, Huchuan Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Nowhere to Disguise: Spot Camouflaged Objects via Saliency Attribute TransferabstractBoth salient object detection (SOD) and camouflaged object detection (COD) are typical object segmentation tasks. They are intuitively contradictory, but are intrinsically related. In this paper, we explore the relationship between SOD and COD, and then borrow successful SOD models to detect camouflaged objects to save the design cost of COD models. The core insight is that both SOD and COD leverage two aspects of information: object semantic representations for distinguishing object and background, and context attributes that decide object category. Specifically, we start by decoupling context attributes and object semantic representations from both SOD and COD datasets through designing a novel decoupling framework with triple measure constraints. Then, we transfer saliency context attributes to the camouflaged images through introducing an attribute transfer network. The generated weakly camouflaged images can bridge the context attribute gap between SOD and COD, thereby improving the SOD models' performances on COD datasets. Comprehensive experiments on three widely-used COD datasets verify the ability of the proposed method. Code and model are available at: https://github.com/wdzhao123/SAT. Wenda Zhao 0003, Shigeng Xie, Fan Zhao 0005, You He 0002, Huchuan Lu |
IEEE Trans. Image Process. | 1 |
| 2023 | Full-Scene Defocus Blur Detection With DeFBD+ via Multi-Level Distillation LearningabstractExisting defocus blur detection (DBD) methods generally perform well on a single type of unfocused blur scene (e.g., foreground focus), thereby suffering from the performance degradation for the other types of unfocused blur scenes. In this paper, we present the first exploration on full-scene DBD, and propose a separate-and-combine framework to achieve excellent performance for diverse defocus blur scenes. We firstly structure full-scene DBD dataset (named as DeFBD+) through collecting more types of unfocused blur scenes (e.g., background focus, full focus and full out of focus) with pixel-level annotations. Then, to avoid performance degradation caused by mutual interference from local feature representation and global content perception, we implement a pixel-level DBD network and an image-level DBD classification network to learn these two abilities separately. After that, we propose an isomeric distillation mechanism to combine these two abilities. Extensive experiments show that the proposed approach achieves superior performance compared with state-of-the-art methods. Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu |
IEEE Trans. Multim. | 1 |
| 2023 | Depth-Distilled Multi-Focus Image FusionabstractHomogeneous regions, which are smooth areas that lack blur clues to discriminate if they are focused or non-focused. Therefore, they bring a great challenge to achieve high accurate multi-focus image fusion (MFIF). Fortunately, we observe that depth maps are highly related to focus and defocus, containing a preponderance of discriminative power to locate homogeneous regions. This offers the potential to provide additional depth cues to assist MFIF task. Taking depth cues into consideration, in this paper, we propose a new depth-distilled multi-focus image fusion framework, namely D2MFIF. In D2MFIF, depth-distilled model (DDM) is designed for adaptively transferring the depth knowledge into MFIF task, gradually improving MFIF performance. Moreover, multi-level fusion mechanism is designed to integrate multi-level decision maps from intermediate outputs for improving the final prediction. Visually and quantitatively experimental results demonstrate the superiority of our method over several state-of-the-art methods. Fan Zhao 0005, Wenda Zhao 0003, Huimin Lu 0001, Yong Liu 0017, Libo Yao, Yu Liu 0005 |
IEEE Trans. Multim. | 2 |
| 2022 | United Defocus Blur Detection and Deblurring via Adversarial Promoting Learning
Wenda Zhao 0003, Fei Wei, You He 0002, Huchuan Lu |
ECCV (30) | 1 |
| 2022 | Image-Scale-Symmetric Cooperative Network for Defocus Blur DetectionabstractDefocus blur detection (DBD) for natural images is a challenging vision task especially in the presence of homogeneous regions and gradual boundaries. In this paper, we propose a novel image-scale-symmetric cooperative network (IS2CNet) for DBD. On one hand, in the process of image scales from large to small, IS2CNet gradually spreads the recept of image content. Thus, the homogeneous region detection map can be optimized gradually. On the other hand, in the process of image scales from small to large, IS2CNet gradually feels the high-resolution image content, thereby gradually refining transition region detection. In addition, we propose a hierarchical feature integration and bi-directional delivering mechanism to transfer the hierarchical feature of previous image scale network to the input and tail of the current image scale network for guiding the current image scale network to better learn the residual. The proposed approach achieves state-of-the-art performance on existing datasets.Codes and results are available at:https://github.com/wdzhao123/IS2CNet. Fan Zhao 0006, Huimin Lu 0001, Wenda Zhao 0003, Libo Yao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Generalizable Crowd Counting via Diverse Context Style LearningabstractExisting crowd counting approaches predominantly perform well on the training-testing protocol. However, due to large style discrepancies not only among images but also within a single image, they suffer from obvious performance degradation when applied to unseen domains. In this paper, we aim to design a generalizable crowd counting framework which is trained on a source domain but can generalize well on the other domains. To reach this, we propose a gated ensemble learning framework. Specifically, we first propose a diverse fine-grained style attention model to help learn discriminative content feature representations, allowing for exploiting diverse features to improve generalization. We then introduce a channel-level binary gating ensemble model, where diverse feature prior, input-dependent guidance and density grade classification constraint are implemented, to optimally select diverse content features to participate in the ensemble, taking advantage of their complementary while avoiding redundancy. Extensive experiments show that our gating ensemble approach achieves superior generalization performance among four public datasets. Codes are publicly available athttps://github.com/wdzhao123/DCSL. Wenda Zhao 0003, Yu Liu 0005, Huimin Lu 0001, Cong'an Xu, Libo Yao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Teaching Teachers First and Then Student: Hierarchical Distillation to Improve Long-Tailed Object Recognition in Aerial ImagesabstractRemote sensing data distribution generally exposes the long-tail characteristic. This will limit the object recognition performance of existing deep models when they are trained with such unbalanced data. In this paper, we propose a novel hierarchical distillation framework to address the long-tailed object recognition in aerial images. Firstly, we notice that not only student model should learn feature representations from teachers, but also teacher models should learn feature representations from each other. Therefore, we build hierarchical teacher-wise distillation to improve the feature representations of the teacher models trained with middle and tail data, which is achieved by distilling the feature representations of the teacher model trained with head data. Secondly, we notice that the feature representations of the middle and tail classes can not be effectively distilled from the teacher to the student, since too little middle and tail data can be used to learn. Thus, we propose self-calibrated sampling learning that enforces the student to strengthen the learning of the middle and tail data, thereby improving the student’ feature learning ability. Extensive experiments on two widely-used DOTA and FGSC-23 datasets demonstrate superior performance of the proposed method compared with state-of-the-art methods. Model and code are publicly available at: https://github.com/wdzhao123/T2FTS. Wenda Zhao 0003, Jiani Liu 0004, Yu Liu 0005, Fan Zhao 0005, You He 0002, Huchuan Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Diversity Consistency Learning for Remote-Sensing Object Recognition With Limited LabelsabstractAnnotating remote sensing object recognition needs high professionalism, and thus limited labeled samples are available. Suffering from this, general remote sensing object recognition methods are facing low recognition accuracy. Addressing this issue, this paper proposes a diversity consistency learning for remote sensing object recognition with limited labels. Specifically, diversity generation model is designed as a teacher model to generate diverse results, which is trained with labeled samples. Then, round consistency distillation model is introduced to distill the knowledge of diverse pseudo labels to a student network, which is trained with unlabeled samples. Especially, diverse pseudo labels are generated by the well-trained diversity generation model, which can improve recognition accuracy since diverse pseudo label errors can cancel each other out. Extensive experiments on two widely-used datasets of FS23 and HRSC2016 demonstrate the superior performance of our method compared with the state of the arts. Wenda Zhao 0003, Tingting Tong, Fan Zhao 0005, You He 0002, Huchuan Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Feature Balance for Fine-Grained Object Classification in Aerial ImagesabstractFine-grained object classification (FGOC) focuses on identifying subcategories of objects, which is crucial in military and civilian. Existing FGOC methods primarily focus on high-resolution aerial images, limiting their application on low-resolution (LR) FGOC that is a more realistic setting, especially on resource-constrained satellite devices. It is more challenging to deal with LR FGOC since objects’ details are blurred or missing. Addressing this issue, we make the first attempt to explore LR FGOC and propose a novel pipeline based on two technical insights: 1) feature balance strategy discriminatively integrates super-resolution weak and strong detailed presentations into coarse features of LR aerial images, achieving a feature balance to avoid that the weak detailed presentations are inhibited by the strong ones and 2) iterative interaction mechanism alternately refines feature details of the discriminative ship regions and optimizes the performance of FGOC. Moreover, we build a low-resolution fine-grained object (LFS) dataset to promote further study and evaluation. Extensive experiments on the proposed LFS dataset and the other three object datasets of DOTA, FS23, and HRSC2016 demonstrate that our method outperforms state-of-the-art algorithms. Dataset and code are publicly available athttps://github.com/wdzhao123/FBNet. Wenda Zhao 0003, Tingting Tong, Libo Yao, Yu Liu 0005, Cong'an Xu, You He 0002, Huchuan Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Self-Generated Defocus Blur Detection via Dual Adversarial DiscriminatorsabstractAlthough existing fully-supervised defocus blur detection (DBD) models significantly improve performance, training such deep models requires abundant pixel-level manual annotation, which is highly time-consuming and error-prone. Addressing this issue, this paper makes an effort to train a deep DBD model without using any pixel-level annotation. The core insight is that a defocus blur region/focused clear area can be arbitrarily pasted to a given realistic full blurred image/full clear image without affecting the judgment of the full blurred image/full clear image. Specifically, we train a generator G in an adversarial manner against dual discriminators Dcand Db. G learns to produce a DBD mask that generates a composite clear image and a composite blurred image through copying the focused area and unfocused region from corresponding source image to another full clear image and full blurred image. Then, Dcand Dbcan not distinguish them from realistic full clear image and full blurred image simultaneously, achieving a self-generated DBD by an implicit manner to define what a defocus blur area is. Besides, we propose a bilateral triplet-excavating constraint to avoid the degenerate problem caused by the case one discriminator defeats the other one. Comprehensive experiments on two widely-used DBD datasets demonstrate the superiority of the proposed approach. Source codes are available at: https://github.com/shangcai1/SG. Wenda Zhao 0003, Cai Shang, Huchuan Lu |
CVPR | 1 |
| 2021 | Defocus Blur Detection via Boosting Diversity of Deep Ensemble NetworksabstractExisting defocus blur detection (DBD) methods usually explore multi-scale and multi-level features to improve performance. However, defocus blur regions normally have incomplete semantic information, which will reduce DBD's performance if it can't be used properly. In this paper, we address the above problem by exploring deep ensemble networks, where we boost diversity of defocus blur detectors to force the network to generate diverse results that some rely more on high-level semantic information while some ones rely more on low-level information. Then, diverse result ensemble makes detection errors cancel out each other. Specifically, we propose two deep ensemble networks (e.g., adaptive ensemble network (AENet) and encoder-feature ensemble network (EFENet)), which focus on boosting diversity while costing less computation. AENet constructs different light-weight sequential adapters for one backbone network to generate diverse results without introducing too many parameters and computation. AENet is optimized only by the self- negative correlation loss. On the other hand, we propose EFENet by exploring the diversity of multiple encoded features and ensemble strategies of features (e.g., group-channel uniformly weighted average ensemble and self-gate weighted ensemble). Diversity is represented by encoded features with less parameters, and a simple mean squared error loss can achieve the superior performance. Experimental results demonstrate the superiority over the state-of-the-arts in terms of accuracy and speed. Codes and models are available at: https://github.com/wdzhao123/DENets. Wenda Zhao 0003, Xueqing Hou, You He 0002, Huchuan Lu |
IEEE Trans. Image Process. | 1 |
| 2021 | Learning Specific and General Realm Feature Representations for Image FusionabstractA universal fusion framework for handling multi-realm image fusion reduces the cost of manual selection in varied applications. Addressing the generality of multiple realms and the sensitivity of specific realm, we propose a novel universal framework for multi-realm image fusion through learning realm-specific and realm-general feature representations. Shared principle network, adaptive realm feature extraction strategy and realm activation mechanism are designed for facilitating high generalization of across-realm and sensitivity of specific-realm simultaneously. In addition, we present realm-specific no-reference perceptual metric losses based on the edge details and contrast for optimizing the learning process, making the fused image exhibit more specific appearance. Moreover, we collect a new multi-realm image fusion dataset (MRIF), consisting of infrared and visual images, medical images and multispectral images, to facilitate our training and testing. Experimental results show that the fused image obtained by the proposed method achieves superior performance compared with the state-of-the-art methods on MRIF and the other three datasets including infrared and visual images, medical images and remote sensing images, respectively. Fan Zhao 0006, Wenda Zhao 0003 |
IEEE Trans. Multim. | 2 |
| 2020 | Defocus Blur Detection via Multi-Stream Bottom-Top-Bottom NetworkabstractDefocus blur detection (DBD) is aimed to estimate the probability of each pixel being in-focus or out-of-focus. This process has been paid considerable attention due to its remarkable potential applications. Accurate differentiation of homogeneous regions and detection of low-contrast focal regions, as well as suppression of background clutter, are challenges associated with DBD. To address these issues, we propose a multi-stream bottom-top-bottom fully convolutional network (BTBNet), which is the first attempt to develop an end-to-end deep network to solve the DBD problems. First, we develop a fully convolutional BTBNet to gradually integrate nearby feature levels of bottom to top and top to bottom. Then, considering that the degree of defocus blur is sensitive to scales, we propose multi-stream BTBNets that handle input images with different scales to improve the performance of DBD. Finally, a cascaded DBD map residual learning architecture is designed to gradually restore finer structures from the small scale to the large scale. To promote further study and evaluation of the DBD models, we construct a new database of 1100 challenging images and their pixel-wise defocus blur annotations. Experimental results on the existing and our new datasets demonstrate that the proposed method achieves significantly better performance than other state-of-the-art algorithms. Wenda Zhao 0003, Fan Zhao 0006, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Towards Weakly-Supervised Focus Region Detection via Recurrent Constraint NetworkabstractRecent state-of-the-art methods on focus region detection (FRD) rely on deep convolutional networks trained with costly pixel-level annotations. In this study, we propose a FRD method that achieves competitive accuracies but only uses easily obtained bounding box annotations. Box-level tags provide important cues of focus regions but lose the boundary delineation of the transition area. A recurrent constraint network (RCN) is introduced for this challenge. In our static training, RCN is jointly trained with a fully convolutional network (FCN) through box-level supervision. The RCN can generate a detailed focus map to locate the boundary of the transition area effectively. In our dynamic training, we iterate between fine-tuning FCN and RCN with the generated pixel-level tags and generate finer new pixel-level tags. To boost the performance further, a guided conditional random field is developed to improve the quality of the generated pixel-level tags. To promote further study of the weakly supervised FRD methods, we construct a new dataset called FocusBox, which consists of 5000 challenging images with bounding box-level labels. Experimental results on existing datasets demonstrate that our method not only yields comparable results than fully supervised counterparts but also achieves a faster speed. Wenda Zhao 0003, Xueqing Hou, You He 0002, Huchuan Lu |
IEEE Trans. Image Process. | 1 |
| 2019 | Enhancing Diversity of Defocus Blur Detectors via Cross-Ensemble NetworkabstractDefocus blur detection (DBD) is a fundamental yet challenging topic, since the homogeneous region is obscure and the transition from the focused area to the unfocused region is gradual. Recent DBD methods make progress through exploring deeper or wider networks with the expense of high memory and computation. In this paper, we propose a novel learning strategy by breaking DBD problem into multiple smaller defocus blur detectors and thus estimate errors can cancel out each other. Our focus is the diversity enhancement via cross-ensemble network. Specifically, we design an end-to-end network composed of two logical parts: feature extractor network (FENet) and defocus blur detector cross-ensemble network (DBD-CENet). FENet is constructed to extract low-level features. Then the features are fed into DBD-CENet containing two parallel-branches for learning two groups of defocus blur detectors. For each individual, we design cross-negative and self-negative correlations and an error function to enhance ensemble diversity and balance individual accuracy. Finally, the multiple defocus blur detectors are combined with a uniformly weighted average to obtain the final DBD map. Experimental results indicate the superiority of our method in terms of accuracy and speed when compared with several state-of-the-art methods. Wenda Zhao 0003, Qiuhua Lin, Huchuan Lu |
CVPR | 1 |
| 2019 | Multi-Focus Image Fusion With a Natural Enhancement via a Joint Multi-Level Deeply Supervised Convolutional Neural NetworkabstractCommon non-focused areas are often present in multi-focus images due to the limitation of the number of focused images. This factor severely degrades the fusion quality of multi-focus images. To address this problem, we propose a novel end-to-end multi-focus image fusion with a natural enhancement method based on deep convolutional neural network (CNN). Several end-to-end CNN architectures that are specifically adapted to this task are first designed and researched. On the basis of the observation that low-level feature extraction can capture low-frequency content, whereas high-level feature extraction effectively captures high-frequency details, we further combine multi-level outputs such that the most visually distinctive features can be extracted, fused, and enhanced. In addition, the multi-level outputs are simultaneously supervised during training to boost the performance of image fusion and enhancement. Extensive experiments show that the proposed method can deliver superior fusion and enhancement performance than the state-of-the-art methods in the presence of multi-focus images with common non-focused areas, anisotropic blur, and misregistration. Wenda Zhao 0003, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Defocus Blur Detection via Multi-Stream Bottom-Top-Bottom Fully Convolutional NetworkabstractDefocus blur detection (DBD) is the separation of in-focus and out-of-focus regions in an image. This process has been paid considerable attention because of its remarkable potential applications. Accurate differentiation of homogeneous regions and detection of low-contrast focal regions, as well as suppression of background clutter, are challenges associated with DBD. To address these issues, we propose a multi-stream bottom-top-bottom fully convolutional network (BTBNet), which is the first attempt to develop an end-to-end deep network for DBD. First, we develop a fully convolutional BTBNet to integrate low-level cues and high-level semantic information. Then, considering that the degree of defocus blur is sensitive to scales, we propose multi-stream BTBNets that handle input images with different scales to improve the performance of DBD. Finally, we design a fusion and recurrent reconstruction network to recurrently refine the preceding blur detection maps. To promote further study and evaluation of the DBD models, we construct a new database of 500 challenging images and their pixel-wise defocus blur annotations. Experimental results on the existing and our new datasets demonstrate that the proposed method achieves significantly better performance than other state-of-the-art algorithms. Wenda Zhao 0003, Fan Zhao 0006, Dong Wang 0004, Huchuan Lu |
CVPR | 1 |
| 2018 | Multisensor Image Fusion and Enhancement in Spectral Total Variation DomainabstractMost existing image fusion methods assume that at least one input image contains high-quality information at any place of an observed scene. Thus, these fusion methods will fail if every input image is degraded. To address this issue, this study proposes a novel fusion framework that integrates image fusion based on spectral total variation (TV) method and image enhancement. For spatially varying multiscale decompositions generated by the spectral TV framework, this study verifies that the decomposition components can be modeled efficiently by tailed α-stable-based random variable distribution (TRD) rather than the commonly used Gaussian distribution. Consequently, salience and match measures based on TRD are proposed to fuse each sub-band decomposition. The spatial intensity information is also adopted to fuse the remainder of the image decomposition components. A sub-band adaptive gain function family based on TV spectrum and space variation is constructed for fused multiscale decompositions to enhance fused image simultaneously. Finally, numerous experiments with various multisensor image pairs are conducted to evaluate the proposed method. Experimental results show that even if the input images are degraded, the fused image obtained by the proposed method achieves significant improvement in terms of edge details and contrast while extracting the main features of the input images, thereby achieving better performance compared with the state-of-the-art methods. Wenda Zhao 0003, Huimin Lu 0001, Dong Wang 0004 |
IEEE Trans. Multim. | 1 |