Wenxu Shi

dblp:321/3392 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 LIIA -Net: A lightweight illumination iterative adjustment network for low-light image enhancement
Chengwan You, Wenxu Shi, Guibin Hu, Bochuan Zheng
Neural Networks2
2025 Enhancing domain adaptation for plant diseases detection through Masked Image Consistency in Multi-Granularity Alignment
Guinan Guo, Songning Lai, Qingyang Wu, Yuntao Shou, Wenxu Shi
Expert Syst. Appl.5
2025 Instance-Level Multitask Learning for 3-D Building Extraction From Monocular Off-Nadir Satellite Sensor Imagery
abstract
Extracting 3D building information from monocular satellite sensor imagery remains a formidable challenge in the field of remote sensing. Multitask frameworks based on deep learning, which particularly for simultaneously predicting 2D building outlines and their respective heights using ortho-rectified satellite imagery, have shown promise in addressing this challenge. Height estimation is notably complex due to the absence of explicit height indicators, limited interaction between semantic-height features, and inadequate representation of building relationships. Moreover, the availability of data sources is a limiting factor for broader application. To overcome these issues, this study introduces an innovative instance-level multitask learning model (named BDH-Net) that leverages off-nadir perspectives and roof-to-footprint offset vectors to enhance modeling. This model comprises four key components: a pixel-wise feature extraction image encoder-decoder, a query transformer decoder, a multitask decoder, and a height decoder that employs intra-instance and inter-instance attention for precise building height estimation. Additionally, we pioneer the use of Google Earth imagery to construct an off-nadir satellite dataset with roof-to-footprint offset vectors specifically designed for building instance segmentation and height prediction, known as the BDH dataset. Comprehensive experiments demonstrate that the proposed BDH-Net significantly improves the accuracy of monocular 3D building data extraction by integrating roof-to-footprint offset vectors and leveraging context specific to each building instance. With the extensive coverage and regular updates of Google Earth imagery, BDH-Net holds substantial potential for wide-ranging and long-term applications. The source code of the proposed BDH-Net and the BDH dataset are publicly available at https://github.com/wishx98/BDHNet.
Wenxu Shi, Qingyan Meng, Linlin Zhang 0007, Maofan Zhao, Guinan Guo, Peter M. Atkinson
IEEE Trans. Geosci. Remote. Sens.1
2024 Building Height Extraction from Monocular Off-Nadir Satellite Sensor Imagery
abstract
The extraction of building heights from monocular satellite sensor imagery poses a significant difficulty in remote sensing. Deep learning-based multi-task frameworks have recently emerged, showing potential in concurrently predicting 2D building shapes and heights from ortho-rectified satellite images. These techniques, however, have several limitations, including a lack of direct height markers, inherent uncertainties in height estimation, a heavy reliance on limited digital surface models, suboptimal integration of semantic-height feature, and inadequate depiction of the spatial relationships between buildings. To address these challenges, we develop a novel instance-level multi-task learning model that uses off-nadir imaging angles, roof-to-footprint as a proxy for height, and integrates certain geometric properties to increase height estimation accuracy. This method was validated using a self-constructed dataset, demonstrating significantly superior performance relative to two benchmarks.
Wenxu Shi, Qingyan Meng, Jian Wang 0138, Tingyuan Zhou, Peter M. Atkinson
IGARSS1
2024 Unsupervised Multimodal Change Detection by Distilling Common and Discrepant Representations
abstract
Change detection (CD) has become increasingly important in remote sensing and Earth observation. Currently, the data in various modalities has rapidly increased, and multimodal change detection is gaining prominence and holds substantial potential for applications demanding high temporal frequency or rapid response. In this research, we proposed a novel CDR-Net architecture for unsupervised multimodal change detection. The CDR-Net fuses the merits of the mask reconstruction and contrastive learning self-supervised paradigm. Within this architecture, the mask reconstruction subnetwcork pays more attention to low-level details, distilling common representations between multimodal remote sensing images to make them comparable, while the CL subnetwork emphasizes high-level semantics, extracting discrepant representations to facilitate the change detection task. Experimental results demonstrated that the CDR-Net achieved outstanding performance. This implies that the CDR-Net is of great value for resource investigation, emergency response and time-series understanding.
Jian Wang 0138, Li Yan 0003, Hong Xie 0002, Tingyuan Zhou, Wenxu Shi, Peter M. Atkinson
IGARSS5
2024 Conflict-Alleviated Gradient Descent for Adaptive Object Detection
Wenxu Shi, Bochuan Zheng
IJCAI1
2024 Alleviating the Equilibrium Challenge with Sample Virtual Labeling for Adversarial Domain Adaptation
abstract
Many domain adaptive object detection (DAOD) methods employ domain adversarial training to align features and mitigate the domain gap. In this approach, a feature extractor is trained to deceive a domain classifier, thereby aligning feature distributions. However, the domain classifier's discrimination capability can easily fall into a local optimum due to the equilibrium challenge, hindering the effective training of the feature extractor. In this work, we propose an efficient optimization strategy called Virtual-label Fooled Domain Discrimination (VFDD), which revitalizes the domain classifier during training using virtual domain labels. Such virtual label makes the separable distributions less separable, and thus leads to a more easily confused domain classifier, which in turn further drives feature alignment. Particularly, we introduce a novel concept of virtual domain label for the unaligned samples and propose the VirtualH -divergence to overcome the problem of falling into local optimum due to the equilibrium challenge. VFDD is orthogonal to most existing DAOD methods and can be integrated as a plug-and-play module to enhance these models. Theoretical insights and experimental analyses demonstrate that VFDD improves many popular baselines and surpasses recent unsupervised DAOD models.
Wenxu Shi, Bochuan Zheng
ACM Multimedia1
2024 DIFNet: Dual-Domain Information Fusion Network for Image Denoising
Zedong Wu, Wenxu Shi, Liming Xu, Zicheng Ding, Bochuan Zheng
PRCV (8)2
2024 A dynamically class-wise weighting mechanism for unsupervised cross-domain object detection under universal scenarios
Wenxu Shi, Dailun Tan, Bochuan Zheng
Knowl. Based Syst.1
2024 Confused and disentangled distribution alignment for unsupervised universal adaptive object detection
Wenxu Shi, Zedong Wu, Bochuan Zheng
Knowl. Based Syst.1
2023 Exploring Implicit Domain-Invariant Features for Domain Adaptive Object Detection
abstract
Recent researches have made a great progress in domain adaptive object detectors. These detectors aim to learn explicit domain-invariant features by adversarially mitigating domain divergence and simultaneously optimizing source risks. However, an inherent problem is that they ignore the informative knowledge implied in domain-specific features, which is recognized as implicit domain-invariant feature. This is mainly caused by the multimode structure underlying target distribution, characterized by various scales and categories of objects in target images. To solve that, we propose the Implicit Domain-invariant Faster R-CNN (IDF) by using non-adversarial domain discriminator, dual attention mechanism and selective feature perception. This idea is implemented on the Faster R-CNN backbone, but with an improved architecture of two branches, i.e. domain-invariant branch and domain-specific branch. The former can clearly learn explicit domain adaptive features w.r.t. easy samples, while the latter aims to learn implicit domain-invariant features w.r.t. hard samples. Experiments on numerous benchmark datasets, including the Cityscapes, Foggy Cityscapes, KITTI and SIM10K, show the superiority of our IDF over other state-of-the-art domain adaptive object detectors. The demo code is released inhttps://github.com/sea123321/IDF.
Qinghai Lang, Lei Zhang 0038, Wenxu Shi, Weijie Chen 0006, Shiliang Pu
IEEE Trans. Circuits Syst. Video Technol.3
2023 PanDiff: A Novel Pansharpening Method Based on Denoising Diffusion Probabilistic Model
abstract
Pansharpening is a crucial image processing technique for numerous remote sensing downstream tasks, aiming to recover high spatial resolution multispectral (HRMS) images by fusing high spatial resolution panchromatic (PAN) images and low spatial resolution multispectral (LRMS) images. Most current mainstream pansharpening fusion frameworks directly learn the mapping relationships from PAN and LRMS images to HRMS images by extracting key features. However, we propose a novel pansharpening method based on the denoising diffusion probabilistic model (DDPM) called PanDiff, which learns the data distribution of the difference maps (DM) between HRMS and interpolated MS (IMS) images from a new perspective. Specifically, PanDiff decomposes the complex fusion process of PAN and LRMS images into a multi-step Markov process, and the U-Net is employed to reconstruct each step of the process from random Gaussian noise. Notably, the PAN and LRMS images serve as the injected conditions to guide the U-Net in PanDiff, rather than being the fusion objects as in other pansharpening methods. Furthermore, we propose a modal intercalibration module (MIM) to enhance the guidance effect of the PAN and LRMS images. The experiments are conducted on a freely available benchmark dataset, including GaoFen-2, QuickBird, and WorldView-3 images. The experimental results from the fusion and generalization tests effectively demonstrate the outstanding fusion performance and high robustness of PanDiff. Fig. 1 depicts the results of the proposed method performed on various scenes. Additionally, the ablation experiments confirm the rationale behind PanDiff’s construction.
Qingyan Meng, Wenxu Shi, Linlin Zhang 0007
IEEE Trans. Geosci. Remote. Sens.2
2022 Universal Domain Adaptive Object Detector
abstract
Universal domain adaptive object detection (UniDAOD) is more challenging than domain adaptive object detection (DAOD) since the label space of the source domain may not be the same as that of the target and the scale of objects in the universal scenarios can vary dramatically (i.e, category shift and scale shift). To this end, we propose US-DAF, namely Universal Scale-Aware Domain Adaptive Faster RCNN with Multi-Label Learning, to reduce the negative transfer effect during training while maximizing transferability as well as discriminability in both domains under a variety of scales. Specifically, our method is implemented by two modules: 1) We facilitate the feature alignment of common classes and suppress the interference of private classes by designing a Filter Mechanism module to overcome the negative transfer caused by category shift. 2) We fill the blank of scale-aware adaptation in object detection by introducing a new Multi-Label Scale-Aware Adapter to perform individual alignment between corresponding scale for two domains. Experiments show that US-DAF achieves state-of-the-art results on three scenarios (\emphi.e, Open-Set, Partial-Set, and Closed-Set) and yields 7.1% and 5.9% relative improvement on benchmark datasets Clipart1k and Watercolor in particular.
Wenxu Shi, Lei Zhang 0038, Weijie Chen 0006, Shiliang Pu
ACM Multimedia1
2022 Multilayer Feature Fusion Network With Spatial Attention and Gated Mechanism for Remote Sensing Scene Classification
abstract
Remote sensing (RS) scene classification has attracted extensive attention due to its large number of applications. Recently, convolutional neural networks (CNNs) methods have shown impressive ability of feature learning in RS scene classification. However, the performance is still limited by large-scale variance and complex background. To address these problems, we present a multilayer feature fusion network with spatial attention and gated mechanism (MLF2Net_SAGM) for RS scene classification. At first, the backbone is employed to extract multilayer convolutional features. Then, a residual spatial attention module (RSAM) is proposed to enhance discriminative regions of the multilayer feature maps, and key areas can be harvested. Finally, the multilayer spatial calibration features are fused to form the final feature map, and a gated fusion module (GFM) is designed to eliminate feature redundancy and mutual exclusion (FRME). To verify the effectiveness of the proposed method, we conduct comparative experiments based on three widely used RS image scene classification benchmarks. The results show that the direct fusion of multilayer features via element-wise addition leads to FRME, whereas our method fuses multilayer features more effectively and improves the performance of scene classification.
Qingyan Meng, Maofan Zhao, Linlin Zhang 0007, Wenxu Shi, Lorenzo Bruzzone
IEEE Geosci. Remote. Sens. Lett.4