Yu Wang 0140

dblp:02/5889-140 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0003-3436-4251ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 From forgotten to pan-sharpening
Jiaming Wang 0001, Yansong Lin, Chuanxi Chen, Xiao Huang 0003, Ruiqian Zhang, Yu Wang 0140, Tao Lu 0001
Pattern Recognit.6
2026 Land-cover prior diffusion probabilistic model for remote sensing image super resolution
Zhizheng Zhang 0009, Jiayi Ma 0001, Jindou Zhang, Yu Wang 0140, Zhenghao Liao, Gui Cheng, Mingqiang Guo, Liang Wu 0005
Pattern Recognit.5
2026 NSBRNet: Non-Local Spatio-Temporal Bidirectional Recurrent Network for Satellite Video Super-Resolution
abstract
In recent years, intelligent processing of satellite videos has emerged as a significant research focus within the field of remote sensing, driven by the growing demand for enhanced spatial resolution. This need has led to increased interest in satellite video super-resolution (SVSR) algorithms, which aim to improve the quality of satellite imagery. However, many existing SVSR methods tend to neglect the global dependencies among frames in satellite videos, resulting in an incomplete utilization of spatio-temporal feature information. To tackle this issue, we propose a novel non-local spatio-temporal bidirectional recurrent network specifically designed for SVSR applications. Our approach employs a gate-guided deformable alignment module that effectively enhances feature alignment and fusion using a dynamic gating mechanism. This allows the network to adaptively focus on relevant features during the reconstruction process. Furthermore, we introduce a non-local spatio-temporal fusion module that integrates both temporal and spatial relationships over long sequences of frames, ensuring a comprehensive extraction of feature information. Through extensive experiments, our proposed method demonstrates superior performance compared to state-of-the-art SVSR techniques in terms of reconstruction quality. Additionally, it demonstrates outstanding performance in downstream satellite video applications, showcasing its potential in satellite video processing tasks. The source code is publicly available at https://github.com/Yu-Wang-0801/NSBRNet.
Yu Wang 0140, Xiaolong Zuo, Tao Lu 0001, Jiaming Wang 0001, Yuankun Wang, Siyuan Wang 0011, Zhizheng Zhang 0009, Xiaojin Zhao
IEEE Trans. Circuits Syst. Video Technol.1
2025 Lightweight remote sensing super-resolution with multi-scale graph attention network
Yu Wang 0140, Tao Lu 0001, Xiao Huang 0003, Jiaming Wang 0001, Zhizheng Zhang 0009, Xiaolong Zuo
Pattern Recognit.1
2025 CANet: A Spatial Structure Constraint and Local Semantic Awareness Based Network for Weakly Supervised Building Extraction
abstract
Benefitting from the easy availability of image-level labels, weakly supervised semantic segmentation (WSSS) methods based on class activation maps (CAMs) have made significant progress in building extraction from remote sensing imagery. However, image-level labels lack precise spatial locations and boundary ranges of buildings, posing challenges in achieving comprehensive and structurally clear building extraction. Furthermore, due to the complex background interference and the diversity of building in high-resolution remote sensing imagery (HRRS), small and sparse buildings suffer from insufficient attention in CAMs. To solve the above problems, this article proposes a spatial structure constraint and local semantic awareness-based WSSS method, CANet, for extracting buildings from HRRS. Specifically, we design a spatial structure constraint module to generate CAMs with finer spatial structural details of buildings, which minimizes feature differences from patches of different granularities and the whole image. Moreover, a local semantic awareness module is designed to address the issue of insufficient coverage of CAMs on sparse and tiny building. This module first strengthens the feature extraction network by embedding discriminative suppression units to force the network to focus on more nondiscriminative regions. Subsequently, visual word learning is introduced to identify additional object categories. Finally, four WSSS datasets are constructed based on public datasets with two simple and two complex scenarios. The results demonstrate that the proposed method outperforms 11 state-of-the-art methods, improving the intersection over union by at least 3.46% and 0.29% in both simple and complex scenarios, respectively.
Siyuan Wang 0011, Dongyang Hou, Yu Wang 0140, Bowen Cai 0002
IEEE Trans. Geosci. Remote. Sens.4
2025 Frequency Decoupling Fusion for Image Super-Resolution of Remote Sensing
abstract
Benefiting from the excellent global expression ability, transformer-based image super-resolution (SR) has made significant progress. However, the existing transformer-based SR methods still have the problem of high-frequency information reconstruction loss when processing remote sensing images due to their wide imaging range, rich high-frequency information and large differences, which affects the characterization ability of the transformer. In addition, the high computational overhead is unacceptable. To alleviate the above problems, we consider the remote sensing image SR from the perspective of the frequency domain. Specifically, we propose an efficient frequency decoupling-fusion remote sensing image SR framework, which is called FDFNet. In particular, we consider that when a large amount of previous work was carried out to extract features in the spatial domain, it was very easy to lose the high-frequency information in the original image. Therefore, we first introduce a frequency decoupling block (FDB), which decouples the image into low-frequency and high-frequency components, processes high-frequency and low-frequency information respectively in a divide-and-conquer manner, and restores high-frequency details before delving deeper. Furthermore, we notice that spatial self-attention is a low-pass filter that tends to have global perception and to demonstrate limitations in reconstructing high-frequency details. Therefore, we meticulously designed a parallel frequency-aware transformer module (PFTM) to extract spatial frequency attention and channel transposition attention, which enables our model to focus more on local texture details to restore high-frequency details. A large number of experimental results on multiple public datasets show that our FDFNet outperforms the state-of-the-art SR methods in quantitative metrics and visual quality, and achieves a balance between performance and efficiency within a limited computing budget.
Kanghui Zhao, Tao Lu 0001, Jiaming Wang 0001, Yu Wang 0140, Yuanzhi Wang, Yanduo Zhang
IEEE Trans. Geosci. Remote. Sens.4
2025 TripleA: An Unsupervised Domain Adaptation Framework for Nighttime VRU Detection
abstract
Detecting vulnerable road users (VRUs) at night presents significant challenges. Numerous methods rely heavily on annotations, yet the low visibility of nighttime images poses difficulties for labeling. To obviate the need for nighttime annotations, unsupervised domain adaptation manifests as a viable solution. However, existing approaches primarily focus on semantic-level domain gaps, often overlooking pixel-level discrepancies caused by inherent degradations in the nighttime domain. These degradations can impair machine vision and limit detection performance. In this paper, we propose TripleA, an unsupervised domain adaptation framework tailored for nighttime VRU detection. TripleA includes triple alignment. First, it aligns daytime and nighttime images to generate synthetic nighttime images, which are then enhanced for illumination and noise. To remove noise, we introduce an illumination difference-aware denoising network, incorporating a novel pseudo-supervised attention to achieve pixel-wise noise distribution alignment. This alignment is driven by pseudo-ground truth generated through a carefully designed exchange-recombination strategy, facilitating self-supervised training of the denoising network. Additionally, we introduce degradation alignment to ensure domain-invariant degradation encoding, which enhances the network’s robustness for real-world nighttime images. Extensive experiments demonstrate the effectiveness of our framework for nighttime VRU detection, all without the need for annotated nighttime data.
Yuankun Wang, Jiaming Wang 0001, Yu Wang 0140, Yulin Ding, Gui Cheng
IEEE Trans. Intell. Transp. Syst.4
2025 Rethinking the Role of Panchromatic Images in Pan-Sharpening
abstract
Recent pan-sharpening methods have predominantly utilized techniques tailored for natural image scenes, often overlooking the unique features arising from non-overlapping spectral responses. In light of this, we have reevaluated the utility of panchromatic (PAN) images and introduced a theory anchored in the spectral response of satellite sensors. This posits that a PAN image is effectively a linear weighted summation of individual bands from its corresponding multi-spectral (MS) image, offset by an error map. We developed a deep unmixing network termed “DUN” that integrates an unmixing network, a fusion mechanism, and a distinctive mutual information contrastive loss function. Notably, the unmixing network is adept at decomposing a PAN image into its MS counterpart and error map. Further, the demixed image alongside the low-resolution MS image is channeled into the fusion network for pan-sharpening. Recognizing the challenges of achieving robust supervised learning directly from the unmixing phase, we have innovated a mutual information contrastive learning loss function, ensuring enhanced separation and minimizing overlap during the unmixing process. Preliminary experiments underscore both the quantitative and qualitative prowess of the proposed method.
Jiaming Wang 0001, Xitong Chen, Xiao Huang 0003, Ruiqian Zhang, Yu Wang 0140, Tao Lu 0001
IEEE Trans. Multim.5
2024 Remote Sensing Pan-Sharpening via Cross-Spectral-Spatial Fusion Network
abstract
Pan-sharpening is a technique used to create high-resolution multispectral (HRMS) images by merging low-resolution multispectral (LRMS) images with corresponding high-resolution panchromatic (PAN) images. Despite achieving state-of-the-art performance, existing panchromatic sharpening networks based on deep-learning (DL) methods suffer from spectral distortion and insufficient spatial texture enhancement. To address this challenge, this letter introduces a novel cross-spectral–spatial fusion network (CSSFN) for pan-sharpening remote sensing images. The network utilizes a cross-spectral–spatial attention block (CSSAB) to extract both the spectral information of the MS branch and the spatial information of the PAN branch. The spectral and spatial feature representations of remote sensing images are then progressively enhanced, which improves the fusion process and generates multispectral images with high spatial resolution. Our network outperforms other pan-sharpening methods on two publicly available datasets, as demonstrated by extensive experiments, yielding state-of-the-art results.
Yu Wang 0140, Tao Lu 0001, Jiaming Wang 0001, Gui Cheng, Xiaolong Zuo, Chaoya Dang
IEEE Geosci. Remote. Sens. Lett.1
2023 Remote Sensing Image Super-Resolution via Multiscale Enhancement Network
abstract
In recent years, remote sensing images have attracted a lot of attention because of their special value. However, images acquired by satellite sensors are usually low-resolution (LR), so remote sensing images are much more difficult to infer high-frequency details from compared with ordinary digital images, which means they cannot meet the needs of certain downstream tasks. In this letter, we propose a multiscale enhancement network (MEN), which uses multiscale features of remote sensing images to enhance the network’s reconstruction capability. Specifically, the network extracts the coarse features of LR remote sensing images using convolutional layers. Then, these features are fed into the multiscale enhancement module (MEM) proposed by this network, which uses a combination of convolutional layers with multiple convolutional kernel sizes to refine the extraction of multiscale features, and finally, the final reconstructed image is generated by the reconstruction module. Extensive experiments show that MEN achieves significant reconstruction advantages in both objective and subjective aspects.
Yu Wang 0140, Tao Lu 0001, Changzhi Wu, Jiaming Wang 0001
IEEE Geosci. Remote. Sens. Lett.1
2023 AERNet: An Attention-Guided Edge Refinement Network and a Dataset for Remote Sensing Building Change Detection
abstract
Advancements in Earth observation technology enable the detection of surface changes in intricate urban environments. Building change detection (BCD) plays a crucial role in urban planning and environmental monitoring. However, existing deep learning-based BCD algorithms exhibit limited capability in feature extraction, feature relationship comprehension, sample imbalance mitigation, and accurate boundary identification for changed objects. To address these challenges, we introduce an attention-guided edge refinement network (AERNet) that employs a global context feature aggregation module (GCFAM) to aggregate information from extracted multi-layer context features. Our approach incorporates an attention decoding block (ADB) guided by enhanced coordinate attention (ECA) to capture channel and location associations between features. Furthermore, we utilize an edge refinement module (ERM) to enhance the network’s capacity to sense and refine the edges of changed areas. To tackle the issue of class imbalance and augment the algorithm’s feature learning ability, we devise a novel self-adaptive weighted binary cross-entropy (SWBCE) loss function, combined with a deep supervision (DS) strategy. Experiments are conducted on two publicly available datasets, GDSCD and LEVIR-CD, as well as our newly developed high-resolution complex urban scene BCD dataset, i.e., HRCUS-CD. The latter dataset comprises 11,388 pairs of images at 0.5-meter resolution and over 12,000 labeled change buildings. Comparative experiments indicate that AERNet surpasses advanced competitive methods, while ablation experiments demonstrate the effectiveness of AERNet’s model components and the SWBCE loss function. Efficiency comparison confirms that AERNet achieves comprehensive detection performance with superior effectiveness and robustness.
Jindou Zhang, Xiao Huang 0003, Yu Wang 0140, Xuechao Zhou, DeRen Li
IEEE Trans. Geosci. Remote. Sens.5
2022 Deep representation learning for face hallucination
Tao Lu 0001, Yu Wang 0140, Ruobo Xu, Wei Liu 0123, Wenhua Fang, Yanduo Zhang
Multim. Tools Appl.2
2021 Face Super-Resolution Through Dual-Identity Constraint
abstract
Recently, most existing face SR methods only focus on generating pleasant texture details, even artifacts. Identity information is an important high-level face attribute, which is often ignored in the low-level super-resolution (SR) task. In view of this, we propose a dual-identity constraint dual-loop network (DIDnet), which employs identity information to constrain the SR model. First, the proposed framework consists of two closed-loop networks: one of the networks is used for generating high resolution (HR) images for exploring identity-preserving in HR feature space and the other one can learn degradation process for utilizing low resolution (LR) identity information. Furthermore, we integrate dual-identity constraints together for rendering characteristic facial images. Extensive experimental results are conducted on face databases and real-world data, which confirmed that the proposed DIDnet consistently and significantly improves both objective and subjective facial image reconstruction performances.
Fangfang Cheng, Tao Lu 0001, Yu Wang 0140, Yanduo Zhang
ICME3
2021 Unsupervised Remoting Sensing Super-Resolution via Migration Image Prior
abstract
Recently, satellites with high temporal resolution have fostered wide attention in various practical applications. Due to limitations of bandwidth and hardware cost, however, the spatial resolution of such satellites is considerably low, largely limiting their potentials in scenarios that require spatially explicit information. To improve image resolution, numerous approaches based on training low-high resolution pairs have been proposed to address the super-resolution (SR) task. De-spite their success, however, low/high spatial resolution pairs are usually difficult to obtain in satellites with a high temporal resolution, making such approaches in SR impractical to use. In this paper, we proposed a new unsupervised learning framework, called "MIP", which achieves SR tasks without low/high resolution image pairs. First, random noise maps are fed into a designed generative adversarial network (GAN) for reconstruction. Then, the proposed method converts the reference image to latent space as the migration image prior. Finally, we update the input noise via an implicit method, and further transfer the texture and structured information from the reference image. Extensive experimental results on the Draper dataset show that MIP achieves significant improvements over state-of-the-art methods both quantitatively and qualitatively. The proposed MIP is open-sourced at https://github.com/jiaming-wang/MIP.
Jiaming Wang 0001, Tao Lu 0001, Xiao Huang 0003, Ruiqian Zhang, Yu Wang 0140
ICME6
2021 Face Hallucination via Split-Attention in Split-Attention Network
abstract
Recently, convolutional neural networks (CNNs) have been widely employed to promote the face hallucination due to the ability to predict high-frequency details from a large number of samples. However, most of them fail to take into account the overall facial profile and fine texture details simultaneously, resulting in reduced naturalness and fidelity of the reconstructed face, and further impairing the performance of downstream tasks (e.g., face detection, facial recognition). To tackle this issue, we propose a novel external-internal split attention group (ESAG), which encompasses two paths responsible for facial structure information and facial texture details, respectively. By fusing the features from these two paths, the consistency of facial structure and the fidelity of facial details are strengthened at the same time. Then, we propose a split-attention in split-attention network (SISN) to reconstruct photorealistic high-resolution facial images by cascading several ESAGs. Experimental results on face hallucination and face recognition unveil that the proposed method not only significantly improves the clarity of hallucinated faces, but also encourages the subsequent face recognition performance substantially. Codes have been released at https://github.com/mdswyz/SISN-Face-Hallucination.
Tao Lu 0001, Yuanzhi Wang, Yanduo Zhang, Yu Wang 0140, Wei Liu 0123, Zhongyuan Wang 0001, Junjun Jiang
ACM Multimedia4
2021 Single Image Super-Resolution via Multi-Scale Information Polymerization Network
abstract
Recently, the performances of deep convolution neural networks (CNNs)-based single-image super-resolution (SISR) have been significantly improved. However, most of the existing CNN-based SISR methods mainly focus on wider or deeper networks and ignore the potential relationship between multi-scale features, leading to the limited representation ability of the reconstructed network. To address this problem, we propose a new multi-scale information polymerization network (MIPN). Specifically, we propose a multi-scale information polymerization block (MIPB), which uses convolution layers of different convolution kernel sizes to extract multi-scale image features, and effectively polymerizate the extracted features together to obtain fine image features. Moreover, we also propose a shallow residual block in MIPB. Compared with the traditional convolution layer, this proposed block can effectively extract image features without increasing the number of parameters. Extensive experiments show that the proposed method performs better than several state-of-the-art methods in quantitative and visual quality indicators.
Tao Lu 0001, Yu Wang 0140, Jiaming Wang 0001, Wei Liu 0123, Yanduo Zhang
IEEE Signal Process. Lett.2
2020 Face Super-Resolution by Learning Multi-view Texture Compensation
Yu Wang 0140, Tao Lu 0001, Ruobo Xu, Yanduo Zhang
MMM (2)1