Wenchao Zhang 0001

dblp:33/6703-1 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-6441-4232ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Image Inpainting in 30 Years: A Survey
abstract
ABSTRACT As a fundamental task in restoring continuous visual signals, image inpainting plays a critical role in autonomous driving perception, medical imaging, video editing and digital heritage preservation. Driven by deep learning and large‐scale generative models, the field has transitioned from low‐level texture synthesis to high‐level semantic generation, yielding major breakthroughs in structural fidelity and visual realism. Centring on the generative paradigm as the architectural trajectory, this survey systematically categorizes the 30‐year evolution of image inpainting into three distinct technological generations: traditional prior‐driven synthesis, deep learning data‐driven reconstruction and modern foundation model‐driven generation. Despite this progress, highly competitive methods still struggle with large‐scale missing regions, global consistency in complex scenes, fine‐grained micro‐details and alignment with human visual perception. To address these gaps, we critically evaluate the technical paradigms and main bottlenecks within each of these evolutionary stages. We categorize and compare mainstream breakthroughs across high‐resolution restoration, text‐guided synthesis and complex scene generation. Furthermore, we compile standard benchmarks, evaluation metrics and quantitative performance comparisons of representative algorithms. Finally, we dissect open challenges—focusing on cross‐scene generalization and evaluation metric alignment—and outline future trajectories, particularly the integration of inpainting with text‐guided foundation models, providing a definitive reference for future theoretical and engineering advancements.
Hengxiang Zhao, Wenchao Zhang 0001, Yu Zheng 0021, Michele Nappi, Junxin Chen 0001
Expert Syst. J. Knowl. Eng.3
2026 Adapting vision foundation models with lightweight trident decoder for remote sensing change detection
Wenhui Ye, Weimin Lei, Wenchao Zhang 0001, Wei Zhang 0033
Expert Syst. Appl.3
2026 A colon polyp segmentation network via collaborative decision-making of mixture of experts
Wenchao Zhang 0001, Wenhui Ye, Zhenhua Yu 0001, Jianguo Ju
Expert Syst. Appl.1
2026 CSANet: Adaptive all-to-all cross-scale alignment for robust remote sensing change detection
Wenchao Zhang 0001, Wenhui Ye
Inf. Sci.1
2025 SRA-VFM: Boosting Remote Sensing Change Detection via Slice-Reassembled Augmentation and Vision Foundation Model-Guided Dual Streams
abstract
To address the dual challenges of inadequate deep semantic feature representation and limited data diversity in bi-temporal remote sensing change detection (RSCD), we propose a collaborative optimization framework (SRA-VFM) integrating customized data augmentation and hierarchical feature interpretation. SRA-VFM comprises three core components: slice-reassemble augmentation (SRA), a dual-stream feature encoding-decoding network, and a multi-task segmentation head. The SRA module synthesizes diverse training samples through random slicing and semantic reassembly while preserving local feature consistency. The dual-stream encoder, based on the FastSAM pre-trained model, incorporates a top-down feature adapter to align pre-trained features with remote sensing data distributions via cross-level fusion and low-dimensional semantic mapping. The bottom-up decoder leverages semantically aligned pyramid features for progressive upsampling, restoring high-resolution spatial details and enhancing multi-scale representation. The multi-task head jointly optimizes change region detection and edge refinement, accelerating convergence and improving boundary localization. Extensive experiments on 5 benchmark datasets (LEVIR-CD, WHU-CD, CLCD, S2Looking, SYSU-CD) demonstrate SRA-VFM’s superiority over state-of-the-art (SOTA) methods, achieving mF1/mIoU of 96.01%/92.55%, 97.13%/94.53%, 88.91%/81.21%, 83.13%/48.91%, and 88.36%/79.67% respectively. Code will be publicly available upon publication.
Wenhui Ye, Weimin Lei, Wenchao Zhang 0001, Wei Zhang 0033
IEEE Trans. Geosci. Remote. Sens.3
2025 Multi-Scale Dynamic Sparse Attention UNet for Medical Image Segmentation
abstract
Transformers have recently gained significant attention in medical image segmentation due to their ability to capture long-range dependencies. However, the presence of excessive background noise in large regions of medical images introduces distractions and increases the computational burden on the fine-grained self-attention (SA) mechanism, which is a key component of the transformer model. Meanwhile, preserving fine-grained details is essential for accurately segmenting complex, blurred medical images with diverse shapes and sizes. Thus, we propose a novel Multi-scale Dynamic Sparse Attention (MDSA) module, which flexibly reduces computational costs while maintaining multi-scale fine-grained interactions with content awareness. Specifically, multi-scale aggregation is first applied to the feature maps to enrich the diversity of interaction information. Then, for each query, irrelevant key-value pairs are filtered out at a coarse-grained level. Finally, fine-grained SA is performed on the remaining key-value pairs. In addition, we design an enhanced downsampling merging (EDM) module and an enhanced upsampling fusion (EUF) module for building pyramid architectures. Using MDSA to construct the basic blocks, combined with EDMs and EUFs, we develop a UNet-like model named MDSA-UNet. Since MDSA-UNet dynamically processes only a small subset of relevant fine-grained features, it achieves strong segmentation performance with high computational efficiency. Extensive experiments on four datasets spanning three different types demonstrate that our MDSA-UNet, without using pre-training, significantly outperforms other non-pretrained methods and even competes with pre-trained models, achieving Dice scores of 82.10% on DDTI, 80.20% on TN3K, 90.75% on ISIC2018, and 91.05% on ACDC. Meanwhile, our model maintains lower complexity, with only 6.65 M parameters and 4.54 G FLOPs at a resolution of 224 × 224, ensuring both effectiveness and efficiency. Code is available at URL.
Chong Fu 0001, Wenchao Zhang 0001, Junxin Chen 0001, Chiu-Wing Sham
IEEE J. Biomed. Health Informatics4
2024 Joint object contour points and semantics for instance segmentation
abstract
Abstract The edges of objects are of great significance to the task of instance segmentation. However, most of the current popular deep neural networks do not pay much attention to the object edge information. More importantly, using the down‐sampling pooling layer in the deep learning network, the edge detail information of the object will be lost. To address this issue, inspired by the manual annotation process, we propose Mask Point R‐CNN aiming at promoting the neural network's attention to the object boundary. Specifically, we introduce the auxiliary task of object contour point detection on the Mask R‐CNN framework, which can effectively improve the gradient flow between different tasks by multi‐task learning and repairing objects' boundary information via feature fusion. Consequently, the model can be more sensitive to the edges of the object and capture more geometric features. Quantitatively, the experimental results show that our Mask Point R‐CNN outperforms vanilla Mask R‐CNN by 3.8% on the Cityscapes dataset and 0.8% on the COCO dataset.
Wenchao Zhang 0001, Chong Fu 0001, Mai Zhu, Lin Cao 0003, Ming Tie, Chiu-Wing Sham
Expert Syst. J. Knowl. Eng.1
2024 DMSA-UNet: Dual Multi-Scale Attention makes UNet more strong for medical image segmentation
Chong Fu 0001, Wenchao Zhang 0001, Chiu-Wing Sham, Junxin Chen 0001
Knowl. Based Syst.4
2024 Attention-based deep supervised hashing for near duplicate video retrieval
Naifei Shi, Chong Fu 0001, Ming Tie, Wenchao Zhang 0001, Xingwei Wang 0001, Chiu-Wing Sham
Neural Comput. Appl.4
2024 Encrypted Video Search with Single/Multiple Writers
abstract
Video-based services have become popular. Clients often outsource their videos to the cloud to relieve local maintenance. However, privacy has become a major concern, since many videos contain sensitive information. Although retrieving (unencrypted) videos has been extensively investigated, retrieving encrypted multimedia has received relatively rare attention, at best in a limitation of image-based similarity searches. We initiate the study of scalable encrypted video search, enabling clients to query videos similar to an image search. Our modular framework leverages intrinsic attributes of videos, such as semantics and visuals, to effectively capture their contents. We propose a two-step approach whereby lightweight searchable encryption techniques are used for pre-screening, followed by an interactive approach for fine-grained search. Furthermore, we present three instantiations, including one centralized-writer instantiation and two distributed-writer instantiations, to effectively cater to varying needs and scenarios: (1) The centralized one employs forward and backward private searchable encryption [CCS 2017] over deep hashing [CVPR 2020]. (2) Motivated by distributed computing, the multi-writer instantiations building atop HSE [Usenix Security 2022] allows searching the relevant videos contributed by multiple intuitions collaboratively. Our experimental results illustrate their practical performance over multiple real-world datasets, whether in a centralized setting or distributed setting.
Yu Zheng 0021, Wenchao Zhang 0001, Xiuhua Wang 0009, Chong Fu 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Remote sensing image instance segmentation network with transformer and multi-scale feature representation
Wenhui Ye, Wei Zhang 0033, Weimin Lei, Wenchao Zhang 0001
Expert Syst. Appl.4
2022 CODH++: Macro-semantic differences oriented instance segmentation network
Wenchao Zhang 0001, Chong Fu 0001, Lin Cao 0003, Chiu-Wing Sham
Expert Syst. Appl.1
2022 A more compact object detector head network with feature enhancement and relational reasoning
Wenchao Zhang 0001, Chong Fu 0001, Xiang shi Chang, Teng fei Zhao, Chiu-Wing Sham
Neurocomputing1
2021 Global context aware RCNN for object detection
Wenchao Zhang 0001, Chong Fu 0001, Haoyu Xie 0002, Mai Zhu, Ming Tie, Junxin Chen 0001
Neural Comput. Appl.1