VLDB 2026 Research / reviewers in the wild / expert
Zhenyu Wang 0008
dblp:22/1486-8
· DBLP profile ↗
13ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-4259-3073ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoCoFR: Collaborative codebooks learning with soft matching strategy for blind face restoration
Teng Feng, Zhenyu Wang 0008, Weisheng Dong, Xin Li 0005, Guangming Shi |
Neural Networks | 4 |
| 2026 | Visible-infrared joint image deraining for harsh rain conditions with cross-modal semantic consistency
Xin Li 0005, Chengpei Xu, Zhenyu Wang 0008, Weisheng Dong |
Pattern Recognit. | 5 |
| 2026 | LoopExpose: An Unsupervised Framework for Arbitrary-Length Exposure CorrectionabstractExposure correction is essential for enhancing image quality under challenging lighting conditions. While supervised learning has achieved significant progress in this area, it relies heavily on large-scale labeled datasets, which are difficult to obtain in practical scenarios. To address this limitation, we propose a pseudo label-based unsupervised method called LoopExpose for arbitrary-length exposure correction. A nested loop optimization strategy is proposed to address the exposure correction problem, where the correction model and pseudo-supervised information are jointly optimized in a two-level framework. Specifically, the upper-level trains a correction model using pseudo-labels generated through multi-exposure fusion at the lower level. A feedback mechanism is introduced where corrected images are fed back into the fusion process to refine the pseudo-labels, creating a self-reinforcing learning loop. Considering the dominant role of luminance calibration in exposure correction, a Luminance Ranking Loss is introduced to leverage the relative luminance ordering across the input sequence as a self-supervised constraint. Extensive experiments on different benchmark datasets demonstrate that LoopExpose achieves superior exposure correction and fusion performance, outperforming existing state-of-the-art unsupervised methods. Code is available at https://github.com/FALALAS/LoopExpose. Zhenyu Wang 0008, Weisheng Dong |
IEEE Trans. Image Process. | 2 |
| 2025 | Bridging Task Boundaries: Remote Sensing Image-Text Retrieval via Dictionary-Driven AdaptationabstractGiven image (or text), remote sensing image-text retrieval (RSITR) aims to retrieve corresponding text (or image) within diverse remote sensing data. However, due to the complex scenes and compact distribution of targets in remote sensing data, existing methods, particularly those leveraging large models like CLIP, often generate features with high intra-modal similarity and insufficient distinctive characteristics, thus resulting in suboptimal retrieval performance. To address these issues, we pro- pose a novel dictionary-based RSITR method that jointly models image and text feature estimation. Specifically, by incorporating a general dictionary and the corresponding sparse coefficients, our method more effectively captures the correlations between the learned features. Furthermore, we introduce adaptive weighted metric learning on a sample-by-sample basis to encourage the model to focus on more challenging negative samples, promoting fine-grained feature alignment. The extensive experiments on the RSICD and RSITMD datasets demonstrate the effectiveness of our method, demonstrating significant improvements in retrieval performance. Zhenyu Wang 0008, Weisheng Dong, Xin Li 0005 |
ICASSP | 3 |
| 2025 | PatternCIR Benchmark and TisCIR: Advancing Zero-Shot Composed Image Retrieval in Remote SensingabstractRemote sensing composed image retrieval (RSCIR) is a new vision-language task that takes a composed query of an image and text, aiming to search for a target remote sensing image satisfying two conditions from intricate remote sensing imagery. However, the existing attribute-based benchmark Patterncom in RSCIR has significant flaws, including the lack of query text sentences and paired triplets, thus making it unable to evaluate the latest methods. To address this, we propose the Zero-Shot Query Text Generator (ZS-QTG) that can generate full query text sentences based on attributes, and then, by capitalizing on ZS-QTG, we develop the PatternCIR benchmark. Pattern CIR rectifies Patterncom’s deficiencies and enables the evaluation of existing methods. Additionally, we explore zero-shot composed image retrieval methods that do not rely on massive pre-collected triplets for training. Existing methods use only the text during retrieval, performing poorly in RSCIR. To improve this, we propose Text-image Sequential Training of Composed Image Retrieval (TisCIR). TisCIR undergoes sequential training of multiple self-masking projection and fine-grained image attention modules, which endows it with the capacity to filter out conflicting information between the image and text, enhancing the retrieval by utilizing both modalities in harmony. TisCIR outperforms existing methods by 12.40% to 62.03% on PatternCIR, achieving state-of-the-art performance in RSCIR. The data and code are available here. Zhechun Liang, Shiwen Xue, Zhenyu Wang 0008, Weisheng Dong, Xin Li 0005, Guangming Shi |
IJCAI | 5 |
| 2025 | Pushing the Limit of Binarized Neural Network for Image Super Resolution with Smooth Information TransmissionabstractLightweight models are currently the focal point in image super-resolution (ISR) research, of which the application on resource-limited devices is constrained by heavy computational requirements. As an efficient approach to enhance the inference efficiency of deep learning models, low-bit quantization has garnered significant interest. In this paper, we emphasize that low-bit ISR is not merely a parody of its full-precision version and explore binary quantization in ISR from the perspective of information transmission, pushing the limits of binarized ISR. Specifically, we propose a Maximum Entropy Routing (MER) mechanism to dynamically control activation distribution, maximizing the information entropy of binarized feature maps. Additionally, a Learnable Deviation Compensation (LDC) and an Adaptive Step-size Estimation (ASE) are introduced to reduce information loss during the forward and backward passes, respectively. By enabling smoother information transmission through more flexible binarized activation representations and more precise gradient estimation, the performance gap between binarized and full-precision models is narrowed to less than 0.3 dB. Extensive experiments demonstrate that our proposed binarization method achieves state-of-the-art results in Peak Signal-to-Noise Ratio (PSNR) across all popular benchmarks. Weimin Cheng, Zhenyu Wang 0008, Weisheng Dong |
ACM Multimedia | 2 |
| 2025 | Compressing Vision Transformer from the View of Model Property in Frequency Domain
Zhenyu Wang 0008, Xuemei Xie, Hao Luo 0004, Weisheng Dong, Yongxu Liu 0001, Fan Wang 0019, Guangming Shi |
Int. J. Comput. Vis. | 1 |
| 2025 | Forgetting the Background: A Masking Approach for Enhanced Infrared Small-Target DetectionabstractInfrared small-target detection (ISTD) in a single frame is an essential, yet challenging task due to its small size of targets, weak energy, and clutter background. Current methods either design complex network architectures to facilitate multilevel information interaction (e.g., DNA-Net and UIU-Net) or introduce structural texture priors to enhance feature discrimination (e.g., SRNet and CSRNet). However, both methods fail to explicitly distinguish or suppress the interference of complex background from infrared small targets, which makes them easy to “get lost” in clutter background with insufficient attention to the targets. In this work, we innovatively propose a novel background-masking approach (denoted as BGM) for ISTD. The proposed BGM aims to force the network to focus exclusively on the target by masking out irrelevant background information, thereby enhancing the network’s ability to detect weak and small infrared targets. Specifically, we present a new ISTD method that leverages a proxy training task with masking, enabling the network to simultaneously predict on both the original input and the masked data, where the background is randomly masked/forgotten. This strategy allows for a better concentration of the model on the shapeless targets rather than the cluttered background. The method is flexible with a simple U-shaped network without complicated manipulation and also computationally efficient without increasing the overall computational burden during inference. Extensive experiments demonstrate that our proposed BGM effectively enhances the detection performance of infrared small targets and achieves 70.8% mean intersection over union (mIoU) on IRSTD-1K. The source code would be available athttps://github.com/ZhihaoMa123/BGM Yongxu Liu 0001, Wenxiang Zhu, Na Li 0040, Chuang Li 0005, Zhenyu Wang 0008, Wei Feng 0004, Junzheng Jiang, Yinghui Quan |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | BVT-IMA: Binary Vision Transformer with Information-Modified AttentionabstractAs a compression method that can significantly reduce the cost of calculations and memories, model binarization has been extensively studied in convolutional neural networks. However, the recently popular vision transformer models pose new challenges to such a technique, in which the binarized models suffer from serious performance drops. In this paper, an attention shifting is observed in the binary multi-head self-attention module, which can influence the information fusion between tokens and thus hurts the model performance. From the perspective of information theory, we find a correlation between attention scores and the information quantity, further indicating that a reason for such a phenomenon may be the loss of the information quantity induced by constant moduli of binarized tokens. Finally, we reveal the information quantity hidden in the attention maps of binary vision transformers and propose a simple approach to modify the attention values with look-up information tables so that improve the model performance. Extensive experiments on CIFAR-100/TinyImageNet/ImageNet-1k demonstrate the effectiveness of the proposed information-modified attention on binary vision transformers. Zhenyu Wang 0008, Hao Luo 0004, Xuemei Xie, Fan Wang 0019, Guangming Shi |
AAAI | 1 |
| 2023 | Filter Clustering for Compressing CNN Model With Better Feature DiversityabstractAs a practical approach for compressing convolutional neural networks (CNNs), network pruning has been rapidly developed in recent years. The conventional methods prune inactive filters permanently from models to reduce the width of each layer and then train the pruned model until convergence. However, such methods have limitations in that: (1) The activation-based pruning criteria ignore the correlation between filters, leading to attenuation in types of features; (2) The permanent filter removal restricts the architecture of models in the subsequent training so that reducing the chances of learning more features; (3) The single-width compression may generate narrow layers that block the information flow, resulting in limited feature capacity in the next layers and hard optimization. These limitations reduce the feature diversity in the pruned model and thus lead to sub-optimal model quality. In this paper, a compression method named filter clustering is proposed to rectify the problem of poor feature diversity in traditional pruning and achieve better model quality from three perspectives. Firstly, to maintain the variety of features after pruning, we treat the model compression as a clustering task and merge filters with similar outputs, rather than removing inactive filters. Specifically, a handy estimation approach is designed to convert the similarity of the output into filter similarity, which liberates the measurement from sampling numerous images. Secondly, to increase the probability of learning more features during training, we propose a periodic training and clustering pipeline, which creates a larger optimization space by dynamically exploring different sub-model architectures. Finally, to prevent the feature capacity from being influenced by the narrow layers, we introduce and leverage a fusible anti-blocking branch to smoothly remove such layers. Extensive experiments demonstrate that the proposed method can achieve compact models with better feature diversity and reduce 1%~15% more calculations than the previous methods while maintaining performance. Zhenyu Wang 0008, Xuemei Xie, Qinghang Zhao, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | VTC-LFC: Vision Transformer Compression with Low-Frequency ComponentsabstractAlthough Vision transformers (ViTs) have recently dominated many vision tasks, deploying ViT models on resource-limited devices remains a challenging problem. To address such a challenge, several methods have been proposed to compress ViTs. Most of them borrow experience in convolutional neural networks (CNNs) and mainly focus on the spatial domain. However, the compression only in the spatial domain suffers from a dramatic performance drop without fine-tuning and is not robust to noise, as the noise in the spatial domain can easily confuse the pruning criteria, leading to some parameters/channels being pruned incorrectly. Inspired by recent findings that self-attention is a low-pass filter and low-frequency signals/components are more informative to ViTs, this paper proposes compressing ViTs with low-frequency components. Two metrics named low-frequency sensitivity (LFS) and low-frequency energy (LFE) are proposed for better channel pruning and token pruning. Additionally, a bottom-up cascade pruning scheme is applied to compress different dimensions jointly. Extensive experiments demonstrate that the proposed method could save 40% ~ 60% of the FLOPs in ViTs, thus significantly increasing the throughput on practical devices with less than 1% performance drop on ImageNet-1K. Zhenyu Wang 0008, Hao Luo 0004, Pichao Wang, Fan Wang 0019, Hao Li 0030 |
NeurIPS | 1 |
| 2020 | A Multi-level Equilibrium Clustering Approach for Unsupervised Person Re-identification
Fangyu Wang, Zhenyu Wang 0008, Xuemei Xie, Guangming Shi |
PRCV (3) | 2 |
| 2020 | Network pruning using sparse learning and genetic algorithm
Zhenyu Wang 0008, Fu Li 0002, Guangming Shi, Xuemei Xie, Fangyu Wang |
Neurocomputing | 1 |