VLDB 2026 Research / reviewers in the wild / expert
Zeyu Wang 0010
dblp:132/7882-10
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-0985-4478ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion TransformersabstractThe Diffusion Transformers Models (DiTs) have transitioned the network architecture from traditional UNets to transformers, demonstrating exceptional capabilities in image generation. Although DiTs have been widely applied to high-definition video generation tasks, their large parameter size hinders inference on edge devices. Vector quantization (VQ) can decompose model weight into a codebook and assignments, allowing extreme weight quantization and significantly reducing memory usage. In this paper, we propose VQ4DiT, a fast post-training vector quantization method for DiTs. We found that traditional VQ methods calibrate only the codebook without calibrating the assignments. This leads to weight sub-vectors being incorrectly assigned to the same assignment, providing inconsistent gradients to the codebook and resulting in a suboptimal result. To address this challenge, VQ4DiT calculates the candidate assignment set for each weight sub-vector based on Euclidean distance and reconstructs the sub-vector based on the weighted average. Then, using the zero-data and block-wise calibration method, the optimal assignment from the set is efficiently selected while calibrating the codebook. VQ4DiT quantizes a DiT XL/2 model on a single NVIDIA A100 GPU within 20 minutes to 5 hours depending on the different quantization settings. Experiments show that VQ4DiT establishes a new state-of-the-art in model size and performance trade-offs, quantizing weights to 2-bit precision while retaining acceptable image generation quality. Juncan Deng, Shuaiting Li, Zeyu Wang 0010, Kedong Xu, Kejie Huang |
AAAI | 3 |
| 2025 | MVQ: Towards Efficient DNN Compression and Acceleration with Masked Vector QuantizationabstractVector quantization(VQ) is a hardware-friendly DNN compression method that can reduce the storage cost and weight-loading datawidth of hardware accelerators. However, conventional VQ techniques lead to significant accuracy loss because the important weights are not well preserved. To tackle this problem, a novel approach called MVQ is proposed, which aims at better approximating important weights with a limited number of codewords. At the algorithm level, our approach removes the less important weights through N:M pruning and then minimizes the vector clustering error between the remaining weights and codewords by the masked k-means algorithm. Only distances between the unpruned weights and the codewords are computed, which are then used to update the codewords. At the architecture level, our accelerator implements vector quantization on an EWS (Enhanced weight stationary) CNN accelerator and proposes a sparse systolic array design to maximize the benefits brought by masked vector quantization. Shuaiting Li, Chengxuan Wang, Juncan Deng, Zeyu Wang 0010, Zewen Ye, Zongsheng Wang, Haibin Shen, Kejie Huang |
ASPLOS (1) | 4 |
| 2025 | ViM-VQ: Efficient Post-Training Vector Quantization for Visual MambaabstractVisual Mamba networks (ViMs) extend the selective state space model (Mamba) to various vision tasks and demonstrate significant potential. As a promising compression technique, vector quantization (VQ) decomposes network weights into codebooks and assignments, significantly reducing memory usage and computational latency, thereby enabling the deployment of ViMs on edge devices. Although existing VQ methods have achieved extremely low-bit quantization (e.g., 3-bit, 2-bit, and 1-bit) in convolutional neural networks and Transformer-based networks, directly applying these methods to ViMs results in unsatisfactory accuracy. We identify several key challenges: 1) The weights of Mamba-based blocks in ViMs contain numerous outliers, significantly amplifying quantization errors. 2) When applied to ViMs, the latest VQ methods suffer from excessive memory consumption, lengthy calibration procedures, and suboptimal performance in the search for optimal codewords. In this paper, we propose ViM-VQ, an efficient post-training vector quantization method tailored for ViMs. ViM-VQ consists of two innovative components: 1) a fast convex combination optimization algorithm that efficiently updates both the convex combinations and the convex hulls to search for optimal codewords, and 2) an incremental vector quantization strategy that incrementally confirms optimal codewords to mitigate truncation errors. Experimental results demonstrate that ViM-VQ achieves state-of-the-art performance in low-bit quantization across various visual tasks. Juncan Deng, Shuaiting Li, Zeyu Wang 0010, Kedong Xu, Kejie Huang |
ICCV | 3 |
| 2025 | AdvSpoofGuard: Optimal transport driven robust face presentation attack detection system
Taha Hasan Masood Siddique, Shujaat Khan, Zeyu Wang 0010, Kejie Huang |
Knowl. Based Syst. | 3 |
| 2025 | UP-Diff: Latent Diffusion Model for Remote Sensing Urban PredictionabstractRemote sensing (RS) technology has become essential for monitoring urban development, including applications like population growth analysis, transportation congestion forecasting, and climate change detection (CD). However, its potential for future urban planning (UP), particularly in predicting future urban layouts remains largely unexplored. This study introduces UP-Diff, a novel method leveraging generative models for UP, to address this gap. UP-Diff leverages information from current urban layouts and planned change maps to predict future urban configurations. Key challenges addressed include the integration of urban layouts and change maps into latent diffusion model (LDM) through careful architecture improvements and the mitigation of limited training data by employing a pretrained stable diffusion (SD) model with fixed weights, trainable ConvNeXt, and trainable cross-attention layers. Our method significantly streamlines the UP process by automating layout predictions, thus reducing the time and effort required compared to traditional manual methods. Comprehensive evaluations on the learning, vision, and RS dataset (LEVIR-CD) and Sun Yat-Sen University dataset (SYSU-CD) validate that UP-Diff achieves high-fidelity predictions of future urban layouts, demonstrating its effectiveness and potential for advancing RS-based UP methodologies. Our code and model weights are available athttps://github.com/zeyuwang-zju/UP-Diff. Zeyu Wang 0010, Zecheng Hao, Yuhan Zhang 0006, Yuchao Feng, Yufei Guo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Cross-Modal Adaptation for Object Detection in Infrared Remote Sensing ImageryabstractModern Thermal InfraRed (TIR) technology has been proven highly significant in Remote Sensing Imagery (RSI). Currently, multimodal RSI object detection based on RGB-TIR image pairs has attracted widespread research. However, capturing features in the TIR domain poses a challenge, as existing object detectors heavily focus on chromatic information in the RGB domain. Furthermore, the quality of RGB images can be influenced by complex environmental conditions, limiting the practicality of multimodal detection. In this paper, we introduce Cross-Modal-YOLO (CM-YOLO), a lightweight yet effective object detector specifically designed for TIR remote sensing images. CM-YOLO employs cross-modal adaptation to enhance the awareness of TIR-RGB modality translation. Specifically, we leverage a Prior Modality Translator (PMT) to learn the InfraRed-Visible (IV) features, which are incorporated into the detection backbone using our IV-Gate modules. Experimental results on the VEDAI dataset demonstrate that CM-YOLO significantly outperforms conventional methods. Moreover, CM-YOLO exhibits a strong generalization ability for TIR-based object detection in urban scenes on the FLIR dataset. Zeyu Wang 0010, Shuaiting Li, Kejie Huang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | A Gray-Box Attack Against Latent Diffusion Model-Based Image Editing by Posterior CollapseabstractRecent advancements in Latent Diffusion Models (LDMs) have revolutionized image synthesis and manipulation, raising significant concerns about data misappropriation and intellectual property infringement. While adversarial attacks have been extensively explored as a protective measure against such misuse of generative AI, current approaches are severely limited by their heavy reliance on model-specific knowledge and substantial computational costs. Drawing inspiration from the posterior collapse phenomenon observed in VAE training, we propose the Posterior Collapse Attack (PCA), a novel framework for protecting images from unauthorized manipulation. Through comprehensive theoretical analysis and empirical validation, we identify two distinct collapse phenomena during VAE inference: diffusion collapse and concentration collapse. Based on this discovery, we design a unified loss function that can flexibly achieve both types of collapse through parameter adjustment, each corresponding to different protection objectives in preventing image manipulation. Our method significantly reduces dependence on model-specific knowledge by requiring access to only the VAE encoder, which constitutes less than 4% of LDM parameters. Notably, PCA achieves prompt-invariant protection by operating on the VAE encoder before text conditioning occurs, eliminating the need for empty prompt optimization required by existing methods. This minimal requirement enables PCA to maintain adequate transferability across various VAE-based LDM architectures while effectively preventing unauthorized image editing. Extensive experiments show PCA outperforms existing techniques in protection effectiveness, computational efficiency (runtime and VRAM), and generalization across VAE-based LDM variants. Our code is available at https://github.com/ZhongliangGuo/PosteriorCollapseAttack. Zhongliang Guo 0001, Chun Tong Lei, Lei Fang 0001, Shuai Zhao 0007, Yifei Qian, Zeyu Wang 0010, Cunjian Chen, Ognjen Arandjelovic, Chun Pong Lau 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | Bridging partial-gated convolution with transformer for smooth-variation image inpainting
Zeyu Wang 0010, Haibin Shen, Kejie Huang |
Multim. Tools Appl. | 1 |
| 2024 | Pair-ID: A Dual Modal Framework for Identity Preserving Image GenerationabstractThe acquisition of large-scale paired visible and thermal images is crucial for enhancing face recognition systems, especially in low-light environments where visible spectrum images fail. However, the task is hindered by the scarcity of thermal images and the need for identity consistency during image generation. In this paper, we propose Pair-ID, an innovative framework that addresses these challenges by creating a shared latent space for simultaneous generation of paired visible and thermal images. Pair-ID integrates identity information into text embeddings and employs fixed templates for diverse facial poses, streamlining the customization process and reducing computational demands. The framework's Joint Learner encodes both modalities, facilitating synchronized image generation and preserving facial details. Extensive evaluations show that Pair-ID surpasses current methods in efficiency and performance for paired data generation, making it a promising solution for face recognition under varying lighting conditions. Yongrong Wu, Zeyu Wang 0010, Xiaode Liu, Yufei Guo 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | Thermal Infrared Image Inpainting Via Edge-Aware GuidanceabstractImage inpainting has achieved fundamental advances with deep learning. However, almost all existing inpainting methods aim to process natural images, while few target Thermal Infrared (TIR) images, which have widespread applications. When applied to TIR images, conventional inpainting methods usually generate distorted or blurry content. In this paper, we propose a novel task—Thermal Infrared Image Inpainting, which aims to reconstruct missing regions of TIR images. Crucially, we propose a novel deep-learning-based model TIR-Fill. We adopt the edge generator to complete the canny edges of broken TIR images. The completed edges are projected to the normalization weights and biases to enhance edge awareness of the model. In addition, a refinement network based on gated convolution is employed to improve TIR image consistency. The experiments demonstrate that our method outperforms state-of-the-art image inpainting approaches on FLIR thermal dataset. Zeyu Wang 0010, Haibin Shen, Changyou Men, Kejie Huang |
ICASSP | 1 |
| 2023 | TIRDet: Mono-Modality Thermal InfraRed Object Detection Based on Prior Thermal-To-Visible TranslationabstractCross-modality images that combine visible-infrared spectra can provide complementary information for object detection. In particular, they are well-suited for autonomous vehicle applications in dark environments with limited illumination. However, it is time-consuming to acquire a large number of pixel-aligned visible-thermal image pairs, and real-time alignment is challenging in practical driving systems. Furthermore, the quality of visible-spectrum images can be adversely affected by complex environmental conditions. In this paper, we propose a novel neural network called TIRDet, which only utilizes Thermal InfraRed (TIR) images for mono-modality object detection. To compensate for the lacked visible-band information, we adopt a prior Thermal-To-Visible (T2V) translation model to obtain the translated visible images and the latent T2V codes. In addition, we introduce a novel attention-based Cross-Modality Aggregation (CMA) module, which can augment the modality-translation awareness of TIRDet by preserving the T2V semantic information. Extensive experiments on FLIR and LLVIP datasets demonstrate that our TIRDet significantly outperforms all mono-modality detection methods based on thermal images, and it even surpasses most State-Of-The-Art (SOTA) multispectral methods using visible-thermal image pairs. Code is available at https://github.com/zeyuwang-zju/TIRDet Zeyu Wang 0010, Fabien Colonnier, Jinghong Zheng 0001, Jyotibdha Acharya, Kejie Huang |
ACM Multimedia | 1 |