Wenzhuo Ma

dblp:176/1080 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0003-8964-0070ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 On Performance of NNVC Inter-Coding
Xinxin Chen, Nianxiang Fu, Junxi Zhang, Ding Ding 0004, Wenzhuo Ma, Zhenzhong Chen 0001
ISCAS6
2026 Learning-enhanced Video Compression with Capability beyond VVC
Luyi Qin, Xinxin Chen, Nianxiang Fu, Haodong Qu, Wenzhuo Ma, Junxi Zhang, Zhenzhong Chen 0001
ISCAS6
2026 Multiscale feature optimization for accurate small object detection in remote sensing imagery
Bingxiang Wang, Mugen Zhou, Wenzhuo Ma
Mach. Vis. Appl.3
2025 Low-Decoding-Complexity Learned Image Compression with Masked Convolutional Layer Re-parameterization
Wenzhuo Ma, Nianxiang Fu, Junxi Zhang, Yuantong Zhang, Zhenzhong Chen 0001
PCS1
2025 Mamba-based Deep Reference Frame Generation for Inter Prediction Enhancement in NNVC
Wenzhuo Ma, Nianxiang Fu, Junxi Zhang, Zhenzhong Chen 0001
PCS2
2025 DiffVC-OSD: One-Step Diffusion-based Perceptual Neural Video Compression Framework
abstract
In this work, we first propose DiffVC-OSD, a One-Step Diffusion-based Perceptual Neural Video Compression framework. Unlike conventional multi-step diffusion-based methods, DiffVC-OSD feeds the reconstructed latent representation directly into a One-Step Diffusion Model, enhancing perceptual quality through a single diffusion step guided by both temporal context and the latent itself. To better leverage temporal dependencies, we design a Temporal Context Adapter that encodes conditional inputs into multi-level features, offering more fine-grained guidance for the Denoising Unet. Additionally, we employ an End-to-End Finetuning strategy to improve overall compression performance. Extensive experiments demonstrate that DiffVC-OSD achieves state-of-the-art perceptual compression performance, offers about 20× faster decoding and an 86.92% bitrate reduction compared to the corresponding multistep diffusion-based variant.
Wenzhuo Ma, Zhenzhong Chen 0001
VCIP1
2025 Diffusion-based Perceptual Neural Video Compression with Temporal Diffusion Information Reuse
abstract
Recently, foundational diffusion models have attracted considerable attention in image compression tasks, whereas their application to video compression remains largely unexplored. In this article, we introduce DiffVC, a diffusion-based perceptual neural video compression framework that effectively integrates foundational diffusion model with the video conditional coding paradigm. This framework uses temporal context from previously decoded frame and the reconstructed latent representation of the current frame to guide the diffusion model in generating high-quality results. To accelerate the iterative inference process of diffusion model, we propose the Temporal Diffusion Information Reuse (TDIR) strategy, which significantly enhances inference efficiency with minimal performance loss by reusing the diffusion information from previous frames. Additionally, to address the challenges posed by distortion differences across various bitrates, we propose the Quantization Parameter-based Prompting (QPP) mechanism, which utilizes quantization parameters as prompts fed into the foundational diffusion model to explicitly modulate intermediate features, thereby enabling a robust variable bitrate diffusion-based neural compression framework. Experimental results demonstrate that our proposed solution delivers excellent performance in both perception metrics and visual quality.
Wenzhuo Ma, Zhenzhong Chen 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Cascade Attention Blend Residual Network For Single Image Super-Resolution
abstract
Nowadays, deep convolutional neural networks are playing an increasingly important role in single-image super-resolution vision applications. Yet, most of the existing deep convolution-based methodologies are insufficiently intelligent to capture targeted information when the distribution of spatial and channel information is uneven for low-resolution images. To address this research issue, we propose a cascade attention blend residual network, with the non-local channel and multi-scale attention being considered for channel-wise dependencies and multi-scale receptive fields, respectively. Cascading both attentions in a potent blend residual block aims to learn more spatial and channel correlations between low- and super-resolution images. Experimental results demonstrate that the proposed method achieves promising performance for super-resolution image reconstruction, as well as gains an average reduction of 50.9% network parameters, compared to some state-of-the-art methods.
Guoqiang Xiao 0001, Xiaoqin Tang, Xian-Feng Han, Wenzhuo Ma, Xinye Gou
ICIP5
2021 Two-Phase Feature Fusion Network For Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a challenging problem that aims to match pedestrians captured by visible and infrared cameras. Prevailing methods in this field mainly focus on learning sharable feature representations from the last layer of deep convolution neural networks(CNNs). However, due to the large intra-modality variations and cross-modality variations, the last layer’s sharable feature representations are less discriminative. To remedy this, we propose a novel Two-Phase Feature Fusion Network(TFFN) to enhance the discriminative feature learning via feature fusion. Specifically, TFFN contains two fusion modules: (1) Multi-Level Fusion Module(MLFM) that reweights and fuses intra-modality multi-level features to utilize high- and low-level information; (2) Graph-Level Fusion Module (GLFM) that mines and fuses rich mutual information across the two modalities by employing a cross-modality graph attention network. Additionally, for effective fusion, we develop a deep supervision method to enhance the discrimination of pre-fusion features and eliminate noise information. Extensive experiments show that TFFN outperforms the state-of-the-art methods on two mainstream VI-ReID datasets: SYSU-MM01 and RegDB.
Yunzhou Cheng, Guoqiang Xiao 0001, Xiaoqin Tang, Wenzhuo Ma, Xinye Gou
ICIP4
2021 PGMANet: Pose-Guided Mixed Attention Network for Occluded Person Re-Identification
abstract
Recently, the person re-identification task becomes increasingly crucial in crowded scenarios, e.g., airports and schools. Many methods with high performance have been proposed to solve this problem. However, the existence of occlusion still challenges the development of person re-identification. In this work, we present the novel Pose - Guided Mixed Attention Network (PGMANet), an end-to-end framework to deal with pedestrian reidentification under occluded situations by fusing posture and second-order information. Specially, we employ two models. The first is Human Part - level Attention Model. We use key point information of a pedestrian to generate a heat map to enhance the pedestrian body part's feature. Simultaneously, we design Second - order Information Attention Model to investigate the correlation among features of different parts. Experimental results show that our method achieves state-of-the-art person re-identification performance on two challenging occlusion datasets Occluded-DukeMTMC and Occluded-Reid.
You Zhai, Xian-Feng Han, Wenzhuo Ma, Xinye Gou, Guoqiang Xiao 0001
IJCNN3
2021 Dual-Path Deep Supervision Network with Self-Attention for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification(VI-RelD) is an emerging but challenging problem that aims to match pedestrians captured by visible and infrared cameras. Existing studies in this field mainly focus on learning sharable feature representations from the last layer of deep convolution neural networks (CNNs) to handle the cross-modality discrepancies. However, due to the huge differences between visible and infrared images, the last layer's feature representations are less discriminative for VI-ReID. To remedy this, we propose a novel deep supervision learning network, namely Dual-path Deep Supervision Network (DDSN), for VI-ReID. Based on the backbone network, DDSN consists of two key modules, (1) a dual-path deep supervision learning (DDSL) module that is plugged into multiple network layers, and (2) a self-attention module that is developed on top of the backbone network. The backbone network extracts multilevel features at lower middle layers, and several DDSL modules utilize these features to generate more discriminative descriptors. Furthermore, we apply the self-attention module for context modeling to capture useful contextual cues as a supplement. By fusing these descriptors, DDSN can utilize both multi-level information and potential contextual information. Despite the apparent simplification, our method outperforms several state-of- the-art methods on two large-scale datasets: RegDB and SYSU-MM01.
Yunzhou Cheng, Guoqiang Xiao 0001, Wenzhuo Ma, Xinye Gou
ISCAS4