VLDB 2026 Research / reviewers in the wild / expert
Zhuowei Xu
dblp:302/9621
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-2222-158XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLIS: Causality-inspired Longitudinal Image Synthesis and its application to Alzheimer's disease characterization
Zhuowei Xu, Shaohua Kevin Zhou |
Medical Image Anal. | 3 |
| 2025 | NeRF-Based CBCT Reconstruction Needs Normalization and Initialization
Zhuowei Xu, Dai Sun, Qingpeng Kong, Nassir Navab, Shaohua Kevin Zhou |
MICCAI (16) | 1 |
| 2025 | Low-rank Content Adaptation for Neural Network-based Video Coding beyond VVC
Zhuowei Xu, Jacek Konieczny, Alexey Filippov, Christopher Hollmann, Vasily Rufitskiy, Tianyu Dong |
PCS | 1 |
| 2025 | Blind Multimodal Quality Assessment of Low-Light Images
Miaohui Wang, Zhuowei Xu, Mai Xu, Weisi Lin |
Int. J. Comput. Vis. | 2 |
| 2025 | Visual Quality Assessment of Composite Images: A Compression-Oriented Database and MeasurementabstractComposite images (CIs) have experienced unprecedented growth, especially with the prosperity of a large number of generative AI technologies. They are usually created by combining multiple visual elements from different sources to form a single cohesive composition, which have an increasing impact on a variety of vision applications. However, transmission of CIs can degrade their visual quality, especially undergoing lossy compression to reduce bandwidth and storage. To facilitate the development of objective measurements for CIs and investigate the influence of compression distortions on their perception, we establish a compression-oriented image quality assessment (CIQA) database for CIs (called ciCIQA) with 30 typical encoding distortions. Compressed with six representative codecs, we have carried out a large-scale subjective experiment that delivered 3,000 encoded CIs with labeled quality scores, making ciCIQA one of the earliest CI databases with the most compression types. ciCIQA enables us to explore the encoding effects on visual quality from the first five just noticeable difference (JND) points, offering insights for perceptual CI compression and related tasks. Moreover, we have proposed a new multi-masked no-reference CIQA method(called mmCIQA), including a multi-masked quality representation module, a self-supervised quality alignment module, and a multi-masked attentive fusion module. Experimental results demonstrate the outstanding performance of our mmCIQA in assessing the quality of CIs, outperforming 17 competitive approaches. The proposed method and database as well as the collected objective metrics are made publicly available on https://charwill.github.io/mmciqa.html. Miaohui Wang, Zhuowei Xu, Yuming Fang 0001, Weisi Lin |
IEEE Trans. Image Process. | 2 |
| 2024 | ReferPose: Distance Optimization-Based Reference Learning for Human Pose Estimation and MonitoringabstractExisting deep learning models for human pose estimation (HPE) have shown satisfactory performance in monitoring human actions. However, they usually face a dilemma between complexity and accuracy. To address this challenge, we propose an effective reference learning method for HPE (namely ReferPose), which is based on a new distance optimization strategy. Specifically, we utilize a reference model for pose learning and representation. The pose representation learned from the entire database is merged into the reference model, providing continuous reference learning guidance for an in-training model. In addition, we design a new cosine annealing-based reference guidance for temporal denoising and further develop a distance optimization strategy to provide joint guidance from pose knowledge, model representation, and temporal experience. Experimental results on two benchmark databases and a human fall monitoring system demonstrate that our ReferPose not only achieves promising accuracy improvement compared with several representative HPE models, but also offers low cost and high efficiency. Miaohui Wang, Zhuowei Xu, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | BinaryFormer: A Hierarchical-Adaptive Binary Vision Transformer (ViT) for Efficient ComputingabstractVision Transformer (ViT) has recently demonstrated impressive nonlinear modeling capabilities and achieved state-of-the-art performance in various industrial applications, such as object recognition, anomaly detection, and robot control. However, their practical deployment can be hindered by high storage requirements and computational intensity. To alleviate these challenges, we propose a binary transformer called BinaryFormer, which quantizes the learned weights of the ViT module from 32-b precision to 1 b. Furthermore, we propose a hierarchical-adaptive architecture that replaces expensive matrix operations with more affordable addition and bit operations by switching between two attention modes. As a result, BinaryFormer is able to effectively compress the model size as well as reduce the computation cost of ViT. Experimental results on the ImageNet-1K benchmark datasets show that BinaryFormer reduces the size of a typical ViT model by an average of 27.7× and converts over 99% of multiplication operations into bit operations while maintaining reasonable accuracy. Miaohui Wang, Zhuowei Xu, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | S-CCR: Super-Complete Comparative Representation for Low-Light Image Quality Inference In-the-wildabstractWith the rapid development of weak-illumination imaging technology, low-light images have brought new challenges to quality of experience and service. However, developing a robust quality indicator for authentic low-light distortions in-the-wild remains a major challenge in practical quality control systems. In this paper, we develop a new super-complete comparative representation (S-CCR) for the region-level quality inference of low-light images. Specifically, we excavate the color, luminance, and detail quality evidence for the feature embedding guidance of comparative representation based on the human visual characteristics. Moreover, we decompose the inputs into a super-complete feature group so that the image quality of each region can be fully represented, which allows to preserve the distinctiveness, distinguishability, and consistency. Finally, we further establish a comparative domain alignment method, so that the comparative representation of an unseen image can be aligned with respect to the quality features of already-seen ones. Extensive experiments on the benchmark dataset validate the superiority of our S-CCR over 11 competing methods on authentic distortions. Miaohui Wang, Zhuowei Xu, Yuanhao Gong, Wuyuan Xie |
ACM Multimedia | 2 |
| 2022 | MSCI: A Multi-Source Composite Image Database for Compression Distortion Quality AssessmentabstractWith the rapid development of multi-sensor fusion technology in various industrial fields, many composite images closely related to human life have been produced. To meet the rapidly growing needs of various image-based applications, we have established the first multi-source composite image (MSCI) database for image quality assessment (IQA). Our MSCI database contains 80 reference images and 1600 distorted images, generated by four advanced compression standards with five distortion levels. In particular, these five distortion levels are determined based on the first five just noticeable difference (JND) levels. Moreover, we verify the IQA performance of some representative methods on our MSCI database. The experimental results show that the performance of the existing methods on the MSCI database needs to be further improved. Zhuowei Xu, Zhiheng Lin, Miaohui Wang |
VCIP | 2 |
| 2022 | Perceptually Quasi-Lossless Compression of Screen Content Data Via Visibility Modeling and Deep ForecastingabstractScreen content data, such as computer-generated photographs, desktop sharing, remote education, video game streaming and screenshot, is one of the most popular visual information carriers in Internet of Video Things. Although lossless compression can guarantee high quality of service for these screen content based industrial applications, it also causes considerable storage space and transmission bandwidth issues. To alleviate these challenges, in this article, we present a visually quasi-lossless coding approach to control the compression distortion belowvisibility thresholdin the human visual system. Specifically, to better quantify the visual redundancy for screen content data, a newvisibility thresholdmethod is designed by incorporating blur sensitivity and oblique correction effects. Then, an end-to-end mapping between thevisibility thresholdand quality control factor is learned and represented as a deep convolutional neural network. The experimental results demonstrate that the proposed method saves the average encoding bits up to 23.15% compared with the latest scheme under the same perceptual quality. Miaohui Wang, Zhuowei Xu, Jian Xiong 0005, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Deep Human Pose Estimation via Self-guided Learning
Zhuowei Xu, Miaohui Wang |
ICIG (2) | 1 |