Wenhao Shen

dblp:10/1683 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MPP-LIC: Mixed Precision Post-Training Quantization for Learned Image Compression
abstract
Learning-based Image Compression (LIC) models are hindered by high computational complexity and encoding-decoding mismatch issues across different hardware platforms. While Post-Training Quantization (PTQ) enables efficient deployment, standard fixed-bit approaches exhibit substantial quality degradation at low precision. To address these limitations, we propose MPP-LIC, an efficient mixed-precision PTQ approach for LIC. Our approach consists of four key steps as illustrated in Fig. 1: •Step 1: acquire initial sensitivity list. We leverage the Hessian trace to estimate each block's sensitivity for following quantization. •Step 2: obtain optimized sensitivity list based on task-constraint metric. We formulate the optimized sensitivity list using a task-constrained metric that considers both quantization error and rate-distortion performance. •Step 3: allocate mixed bit-width by the sensitivity list. Under the manually specified model size constraint, we assign mixed bit-widths to each block according to the sensitivity list obtained from Step 2. •Step 4: conduct task-constraint block-wise optimization. We perform blockwise optimization on mixed- precision model using minimal calibration images.
Yuefeng Zhang, Wenhao Shen
DCC2
2026 MultiGO++: Monocular 3D Clothed Human Reconstruction via Geometry-Texture Collaboration
abstract
Monocular 3D clothed human reconstruction aims to generate a complete and realistic textured 3D avatar from a single image. Existing methods are commonly trained under multi-view supervision with annotated geometric priors, and during inference, these priors are estimated by the pre-trained network from the monocular input. These methods are constrained by three key limitations: texturally by the unavailability of training data, geometrically by inaccurate external priors, and systematically by biased single-modality supervision, all leading to suboptimal reconstruction. To address these issues, we propose a novel reconstruction framework, named MultiGO++, which achieves effective systematic geometry-texture collaboration. It consists of three core parts: (1) a multi-source texture synthesis strategy that constructs more than 15,000 3D textured human scans to improve the performance of texture quality estimation in challenging scenarios; (2) a region-aware shape extraction module that extracts features and enables feature interactions from each body region to obtain geometry information and a Fourier geometry encoder that mitigates the modality gap to achieve effective geometry learning; (3) a dual reconstruction U-Net that leverages geometry-texture collaborative features to refine and generate high-fidelity textured 3D human meshes. Extensive experiments on two benchmarks and numerous in-the-wild cases show the superiority of our method over state-of-the-art approaches.
Nanjie Yao, Gangjian Zhang, Wenhao Shen, Jian Shu 0001, Hao Wang 0094
IEEE Trans. Vis. Comput. Graph.3
2025 Image Quality Assessment: Investigating Causal Perceptual Effects with Abductive Counterfactual Inference
abstract
Existing full-reference image quality assessment (FR-IQA) methods often fail to capture the complex causal mechanisms that underlie human perceptual responses to image distortions, limiting their ability to generalize across diverse scenarios. In this paper, we propose an FR-IQA method based on abductive counterfactual inference to investigate the causal relationships between deep network features and perceptual distortions. First, we explore the causal effects of deep features on perception and integrate causal reasoning with feature comparison, constructing a model that effectively handles complex distortion types across different IQA scenarios. Second, the analysis of the perceptual causal correlations of our proposed method is independent of the backbone architecture and thus can be applied to a variety of deep networks. Through abductive counterfactual experiments, we validate the proposed causal relationships, confirming the model’s superior perceptual relevance and interpretability of quality scores. The experimental results demonstrate the robustness and effectiveness of the method, providing competitive quality predictions across multiple benchmarks. The source code is available at https://anonymous.4open.science/r/DeepCausalQuality-25BC.
Wenhao Shen, Mingliang Zhou 0001, Xuekai Wei, Yong Feng 0002, Huayan Pu, Weijia Jia 0001
CVPR1
2025 SMPL Normal Map Is All You Need for Single-view Textured Human Reconstruction
abstract
Single-view textured human reconstruction aims to reconstruct a clothed 3D digital human by inputting a monocular 2D image. Existing approaches include feed-forward methods, limited by scarce 3D human data, and diffusion-based methods, prone to erroneous 2D hallucinations. To address these issues, we propose a novel SMPL normal map Equipped 3D Human Reconstruction (SEHR) framework, integrating a pretrained large 3D reconstruction model with human geometry prior. SEHR performs single-view human reconstruction without using a preset diffusion model in one forward propagation. Concretely, SEHR consists of two key components: SMPL Normal Map Guidance (SNMG) and SMPL Normal Map Constraint (SNMC). SNMG incorporates SMPL normal maps into an auxiliary network to provide improved body shape guidance. SNMC enhances invisible body parts by constraining the model to predict an extra SMPL normal Gaussians. Extensive experiments on two benchmark datasets demonstrate that SEHR outperforms existing state-of-the-art methods.
Wenhao Shen, Gangjian Zhang, Nanjie Yao, Xuanmeng Zhang, Hao Wang 0094
ICME1
2025 ADHMR: Aligning Diffusion-based Human Mesh Recovery via Direct Preference Optimization
abstract
Human mesh recovery (HMR) from a single image is inherently ill-posed due to depth ambiguity and occlusions. Probabilistic methods have tried to solve this by generating numerous plausible 3D human mesh predictions, but they often exhibit misalignment with 2D image observations and weak robustness to in-the-wild images. To address these issues, we propose ADHMR, a framework that Aligns a Diffusion-based HMR model in a preference optimization manner. First, we train a human mesh prediction assessment model, HMR-Scorer, capable of evaluating predictions even for in-the-wild images without 3D annotations. We then use HMR-Scorer to create a preference dataset, where each input image has a pair of winner and loser mesh predictions. This dataset is used to finetune the base model using direct preference optimization. Moreover, HMR-Scorer also helps improve existing HMR models by data cleaning, even with fewer training samples. Extensive experiments show that ADHMR outperforms current state-of-the-art methods. Code is available at: https://github.com/shenwenhao01/ADHMR.
Wenhao Shen, Wanqi Yin, Chaoyue Song, Zhongang Cai, Lei Yang 0045, Hao Wang 0094, Guosheng Lin
ICML1
2025 3D Cartoon Face Generation with Controllable Expressions from a Single GAN Image
abstract
In this paper, we investigate an open research task of generating 3D cartoon face shapes from single 2D GAN generated human faces and without 3D supervision, where we can also manipulate the facial expressions of the 3D shapes. To this end, we discover the semantic meanings of StyleGAN latent space, such that we are able to produce face images of various expressions, poses, and lighting conditions by controlling the latent codes. Specifically, we first finetune the pretrained StyleGAN face model on the cartoon datasets. By feeding the same latent codes to face and cartoon generation models, we aim to realize the translation from 2D human face images to cartoon styled avatars. We then discover semantic directions of the GAN latent space, in an attempt to change the facial expressions while preserving the original identity. As we do not have any 3D annotations for cartoon faces, we manipulate the latent codes to generate images with different poses and lighting conditions, such that we can reconstruct the 3D cartoon face shapes. We validate the efficacy of our method on three cartoon datasets qualitatively and quantitatively.
Hao Wang 0094, Wenhao Shen, Guosheng Lin, Steven C. H. Hoi, Chunyan Miao
IJCNN2
2025 Blind Image Quality Assessment: Exploring Content Fidelity Perceptibility via Quality Adversarial Learning
Mingliang Zhou 0001, Wenhao Shen, Xuekai Wei, Jun Luo 0003, Fan Jia 0005, Xu Zhuang, Weijia Jia 0001
Int. J. Comput. Vis.2
2024 HMR-Adapter: A Lightweight Adapter with Dual-Path Cross Augmentation for Expressive Human Mesh Recovery
abstract
Expressive Human Mesh Recovery (HMR) involves reconstructing the 3D human body, including hands and face, from RGB images. It is difficult because humans are highly deformable, and hands are small and frequently occluded. Recent approaches have attempted to mitigate these issues using large datasets and models, but these solutions remain imperfect. Specifically, whole-body estimation models often inaccurately estimate hand poses, while hand expert models struggle with severe occlusions. To overcome these limitations, we introduce a dual-path cross augmentation framework with a novel adaptation approach called HMR-Adapter that enhances existing large HMR models. HMR-Adapter significantly improves expressive HMR performance by injecting additional guidance from other body parts. This approach refines hand pose predictions by incorporating body pose information and uses additional hand features to enhance body pose estimation in whole-body models. Remarkably, an HMR-Adapter with about 30M parameters significantly improves expressive HMR results by combining the adapted large whole-body and hand expert models. We show extensive experiments and analysis to demonstrate the efficacy of our method.
Wenhao Shen, Wanqi Yin, Hao Wang 0094, Zhongang Cai, Lei Yang 0045, Guosheng Lin
ACM Multimedia1
2024 Graph-Represented Distribution Similarity Index for Full-Reference Image Quality Assessment
abstract
In this paper, we propose a graph-represented image distribution similarity (GRIDS) index for full-reference (FR) image quality assessment (IQA), which can measure the perceptual distance between distorted and reference images by assessing the disparities between their distribution patterns under a graph-based representation. First, we transform the input image into a graph-based representation, which is proven to be a versatile and effective choice for capturing visual perception features. This is achieved through the automatic generation of a vision graph from the given image content, leading to holistic perceptual associations for irregular image regions. Second, to reflect the perceived image distribution, we decompose the undirected graph into cliques and then calculate the product of the potential functions for the cliques to obtain the joint probability distribution of the undirected graph. Finally, we compare the distances between the graph feature distributions of the distorted and reference images at different stages; thus, we combine the distortion distribution measurements derived from different graph model depths to determine the perceived quality of the distorted images. The empirical results obtained from an extensive array of experiments underscore the competitive nature of our proposed method, which achieves performance on par with that of the state-of-the-art methods, demonstrating its exceptional predictive accuracy and ability to maintain consistent and monotonic behaviour in image quality prediction tasks. The source code is publicly available at the following website https://github.com/Land5cape/GRIDS.
Wenhao Shen, Mingliang Zhou 0001, Jun Luo 0003, Zhengguo Li, Sam Kwong
IEEE Trans. Image Process.1
2023 A Module for Enhancing Accuracy of Building Damage Detection by Fusing Features from Pre and Post Disaster Remote Sensing Images
abstract
In the aftermath of large-scale natural disasters, the accuracy of building damage detection (BDD) is of critical importance. Post-disaster high-resolution (post-HR) remote sensing imagery is fundamental for BDD; however, prompt acquisition of such imagery remains a significant challenge. To address this issue, we introduce a novel plug-and-play feature fusion (FF) module. This module, strategically situated between a pre-trained super-resolution (SR) model and a BDD model, ingeniously combines features from both pre-disaster high-resolution (pre-HR) and super-resolved post-disaster remote sensing imagery. The proposed approach is designed to maximize the utilization of pre-HR images, thereby enhancing BDD accuracy. Experimental validation confirms that this improvement in accuracy is attributable to the pragmatic extraction of features from pre-HR imagery, not just an increase in model complexity. Consequently, our approach holds substantial promise for real-world post-disaster scenarios and lays a solid foundation for future BDD research, demonstrating potential improvements in both efficacy and practicality.
Xuanchao Fu, Toru Kouyama, Wenhao Shen, Suomi Seki, Ryosuke Nakamura, Ichiro Yoshikawa
IGARSS3
2018 Surface Deformation Detection by Small Baseline Sar Interferometry in Cangzhou Coastal Zone
abstract
We collaborate with Department Of Land And Resources of Hebei Province and detect the Surface deformation by InSAR SBAS processing. SAR data covered the coast of Hebei Cangzhou (ENVISAT IS2 & RADARSAT-2 W1 beam,2003-2013) are acquired. Image interpretation and change detection of buildings base on middle and high resolution optical images, provides prior background knowledge for InSAR processing in this case. InSAR results and fieldwork show there is a strong relationship between the effect of industrial expansion and the location of land subsidence.
Jingfa Zhang, Wenhao Shen, Qisong Jiao, Liming Zuo
IGARSS4
2007 Stability Analysis of Particle Swarm Optimization
Jin-Xing Liu 0001, Huanbin Liu, Wenhao Shen
ICIC (2)3
2006 A Fuzzy PID Controller for Controlling Flotation De-inking Column
Jin-Xing Liu 0001, Huanbin Liu, Wenhao Shen, Yonggen Xu, Shuangchun Yang
ICIC (2)3