Wenhan Zhu

dblp:57/7648 · DBLP profile ↗
← Back
37ranked-venue papers
9as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 13 · 13 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 De-biasing Facial Albedo Estimation via Visual-Textual Cues
Xingyu Ren, Jiankang Deng, Chao Ma 0004, Yichao Yan, Wenhan Zhu, Xiaokang Yang 0001
J. Artif. Intell. Res.5
2026 Underwater image enhancement with frequency-spatial cross-domain transformer and hybrid collaborative representation
Jing Yang 0041, Wenhan Zhu, Fengling Jiang, Amir Hussain 0001
Knowl. Based Syst.2
2026 Topo4D++: Realistic Physically Based 4D Head Capture With Topology-Preserving Gaussian Splatting and Expression Priors
abstract
4D head capture aims to generate dynamic facial meshes in the same topology with corresponding UV maps, which requires temporal correspondence between 3D head models. Existing pipelines either involve manual processing of artists or employ constraints such as landmark tracking and optical flow, failing to achieve a trade-off between accuracy and efficiency. To enhance this process, we propose Topo4D++, a novel framework for automatic geometry and texture reconstruction that optimizes densely aligned 4D heads and 8 K BRDF maps directly from calibrated multi-view videos. Our key insight is to represent facial models as a set of dynamic 3D Gaussians with fixed topology, where the Gaussian centers are bound to the mesh vertices. This enables tracking all vertices rather than sparse vertices on the face accurately by leveraging the inverse rendering capabilities of 3D Gaussian Splatting (3DGS), while also enabling ultra-high-resolution texture generation. To maintain face structure during dynamic 3DGS optimization, we propose to optimize geometry and texture alternatively under physical and topological constraints frame-by-frame and employ blendshape-based expression priors to address extreme expressions. Then, we propose to extract dynamic facial meshes in a regular wiring arrangement and high-fidelity textures with pore-level details from the learned Gaussians. Finally, we train a diffusion-based model to generate BRDF texture maps to achieve physically based rendering. Given the absence of a universal benchmark, we construct JHead, a novel benchmark for the comprehensive evaluation of 4D head capture methods. Extensive experiments on different datasets demonstrate that our method is generalized to different capture systems, identities, and expressions, outperforming current state-of-the-art head reconstruction methods in both mesh and texture qualitatively and quantitatively.
Yuhao Cheng, Xuanchen Li, Xingyu Ren, Haozhe Jia, Di Xu 0012, Wenhan Zhu, Bingbing Ni, Yichao Yan
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 Towards High-fidelity 3D Talking Avatar with Personalized Dynamic Texture
abstract
Significant progress has been made for speech-driven 3D face animation, but most works focus on learning the motion of mesh/geometry, ignoring the impact of dynamic texture. In this work, we reveal that dynamic texture plays a key role in rendering high-fidelity talking avatars, and introduce a high-resolution 4D dataset TexTalk4D, consisting of 100 minutes of audio-synced scan-level meshes with detailed 8K dynamic textures from 100 subjects. Based on the dataset, we explore the inherent correlation between motion and texture, and propose a diffusion-based framework TexTalker to simultaneously generate facial motions and dynamic textures from speech. Furthermore, we propose a novel pivot-based style injection strategy to capture the complicity of different texture and motion styles, which allows disentangled control. TexTalker, as the first method to generate audio-synced facial motion with dynamic texture, not only outperforms the prior arts in synthesising facial motions, but also produces realistic textures that are consistent with the underlying facial movements. Project page: https://xuanchenli.github.io/TexTalk/.
Xuanchen Li, Yuhao Cheng, Yikun Zeng, Xingyu Ren, Wenhan Zhu, Weiming Zhao, Yichao Yan
CVPR6
2025 S^3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion Priors
abstract
Recent 3D face reconstruction methods have made remarkable advancements, yet achieving high-quality facial reflectance from monocular input remains challenging. Existing methods rely on the light-stage captured data to learn facial reflectance models. However, limited subject diversity in these datasets poses challenges in achieving good generalization and broad applicability. This motivates us to explore whether the extensive priors captured in recent generative diffusion models (e.g., Stable Diffusion) can enable more generalizable facial reflectance estimation as these models have been pre-trained on large-scale internet image collections containing rich visual patterns. In this paper, we introduce the use of Stable Diffusion as a prior for facial reflectance estimation, achieving robust results with minimal captured data for fine-tuning. We present S3-Face, a comprehensive framework capable of producing SSS-compliant skin reflectance from in-the-wild images. Our method adopts a two-stage training approach: in the first stage, DSN-Net is trained to predict diffuse albedo, specular albedo, and normal maps from in-the-wild images using a novel joint reflectance attention module. In the second stage, HM-Net is trained to generate hemoglobin and melanin maps based on the diffuse albedo predicted in the first stage, yielding SSS-compliant and detailed reflectance maps. Extensive experiments demonstrate that our method achieves strong generalization and produces high-fidelity, SSS-compliant facial reflectance estimation.
Xingyu Ren, Jiankang Deng, Yuhao Cheng, Wenhan Zhu, Yichao Yan, Xiaokang Yang 0001, Stefanos Zafeiriou, Chao Ma 0004
CVPR4
2025 Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation
abstract
Generating sewing patterns in garment design is receiving increasing attention due to its CG-friendly and flexible-editing nature. Previous sewing pattern generation methods have been able to produce exquisite clothing, but struggle to design complex garments with detailed control. To address these issues, we propose SewingLDM, a multi-modal generative model that generates sewing patterns controlled by text prompts, body shapes, and garment sketches. Initially, we extend the original vector of sewing patterns into a more comprehensive representation to cover more intricate details and then compress them into a compact latent space. To learn the sewing pattern distribution in the latent space, we design a two-step training strategy to inject the multi-modal conditions, \ie, body shapes, text prompts, and garment sketches, into a diffusion model, ensuring the generated garments are body-suited and detail-controlled. Comprehensive qualitative and quantitative experiments show the effectiveness of our proposed method, significantly surpassing previous approaches in terms of complex garment design and various body adaptability. Our project page: https://shengqiliu1.github.io/SewingLDM.
Shengqi Liu, Yuhao Cheng, Zhuo Chen 0060, Xingyu Ren, Wenhan Zhu, Lincheng Li, Mengxiao Bi, Xiaokang Yang 0001, Yichao Yan
ICCV5
2024 3D-Aware Face Editing via Warping-Guided Latent Direction Learning
abstract
3D facial editing, a longstanding task in computer vision with broad applications, is expected to fast and intuitively manipulate any face from arbitrary viewpoints following the user's will. Existing works have limitations in terms of intuitiveness, generalization, and efficiency. To overcome these challenges, we propose FaceEdit3D, which allows users to directly manipulate 3D points to edit a 3D face, achieving natural and rapid face editing. After one or several points are manipulated by users, we propose the tri-plane warping to directly deform the view-independent 3D representation. To address the problem of distortion caused by tri-plane warping, we train a warp-aware encoder to project the warped face onto a standardized latent space. In this space, we further propose directional latent editing to mitigate the identity bias caused by the encoder and realize the disentangled editing of various attributes. Extensive experiments show that our method achieves superior results with rich facial details and nice identity preservation. Our approach also supports general applications like multi-attribute continuous editing and cat/car editing. The project website is https://cyh-sj.github.io/FaceEdit3DI.
Yuhao Cheng, Zhuo Chen 0060, Xingyu Ren, Wenhan Zhu, Zhengqin Xu, Di Xu 0012, Changpeng Yang, Yichao Yan
CVPR4
2024 Monocular Identity-Conditioned Facial Reflectance Reconstruction
abstract
Recent 3D face reconstruction methods have made re-markable advancements, yet there remain huge challenges in monocular high-quality facial reflectance reconstruction. Existing methods rely on a large amount of light-stage captured data to learn facial reflectance models. However, the lack of subject diversity poses challenges in achieving good generalization and widespread applicability. In this paper, we learn the reflectance prior in image space rather than UV space and present a framework named ID2Reflectance. Our framework can directly estimate the reflectance maps of a single image while using limited reflectance data for training. Our key insight is that reflectance data shares facial structures with RGB faces, which enables obtaining expressive facial prior from inexpensive RGB data thus re-ducing the dependency on reflectance data. We first learn a high-quality prior for facial reflectance. Specifically, we pretrain multi-domain facial feature code books and design a codebook fusion method to align the reflectance and RGB domains. Then, we propose an identity-conditioned swapping module that injects facial identity from the target image into the pre-trained autoencoder to modify the identity of the source reflectance image. Finally, we stitch multi-view swapped reflectance images to obtain renderable assets. Extensive experiments demonstrate that our method exhibits excellent generalization capability and achieves state-of-the-art facial reflectance reconstruction results for in-the-wild faces. Our project page is https://xingyuren.github.io/id2reflectance.
Xingyu Ren, Jiankang Deng, Yuhao Cheng, Jia Guo 0003, Chao Ma 0004, Yichao Yan, Wenhan Zhu, Xiaokang Yang 0001
CVPR7
2024 ReGenNet: Towards Human Action-Reaction Synthesis
abstract
Humans constantly interact with their surrounding environments. Current human-centric generative models mainly focus on synthesizing humans plausibly interacting with static scenes and objects, while the dynamic human action-reaction synthesis for ubiquitous causal human-human interactions is less explored. Human-human interactions can be regarded as asymmetric with actors and reactors in atomic interaction periods. In this paper, we compre-hensively analyze the asymmetric, dynamic, synchronous, and detailed nature of human-human interactions and propose the first multi-setting human action-reaction synthe-sis benchmark to generate human reactions conditioned on given human actions. To begin with, we propose to an-notate the actor-reactor order of the interaction sequences for the NTU120, InterHuman, and Chi3D datasets. Based on them, a diffusion-based generative model with a Trans-former decoder architecture called ReGenNet together with an explicit distance-based interaction loss is proposed to predict human reactions in an online manner, where the future states of actors are unavailable to reactors. Quantitative and qualitative results show that our method can gener-ate instant and plausible human reactions compared to the baselines, and can generalize to unseen actor motions and viewpoint changes.
Liang Xu 0012, Yizhou Zhou, Yichao Yan, Xin Jin 0014, Wenhan Zhu, Fengyun Rao, Xiaokang Yang 0001, Wenjun Zeng 0001
CVPR5
2024 Topo4D: Topology-Preserving Gaussian Splatting for High-fidelity 4D Head Capture
Xuanchen Li, Yuhao Cheng, Xingyu Ren, Haozhe Jia, Di Xu 0012, Wenhan Zhu, Yichao Yan
ECCV (34)6
2024 HQ-Avatar: Towards High-Quality 3D Avatar Generation via Point-based Representation
abstract
Despite the flourishing of 3D object generation, generating high-quality digital avatars with detailed geometry and texture that are free to animate remains a challenging task. Existing avatar generation techniques often suffer from limitations such as low-quality geometry and blurry texture. Thus, we propose HQ-Avatar, a novel method for generating animatable avatars with high-quality geometry and texture. We enhance the geometry quality by proposing an importance sampling strategy and capturing intricate details through learned normal maps. To achieve high-quality texture, we present a neural point-based avatar representation, which enables high-resolution rendering results at a resolution of 10242, allowing detailed supervision. Extensive experiments on THuman2.0 dataset demonstrate the superiority of our method over state-of-the-art techniques in generating high-quality avatars. Furthermore, we show the applicability of our method by employing it as 3D priors to simplify the human avatar reconstruction process from scans or even single images. Code is available at https://github.com/olivia23333/HQ-Avatar
Weitian Zhang, Sijing Wu, Yichao Yan, Ben Xue, Wenhan Zhu, Xiaokang Yang 0001
ICME5
2024 Directional Texture Editing for 3D Models
abstract
Abstract Texture editing is a crucial task in 3D modelling that allows users to automatically manipulate the surface materials of 3D models. However, the inherent complexity of 3D models and the ambiguous text description lead to the challenge of this task. To tackle this challenge, we propose ITEM3D, a Texture Editing Model designed for automatic 3D object editing according to the text Instructions. Leveraging the diffusion models and the differentiable rendering, ITEM3D takes the rendered images as the bridge between text and 3D representation and further optimizes the disentangled texture and environment map. Previous methods adopted the absolute editing direction, namely score distillation sampling (SDS) as the optimization objective, which unfortunately results in noisy appearances and text inconsistencies. To solve the problem caused by the ambiguous text, we introduce a relative editing direction, an optimization objective defined by the noise difference between the source and target texts, to release the semantic ambiguity between the texts and images. Additionally, we gradually adjust the direction during optimization to further address the unexpected deviation in the texture domain. Qualitative and quantitative experiments show that our ITEM3D outperforms the state‐of‐the‐art methods on various 3D objects. We also perform text‐guided relighting to show explicit control over lighting. Our project page: https://shengqiliu1.github.io/ITEM3D/ .
Shengqi Liu, Zhuo Chen 0060, Jingnan Gao, Yichao Yan, Wenhan Zhu, Jiangjing Lyu, Xiaokang Yang 0001
Comput. Graph. Forum5
2024 What is an app store? The software engineering perspective
Wenhan Zhu, Sebastian Proksch 0001, Daniel M. Germán, Michael W. Godfrey, Li Li 0029, Shane McIntosh
Empir. Softw. Eng.1
2024 HyperStyle3D: Text-Guided 3D Portrait Stylization via Hypernetworks
abstract
Portrait stylization is a long-standing task enabling extensive applications. Although 2D-based methods have made great progress in recent years, real-world applications such as metaverse and games often demand 3D content. On the other hand, the requirement of 3D data, which is costly to acquire, significantly impedes the development of 3D portrait stylization methods. In this paper, inspired by the success of 3D-aware GANs that bridge 2D and 3D domains with 3D fields as the intermediate representation for rendering 2D images, we propose a novel method, dubbed HyperStyle3D, based on 3D-aware GANs for 3D portrait stylization. At the core of our method is a hyper-network learned to manipulate the parameters of the generator in a single forward pass. It not only offers a strong capacity to handle multiple styles with a single model, but also enables flexible fine-grained stylization that affects only texture, shape, or local part of the portrait. While the use of 3D-aware GANs bypasses the requirement of 3D data, we further alleviate the necessity of style images with the CLIP model being the style guidance. We conduct an extensive set of experiments across the style, attribute, and shape, and meanwhile, measure the 3D consistency. These experiments demonstrate the superior capability of our HyperStyle3D model in rendering 3D-consistent images in diverse styles, deforming the face shape, and editing various attributes.
Zhuo Chen 0060, Xudong Xu, Yichao Yan, Wenhan Zhu, Wayne Wu, Bo Dai 0002, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Head3D: Complete 3D Head Generation via Tri-plane Feature Distillation
abstract
Head generation with diverse identities is an important task in computer vision and computer graphics, widely used in multimedia applications. However, current full-head generation methods require a large number of three-dimensional (3D) scans or multi-view images to train the model, resulting in expensive data acquisition costs. To address this issue, we propose Head3D, a method to generate full 3D heads with limited multi-view images. Specifically, our approach first extracts facial priors represented by tri-planes learned in EG3D, a 3D-aware generative model, and then proposes feature distillation to deliver the 3D frontal faces within complete heads without compromising head integrity. To mitigate the domain gap between the face and head models, we present a dual-discriminator to guide the frontal and back head generation. Our model achieves cost-efficient and diverse complete head generation with photo-realistic renderings and high-quality geometry representations. Extensive experiments demonstrate the effectiveness of our proposed Head3D, both qualitatively and quantitatively.
Yuhao Cheng, Yichao Yan, Wenhan Zhu, Bowen Pan, Xiaokang Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2023 3D-Aware Face Swapping
abstract
Face swapping is an important research topic in computer vision with wide applications in entertainment and privacy protection. Existing methods directly learn to swap 2D facial images, taking no account of the geometric information of human faces. In the presence of large pose variance between the source and the target faces, there always exist undesirable artifacts on the swapped face. In this paper, we present a novel 3D-aware face swapping method that generates high-fidelity and multi-view-consistent swapped faces from single-view source and target images. To achieve this, we take advantage of the strong geometry and texture prior of 3D human faces, where the 2D faces are projected into the latent space of a 3D generative model. By disentangling the identity and attribute features in the latent space, we succeed in swapping faces in a 3D-aware manner, being robust to pose variations while transferring fine-grained facial details. Extensive experiments demonstrate the superiority of our 3D-aware face swapping framework in terms of visual quality, identity similarity, and multi-view consistency. Code is available at https://lyx0208.github.io/3dSwap.
Chao Ma 0004, Yichao Yan, Wenhan Zhu, Xiaokang Yang 0001
CVPR4
2023 GANHead: Towards Generative Animatable Neural Head Avatars
abstract
To bring digital avatars into people's lives, it is highly demanded to efficiently generate complete, realistic, and animatable head avatars. This task is challenging, and it is difficult for existing methods to satisfy all the requirements at once. To achieve these goals, we propose GANHead (Generative Animatable Neural Head Avatar), a novel generative head model that takes advantages of both the fine-grained control over the explicit expression parameters and the realistic rendering results of implicit representations. Specifically, GANHead represents coarse geometry, fine-gained details and texture via three networks in canonical space to obtain the ability to generate complete and realistic head avatars. To achieve flexible animation, we define the deformation filed by standard linear blend skinning (LBS), with the learned continuous pose and expression bases and LBS weights. This allows the avatars to be directly animated by FLAME [22] parameters and generalize well to unseen poses and expressions. Compared to state-of-the-art (SOTA) methods, GANHead achieves superior performance on head avatar generation and raw scan fitting.
Sijing Wu, Yichao Yan, Yuhao Cheng, Wenhan Zhu, Ke Gao 0012, Guangtao Zhai
CVPR5
2023 NeRF-IBVS: Visual Servo Based on NeRF for Visual Localization and Navigation
abstract
Visual localization is a fundamental task in computer vision and robotics. Training existing visual localization methods requires a large number of posed images to generalize to novel views, while state-of-the-art methods generally require dense ground truth 3D labels for supervision. However, acquiring a large number of posed images and dense 3D labels in the real world is challenging and costly. In this paper, we present a novel visual localization method that achieves accurate localization while using only a few posed images compared to other localization methods. To achieve this, we first use a few posed images with coarse pseudo-3D labels provided by NeRF to train a coordinate regression network. Then a coarse pose is estimated from the regression network with PNP. Finally, we use the image-based visual servo (IBVS) with the scene prior provided by NeRF for pose optimization. Furthermore, our method can provide effective navigation prior, which enable navigation based on IBVS without using custom markers and depth sensor. Extensive experiments on 7-Scenes and 12-Scenes datasets demonstrate that our method outperforms state-of-the-art methods under the same setting, with only 5\% to 25\% training data. Furthermore, our framework can be naturally extended to the visual navigation task based on IBVS, and its effectiveness is verified in simulation experiments.
Yuanze Wang, Yichao Yan, Dian-xi Shi, Wenhan Zhu, Jianqiang Xia, Jeff Tan, Songchang Jin, Ke Gao 0012, Xiaokang Yang 0001
NeurIPS4
2023 Image Quality Score Distribution Prediction via Alpha Stable Model
abstract
Based on potentially subjective and diverse image quality scores given by a group of subjects, we propose to predict the distribution of image quality scores rather than the mean opinion score (MOS) of image quality. Therefore, in this paper, we use an alpha stable model to parameterize the image quality score distribution (IQSD), and propose an objective method to predict the alpha-stable-model-based IQSD. First, the LIVE database is re-recorded. Specifically, we invite a large group of subjects (187 valid subjects) to evaluate the quality of all 808 images in the LIVE database, with their scores forming reliable IQSDs. All images in the LIVE database and their collected subjective quality scores form a new image quality assessment database, named the SJTU IQSD database. We then propose a framework and algorithm to predict the alpha-stable-model-based IQSD, in which quality features are extracted from the structural and natural statistical information of each image, and support vector regressors are trained to predict the alpha stable model parameters. Experiments carried out on the SJTU IQSD database verify the feasibility of using the alpha stable model to describe the IQSD, and the experimental results show that the alpha-stable-model-based IQSD can reflect a large amount of subjective information on image quality. We also prove that the objective alpha-stable-model-based IQSD prediction method is effective. The code and the SJTU IQSD database can be downloaded at ‘https://github.com/YixuanGao98/Image-Quality-Score-Distribution-Prediction-via-Alpha-Stable-Model.git’.
Xiongkuo Min, Wenhan Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.3
2022 A No-Reference Deep Learning Quality Assessment Method for Super-Resolution Images Based on Frequency Maps
abstract
To support the application scenarios where high-resolution (HR) images are urgently needed, various single image super-resolution (SISR) algorithms are developed. However, SISR is an ill-posed inverse problem, which may bring artifacts like texture shift, blur, etc. to the reconstructed images, thus it is necessary to evaluate the quality of super-resolution images (SRIs). Note that most existing image quality assessment (IQA) methods were developed for synthetically distorted images, which may not work for SRIs since their distortions are more diverse and complicated. Therefore, in this paper, we propose a no-reference deep-learning image quality assessment method based on frequency maps because the artifacts caused by SISR algorithms are quite sensitive to frequency information. Specifically, we first obtain the high-frequency map (HM) and low-frequency map (LM) of SRI by using Sobel operator and piecewise smooth image approximation. Then, a two-stream network is employed to extract the quality-aware features of both frequency maps. Finally, the features are regressed into a single quality value using fully connected layers. The experimental results show that our method outperforms all compared IQA models on the selected three super-resolution quality assessment (SRQA) databases.
Wei Sun 0029, Xiongkuo Min, Wenhan Zhu, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai
ISCAS4
2022 An empirical study of question discussions on Stack Overflow
Wenhan Zhu, Haoxiang Zhang 0001, Ahmed E. Hassan, Michael W. Godfrey
Empir. Softw. Eng.1
2021 Perceptual Quality Assessment for Recognizing True and Pseudo 4k Content
abstract
To meet the imperative demand for monitoring the quality of Ultra High-Definition (UHD) content in multimedia industries, we propose an efficient no-reference (NR) image quality assessment (IQA) metric to distinguish original and pseudo 4K contents and measure the quality of their quality in this paper. First, we establish a database including more than 3000 4K images composed of natural 4K images together with upscaled versions interpolated from 1080p and 720p images by fourteen algorithms. To improve computing efficiency, our model segments the input image and selects three representative patches by local variances. Then, we extract the histogram features and cut-off frequency features in the frequency domain as well as the natural scenes statistic (NSS) based features from the representative patches. Finally, we employ support vector regressor (SVR) to aggregate these extracted features as an overall quality metric to predict the quality score of the target image. Extensive experimental comparisons using seven common evaluation indicators demonstrate that the proposed model outperforms the competitive NR IQA methods and has a great ability to distinguish true and pseudo 4K images.
Wenhan Zhu, Guangtao Zhai, Xiongkuo Min, Xiaokang Yang 0001, Xiao-Ping Zhang 0002
ICASSP1
2021 Modeling Image Quality Score Distribution Using Alpha Stable Model
abstract
In recent years, image quality is generally described by a mean opinion score (MOS). However, we observe that an image’s quality ratings given by a group of subjects may not follow a Gaussian distribution and the image quality can not be fully described by a MOS. In this paper, we propose to describe the image quality using a parameterized distribution rather than a MOS, and an objective method is also proposed to predict the image quality score distribution (IQSD). Specifically, we selected 100 images from the LIVE database and invited a large group of subjects to evaluate the quality of these images. By analyzing the subjective quality ratings, we find that the IQSD can be well modeled by an alpha stable model and this model can reflect much more information than MOS. Therefore, we propose an algorithm to model the IQSD described by an alpha stable model. Features are extracted from images based on natural scene statistics and support vector regressors are trained to predict the IQSD described by an alpha stable model. We validate the proposed IQSD prediction model on the collected subjective quality ratings. Experimental results verify the effectiveness of the proposed algorithm in modeling the IQSD.
Xiongkuo Min, Wenhan Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai
ICIP3
2021 A No-Reference Evaluation Metric for Low-Light Image Enhancement
abstract
Low-light images, which are usually taken in dark or back-lighting conditions, are hard to perceive due to the low visibility and low contrast. To improve viewers’ Quality of Experience (QoE) and support the application of vision-based systems, various low-light image enhancement algorithms (LIEAs) have been proposed to lighten low-light images. However, some LIEAs may amplify the hidden distortions in the dark like noise and even further, introduce new distortions such as structural damage, color shift, etc, which severely affect the quality of light-enhanced images and need to be evaluated quantificationally. However, in the literature, few measures are proposed to assess the quality of light-enhanced images. Therefore, in this paper, we develop a no-reference low-light image enhancement evaluation (NLIEE) metric to predict the quality of light-enhanced images. The image quality is mainly assessed from four key aspects: light enhancement, color comparison, noise measurement, and structure evaluation. The experiment results show that NLIEE achieves the best performance among the general no-reference image quality assessment (NR IQA) models and quality descriptors for light enhancement.
Wei Sun 0029, Xiongkuo Min, Wenhan Zhu, Tao Wang 0078, Wei Lu 0021, Guangtao Zhai
ICME4
2021 Mea culpa: How developers fix their own simple bugs differently from other developers
abstract
In this work, we study how the authorship of code affects bug-fixing commits using the SStuBs dataset, a collection of single-statement bug fix changes in popular Java Maven projects. More specifically, we study the differences in characteristics between simple bug fixes by the original author - that is, the developer who submitted the bug-inducing commit - and by different developers (i.e., non-authors). Our study shows that nearly half (i.e., 44.3%) of simple bugs are fixed by a different developer. We found that bug fixes by the original author and by different developers differed qualitatively and quantitatively. We observed that bug-fixing time by authors is much shorter than that of other developers. We also found that bug-fixing commits by authors tended to be larger in size and scope, and address multiple issues, whereas bug-fixing commits by other developers tended to be smaller and more focused on the bug itself. Future research can further study the different patterns in bug-fixing and create more tailored tools based on the developer's needs.
Wenhan Zhu, Michael W. Godfrey
MSR1
2020 Automatic Region Selection For Objective Sharpness Assessment Of Mobile Device Photos
abstract
Mobile devices are the source of a vast majority of digital photos today. Photos taken by mobile devices generally have fairly good visual quality. When evaluating high-quality mobile device photos, people have to manually zoom in to local regions to discern the subtle difference. Understandably, a global objective quality assessment method cannot perform well on such task. Therefore, local region selection is widely recognized as a prerequisite for the following quality evaluation. Clearly, subjective regions selection suffers from the drawbacks in terms of productivity, reproducibility and optimality. In this paper, we propose an automatic local region selection algorithm for sharpness measurement of mobile device photos. Specifically, local texture statistics, depth, saliency, as well as inter-pictures difference, are used as main features to select an optimal local region, in which the sharpness is then measured. For validation, we have built a largescale database for sharpness evaluation of mobile device photos, with 100 different scenes shot by several flagship mobile phones. The experimental results show that the performance of classic sharpness evaluation algorithms can be substantially improved with the region selected by the proposed algorithm.
Guangtao Zhai, Wenhan Zhu, Yucheng Zhu, Xiongkuo Min, Xiao-Ping Zhang 0002, Hua Yang 0001
ICIP3
2020 A Multiple Attributes Image Quality Database for Smartphone Camera Photo Quality Assessment
abstract
Smartphone is the superstar product in digital device market and the quality of smartphone camera photos (SCPs) is becoming one of the dominant considerations when consumers purchase smartphones. How to evaluate the quality of smartphone cameras and the taken photos is urgent issue to be solved. To bridge the gap between academic research accomplishment and industrial needs, in this paper, we establish a new Smartphone Camera Photo Quality Database (SCPQD2020) including 1800 images with 120 scenes taken by 15 smartphones. Exposure, color, noise and texture which are four dominant factors influencing the quality of SCP are evaluated in the subjective study, respectively. Ten popular no-reference (NR) image quality assessment (IQA) algorithms are tested and analyzed on our database. Experimental results demonstrate that the current objective models are not suitable for SCPs, and quality metrics having high correlation with human visual perception are highly needed.
Wenhan Zhu, Guangtao Zhai, Zongxi Han, Xiongkuo Min, Tao Wang 0078, Xiaokang Yang 0001
ICIP1
2019 Multi-Channel Decomposition in Tandem With Free-Energy Principle for Reduced-Reference Image Quality Assessment
abstract
The visual quality of perceptions is highly correlated with the mechanisms of the human brain and visual system. Recently, the free-energy principle, which has been widely researched in brain theory and neuroscience, is introduced to quantize the perception, action, and learning in human brain. In the field of image quality assessment (IQA), on one hand, the free-energy principle can resort to the internal generative model to simulate the visual stimulus of the human beings. On the other hand, abundant psychological and neurobiological studies reveal that different frequency and orientation components of one visual stimulus arouse different neurons in the striate cortex, and the striate cortex processes visual information in the cerebral cortex. Motivated by these two aspects, a novel reduce-reference IQA metric called the multi-channel free-energy based reduced-reference quality metric is proposed in this paper. First, a two-level discrete Haar wavelet transform is used to decompose the input reference and distorted images. Next, to simulate the generative model in the human brain, the sparse representation is leveraged to extract the free-energy-based features in subband images. Finally, the overall quality metric is obtained through the support vector regressor. Extensive experimental comparisons on four benchmark image quality databases (LIVE, CSIQ, TID2008, and TID2013) demonstrate that the proposed method is highly competitive with the representative reduced-reference and classical full-reference models.
Wenhan Zhu, Guangtao Zhai, Xiongkuo Min, Menghan Hu, Jing Liu 0002, Guodong Guo, Xiaokang Yang 0001
IEEE Trans. Multim.1
2018 Modeling Thermal Sequence Signal Decreasing for Dual Modal Password Breaking
abstract
The thermal camera records trace of users' touch a while after they type in the password. People's password may be stolen through a thermal camera due to this phenomenon. In this paper, we model the procedure of the thermal sequence from the view point of physical process. Based on the Newton's Law of Cooling, we set up a physical model to describe the process of the keys' temperature decreasing. Then the model is used to estimate keystroke time instants by maximizing likelihood function method. We estimate the password after getting keystroke time instants of each key. The possibility that cheaters have to steal people's password is also explored in the experiment. Based on our findings in the experiment, we give several pieces of practical advice for people to protect their password.
Xiao-Ping Zhang 0002, Guangtao Zhai, Xiaokang Yang 0001, Wenhan Zhu, Xiao Gu 0001
ICIP5
2018 An Image Augmentation Method for Quality Assessment Database
abstract
Image databases for quality assessment are helpful to evaluate the performance of objective assessment methods. Recommendations in regard to the constitution of databases and experimental methods of the subjective assessment have been proposed to ensure the database a good ground truth for the validation of objective quality assessment methods. However, these restrictions make databases scale-limited by covering small number of scenes distorted by few levels. To enrich IQA databases and increase the generalization capability of IQA models, we devise an effective image augmentation method. The two-stages scheme consists of the image-label pairs generation by minimizing the free energy between the pristine image and its augmentation as well as the distortion level interpolation which is based on the monotonicity of the perceptual quality with the severity of distortion. The experimental results show the ability of the augmented database to improve the prediction accuracy of learning-based no-reference image quality assessment metrics which in turn demonstrates the effectiveness of our method.
Yucheng Zhu, Guangtao Zhai, Wenhan Zhu, Jiantao Zhou 0001
ISCAS3
2018 A Large-Scale Compressed 360-Degree Spherical Image Database: From Subjective Quality Evaluation to Objective Model Comparison
abstract
360-degree images/videos have been dramatically increasing in recent years. But the high resolution makes it difficult to be transported, compressed and stored, and thus constrains the development of 360-degree images/videos. Therefore, it is important to study how popular coding technologies influence the quality of 360-degree images. In this paper, we present a study on subjective assessment of compressed 360-degree images and investigate whether existing objective image quality assessment (IQA) methods can effectively evaluate the quality of compressed 360-degree images. We first construct the largest compressed 360-degree image database (CVIQD2018) including 16 source images and 528 compressed ones with three prevailing coding technologies. Then, we implement 16 full reference (FR) IQA metrics, which include 10 traditional IQA metrics for 2D images and 3 PSNR-based metrics for 360-degree images, as well as 5 no reference (NR) IQA metrics and calculate the correlation between each above metric and subjective assessment in terms of three commonly used performance indices. The experiment results reveal structure information, visual saliency information and compensation for geometric distortion are crucial for evaluating the quality of compressed 360-degree images.
Wei Sun 0029, Ke Gu 0001, Siwei Ma 0001, Wenhan Zhu, Guangtao Zhai
MMSP4
2018 Reduced-Reference Image Quality Assessment Based on Free-Energy Principle with Multi-Channel Decomposition
abstract
The free-energy principle studied in brain theory and neuroscience accounts for the mechanism of perception and understanding in human brain, which is highly adapted for measuring the visual quality of perceptions. On the other hand, psychologists and neurologists report that different frequency and orientation components of one stimulus arouse different neurons in striate cortex. In this paper, a novel reduce-reference (RR) image quality assessment (IQA) metric based on free-energy principle in multi-channel is proposed, which is called MCFEM (Multi-Channel Free-Energy principle Metric). We first decompose the input reference image and distorted image via a two-level discrete Haar wavelet transform (DHWT). Next, the free-energy features of each subband images are computed based on sparse representation. Finally, an overall quality index is received through the support vector regressor (SVR). Extensive experimental comparisons on four (LIVE, CSIQ, TID2008 and TID2013) benchmark image databases demonstrate that the proposed method is highly competitive with the representative RR and no-reference models as well as full-reference ones.
Wenhan Zhu, Guangtao Zhai, Yutao Liu 0002, Ning Lin, Xiaokang Yang 0001
MMSP1
2018 SIQD: Surveillance Image Quality Database and Performance Evaluation for Objective Algorithms
abstract
The surveillance camera is an important security device for police to maintain public order and provides clues to trace suspects. The quality of surveillance videos/images may be degraded during acquisition, compression and communication. It is difficult to acquire key information, such as human faces, with the bad quality surveillance videos/images. So the research of the quality assessment of surveillance images is quite necessary. In this paper, we perform a study on subjective quality evaluation of surveillance images and investigate whether the existing objective quality measures can be applied to the surveillance images. Concretely, we establish a new surveillance image quality database (SIQD) including 500 surveillance images with different degrees of quality through subjective study. Next, we investigate the prediction performance of the existing popular image quality assessment (IQA) algorithms on the SIQD database, which include eleven NR methods and four sharpness models. Experimental results demonstrate that the present objective models do not work well and quality measures having high correlation with human visual perception are high needed. Some meaningful findings are also given, which may enlighten the design of objective surveillance IQA models. The SIQD will be made publicly available to facilitate further surveillance IQA researches.
Wenhan Zhu, Guangtao Zhai, Xiaokang Yang 0001
VCIP1
2018 Arrow's Impossibility Theorem inspired subjective image quality assessment approach
Wenhan Zhu, Guangtao Zhai, Menghan Hu, Jing Liu 0002, Xiaokang Yang 0001
Signal Process.1
2017 No-reference quality assessment for JPEG compressed images
abstract
JPEG is a most commonly used standard of compression for digital images. Quality factor (Qfactor) for JPEG compressed image is actually a suitable indicator to the perceptual quality. However, the information of the compressor might be unknown due to various reasons. To evaluate the Qfactor, we recompress the formerly compressed image and measure the consistency between them. Then we define the fixed points (the points on the Qfactor-axis where the content of recompressed images are almost the same with that of directly compressed images) by following the Qfactor based specifications and form the image set. The quality of JPEG compressed images are measured by combining the estimated Qfactor with the features extracted from the image set. The experimental results confirm that the proposed image quality assessment technique, which is no-reference, is able to faithfully predict the visual quality of JPEG compressed images.
Yucheng Zhu, Guangtao Zhai, Ke Gu 0001, Wenhan Zhu
QoMEX4
2010 Bacillus anthracis genome organization in light of whole transcriptome sequencing
abstract
Emerging knowledge of whole prokaryotic transcriptomes could validate a number of theoretical concepts introduced in the early days of genomics. What are the rules connecting gene expression levels with sequence determinants such as quantitative scores of promoters and terminators? Are translation efficiency measures, e.g. codon adaptation index and RBS score related to gene expression? We used the whole transcriptome shotgun sequencing of a bacterial pathogen Bacillus anthracis to assess correlation of gene expression level with promoter, terminator and RBS scores, codon adaptation index, as well as with a new measure of gene translational efficiency, average translation speed. We compared computational predictions of operon topologies with the transcript borders inferred from RNA-Seq reads. Transcriptome mapping may also improve existing gene annotation. Upon assessment of accuracy of current annotation of protein-coding genes in the B. anthracis genome we have shown that the transcriptome data indicate existence of more than a hundred genes missing in the annotation though predicted by an ab initio gene finder. Interestingly, we observed that many pseudogenes possess not only a sequence with detectable coding potential but also promoters that maintain transcriptional activity.
Jeffrey Martin, Wenhan Zhu, Karla D. Passalacqua, Nicholas H. Bergman, Mark Borodovsky
BMC Bioinform.2
2009 Assessment of Gene Annotation Accuracy by Inferring Transcripts from RNA-Seq
abstract
Next generation sequencing is quickly changing long standing paradigms of genomics in terms of what is feasible to accomplish within a ldquoresearch life timerdquo and what is supposed to remain beyond limits of reliable experimental analysis. Sequencing and mapping of a prokaryote transcriptome can provide experimental validation for computationally predicted genes annotated in a prokaryotic genome. In this study we use the transcriptome sequencing to assess the accuracy of the public annotation of genes in a genome of important pathogen Bacillus anthracis. We show that the transcriptome data confirm existence of a number of genes missing in the annotation. Also, the transcriptome data indicate that pseudogenes, genes with frameshifts and in frame stop codons, do keep the transcriptional activity (while possessing a coding potential as well) and continue to present a challenge for accurate annotation.
Jeffrey Martin, Wenhan Zhu, Nicholas H. Bergman, Mark Borodovsky
BIBM2