Wenhui Zhou 0001

dblp:58/2694-1 · DBLP profile ↗
← Back
38ranked-venue papers
16as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 10 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GA-LLMRec: Recommender Systems with Graph-Augmented Large Language Models
Ding Luo, Wenhui Zhou 0001, Zhengliang Ding, Guojun Dai
KSEM (3)2
2026 FCTL: Feature-level contrastive transfer learning for open set recognition
Wenhui Zhou 0001, Zhenglei Yang, Xinke Yang, Lili Lin, Ercan E. Kuruoglu
Comput. Vis. Image Underst.1
2026 AI-driven remanufacturing supply chains: Greening and intelligence diffusion
Lei Yang 0022, Wenhui Zhou 0001
Inf. Process. Manag.3
2026 Pseudo-4D-DCT representation learning for no-reference light field image quality assessment
Xuehui Wei, Feifan Wu, Wenhui Zhou 0001, Lili Lin
Image Vis. Comput.7
2026 EEG-driven natural image reconstruction with regional semantic awareness
Wenhui Zhou 0001, Yunrui Li, Guojun Dai, Lili Lin
Pattern Recognit.2
2026 Latent EEG-Vision Alignment for EEG-Driven 3D Object Reconstruction With Multi-View Stylistic Consistency
Wenhui Zhou 0001, Guojun Dai, Shanggui Zhan
IEEE Signal Process. Lett.2
2025 Electroencephalography-driven three-dimensional object decoding with multi-view perception diffusion
Wenhui Zhou 0001, Guojun Dai
Eng. Appl. Artif. Intell.2
2025 PG-VTON: Front-And-Back Garment Guided Panoramic Gaussian Virtual Try-On With Diffusion Modeling
abstract
ABSTRACT Virtual try‐on (VTON) technology enables the rapid creation of realistic try‐on experiences, which makes it highly valuable for the metaverse and e‐commerce. However, 2D VTON methods struggle to convey depth and immersion, while existing 3D methods require multi‐view garment images and face challenges in generating high‐fidelity garment textures. To address the aforementioned limitations, this paper proposes a panoramic Gaussian VTON framework guided solely by front‐and‐back garment information, named PG‐VTON, which uses an adapted local controllable diffusion model for generating virtual dressing effects in specific regions. Specifically, PG‐VTON adopts a coarse‐to‐fine architecture consisting of two stages. The coarse editing stage employs a local controllable diffusion model with a score distillation sampling (SDS) loss to generate coarse garment geometries with high‐level semantics. Meanwhile, the refinement stage applies the same diffusion model with a photometric loss not only to enhance garment details and reduce artifacts but also to correct unwanted noise and distortions introduced during the coarse stage, thereby effectively enhancing realism. To improve training efficiency, we further introduce a dynamic noise scheduling (DNS) strategy, which ensures stable training and high‐fidelity results. Experimental results demonstrate the superiority of our method, which achieves geometrically consistent and highly realistic 3D virtual try‐on generation.
Shengwei Sang, Guojun Dai, Xiaoyang Mao, Wenhui Zhou 0001
Comput. Animat. Virtual Worlds6
2025 Cascade residual learning based adaptive feature aggregation for light field super-resolution
Hao Zhang 0146, Wenhui Zhou 0001, Lili Lin, Andrew Lumsdaine
Pattern Recognit.2
2024 Multi-Target Multi-Camera Tracking based on lightweight detector
abstract
Multi-target multi-camera tracking (MTMCT) aims to associate the multiple targets in consecutive frames to obtain trajectories under multiple cameras. This paper proposes a practical MTMCT framework, which is mainly composed of a lightweight detector, ReID fusion module, tracker, and cross-camera matching module. We design a lightweight detector of gate recursive convolution to improve detection efficiency while ensuring detection accuracy. To accurately track the targets in low confidence under occlusion, we propose a feature compensation strategy to search and compensate corresponding target features in previous frames. To match targets between cameras, we utilize the temporal constraints of the region to filter out false positive trajectories and merge all candidate trajectories according to the similarity matrix. In the post-processing stage, we incorporate the Re-rank, k-mutual nearest neighbor, and the hierarchical clustering algorithm to generate the final tracking results. Experiments conducted on highway scenarios and city-flow datasets demonstrate the competitive performance of the proposed method.
Zhuozhen Xu, Wenhui Zhou 0001
CSCWD4
2024 LF-SAET: Cascaded Spatial-Angular-EPI Transformers for Light Field Image Super-Resolution
Hao Zhang 0146, Junle Yu, Jiahan Meng, Wenhui Zhou 0001
PRCV (9)5
2024 Temporal-channel cascaded transformer for imagined handwriting character recognition
Wenhui Zhou 0001, Liangyan Mo, Wanzeng Kong, Guojun Dai
Neurocomputing1
2024 Beyond Photometric Consistency: Geometry-Based Occlusion-Aware Unsupervised Light Field Disparity Estimation
abstract
Although learning-based light field disparity estimation has achieved great progress in the most recent years, the performance of unsupervised light field learning is still hindered by occlusions and noises. By analyzing the overall strategy underlying the unsupervised methodology and the light field geometry implied in epipolar plane images (EPIs), we look beyond the photometric consistency assumption, and design an occlusion-aware unsupervised framework to deal with the situations of photometric consistency conflict. Specifically, we present a geometry-based light field occlusion modeling, which predicts a group of visibility masks and occlusion maps, respectively, by forward warping and backward EPI-line tracing. In order to learn better the noise- and occlusion-invariant representations of the light field, we propose two occlusion-aware unsupervised losses: occlusion-aware SSIM and statistics-based EPI loss. Experiment results demonstrate that our method can improve the estimation accuracy of light field depth over the occluded and noisy regions, and preserve the occlusion boundaries better.
Wenhui Zhou 0001, Lili Lin, Yongjie Hong, Qiujian Li, Xingfa Shen, Ercan E. Kuruoglu
IEEE Trans. Neural Networks Learn. Syst.1
2023 PEAL: Prior-embedded Explicit Attention Learning for Low-overlap Point Cloud Registration
abstract
Learning distinctive point-wise features is critical for low-overlap point cloud registration. Recently, it has achieved huge success in incorporating Transformer into point cloud feature representation, which usually adopts a self-attention module to learn intra-point-cloud features first, then utilizes a cross-attention module to perform feature exchange between input point clouds. The advantage of Transformer models mainly benefits from the use of self-attention to capture the global correlations in feature space. However, these global correlations may involve ambiguity for point cloud registration task, especially in indoor low-overlap scenarios, because the correlations with an extensive range of non-overlapping points may degrade the feature distinctiveness. To address this issue, we present PEAL, a Prior-embedded Explicit Attention Learning model. By incorporating prior knowledge into the learning process, the points are divided into two parts. One includes points lying in the putative overlapping region and the other includes points located in the putative non-overlapping region. Then PEAL explicitly learns one-way attention with the putative overlapping points. This simplistic design attains surprising performance, significantly relieving the aforementioned feature ambiguity. Our method improves the Registration Recall by 6+% on the challenging 3DLoMatch benchmark and achieves state-of-the-art performance on Feature Matching Recall, Inlier Ratio, and Registration Recall on both 3DMatch and 3DLoMatch.
Junle Yu, Luwei Ren, Wenhui Zhou 0001, Yu Zhang 0280, Lili Lin, Guojun Dai
CVPR3
2023 DTBS: Dual-Teacher Bi-Directional Self-Training for Domain Adaptation in Nighttime Semantic Segmentation
abstract
Due to the poor illumination and the difficulty in annotating, nighttime conditions pose a significant challenge for autonomous vehicle perception systems. Unsupervised domain adaptation (UDA) has been widely applied to semantic segmentation on such images to adapt models from normal conditions to target nighttime-condition domains. Self-training (ST) is a paradigm in UDA, where a momentum teacher is utilized for pseudo-label prediction, but a confirmation bias issue exists. Because the one-directional knowledge transfer from a single teacher is insufficient to adapt to a large domain shift. To mitigate this issue, we propose to alleviate domain gap by incrementally considering style influence and illumination change. Therefore, we introduce a one-stage Dual-Teacher Bi-directional Self-training (DTBS) framework for smooth knowledge transfer and feedback. Based on two teacher models, we present a novel pipeline to respectively decouple style and illumination shift. In addition, we propose a new Re-weight exponential moving average (EMA) to merge the knowledge of style and illumination factors, and provide feedback to the student model. In this way, our method can be embedded in other UDA methods to enhance their performance. For example, the Cityscapes to ACDC night task yielded 53.8 mIoU (%), which corresponds to an improvement of +5% over the previous state-of-the-art. The code is available at https://github.com/hf618/DTBS.
Fanding Huang, Zihao Yao, Wenhui Zhou 0001
ECAI3
2023 DecoupledPoseNet: Cascade Decoupled Pose Learning for Unsupervised Camera Ego-Motion Estimation
abstract
Although many impressive works on learning-based camera ego-motion estimation methods have been proposed recently, most of them promote the accuracy of camera pose estimation by various sequential learning with loop closure optimization, while neglecting the improvement of PoseNet itself. In this paper, we focus on the coupling of rotation and translation in ego-motion estimation, and design a cascade decoupling structure to separately learn the rotation and translation of camera relative motion between adjacent frames. Meanwhile, a rigid-aware unsupervised learning framework with iterative pose refinement scheme is proposed for camera ego-motion estimation. It can disambiguate rigid motion and deformations in dynamic scenarios by jointly learning of optical flow, stereo disparity and camera pose. Validated with evaluation experiments on the public available datasets, our method is superior to the state-of-the-art unsupervised methods, and can achieve comparable results with the supervised ones.
Wenhui Zhou 0001, Hua Zhang 0011, Zhengmao Yan, Weisheng Wang, Lili Lin
IEEE Trans. Multim.1
2022 PCR-CG: Point Cloud Registration via Deep Explicit Color and Geometry
Yu Zhang 0280, Junle Yu, Xiaolin Huang, Wenhui Zhou 0001, Ji Hou
ECCV (10)4
2022 Unsupervised learning of light field depth estimation with spatial and angular consistencies
Lili Lin, Qiujian Li, Yuxiang Yan 0003, Wenhui Zhou 0001, Ercan E. Kuruoglu
Neurocomputing5
2022 Robust structural similarity index measure for images with non-Gaussian distortions
Lili Lin, Ercan E. Kuruoglu, Wenhui Zhou 0001
Pattern Recognit. Lett.4
2021 Robust dense light field reconstruction from sparse noisy sampling
Wenhui Zhou 0001, Jiangwei Shi, Yongjie Hong, Lili Lin, Ercan E. Kuruoglu
Signal Process.1
2020 Flexible Spatial and Angular Light Field Super Resolution
abstract
A light field contains information in four dimensions, two spatial and two angular. Representing a light field by sampling it with a fixed number of pixels implies an inherent trade-off between angular resolution and spatial resolution- one apparently fixed at the time of capture. To enable flexible trade-offs in spatial and angular resolution after the fact, in this paper we apply techniques from super resolution in an integrated fashion. Our approach explores the similarity between light field super resolution (LFSR) and single image super resolution (SISR) and proposes a neural network framework that can carry out flexible super resolution tasks. We present concrete instances of the framework for center-view spatial LFSR, full-view spatial LFSR, and combined spatial and angular LFSR. Experiments with synthetic and real-world data sets show the center-view and full-views approaches outperform state-of-the-art spatial LFSR by over 1dB in PSNR and that the combined approach achieves comparable performance to state-of-the-art spatial LFSR algorithms. Visual results for images rendered from the combined approach show improved resolution of detail, without rendering artifacts.
Dizhi Ma, Andrew Lumsdaine, Wenhui Zhou 0001
ICIP3
2020 Revisiting Rubik's Cube: Self-supervised Learning with Volume-Wise Transformation for 3D Medical Image Segmentation
Xing Tao, Yuexiang Li, Wenhui Zhou 0001, Kai Ma 0002, Yefeng Zheng 0001
MICCAI (4)3
2020 Depth-guided view synthesis for light field reconstruction from a single image
Wenhui Zhou 0001, Gaomin Liu, Jiangwei Shi, Hua Zhang 0011, Guojun Dai
Image Vis. Comput.1
2020 Unsupervised Monocular Depth Estimation From Light Field Image
abstract
Learning based depth estimation from light field has made significant progresses in recent years. However, most existing approaches are under the supervised framework, which requires vast quantities of ground-truth depth data for training. Furthermore, accurate depth maps of light field are hardly available except for a few synthetic datasets. In this paper, we exploit the multi-orientation epipolar geometry of light field and propose an unsupervised monocular depth estimation network. It predicts depth from the central view of light field without any ground-truth information. Inspired by the inherent depth cues and geometry constraints of light field, we then introduce three novel unsupervised loss functions: photometric loss, defocus loss and symmetry loss. We have evaluated our method on a public 4D light field synthetic dataset. As the first unsupervised method published in the 4D Light Field Benchmark website, our method can achieve satisfactory performance in most error metrics. Comparison experiments with two state-of-the-art unsupervised methods demonstrate the superiority of our method. We also prove the effectiveness and generality of our method on real-world light-field images.
Wenhui Zhou 0001, Enci Zhou, Gaomin Liu, Lili Lin, Andrew Lumsdaine
IEEE Trans. Image Process.1
2019 Learning Depth Cues from Focal Stack for Light Field Depth Estimation
abstract
Deep neural networks have shown their excellent abilities in light field depth estimation. Most of learning based approaches focus on the depth feature extraction from the epipolar plane images (EPIs) or sub-apertures of light field, while pay less attention to the focal stack which is also one of the most distinctive characteristics of light field. In this paper, we propose a FocalStackNet which learns depth semantic features and local structure information from the focal stack for light field depth estimation. Specifically, we formulate the disparity estimation as a pixel-wise classification task, and discretize the continuous disparity range into 115 bins. Then we generate a discrete focal stack and extract a set of focal stack patches as training data. Finally, we train a two-pathway convolutional neural networks (CNN) to predict the disparity label of each pixel. Evaluation experiments are carried on the public 4D light field synthetic dataset. Our method achieves state-of-the-art performance. It ranks first among the published methods on the aspects of average and median error scores of Bad Pixel Ratio 0.03.
Wenhui Zhou 0001, Enci Zhou, Yuxiang Yan 0003, Lili Lin, Andrew Lumsdaine
ICIP1
2019 N-Net: 3D Fully Convolution Network-Based Vertebrae Segmentation from CT Spinal Images
abstract
Accurate vertebrae segmentation from CT spinal images is crucial for the clinical tasks of diagnosis, surgical planning, and post-operative assessment. This paper describes an [Formula: see text]-shaped 3D fully convolution network (FCN) for vertebrae segmentation: [Formula: see text]-net. In this network, a global structure guidance pathway is designed for fusing the high-level semantic features with the global structure information. Moreover, the residual structure and the skip connection are introduced into traditional 3D FCN framework. These schemes can significantly improve the accuracy of vertebrae segmentation. Experimental results demonstrate the effectiveness and robustness of our method. A high average DICE score of 0.9499 [Formula: see text] 0.02 can be obtained, which is better than those of existing methods.
Wenhui Zhou 0001, Lili Lin, Guangtao Ge
Int. J. Pattern Recognit. Artif. Intell.1
2018 Scale and Orientation Aware EPI-Patch Learning for Light Field Depth Estimation
abstract
Epipolar Plane Image (EPI) implies some important depth cues for light field depth estimation. Intuitively, the EPI patches with different spatial scales and orientations may exhibit different features and result in different estimation precision. In this paper, we discuss this issue and present a scale and orientation aware EPI-Patch learning model for depth estimation. We take the multi-orientation EPI patches of each pixel as input, and design two types of network structures for adaptive scale selection and orientation fusion. One type is a scale-aware structure, which feeds one orientation patch into a multi-layer feed-forward network with long and short skip connections. The other type is a shared-weight network for fusing the multi-orientation features. We demonstrate the effectiveness of our model by experiments on 4D Light Field Benchmark.
Wenhui Zhou 0001, Linkai Liang, Hua Zhang 0011, Andrew Lumsdaine, Lili Lin
ICPR1
2018 Fast Depth Intra Mode Decision Based on DCT in 3D-HEVC
Renbin Yang, Guojun Dai, Hua Zhang 0011, Wenhui Zhou 0001, Shifang Yu, Jie Feng 0010
PRCV (1)4
2017 Light-field flow: A subpixel-accuracy depth flow estimation with geometric occlusion model from a single light-field image
abstract
Light-field cameras capture not only 2D images, but also the angles of the incoming light. These additional light angles bring the benefit of getting a sub-aperture image array from a single light-field image. Inspired by the traditional optical flow with occlusion detection, this paper focuses on the correlation analysis and the occlusion modeling for the sub-aperture array, and unifies them into a light-field flow framework. The main challenges faced are subpixel displacements and occlusion handling among the sub-aperture images. We build a light-field flow for joint depth estimation and occlusion detection, and develop a geometric occlusion model. More specifically, we firstly estimate subpixel-accuracy optical flows from each two sub-aperture images by the phase shift theorem, then a forward-backward consistency checking is adopted to detect the occluded regions. According to the geometric complementary character of occlusion in a light-field image, an occlusion filling strategy is proposed to refine depth estimation in the occluded regions. Experimental results on the synthetic scenes and Lytro Illum camera data both demonstrate the effectiveness and robustness of our method which has excellent performance in handling occlusions.
Wenhui Zhou 0001, Andrew Lumsdaine, Lili Lin
ICIP1
2017 EPI-Patch Based Convolutional Neural Network for Depth Estimation on 4D Light Field
Yaoxiang Luo, Wenhui Zhou 0001, Junpeng Fang, Linkai Liang, Hua Zhang 0011, Guojun Dai
ICONIP (3)2
2016 Depth estimation with cascade occlusion culling filter for light-field cameras
abstract
Depth recovery from a light-field camera is an essential and interesting problem. One of its most challenges is to get accurate estimation for the depth discontinuities and occluded regions. We propose a simple and efficient solution with a cascade occlusion culling filter. It is a cascade processing corresponding to the different manifestations of occlusions at ray-level, pixel-level and image-level. (i) At ray-level, any potential occluded ray will be filtered out in depth-cue responses computation, and then reliable multiple-cue cost volumes are constructed. (ii) At pixel-level, occlusions generally result in weak depth discontinuity ramp edges. These discontinuities will be preserved and enhanced by a multiple-cue cost-volume filter with edge preserving property. (iii) At image-level, occlusion is embodied in the uncertain regions. In order to obtain optimal depth estimation of uncertain regions, an iterative depth optimization framework is applied to integrate the aforementioned filtered multiple-cue cost volumes with their confidences. We show that our method has good depth-discontinuity preserving property, and is insensitive to the surface color / texture discontinuities at the same time. Experimental results on Lytro Illum camera data demonstrate the effectiveness and robustness of our method which has excellent performance in handling discontinuities and occlusions.
Wenhui Zhou 0001, Andrew Lumsdaine, Lili Lin
ICPR1
2014 Multi-scale contrast-based saliency enhancement for salient object detection
abstract
To achieve more complete and more uniformly highlighted salient object regions, this study presents a computational saliency enhancement model that incorporates the properties of multi‐scale and logarithmic response into the local and global contrasts. A distinct feature of the authors model is a novel saliency enhancement operator. This operator can effectively enhance the saliency of object interior regions while simultaneously reducing blur on object boundaries caused by multiple scales. Their model is a general one that can make flexible tradeoffs between precision and recall. Detailed comparisons with 12 state‐of‐the‐art methods show that their method can obtain satisfactory salient object regions that are closer to the human‐labelled results. In addition, their method provides superior results in precision–recall, F ‐measure and mean absolute error.
Wenhui Zhou 0001, Teng Song, Lili Lin, Andrew Lumsdaine
IET Comput. Vis.1
2013 Guided depth enhancement via a fast marching method
Xiaojin Gong, Wenhui Zhou 0001, Jilin Liu
Image Vis. Comput.3
2011 Weber's Law Based Center-Surround Hypothesis for Bottom-Up Saliency Detection
Lili Lin, Wenhui Zhou 0001, Hua Zhang 0011
ICONIP (3)2
2010 Combining dark channel prior and color cues for road following in outdoor environments
abstract
This paper proposes a road following method in outdoor environments based on dark channel prior and color cues, which integrates the atmospheric transmission estimation, mean shift filtering and graph cuts algorithm together. The main idea is fusing scene depth cue with color information in six-dimensional feature space for road image following. Specifically, the dark channel prior based transmission estimation is firstly employed to recover depth cue. Then the mean shift filtering in the weighted color-depth space is proposed. Finally, graph cuts algorithm is applied to achieve the final road detection. Experimental results indicate the proposed method has excellent performance in complex outdoor environments.
Wenhui Zhou 0001, Lili Lin, Xuehui Wei, Bin Lou
ICIP1
2010 Monocular depth cue fusion for image segmentation and grouping in outdoor navigation
abstract
This paper proposes an efficient fusion strategy of monocular depth cue and other image features for natural image segmentation and grouping. The main idea is to improve the performance of image clustering via fusing depth cue, color, spatial location, and edge confidence in six-dimensional color-depth feature space. It integrates the monocular depth cue estimation, mean shift filtering and graph cuts algorithm together. Firstly, the dark channel prior based atmospheric transmission estimation is employed to recover monocular depth cue. Then the mean shift filtering in the weighted color-depth space is proposed to obtain cluster regions with correct boundaries. Finally, graph cuts algorithm is applied to achieve the final regional grouping. Experimental results indicate the proposed method has excellent performance in outdoor natural environments.
Wenhui Zhou 0001, Lili Lin, Bin Lou, Xuehui Wei
IROS1
2007 A Robust and Adaptive Road Following Algorithm for Video Image Sequence
Lili Lin, Wenhui Zhou 0001
ICIC (1)2
2006 A Swarm Optimization Model for Energy Minimization Problem of Early Vision
Wenhui Zhou 0001, Lili Lin, Weikang Gu
ICONIP (2)1