Guoqing Zhou 0003

dblp:57/3780-3 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0003-3336-708XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Generalizable Occlusion-Aware Human Novel View Synthesis From Image Pairs
Kaijin Zhao, Xin Huang 0021, Guoqing Zhou 0003, Qing Wang 0006
IEEE Trans. Circuits Syst. Video Technol.3
2026 Zero-Pose-Prior NeRF: Recursive Radiance Field Reconstruction From Unposed and Unordered Images
abstract
The dependence of neural radiance fields (NeRF) on accurate camera poses has emerged as a critical obstacle to their widespread real-world applications. While recent advances have demonstrated the potential for simultaneously addressing camera registration and scene reconstruction, these methods inherently rely on reasonable initialization derived from pose or scene priors and struggle with complex scenes involving large camera motions, particularly in unordered 360-degree scenes. In this work, we propose Zero-Pose-Prior NeRF to recover radiance fields from unposed and unordered image collections without any prior knowledge. Our key insight is to decompose this complex problem into smaller sub-problems, wherein the sub-problems' camera poses are initially estimated to provide self-bootstrapping priors for the global pose estimation, followed by a recursive registration and reconstruction. To achieve this, we first perform scene partitioning to establish a hierarchical structure that describes registration order from local to global. Thereafter, we devise a conditionally-decoupled positional encoding for NeRFs, which serves as the basic model for camera pose estimation and scene representation. Following this, we develop a recursive registration to recursively estimate the poses of local scenes and register them into a unified global pose space, ultimately enabling the reconstruction of the entire scene. Experiments on real-world scenes show that our approach outperforms the state-of-the-art pose-free methods in terms of accurate camera poses and robust radiance field reconstruction, resulting in high-fidelity view synthesis.
Xinxin Liu 0020, Qi Zhang 0029, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006
IEEE Trans. Image Process.4
2025 Bright-NeRF: Brightening Neural Radiance Field with Color Restoration from Low-Light RAW Images
abstract
Neural Radiance Fields (NeRF) have demonstrated prominent performance in novel view synthesis tasks. However, their input heavily relies on image acquisition under normal light conditions, making it challenging to learn accurate scene contents in low-light environments where images typically exhibit significant noise and severe color distortion. To address these challenges, we propose a novel approach, Bright-NeRF, which learns enhanced and high-quality radiance fields from multi-view low-light RAW images in an unsupervised manner. Our method simultaneously achieves color restoration, denoising, and enhanced novel view synthesis. Specifically, we leverage a physically-inspired model of the sensor's response to illumination and introduce a chromatic adaptation loss to constrain the learning of response, enabling consistent color perception of objects regardless of lighting conditions. We further utilize the RAW data's properties to expose the scene's intensity automatically. Additionally, we have collected a multi-view low-light RAW image dataset of real-world scenes to advance research in this field. Experimental results demonstrate that our proposed method significantly outperforms existing 2D and 3D approaches. Our code and dataset will be made publicly available.
Xin Huang 0021, Guoqing Zhou 0003, Qifeng Guo, Qing Wang 0006
AAAI3
2025 Phase shift guided dynamic view synthesis from monocular video
Chuyue Zhao, Xin Huang 0021, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006
Image Vis. Comput.4
2025 ${\rm{H}}_{2}{\rm{O}}$H2O-NeRF: Radiance Fields Reconstruction for Two-Hand-Held Objects
abstract
Our work aims to reconstruct the appearance and geometry of the two-hand-held object from a sequence of color images. In contrast to traditional single-hand-held manipulation, two-hand-holding allows more flexible interaction, thereby providing back views of the object, which is particularly convenient for reconstruction but generates complex view-dependent occlusions. The recent development of neural rendering provides new potential for hand-held object reconstruction. In this paper, we propose a novel neural representation-based framework to recover radiance fields of the two-hand-held object, named ${\rm{H}}_{2}{\rm{O}}$H2O-NeRF. We first design an object-centric semantic module based on the geometric signed distance function cues to predict 3D object-centric regions and develop the view-dependent visible module based on the image-related cues to label 2D occluded regions. We then combine them to obtain a 2D visible mask that adaptively guides ray sampling on the object for optimization. We also provide a newly collected ${\rm{H}}_{2}{\rm{O}}$H2O dataset to validate the proposed method. Experiments show that our method achieves superior performance on reconstruction completeness and view-consistency synthesis compared to the state-of-the-art methods.
Xinxin Liu 0020, Qi Zhang 0029, Xin Huang 0021, Guoqing Zhou 0003, Qing Wang 0006
IEEE Trans. Vis. Comput. Graph.5
2024 A two-stage substation equipment classification method based on dual-scale attention
abstract
Abstract Accurate classification of substation equipment images remains challenging due to various factors such as unexpected illumination, viewing angles, scale variations, shadows, surface contaminants, and different elements sharing similar appearances. This paper presents a novel two‐stage substation equipment classification method based on dual‐scale attention. Leveraging the region proposal technique from Faster‐regions with CNN features (RCNN), the input images are initially decomposed into multiple scales to capture latent features. A dual‐scale attention module is introduced to enhance the precision of feature extraction. Furthermore, a two‐stage network is proposed to address the challenge of classifying closely similar substation equipment. A multi‐layer perceptron performs a coarse classification to categorize the equipment into broad categories. Then, a lightweight classifier is employed for fine‐grained subclassification, further distinguishing equipment within the same broad category. To mitigate the issue of limited training data, a specialized dataset is collected and annotated for the substation equipment classification. Experimental results demonstrate that the proposed method achieves remarkable accuracy, recall, and F1‐score surpassing 0.91, outperforming mainstream approaches in terms of recall and F1 scores. Ablation experiments further validate the significant contributions of both the dual‐scale attention and the two‐stage classification module in improving the overall performance of the classification network.
Yiyang Yao, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006
IET Image Process.3
2024 Sheared Epipolar Focus Spectrum for Dense Light Field Reconstruction
abstract
This paper presents a novel technique for the dense reconstruction of light fields (LFs) from sparse input views. Our approach leverages the Epipolar Focus Spectrum (EFS) representation, which models the LF in the transformed spatial-focus domain, avoiding the dependence on the scene depth and providing a high-quality basis for dense LF reconstruction. Previous EFS-based LF reconstruction methods learn the cross-view, occlusion, depth and shearing terms simultaneously, which makes the training difficult due to stability and convergence problems and further results in limited reconstruction performance for challenging scenarios. To address this issue, we conduct a theoretical study on the transformation between the EFSs derived from one LF with sparse and dense angular samplings, and propose that a dense EFS can be decomposed into a linear combination of the EFS of the sparse input, the sheared EFS, and a high-order occlusion term explicitly. The devised learning-based framework with the input of the under-sampled EFS and its sheared version provides high-quality reconstruction results, especially in large disparity areas. Comprehensive experimental evaluations show that our approach outperforms state-of-the-art methods, especially achieves at most dB advantages in reconstructing scenes containing thin structures.
Xue Wang 0006, Guoqing Zhou 0003, Hao Zhu 0005, Qing Wang 0006
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Dense light field reconstruction based on epipolar focus spectrum
abstract
Existing light field (LF) representations, such as epipolar plane image (EPI) and sub-aperture images, do not consider the structural characteristics across the views, so they usually require additional disparity and spatial structure cues for follow-up tasks. Besides, they have difficulties dealing with occlusions or large disparity scenes. To this end, this paper proposes a novel Epipolar Focus Spectrum (EFS) representation by rearranging the EPI spectrum. Different from the classical EPI representation where an EPI line corresponds to a specific depth, there is a one-to-one mapping from the EFS line to the view. By exploring the EFS sampling task, the analytical function is derived for constructing a non-aliasing EFS. To demonstrate its effectiveness, we develop a trainable EFS-based pipeline for light field reconstruction, where a dense light field can be reconstructed by compensating the missing EFS lines given a sparse light field, yielding promising results with cross-view consistency, especially in the presence of severe occlusion and large disparity. Experimental results on both synthetic and real-world datasets demonstrate the validity and superiority of the proposed method over SOTA methods.
Xue Wang 0006, Hao Zhu 0005, Guoqing Zhou 0003, Qing Wang 0006
Pattern Recognit.4
2023 Learning Reliable Gradients From Undersampled Circular Light Field for 3D Reconstruction
abstract
The paper presents a 3D reconstruction algorithm from an undersampled circular light field (LF). With an ultra-dense angular sampling rate, every scene point captured by a circular LF corresponds to a smooth trajectory in the circular epipolar plane volume (CEPV). Thus per-pixel disparities can be calculated by retrieving the local gradients of the CEPV-trajectories. However, the continuous curve will be broken up into discrete segments in an undersampled circular LF, which leads to a noticeable deterioration of the 3D reconstruction accuracy. We observe that the coherent structure is still embedded in the discrete segments. With less noise and ambiguity, the scene points can be reconstructed using gradients from reliable epipolar plane image (EPI) regions. By analyzing the geometric characteristics of the coherent structure in the CEPV, both the trajectory itself and its gradients could be modeled as 3D predictable series. Thus a mask-guided CNN+LSTM network is proposed to learn the mapping from the CEPV with a lower angular sampling rate to the gradients under a higher angular sampling rate. To segment the reliable regions, the reliable-mask-based loss that assesses the difference between learned gradients and ground truth gradients is added to the loss function. We construct a synthetic circular LF dataset with ground truth for depth and foreground/background segmentation to train the network. Moreover, a real-scene circular LF dataset is collected for performance evaluation. Experimental results on both public and self-constructed datasets demonstrate the superiority of the proposed method over existing state-of-the-art methods.
Zhengxi Song, Xue Wang 0006, Hao Zhu 0005, Guoqing Zhou 0003, Qing Wang 0006
IEEE Trans. Vis. Comput. Graph.4
2022 Fast and Unsupervised Action Boundary Detection for Action Segmentation
abstract
To deal with the great number of untrimmed videos produced every day, we propose an efficient unsupervised action segmentation method by detecting boundaries, named action boundary detection (ABD). In particular, the proposed method has the following advantages: no training stage and low-latency inference. To detect action boundaries, we estimate the similarities across smoothed frames, which inherently have the properties of internal consistency within actions and external discrepancy across actions. Under this circumstance, we successfully transfer the boundary detection task into the change point detection based on the similarity. Then, non-maximum suppression (NMS) is conducted in local windows to select the smallest points as candidate boundaries. In addition, a clustering algorithm is followed to refine the initial proposals. Moreover, we also extend ABD to the online setting, which enables real-time action segmentation in long untrimmed videos. By evaluating on four challenging datasets, our method achieves state-of-the-art performance. Moreover, thanks to the efficiency of ABD, we achieve the best trade-off between the accuracy and the inference time compared with existing unsupervised approaches.
Zexing Du, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006
CVPR3
2021 Behavioral features fusion for ethological CNN classification of open field test videos
Zhaolin Xiao, Guoqing Zhou 0003, Haiyan Jin
Multim. Tools Appl.3
2018 Dense Light Field Reconstruction from Sparse Sampling Using Residual Network
Mantang Guo, Hao Zhu 0005, Guoqing Zhou 0003, Qing Wang 0006
ACCV (6)3
2017 High angular resolution light field reconstruction with coded-aperture mask
abstract
In the past decade, light field imaging has greatly extended the imaging capabilities of traditional photography. However, the applications of light field imaging are limited by the aliasing artifacts due to the plenoptic sampling trade-off between angular and spatial domains. We propose to use a coded aperture light field camera instead of the traditional one, which can get more angular information without losing spatial resolution. To that end, we exploit a theoretical model to explain the relationship between light field and the raw data captured by the sensor. Then, we design a mask to code the rays using compressive sensing. Last, the sparse characteristic of light field in gradient domain and the corresponding optimization methods are utilized to reconstruct the high angular resolution light field. Experimental results on synthetic data and real data demonstrate that our system can obtain high angular resolution light field by producing a low-aliasing refocused image and high PSNR multi-view images.
Wanxin Qu, Guoqing Zhou 0003, Hao Zhu 0005, Zhaolin Xiao, Qing Wang 0006, René Vidal
ICIP2
2017 Robust outlier removal using penalized linear regression in multiview geometry
Guoqing Zhou 0003, Qing Wang 0006, Zhaolin Xiao
Neurocomputing1
2017 Aliasing Detection and Reduction Scheme on Angularly Undersampled Light Fields
abstract
When using plenoptic camera for digital refocusing, angular undersampling can cause severe (angular) aliasing artifacts. Previous approaches have focused on avoiding aliasing by pre-processing the acquired light field via prefiltering, demosaicing, reparameterization, and so on. In this paper, we present a different solution that first detects and then removes angular aliasing at the light field refocusing stage. Different from previous frequency domain aliasing analysis, we carry out a spatial domain analysis to reveal whether the angular aliasing would occur and uncover where in the image it would occur. The spatial analysis also facilitates easy separation of the aliasing versus non-aliasing regions and angular aliasing removal. Experiments on both synthetic scene and real light field data sets (camera array and Lytro camera) demonstrate that our approach has a number of advantages over the classical prefiltering and depth-dependent light field rendering techniques.
Zhaolin Xiao, Qing Wang 0006, Guoqing Zhou 0003, Jingyi Yu 0001
IEEE Trans. Image Process.3
2014 Aliasing Detection and Reduction in Plenoptic Imaging
abstract
When using plenoptic camera for digital refocusing, angular undersampling can cause severe (angular) aliasing artifacts. Previous approaches have focused on avoiding aliasing by pre-processing the acquired light field via prefiltering, demosaicing, reparameterization, etc. In this paper, we present a different solution that first detects and then removes aliasing at the light field refocusing stage. Different from previous frequency domain aliasing analysis, we carry out a spatial domain analysis to reveal whether the aliasing would occur and uncover where in the image it would occur. The spatial analysis also facilitates easy separation of the aliasing vs. non-aliasing regions and aliasing removal. Experiments on both synthetic scene and real light field camera array data sets demonstrate that our approach has a number of advantages over the classical prefiltering and depth-dependent light field rendering techniques.
Zhaolin Xiao, Qing Wang 0006, Guoqing Zhou 0003, Jingyi Yu 0001
CVPR3
2014 Reconstructing scene depth and appearance behind foreground occlusion using camera array
abstract
Foreground occlusion is a significant challenge in 3D reconstruction. In the paper, we first characterize the differences between multiview reconstruction with and without foreground occlusion. Considering both scene depth and appearance are unknown, we propose a generalized model for scene reconstruction. Then, we propose an iterative reconstruction approach in the global optimization framework, which is well performed on the camera array system. Even when all views are partially occluded, our approach can recover accurate depth map as well as scene appearance. Experimental results have indicated that our approach is more robust to foreground occlusions and outperforms state-of-the-art approaches.
Zhaolin Xiao, Qing Wang 0006, Lipeng Si, Guoqing Zhou 0003
ICIP4
2013 Enhanced Continuous Tabu Search for Parameter Estimation in Multiview Geometry
abstract
Optimization using the L_infty norm has been becoming an effective way to solve parameter estimation problems in multiview geometry. But the computational cost increases rapidly with the size of measurement data. Although some strategies have been presented to improve the efficiency of L_infty optimization, it is still an open issue. In the paper, we propose a novel approach under the framework of enhanced continuous tabu search (ECTS) for generic parameter estimation in multiview geometry. ECTS is an optimization method in the domain of artificial intelligence, which has an interesting ability of covering a wide solution space by promoting the search far away from current solution and consecutively decreasing the possibility of trapping in the local minima. Taking the triangulation as an example, we propose the corresponding ways in the key steps of ECTS, diversification and intensification. We also present theoretical proof to guarantee the global convergence of search with probability one. Experimental results have validated that the ECTS based approach can obtain global optimum efficiently, especially for large scale dimension of parameter. Potentially, the novel ECTS based algorithm can be applied in many applications of multiview geometry.
Guoqing Zhou 0003, Qing Wang 0006
ICCV1