Xue Wang 0006

dblp:39/2811-6 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Zero-Pose-Prior NeRF: Recursive Radiance Field Reconstruction From Unposed and Unordered Images
abstract
The dependence of neural radiance fields (NeRF) on accurate camera poses has emerged as a critical obstacle to their widespread real-world applications. While recent advances have demonstrated the potential for simultaneously addressing camera registration and scene reconstruction, these methods inherently rely on reasonable initialization derived from pose or scene priors and struggle with complex scenes involving large camera motions, particularly in unordered 360-degree scenes. In this work, we propose Zero-Pose-Prior NeRF to recover radiance fields from unposed and unordered image collections without any prior knowledge. Our key insight is to decompose this complex problem into smaller sub-problems, wherein the sub-problems' camera poses are initially estimated to provide self-bootstrapping priors for the global pose estimation, followed by a recursive registration and reconstruction. To achieve this, we first perform scene partitioning to establish a hierarchical structure that describes registration order from local to global. Thereafter, we devise a conditionally-decoupled positional encoding for NeRFs, which serves as the basic model for camera pose estimation and scene representation. Following this, we develop a recursive registration to recursively estimate the poses of local scenes and register them into a unified global pose space, ultimately enabling the reconstruction of the entire scene. Experiments on real-world scenes show that our approach outperforms the state-of-the-art pose-free methods in terms of accurate camera poses and robust radiance field reconstruction, resulting in high-fidelity view synthesis.
Xinxin Liu 0020, Qi Zhang 0029, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006
IEEE Trans. Image Process.3
2025 Phase shift guided dynamic view synthesis from monocular video
Chuyue Zhao, Xin Huang 0021, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006
Image Vis. Comput.3
2025 Generalizable 3D Gaussian Splatting for novel view synthesis
Chuyue Zhao, Xin Huang 0021, Xue Wang 0006, Qing Wang 0006
Pattern Recognit.4
2024 Dual-Scale Temporal Dependency Learning for Unsupervised Video Anomaly Detection
Xue Wang 0006, Zexing Du, Qing Wang 0006
PRCV (10)2
2024 A two-stage substation equipment classification method based on dual-scale attention
abstract
Abstract Accurate classification of substation equipment images remains challenging due to various factors such as unexpected illumination, viewing angles, scale variations, shadows, surface contaminants, and different elements sharing similar appearances. This paper presents a novel two‐stage substation equipment classification method based on dual‐scale attention. Leveraging the region proposal technique from Faster‐regions with CNN features (RCNN), the input images are initially decomposed into multiple scales to capture latent features. A dual‐scale attention module is introduced to enhance the precision of feature extraction. Furthermore, a two‐stage network is proposed to address the challenge of classifying closely similar substation equipment. A multi‐layer perceptron performs a coarse classification to categorize the equipment into broad categories. Then, a lightweight classifier is employed for fine‐grained subclassification, further distinguishing equipment within the same broad category. To mitigate the issue of limited training data, a specialized dataset is collected and annotated for the substation equipment classification. Experimental results demonstrate that the proposed method achieves remarkable accuracy, recall, and F1‐score surpassing 0.91, outperforming mainstream approaches in terms of recall and F1 scores. Ablation experiments further validate the significant contributions of both the dual‐scale attention and the two‐stage classification module in improving the overall performance of the classification network.
Yiyang Yao, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006
IET Image Process.2
2024 Spatially-Varying Illumination-Aware Indoor Harmonization
Zhongyun Hu, Xue Wang 0006, Qing Wang 0006
Int. J. Comput. Vis.3
2024 Sheared Epipolar Focus Spectrum for Dense Light Field Reconstruction
abstract
This paper presents a novel technique for the dense reconstruction of light fields (LFs) from sparse input views. Our approach leverages the Epipolar Focus Spectrum (EFS) representation, which models the LF in the transformed spatial-focus domain, avoiding the dependence on the scene depth and providing a high-quality basis for dense LF reconstruction. Previous EFS-based LF reconstruction methods learn the cross-view, occlusion, depth and shearing terms simultaneously, which makes the training difficult due to stability and convergence problems and further results in limited reconstruction performance for challenging scenarios. To address this issue, we conduct a theoretical study on the transformation between the EFSs derived from one LF with sparse and dense angular samplings, and propose that a dense EFS can be decomposed into a linear combination of the EFS of the sparse input, the sheared EFS, and a high-order occlusion term explicitly. The devised learning-based framework with the input of the under-sampled EFS and its sheared version provides high-quality reconstruction results, especially in large disparity areas. Comprehensive experimental evaluations show that our approach outperforms state-of-the-art methods, especially achieves at most dB advantages in reconstructing scenes containing thin structures.
Xue Wang 0006, Guoqing Zhou 0003, Hao Zhu 0005, Qing Wang 0006
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Learning Semantics-Guided Representations for Scoring Figure Skating
abstract
This paper explores semantic-aware representations for scoring figure skating videos. Most existing approaches to sports video analysis only focus on reasoning action scores based on visual input, limiting their ability to depict high-level semantic representations. Here, we propose a teacher-student-based network with an attention mechanism to realize an adaptive knowledge transfer from the semantic domain to the visual domain, which is termed semantics-guided network (SGN). Specifically, we use a set of learnable atomic queries in the student branch to mimic the semantic-aware distribution in the teacher branch, which is represented by the visual and semantic inputs. In addition, we propose three auxiliary losses to align features in different domains. With aligned feature representations, the adapted teacher is capable of transferring the semantic knowledge to the student. To verify the effectiveness of our method, we collect a new dataset OlympicFS for scoring figure skating. Besides action scores, OlympicFS also provides professional comments on actions for learning semantic representations. By evaluating four challenging datasets, our method achieves state-of-the-art performance.
Zexing Du, Di He 0010, Xue Wang 0006, Qing Wang 0006
IEEE Trans. Multim.3
2023 InterFormer: Human Interaction Understanding with Deformed Transformer
Di He 0010, Zexing Du, Xue Wang 0006, Qing Wang 0006
ICIC (5)3
2023 Perceiving local relative motion and global correlations for weakly supervised group activity recognition
Zexing Du, Xue Wang 0006, Qing Wang 0006
Image Vis. Comput.2
2023 Dense light field reconstruction based on epipolar focus spectrum
abstract
Existing light field (LF) representations, such as epipolar plane image (EPI) and sub-aperture images, do not consider the structural characteristics across the views, so they usually require additional disparity and spatial structure cues for follow-up tasks. Besides, they have difficulties dealing with occlusions or large disparity scenes. To this end, this paper proposes a novel Epipolar Focus Spectrum (EFS) representation by rearranging the EPI spectrum. Different from the classical EPI representation where an EPI line corresponds to a specific depth, there is a one-to-one mapping from the EFS line to the view. By exploring the EFS sampling task, the analytical function is derived for constructing a non-aliasing EFS. To demonstrate its effectiveness, we develop a trainable EFS-based pipeline for light field reconstruction, where a dense light field can be reconstructed by compensating the missing EFS lines given a sparse light field, yielding promising results with cross-view consistency, especially in the presence of severe occlusion and large disparity. Experimental results on both synthetic and real-world datasets demonstrate the validity and superiority of the proposed method over SOTA methods.
Xue Wang 0006, Hao Zhu 0005, Guoqing Zhou 0003, Qing Wang 0006
Pattern Recognit.2
2023 Self-Supervised Global Spatio-Temporal Interaction Pre-Training for Group Activity Recognition
abstract
This paper focuses on exploring distinctive spatio-temporal representation in a self-supervised manner for group activity recognition. Firstly, previous networks treat spatial- and temporal-aware information as a whole, limiting their abilities to represent complex spatio-temporal correlations for group activity. Here, we propose the Spatial and Temporal Attention Heads (STAHs) to extract spatial- and temporal-aware representations independently, which generate complementary contexts for boosting group activity understanding. Then, we propose the Global Spatio-Temporal Contrastive (GSTCo) loss to aggregate these two kinds of features. Unlike previous works focusing on the individual temporal consistency while overlooking the correlations between actors, i.e., in a local perspective, we explore the global spatial and temporal dependency. Moreover, GSTCo could effectively avoid the trivial solution faced in contrastive learning by achieving the right balance between spatial and temporal representations. Furthermore, our method imports affordable overhead during pre-training, without additional parameters or computational costs in inference, guaranteeing efficiency. By evaluating on widely-used datasets for group activity recognition, our method achieves good performance. State-of-the-art performance is achieved when applying our pre-trained backbone to existing networks. Extensive experiments verify the generalizability of our method.
Zexing Du, Xue Wang 0006, Qing Wang 0006
IEEE Trans. Circuits Syst. Video Technol.2
2023 Learning Reliable Gradients From Undersampled Circular Light Field for 3D Reconstruction
abstract
The paper presents a 3D reconstruction algorithm from an undersampled circular light field (LF). With an ultra-dense angular sampling rate, every scene point captured by a circular LF corresponds to a smooth trajectory in the circular epipolar plane volume (CEPV). Thus per-pixel disparities can be calculated by retrieving the local gradients of the CEPV-trajectories. However, the continuous curve will be broken up into discrete segments in an undersampled circular LF, which leads to a noticeable deterioration of the 3D reconstruction accuracy. We observe that the coherent structure is still embedded in the discrete segments. With less noise and ambiguity, the scene points can be reconstructed using gradients from reliable epipolar plane image (EPI) regions. By analyzing the geometric characteristics of the coherent structure in the CEPV, both the trajectory itself and its gradients could be modeled as 3D predictable series. Thus a mask-guided CNN+LSTM network is proposed to learn the mapping from the CEPV with a lower angular sampling rate to the gradients under a higher angular sampling rate. To segment the reliable regions, the reliable-mask-based loss that assesses the difference between learned gradients and ground truth gradients is added to the loss function. We construct a synthetic circular LF dataset with ground truth for depth and foreground/background segmentation to train the network. Moreover, a real-scene circular LF dataset is collected for performance evaluation. Experimental results on both public and self-constructed datasets demonstrate the superiority of the proposed method over existing state-of-the-art methods.
Zhengxi Song, Xue Wang 0006, Hao Zhu 0005, Guoqing Zhou 0003, Qing Wang 0006
IEEE Trans. Vis. Comput. Graph.2
2022 Fast and Unsupervised Action Boundary Detection for Action Segmentation
abstract
To deal with the great number of untrimmed videos produced every day, we propose an efficient unsupervised action segmentation method by detecting boundaries, named action boundary detection (ABD). In particular, the proposed method has the following advantages: no training stage and low-latency inference. To detect action boundaries, we estimate the similarities across smoothed frames, which inherently have the properties of internal consistency within actions and external discrepancy across actions. Under this circumstance, we successfully transfer the boundary detection task into the change point detection based on the similarity. Then, non-maximum suppression (NMS) is conducted in local windows to select the smallest points as candidate boundaries. In addition, a clustering algorithm is followed to refine the initial proposals. Moreover, we also extend ABD to the online setting, which enables real-time action segmentation in long untrimmed videos. By evaluating on four challenging datasets, our method achieves state-of-the-art performance. Moreover, thanks to the efficiency of ABD, we achieve the best trade-off between the accuracy and the inference time compared with existing unsupervised approaches.
Zexing Du, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006
CVPR2
2022 PNRNet: Physically-Inspired Neural Rendering for Any-to-Any Relighting
abstract
Existing any-to-any relighting methods suffer from the task-aliasing effects and the loss of local details in the image generation process, such as shading and attached-shadow. In this paper, we present PNRNet, a novel neural architecture that decomposes the any-to-any relighting task into three simpler sub-tasks, i.e. lighting estimation, color temperature transfer, and lighting direction transfer, to avoid the task-aliasing effects. These sub-tasks are easy to learn and can be trained with direct supervisions independently. To better preserve local shading and attached-shadow details, we propose a parallel multi-scale network that incorporates multiple physical attributes to model local illuminations for lighting direction transfer. We also introduce a simple yet effective color temperature transfer network to learn a pixel-level non-linear function which allows color temperature adjustment beyond the predefined color temperatures and generalizes well to real images. Extensive experiments demonstrate that our proposed approach achieves better results quantitatively and qualitatively than prior works.
Zhongyun Hu, Ntumba Elie Nsampi, Xue Wang 0006, Qing Wang 0006
IEEE Trans. Image Process.3
2021 3D Scene Reconstruction with an Un-calibrated Light Field Camera
Qi Zhang 0029, Hongdong Li, Xue Wang 0006, Qing Wang 0006
Int. J. Comput. Vis.3
2021 Region-based depth feature descriptor for saliency detection on light field
Xue Wang 0006, Yingying Dong, Qi Zhang 0029, Qing Wang 0006
Multim. Tools Appl.1
2021 4D Light Field Segmentation From Light Field Super-Pixel Hypergraph Representation
abstract
Efficient and accurate segmentation of full 4D light fields is an important task in computer vision and computer graphics. The massive volume and the redundancy of light fields make it an open challenge. In this article, we propose a novel light field hypergraph (LFHG) representation using the light field super-pixel (LFSP) for interactive light field segmentation. The LFSPs not only maintain the light field spatio-angular consistency, but also greatly contribute to the hypergraph coarsening. These advantages make LFSPs useful to improve segmentation performance. Based on the LFHG representation, we present an efficient light field segmentation algorithm via graph-cut optimization. Experimental results on both synthetic and real scene data demonstrate that our method outperforms state-of-the-art methods on the light field segmentation task with respect to both accuracy and efficiency.
Xianqiang Lv, Xue Wang 0006, Qing Wang 0006, Jingyi Yu 0001
IEEE Trans. Vis. Comput. Graph.2
2020 DGAN: Disentangled Representation Learning for Anisotropic BRDF Reconstruction
abstract
Accurate reconstruction of real-world materials' appearance from a very limited number of samples is still a huge challenge in computer vision and graphics. In this paper, we present a novel deep architecture, Disentangled Generative Adversarial Network (DGAN), which performs anisotropic Bidirectional Reflectance Distribution Function (BRDF) reconstruction from single BRDF subspace with the maximum entropy. In contrast to previous approaches that directly map known samples to a full BRDF using a CNN, a disentangled representation learning is applied to guide the reconstruction process. In order to learn different physical factors of the BRDF, the generator of the DGAN mainly consists of a fresnel estimator module (FEM) and a directional module (DM). Considering the fact that the entropy of different BRDF subspace varies, we further divide the BRDF into He-BRDF and Le-BRDF to reconstruct the interior part and the exterior part of the directional factor. Experimental results show that our approach outperforms state-of-the-art methods.
Zhongyun Hu, Xue Wang 0006, Qing Wang 0006
ICASSP2
2020 Accurate 3D Reconstruction from Circular Light Field Using CNN-LSTM
abstract
A light field is formed by densely capturing images on a regular sub-aperture grid. Geometry information endowed in the epipolar plane images(EPI) can only lead to a 2. 5D reconstruction. In order to obtain a full 360°view of an object, we focus on light fields captured by a circularly moving camera, resulting in circular light fields (or Cir-LFs in short). Compared with traditional EPIs, Circular EPIs(CEPIs) provide unique advantages, such as that corresponding points forming a 3D sinusoid like curve instead of a 2D straight line and geometry information encoded sequentially in multiple adjacent views along the curve. However, current reconstruction methods only focus on the 2D projection of 3D curve, leading to distortions in the reconstructed upper and lower surfaces. We propose to analyze 3D features contained in the 3D CEPI volume and we develop a deep CNN-LSTM network to model the gradient map in the CEPI volume. Additionally, a large scale Cir-LF dataset is constructed for research purpose. Experiments on both synthetic and real scenes demonstrate the effectiveness and generaliability of the proposed method.
Zhengxi Song, Hao Zhu 0005, Xue Wang 0006, Hongdong Li, Qing Wang 0006
ICME4
2017 Motion-Based Temporal Alignment of Independently Moving Cameras
abstract
This paper presents a method to establish a nonlinear temporal correspondence between two video sequences captured by cameras independently moving in a dynamic 3D scene. We assume that the 3D spatial poses of the cameras are known for each frame. With predefined trajectory basis, the coefficients of the reconstructed trajectory of a moving scene point reflect the rhythm in motion. A robust rank constraint from the coefficient matrices is exploited to measure the spatiotemporal alignment quality for every feasible pair of video fragments. Point correspondences across sequences are not required or even it is possible that different points are tracked in different sequences, only if they satisfy the assumption that every 3D point tracked in the observed sequence can be described as a linear combination of a subset of the 3D points tracked in the reference sequence. Synchronization is then performed using a graph-based search algorithm to find the globally optimal path that minimizes both spatial and temporal misalignments. Our algorithm can use both complete and incomplete feature trajectories along time, and is robust to mild outliers. We verify the robustness and performance of the proposed approach on synthetic data as well as on challenging real video sequences.
Xue Wang 0006, Jianbo Shi, Hyun Soo Park, Qing Wang 0006
IEEE Trans. Circuits Syst. Video Technol.1