Zhou Xue

dblp:94/10764 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
14since 2021 · last 2025
0000-0002-3157-6606ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021
YearPublicationVenuePosition
2025 Hyper-3DG: Text-to-3D Gaussian Generation via Hypergraph
Donglin Di, Chaofan Luo, Zhou Xue, Wei Chen 0089, Xun Yang 0001, Yue Gao 0002
Int. J. Comput. Vis.4
2025 RGB-D Visual Perception for Occluded Scenes via Event Camera
Siqi Li 0001, Zongze Wu 0001, Zhou Xue, Yu-Shen Liu, Yue Gao 0002
Int. J. Comput. Vis.4
2025 TrAME: Trajectory-Anchored Multi-View Editing for Text-Guided 3D Gaussian Manipulation
abstract
Despite significant strides in the field of 3D scene editing, current methods encounter substantial challenge, particularly in preserving 3D consistency during the multi-view editing process. To tackle this challenge, we propose a progressive 3D editing strategy that ensures multi-view consistency via a Trajectory-Anchored Scheme (TAS) with a dual-branch editing mechanism. Specifically, TAS facilitates a tightly coupled iterative process between 2D view editing and 3D updating, preventing error accumulation yielded from the text-to-image process. Additionally, we explore the connection between optimization-based methods and reconstruction-based methods, offering a unified perspective for selecting superior design choices, supporting the rationale behind the designed TAS. We further present a tuning-free View-Consistent Attention Control (VCAC) module that leverages cross-view semantic and geometric reference from the source branch to yield aligned views from the target branch during the editing of 2D views. To validate the effectiveness of our method, we analyze 2D examples to demonstrate the improved consistency with the VCAC module. Extensive quantitative and qualitative results in text-guided 3D scene editing clearly indicate that our method can achieve superior editing quality compared with state-of-the-art 3D scene editing methods. Our project site is athttps://fkcptlst.github.io/TrAME/
Chaofan Luo, Donglin Di, Xun Yang 0001, Yongjia Ma, Zhou Xue, Wei Chen 0089, Xiaofei Gou, Yebin Liu
IEEE Trans. Multim.5
2024 3D Feature Tracking via Event Camera
abstract
This paper presents the first 3D feature tracking method with the corresponding dataset. Our proposed method takes event streams from stereo event cameras as input to pre-dict 3D trajectories of the target features with high-speed motion. To achieve this, our method leverages a joint framework to predict the 2D feature motion offsets and the 3D feature spatial position simultaneously. A motion compensation module is leveraged to overcome the feature deformation. A patch matching module based on bi-polarity hypergraph modeling is proposed to robustly es-timate the feature spatial position. Meanwhile, we collect the first 3D feature tracking dataset with high-speed moving objects and ground truth 3D feature trajectories at 250 FPS, named E-3DTrack, which can be used as the first high-speed 3D feature tracking benchmark. Our code and dataset could be found at: https://github.com/lisiqi19971013/E-3DTrack.
Siqi Li 0001, Zhikuan Zhou, Zhou Xue, Shaoyi Du, Yue Gao 0002
CVPR3
2024 iToF-Flow-Based High Frame Rate Depth Imaging
abstract
iToF is a prevalent, cost-effective technology for 3D perception. While its reliance on multi-measurement commonly leads to reduced performance in dynamic environments. Based on the analysis of the physical iToF imaging process, we propose the iToF flow, composed of crossmode transformation and uni-mode photometric correction, to model the variation of measurements caused by different measurement modes and 3D motion, respectively. We propose a local linear transform (LLT) based cross-mode transfer module (LCTM) for mode-varying and pixel shift compensation of cross-mode flow, and uni-mode photometric correct module (UPCM) for estimating the depth-wise motion caused photometric residual of uni-mode flow. The iToF flow-based depth extraction network is proposed which could facilitate the estimation of the 4-phase measurements at each individual time for high framerate and accurate depth estimation. Extensive experiments, including both simulation and real-world experiments, are conducted to demonstrate the effectiveness of the proposed methods. Compared with the SOTA method, our approach reduces the computation time by 75% while improving the performance by 38%. The code and database are available at https://github.com/ComputationalPerceptionLab/iToF_flow.
Zhou Xue, Tao Yue 0003
CVPR2
2024 Joint2Human: High-quality 3D Human Generation via Compact Spherical Embedding of 3D Joints
abstract
3D human generation is increasingly significant in var-ious applications. However, the direct use of 2D genera-tive methods in 3D generation often results in losing lo-cal details, while methods that reconstruct geometry from generated images struggle with global view consistency. In this work, we introduce joint2Human, a novel method that leverages 2D diffusion models to generate detailed 3D human geometry directly, ensuring both global structure and local details. To achieve this, we employ the Fourier occupancy field (FOF) representation, enabling the direct generation of 3D shapes as preliminary results with 2D generative models. With the proposed high-frequency enhancer and the multi-view recarving strategy, our method can seamlessly integrate the details from different views into a uniform global shape. To better utilize the 3D human prior and enhance control over the generated geometry, we introduce a compact spherical embedding of 3D joints. This allows for an effective guidance of pose during the gener-ation process. Additionally, our method can generate 3D humans guided by textual inputs. Our experimental results demonstrate the capability of our method to ensure global structure, local details, high resolution, and low computational cost simultaneously. More results and the code can be found on our project page at http://cic.tju.edu.cn/faculty/likun/projects/Joint2Human.
Muxin Zhang, Qiao Feng 0001, Zhuo Su 0006, Zhou Xue, Kun Li 0001
CVPR5
2024 OHTA: One-shot Hand Avatar via Data-driven Implicit Priors
abstract
In this paper, we delve into the creation of one-shot hand avatars, attaining high-fidelity and drivable hand represen-tations swiftly from a single image. With the burgeoning domains of the digital human, the need for quick and per-sonalized hand avatar creation has become increasingly critical. Existing techniques typically require extensive in-put data and may prove cumbersome or even impractical in certain scenarios. To enhance accessibility, we present a novel method OHTA (One-shot Hand avaTAr) that en-ables the creation of detailed hand avatars from merely one image. OHTA tackles the inherent difficulties of this data-limited problem by learning and utilizing data-driven hand priors. Specifically, we design a hand prior model initially employed for 1) learning various hand priors with available data and subsequently for 2) the inversion and fitting of the target identity with prior knowledge. OHTA demonstrates the capability to create high-fidelity hand avatars with con-sistent animatable quality, solely relying on a single image. Furthermore, we illustrate the versatility of OHTA through diverse applications, encompassing text-to-avatar conver-sion, hand editing, and identity latent space manipulation.
Xiaozheng Zheng, Zhuo Su 0006, Zeran Xu, Zhaohu Li, Yang Zhao 0025, Zhou Xue
CVPR7
2023 Decoupled Iterative Refinement Framework for Interacting Hands Reconstruction from a Single RGB Image
abstract
Reconstructing interacting hands from a single RGB image is a very challenging task. On the one hand, severe mutual occlusion and similar local appearance between two hands confuse the extraction of visual features, resulting in the misalignment of estimated hand meshes and the image. On the other hand, there are complex spatial relationship between interacting hands, which significantly increases the solution space of hand poses and increases the difficulty of network learning. In this paper, we propose a decoupled iterative refinement framework to achieve pixel-alignment hand reconstruction while efficiently modeling the spatial relationship between hands. Specifically, we define two feature spaces with different characteristics, namely 2D visual feature space and 3D joint feature space. First, we obtain joint-wise features from the visual feature map and utilize a graph convolution network and a transformer to perform intra- and inter-hand information interaction in the 3D joint feature space, respectively. Then, we project the joint features with global information back into the 2D visual feature space in an obfuscation-free manner and utilize the 2D convolution for pixel-wise enhancement. By performing multiple alternate enhancements in the two feature spaces, our method can achieve an accurate and robust reconstruction of interacting hands. Our method outperforms all existing two-hand reconstruction methods by a large margin on the InterHand2.6M dataset.
Pengfei Ren 0001, Xiaozheng Zheng, Zhou Xue, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao
ICCV4
2023 Realistic Full-Body Tracking from Sparse Observations via Joint-Level Modeling
abstract
To bridge the physical and virtual worlds for rapidly developed VR/AR applications, the ability to realistically drive 3D full-body avatars is of great significance. Although real-time body tracking with only the head-mounted displays (HMDs) and hand controllers is heavily under-constrained, a carefully designed end-to-end neural network is of great potential to solve the problem by learning from large-scale motion data. To this end, we propose a two-stage framework that can obtain accurate and smooth full-body motions with the three tracking signals of head and hands only. Our framework explicitly models the joint-level features in the first stage and utilizes them as spatiotemporal tokens for alternating spatial and temporal transformer blocks to capture joint-level correlations in the second stage. Furthermore, we design a set of loss terms to constrain the task of a high degree of freedom, such that we can exploit the potential of our joint-level modeling. With extensive experiments on the AMASS motion dataset and real-captured data, we validate the effectiveness of our designs and show our proposed method can achieve more accurate and smooth motion compared to existing approaches.
Xiaozheng Zheng, Zhuo Su 0006, Zhou Xue, Xiaojie Jin 0004
ICCV4
2023 HaMuCo: Hand Pose Estimation via Multiview Collaborative Self-Supervised Learning
abstract
Recent advancements in 3D hand pose estimation have shown promising results, but its effectiveness has primarily relied on the availability of large-scale annotated datasets, the creation of which is a laborious and costly process. To alleviate the label-hungry limitation, we propose a self-supervised learning framework, HaMuCo, that learns a single-view hand pose estimator from multi-view pseudo 2D labels. However, one of the main challenges of self-supervised learning is the presence of noisy labels and the "groupthink" effect from multiple views. To overcome these issues, we introduce a cross-view interaction network that distills the single-view estimator by utilizing the cross-view correlated features and enforcing multi-view consistency to achieve collaborative learning. Both the single-view estimator and the cross-view interaction network are trained jointly in an end-to-end manner. Extensive experiments show that our method can achieve state-of-the-art performance on multi-view self-supervised hand pose estimation. Furthermore, the proposed cross-view interaction network can also be applied to hand pose estimation from multi-view input and outperforms previous methods under the same settings.
Xiaozheng Zheng, Zhou Xue, Pengfei Ren 0001, Jingyu Wang 0001
ICCV3
2023 Reconstructing Interacting Hands with Interaction Prior from Monocular Images
abstract
Reconstructing interacting hands from monocular images is indispensable in AR/VR applications. Most existing solutions rely on the accurate localization of each skeleton joint. However, these methods tend to be unreliable due to the severe occlusion and confusing similarity among adjacent hand parts. This also defies human perception because humans can quickly imitate an interaction pattern without localizing all joints. Our key idea is to first construct a two-hand interaction prior and recast the interaction reconstruction task as the conditional sampling from the prior. To expand more interaction states, a large-scale multimodal dataset with physical plausibility is proposed. Then a VAE is trained to further condense these interaction patterns as latent codes in a prior distribution. When looking for image cues that contribute to interaction prior sampling, we propose the interaction adjacency heatmap (IAH). Compared with a joint-wise heatmap for localization, IAH assigns denser visible features to those invisible joints. Compared with an all-in-one visible heatmap, it provides more fine-grained local interaction information in each interaction region. Finally, the correlations between the extracted features and corresponding interaction codes are linked by the ViT module. Comprehensive evaluations on benchmark datasets have verified the effectiveness of this framework. The code and dataset are publicly available at https://github.com/binghui-z/InterPrior_pytorch.
Binghui Zuo, Zimeng Zhao, Wenqian Sun, Wei Xie 0012, Zhou Xue, Yangang Wang 0001
ICCV5
2022 SHRAG: Semantic Hierarchical Graph for Floorplan Representation
abstract
Representation learning from a floorplan is a fundamental step for various floorplan-related applications, such as retrieval, reconstruction, and generation. Previous works often use single-layer graphs to model room categories and adjacencies as nodes and edges, which can only represent a floorplan in a coarse level without detailed information such as room shape and door/window locations. Thus, we propose SHRAG, a hierarchical semantic graph for floor-plan representation, with two hierarchical graph modules to semantically encode both coarse and detailed information of a floorplan. First, a Detail Graph Module(DGM) was designed to learn detailed contour and attachable elements embedding for each room. Then, a Global Graph Module(GGM) was applied to concatenate and encode both detailed embeddings from DGM and coarse information including room categories and room relations. Results show that the representation learned from SHRAG achieved SOTA on floorplan retrieval tasks. Further, SHRAG performs more satisfying than other floorplan similarity metrics in a comprehensive user study. We also demonstrate that SHRAG can be easily adapted to real-world complex indoor scenes even with furniture.
Jiongchao Jin, Zhou Xue, Biao Leng
3DV2
2022 3D Room Layout Estimation from a Cubemap of Panorama Image via Deep Manhattan Hough Transform
Chao Wen 0001, Zhou Xue, Yue Gao 0002
ECCV (1)3
2021 Fast Light-field Disparity Estimation with Multi-disparity-scale Cost Aggregation
abstract
Light field images contain both angular and spatial information of captured light rays. The rich information of light fields enables straightforward disparity recovery capability but demands high computational cost as well. In this paper, we design a lightweight disparity estimation model with physical-based multi-disparity-scale cost volume aggregation for fast disparity estimation. By introducing a sub-network of edge guidance, we significantly improve the recovery of geometric details near edges and improve the overall performance. We test the proposed model extensively on both synthetic and real-captured datasets, which provide both densely and sparsely sampled light fields. Finally, we significantly reduce computation cost and GPU memory consumption, while achieving comparable performance with state-of-the-art disparity estimation methods for light fields. Our source code is available at https://github.com/zcong17huang/FastLFnet.
Zhou Xue, Weizhu Xu, Tao Yue 0003
ICCV3