Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhaoxuan Zhang

dblp:237/9860 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-9366-859XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 55% Generative modeling · 18% Reinforcement learning · 11%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d scene reconstruction
1.022023
Point Cloud Scene Completion With Joint Color and Semantic Estimation From Single RGB-D Image · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Deep Reinforcement Learning of Volume-Guided Progressive View Inpainting for 3D Point Scene Completion From a Single Depth Image · CVPR 2019
Machine learning › Generative modeling › diffusion model
3d diffusion models
0.912025
Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and Reconstruction · CVPR 2025
Computer vision › 3D vision
3d shape reconstruction
0.912025
Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and Reconstruction · CVPR 2025
Machine learning › Generative modeling
diffusion model
0.912025
Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and Reconstruction · CVPR 2025
Computer vision › 3D vision
3d scene understanding
0.922023
Point Cloud Scene Completion With Joint Color and Semantic Estimation From Single RGB-D Image · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Explore Contextual Information for 3D Scene Graph Generation · IEEE Trans. Vis. Comput. Graph. 2023
Machine learning › Reinforcement learning
deep reinforcement learning
0.822023
Point Cloud Scene Completion With Joint Color and Semantic Estimation From Single RGB-D Image · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Deep Reinforcement Learning of Volume-Guided Progressive View Inpainting for 3D Point Scene Completion From a Single Depth Image · CVPR 2019
Computer vision › 3D vision › 3d scene understanding
3d scene graph
0.712023
Explore Contextual Information for 3D Scene Graph Generation · IEEE Trans. Vis. Comput. Graph. 2023
Robotics › Robot navigation and mapping › view planning
next-best-view planning
0.712023
Point Cloud Scene Completion With Joint Color and Semantic Estimation From Single RGB-D Image · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › Vision and language › visual reasoning
scene graph reasoning
0.712023
Explore Contextual Information for 3D Scene Graph Generation · IEEE Trans. Vis. Comput. Graph. 2023
Computer vision › 3D vision › 3d scene understanding
semantic scene completion
0.712023
Point Cloud Scene Completion With Joint Color and Semantic Estimation From Single RGB-D Image · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › 3D vision › 3d shape reconstruction
shape completion
0.712023
Single Depth-image 3D Reflection Symmetry and Shape Prediction · ICCV 2023
Computer vision › 3D vision › 3d shape analysis › symmetry analysis
symmetry detection
0.712023
Single Depth-image 3D Reflection Symmetry and Shape Prediction · ICCV 2023
Robotics › Robot manipulation
tactile sensing
0.312025
Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and Reconstruction · CVPR 2025
Machine learning › Reinforcement learning › deep reinforcement learning
deep q-network
0.112019
Deep Reinforcement Learning of Volume-Guided Progressive View Inpainting for 3D Point Scene Completion From a Single Depth Image · CVPR 2019

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.5touch embedding · 0.9diffusion model · 0.9view inpainting · 0.7iterative symmetry completion · 0.7graph feature extraction · 0.7graph contextual reasoning · 0.7deep reinforcement learning · 0.7a3c · 0.7deep q-network · 0.4
YearPublicationVenuePosition
2025 Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and Reconstruction
abstract
Diffusion models have made breakthroughs in 3D generation tasks. Current 3D diffusion models focus on reconstructing target shape from images or a set of partial observations. While excelling in global context understanding, they struggle to capture the local details of complex shapes and limited to the occlusion and lighting conditions. To overcome these limitations, we utilize tactile images to capture the local 3D information and propose a Touch2Shape model, which leverages a touch-conditioned diffusion model to explore and reconstruct the target shape from touch. For shape reconstruction, we have developed a touch embedding module to condition the diffusion model in creating a compact representation and a touch shape fusion module to refine the reconstructed shape. For shape exploration, we combine the diffusion model with reinforcement learning to train a policy. This involves using the generated latent vector from the diffusion model to guide the touch exploration policy training through a novel reward design. Experiments validate the reconstruction quality thorough both qualitatively and quantitative analysis, and our touch exploration policy further boosts reconstruction performance.
Zhaoxuan Zhang, Jiajin Qiu, Dilong Sun, Zhengyu Meng, Xiaopeng Wei
CVPR2
2025 Automatic Modulation Classification Based on Efficient Multimodal Feature Fusion
Zhaoxuan Zhang
Mob. Networks Appl.5
2025 Self-supervised indoor scene point cloud completion from a single panorama
Zhaoxuan Zhang
Vis. Comput.2
2023 Single Depth-image 3D Reflection Symmetry and Shape Prediction
abstract
In this paper, we present Iterative Symmetry Completion Network (ISCNet), a single depth-image shape completion method that exploits reflective symmetry cues to obtain more detailed shapes. The efficacy of single depth-image shape completion methods is often sensitive to the accuracy of the symmetry plane. ISCNet therefore jointly estimates the symmetry plane and shape completion iteratively; more complete shapes contribute to more robust symmetry plane estimates and vice versa. Furthermore, our shape completion method operates in the image domain, enabling more efficient high-resolution, detailed geometry reconstruction. We perform the shape completion from pairs of viewpoints, reflected across the symmetry plane, predicted by a reinforcement learning agent to improve robustness and to simultaneously explicitly leverage symmetry. We demonstrate the effectiveness of ISCNet on a variety of object categories on both synthetic and real-scanned datasets.
Zhaoxuan Zhang, Bo Dong 0004, Felix Heide, Pieter Peers, Xin Yang 0011
ICCV1
2023 Point Cloud Scene Completion With Joint Color and Semantic Estimation From Single RGB-D Image
abstract
We present a deep reinforcement learning method of progressive view inpainting for colored semantic point cloud scene completion under volume guidance, achieving high-quality scene reconstruction from only a single RGB-D image with severe occlusion. Our approach is end-to-end, consisting of three modules: 3D scene volume reconstruction, 2D RGB-D and segmentation image inpainting, and multi-view selection for completion. Given a single RGB-D image, our method first predicts its semantic segmentation map and goes through the 3D volume branch to obtain a volumetric scene reconstruction as a guide to the next view inpainting step, which attempts to make up the missing information; the third step involves projecting the volume under the same view of the input, concatenating them to complete the current view RGB-D and segmentation map, and integrating all RGB-D and segmentation maps into the point cloud. Since the occluded areas are unavailable, we resort to a A3C network to glance around and pick the next best view for large hole completion progressively until a scene is adequately reconstructed while guaranteeing validity. All steps are learned jointly to achieve robust and consistent results. We perform qualitative and quantitative evaluations with extensive experiments on the 3D-FUTURE data, obtaining better results than state-of-the-arts.
Zhaoxuan Zhang, Xiaoguang Han 0001, Bo Dong 0004, Xin Yang 0011
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Explore Contextual Information for 3D Scene Graph Generation
abstract
3D scene graph generation (SGG) has been of high interest in computer vision. Although the accuracy of 3D SGG on coarse classification and single relation label has been gradually improved, the performance of existing works is still far from being perfect for fine-grained and multi-label situations. In this article, we propose a framework fully exploring contextual information for the 3D SGG task, which attempts to satisfy the requirements of fine-grained entity class, multiple relation labels, and high accuracy simultaneously. Our proposed approach is composed of a Graph Feature Extraction module and a Graph Contextual Reasoning module, achieving appropriate information-redundancy feature extraction, structured organization, and hierarchical inferring. Our approach achieves superior or competitive performance over previous methods on the 3DSSG dataset, especially on the relationship prediction sub-task.
Chengjiang Long, Zhaoxuan Zhang, Bokai Liu, Qiang Zhang 0008, Xin Yang 0011
IEEE Trans. Vis. Comput. Graph.3
2020 Point cloud semantic scene segmentation based on coordinate convolution
abstract
Abstract Point cloud semantic segmentation, a crucial research area in the 3D computer vision, lies at the core of many vision and robotics applications. Due to the irregular and disordered of the point cloud, however, the application of convolution on point clouds is challenging. In this article, we propose the “coordinate convolution,” which can effectively extract local structural information of the point cloud, to solve the inapplicability of conventional convolution neural network (CNN) structures on the 3D point cloud. The “coordinate convolution” is a projection operation of three planes based on the local coordinate system of each point. Specifically, we project the point cloud on three planes in the local coordinate system with a joint 2D convolution operation to extract its features. Additionally, we leverage a self‐encoding network based on image semantic segmentation U‐Net structure as the overall architecture of the point cloud semantic segmentation algorithm. The results demonstrate that the proposed method exhibited excellent performances for point cloud data sets corresponding to various scenes.
Zhaoxuan Zhang, Xuefeng Yin, Xinglin Piao, Yuxin Wang 0001, Xin Yang 0011
Comput. Animat. Virtual Worlds1
2019 Deep Reinforcement Learning of Volume-Guided Progressive View Inpainting for 3D Point Scene Completion From a Single Depth Image
abstract
We present a deep reinforcement learning method of progressive view inpainting for 3D point scene completion under volume guidance, achieving high-quality scene reconstruction from only a single depth image with severe occlusion. Our approach is end-to-end, consisting of three modules: 3D scene volume reconstruction, 2D depth map inpainting, and multi-view selection for completion. Given a single depth image, our method first goes through the 3D volume branch to obtain a volumetric scene reconstruction as a guide to the next view inpainting step, which attempts to make up the missing information; the third step involves projecting the volume under the same view of the input, concatenating them to complete the current view depth, and integrating all depth into the point cloud. Since the occluded areas are unavailable, we resort to a deep Q-Network to glance around and pick the next best view for large hole completion progressively until a scene is adequately reconstructed while guaranteeing validity. All steps are learned jointly to achieve robust and consistent results. We perform qualitative and quantitative evaluations with extensive experiments on the SUNCG data, obtaining better results than the state of the art.
Xiaoguang Han 0001, Zhaoxuan Zhang, Dong Du 0002, Mingdai Yang, Jingming Yu, Xin Yang 0011, Ligang Liu 0001, Zixiang Xiong, Shuguang Cui
CVPR2