Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xin Wang 0178

dblp:10/5630-178 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0002-6789-4569ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 38% Segmentation and scene understanding · 36% Generative modeling · 12%
Computer graphics and multimedia
3 papers
Geometric modeling and processing · 54% Visual content generation and editing · 30% Rendering · 16%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
3d semantic segmentation
1.932025
Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation of Indoor Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2025
LiDAL: Inter-frame Uncertainty Based Active Learning for 3D LiDAR Semantic Segmentation · ECCV (27) 2022
VMNet: Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation · ICCV 2021
Computer vision › Segmentation and scene understanding › semantic segmentation
indoor scene segmentation
1.422025
Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation of Indoor Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2025
VMNet: Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation · ICCV 2021
Visual content generation and editing
3d content creation
1.322026
MuMA: 3D PBR Texturing via Multi-Channel Multi-View Generation and Albedo Post-Processing · IEEE Trans. Image Process. 2026
Light-SQ: Structure-aware Shape Abstraction with Superquadrics for Generated Meshes · SIGGRAPH Asia 2025
Rendering
physically based rendering
1.012026
MuMA: 3D PBR Texturing via Multi-Channel Multi-View Generation and Albedo Post-Processing · IEEE Trans. Image Process. 2026
Computer vision › 3D vision
3d content generation
0.912025
MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generation · CVPR 2025
Machine learning › Generative modeling
autoregressive model
0.912025
MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generation · CVPR 2025
Geometric modeling and processing
mesh processing
0.912025
Light-SQ: Structure-aware Shape Abstraction with Superquadrics for Generated Meshes · SIGGRAPH Asia 2025
Geometric modeling and processing › shape representation
shape abstraction
0.912025
Light-SQ: Structure-aware Shape Abstraction with Superquadrics for Generated Meshes · SIGGRAPH Asia 2025
Geometric modeling and processing
shape decomposition
0.912025
Light-SQ: Structure-aware Shape Abstraction with Superquadrics for Generated Meshes · SIGGRAPH Asia 2025
Geometric modeling and processing › surface fitting
superquadric fitting
0.912025
Light-SQ: Structure-aware Shape Abstraction with Superquadrics for Generated Meshes · SIGGRAPH Asia 2025
Computer vision › 3D vision
3d scene understanding
0.822025
VMNet: Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation · ICCV 2021
Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation of Indoor Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › 3D vision
3d face reconstruction
0.712023
JR2Net: Joint Monocular 3D Face Reconstruction and Reenactment · AAAI 2023
Computer vision › Face, body and person analysis › facial expression analysis
facial expression tracking
0.712023
JR2Net: Joint Monocular 3D Face Reconstruction and Reenactment · AAAI 2023
Computer vision › 3D vision › 3d face reconstruction
single-image 3d face reconstruction
0.712023
JR2Net: Joint Monocular 3D Face Reconstruction and Reenactment · AAAI 2023
Visual content generation and editing › face editing
face reenactment
0.712023
JR2Net: Joint Monocular 3D Face Reconstruction and Reenactment · AAAI 2023
Machine learning › Efficient and distributed learning
active learning
0.612022
LiDAL: Inter-frame Uncertainty Based Active Learning for 3D LiDAR Semantic Segmentation · ECCV (27) 2022
Computer vision › 3D vision › point cloud segmentation
point cloud semantic segmentation
0.612022
LiDAL: Inter-frame Uncertainty Based Active Learning for 3D LiDAR Semantic Segmentation · ECCV (27) 2022
Machine learning › Generative modeling
variational autoencoder
0.312025
MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generation · CVPR 2025

Methods — techniques the papers use, named apart from their topics

feature fusion · 1.4self-supervised learning · 1.3cascade framework · 1.3multimodal large language model · 1.0multi-view generation · 1.0intrinsic decomposition · 1.0volumetric decomposition · 0.9vector quantization · 0.9sparse voxel CNN · 0.9signed distance field carving · 0.9residual pruning · 0.9optimization · 0.9mesh representation · 0.9masked autoregressive transformer · 0.9cascaded training · 0.9attention module · 0.9active learning · 0.6
YearPublicationVenuePosition
2026 MuMA: 3D PBR Texturing via Multi-Channel Multi-View Generation and Albedo Post-Processing
abstract
Current methods for 3D generation still fall short in physically based rendering (PBR) texturing, primarily due to limited data and challenges in modeling multi-channel materials. In this work, we propose MuMA, a method for 3D PBR texturing through Multi-channel Multi-view generation and Albedo post-processing. Our approach features two key innovations: 1) we opt to model shaded and albedo appearance channels, where the shaded channels enables the integration intrinsic decomposition modules for material properties; and 2) leveraging multimodal large language models, we emulate artists' techniques for material assessment and selection. Experiments demonstrate that MuMA achieves superior results in visual quality and material fidelity compared to existing methods.
Lingting Zhu, Jingrui Ye, Zeyu Hu, Yingda Yin, Lanjiong Li, Jinnan Chen, Shengju Qian, Xin Wang 0178, Qingmin Liao, Lequan Yu
IEEE Trans. Image Process.9
2025 MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generation
abstract
Recent advances in auto-regressive transformers have revolutionized generative modeling across different domains, from language processing to visual generation, demonstrating remarkable capabilities. However, applying these advances to 3D generation presents three key challenges: the unordered nature of 3D data conflicts with sequential next-token prediction paradigm, conventional vector quantization approaches incur substantial compression loss when applied to 3D meshes, and the lack of efficient scaling strategies for higher resolution latent prediction. To address these challenges, we introduce MAR-3D, which integrates a pyramid variational autoencoder with a cascaded masked auto-regressive transformer (Cascaded MAR) for progressive latent upscaling in the continuous space. Our architecture employs random masking during training and auto-regressive denoising in random order during inference, naturally accommodating the unordered property of 3D latent tokens. Additionally, we propose a cascaded training strategy with condition augmentation that enables efficiently up-scale the latent token resolution with fast convergence. Extensive experiments demonstrate that MAR-3D not only achieves superior performance and generalization capabilities compared to existing methods but also exhibits enhanced scaling capabilities compared to joint distribution modeling approaches (e.g., diffusion transformers).
Jinnan Chen, Lingting Zhu, Zeyu Hu, Shengju Qian, Yugang Chen, Xin Wang 0178, Gim Hee Lee
CVPR6
2025 Light-SQ: Structure-aware Shape Abstraction with Superquadrics for Generated Meshes
abstract
In user-generated-content (UGC) applications, non-expert users often rely on image-to-3D generative models to create 3D assets. In this context, primitive-based shape abstraction offers a promising solution for UGC scenarios by compressing high-resolution meshes into compact, editable representations. Towards this end, effective shape abstraction must therefore be structure-aware, characterized by low overlap between primitives, part-aware alignment, and primitive compactness. We present Light-SQ, a novel superquadric-based optimization framework that explicitly emphasizes structure-awareness from three aspects. (a) We introduce SDF carving to iteratively udpate the target signed distance field, discouraging overlap between primitives. (b) We propose a block-regrow-fill strategy guided by structure-aware volumetric decomposition, enabling structural partitioning to drive primitive placement. (c) We implement adaptive residual pruning based on SDF update history to surpress over-segmentation and ensure compact results. In addition, Light-SQ supports multiscale fitting, enabling localized refinement to preserve fine geometric details. To evaluate our method, we introduce 3DGen-Prim, a benchmark extending 3DGen-Bench with new metrics for both reconstruction quality and primitive-level editability. Extensive experiments demonstrate that Light-SQ enables efficient, high-fidelity, and editable shape abstraction with superquadrics for complex generated geometry, advancing the feasibility of 3D UGC creation. Project Page: https://johann.wang/Light-SQ/ .
Yuhan Wang 0002, Weikai Chen 0001, Zeyu Hu, Yingda Yin, Keyang Luo, Shengju Qian, Yiyan Ma, Yuhuan Zhou, Hao Luo 0001, Wan Wang, Xiaobin Shen 0004, Kuixin Zhu, Chuanlang Hong, Lijie Feng, Xin Wang 0178, Chen Change Loy
SIGGRAPH Asia21
2025 Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation of Indoor Scenes
abstract
In recent years, sparse voxel-based methods have become the state-of-the-arts for 3D semantic segmentation of indoor scenes, thanks to the powerful 3D CNNs. Nevertheless, being oblivious to the underlying geometry, voxel-based methods suffer from ambiguous features on spatially close objects and struggle with handling complex and irregular geometries due to the lack of geodesic information. In view of this, we present Voxel-Mesh Network (VMNet), a novel 3D deep architecture that operates on the voxel and mesh representations leveraging both the euclidean and geodesic information. Intuitively, the euclidean information extracted from voxels can offer contextual cues representing interactions between nearby objects, while the geodesic information extracted from meshes can help separate objects that are spatially close but have disconnected surfaces. To incorporate such information from the two domains, we design an intra-domain attentive module for effective feature aggregation and an inter-domain attentive module for adaptive feature fusion. Experimental results validate the effectiveness of VMNet: specifically, on the challenging ScanNet dataset for large-scale segmentation of indoor scenes, it outperforms the state-of-the-art SparseConvNet and MinkowskiNet (74.6% versus 72.5% and 73.6% in mIoU) with a simpler network structure (17M versus 30M and 38M parameters).
Zeyu Hu, Xuyang Bai, Jiaxiang Shang, Jiayu Dong, Xin Wang 0178, Guangyuan Sun, Hongbo Fu 0001, Chiew-Lan Tai
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 JR2Net: Joint Monocular 3D Face Reconstruction and Reenactment
abstract
Face reenactment and reconstruction benefit various applications in self-media, VR, etc. Recent face reenactment methods use 2D facial landmarks to implicitly retarget facial expressions and poses from driving videos to source images, while they suffer from pose and expression preservation issues for cross-identity scenarios, i.e., when the source and the driving subjects are different. Current self-supervised face reconstruction methods also demonstrate impressive results. However, these methods do not handle large expressions well, since their training data lacks samples of large expressions, and 2D facial attributes are inaccurate on such samples. To mitigate the above problems, we propose to explore the inner connection between the two tasks, i.e., using face reconstruction to provide sufficient 3D information for reenactment, and synthesizing videos paired with captured face model parameters through face reenactment to enhance the expression module of face reconstruction. In particular, we propose a novel cascade framework named JR2Net for Joint Face Reconstruction and Reenactment, which begins with the training of a coarse reconstruction network, followed by a 3D-aware face reenactment network based on the coarse reconstruction results. In the end, we train an expression tracking network based on our synthesized videos composed by image-face model parameter pairs. Such an expression tracking network can further enhance the coarse face reconstruction. Extensive experiments show that our JR2Net outperforms the state-of-the-art methods on several face reconstruction and reenactment benchmarks.
Jiaxiang Shang, Xin Wang 0178, Guangyuan Sun, Hongbo Fu 0001
AAAI4
2022 LiDAL: Inter-frame Uncertainty Based Active Learning for 3D LiDAR Semantic Segmentation
Zeyu Hu, Xuyang Bai, Xin Wang 0178, Guangyuan Sun, Hongbo Fu 0001, Chiew-Lan Tai
ECCV (27)4
2021 VMNet: Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation
abstract
In recent years, sparse voxel-based methods have be-come the state-of-the-arts for 3D semantic segmentation of indoor scenes, thanks to the powerful 3D CNNs. Nevertheless, being oblivious to the underlying geometry, voxel-based methods suffer from ambiguous features on spatially close objects and struggle with handling complex and irregular geometries due to the lack of geodesic information. In view of this, we present Voxel-Mesh Network (VMNet), a novel 3D deep architecture that operates on the voxel and mesh representations leveraging both the Euclidean and geodesic information. Intuitively, the Euclidean information extracted from voxels can offer contextual cues representing interactions between nearby objects, while the geodesic information extracted from meshes can help separate objects that are spatially close but have disconnected surfaces. To incorporate such information from the two domains, we design an intra-domain attentive module for effective feature aggregation and an inter-domain attentive module for adaptive feature fusion. Experimental results validate the effectiveness of VMNet: specifically, on the challenging ScanNet dataset for large-scale segmentation of indoor scenes, it outperforms the state-of-the-art SparseConvNet and MinkowskiNet (74.6% vs 72.5% and 73.6% in mIoU) with a simpler network structure (17M vs 30M and 38M parameters). Code release: https://github.com/hzykent/VMNet
Zeyu Hu, Xuyang Bai, Jiaxiang Shang, Jiayu Dong, Xin Wang 0178, Guangyuan Sun, Hongbo Fu 0001, Chiew-Lan Tai
ICCV6