Wei Cao 0015

dblp:54/6265-15 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0005-5163-6484ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 77% Generative modeling · 14% Video understanding and tracking · 10%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d reconstruction
non-rigid reconstruction
1.822026
Motion2VecSets: Non-Rigid Shape Reconstruction and Tracking With 4D Latent Set Diffusion · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Motion2VecSets: 4D Latent Vector Set Diffusion for Non-Rigid Shape Reconstruction and Tracking · CVPR 2024
Computer vision › 3D vision › 3d reconstruction › dynamic 3d reconstruction
dynamic mesh reconstruction
1.012026
Motion2VecSets: Non-Rigid Shape Reconstruction and Tracking With 4D Latent Set Diffusion · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › 3D vision
3d reconstruction
0.812024
Motion2VecSets: 4D Latent Vector Set Diffusion for Non-Rigid Shape Reconstruction and Tracking · CVPR 2024
Machine learning › Generative modeling
diffusion model
0.812024
Motion2VecSets: 4D Latent Vector Set Diffusion for Non-Rigid Shape Reconstruction and Tracking · CVPR 2024
Computer vision › 3D vision › 3d reconstruction › dynamic 3d reconstruction
dynamic surface reconstruction
0.812024
Motion2VecSets: 4D Latent Vector Set Diffusion for Non-Rigid Shape Reconstruction and Tracking · CVPR 2024
Computer vision › Video understanding and tracking › object tracking
non-rigid object tracking
0.312026
Motion2VecSets: Non-Rigid Shape Reconstruction and Tracking With 4D Latent Set Diffusion · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Video understanding and tracking
object tracking
0.212024
Motion2VecSets: 4D Latent Vector Set Diffusion for Non-Rigid Shape Reconstruction and Tracking · CVPR 2024

Methods — techniques the papers use, named apart from their topics

interleaved spatial-temporal attention · 1.04d latent set diffusion · 1.0latent vector set diffusion · 0.8iterative denoising · 0.8interleaved space-time attention · 0.8
YearPublicationVenuePosition
2026 Motion2VecSets: Non-Rigid Shape Reconstruction and Tracking With 4D Latent Set Diffusion
abstract
We introduce Motion2VecSets, a 4D diffusion model for dynamic surface mesh generation from various ambiguous observations, including a sequence of RGB images, sparse and partial point clouds, and low-resolution voxel grids. While recent methods using neural field representations have shown success in modeling non-rigid objects, conventional feed-forward architectures struggle with noisy, partial, or sparse observations due to their deterministic nature. To address the inherent one-to-many mapping problem, we introduce a diffusion model that explicitly learns the shape and motion distribution of non-rigid objects through an iterative denoising process of compressed latent representations. The diffusion-based priors provide more plausible and diverse reconstructions under ambiguous conditions. Instead of relying on global latent codes, we represent 4D dynamics using latent sets. This novel 4D representation captures local shape and deformation patterns, leading to more accurate non-linear motion capture and significantly improving generalization capacity to unseen motions and identities. For temporally coherent tracking, we jointly denoise latent sets across frames and enable cross-frame information exchange. To reduce computational cost, we design an interleaved spatial-temporal attention block that alternately aggregates deformation latents along spatial and temporal dimensions. Extensive experiments on datasets of humans, animals, and articulated objects demonstrate that Motion2VecSets outperforms prior methods in reconstructing and tracking non-rigid deformations from various imperfect observations.
Jiapeng Tang, Wei Cao 0015, Biao Zhang 0005, Chang Luo, Yaoyao Liu 0001, Matthias Nießner
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Predicting the Road Ahead: A Knowledge Graph Based Foundation Model for Scene Understanding in Autonomous Driving
Stefan Schimid, Yicong Li 0013, Lavdim Halilaj, Xiangtong Yao, Wei Cao 0015
ESWC (1)6
2024 Motion2VecSets: 4D Latent Vector Set Diffusion for Non-Rigid Shape Reconstruction and Tracking
abstract
We introduce Motion2VecSets, a 4D diffusion model for dynamic surface reconstruction from point cloud sequences. While existing state-of-the-art methods have demonstrated success in reconstructing non-rigid objects using neural field representations, conventional feed-forward networks encounter challenges with ambiguous observations from noisy, partial, or sparse point clouds. To address these challenges, we introduce a diffusion model that explicitly learns the shape and motion distribution of non-rigid objects through an iterative denoising process of compressed latent representations. The diffusion-based priors enable more plausible and probabilistic reconstructions when handling ambiguous inputs. We parameterize 4D dynamics with latent sets instead of using global latent codes. This novel 4D representation allows us to learn local shape and deformation patterns, leading to more accurate nonlinear motion capture and significantly improving generalizability to unseen motions and identities. For more temporally-coherent object tracking, we synchronously denoise deformation latent sets and exchange information across multiple frames. To avoid computational overhead, we designed an interleaved space and time attention block to alternately aggregate deformation latents along spatial and temporal domains. Extensive comparisons against state-of-the-art methods demonstrate the superiority of our Motion2VecSets in 4D reconstruction from various imperfect observations.
Wei Cao 0015, Chang Luo, Biao Zhang 0005, Matthias Nießner, Jiapeng Tang
CVPR1