Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Daoyi Gao

dblp:304/3293 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-0458-8107ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 100%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 50% Visual content generation and editing · 50%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing
3d content creation
0.912025
MeshArt: Generating Articulated Meshes with Structure-Guided Transformers · CVPR 2025
Geometric modeling and processing
mesh generation
0.912025
MeshArt: Generating Articulated Meshes with Structure-Guided Transformers · CVPR 2025
Computer vision › 3D vision › 3d shape analysis › 3d shape understanding
CAD model alignment
0.812024
DiffCAD: Weakly-Supervised Probabilistic CAD Model Retrieval and Alignment from an RGB Image · ACM Trans. Graph. 2024
Computer vision › 3D vision › 3d shape analysis › 3d shape retrieval
CAD model retrieval
0.812024
DiffCAD: Weakly-Supervised Probabilistic CAD Model Retrieval and Alignment from an RGB Image · ACM Trans. Graph. 2024
Computer vision › 3D vision
object pose estimation
0.812024
S2P3: Self-Supervised Polarimetric Pose Prediction · Int. J. Comput. Vis. 2024
Computer vision › 3D vision › pose estimation › learning-based pose estimation
self-supervised pose estimation
0.812024
S2P3: Self-Supervised Polarimetric Pose Prediction · Int. J. Comput. Vis. 2024
Computer vision › 3D vision
pose estimation
0.612022
Polarimetric Pose Prediction · ECCV (9) 2022
Computer vision › 3D vision
3d reconstruction
0.212024
S2P3: Self-Supervised Polarimetric Pose Prediction · Int. J. Comput. Vis. 2024
Computer vision › 3D vision › neural rendering
differentiable rendering
0.212024
S2P3: Self-Supervised Polarimetric Pose Prediction · Int. J. Comput. Vis. 2024

Methods — techniques the papers use, named apart from their topics

vector quantization · 0.9transformer · 0.9autoregressive generation · 0.9weakly supervised learning · 0.8self-supervised learning · 0.8monocular depth estimation · 0.8knowledge distillation · 0.8diffusion model · 0.8differentiable rendering · 0.8polarimetry · 0.6
YearPublicationVenuePosition
2025 MeshArt: Generating Articulated Meshes with Structure-Guided Transformers
abstract
Articulated 3D object generation is fundamental for creating realistic, functional, and interactable virtual assets which are not simply static. We introduce MeshArt, a hierarchical transformer-based approach to generate articulated 3D meshes with clean, compact geometry, reminiscent of human-crafted 3D models. We approach articulated mesh generation in a part-by-part fashion across two stages. First, we generate a high-level articulation-aware object structure; then, based on this structural information, we synthesize each part’s mesh faces. Key to our approach is modeling both articulation structures and part meshes as sequences of quantized triangle embeddings, leading to a unified hierarchical framework with transformers for autoregressive generation. Object part structures are first generated as their bounding primitives and articulation modes; a second transformer, guided by these articulation structures, then generates each part’s mesh triangles. To ensure coherency among generated parts, we introduce structure-guided conditioning that also incorporates local part mesh connectivity. MeshArt shows significant improvements over state of the art, with 57.1% improvement in structure coverage and a 209-point improvement in mesh generation FID.
Daoyi Gao, Yawar Siddiqui, Lei Li 0038, Angela Dai
CVPR1
2024 S2P3: Self-Supervised Polarimetric Pose Prediction
abstract
Abstract This paper proposes the first self-supervised 6D object pose prediction from multimodal RGB + polarimetric images. The novel training paradigm comprises (1) a physical model to extract geometric information of polarized light, (2) a teacher–student knowledge distillation scheme and (3) a self-supervised loss formulation through differentiable rendering and an invertible physical constraint. Both networks leverage the physical properties of polarized light to learn robust geometric representations by encoding shape priors and polarization characteristics derived from our physical model. Geometric pseudo-labels from the teacher support the student network without the need for annotated real data. Dense appearance and geometric information of objects are obtained through a differentiable renderer with the predicted pose for self-supervised direct coupling. The student network additionally features our proposed invertible formulation of the physical shape priors that enables end-to-end self-supervised training through physical constraints of derived polarization characteristics compared against polarimetric input images. We specifically focus on photometrically challenging objects with texture-less or reflective surfaces and transparent materials for which the most prominent performance gain is reported.
Patrick Ruhkamp, Daoyi Gao, Nassir Navab, Benjamin Busam
Int. J. Comput. Vis.2
2024 DiffCAD: Weakly-Supervised Probabilistic CAD Model Retrieval and Alignment from an RGB Image
abstract
Perceiving 3D structures from RGB images based on CAD model primitives can enable an effective, efficient 3D object-based representation of scenes. However, current approaches rely on supervision from expensive yet imperfect annotations of CAD models associated with real images, and encounter challenges due to the inherent ambiguities in the task - both in depth-scale ambiguity in monocular perception, as well as inexact matches of CAD database models to real observations. We thus propose DiffCAD, the first weakly-supervised probabilistic approach to CAD retrieval and alignment from an RGB image. We learn a probabilistic model through diffusion, modeling likely distributions of shape, pose, and scale of CAD objects in an image. This enables multi-hypothesis generation of different plausible CAD reconstructions, requiring only a few hypotheses to characterize ambiguities in depth/scale and inexact shape matches. Our approach is trained only on synthetic data, leveraging monocular depth and mask estimates to enable robust zero-shot adaptation to various real target domains. Despite being trained solely on synthetic data, our multi-hypothesis approach can even surpass the supervised state-of-the-art on the Scan2CAD dataset by 5.9% with 8 hypotheses.
Daoyi Gao, Dávid Rozenberszki, Stefan Leutenegger, Angela Dai
ACM Trans. Graph.1
2022 Polarimetric Pose Prediction
Daoyi Gao, Patrick Ruhkamp, Iuliia Skobleva, Magdalena Wysocki, Pengyuan Wang 0002, Arturo Guridi, Benjamin Busam
ECCV (9)1
2021 Attention meets Geometry: Geometry Guided Spatial-Temporal Attention for Consistent Self-Supervised Monocular Depth Estimation
abstract
Inferring geometrically consistent dense 3D scenes across a tuple of temporally consecutive images remains challenging for self-supervised monocular depth prediction pipelines. This paper explores how the increasingly popular transformer architecture, together with novel regularized loss formulations, can improve depth consistency while preserving accuracy. We propose a spatial attention module that correlates coarse depth predictions to aggregate local geometric information. A novel temporal attention mechanism further processes the local geometric information in a global context across consecutive images. Additionally, we introduce geometric constraints between frames regularized by photometric cycle consistency. By combining our proposed regularization and the novel spatial-temporal-attention module we fully leverage both the geometric and appearance-based consistency across monocular frames. This yields geometrically meaningful attention and improves temporal depth stability and accuracy compared to previous methods.
Patrick Ruhkamp, Daoyi Gao, Hanzhi Chen, Nassir Navab, Benjamin Busam
3DV2