VLDB 2026 Research / reviewers in the wild / expert
Daoyi Gao
dblp:304/3293
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-0458-8107ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
3D vision · 100% | |
| Computer graphics and multimedia
1 paper |
Geometric modeling and processing · 50% Visual content generation and editing · 50% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visual content generation and editing
3d content creation |
0.9 | 1 | 2025 | MeshArt: Generating Articulated Meshes with Structure-Guided Transformers · CVPR 2025 |
Geometric modeling and processing
mesh generation |
0.9 | 1 | 2025 | MeshArt: Generating Articulated Meshes with Structure-Guided Transformers · CVPR 2025 |
Computer vision › 3D vision › 3d shape analysis › 3d shape understanding
CAD model alignment |
0.8 | 1 | 2024 | DiffCAD: Weakly-Supervised Probabilistic CAD Model Retrieval and Alignment from an RGB Image · ACM Trans. Graph. 2024 |
Computer vision › 3D vision › 3d shape analysis › 3d shape retrieval
CAD model retrieval |
0.8 | 1 | 2024 | DiffCAD: Weakly-Supervised Probabilistic CAD Model Retrieval and Alignment from an RGB Image · ACM Trans. Graph. 2024 |
Computer vision › 3D vision
object pose estimation |
0.8 | 1 | 2024 | S2P3: Self-Supervised Polarimetric Pose Prediction · Int. J. Comput. Vis. 2024 |
Computer vision › 3D vision › pose estimation › learning-based pose estimation
self-supervised pose estimation |
0.8 | 1 | 2024 | S2P3: Self-Supervised Polarimetric Pose Prediction · Int. J. Comput. Vis. 2024 |
Computer vision › 3D vision
pose estimation |
0.6 | 1 | 2022 | Polarimetric Pose Prediction · ECCV (9) 2022 |
Computer vision › 3D vision
3d reconstruction |
0.2 | 1 | 2024 | S2P3: Self-Supervised Polarimetric Pose Prediction · Int. J. Comput. Vis. 2024 |
Computer vision › 3D vision › neural rendering
differentiable rendering |
0.2 | 1 | 2024 | S2P3: Self-Supervised Polarimetric Pose Prediction · Int. J. Comput. Vis. 2024 |
Methods — techniques the papers use, named apart from their topics
vector quantization · 0.9transformer · 0.9autoregressive generation · 0.9weakly supervised learning · 0.8self-supervised learning · 0.8monocular depth estimation · 0.8knowledge distillation · 0.8diffusion model · 0.8differentiable rendering · 0.8polarimetry · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MeshArt: Generating Articulated Meshes with Structure-Guided TransformersabstractArticulated 3D object generation is fundamental for creating realistic, functional, and interactable virtual assets which are not simply static. We introduce MeshArt, a hierarchical transformer-based approach to generate articulated 3D meshes with clean, compact geometry, reminiscent of human-crafted 3D models. We approach articulated mesh generation in a part-by-part fashion across two stages. First, we generate a high-level articulation-aware object structure; then, based on this structural information, we synthesize each part’s mesh faces. Key to our approach is modeling both articulation structures and part meshes as sequences of quantized triangle embeddings, leading to a unified hierarchical framework with transformers for autoregressive generation. Object part structures are first generated as their bounding primitives and articulation modes; a second transformer, guided by these articulation structures, then generates each part’s mesh triangles. To ensure coherency among generated parts, we introduce structure-guided conditioning that also incorporates local part mesh connectivity. MeshArt shows significant improvements over state of the art, with 57.1% improvement in structure coverage and a 209-point improvement in mesh generation FID. Daoyi Gao, Yawar Siddiqui, Lei Li 0038, Angela Dai |
CVPR | 1 |
| 2024 | S2P3: Self-Supervised Polarimetric Pose PredictionabstractAbstract This paper proposes the first self-supervised 6D object pose prediction from multimodal RGB + polarimetric images. The novel training paradigm comprises (1) a physical model to extract geometric information of polarized light, (2) a teacher–student knowledge distillation scheme and (3) a self-supervised loss formulation through differentiable rendering and an invertible physical constraint. Both networks leverage the physical properties of polarized light to learn robust geometric representations by encoding shape priors and polarization characteristics derived from our physical model. Geometric pseudo-labels from the teacher support the student network without the need for annotated real data. Dense appearance and geometric information of objects are obtained through a differentiable renderer with the predicted pose for self-supervised direct coupling. The student network additionally features our proposed invertible formulation of the physical shape priors that enables end-to-end self-supervised training through physical constraints of derived polarization characteristics compared against polarimetric input images. We specifically focus on photometrically challenging objects with texture-less or reflective surfaces and transparent materials for which the most prominent performance gain is reported. Patrick Ruhkamp, Daoyi Gao, Nassir Navab, Benjamin Busam |
Int. J. Comput. Vis. | 2 |
| 2024 | DiffCAD: Weakly-Supervised Probabilistic CAD Model Retrieval and Alignment from an RGB ImageabstractPerceiving 3D structures from RGB images based on CAD model primitives can enable an effective, efficient 3D object-based representation of scenes. However, current approaches rely on supervision from expensive yet imperfect annotations of CAD models associated with real images, and encounter challenges due to the inherent ambiguities in the task - both in depth-scale ambiguity in monocular perception, as well as inexact matches of CAD database models to real observations. We thus propose DiffCAD, the first weakly-supervised probabilistic approach to CAD retrieval and alignment from an RGB image. We learn a probabilistic model through diffusion, modeling likely distributions of shape, pose, and scale of CAD objects in an image. This enables multi-hypothesis generation of different plausible CAD reconstructions, requiring only a few hypotheses to characterize ambiguities in depth/scale and inexact shape matches. Our approach is trained only on synthetic data, leveraging monocular depth and mask estimates to enable robust zero-shot adaptation to various real target domains. Despite being trained solely on synthetic data, our multi-hypothesis approach can even surpass the supervised state-of-the-art on the Scan2CAD dataset by 5.9% with 8 hypotheses. Daoyi Gao, Dávid Rozenberszki, Stefan Leutenegger, Angela Dai |
ACM Trans. Graph. | 1 |
| 2022 | Polarimetric Pose Prediction
Daoyi Gao, Patrick Ruhkamp, Iuliia Skobleva, Magdalena Wysocki, Pengyuan Wang 0002, Arturo Guridi, Benjamin Busam |
ECCV (9) | 1 |
| 2021 | Attention meets Geometry: Geometry Guided Spatial-Temporal Attention for Consistent Self-Supervised Monocular Depth EstimationabstractInferring geometrically consistent dense 3D scenes across a tuple of temporally consecutive images remains challenging for self-supervised monocular depth prediction pipelines. This paper explores how the increasingly popular transformer architecture, together with novel regularized loss formulations, can improve depth consistency while preserving accuracy. We propose a spatial attention module that correlates coarse depth predictions to aggregate local geometric information. A novel temporal attention mechanism further processes the local geometric information in a global context across consecutive images. Additionally, we introduce geometric constraints between frames regularized by photometric cycle consistency. By combining our proposed regularization and the novel spatial-temporal-attention module we fully leverage both the geometric and appearance-based consistency across monocular frames. This yields geometrically meaningful attention and improves temporal depth stability and accuracy compared to previous methods. Patrick Ruhkamp, Daoyi Gao, Hanzhi Chen, Nassir Navab, Benjamin Busam |
3DV | 2 |