Dave Zhenyu Chen

dblp:255/5913 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-3883-1905ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs
abstract
Despite recent successes in novel view synthesis using 3D Gaussian Splatting (3DGS), modeling scenes with sparse inputs remains a challenge. In this work, we address two critical yet overlooked issues in real-world sparse-input modeling: extrapolation and occlusion. To tackle these issues, we propose to use a reconstruction by generation pipeline that leverages learned priors from video diffusion models to provide plausible interpretations for regions outside the field of view or occluded. However, the generated sequences exhibit inconsistencies that do not fully benefit subsequent 3DGS modeling. To address the challenge of inconsistencies, we introduce a novel scene-grounding guidance based on rendered sequences from an optimized 3DGS, which tames the diffusion model to generate consistent sequences. This guidance is training-free and does not require any fine-tuning of the diffusion model. To facilitate holistic scene modeling, we also propose a trajectory initialization method. It effectively identifies regions that are outside the field of view and occluded. We further design a scheme tailored for 3DGS optimization with generated sequences. Experiments demonstrate that our method significantly improves upon the baseline and achieves state-of-the-art performance on challenging benchmarks.
Yingji Zhong, Zhihao Li 0002, Dave Zhenyu Chen, Lanqing Hong, Dan Xu 0002
CVPR3
2024 SceneTex: High-Quality Texture Synthesis for Indoor Scenes via Diffusion Priors
abstract
We propose SceneTex, a novel method for effectively gen-erating high-quality and style-consistent textures for indoor scenes using depth-to-image diffusion priors. Unlike pre-vious methods that either iteratively warp 2D views onto a mesh surface or distillate diffusion latent features with-out accurate geometric and style cues, SceneTexformulates the texture synthesis task as an optimization problem in the RGB space where style and geometry consistency are prop-erly reflected. At its core, SceneTex proposes a multires-olution texture field to implicitly encode the mesh appear-ance. We optimize the target texture via a score-distillation-based objective function in respective RGB renderings. To further secure the style consistency across views, we introduce a cross-attention decoder to predict the RGB values by cross-attending to the pre-sampled reference locations in each instance. SceneTex enables various and accurate texture synthesis for 3D-FRONT scenes, demonstrating sig-nificant improvements in visual quality and prompt fidelity over the prior texture generation methods.
Dave Zhenyu Chen, Hsin-Ying Lee 0001, Sergey Tulyakov, Matthias Nießner
CVPR1
2024 EchoScene: Indoor Scene Generation via Information Echo Over Scene Graph Diffusion
Guangyao Zhai, Evin Pinar Örnek, Dave Zhenyu Chen, Ruotong Liao, Yan Di, Nassir Navab, Federico Tombari, Benjamin Busam
ECCV (21)3
2023 Generating Context-Aware Natural Answers for Questions in 3D Scenes
Mohammed Munzer Dwedari, Matthias Nießner, Dave Zhenyu Chen
BMVC3
2023 UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding
abstract
Performing 3D dense captioning and visual grounding requires a common and shared understanding of the underlying multimodal relationships. However, despite some previous attempts on connecting these two related tasks with highly task-specific neural modules, it remains understudied how to explicitly depict their shared nature to learn them simultaneously. In this work, we propose UniT3D, a simple yet effective fully unified transformer-based architecture for jointly solving 3D visual grounding and dense captioning. UniT3D enables learning a strong multimodal representation across the two tasks through a supervised joint pre-training scheme with bidirectional and seq-to-seq objectives. With a generic architecture design, UniT3D allows expanding the pre-training scope to more various training sources such as the synthesized data from 2D prior knowledge to benefit 3D vision-language tasks. Extensive experiments and analysis demonstrate that UniT3D obtains significant gains for 3D dense captioning and visual grounding.
Dave Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner, Angel X. Chang
ICCV1
2023 Text2Tex: Text-driven Texture Synthesis via Diffusion Models
abstract
We present Text2Tex, a novel method for generating high-quality textures for 3D meshes from the given text prompts. Our method incorporates inpainting into a pre-trained depth-aware image diffusion model to progressively synthesize high resolution partial textures from multiple viewpoints. To avoid accumulating inconsistent and stretched artifacts across views, we dynamically segment the rendered view into a generation mask, which represents the generation status of each visible texel. This partitioned view representation guides the depth-aware inpainting model to generate and update partial textures for the corresponding regions. Furthermore, we propose an automatic view sequence generation scheme to determine the next best view for updating the partial texture. Extensive experiments demonstrate that our method significantly outperforms the existing text-driven approaches and GAN-based methods.
Dave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee 0001, Sergey Tulyakov, Matthias Nießner
ICCV1
2023 Federated Learning via Decentralized Dataset Distillation in Resource-Constrained Edge Environments
abstract
In federated learning, all networked clients contribute to the model training cooperatively. However, with model sizes increasing, even sharing the trained partial models often leads to severe communication bottlenecks in underlying networks, especially when communicated iteratively. In this paper, we introduce a federated learning framework FedD3 requiring only one-shot communication by integrating dataset distillation instances. Instead of sharing model updates in other federated learning approaches, FedD3 allows the connected clients to distill the local datasets independently, and then aggregates those decentralized distilled datasets (e.g. a few unrecognizable images) from networks for model training. Our experimental results show that FedD3 significantly outperforms other federated learning frameworks in terms of needed communication volumes, while it provides the additional benefit to be able to balance the trade-off between accuracy and communication cost, depending on usage scenario or target dataset. For instance, for training an AlexNet model on CIFAR-10 with 10 clients under non-independent and identically distributed (Non-IID) setting, FedD3 can either increase the accuracy by over 71% with a similar communication volume, or save 98% of communication volume, while reaching the same accuracy, compared to other one-shot federated learning approaches.
Rui Song 0007, Dai Liu, Dave Zhenyu Chen, Andreas Festag, Carsten Trinitis, Martin Schulz 0001, Alois C. Knoll
IJCNN3
2022 D3Net: A Unified Speaker-Listener Architecture for 3D Dense Captioning and Visual Grounding
Dave Zhenyu Chen, Qirui Wu, Matthias Nießner, Angel X. Chang
ECCV (32)1
2021 Scan2Cap: Context-Aware Dense Captioning in RGB-D Scans
abstract
We introduce the task of dense captioning in 3D scans from commodity RGB-D sensors. As input, we assume a point cloud of a 3D scene; the expected output is the bounding boxes along with the descriptions for the underlying objects. To address the 3D object detection and description problems, we propose Scan2Cap, an end-to-end trained method, to detect objects in the input scene and describe them in natural language. We use an attention mechanism that generates descriptive tokens while referring to the related components in the local context. To reflect object relations (i.e. relative spatial relations) in the generated captions, we use a message passing graph module to facilitate learning object relation features. Our method can effectively localize and describe 3D objects in scenes from the ScanRefer dataset, outperforming 2D baseline methods by a significant margin (27.61% [email protected] improvement).
Dave Zhenyu Chen, Matthias Nießner, Angel X. Chang
CVPR1
2020 ScanRefer: 3D Object Localization in RGB-D Scans Using Natural Language
Dave Zhenyu Chen, Angel X. Chang, Matthias Nießner
ECCV (20)1