Chandan Yeshwanth

dblp:195/5814 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 44% Vision and language · 33% Graph learning · 14%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d scene understanding
1.522025
ExCap3d: Expressive 3D Scene Understanding via Object Captioning with Varying Detail · ICCV 2025
ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes · ICCV 2023
Computer vision › Vision and language › 3d vision and language
3d captioning
0.912025
ExCap3d: Expressive 3D Scene Understanding via Object Captioning with Varying Detail · ICCV 2025
Computer vision › Vision and language › image captioning › grounded image captioning
object captioning
0.912025
ExCap3d: Expressive 3D Scene Understanding via Object Captioning with Varying Detail · ICCV 2025
Computer vision › Vision and language
vision-language model
0.912025
ExCap3d: Expressive 3D Scene Understanding via Object Captioning with Varying Detail · ICCV 2025
Computer vision › 3D vision
3d reconstruction
0.712023
ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes · ICCV 2023
Computer vision › 3D vision
3d scene reconstruction
0.712023
ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes · ICCV 2023
Computer vision › Segmentation and scene understanding › scene understanding
indoor scene understanding
0.712023
ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes · ICCV 2023
Computer vision › 3D vision
novel view synthesis
0.712023
ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes · ICCV 2023
Machine learning › Graph learning
graph neural network
0.512021
Directional Message Passing on Molecular Graphs via Synthetic Coordinates · NeurIPS 2021
Machine learning › Graph learning › graph neural network
message passing
0.512021
Directional Message Passing on Molecular Graphs via Synthetic Coordinates · NeurIPS 2021
Bioinformatics and computational biology
molecular property prediction
0.512021
Directional Message Passing on Molecular Graphs via Synthetic Coordinates · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

synthetic coordinates · 1.0distance bounds · 1.0visual language model · 0.9semantic consistency · 0.9part-level description generation · 0.9semantic annotation · 0.7laser scanning · 0.7RGB-D capture · 0.7
YearPublicationVenuePosition
2025 ExCap3d: Expressive 3D Scene Understanding via Object Captioning with Varying Detail
abstract
Generating text descriptions of objects in 3D indoor scenes is an important building block of embodied understanding. Existing methods do this by describing objects at a single level of detail, which often does not capture fine-grained details such as varying textures, materials, and shapes of the parts of objects. We propose the task of expressive 3D captioning: given an input 3D scene, describe objects at multiple levels of detail: a high-level object description, and a low-level description of the properties of its parts. To produce such captions, we present ExCap3D, an expressive 3D captioning model which takes as input a 3D scan, and for each detected object in the scan, generates a fine-grained collective description of the parts of the object, along with an object-level description conditioned on the part-level description. We design ExCap3D to encourage semantic consistency between the generated text descriptions, as well as textual similarity in the latent space, to further increase the quality of the generated captions. To enable this task, we generated the ExCap3D Dataset by leveraging a visual-language model (VLM) for multi-view captioning. The ExCap3D Dataset contains captions on the ScanNet++ dataset with varying levels of detail, comprising 190k text descriptions of 34k 3D objects in 947 indoor scenes. Our experiments show that the object- and part-level of detail captions generated by ExCap3D are of higher quality than those produced by state-of-the-art methods, with a Cider score improvement of 17% and 124% for object- and part-level details respectively. Our code, dataset and models will be made publicly available.
Chandan Yeshwanth, Dávid Rozenberszki, Angela Dai
ICCV1
2023 ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes
abstract
We present ScanNet++, a large-scale dataset that couples together capture of high-quality and commodity-level geometry and color of indoor scenes. Each scene is captured with a high-end laser scanner at sub-millimeter resolution, along with registered 33-megapixel images from a DSLR camera, and RGB-D streams from an iPhone. Scene reconstructions are further annotated with an open vocabulary of semantics, with label-ambiguous scenarios explicitly annotated for comprehensive semantic understanding. ScanNet++ enables a new real-world benchmark for novel view synthesis, both from high-quality RGB capture, and importantly also from commodity-level images, in addition to a new benchmark for 3D semantic scene understanding that comprehensively encapsulates diverse and ambiguous semantic labeling scenarios. Currently, ScanNet++ contains 460 scenes, 280,000 captured DSLR images, and over 3.7M iPhone RGBD frames.
Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela Dai
ICCV1
2021 SceneFormer: Indoor Scene Generation with Transformers
abstract
We address the task of indoor scene generation by generating a sequence of objects, along with their locations and orientations conditioned on a room layout. Large-scale indoor scene datasets allow us to extract patterns from user-designed indoor scenes, and generate new scenes based on these patterns. Existing methods rely on the 2D or 3D appearance of these scenes in addition to object positions, and make assumptions about the possible relations between objects. In contrast, we do not use any appearance information, and implicitly learn object relations using the self-attention mechanism of transformers. We show that our model design leads to faster scene generation with similar or improved levels of realism compared to previous methods. Our method is also flexible, as it can be conditioned not only on the room layout but also on text descriptions of the room, using only the cross-attention mechanism of transformers. Our user study shows that our generated scenes are preferred to the state-of-the-art FastSynth scenes 53.9% and 56.7% of the time for bedroom and living room scenes, respectively. At the same time, we generate a scene in 1.48 seconds on average, 20% faster than FastSynth.
Xinpeng Wang 0003, Chandan Yeshwanth, Matthias Nießner
3DV2
2021 Directional Message Passing on Molecular Graphs via Synthetic Coordinates
abstract
Graph neural networks that leverage coordinates via directional message passing have recently set the state of the art on multiple molecular property prediction tasks. However, they rely on atom position information that is often unavailable, and obtaining it is usually prohibitively expensive or even impossible. In this paper we propose synthetic coordinates that enable the use of advanced GNNs without requiring the true molecular configuration. We propose two distances as synthetic coordinates: Distance bounds that specify the rough range of molecular configurations, and graph-based distances using a symmetric variant of personalized PageRank. To leverage both distance and angular information we propose a method of transforming normal graph neural networks into directional MPNNs. We show that with this transformation we can reduce the error of a normal graph neural network by 55% on the ZINC benchmark. We furthermore set the state of the art on ZINC and coordinate-free QM9 by incorporating synthetic coordinates in the SMP and DimeNet++ models. Our implementation is available online.
Johannes Gasteiger, Chandan Yeshwanth, Stephan Günnemann
NeurIPS2