Shengyi Qian 0001

dblp:250/4431-1 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0003-0262-2412ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LinkGPT: Leveraging Large Language Models for Enhanced Link Prediction in Text-Attributed Graphs
abstract
Inspired by the success of Large Language Models (LLMs) in language and vision tasks, there has been growing interest in applying LLMs to graph tasks, particularly on Text-Attributed Graphs (TAGs). However, most prior work tackles the node classification task. In this work, we evaluate an LLM's ability to reason over structured data and infer new facts based on learned patterns by focusing on link prediction (LP)-the task of predicting missing links between nodes-that is understudied in the literature. This task poses two key challenges: (1) How to effectively integrate pairwise structural information, which is crucial for LP performance, into LLMs, and (2) how to address the computational bottleneck during inference. To tackle these challenges, we propose LinkGPT, the first LLM-based training and inference framework specifically designed for LP on homogeneous TAGs. To enhance the LLM's ability to understand the underlying structure, we carefully design a node encoder and pairwise encoder, and leverage a two-stage instruction tuning to effectively incorporate the nodewise and pairwise information into LLMs. For inference efficiency, we introduce a retrieval-reranking scheme. Extensive experiments show that LinkGPT achieves state-of-the-art performance on real-world graphs and demonstrates superior zero-shot and few-shot generalization. At inference time, it achieves a 10× speedup while maintaining high LP accuracy.
Zhongmou He, Jing Zhu 0005, Shengyi Qian 0001, Joyce Y. Chai, Danai Koutra
CIKM3
2025 3D-MVP: 3D Multiview Pretraining for Manipulation
abstract
Recent works have shown that visual pretraining on ego-centric datasets using masked autoencoders (MAE) can improve generalization for downstream robotics tasks. However, these approaches pretrain only on 2D images, while many robotics applications require 3D scene understanding. In this work, we propose 3D-MVP, a novel approach for 3D Multi-View Pretraining using masked autoencoders. We leverage Robotic View Transformer (RVT), which uses a multi-view transformer to understand the 3D scene and predict gripper pose actions. We split RVT’s multi-view transformer into visual encoder and action decoder, and pretrain its visual encoder using masked autoencoding on large-scale 3D datasets such as Objaverse. We evaluate 3D-MVP on a suite of virtual robot manipulation tasks and demonstrate improved performance over baselines. Our results suggest that 3D-aware pretraining is a promising approach to improve generalization of vision-based robotic manipulation policies.
Shengyi Qian 0001, Kaichun Mo, Valts Blukis, David F. Fouhey, Dieter Fox, Ankit Goyal 0001
CVPR1
2025 Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph Learning
abstract
Graph machine learning has made significant strides in recent years, yet the integration of visual information with graph structure and its potential for improving performance in downstream tasks remains an underexplored area. To address this critical gap, we introduce the Multimodal Graph Benchmark (MM-Graph), a pioneering benchmark that incorporates both visual and textual information into graph learning tasks. MM-Graph extends beyond existing text-attributed graph benchmarks, offering a more comprehensive evaluation framework for multimodal graph learning. Our benchmark comprises seven diverse datasets of varying scales (ranging from thousands to millions of edges), designed to assess algorithms across different tasks in real-world scenarios. These datasets feature rich multimodal node attributes, including visual data, which enables a more holistic evaluation of various graph learning frameworks in complex, multimodal environments. To support advancements in this emerging field, we provide an extensive empirical study on various graph learning frameworks when presented with features from multiple modalities, particularly emphasizing the impact of visual information. This study offers valuable insights into the challenges and opportunities of integrating visual data into graph learning.
Jing Zhu 0005, Shengyi Qian 0001, Zhongmou He, Tong Zhao 0003, Neil Shah, Danai Koutra
CVPR3
2025 3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination
abstract
The integration of language and 3D perception is crucial for embodied agents and robots that comprehend and interact with the physical world. While large language models (LLMs) have demonstrated impressive language understanding and generation capabilities, their adaptation to 3D environments (3D-LLMs) remains in its early stages. A primary challenge is a lack of large-scale datasets with dense grounding between language and 3D scenes. We introduce 3D-GRAND, a pioneering large-scale dataset comprising 40,087 household scenes paired with 6.2 million densely-grounded scene-language instructions. Our results show that instruction tuning with 3D-GRAND significantly enhances grounding capabilities and reduces hallucinations in 3D-LLMs. As part of our contributions, we propose a comprehensive benchmark 3D-POPE to systematically evaluate hallucination in 3D-LLMs, enabling fair comparisons of models. Our experiments highlight a scaling effect between dataset size and 3D-LLM performance, emphasizing the importance of large-scale 3D-text datasets for embodied AI research. Our results demonstrate early signals for effective sim-to-real transfer, indicating that models trained on large synthetic data can perform well on real-world 3D scans. Through 3D-GRAND and 3D-POPE, we aim to equip the embodied AI community with resources and insights to lead to more reliable and better-grounded 3D-LLMs.
Xuweiyi Chen, Nikhil Madaan, Madhavan Iyengar, Shengyi Qian 0001, David F. Fouhey, Joyce Y. Chai
CVPR5
2024 LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agent
abstract
3D visual grounding is a critical skill for household robots, enabling them to navigate, manipulate objects, and answer questions based on their environment. While existing approaches often rely on extensive labeled data or exhibit limitations in handling complex language queries, we propose LLM-Grounder, a novel zero-shot, open-vocabulary, Large Language Model (LLM)-based 3D visual grounding pipeline. LLM-Grounder utilizes an LLM to decompose complex natural language queries into semantic constituents and employs a visual grounding tool, such as OpenScene or LERF, to identify objects in a 3D scene. The LLM then evaluates the spatial and commonsense relations among the proposed objects to make a final grounding decision. Our method does not require any labeled training data and can generalize to novel 3D scenes and arbitrary text queries. We evaluate LLM-Grounder on the ScanRefer benchmark and demonstrate state-of-the-art zero-shot grounding accuracy. Our findings indicate that LLMs significantly improve the grounding capability, especially for complex language queries, making LLM-Grounder an effective approach for 3D vision-language tasks in robotics.
Xuweiyi Chen, Shengyi Qian 0001, Nikhil Madaan, Madhavan Iyengar, David F. Fouhey, Joyce Y. Chai
ICRA3
2024 Multi-Object Hallucination in Vision Language Models
abstract
Large vision language models (LVLMs) often suffer from object hallucination, producing objects not present in the given images. While current benchmarks for object hallucination primarily concentrate on the presence of a single object class rather than individual entities, this work systematically investigates multi-object hallucination, examining how models misperceive (e.g., invent nonexistent objects or become distracted) when tasked with focusing on multiple objects simultaneously. We introduce Recognition-based Object Probing Evaluation (ROPE), an automated evaluation protocol that considers the distribution of object classes within a single image during testing and uses visual referring prompts to eliminate ambiguity. With comprehensive empirical studies and analysis of potential factors leading to multi-object hallucination, we found that (1) LVLMs suffer more hallucinations when focusing on multiple objects compared to a single object. (2) The tested object class distribution affects hallucination behaviors, indicating that LVLMs may follow shortcuts and spurious correlations. (3) Hallucinatory behaviors are influenced by data-specific factors, salience and frequency, and model intrinsic behaviors. We hope to enable LVLMs to recognize and reason about multiple objects that often occur in realistic visual scenes, provide insights, and quantify our progress towards mitigating the issues.
Xuweiyi Chen, Ziqiao Ma 0001, Xuejun Zhang 0003, Sihan Xu, Shengyi Qian 0001, David F. Fouhey, Joyce Y. Chai
NeurIPS5
2024 Pitfalls in Link Prediction with Graph Neural Networks: Understanding the Impact of Target-link Inclusion & Better Practices
abstract
While Graph Neural Networks (GNNs) are remarkably successful in a variety of high-impact applications, we demonstrate that, in link prediction, the common practices of including the edges being predicted in the graph at training and/or test have outsized impact on the performance of low-degree nodes. We theoretically and empirically investigate how these practices impact node-level performance across different degrees. Specifically, we explore three issues that arise: (I1) overfitting; (I2) distribution shift; and (I3) implicit test leakage. The former two issues lead to poor generalizability to the test data, while the latter leads to overestimation of the model's performance and directly impacts the deployment of GNNs. To address these issues in a systematic way, we introduce an effective and efficient GNN training framework, SpotTarget, which leverages our insight on low-degree nodes: (1) at training time, it excludes a (training) edge to be predicted if it is incident to at least one low-degree node; and (2) at test time, it excludes all test edges to be predicted (thus, mimicking real scenarios of using GNNs, where the test data is not included in the graph). SpotTarget helps researchers and practitioners adhere to best practices for learning from graph data, which are frequently overlooked even by the most widely-used frameworks. Our experiments on various real-world datasets show that SpotTarget makes GNNs up to 15× more accurate in sparse graphs, and significantly improves their performance for low-degree nodes in dense graphs.
Jing Zhu 0005, Vassilis N. Ioannidis, Shengyi Qian 0001, Wei Ai 0002, Xiang Song 0003, Danai Koutra
WSDM4
2023 Understanding 3D Object Interaction from a Single Image
abstract
Humans can easily understand a single image as depicting multiple potential objects permitting interaction. We use this skill to plan our interactions with the world and accelerate understanding new objects without engaging in interaction. In this paper, we would like to endow machines with the similar ability, so that intelligent agents can better explore the 3D scene or manipulate objects. Our approach is a transformer-based model that predicts the 3D location, physical properties and affordance of objects. To power this model, we collect a dataset with Internet videos, egocentric videos and indoor images to train and validate our approach. Our model yields strong performance on our data, and generalizes well to robotics data.
Shengyi Qian 0001, David F. Fouhey
ICCV1
2023 Sound Localization from Motion: Jointly Learning Sound Direction and Camera Rotation
abstract
The images and sounds that we perceive undergo subtle but geometrically consistent changes as we rotate our heads. In this paper, we use these cues to solve a problem we call Sound Localization from Motion (SLfM): jointly estimating camera rotation and localizing sound sources. We learn to solve these tasks solely through self-supervision. A visual model predicts camera rotation from a pair of images, while an audio model predicts the direction of sound sources from binaural sounds. We train these models to generate predictions that agree with one another. At test time, the models can be deployed independently. To obtain a feature representation that is well-suited to solving this challenging problem, we also propose a method for learning an audio-visual representation through cross-view binauralization: estimating binaural sound from one view, given images and sound from another. Our model can successfully estimate accurate rotations on both real and synthetic scenes, and localize sound sources with accuracy competitive with state-of-the-art self-supervised approaches. Project site: https://ificl.github.io/SLfM.
Shengyi Qian 0001, Andrew Owens
ICCV2
2022 Understanding 3D Object Articulation in Internet Videos
abstract
We propose to investigate detecting and characterizing the 3D planar articulation of objects from ordinary RGB videos. While seemingly easy for humans, this problem poses many challenges for computers. Our approach is based on a top-down detection system that finds planes that can be articulated. This approach is followed by optimizing for a 3D plane that explains a sequence of detected articulations. We show that this system can be trained on a combination of videos and 3D scan datasets. When tested on a dataset of challenging Internet videos and the Charades dataset, our approach obtains strong performance.
Shengyi Qian 0001, Linyi Jin, Chris Rockwell 0001, Siyi Chen 0003, David F. Fouhey
CVPR1
2021 Planar Surface Reconstruction from Sparse Views
abstract
The paper studies planar surface reconstruction of indoor scenes from two views with unknown camera poses. While prior approaches have successfully created object-centric reconstructions of many scenes, they fail to exploit other structures, such as planes, which are typically the dominant components of indoor scenes. In this paper, we re-construct planar surfaces from multiple views, while jointly estimating camera pose. Our experiments demonstrate that our method is able to advance the state of the art of re-construction from sparse views, on challenging scenes from Matterport3D.
Linyi Jin, Shengyi Qian 0001, Andrew Owens, David F. Fouhey
ICCV2
2020 OASIS: A Large-Scale Dataset for Single Image 3D in the Wild
abstract
Single-view 3D is the task of recovering 3D properties such as depth and surface normals from a single image. We hypothesize that a major obstacle to single-image 3D is data. We address this issue by presenting Open Annotations of Single Image Surfaces (OASIS), a dataset for single-image 3D in the wild consisting of annotations of detailed 3D geometry for 140,000 images. We train and evaluate leading models on a variety of single-image 3D tasks. We expect OASIS to be a useful resource for 3D vision research. Project site: https://pvl.cs.princeton.edu/OASIS.
Shengyi Qian 0001, David Fan 0001, Noriyuki Kojima, Max Hamilton, Jia Deng 0001
CVPR2
2020 Associative3D: Volumetric Reconstruction from Sparse Views
Shengyi Qian 0001, Linyi Jin, David F. Fouhey
ECCV (15)1
2019 Learning Single-Image Depth From Videos Using Quality Assessment Networks
abstract
Depth estimation from a single image in the wild remains a challenging problem. One main obstacle is the lack of high-quality training data for images in the wild. In this paper we propose a method to automatically generate such data through Structure-from-Motion (SfM) on Internet videos. The core of this method is a Quality Assessment Network that identifies high-quality reconstructions obtained from SfM. Using this method, we collect single-view depth training data from a large number of YouTube videos and construct a new dataset called YouTube3D. Experiments show that YouTube3D is useful in training depth estimation networks and advances the state of the art of single-view depth estimation in the wild.
Shengyi Qian 0001, Jia Deng 0001
CVPR2