Yinghao Shuai

dblp:400/4139 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 61% Robot manipulation · 15% Segmentation and scene understanding · 15%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.912025
PUGS: Zero-Shot Physical Understanding with Gaussian Splatting · ICRA 2025
Computer vision › 3D vision › 3d scene understanding
3d instance segmentation
0.912025
RE0: Recognize Everything with 3D Zero-Shot Instance Segmentation · ICRA 2025
Computer vision › 3D vision
3d reconstruction
0.912025
PUGS: Zero-Shot Physical Understanding with Gaussian Splatting · ICRA 2025
Robotics › Robot manipulation
grasping
0.912025
PUGS: Zero-Shot Physical Understanding with Gaussian Splatting · ICRA 2025
Computer vision › 3D vision
physical property estimation
0.912025
PUGS: Zero-Shot Physical Understanding with Gaussian Splatting · ICRA 2025
Computer vision › Segmentation and scene understanding › instance segmentation
zero-shot instance segmentation
0.912025
RE0: Recognize Everything with 3D Zero-Shot Instance Segmentation · ICRA 2025
Computer vision › Vision and language › vision-language model
CLIP
0.312025
RE0: Recognize Everything with 3D Zero-Shot Instance Segmentation · ICRA 2025
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
vision foundation model
0.312025
RE0: Recognize Everything with 3D Zero-Shot Instance Segmentation · ICRA 2025

Methods — techniques the papers use, named apart from their topics

volume integration · 0.9projection relationship · 0.9cropformer · 0.9contrastive loss · 0.9CLIP · 0.9
YearPublicationVenuePosition
2025 PUGS: Zero-Shot Physical Understanding with Gaussian Splatting
abstract
Current robotic systems can understand the categories and poses of objects well. But understanding physical properties like mass, friction, and hardness, in the wild, remains challenging. We propose a new method that reconstructs 3D objects using the Gaussian splatting representation and predicts various physical properties in a zero-shot manner. We propose two techniques during the reconstruction phase: a geometryaware regularization loss function to improve the shape quality and a region-aware feature contrastive loss function to promote region affinity. Two other new techniques are designed during inference: a feature-based property propagation module and a volume integration module tailored for the Gaussian representation. Our framework is named as zero-shot physical understanding with Gaussian splatting, or PUGS. PUGS achieves new state-of-the-art results on the standard benchmark of ABO-500 mass prediction. We provide extensive quantitative ablations and qualitative visualization to demonstrate the mechanism of our designs. We show the proposed methodology can help address challenging real-world grasping tasks. Our codes, data, and models are available at https://github.com/EverNorif/PUGS
Yinghao Shuai, Yuantao Chen, Zijian Jiang, Nan Wang 0041, Jv Zheng, Jianzhu Ma, Meng Yang 0035, Zhicheng Wang 0022, Wenbo Ding 0001, Hao Zhao 0002
ICRA1
2025 RE0: Recognize Everything with 3D Zero-Shot Instance Segmentation
abstract
Recognizing objects in the 3D world is a significant challenge for robotics. Due to the lack of high-quality 3D data, directly training a general-purpose segmentation model in 3D is almost infeasible. Meanwhile, vision foundation models (VFM) have revolutionized the 2D computer vision field with outstanding performance, making the use of VFM to assist 3D perception a promising direction. However, most existing VFM-assisted methods do not effectively address the 2D-3D inconsistency problem or adequately provide corresponding semantic information for 3D instance objects. To address these two issues, this paper introduces a novel framework for 3D zero-shot instance segmentation called RE0. For the given 3D point clouds and multi-view RGB-D images with poses, we leverage the 3D geometric information, projection relationships, and CLIP semantic features. Specifically, we utilize CropFormer to extract mask information from multi-view posed images, combined with projection relationships to assign point-level labels to each point in the point cloud, and achieve instance-level consistency through inter-frame information interaction. Then, we employ projection relationships again to assign CLIP semantic features to the point cloud and achieve aggregation of small-scale point clouds. Notably, RE0 does not require any additional training and can be implemented by supporting only one inference of CropFormer and one inference of CLIP. Experiments on ScanNet200 and ScanNet++ show that our method achieves higher quality segmentation than the previous zero-shot methods. Our codes and demos are available at https://recognizeeverything.github.io/, with only one RTX 3090 GPU required.
Xiaohan Yan, Zijian Jiang, Yinghao Shuai, Nan Wang 0041, Wenbo Ji, Jinyu He, Zhicheng Wang 0022
ICRA3