Shengjie Lin

dblp:348/0020 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 35% Efficient and distributed learning · 23% Segmentation and scene understanding · 12%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
0.912025
SPLART: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian Splatting · ICCV 2025
Computer vision › 3D vision
3d scene understanding
0.912025
SPLART: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian Splatting · ICCV 2025
Computer vision › 3D vision › 3d reconstruction › object reconstruction
articulated object reconstruction
0.912025
SPLART: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian Splatting · ICCV 2025
Machine learning › Efficient and distributed learning › distributed training › communication-efficient training
communication optimization
0.912025
Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep Learning · ASPLOS (1) 2025
Machine learning › Efficient and distributed learning
distributed training
0.912025
Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep Learning · ASPLOS (1) 2025
Computer vision › Segmentation and scene understanding
part segmentation
0.912025
SPLART: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian Splatting · ICCV 2025
Compilers and program optimization
deep learning compiler
0.912025
Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep Learning · ASPLOS (1) 2025
Computer vision › Vision and language › multimodal reasoning
embodied reasoning
0.812024
Statler: State-Maintaining Language Models for Embodied Reasoning · ICRA 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
long-horizon planning
0.812024
Statler: State-Maintaining Language Models for Embodied Reasoning · ICRA 2024
Robotics › Motion planning and robot control
robot planning
0.812024
Statler: State-Maintaining Language Models for Embodied Reasoning · ICRA 2024

Methods — techniques the papers use, named apart from their topics

resource-constrained project scheduling · 1.7auto-decomposition · 1.7self-supervised learning · 0.9differentiable rendering · 0.93d gaussian splatting · 0.9world state estimation · 0.8large language model prompting · 0.8
YearPublicationVenuePosition
2026 Fine-grained mutual geometric features enhanced human-object interaction detection
Ri Liu, Shengjie Lin
Pattern Recognit.3
2025 Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep Learning
abstract
With the exponential growth of deep learning (DL), there arises an escalating need for scalability. Despite significant advancements in communication hardware capabilities, the time consumed by communication remains a bottleneck during training. The existing various optimizations are coupled within parallel systems to implement specific computation-communication overlap. These approaches pose challenges in terms of performance, programmability, and generality. In this paper, we introduce Concerto, a compiler framework designed to address these challenges by automatically optimizing and scheduling communication. We formulate the scheduling problem as a resource-constrained project scheduling problem and use off-the-shelf solver to get the near-optimal scheduling. And use auto-decomposition to create overlap opportunity for critical (synchronous) communication. Our evaluation shows Concerto can match or outperform state-of-the-art parallel frameworks, including Megatron-LM, JAX/XLA, DeepSpeed, and Alpa, all of which include extensive hand-crafted optimization. Unlike previous works, Concerto decouples the parallel approach and communication optimization, then can generalize to a wide variety of parallelisms without manual optimization.
Shenggan Cheng, Shengjie Lin, Lansong Diao, Hao Wu 0077, Siyu Wang 0006, Chang Si, Xuanlei Zhao, Jiangsu Du, Wei Lin 0016, Yang You 0001
ASPLOS (1)2
2025 SPLART: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian Splatting
abstract
Reconstructing articulated objects prevalent in daily environments is crucial for applications in augmented/virtual reality and robotics. However, existing methods face scalability limitations (requiring 3D supervision or costly annotations), robustness issues (being susceptible to local optima), and rendering shortcomings (lacking speed or photorealism). We introduce SplArt, a self-supervised, category-agnostic framework that leverages 3D Gaussian Splatting (3DGS) to reconstruct articulated objects and infer kinematics from two sets of posed RGB images captured at different articulation states, enabling real-time photorealistic rendering for novel viewpoints and articulations. SplArt augments 3DGS with a differentiable mobility parameter per Gaussian, achieving refined part segmentation. A multi-stage optimization strategy is employed to progressively handle reconstruction, part segmentation, and articulation estimation, significantly enhancing robustness and accuracy. SplArt exploits geometric self-supervision, effectively addressing challenging scenarios without requiring 3D annotations or category-specific priors. Evaluations on established and newly proposed benchmarks, along with applications to real-world scenarios using a handheld RGB camera, demonstrate SplArt's state-of-the-art performance and real-world practicality. Code is publicly available at https://github.com/ripl/splart.
Shengjie Lin, Jiading Fang, Muhammad Zubair Irshad, Vitor Campagnolo Guizilini, Rares Ambrus, Gregory Shakhnarovich, Matthew R. Walter
ICCV1
2024 Statler: State-Maintaining Language Models for Embodied Reasoning
abstract
There has been a significant research interest in employing large language models to empower intelligent robots with complex reasoning. Existing work focuses on harnessing their abilities to reason about the histories of their actions and observations. In this paper, we explore a new dimension in which large language models may benefit robotics planning. In particular, we propose Statler, a framework in which large language models are prompted to maintain an estimate of the world state, which are often unobservable, and track its transition as new actions are taken. Our framework then conditions each action on the estimate of the current world state. Despite being conceptually simple, our Statler framework significantly outperforms strong competing methods (e.g., Code-as-Policies) on several robot planning tasks. Additionally, it has the potential advantage of scaling up to more challenging long-horizon planning tasks. We release our code here.
Takuma Yoneda, Jiading Fang, Tianchong Jiang, Shengjie Lin, Ben Picker, David Yunis, Hongyuan Mei, Matthew R. Walter
ICRA6
2024 Transcrib3D: 3D Referring Expression Resolution through Large Language Models
abstract
If robots are to work effectively alongside people, they must be able to interpret natural language references to objects in their 3D environment. Understanding 3D referring expressions is challenging—it requires the ability to both parse the 3D structure of the scene and correctly ground free-form language in the presence of distraction and clutter. We introduce Transcrib3D, an approach that brings together 3D detection methods and the emergent reasoning capabilities of large language models (LLMs). Transcrib3D uses text as the unifying medium, which allows us to sidestep the need to learn shared representations connecting multi-modal inputs, which would require massive amounts of annotated 3D data. As a demonstration of its effectiveness, Transcrib3D achieves state-of-the-art results on 3D reference resolution benchmarks, with a great leap in performance from previous multi-modality baselines. To improve upon zero-shot performance and facilitate local deployment on edge computers and robots, we propose self-correction for fine-tuning that trains smaller models, resulting in performance close to that of large models. We show that our method enables a real robot to perform pick-and-place tasks given queries that contain challenging referring expressions. Code will be available at https://ripl.github.io/Transcrib3D.
Jiading Fang, Xiangshan Tan, Shengjie Lin, Igor Vasiljevic, Vitor Campagnolo Guizilini, Hongyuan Mei, Rares Ambrus, Gregory Shakhnarovich, Matthew R. Walter
IROS3
2024 Joint Specifics and Dual-Semantic Hashing Learning for Cross-Modal Retrieval
Shaohua Teng, Shengjie Lin, Luyao Teng, Zefeng Zheng, Lunke Fei, Wei Zhang 0005
Neurocomputing2