Zhao Mandi

dblp:336/3180 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Motion planning and robot control · 42% 3D vision · 24% Multi-agent systems · 21%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d reconstruction › object reconstruction
articulated object reconstruction
0.912025
Real2Code: Reconstruct Articulated Objects via Code Generation · ICLR 2025
Program synthesis and code generation
code generation with language models
0.912025
Real2Code: Reconstruct Articulated Objects via Code Generation · ICLR 2025
Robotics › Motion planning and robot control › motion planning › multi-robot motion planning
multi-arm motion planning
0.812024
RoCo: Dialectic Multi-Robot Collaboration with Large Language Models · ICRA 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-robot coordination
0.812024
RoCo: Dialectic Multi-Robot Collaboration with Large Language Models · ICRA 2024
Robotics › Motion planning and robot control
trajectory optimization
0.812024
RoCo: Dialectic Multi-Robot Collaboration with Large Language Models · ICRA 2024
Robotics › Robot manipulation
grasping
0.312025
Real2Code: Reconstruct Articulated Objects via Code Generation · ICLR 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.212024
RoCo: Dialectic Multi-Robot Collaboration with Large Language Models · ICRA 2024

Methods — techniques the papers use, named apart from their topics

shape completion · 1.7large language model · 1.7image segmentation · 1.7large language model prompting · 0.8in-context learning · 0.8
YearPublicationVenuePosition
2025 Real2Code: Reconstruct Articulated Objects via Code Generation
abstract
We present Real2Code, a novel approach to reconstructing articulated objects via code generation. Given visual observations of an object, we first reconstruct its part geometry using image segmentation and shape completion. We represent these object parts with oriented bounding boxes, from which a fine-tuned large language model (LLM) predicts joint articulation as code. By leveraging pre-trained vision and language models, our approach scales elegantly with the number of articulated parts, and generalizes from synthetic training data to real world objects in unstructured environments. Experimental results demonstrate that Real2Code significantly outperforms the previous state-of-the-art in terms of reconstruction accuracy, and is the first approach to extrapolate beyond objects' structural complexity in the training set, as we show for objects with up to 10 articulated parts. When incorporated with a stereo reconstruction model, Real2Code moreover generalizes to real-world objects, given only a handful of multi-view RGB images and without the need for depth or camera information.
Zhao Mandi, Yijia Weng, Dominik Bauer, Shuran Song
ICLR1
2024 RoCo: Dialectic Multi-Robot Collaboration with Large Language Models
abstract
We propose a novel approach to multi-robot collaboration that harnesses the power of pre-trained large language models (LLMs) for both high-level communication and low-level path planning. Robots are equipped with LLMs to discuss and collectively reason task strategies. They generate sub-task plans and task space waypoint paths, which are used by a multi-arm motion planner to accelerate trajectory planning. We also provide feedback from the environment, such as collision checking, and prompt the LLM agents to improve their plan and waypoints in-context. For evaluation, we introduce RoCoBench, a 6-task benchmark covering a wide range of multi-robot collaboration scenarios, accompanied by a text-only dataset that evaluates LLMs’ agent representation and reasoning capability. We experimentally demonstrate the effectiveness of our approach — it achieves high success rates across all tasks in RoCoBench and adapts to variations in task semantics. Our dialog setup offers high interpretability and flexibility — in real world experiments, we show RoCo easily incorporates human-in-the-loop, where a user can communicate and collaborate with a robot agent to complete tasks together.
Zhao Mandi, Shreeya Jain, Shuran Song
ICRA1