Shuo Yang 0007

dblp:78/1102-7 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0003-1847-6413ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Multi-agent systems · 38% Reinforcement learning · 38% Motion planning and robot control · 23%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems
multi-agent safety
0.812024
Learning Adaptive Safety for Multi-Agent Systems · ICRA 2024
Machine learning › Reinforcement learning
safe reinforcement learning
0.812024
Learning Adaptive Safety for Multi-Agent Systems · ICRA 2024
Robotics › Motion planning and robot control › robot control › safe control
control barrier functions
0.212024
Learning Adaptive Safety for Multi-Agent Systems · ICRA 2024
Robotics › Motion planning and robot control
robot control
0.212024
Learning Adaptive Safety for Multi-Agent Systems · ICRA 2024

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.8control barrier functions · 0.8adaptive safe RL · 0.8
YearPublicationVenuePosition
2025 Multi-Agent Reinforcement Learning Guided by Signal Temporal Logic Specifications
abstract
Reward design is a key component of deep reinforcement learning (DRL), yet some tasks and designer’s objectives may be unnatural to define as a scalar cost function. Among the various techniques, formal methods integrated with DRL have garnered considerable attention due to their expressiveness and flexibility in defining the reward and requirements for different states and actions of the agent. Nevertheless, the exploration of leveraging Signal Temporal Logic (STL) for guiding multi-agent reinforcement learning (MARL) reward design is still limited. The presence of complex interactions, heterogeneous goals, and critical safety requirements in multi-agent systems exacerbates this challenge. In this paper, we propose a novel STL-guided multi-agent reinforcement learning framework. The STL requirements are designed to include both task specifications according to the objective of each agent and safety specifications. The robustness values from checking the states against STL specifications are leveraged to generate rewards. We validate our approach by conducting experiments across various testbeds. The experimental results demonstrate significant performance improvements compared to MARL without STL guidance, along with a remarkable increase in the overall safety rate of the multi-agent systems.
Jiangwei Wang, Shuo Yang 0007, Ziyan An, Songyang Han, Rahul Mangharam, Meiyi Ma, Fei Miao
IROS2
2024 Learning Adaptive Safety for Multi-Agent Systems
abstract
Ensuring safety in dynamic multi-agent systems is challenging due to limited information about the other agents. Control Barrier Functions (CBFs) are showing promise for safety assurance but current methods make strong assumptions about other agents and often rely on manual tuning to balance safety, feasibility, and performance. In this work, we delve into the problem of adaptive safe learning for multi-agent systems with CBF. We show how emergent behaviour can be profoundly influenced by the CBF configuration, highlighting the necessity for a responsive and dynamic approach to CBF design. We present ASRL, a novel adaptive safe RL framework, to fully automate the optimization of policy and CBF coefficients, to enhance safety and long-term performance through reinforcement learning. By directly interacting with the other agents, ASRL learns to cope with diverse agent behaviours and maintains the cost violations below a desired limit. We evaluate ASRL in a multi-robot system and competitive multi-agent racing, against learning-based and control-theoretic approaches. We empirically demonstrate the efficacy of ASRL, and assess generalization and scalability to out-of-distribution scenarios.
Luigi Berducci, Shuo Yang 0007, Rahul Mangharam, Radu Grosu
ICRA2
2023 A Benchmark Comparison of Imitation Learning-based Control Policies for Autonomous Racing
abstract
Autonomous racing with scaled race cars has gained increasing attention as an effective approach for developing perception, planning and control algorithms for safe autonomous driving at the limits of the vehicle’s handling. To train agile control policies for autonomous racing, learning-based approaches largely utilize reinforcement learning, albeit with mixed results. In this study, we benchmark a variety of imitation learning policies for racing vehicles that are applied directly or for bootstrapping reinforcement learning both in simulation and on scaled real-world environments. We show that interactive imitation learning techniques outperform traditional imitation learning methods and can greatly improve the performance of reinforcement learning policies by bootstrapping thanks to its better sample efficiency. Our benchmarks provide a foundation for future research on autonomous racing using Imitation Learning and Reinforcement Learning.
Xiatao Sun, Mingyan Zhou, Zhijun Zhuang, Shuo Yang 0007, Johannes Betz, Rahul Mangharam
IV4
2023 Towards Explainability in Modular Autonomous System Software
abstract
Safety-critical Autonomous Systems require trustworthy and transparent decision-making process to be deployable in the real world. The advancement of Machine Learning introduces high performance but largely through black-box algorithms. In this position paper, we focus the discussion of explainability specifically with Autonomous Vehicles (AVs). As a safety-critical system, AVs provide the unique opportunity to utilize cutting-edge Machine Learning techniques while requiring transparency in decision making. Interpretability in every action the AV takes becomes crucial in post-hoc analysis where blame assignment might be necessary. In this paper, we provide positioning on how researchers could consider incorporating explainability and interpretability into design and optimization of separate Autonomous Vehicle modules including Perception, Planning, and Control.
Hongrui Zheng, Zirui Zang, Shuo Yang 0007, Rahul Mangharam
IV3