Yudong Luo

dblp:161/8157 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 12 since 2021Systems, architecture and hardware · 10 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 Efficient Cross-Boundary Grasping in Stacked Clutter with Single-Visual Mapping Multi-Step
abstract
In logistics applications, the vision-based technology for grasping target objects in the air is relatively mature. However, when operating across the air and water such as grasping marine products from the water, the visual information collected by the camera will be disturbed by ripples and bubbles on the water surface, resulting in low grasping efficiency. Therefore, we introduce a grasping strategy based on single-visual mapping for multi-step (SVMMS) strategy to achieve cross-medium operations involving stacked objects. Specifically, we design a multifunctional integrated Deep Q-learning-based network model to extract visual features from the scene to effectively detect stacked objects and outputs their hierarchical relationships. Moreover, we quantify the underlying relationship between motion logic during action execution and changes in RGB-D during action execution to help the robot achieve efficient and collision-free operations. Our approach also incorporates a time-series design with prioritized experience replay to globally optimize the action sequence. Additionally, we propose a novel sim2real method by combining domain randomization to address the difference in object sizes between the simulation and the real world. Extensive experiments in both simulation and physical environments show that SVMMS-Grasp significantly outperforms existing methods in terms of task success rate, stability, and operational efficiency.
Yudong Luo, Feiyu Xie, Na Zhao 0008, Xianping Fu, Yantao Shen 0001
ICRA1
2025 Data-Driven MPC for Attitude Control of Autonomous Underwater Robot
abstract
High maneuverability is essential to the autonomous operation of underwater robots. To achieve real-time maneuvering motion, the control strategy must take into account nonlinear hydrodynamic effects, which are extremely difficult to accurately capture during motion and therefore a balance must be struck between accuracy and real-time computational efficiency. Therefore, this paper proposes a data-driven approach to model the dynamics of the underwater robot using Sparse Identification of Nonlinear Dynamics (SINDy). Compared with existing works, our method does not require any physical prior knowledge and only uses a short period of onboard sensor data. Subsequently, the learned dynamic model is incorporated into a model predictive controller (MPC) to enable precise attitude control. Finally, the proposed method is implemented on our developed fully vectored propulsion underwater robot, and a series of attitude tracking experiments are conducted in an indoor water tank. Experimental results reveal that our approach significantly improves the model accuracy and reduces the attitude tracking errors by over 79% at a control frequency of 20 Hz, which proves the effectiveness and real-time performance of the method.
Tianzhu Gao, Yudong Luo, Na Zhao 0008, Yuanchu Yan, Xianping Fu, Yantao Shen 0001
IROS2
2025 Reevaluation of Large Neighborhood Search for MAPF: Findings and Opportunities
abstract
Multi-Agent Path Finding (MAPF) aims to arrange collision-free goal-reaching paths for a group of agents. Anytime MAPF solvers based on large neighborhood search (LNS) have gained prominence recently due to their flexibility and scalability, leading to a surge of methods, especially those leveraging machine learning, to enhance neighborhood selection. However, several pitfalls exist and hinder a comprehensive evaluation of these new methods, which mainly include: 1) Lower than actual or incorrect baseline performance; 2) Lack of a unified evaluation setting and criterion; 3) Lack of a codebase or executable model for supervised learning methods. To address these challenges, we introduce a unified evaluation framework, implement prior methods, and conduct an extensive comparison of prominent methods. Our evaluation reveals that rule-based heuristics serve as strong baselines, while current learning-based methods show no clear advantage on time efficiency or improvement capacity. Our extensive analysis also opens up new research opportunities for improving MAPF-LNS, such as targeting high-delayed agents, applying contextual algorithms, optimizing replan order and neighborhood size, where machine learning can potentially be integrated.
Jiaqi Tan 0005, Yudong Luo, Jiaoyang Li 0001, Hang Ma 0001
SOCS2
2024 Attitude Control for Morphing Quadrotor through Model Predictive Control with Constraints
abstract
Morphing quadrotors that can be potentially applied to confined spaces such as warehouses, tanks, and pipelines have flourished in recent years. Most work has focused on the mechanical feasibility of the morphing systems and high-level flight controller design, with limited discussions on low-level control. In this paper, a constrained model predictive control (MPC) is proposed and applied to solve the attitude control problem of a morphing quadrotor. Prior to controller design, a custom-built morphing quadrotor is introduced with the kinematic and dynamic models established and corresponding issues and challenges presented. In the controller, to eliminate the steady-state error, an embedded integrator is adopted by exploiting the differential variables; then, the constraints of the morphing quadrotor are incorporated into the MPC formulation to simulate real flight conditions, and an orthonormal function is employed to approximate the control input sequences in the controller to alleviate the computational burden. In the comparative studies, several scenarios are considered to demonstrate the effectiveness of the proposed control strategy in attitude control.
Na Zhao 0008, Yudong Luo, Chaojun Qin, Yantao Shen 0001
ICRA2
2024 Model Predictive Control for an Autonomous Underwater Robot with Fully Vectored Propulsion
abstract
Due to the low motion efficiency and maneuver-ability of underwater robots with six degrees of freedom, it is challenging for them to respond quickly to the attitude requirements during underwater autonomous manipulation. This paper presents a novel autonomous underwater robot with fully vectored propulsion and a model predictive control method to achieve more agile and efficient movements autonomously. In detail, we first design a robot with eight vector-distributed thruster layouts for fully vectored propulsion and construct the software architecture based on the robot operating system (ROS). Then, we establish the hydrodynamic model by adopting the Fossen approach and construct a 13-dimensional system state-space equation, which is discretized using the explicit fourth-order Runge-Kutta method. To achieve autonomous manipulation, model predictive control is employed along with physical constraints of the custom-built robot to enable real-time prediction and optimization of the robot’s states for control purposes. Finally, numerical simulations and experiments of the Point-to-Point Motion are conducted to test the robot’s performance. Experimental results reveal that the average error of each direction is 0.0027 m, 0.0031 m, and 0.0368 m in the x-axis, y-axis, and z-axis, respectively, and 0.8502°, 2.1941°, 0.2408° corresponding to three attitude angles, which verify the performance of employing MPC to control an autonomous underwater robot with fully vectored propulsion.
Tianzhu Gao, Yudong Luo, Weirong Luo, Xianping Fu, Na Zhao 0008, Yantao Shen 0001
ICRA2
2024 A Multi-modal Hybrid Robot with Enhanced Traversal Performance
abstract
Current multi-modal hybrid robots with flight and wheeled modes have fallen into the dilemma that they can only avoid obstacles by re-taking off when encountering obstacles due to the poor performance of wheeled obstacle-crossing. To tackle this problem, this paper presents a novel multi-modal hybrid robot with the ability to actively adjust the wheel’s size, which is inspired by the behavior of the turtle’s legs when it encounters obstacles, to enhance the traversal performance. In detail, we describe the hardware design that allows the robot to achieve a modal switch between flight and wheeled modes through foldable structures and variable wheel diameters; then, we present the architecture to control these two morphing mechanisms. After that, we establish the theoretical kinematic models for both the foldable arm and variable wheel and carry out extensive experiments to test the performance of the foldable arm, the variable-diameter wheel, as well as the traversal performance of the robot. Experimental results show that the proposed multimodal robot can realize the function of a quadrotor, respond quickly with full-scale folding within 0.9 s, climb a maximum slope of 36°, and traverse narrow passageways, which exhibit superior mobility and environmental adaptability.
Zhipeng He 0011, Na Zhao 0008, Yudong Luo, Sian Long, Hongbin Deng
ICRA3
2024 PESI: Personalized Explanation recommendation with Sentiment Inconsistency between ratings and reviews
Huiqiong Wu, Guibing Guo, Enneng Yang, Yudong Luo, Yabo Chu, Linying Jiang, Xingwei Wang 0001
Knowl. Based Syst.4
2023 Benchmarking Constraint Inference in Inverse Reinforcement Learning
Guiliang Liu, Yudong Luo, Ashish Gaurav, Kasra Rezaee, Pascal Poupart
ICLR2
2023 An Alternative to Variance: Gini Deviation for Risk-averse Policy Gradient
abstract
Restricting the variance of a policy’s return is a popular choice in risk-averse Reinforcement Learning (RL) due to its clear mathematical definition and easy interpretability. Traditional methods directly restrict the total return variance. Recent methods restrict the per-step reward variance as a proxy. We thoroughly examine the limitations of these variance-based methods, such as sensitivity to numerical scale and hindering of policy learning, and propose to use an alternative risk measure, Gini deviation, as a substitute. We study various properties of this new risk measure and derive a policy gradient algorithm to minimize it. Empirical evaluation in domains where risk-aversion can be clearly defined, shows that our algorithm can mitigate the limitations of variance-based risk measures and achieves high return with low risk in terms of variance and Gini deviation when others fail to learn a reasonable policy.
Yudong Luo, Guiliang Liu, Pascal Poupart, Yangchen Pan
NeurIPS1
2022 Distributional Reinforcement Learning with Monotonic Splines
Yudong Luo, Guiliang Liu, Haonan Duan 0002, Oliver Schulte, Pascal Poupart
ICLR1
2022 Uncertainty-Aware Reinforcement Learning for Risk-Sensitive Player Evaluation in Sports Game
abstract
A major task of sports analytics is player evaluation. Previous methods commonly measured the impact of players' actions on desirable outcomes (e.g., goals or winning) without considering the risk induced by stochastic game dynamics. In this paper, we design an uncertainty-aware Reinforcement Learning (RL) framework to learn a risk-sensitive player evaluation metric from stochastic game dynamics. To embed the risk of a player’s movements into the distribution of action-values, we model their 1) aleatoric uncertainty, which represents the intrinsic stochasticity in a sports game, and 2) epistemic uncertainty, which is due to a model's insufficient knowledge regarding Out-of-Distribution (OoD) samples. We demonstrate how a distributional Bellman operator and a feature-space density model can capture these uncertainties. Based on such uncertainty estimation, we propose a Risk-sensitive Game Impact Metric (RiGIM) that measures players' performance over a season by conditioning on a specific confidence level. Empirical evaluation, based on over 9M play-by-play ice hockey and soccer events, shows that RiGIM correlates highly with standard success measures and has a consistent risk sensitivity.
Guiliang Liu, Yudong Luo, Oliver Schulte, Pascal Poupart
NeurIPS2
2021 Distributed Heuristic Multi-Agent Path Finding with Communication
abstract
Multi-Agent Path Finding (MAPF) is essential to large-scale robotic systems. Recent methods have applied reinforcement learning (RL) to learn decentralized polices in partially observable environments. A fundamental challenge of obtaining collision-free policy is that agents need to learn co-operation to handle congested situations. This paper combines communication with deep Q-learning to provide a novel learning based method for MAPF, where agents achieve cooperation via graph convolution. To guide RL algorithm on long-horizon goal-oriented tasks, we embed the potential choices of shortest paths from single source as heuristic guidance instead of using a specific path as in most existing works. Our method treats each agent independently and trains the model from a single agent’s perspective. The final trained policy is applied to each agent for decentralized execution. The whole system is distributed during training and is trained under a curriculum learning strategy. Empirical evaluation in obstacle-rich environment indicates the high success rate with low average step of our method.
Ziyuan Ma, Yudong Luo, Hang Ma 0001
ICRA2
2020 Inverse Reinforcement Learning for Team Sports: Valuing Actions and Players
abstract
A major task of sports analytics is to rank players based on the impact of their actions. Recent methods have applied reinforcement learning (RL) to assess the value of actions from a learned action value or Q-function. A fundamental challenge for estimating action values is that explicit reward signals (goals) are very sparse in many team sports, such as ice hockey and soccer. This paper combines Q-function learning with inverse reinforcement learning (IRL) to provide a novel player ranking method. We treat professional play as expert demonstrations for learning an implicit reward function. Our method alternates single-agent IRL to learn a reward function for multiple agents; we provide a theoretical justification for this procedure. Knowledge transfer is used to combine learned rewards and observed rewards from goals. Empirical evaluation, based on 4.5M play-by-play events in the National Hockey League (NHL), indicates that player ranking using the learned rewards achieves high correlations with standard success measures and temporal consistency throughout a season.
Yudong Luo, Oliver Schulte, Pascal Poupart
IJCAI1
2020 Deep soccer analytics: learning an action-value function for evaluating soccer players
Guiliang Liu, Yudong Luo, Oliver Schulte, Tarak Kharrat
Data Min. Knowl. Discov.2
2018 Inchworm Locomotion Mechanism Inspired Self-Deformable Capsule-Like Robot: Design, Modeling, and Experimental Validation
abstract
Inspired by the inchworm locomotion mechanism, this paper presents our recently developed self-deformable capsule-like robot. The robot has the actuated deformation capability that relies on a novel rigid elements-based morphing structure (REMS) and its soft actuation mechanisms. When the robot deforms, it generates the crawling locomotion behavior and thus friction waves between the robot and contact surface to facilitate the inchworm-like crawling movement. The paper starts reviewing the deformable properties of natural biological entities like capsules, presents state of the art of the current capsule-like robots, and details the bio-inspired design of the self-deformable capsule-like robot by describing the model of robot kinematics and its locomotion mechanism. Both simulation and experimental results validate the excellent performance of this capsule-like robot. The developed self-deformable capsule-like robot has the advantage of crawling on varied surfaces and it also has the capabilities to crawl in a variety of narrow pipes based on the deformation elicited locomotion nature of the robot.
Yudong Luo, Na Zhao 0008, Kwang J. Kim, Jingang Yi, Yantao Shen 0001
ICRA1
2018 The Deformable Quad-Rotor Enabled and Wasp-Pedal-Carrying Inspired Aerial Gripper
abstract
The paper presents the development of a novel deformable quad-rotor enabled aerial gripper. The mechanism of our deformable quad-rotor is based on simultaneous expansion or contraction of the quad-rotor body, which is generated by controlling a rigid elements based morphing structure (REMS). Such deformation results in a highly deformable quad-rotor that can not only perform morphological adaptation in response to environmental changes and obstacles, but also improve the flight performance by contracting to facilitate the agility/maneuverability or by expanding to enhance the stability. Meanwhile, inspired by the wasp grasping behavior, such controllable expansion and contraction from the REMS ingeniously enable a new function of aerial gripper. In this paper, we start to detail the mechanism and design of the REMS based deformable quad-rotor, then present the quad-rotor deformation enabled aerial gripper design, its dynamics modeling, the grasping function and analysis. The simulation was conducted in order to graphically show the elicited aerodynamic flow situation during expansion or contraction of the quad-rotor with and without carrying payload. Experiments were further implemented to validate the grasping function of the gripper and the flight performance of the quad-rotor. Finally, two case studies on the new aerial gripper were performed. All results demonstrate the excellent performance of the deformable quad-rotor enabled aerial gripper, that is, it has the advantages of both flight maneuverability and grasping capability during performing tasks.
Na Zhao 0008, Yudong Luo, Hongbin Deng, Yantao Shen 0001, Hao Xu 0002
IROS2
2017 Design, modeling and experimental validation of a scissor mechanisms enabled compliant modular earthworm-like robot
abstract
Inspired by natural earthworm locomotion behavior and segmental muscle motion mechanism, this paper presents our recently developed compliant modular earthwormlike robot with the novel segmental muscle-mimetic design unit that is capable of efficiently mimicking earthworms' segmental muscle contraction and extension functions. The new class of segmental muscle-mimetic design unit relies on the curvature of scissor mechanisms that can be extended and contracted smoothly through controlled servo motors. The paper starts reviewing natural earthworm locomotion behavior, details the bio-inspired concept and design of both the segmental muscle-mimetic unit and the multi-segment earthworm-like robot prototype, and then presents the robot's locomotion models and analysis of the locomotion efficiency of the robot. Simulation and experimental results validate that both the design and the prototyped multi-segment earthworm-like robot have the excellent performance such as its muscle-like contractions behavior, peristaltic locomotion behavior, and highly competitive moving speed.
Yudong Luo, Na Zhao 0008, Hesheng Wang 0001, Kwang J. Kim, Yantao Shen 0001
IROS1
2017 The deformable quad-rotor: Design, kinematics and dynamics characterization, and flight performance validation
abstract
To improve the obstacle surmounting performance of the quad-rotor vehicle, this paper focuses on designing, kinematically and dynamically characterizing a novel deformable quad-rotor that is based on the scissor-like foldable structures. The foldable structure allows that the volume of the quad-rotor can be tuned to dynamically adapt variously sized obstacles and small spaces. To generate the controllable deformation, the actuated angulated elements that are the essential components of the scissor-like foldable structure play an important role. The element design, its actuation mechanism and the corresponding configuration patterns for the new quad-rotor are presented in the paper in detail. The simulations on deformation properties and obstacle surmounting ability are then performed to verify the deformation capability of the structure. In addition, experiments were extensively conducted to test the controlled deformation of the structure as well to investigate the deformation induced effects to the activated quad-rotor airframe and its aerodynamics. All implementation results validate the effectiveness of the proposed deformable quad-rotor design, that is, it enables the new quad-rotor having excellent obstacle surmounting performance, adaptability, flight maneuverability, as well as minimal aerodynamics influences during deforming.
Na Zhao 0008, Yudong Luo, Hongbin Deng, Yantao Shen 0001
IROS2