EDBT 2026 Demo / reviewers in the wild / expert
Shiyu Zhao 0002
dblp:120/7474-2
· DBLP profile ↗
19ranked-venue papers
0as first author
19since 2021 · last 2026
0000-0003-3098-8059ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Systems, architecture and hardware · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Observability-Enhanced Target Motion Estimation via Bearing-Box: Theory and MAV Applications
Zian Ning, Shiyu Zhao 0002 |
IEEE Trans. Robotics | 3 |
| 2025 | Predictive Kinematic Coordinate Control for Aerial Manipulators Based on Modified Kinematics LearningabstractHigh-precision manipulation has always been a developmental goal for aerial manipulators. This paper investigates the kinematic coordinate control issue in aerial manipulators. We propose a predictive kinematic coordinate control method, which includes a learning-based modified kinematic model and a model predictive control (MPC) scheme based on weight allocation. Compared to existing methods, our proposed approach offers several attractive features. First, the kinematic model incorporates closed-loop dynamics characteristics and online residual learning. Compared to methods that do not consider closed-loop dynamics and residuals, our proposed method has improved accuracy by 59.6%. Second, a MPC scheme that considers weight allocation has been proposed, which can coordinate the motion strategies of quadcopters and manipulators. Compared to methods that do not consider weight allocation, the proposed method can meet the requirements of more tasks. The proposed approach is verified through complex trajectory tracking and moving target tracking experiments. The results validate the effectiveness of the proposed method. Zhengzhen Li, Mengyu Ji, Huazi Cao, Shiyu Zhao 0002 |
ICRA | 5 |
| 2025 | A Cooperative Bearing-Rate Approach for Observability-Enhanced Target Motion EstimationabstractVision-based target motion estimation is a fundamental problem in many robotic tasks. The existing methods have the limitation of low observability and, hence, face challenges in tracking highly maneuverable targets. Motivated by the aerial target pursuit task where a target may maneuver in 3D space, this paper studies how to further enhance observability by incorporating the bearing rate information that has not been well explored in the literature. The main contribution of this paper is to propose a new cooperative estimator called STT-R (Spatial-Temporal Triangulation with bearing Rate), which is designed under the framework of distributed recursive least squares. This theoretical result is further verified by numerical simulation and real-world experiments. It is shown that the proposed STT-R algorithm can effectively generate more accurate estimations and effectively reduce the lag in velocity estimation, enabling tracking of more maneuverable targets. Canlun Zheng, Hanqing Guo, Shiyu Zhao 0002 |
ICRA | 3 |
| 2025 | Collective Behavior Clone with Visual Attention via Neural Interaction Graph PredictionabstractIn this paper, we propose a framework, collective behavioral cloning (CBC), to learn the underlying interaction mechanism and control policy of a swarm system. Given the trajectory data of a swarm system, we propose a graph variational autoencoder (GVAE) to learn the local interaction graph. Based on the interaction graph and swarm trajectory, we use behavioral cloning to learn the control policy of the swarm system. To demonstrate the practicality of CBC, we deploy it on a real-world decentralized vision-based robot swarm system. A visual attention network is trained based on the learned interaction graph for online neighbor selection. Experimental results show that our method outperforms previous approaches in predicting both the interaction graph and swarm actions with higher accuracy. This work offers a promising approach for understanding interaction mechanisms and swarm dynamics in future swarm robotics research. Code and data are available6. Zhao Ma, Shiyu Zhao 0002 |
IROS | 4 |
| 2025 | Cooperative Bearing-Only Target Pursuit via Multiagent Reinforcement Learning: Design and ExperimentabstractThis paper addresses the multi-robot pursuit problem for an unknown target, encompassing both target state estimation and pursuit control. First, in state estimation, we focus on using only bearing information, as it is readily available from vision sensors and effective for small, distant targets. Challenges such as instability due to the nonlinearity of bearing measurements and singularities in the two-angle representation are addressed through a proposed uniform bearing-only information filter. This filter integrates multiple 3D bearing measurements, provides a concise formulation, and enhances stability and resilience to target loss caused by limited field of view (FoV). Second, in target pursuit control within complex environments, where challenges such as heterogeneity and limited FoV arise, conventional methods like differential games or Voronoi partitioning often prove inadequate. To address these limitations, we propose a novel multiagent reinforcement learning (MARL) framework, enabling multiple heterogeneous vehicles to search, localize, and follow a target while effectively handling those challenges. Third, to bridge the sim-to-real gap, we propose two key techniques: incorporating adjustable low-level control gains in training to replicate the dynamics of real-world autonomous ground vehicles (AGVs), and proposing spectral-normalized RL algorithms to enhance policy smoothness and robustness. Finally, we demonstrate the successful zero-shot transfer of the MARL controllers to AGVs, validating the effectiveness and practical feasibility of our approach. The accompanying video is available at https://youtu.be/HO7FJyZiJ3E. Susheng Ding, Shiliang Guo, Shiyu Zhao 0002 |
IROS | 5 |
| 2025 | TACO: General Acrobatic Flight Control via Target-and-Command-Oriented Reinforcement LearningabstractAlthough acrobatic flight control has been studied extensively, one key limitation of the existing methods is that they are usually restricted to specific maneuver tasks and cannot change flight pattern parameters online. In this work, we propose a target-and-command-oriented reinforcement learning (TACO) framework, which can handle different maneuver tasks in a unified way and allows online parameter changes. We also propose a spectral normalization method with input-output rescaling to enhance the policy’s temporal and spatial smoothness, independence, and symmetry, thereby overcoming the sim-to-real gap. We validate the TACO approach through extensive simulation and real-world experiments, demonstrating its ability to achieve high-speed, high-accuracy circular flights and continuous multi-flips. The code is available at https://github.com/yinzikang/TACO. Zikang Yin, Canlun Zheng, Shiliang Guo, Shiyu Zhao 0002 |
IROS | 5 |
| 2025 | Vision-Based Cooperative MAV-Capturing-MAVabstractMAV-capturing-MAV (MCM) is one of the few effective methods for physically countering misused or malicious MAVs. This paper presents a vision-based cooperative MCM system, where multiple pursuer MAVs equipped with onboard vision systems detect, localize, and pursue a target MAV. To enhance robustness, a distributed state estimation and control framework enables the pursuer MAVs to autonomously coordinate their actions. Pursuer trajectories are optimized using Model Predictive Control (MPC) and executed via a low-level SO(3) controller, ensuring smooth and stable pursuit. Once the capture conditions are satisfied, the pursuer MAVs automatically deploy a flying net to intercept the target. These capture conditions are determined based on the predicted motion of the net. To enable real-time decision-making, we propose a lightweight computational method to approximate the net’s motion, avoiding the prohibitive cost of solving the full net dynamics. The effectiveness of the proposed system is validated through simulations and real-world experiments. In real-world tests, our approach successfully captures a moving target traveling at 4 m/s with an acceleration of 1 m/s2, achieving a success rate of 64.7%. Canlun Zheng, Yize Mi, Hanqing Guo, Huaben Chen, Shiyu Zhao 0002 |
IROS | 5 |
| 2025 | Casual3DHDR: High Dynamic Range 3D Gaussian Splatting from Casually Captured VideosabstractPhoto-realistic novel view synthesis from multi-view images, such as neural radiance field (NeRF) and 3D Gaussian Splatting (3DGS), has gained significant attention for its superior performance. However, most existing methods rely on low dynamic range (LDR) images, limiting their ability to capture detailed scenes in high-contrast environments. While some prior works address high dynamic range (HDR) scene reconstruction, they typically require multi-view sharp images with varying exposure times captured at fixed camera positions-a process that is time-consuming and impractical. To make data acquisition more flexible, we propose Casual3DHDR, a robust one-stage method that reconstructs 3D HDR scenes from casually-captured auto-exposure (AE) videos, even under severe motion blur and unknown, varying exposure times. Our approach integrates a continuous camera trajectory into a unified physical imaging model, jointly optimizing exposure times, camera poses, and the camera response function (CRF). Extensive experiments on synthetic and real-world datasets demonstrate that Casual3DHDR outperforms existing methods in robustness and rendering quality. Shucheng Gong, Lingzhe Zhao, Wenpu Li, Hong Xie 0002, Shiyu Zhao 0002, Peidong Liu 0001 |
ACM Multimedia | 6 |
| 2025 | MBA-SLAM: Motion Blur Aware Dense Visual SLAM With Radiance Fields RepresentationabstractEmerging 3D scene representations, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), have demonstrated their effectiveness in Simultaneous Localization and Mapping (SLAM) for photo-realistic rendering, particularly when using high-quality video sequences as input. However, existing methods struggle with motion-blurred frames, which are common in real-world scenarios like low-light or long-exposure conditions. This often results in a significant reduction in both camera localization accuracy and map reconstruction quality. To address this challenge, we propose a dense visual SLAM pipeline (i.e., MBA-SLAM) to handle severe motion-blurred inputs. Our approach integrates an efficient motion blur-aware tracker with either neural radiance fields or Gaussian Splatting based mapper. By accurately modeling the physical image formation process of motion-blurred images, our method simultaneously learns 3D scene representation and estimates the cameras' local trajectory during exposure time, enabling proactive compensation for motion blur caused by camera movement. In our experiments, we demonstrate that MBA-SLAM surpasses previous state-of-the-art methods in both camera localization and map reconstruction, showcasing superior performance across a range of datasets, including synthetic and real datasets featuring sharp images as well as those affected by motion blur, highlighting the versatility and robustness of our approach. Peng Wang 0141, Lingzhe Zhao, Shiyu Zhao 0002, Peidong Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Motion Planning for Aerial Pick-and-Place With Geometric Feasibility ConstraintsabstractThis paper studies the motion planning problem of the pick-and-place of an aerial manipulator that consists of a quadcopter flying base and a Delta arm. We propose a novel partially decoupled motion planning framework to solve this problem. Compared to the state-of-the-art approaches, the proposed one has two novel features. First, it does not suffer from increased computation in high-dimensional configuration spaces. That is because it calculates the trajectories of the quadcopter base and the end-effector separately in Cartesian space based on proposed geometric feasibility constraints. The geometric feasibility constraints can ensure the resulting trajectories satisfy the aerial manipulator’s geometry. Second, collision avoidance for the Delta arm is achieved through an iterative approach based on a pinhole mapping method, so that the feasible trajectory can be found in an efficient manner. The proposed approach is verified by five experiments on a real aerial manipulation platform. The experimental results show the effectiveness of the proposed method for the aerial pick-and-place task.Note to Practitioners—Aerial manipulators have attracted increasing research interest in recent years due to their potential applications in various domains. In this paper, we particularly focus on the motion planning problem of the pick-and-place of aerial manipulators. We propose a novel partially decoupled motion planning framework, which calculates the trajectories of the quadcopter base and the end-effector in Cartesian space, respectively. Geometric feasibility constraints are proposed to coordinate the trajectories to ensure successful execution. Five experiments on a real aerial manipulator platform demonstrate the effectiveness of the approach. In future research, we will address the motion planning problem of aerial manipulators in complex environments. Huazi Cao, Cunjia Liu, Bo Zhu 0005, Shiyu Zhao 0002 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Domain Adaptive Detection of MAVs: A Benchmark and Noise Suppression NetworkabstractVisual detection of Micro Air Vehicles (MAVs) has attracted increasing attention in recent years due to its important application in various tasks. The existing methods for MAV detection assume that the training set and testing set have the same distribution. As a result, when deployed in new domains, the detectors would have a significant performance degradation due to domain discrepancy. In this paper, we study the problem of cross-domain MAV detection. The contributions of this paper are threefold. 1) We propose a Multi-MAV-Multi-Domain (M3D) dataset consisting of both simulation and realistic images. Compared to other existing datasets, the proposed one is more comprehensive in the sense that it covers rich scenes, diverse MAV types, and various viewing angles. A new benchmark for cross-domain MAV detection is proposed based on the proposed dataset. 2) We propose a Noise Suppression Network (NSN) based on the framework of pseudo-labeling and a large-to-small training procedure. To reduce the challenging pseudo-label noises, two novel modules are designed in this network. The first is a prior-based curriculum learning module for allocating adaptive thresholds for pseudo labels with different difficulties. The second is a masked copy-paste augmentation module for pasting truly-labeled MAVs on unlabeled target images and thus decreasing pseudo-label noises. 3) Extensive experimental results verify the superior performance of the proposed method compared to the state-of-the-art ones. In particular, it achieves mAP of 46.9%(+5.8%), 50.5%(+3.7%), and 61.5%(+11.3%) on the tasks of simulation-to-real adaptation, cross-scene adaptation, and cross-camera adaptation, respectively.Note to Practitioners— To study the cross-domain MAV detection problem, this paper establishes a novel benchmark that consists of three domain adaptation tasks: simulation-to-real adaptation, cross-scene adaptation, and cross-camera adaptation, respectively. The benchmark is based on a novel MAV dataset called Multi-MAV-Multi-Domain (M3D), which is available at: https://github.com/WestlakeAerialRobotics/M3D. To reduce the noises caused by pseudo labels, a noise suppression network is proposed to overcome the error accumulation. Extensive experiments are conducted to prove the effectiveness of the proposed approach. Jinhong Deng, Peidong Liu 0001, Wen Li 0001, Shiyu Zhao 0002 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Uncertainty-Aware Semi-Supervised Semantic Key Point Detection via Bundle AdjustmentabstractVisual relative localization is widely used in multi-robot systems. While semantic key points offer a promising solution for 6DoF pose estimation, manual data labeling for network training remains unavoidable. In this paper, we introduce a novel method that jointly estimates the semantic key point detection model and 6DoF camera pose. Our key idea is to leverage the 3D-2D projection to produce pseudo labels for detection model training while taking the key point predictions as landmarks for 6DoF camera pose estimation. Compared with state-of-the-art works, our method eliminates the need for calibration and time synchronization of multi-camera systems, requiring only a handful of manually labeled data, which significantly improves the training efficiency. The experiment validates the effectiveness and practicality of our method in public datasets and real-world robotic applications. Code and data are made available3. Shiyu Zhao 0002 |
IROS | 3 |
| 2024 | Motion-guided small MAV detection in complex and non-planar scenesabstractIn recent years, there has been a growing interest in the visual detection of micro aerial vehicles (MAVs) due to its importance in numerous applications. However, the existing methods based on either appearance or motion features encounter difficulties when the background is complex or the MAV is too small. In this paper, we propose a novel motion-guided MAV detector that can accurately identify small MAVs in complex and non-planar scenes. This detector first exploits a motion feature enhancement module to capture the motion features of small MAVs. Then it uses multi-object tracking and trajectory filtering to eliminate false positives caused by motion parallax . Finally, an appearance-based classifier and an appearance-based detector that operates on the cropped regions are used to achieve precise detection results. Our proposed method can effectively and efficiently detect extremely small MAVs from dynamic and complex backgrounds because it aggregates pixel-level motion features and eliminates false positives based on the motion and appearance features of MAVs. Experiments on the ARD-MAV dataset demonstrate that the proposed method could achieve high performance in small MAV detection under challenging conditions and outperform other state-of-the-art methods across various metrics. Hanqing Guo, Canlun Zheng, Shiyu Zhao 0002 |
Pattern Recognit. Lett. | 3 |
| 2024 | ESO-Based Robust and High-Precision Tracking Control for Aerial ManipulationabstractThis paper studies the tracking control problem of an aerial manipulator that consists of a quadcopter flying base and a Delta robotic arm. We propose a novel control approach that consists of extended state observers (ESOs) for dynamic coupling estimation, ESO-based flight controllers, and a cooperative trajectory planner. Compared to the state-of-the-art approaches, the proposed one has some attractive features. First, it requires much less measurement information as opposed to the full-body control approaches and hence can be implemented conveniently and efficiently in practice. Second, while the existing approaches estimate the coupling effect based on precise models, the proposed ESOs can do that based on much less information about the system model. The proposed approach is verified by four experiments on a real aerial manipulation platform. The experimental results show that the average tracking error can reach 1 cm by the proposed approach as opposed to 10 cm by the PX4 baseline controller. Although force control is not considered specifically in the approach, the system can complete aerial weaving tasks thanks to the ESOs in the presence of drag forces applied to the end-effector during manipulation.Note to Practitioners—Aerial manipulators have received increasing research attention in recent years due to their wide range of applications. In this paper, we particularly focus on the high-precision and robust control of aerial manipulators. We propose a novel control approach that consists of extended state observers (ESOs) for dynamic coupling estimation, ESO-based flight controllers, and a cooperative trajectory planner. Four experiments on a real aerial manipulation platform demonstrate the effectiveness of the approach. In future research, we will address the control problem when the aerial manipulator contacts the environment. Huazi Cao, Yongqi Li 0007, Cunjia Liu, Shiyu Zhao 0002 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Global-Local MAV Detection Under Challenging Conditions Based on Appearance and MotionabstractVisual detection of micro aerial vehicles (MAVs) has received increasing research attention in recent years due to its importance in many applications. However, the existing approaches based on either appearance or motion features of MAVs still face challenges when the background is complex, the MAV target is small, or the computation resource is limited. In this paper, we propose a global-local MAV detector that can fuse both motion and appearance features for MAV detection under challenging conditions. This detector first searches MAV targets using a global detector and then switches to a local detector which works in an adaptive search region to enhance accuracy and efficiency. Additionally, a detector switcher is applied to coordinate the global and local detectors. A new dataset is created to train and verify the effectiveness of the proposed detector. This dataset contains more challenging scenarios that can occur in practice. Extensive experiments on three challenging datasets show that the proposed detector outperforms the state-of-the-art ones in terms of detection accuracy and computational efficiency. In particular, this detector can run with near real-time frame rate on NVIDIA Jetson NX Xavier, which demonstrates the usefulness of our approach for real-world applications. The dataset is available at https://github.com/WestlakeIntelligentRobotics/GLAD. In addition, A video summarizing this work is available at https://youtu.be/Tv473mAzHbU. Hanqing Guo, Zhi Gao 0005, Shiyu Zhao 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Keypoint-Guided Efficient Pose Estimation and Domain Adaptation for Micro Aerial VehiclesabstractVisual detection of micro aerial vehicles (MAVs) is an important problem in many tasks such as vision-based swarming of MAVs. This paper studies vision-based 6D pose estimation to detect a 3D bounding box of a target MAV and then estimate its 3D position and 3D attitude. The 3D attitude information is critical to better estimate the target's velocity since the attitude and motion are dynamically coupled. In this paper, we propose a novel 6D pose estimation method, whose novelties are threefold. First, we propose a novel centroid point-guided keypoint localization network that outperforms the state-of-the-art methods in terms of both accuracy and efficiency. Second, while there are no publicly available real-world datasets for 6D pose estimation for MAVs up to now, we propose a high-quality dataset based on an automatic dataset collection method. Third, since the dataset is collected in an indoor environment but detection tasks are usually in outdoor environments, we propose a self-training-based unsupervised domain adaption method to transfer the method from indoor to outdoor. Finally, we show that the estimated 6D pose especially the 3D attitude can significantly help improve the target's velocity estimation. Canlun Zheng, Peidong Liu 0001, Shiyu Zhao 0002 |
IEEE Trans. Robotics | 5 |
| 2023 | Detection, Localization, and Tracking of Multiple MAVs With Panoramic Stereo Camera NetworksabstractMalicious use of micro aerial vehicles (MAVs) has become a serious threat to public safety and personal privacy in recent years. Motivated by this problem, we propose a systematic approach to monitor the intrusion of malicious MAVs based on a novel type of panoramic stereo camera networks. Each sensing node of such a network consists of 16 lenses that can form a 360-degree panoramic vision system. The 16 lenses further form 8 pairs of stereo cameras that can directly localize aerial targets. The effective range for a sensing node localizing a MAV like DJI M300 could reach 80 meters, which is much farther than existing commercial stereo cameras. In terms of algorithms, we propose i) a novel visual MAV detection algorithm based primarily on motion features of MAVs, ii) an efficient stereo localization algorithm based on sparse feature points, and iii) robust multi-target tracking and trajectory fusion algorithm to fuse the observations of different sensing nodes. The effectiveness, robustness, and accuracy of the proposed algorithms together with the overall system have been verified by extensive experimental tests. To the best of our knowledge, this is the first systematic approach to detect, localize, and track unknown MAVs in the literature. Our approach provides a scalable solution to securely cover large areas of interest against malicious MAV intrusion. Note to Practitioners—Micro aerial vehicles (MAVs) have been widely used in many domains nowadays. However, they have also brought many safety problems. To monitor the intrusion of malicious MAVs, this paper proposes a novel type of panoramic stereo camera networks that can detect, localize, and track multiple MAVs simultaneously. Such a network consists of a number of sensing nodes and a central node. Each sensing node is able to detect, localize, and track multiple MAV targets. The role of the central node is to fuse the observations from multiple sensing nodes to generate more accurate trajectories of the MAV targets and in the meantime secure a large area in a coordinated way. This paper presents the details of the prototype of the system and the key algorithms therein. Canlun Zheng, Xiaoyu Zhang 0017, Fei Chen 0007, Shiyu Zhao 0002 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2023 | Three-Dimensional Bearing-Only Target Following via Observability-Enhanced Helical GuidanceabstractThis paper studies the problem of air-to-air target following of micro aerial vehicles (MAVs) motivated by the application of defense against malicious MAVs. When the bearing of the target MAV has been measured by the onboard visual sensor of the pursuer MAV, the problem becomes three-dimensional (3-D) bearing-only target following, which has been rarely studied in the literature and faces some unique challenges. To solve this problem, we propose the following novel results. First, to estimate the motion of the target MAV from 3-D bearing measurements, we propose a new pseudo-linear Kalman filter, which has a concise expression and superior stability compared to the classic ones such as the extended Kalman filter and modified polar coordinate filter. Second, we propose a novel approach to analyze the observability of state estimation when only bearing information is available. While the existing approaches are applicable to 2-D and single-step time-horizon cases, ours can handle more general 3-D and multiple-step time-horizon cases. Third, based on the theoretical conclusion of our observability analysis, we design a new 3-D helical guidance law that can better exploit the additional degree of freedom in 3D. The guidance law is adapted to the quadcopter's dynamics and a low-level flight controller is designed based on geometric control. Numerical simulation results verify the superior performance of the proposed algorithms compared to the state-of-the-art ones. Flight experiments on real quadrotor platforms further show the effectiveness and robustness of the proposed algorithms in practice. Zian Ning, Shaoming He, Shiyu Zhao 0002 |
IEEE Trans. Robotics | 5 |
| 2022 | Bearing-Only Formation Control With Prespecified Convergence TimeabstractThis article considers the bearing-only formation control problem, where the control of each agent only relies on relative bearings of their neighbors. A new control law is proposed to achieve target formations in finite time. Different from the existing results, the control law is based on a time-varying scaling gain. Hence, the convergence time can be arbitrarily chosen by users, and the derivative of the control input is continuous. Furthermore, sufficient conditions are given to guarantee almost global convergence and interagent collision avoidance. Then, a leader-follower control structure is proposed to achieve global convergence. By exploring the properties of the bearing Laplacian matrix, the collision avoidance and smooth control input are preserved. A multirobot hardware platform is designed to validate the theoretical results. Both simulation and experimental results demonstrate the effectiveness of our design. Zhenhong Li 0002, Hilton Tnunay, Shiyu Zhao 0002, Wei Meng 0003, Shengquan Xie, Zhengtao Ding |
IEEE Trans. Cybern. | 3 |