VLDB 2026 Research / reviewers in the wild / expert
Yunlong Song
dblp:83/10696
· DBLP profile ↗
21ranked-venue papers
10as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Systems, architecture and hardware · 12 · 4 first-author · 11 since 2021Computer networks · 5 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Actor-Critic Model Predictive Control: Differentiable Optimization Meets Reinforcement Learning for Agile Flight
Angel Romero, Elie Aljalbout, Yunlong Song, Davide Scaramuzza 0001 |
IEEE Trans. Robotics | 3 |
| 2025 | Learning Quadrotor Control from Visual Features Using Differentiable SimulationabstractThe sample inefficiency of reinforcement learning (RL) remains a significant challenge in robotics. RL requires large-scale simulation and can still cause long training times, slowing research and innovation. This issue is particularly pronounced in vision-based control tasks where reliable state estimates are not accessible Differentiable simulation offers an alternative by enabling gradient back-propagation through the dynamics model, providing low-variance analytical policy gradients and, hence, higher sample efficiency. However, its usage for real-world robotic tasks has yet been limited. This work demonstrates the great potential of differentiable simulation for learning quadrotor control. We show that training in differentiable simulation significantly outperforms model-free RL in terms of both sample efficiency and training time, allowing a policy to learn to recover a quadrotor in seconds when providing vehicle states and in minutes when relying solely on visual features. The key to our success is two-fold. First, the use of a simple surrogate model for gradient computation greatly accelerates training without sacrificing control performance. Second, combining state representation learning with policy learning enhances convergence speed in tasks where only visual features are observable. These findings highlight the potential of differentiable simulation for real-world robotics and offer a compelling alternative to conventional RL approaches. Video: https://youtu.be/LdgvGCLB9do Code: https://github.com/uzh-rpg/rpgflightning Johannes Heeg, Yunlong Song, Davide Scaramuzza 0001 |
ICRA | 2 |
| 2025 | Residual Policy Learning for Perceptive Quadruped Control Using Differentiable SimulationabstractFirst-order Policy Gradient (FoPG) algorithms such as Backpropagation through Time and Analytical Policy Gradients leverage local simulation physics to accelerate policy search, significantly improving sample efficiency in robot control compared to standard model-free reinforcement learning. However, FoPG algorithms can exhibit poor learning dynamics in contact-rich tasks like locomotion. Previous approaches address this issue by alleviating contact dynamics via algorithmic or simulation innovations. In contrast, we propose guiding the policy search by learning a residual over a simple baseline policy. For quadruped locomotion, we find that the role of residual policy learning in FoPG-based training (FoPG RPL) is primarily to improve asymptotic rewards, compared to improving sample efficiency for model-free RL. Additionally, we provide insights on applying FoPG's to pixel-based local navigation, training a point-mass robot to convergence within seconds. Finally, we showcase the versatility of FoPG RPL by using it to train locomotion and perceptive navigation end-toend on a quadruped in minutes. Jing Yuan Luo, Yunlong Song, Victor Klemm, Fan Shi 0002, Davide Scaramuzza 0001, Marco Hutter 0001 |
ICRA | 2 |
| 2024 | Contrastive Initial State Buffer for Reinforcement LearningabstractIn Reinforcement Learning, the trade-off between exploration and exploitation poses a complex challenge for achieving efficient learning from limited samples. While recent works have been effective in leveraging past experiences for policy updates, they often overlook the potential of reusing past experiences for data collection. Independent of the underlying RL algorithm, we introduce the concept of a Contrastive Initial State Buffer, which strategically selects states from past experiences and uses them to initialize the agent in the environment in order to guide it toward more informative states. We validate our approach on two complex robotic tasks without relying on any prior information about the environment: (i) locomotion of a quadruped robot traversing challenging terrains and (ii) a quadcopter drone racing through a track. The experimental results show that our initial state buffer achieves higher task performance than the nominal baseline while also speeding up training convergence. Nico Messikommer, Yunlong Song, Davide Scaramuzza 0001 |
ICRA | 2 |
| 2024 | Actor-Critic Model Predictive ControlabstractAn open research question in robotics is how to combine the benefits of model-free reinforcement learning (RL)—known for its strong task performance and flexibility in optimizing general reward formulations—with the robustness and online replanning capabilities of model predictive control (MPC). This paper provides an answer by introducing a new framework called Actor-Critic Model Predictive Control. The key idea is to embed a differentiable MPC within an actor-critic RL framework. The proposed approach leverages the short-term predictive optimization capabilities of MPC with the exploratory and end-to-end training properties of RL. The resulting policy effectively manages both short-term decisions through the MPC-based actor and long-term prediction via the critic network, unifying the benefits of both model-based control and end-to-end learning. We validate our method in both simulation and the real world with a quadcopter platform across various high-level tasks. We show that the proposed architecture can achieve real-time control performance, learn complex behaviors via trial and error, and retain the predictive properties of the MPC to better handle out of distribution behaviour. Angel Romero, Yunlong Song, Davide Scaramuzza 0001 |
ICRA | 2 |
| 2024 | Contrastive Learning for Enhancing Robust Scene Transfer in Vision-based Agile FlightabstractScene transfer for vision-based mobile robotics applications is a highly relevant and challenging problem. The utility of a robot greatly depends on its ability to perform a task in the real world, outside of a well-controlled lab environment. Existing scene transfer end-to-end policy learning approaches often suffer from poor sample efficiency or limited generalization capabilities, making them unsuitable for mobile robotics applications. This work proposes an adaptive multi-pair contrastive learning strategy for visual representation learning that enables zero-shot scene transfer and real-world deployment. Control policies relying on the embedding are able to operate in unseen environments without the need for finetuning in the deployment environment. We demonstrate the performance of our approach on the task of agile, vision-based quadrotor flight. Extensive simulation and real-world experiments demonstrate that our approach successfully generalizes beyond the training domain and outperforms all baselines. Video: https://youtu.be/4A4YyPgEWD8 Jiaxu Xing, Leonard Bauersfeld, Yunlong Song, Chunwei Xing, Davide Scaramuzza 0001 |
ICRA | 3 |
| 2024 | Learning to Walk and Fly with Adversarial Motion PriorsabstractRobot multimodal locomotion encompasses the ability to transition between walking and flying, representing a significant challenge in robotics. This work presents an approach that enables automatic smooth transitions between legged and aerial locomotion. Leveraging the concept of Adversarial Motion Priors, our method allows the robot to imitate motion datasets and accomplish the desired task without the need for complex reward functions. The robot learns walking patterns from human-like gaits and aerial locomotion patterns from motions obtained using trajectory optimization. Through this process, the robot adapts the locomotion scheme based on environmental feedback using reinforcement learning, with the spontaneous emergence of mode-switching behavior. The results highlight the potential for achieving multimodal locomotion in aerial humanoid robotics through automatic control of walking and flying modes, paving the way for applications in diverse domains such as search and rescue, surveillance, and exploration missions. This research contributes to advancing the capabilities of aerial humanoid robots in terms of versatile locomotion in various environments. Video: https://youtu.be/mi6Do-x67CM Giuseppe L'Erario, Drew Hanover, Angel Romero, Yunlong Song, Gabriele Nava, Paolo Maria Viceconte, Daniele Pucci, Davide Scaramuzza 0001 |
IROS | 4 |
| 2024 | Autonomous Drone Racing: A SurveyabstractOver the last decade, the use of autonomous drone systems for surveying, search and rescue, or last-mile delivery has increased exponentially. With the rise of these applications comes the need for highly robust, safety-critical algorithms that can operate drones in complex and uncertain environments. Additionally, flying fast enables drones to cover more ground, increasing productivity and further strengthening their use case. One proxy for developing algorithms used in high-speed navigation is the task of autonomous drone racing, where researchers program drones to fly through a sequence of gates and avoid obstacles as quickly as possible using onboard sensors and limited computational power. Speeds and accelerations exceed over 80 kph and 4 g, respectively, raising significant challenges across perception, planning, control, and state estimation. To achieve maximum performance, systems require real-time algorithms that are robust to motion blur, high dynamic range, model uncertainties, aerodynamic disturbances, and often unpredictable opponents. This survey covers the progression of autonomous drone racing across model-based and learning-based approaches. We provide an overview of the field, its evolution over the years, and conclude with the biggest challenges and open questions to be faced in the future. Drew Hanover, Antonio Loquercio, Leonard Bauersfeld, Angel Romero, Robert Penicka, Yunlong Song, Giovanni Cioffi, Elia Kaufmann, Davide Scaramuzza 0001 |
IEEE Trans. Robotics | 6 |
| 2023 | Weighted Maximum Likelihood for Controller TuningabstractRecently, Model Predictive Contouring Control (MPCC) has arisen as the state-of-the-art approach for model-based agile flight. MPCC benefits from great flexibility in trading-off between progress maximization and path following at runtime without relying on globally optimized trajectories. However, finding the optimal set of tuning parameters for MPCC is challenging because (i) the full quadrotor dynamics are non-linear, (ii) the cost function is highly non-convex, and (iii) of the high dimensionality of the hyperparameter space. This paper leverages a probabilistic Policy Search method—Weighted Maximum Likelihood (WML)—to automatically learn the optimal objective for MPCC. WML is sample-efficient due to its closed-form solution for updating the learning parameters. Additionally, the data efficiency provided by the use of a model-based approach allows us to directly train in a high-fidelity simulator, which in turn makes our approach able to transfer zero-shot to the real world. We validate our approach in the real world, where we show that our method outperforms both the previous manually tuned controller and the state-of-the-art auto-tuning baseline reaching speeds of 75 km/h. Angel Romero, Shreedhar Govil, Gonca Yilmaz, Yunlong Song, Davide Scaramuzza 0001 |
ICRA | 4 |
| 2023 | Learning Perception-Aware Agile Flight in Cluttered EnvironmentsabstractRecently, neural control policies have outperformed existing model-based planning-and-control methods for autonomously navigating quadrotors through cluttered environments in minimum time. However, they are not perception aware, a crucial requirement in vision-based navigation due to the camera's limited field of view and the underactuated nature of a quadrotor. We propose a learning-based system that achieves perception-aware, agile flight in cluttered environments. Our method combines imitation learning with reinforcement learning (RL) by leveraging a privileged learning-by-cheating framework. Using RL, we first train a perception-aware teacher policy with full-state information to fly in minimum time through cluttered environments. Then, we use imitation learning to distill its knowledge into a vision-based student policy that only perceives the environment via a camera. Our approach tightly couples perception and control, showing a significant advantage in computation speed (10×faster) and success rate. We demonstrate the closed-loop control performance using hardware-in-the-loop simulation. Video: https://youtu.be/9q059CFGcVA Yunlong Song, Robert Penicka, Davide Scaramuzza 0001 |
ICRA | 1 |
| 2023 | Learning Deep Sensorimotor Policies for Vision-Based Autonomous Drone RacingabstractThe development of effective vision-based algorithms has been a significant challenge in achieving autonomous drones, which promise to offer immense potential for many real-world applications. This paper investigates learning deep sensorimotor policies for vision-based drone racing, which is a particularly demanding setting for testing the limits of an algorithm. Our method combines feature representation learning to extract task-relevant feature representations from high-dimensional image inputs with a learning-by-cheating framework to train a deep sensorimotor policy for vision-based drone racing. This approach eliminates the need for globally-consistent state estimation, trajectory planning, and handcrafted control design, allowing the policy to directly infer control commands from raw images, similar to human pilots. We conduct experiments using a realistic simulator and show that our vision-based policy can achieve state-of-the-art racing performance while being robust against unseen visual disturbances. Our study suggests that consistent feature embeddings are essential for achieving robust control performance in the presence of visual disturbances. The key to acquiring consistent feature embeddings is utilizing contrastive learning along with data augmentation. Video: https://youtu.be/AX_fcnW9yqE Jiawei Fu 0003, Yunlong Song, Fisher Yu 0001, Davide Scaramuzza 0001 |
IROS | 2 |
| 2022 | Light-GC: a lightweight and efficient garbage collection scheme for embedded file systemsabstractRaw flash file systems are essential in today's cost-sensitive embedded systems and legacy embedded devices. The performance of raw flash file systems is often limited by their inefficient garbage collection (GC) due to the tight capacity of the onboard flash memory and CPU computing power. However, in real-world scenarios, we observed that the GC overhead can account for more than 85% of the total I/O time, which not only degrades the performance of the file system but also shortens the endurance of the flash chip. However, existing GC optimization used in high-end SSD or server-side file systems usually leads to high CPU and memory usage, which is not suitable for the embedded environment. In this paper, we propose LightGC, a low-complexity and high-efficiency GC scheme. We design an optimized clustering algorithm for page hotness measurement and a hot-delay GC victim selection strategy. LightGC can also be adapted to locality changes in workloads to maximize the effect of cold and heat separation. Our experimental results show that LightGC offers a relative improvement of up to 19% in GC efficiency, a 5%-15% reduction in write amplification, and a 10%-17% increase in flash memory endurance compared to other existing GC methods. In addition, driven by LightGC, we implement the UBIFS2 file system, an upgraded version of the UBIFS flash file system. We have done careful engineering design so that UBIFS2 and the original UBIFS are compatible with each other and users can switch between them freely. Diansen Sun, Yunlong Song, Yunpeng Chai, Baoling Peng, Fangzhou Lu |
Middleware | 2 |
| 2022 | Policy Search for Model Predictive Control With Application to Agile Drone FlightabstractPolicy search and model predictive control (MPC) are two different paradigms for robot control: policy search has the strength of automatically learning complex policies using experienced data, and MPC can offer optimal control performance using models and trajectory optimization. An open research question is how to leverage and combine the advantages of both approaches. In this article, we provide an answer by using policy search for automatically choosing high-level decision variables for MPC, which leads to a novelpolicy-search-for-model-predictive-control framework. Specifically, we formulate the MPC as a parameterized controller, where the hard-to-optimize decision variables are represented as high-level policies. Such a formulation allows optimizing policies in a self-supervised fashion. We validate this framework by focusing on a challenging problem in agile drone flight: flying a quadrotor through fast-moving gates. Experiments show that our controller achieves robust and real-time control performance in both simulation and the real world. The proposed framework offers a new perspective for merging learning and control. Yunlong Song, Davide Scaramuzza 0001 |
IEEE Trans. Robotics | 1 |
| 2021 | Autonomous Overtaking in Gran Turismo Sport Using Curriculum Reinforcement LearningabstractProfessional race-car drivers can execute extreme overtaking maneuvers. However, existing algorithms for autonomous overtaking either rely on simplified assumptions about the vehicle dynamics or try to solve expensive trajectory-optimization problems online. When the vehicle approaches its physical limits, existing model-based controllers struggle to handle highly nonlinear dynamics, and cannot leverage the large volume of data generated by simulation or real-world driving. To circumvent these limitations, we propose a new learning-based method to tackle the autonomous overtaking problem. We evaluate our approach in the popular car racing game Gran Turismo Sport, which is known for its detailed modeling of various cars and tracks. By leveraging curriculum learning, our approach leads to faster convergence as well as increased performance compared to vanilla reinforcement learning. As a result, the trained controller outperforms the built-in model-based game AI and achieves comparable over-taking performance with an experienced human driver. Yunlong Song, HaoChih Lin, Elia Kaufmann, Peter Dürr, Davide Scaramuzza 0001 |
ICRA | 1 |
| 2021 | Autonomous Drone Racing with Deep Reinforcement LearningabstractIn many robotic tasks, such as autonomous drone racing, the goal is to travel through a set of waypoints as fast as possible. A key challenge for this task is planning the timeoptimal trajectory, which is typically solved by assuming perfect knowledge of the waypoints to pass in advance. The resulting solution is either highly specialized for a single-track layout, or suboptimal due to simplifying assumptions about the platform dynamics. In this work, a new approach to near-time-optimal trajectory generation for quadrotors is presented. Leveraging deep reinforcement learning and relative gate observations, our approach can compute near-time-optimal trajectories and adapt the trajectory to environment changes. Our method exhibits computational advantages over approaches based on trajectory optimization for non-trivial track configurations. The proposed approach is evaluated on a set of race tracks in simulation and the real world, achieving speeds of up to 60kmh−1with a physical quadrotor. Yunlong Song, Mats Steinweg, Elia Kaufmann, Davide Scaramuzza 0001 |
IROS | 1 |
| 2020 | Learning High-Level Policies for Model Predictive ControlabstractThe combination of policy search and deep neural networks holds the promise of automating a variety of decision-making tasks. Model Predictive Control (MPC) provides robust solutions to robot control tasks by making use of a dynamical model of the system and solving an optimization problem online over a short planning horizon. In this work, we leverage probabilistic decision-making approaches and the generalization capability of artificial neural networks to the powerful online optimization by learning a deep high-level policy for the MPC (High-MPC). Conditioning on robot's local observations, the trained neural network policy is capable of adaptively selecting high-level decision variables for the low-level MPC controller, which then generates optimal control commands for the robot. First, we formulate the search of high-level decision variables for MPC as a policy search problem, specifically, a probabilistic inference problem. The problem can be solved in a closed-form solution. Second, we propose a self-supervised learning algorithm for learning a neural network high-level policy, which is useful for online hyperparameter adaptations in highly dynamic environments. We demonstrate the importance of incorporating the online adaption into autonomous robots by using the proposed method to solve a challenging control problem, where the task is to control a simulated quadrotor to fly through a swinging gate. We show that our approach can handle situations that are difficult for standard MPC. Yunlong Song, Davide Scaramuzza 0001 |
IROS | 1 |
| 2014 | Greedy-based distributed algorithms for green traffic routingabstractGreen networking has become a hot research area in recent years. Most works just focus on centralized algorithms, where a central controller is used to collect all the information of the network and compute the energy aware traffic routing. However, these centralized algorithms only fit to the scenario when the network is not large and assume that all the network's information can be timely collected. When the above condition is not satisfied, the centralized algorithms will not work any more. In this paper, we study the basic principles of energy aware traffic routing and use them to explore a greedy-based distributed algorithm. The evaluation shows that our algorithm can reduce more energy consumption (almost 15%) than the existing distributed algorithms. Yunlong Song |
ICCCN | 1 |
| 2014 | Towards the emergency for Green Networking with robust algorithmabstractGreen Networking improves the energy efficiency by reducing the redundant resources of the network. However, these redundant resources are originally prepared in the network to make the network more robust. In this case, robust issue appears as a problem in green networking, especially when the emergency happens. This problem is usually less considered or ignored in existing research of green networking. In this paper, we explore and provide robust algorithms towards network emergency including traffic burst and link failure. The evaluation shows that our algorithm can cost less extra energy to handle the network emergency. Yunlong Song |
ISCC | 1 |
| 2012 | Understanding the basic principles implied in green traffic routingabstractThe aim of green traffic routing is to reduce the extra energy cost in a network while still satisfying the traffic demand. The extra energy cost can be found from the line card, a device used to switch the traffic of the links. We study the mechanism of green traffic routing and find out basic principles behind it: first, traffic should be routed along paths with small hops; second, the ratio of the power consumption of line cards to the traffic amount though the link should be as small as possible. Based on these principles, we suggest that the line card manufacturing should be more flexible and various, which can also produce some line cards with small capacity and small power to adapt for the goal of green networking. Besides, we propose that the network can dynamically configure all the different line cards in need rather than only make a sleeping of pre-positioned line cards without changing their installation any more. Yunlong Song |
ISCC | 1 |
| 2012 | Time series matrix factorization prediction of internet traffic matricesabstractTraffic matrices (TMs) are very important for traffic engineering and if they can be predicted, the network operations can be made beforehand. However, existing prediction methods are neither accurate nor efficient in practice. In this paper, we utilize the spatio-temporal property and low rank nature to directly predict the total TMs. The problem is that conventional matrix interpolation only works well when elements are missing uniformly and randomly. But in the case of TMs prediction, an entire part of the matrix is unknown. To solve this problem, we utilize some essential properties of TMs and add the time series forecasting into the matrix interpolation. We analyze our algorithm and evaluate its performance. The experiment result shows that our method can predict TMs under an NMAE of 30% in most cases, even predicting all the elements of next 3 weeks. Yunlong Song, Min Liu 0001, Shaojie Tang 0001, Xufei Mao |
LCN | 1 |
| 2011 | Power-Aware Traffic Engineering with Named Data NetworkingabstractPower-aware traffic engineering puts some links to sleep by moving their traffic to other links. However, this can make the utilization of remaining links higher, especially when the traffic amount is large. There is a tradeoff between the number of sleeping links and the utilization of links. To solve this problem, we propose to use a state-of-the-art networking called Named Data Networking (NDN), which can cache and retrieve the content in the storable routers. This can facilitate power-aware traffic engineering, because some traffic does not need to travel through the core network any more, and it only ends up at the edge routers which have already cached the required content. We use NDN to balance the traffic demand between origin-destination core routers so that traffic demand through the network can be adjusted to satisfy the requirement of power-aware traffic engineering. We evaluate power-aware traffic engineering with NDN and show its advantage compared to the one with conventional networking. Yunlong Song, Min Liu 0001, Yuwei Wang 0003 |
MSN | 1 |