EDBT 2026 Demo / reviewers in the wild / expert
Koushil Sreenath
dblp:20/3046
· DBLP profile ↗
46ranked-venue papers
4as first author
31since 2021 · last 2026
0000-0002-5346-3637ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 3 first-author · 28 since 2021Systems, architecture and hardware · 37 · 3 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Traversability-Aware Legged Navigation by Learning From Real-World Visual DataabstractThe enhanced mobility brought by legged locomotion empowers quadrupedal robots to navigate through complex and unstructured environments. However, optimizing agile locomotion while accounting for the varying energy costs of traversing different terrains remains an open challenge. Most previous work focuses on planning trajectories with traversability cost estimation based on human-labeled environmental features. This human-centric approach is insufficient because it does not account for the varying capabilities of the robot locomotion controllers over challenging terrains. To address this, we introduce a novel real-world learning pipeline that unifies offline demonstrations, online reinforcement learning, and multi-modal perception to achieve robust legged navigation. The framework employs multiple training stages to develop a planner that guides the robot in avoiding obstacles and hardto- traverse terrains while reaching its goals. We first develop a novel traversability estimator in a robot-centric manner. The training of the navigation planner is directly performed in the real world using a sample efficient reinforcement learning method. With the proposed method, a quadrupedal robot learns to perform traversability-aware navigation through realworld interactions in diverse offroad and unstructured environments. Moreover, the robot demonstrates the ability to generalize the learned navigation skills to unseen scenarios. Zhongyu Li 0003, Xuanqi Zeng, Laura Smith 0001, Kyle Stachowicz, Dhruv Shah, Linzhu Yue, Zhitao Song, Weipeng Xia, Sergey Levine, Koushil Sreenath, Yun-Hui Liu 0001 |
IEEE Trans. Robotics | 11 |
| 2025 | Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsabstractVisual representations play a crucial role in developing generalist robotic policies. Previous vision encoders, typically pre-trained with single-image reconstruction or two-image contrastive learning, tend to capture static information, often neglecting the dynamic aspects vital for embodied tasks. Recently, video diffusion models (VDMs) demonstrate the ability to predict future frames and showcase a strong understanding of physical world. We hypothesize that VDMs inherently produce visual representations that encompass both current static information and predicted future dynamics, thereby providing valuable guidance for robot action learning. Based on this hypothesis, we propose the Video Prediction Policy (VPP), which learns implicit inverse dynamics model conditioned on predicted future representations inside VDMs. To predict more precise future, we fine-tune pre-trained video foundation model on robot datasets along with internet human manipulation data. In experiments, VPP achieves a 18.6% relative improvement on the Calvin ABC-D generalization benchmark compared to the previous state-of-the-art, and demonstrates a 31.6% increase in success rates for complex real-world dexterous manipulation tasks. For your convenience, videos can be found at https://video-prediction-policy.github.io/ Yanjiang Guo, Pengchao Wang, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, Jianyu Chen 0002 |
ICML | 7 |
| 2025 | Adaptive Energy Regularization for Autonomous Gait Transition and Energy-Efficient Quadruped LocomotionabstractIn reinforcement learning for legged robot locomotion, crafting effective reward strategies is crucial. Predefined gait patterns and complex reward systems are widely used to stabilize policy training. Drawing from the natural locomotion behaviors of humans and animals, which adapt their gaits to minimize energy consumption, we investigate the impact of incorporating an energy-efficient reward term that prioritizes distance-averaged energy consumption into the reinforcement learning framework. Our findings demonstrate that this simple addition enables quadruped robots to autonomously select appropriate gaits-such as four-beat walking at lower speeds and trotting at higher speeds-without the need for explicit gait regularizations. Furthermore, we provide a guideline for tuning the weight of this energy-efficient reward, facilitating its application in real-world scenarios. The effectiveness of our approach is validated through simulations and on a real Unitree Gol robot. This research highlights the potential of energy-centric reward functions to simplify and enhance the learning of adaptive and efficient locomotion in quadruped robots. Videos and more details are at https://sites.google.com/berkeley.edu/efficient-locomotion Boyuan Liang, Lingfeng Sun, Xinghao Zhu, Bike Zhang, Ziyin Xiong, Chenran Li, Koushil Sreenath, Masayoshi Tomizuka |
ICRA | 8 |
| 2025 | Berkeley Humanoid: A Research Platform for Learning-Based ControlabstractWe introduce Berkeley Humanoid, a reliable and low-cost mid-scale humanoid research platform for learningbased control. Our lightweight, in-house-built robot is designed specifically for learning algorithms with accurate simulation, low simulation complexity, anthropomorphic motion, and high reliability against falls. The narrow sim-to-real gap enables agile and robust locomotion across various terrains in outdoor environments, achieved with a simple reinforcement learning controller using light domain randomization. Furthermore, we demonstrate the robot traversing for hundreds of meters, walking on a steep unpaved trail, and hopping with single and double legs as a testimony to its high performance in dynamic walking. Capable of omnidirectional locomotion and withstanding large perturbations with a compact setup, our system aims for rapid sim-to-real deployment of learningbased humanoid systems. Please check our website https:// berkeley-humanoid.com/ and code https://github. com/HybridRobotics/isaac_berkeley_humanoid/. Qiayuan Liao, Bike Zhang, Xuanyu Huang, Zhongyu Li 0003, Koushil Sreenath |
ICRA | 6 |
| 2025 | CurricuLLM: Automatic Task Curricula Design for Learning Complex Robot Skills Using Large Language ModelsabstractCurriculum learning is a training mechanism in reinforcement learning (RL) that facilitates the achievement of complex policies by progressively increasing the task difficulty during training. However, designing effective curricula for a specific task often requires extensive domain knowledge and human intervention, which limits its applicability across various domains. Our core idea is that large language models (LLMs), with their extensive training on diverse language data and ability to encapsulate world knowledge, present significant potential for efficiently breaking down tasks and decomposing skills across various robotics environments. Additionally, the demonstrated success of LLMs in translating natural language into executable code for RL agents strengthens their role in generating task curricula. In this work, we propose CurricuLLM, which leverages the high-level planning and programming capabilities of LLMs for curriculum design, thereby enhancing the efficient learning of complex target tasks. CurricuLLM consists of: (Step 1) Generating a sequence of subtasks that aid target task learning in natural language form, (Step 2) Translating natural language description of subtasks in executable task code, including the reward code and goal distribution code, and (Step 3) Evaluating trained policies based on trajectory rollout and subtask description. We evaluate Cur-ricuLLM in various robotics simulation environments, ranging from manipulation, navigation, and locomotion, to show that CurricuLLM can aid learning complex robot control tasks. In addition, we validate humanoid locomotion policy learned through CurricuLLM in the real-world. Project website is https://iconlab.negarmehr.com/CurricuLLM/ Kanghyun Ryu, Qiayuan Liao, Zhongyu Li 0003, Payam Delgosha, Koushil Sreenath, Negar Mehr |
ICRA | 5 |
| 2025 | Learning Smooth Humanoid Locomotion through Lipschitz-Constrained PoliciesabstractReinforcement learning combined with sim-to-real transfer offers a general framework for developing locomotion controllers for legged robots. To facilitate successful deployment in the real world, smoothing techniques, such as low-pass filters and smoothness rewards, are often employed to develop policies with smooth behaviors. However, because these techniques are non-differentiable and usually require tedious tuning of a large set of hyperparameters, they tend to require extensive manual tuning for each robotic platform. To address this challenge and establish a general technique for enforcing smooth behaviors, we propose a simple and effective method that imposes a Lipschitz constraint on a learned policy, which we refer to as Lipschitz-Constrained Policies (LCP). We show that the Lipschitz constraint can be implemented in the form of a gradient penalty, which provides a differentiable objective that can be easily incorporated with automatic differentiation frameworks. We demonstrate that LCP effectively replaces the need for smoothing rewards or low-pass filters and can be easily integrated into training frameworks for many distinct humanoid robots. We extensively evaluate LCP in both simulation and real-world humanoid robots, producing smooth and robust locomotion controllers. All simulation and deployment code, along with complete checkpoints, is available on our project page: https://lipschitz-constrained-policy.github.io. Xialin He, Yen-Jen Wang, Qiayuan Liao, Yanjie Ze, Zhongyu Li 0003, S. Shankar Sastry, Jiajun Wu 0001, Koushil Sreenath, Xue Bin Peng |
IROS | 9 |
| 2025 | Estimation of Aerodynamics Forces in Dynamic Morphing Wing FlightabstractAccurate estimation of aerodynamic forces is essential for advancing the control, modeling, and design of flapping-wing aerial robots with dynamic morphing capabilities. In this paper, we investigate two distinct methodologies for force estimation on Aerobat, a bio-inspired flapping-wing platform designed to emulate the inertial and aerodynamic behaviors observed in bat flight. Our goal is to quantify aerodynamic force contributions during tethered flight, a crucial step toward closed-loop flight control. The first method is a physics-based observer derived from Hamiltonian mechanics that leverages the concept of conjugate momentum to infer external aerodynamic forces acting on the robot. This observer builds on the system’s reduced-order dynamic model and utilizes real-time sensor data to estimate forces without requiring training data. The second method employs a neural network-based regression model, specifically a multi-layer perceptron (MLP), to learn a mapping from joint kinematics, flapping frequency, and environmental parameters to aerodynamic force outputs. We evaluate both estimators using a 6-axis load cell in a high-frequency data acquisition setup that enables fine-grained force measurements during periodic wingbeats. The conjugate momentum observer and the regression model demonstrate strong agreement across three force components (Fx, Fy, Fz). Bibek Gupta, Albert Park, Eric Sihite, Koushil Sreenath, Alireza Ramezani |
IROS | 5 |
| 2025 | Long-horizon Locomotion and Manipulation on a Quadrupedal Robot with Large Language ModelsabstractWe present a large language model (LLM) based system to empower quadrupedal robots with problem-solving abilities for long-horizon tasks beyond short-term motions. Long-horizon tasks for quadrupeds are challenging since they require both a high-level understanding of the semantics of the problem for task planning and a broad range of locomotion and manipulation skills to interact with the environment. Our system builds a high-level reasoning layer with large language models, which generates hybrid discrete-continuous plans as robot code from task descriptions. It comprises multiple LLM agents: a semantic planner that sketches a plan, a parameter calculator that predicts arguments in the plan, a code generator that converts the plan into executable robot code, and a replanner that handles execution failures or human interventions. At the low level, we adopt reinforcement learning to train a set of motion planning and control skills to unleash the flexibility of quadrupeds for rich environment interactions. Our system is tested on long-horizon tasks that are infeasible to complete with one single skill. Simulation and real-world experiments show that it successfully figures out multi-step strategies and demonstrates non-trivial behaviors, including building tools or notifying a human for help. Demos are available on our project page: https://sites.google.com/view/long-horizon-robot. Yutao Ouyang, Jinhan Li, Yunfei Li 0005, Zhongyu Li 0003, Chao Yu 0005, Koushil Sreenath, Yi Wu 0013 |
IROS | 6 |
| 2025 | Diffuse-CLoC: Guided Diffusion for Physics-based Character Look-ahead ControlabstractWe present Diffuse-CLoC, a guided diffusion framework for physics-based look-ahead control that enables intuitive, steerable, and physically realistic motion generation. While existing kinematics motion generation with diffusion models offer intuitive steering capabilities with inference-time conditioning, they often fail to produce physically viable motions. In contrast, recent diffusion-based control policies have shown promise in generating physically realizable motion sequences, but the lack of kinematics prediction limits their steerability. Diffuse-CLoC addresses these challenges through a key insight: modeling the joint distribution of states and actions within a single diffusion model makes action generation steerable by conditioning it on the predicted states. This approach allows us to leverage established conditioning techniques from kinematic motion generation while producing physically realistic motions. As a result, we achieve planning capabilities without the need for a high-level planner. Our method handles a diverse set of unseen long-horizon downstream tasks through a single pre-trained model, including static and dynamic obstacle avoidance, motion in-betweening, and task-space control. Experimental results show that our method significantly outperforms the traditional hierarchical framework of high-level motion diffusion and low-level tracking. Takara E. Truong, Fangzhou Yu, Jean-Pierre Sleiman, Jessica K. Hodgins, Koushil Sreenath, Farbod Farshidian |
ACM Trans. Graph. | 7 |
| 2025 | Constraint-Guided Online Data Selection for Scalable Data-Driven Safety Filters in Uncertain Robotic SystemsabstractAs the use of autonomous robots expands in tasks that are complex and challenging to model, the demand for robust data-driven control methods that can certify safety and stability in uncertain conditions is increasing. However, the practical implementation of these methods often faces scalability issues due to the growing amount of data points with system complexity, and a significant reliance on high-quality training data. In response to these challenges, this study presents a scalable data-driven controller that efficiently identifies and infers from the most informative data points for implementing data-driven safety filters. Our approach is grounded in the integration of a model-based certificate function-based method and Gaussian Process (GP) regression, reinforced by a novel online data selection algorithm that reduces time complexity from quadratic to linear relative to dataset size. Empirical evidence, gathered from successful real-world cart-pole swing-up experiments and simulated locomotion of a five-link bipedal robot, demonstrates the efficacy of our approach. Our findings reveal that our efficient online data selection algorithm, which strategically selects key data points, enhances the practicality and efficiency of data-driven certifying filters in complex robotic systems, significantly mitigating scalability concerns inherent in nonparametric learning-based control methods. Jason J. Choi, Fernando Castañeda, Wonsuhk Jung, Bike Zhang, Claire J. Tomlin, Koushil Sreenath |
IEEE Trans. Robotics | 6 |
| 2024 | Point Cloud-Based Control Barrier Function Regression for Safe and Efficient Vision-Based ControlabstractControl barrier functions have become an increasingly popular framework for safe real-time control. In this work, we present a computationally low-cost framework for synthesizing barrier functions over point cloud data for safe vision-based control. We take advantage of surface geometry to locally define and synthesize a quadratic CBF over a point cloud. This CBF is used in a CBF-QP for control and verified in simulation on quadrotors and in hardware on quadrotors and the TurtleBot3. This technique enables safe navigation through unstructured and dynamically changing environments and is shown to be significantly more efficient than current methods. Massimiliano de Sa, Prasanth Kotaru, Koushil Sreenath |
ICRA | 3 |
| 2024 | Learning Visual Quadrupedal Loco-Manipulation from DemonstrationsabstractQuadruped robots are progressively being integrated into human environments. Despite the growing locomotion capabilities of quadrupedal robots, their interaction with objects in realistic scenes is still limited. While additional robotic arms on quadrupedal robots enable manipulating objects, they are sometimes redundant given that a quadruped robot is essentially a mobile unit equipped with four limbs, each possessing 3 degrees of freedom (DoFs). Hence, we aim to empower a quadruped robot to execute real-world manipulation tasks using only its legs. We decompose the loco-manipulation process into a low-level reinforcement learning (RL)-based controller and a high-level Behavior Cloning (BC)-based planner. By parameterizing the manipulation trajectory, we synchronize the efforts of the upper and lower layers, thereby leveraging the advantages of both RL and BC. Our approach is validated through simulations and real-world experiments, demonstrating the robot’s ability to perform tasks that demand mobility and high precision, such as lifting a basket from the ground while moving, closing a dishwasher, pressing a button, and pushing a door. Zhengmao He, Kun Lei, Yanjie Ze, Koushil Sreenath, Zhongyu Li 0003, Huazhe Xu |
IROS | 4 |
| 2024 | HiLMa-Res: A General Hierarchical Framework via Residual RL for Combining Quadrupedal Locomotion and ManipulationabstractThis work presents HiLMa-Res, a hierarchical framework leveraging reinforcement learning to tackle manipulation tasks while performing continuous locomotion using quadrupedal robots. Unlike most previous efforts that focus on solving a specific task, HiLMa-Res is designed to be general for various loco-manipulation tasks that require quadrupedal robots to maintain sustained mobility. The novel design of this framework tackles the challenges of integrating continuous locomotion control and manipulation using legs. It develops an operational space locomotion controller that can track arbitrary robot end-effector (toe) trajectories while walking at different velocities. This controller is designed to be generic to different downstream tasks, and therefore, can be utilized in high-level manipulation planning policy to address specific tasks. To demonstrate the versatility of this framework, we utilize HiLMa-Res to tackle several challenging loco-manipulation tasks using a quadrupedal robot in the real world. These tasks span from leveraging state-based policy to vision-based policy, from training purely from the simulation data to learning from real-world data. In these tasks, HiLMa-Res shows better performance than other methods. Qiayuan Liao, Yiming Ni, Zhongyu Li 0003, Laura Smith 0001, Sergey Levine, Xue Bin Peng, Koushil Sreenath |
IROS | 8 |
| 2024 | Deep Geometric Potential Functions for Tracking on ManifoldsabstractIn this paper, we introduce a novel approach for designing invariant control laws through potential functions for fully actuated dynamical systems evolving on manifolds by leveraging the power of neural networks. The geometry and non-linearity inherent to manifold-based dynamical systems pose challenges for traditional control law design, necessitating techniques with the interplay of differential geometry and dynamical systems for ensuring stability. Apart from stability, performance and optimality are other challenging areas to address for dynamical systems evolving on manifolds. On top of these, the concept of invariance helps us improve learning transferability skills from one scene to another scene. We propose invariant potential functions on manifolds defined by neural networks that can be used to generate elastic forces for asymptotic tracking of trajectories. The weights of the potential function can be tuned to shape the potential functions according to the performance requirements through minimizing a loss function. Nikhil Potu Surya Prakash, Joo-Hwan Seo, Koushil Sreenath, Jongeun Choi, Roberto Horowitz |
IROS | 3 |
| 2024 | Leveraging Symmetry in RL-based Legged Locomotion ControlabstractModel-free reinforcement learning is a promising approach for autonomously solving challenging robotics control problems, but faces exploration difficulty without information about the robot’s morphology. The under-exploration of multiple modalities with symmetric states leads to behaviors that are often unnatural and sub-optimal. This issue becomes particularly pronounced in the context of robotic systems with morphological symmetries, such as legged robots for which the resulting asymmetric and aperiodic behaviors compromise performance, robustness, and transferability to real hardware. To mitigate this challenge, we can leverage symmetry to guide and improve the exploration in policy learning via equivariance / invariance constraints. We investigate the efficacy of two approaches to incorporate symmetry: modifying the network architectures to be strictly equivariant / invariant, and leveraging data augmentation to approximate equivariant / invariant actor-critics. We implement the methods on challenging loco-manipulation and bipedal locomotion tasks and compare with an unconstrained baseline. We find that the strictly equivariant policy consistently outperforms other methods in sample efficiency and task performance in simulation. Additionaly, symmetry-incorporated approaches exhibit better gait quality, higher robustness and can be deployed zero-shot to hardware. Zhi Su, Daniel Felipe Ordoñez Apraez, Yunfei Li 0005, Zhongyu Li 0003, Qiayuan Liao, Giulio Turrisi, Massimiliano Pontil, Claudio Semini, Yi Wu 0013, Koushil Sreenath |
IROS | 11 |
| 2024 | Humanoid Locomotion as Next Token PredictionabstractWe cast real-world humanoid control as a next token prediction problem, akin to predicting the next word in language. Our model is a causal transformer trained via autoregressive prediction of sensorimotor sequences. To account for the multi-modal nature of the data, we perform prediction in a modality-aligned way, and for each input token predict the next token from the same modality. This general formulation enables us to leverage data with missing modalities, such as videos without actions. We train our model on a dataset of sequences from a prior neural network policy, a model-based controller, motion capture, and YouTube videos of humans. We show that our model enables a real humanoid robot to walk in San Francisco zero-shot. Our model can transfer to the real world even when trained on only 27 hours of walking data, and can generalize to commands not seen during training. These findings suggest a promising path toward learning challenging real-world control tasks by generative modeling of sensorimotor sequences. Ilija Radosavovic, Bike Zhang, Baifeng Shi, Jathushan Rajasegaran, Sarthak Kamat, Trevor Darrell, Koushil Sreenath, Jitendra Malik |
NeurIPS | 7 |
| 2023 | Creating a Dynamic Quadrupedal Robotic Goalkeeper with Reinforcement LearningabstractWe present a reinforcement learning (RL) framework that enables quadrupedal robots to perform soccer goalkeeping tasks in the real world. Soccer goalkeeping with quadrupeds is a challenging problem, that combines highly dynamic locomotion with precise and fast non-prehensile object (ball) manipulation. The robot needs to react to and intercept a potentially flying ball using dynamic locomotion maneuvers in a very short amount of time, usually less than one second. In this paper, we propose to address this problem using a hierarchical model-free RL framework. The first component of the framework contains multiple control policies for distinct locomotion skills, which can be used to cover different regions of the goal. Each control policy enables the robot to track random parametric end-effector trajectories while performing one specific locomotion skill, such as jump, dive, and sidestep. These skills are then utilized by the second part of the framework which is a high-level planner to determine a desired skill and end-effector trajectory in order to intercept a ball flying to different regions of the goal. We deploy the proposed framework on a Mini Cheetah quadrupedal robot and demonstrate the effectiveness of our framework for various agile interceptions of a fast-moving ball in the real world. Zhongyu Li 0003, Yanzhen Xiang, Yiming Ni, Yufeng Chi, Lizhi Yang, Xue Bin Peng, Koushil Sreenath |
IROS | 9 |
| 2023 | Walking in Narrow Spaces: Safety-Critical Locomotion Control for Quadrupedal Robots with Duality-Based OptimizationabstractThis paper presents a safety-critical locomotion control framework for quadrupedal robots. Our goal is to enable quadrupedal robots to safely navigate in cluttered environments. To tackle this, we introduce exponential Discrete Control Barrier Functions (exponential DCBFs) with duality-based obstacle avoidance constraints into a Non-linear Model Predictive Control (NMPC) with Whole-Body Control (WBC) framework for quadrupedal locomotion control. This enables us to use polytopes to describe the shapes of the robot and obstacles for collision avoidance while doing locomotion control of quadrupedal robots. Compared to most prior work, especially using CBFs, that utilize spherical and conservative approximation for obstacle avoidance, this work demonstrates a quadrupedal robot autonomously and safely navigating through very tight spaces in the real world. (Our open-source code is available at https://github.com/HybridRobotics/quadruped_nmpc_dcbf_duality, and the video is available at https://youtu.be/plgSQjwXm1Q.) Qiayuan Liao, Zhongyu Li 0003, Akshay Thirugnanam, Jun Zeng 0002, Koushil Sreenath |
IROS | 5 |
| 2022 | Vision-Aided Dynamic Quadrupedal Locomotion on Discrete Terrain Using Motion LibrariesabstractIn this paper, we present a framework rooted in control and planning that enables quadrupedal robots to traverse challenging terrains with discrete footholds using visual feedback. Navigating discrete terrain is challenging for quadrupeds because the motion of the robot can be aperiodic, highly dynamic, and blind for the hind legs of the robot. Additionally, the robot needs to reason over both the feasible footholds as well as the base velocity in order to speed up or slow down at different parts of the discrete terrain. To address these challenges, we build an offline library of periodic gaits which span two trotting steps, and switch between different motion primitives to achieve aperiodic motions of different step lengths on a quadrupedal robot. The motion library is used to provide targets to a geometric model predictive controller which outputs the contact forces at the stance feet. To incorporate visual feedback, we use terrain mapping tools and a forward facing depth camera to build a local height map of the terrain around the robot, and extract feasible foothold locations around both the front and hind legs of the robot. Our experiments show a small scale quadruped robot navigating multiple unknown, challenging and discrete terrains in the real world. Shuxiao Chen, Akshara Rai, Koushil Sreenath |
ICRA | 4 |
| 2022 | Autonomous Racing with Multiple Vehicles using a Parallelized Optimization with Safety Guarantee using Control Barrier FunctionsabstractThis paper presents a novel planning and control strategy for competing with multiple vehicles in a car racing scenario. The proposed racing strategy switches between two modes. When there are no surrounding vehicles, a learning-based model predictive control (MPC) trajectory planner is used to guarantee that the ego vehicle achieves better lap timing performance. When the ego vehicle is competing with other surrounding vehicles to overtake, an optimization-based planner generates multiple dynamically-feasible trajectories through parallel computation. Each trajectory is optimized under a MPC formulation with different homotopic Bezier-curve reference paths lying laterally between surrounding vehicles. The time-optimal trajectory among these different homotopic trajectories is selected and a low-level MPC controller with control barrier function constraints for obstacle avoidance is used to guarantee the system's safety-critical performance. The proposed algorithm has the capability to generate collision-free trajectories and track them while enhancing the lap timing performance with steady low computational complexity, outper-forming existing approaches in both timing and performance for an autonomous racing environment. To demonstrate the performance of our racing strategy, we simulate with multiple randomly generated moving vehicles on the track and test the ego vehicle's overtaking maneuvers. Suiyi He, Jun Zeng 0002, Koushil Sreenath |
ICRA | 3 |
| 2022 | Safety-Critical Control and Planning for Obstacle Avoidance between Polytopes with Control Barrier FunctionsabstractObstacle avoidance between polytopes is a chal-lenging topic for optimal control and optimization-based tra-jectory planning problems. Existing work either solves this problem through mixed-integer optimization, relying on simpli-fication of system dynamics, or through model predictive control with dual variables using distance constraints, requiring long horizons for obstacle avoidance. In either case, the solution can only be applied as an offline planning algorithm. In this paper, we exploit the property that a smaller horizon is sufficient for obstacle avoidance by using discrete-time control barrier function (DCBF) constraints and we propose a novel optimization formulation with dual variables based on DCBFs to generate a collision-free dynamically-feasible trajectory. The proposed optimization formulation has lower computational complexity compared to existing work and can be used as a fast online algorithm for control and planning for general nonlinear dynamical systems. We validate our algorithm on different robot shapes using numerical simulations with a kinematic bicycle model, resulting in successful navigation through maze environments with polytopic obstacles. Akshay Thirugnanam, Jun Zeng 0002, Koushil Sreenath |
ICRA | 3 |
| 2022 | Bayesian Optimization Meets Hybrid Zero Dynamics: Safe Parameter Learning for Bipedal Locomotion ControlabstractIn this paper, we propose a multi-domain control parameter learning framework that combines Bayesian Optimization (BO) and Hybrid Zero Dynamics (HZD) for locomotion control of bipedal robots. We leverage BO to learn the control parameters used in the HZD-based controller. The learning process is firstly deployed in simulation to optimize different control parameters for a large repertoire of gaits. Next, to tackle the discrepancy between the simulation and the real world, the learning process is applied on the physical robot to learn for corrections to the control parameters learned in simulation while also respecting a safety constraint for gait stability. This method empowers an efficient sim-to-real transition with a small number of samples in the real world, and does not require a valid controller to initialize the training in simulation. Our proposed learning framework is experimentally deployed and validated on a bipedal robot Cassie to perform versatile locomotion skills with improved performance on smoothness of walking gaits and reduction of steady-state tracking errors. Lizhi Yang, Zhongyu Li 0003, Jun Zeng 0002, Koushil Sreenath |
ICRA | 4 |
| 2022 | Hierarchical Reinforcement Learning for Precise Soccer Shooting Skills using a Quadrupedal RobotabstractWe address the problem of enabling quadrupedal robots to perform precise shooting skills in the real world using reinforcement learning. Developing algorithms to enable a legged robot to shoot a soccer ball to a given target is a challenging problem that combines robot motion control and planning into one task. To solve this problem, we need to consider the dynamics limitation and motion stability during the control of a dynamic legged robot. Moreover, we need to consider motion planning to shoot the hard-to-model deformable ball rolling on the ground with uncertain friction to a desired location. In this paper, we propose a hierarchical framework that leverages deep reinforcement learning to train (a) a robust motion control policy that can track arbitrary motions and (b) a planning policy to decide the desired kicking motion to shoot a soccer ball to a target. We deploy the proposed framework on an A1 quadrupedal robot and enable it to accurately shoot the ball to random targets in the real world. Yandong Ji, Zhongyu Li 0003, Xue Bin Peng, Sergey Levine, Glen Berseth, Koushil Sreenath |
IROS | 7 |
| 2022 | Adapting Rapid Motor Adaptation for Bipedal RobotsabstractRecent advances in legged locomotion have en-abled quadrupeds to walk on challenging terrains. However, bipedal robots are inherently more unstable and hence it's harder to design walking controllers for them. In this work, we leverage recent advances in rapid adaptation for locomotion control, and extend them to work on bipedal robots. Similar to existing works, we start with a base policy which produces actions while taking as input an estimated extrinsics vector from an adaptation module. This extrinsics vector contains information about the environment and enables the walking controller to rapidly adapt online. However, the extrinsics estimator could be imperfect, which might lead to poor performance of the base policy which expects a perfect estimator. In this paper, we propose A-RMA (Adapting RMA), which additionally adapts the base policy for the imperfect extrinsics estimator by finetuning it using model-free RL. We demonstrate that A-RMA outperforms a number of RL-based baseline controllers and model-based controllers in simulation, and show zero-shot deployment of a single A-RMA policy to enable a bipedal robot, Cassie, to walk in a variety of different scenarios in the real world beyond what it has seen during training. Videos and results at https: //ashish-kmr.github.io/a-rma/ Ashish Kumar 0007, Zhongyu Li 0003, Jun Zeng 0002, Deepak Pathak, Koushil Sreenath, Jitendra Malik |
IROS | 5 |
| 2022 | Teaching Robots to Span the Space of Functional Expressive MotionabstractOur goal is to enable robots to perform functional tasks in emotive ways, be it in response to their users' emotional states, or expressive of their confidence levels. Prior work has proposed learning independent cost functions from user feedback for each target emotion, so that the robot may optimize it alongside task and environment specific objectives for any situation it encounters. However, this approach is inefficient when modeling multiple emotions and unable to generalize to new ones. In this work, we leverage the fact that emotions are not independent of each other: they are related through a latent space of Valence-Arousal-Dominance (VAD). Our key idea is to learn a model for how trajectories map onto VAD with user labels. Considering the distance between a trajectory's mapping and a target VAD allows this single model to represent cost functions for all emotions. As a result 1) all user feedback can contribute to learning about every emotion; 2) the robot can generate trajectories for any emotion in the space instead of only a few predefined ones; and 3) the robot can respond emotively to user-generated natural language by mapping it to a target VAD. We introduce a method that interactively learns to map trajectories to this latent space and test it in simulation and in a user study. In experiments, we use a simple vacuum robot as well as the Cassie biped. Arjun Sripathy, Andreea Bobu, Zhongyu Li 0003, Koushil Sreenath, Daniel S. Brown, Anca D. Dragan |
IROS | 4 |
| 2021 | Motion Planning and Feedback Control for Bipedal Robots Riding a Snakeboard
Jonathan Anglingdarma, Joshua Morey, Koushil Sreenath |
ICRA | 4 |
| 2021 | Scalable Learning of Safety Guarantees for Autonomous Systems using Hamilton-Jacobi ReachabilityabstractAutonomous systems like aircraft and assistive robots often operate in scenarios where guaranteeing safety is critical. Methods like Hamilton-Jacobi reachability can provide guaranteed safe sets and controllers for such systems. However, often these same scenarios have unknown or uncertain environments, system dynamics, or predictions of other agents. As the system is operating, it may learn new knowledge about these uncertainties and should therefore update its safety analysis accordingly. However, work to learn and update safety analysis is limited to small systems of about two dimensions due to the computational complexity of the analysis. In this paper we synthesize several techniques to speed up computation: decomposition, warm-starting, and adaptive grids. Using this new framework we can update safe sets by one or more orders of magnitude faster than prior work, making this technique practical for many realistic systems. We demonstrate our results on simulated 2D and 10D near-hover quadcopters operating in a windy environment. Sylvia L. Herbert, Jason J. Choi, Suvansh Sanjeev, Marsalis T. Gibson, Koushil Sreenath, Claire J. Tomlin |
ICRA | 5 |
| 2021 | Reinforcement Learning for Robust Parameterized Locomotion Control of Bipedal RobotsabstractDeveloping robust walking controllers for bipedal robots is a challenging endeavor. Traditional model-based locomotion controllers require simplifying assumptions and careful modelling; any small errors can result in unstable control. To address these challenges for bipedal locomotion, we present a model-free reinforcement learning framework for training robust locomotion policies in simulation, which can then be transferred to a real bipedal Cassie robot. To facilitate sim-to-real transfer, domain randomization is used to encourage the policies to learn behaviors that are robust across variations in system dynamics. The learned policies enable Cassie to perform a set of diverse and dynamic behaviors, while also being more robust than traditional controllers and prior learning-based methods that use residual control. We demonstrate this on versatile walking behaviors such as tracking a target walking velocity, walking height, and turning yaw. (Video1) Zhongyu Li 0003, Xuxin Cheng, Xue Bin Peng, Pieter Abbeel, Sergey Levine, Glen Berseth, Koushil Sreenath |
ICRA | 7 |
| 2021 | Legged Robot State Estimation in Slippery Environments Using Invariant Extended Kalman Filter with Velocity UpdateabstractThis paper proposes a state estimator for legged robots operating in slippery environments. An Invariant Extended Kalman Filter (InEKF) is implemented to fuse inertial and velocity measurements from a tracking camera and leg kinematic constraints. The misalignment between the camera and the robot-frame is also modeled thus enabling auto-calibration of camera pose. The leg kinematics based velocity measurement is formulated as a right-invariant observation. Nonlinear observability analysis shows that other than the rotation around the gravity vector and the absolute position, all states are observable except for some singular cases. Discrete observability analysis demonstrates that our filter is consistent with the underlying nonlinear system. An online noise parameter tuning method is developed to adapt to the highly time-varying camera measurement noise. The proposed method is experimentally validated on a Cassie bipedal robot walking over slippery terrain. A video for the experiment can be found at https://youtu.be/VIqJL0cUr7s. Sangli Teng, Mark W. Mueller, Koushil Sreenath |
ICRA | 3 |
| 2021 | Robotic Guide Dog: Leading a Human with Leash-Guided Hybrid Physical InteractionabstractAn autonomous robot that is able to physically guide humans through narrow and cluttered spaces could be a big boon to the visually-impaired. Most prior robotic guiding systems are based on wheeled platforms with large bases with actuated rigid guiding canes. The large bases and the actuated arms limit these prior approaches from operating in narrow and cluttered environments. We propose a method that introduces a quadrupedal robot with a leash to enable the robot-guidinghuman system to change its intrinsic dimension (by letting the leash go slack) in order to fit into narrow spaces. We propose a hybrid physical Human Robot Interaction model that involves leash tension to describe the dynamical relationship in the robot-guiding-human system. This hybrid model is utilized in a mixed-integer programming problem to develop a reactive planner that is able to utilize slack-taut switching to guide a blind-folded person to safely travel in a confined space. The proposed leash-guided robot framework is deployed on a Mini Cheetah quadrupedal robot and validated in experiments (Video1) Anxing Xiao, Wenzhe Tong, Lizhi Yang, Jun Zeng 0002, Zhongyu Li 0003, Koushil Sreenath |
ICRA | 6 |
| 2021 | Real-time Geo-localization Using Satellite Imagery and Topography for Unmanned Aerial VehiclesabstractThe capabilities of autonomous flight with unmanned aerial vehicles (UAVs) have significantly increased in recent times. However, basic problems such as fast and robust geo-localization in GPS-denied environments still remain unsolved. Existing research has primarily concentrated on improving the accuracy of localization at the cost of long and varying computation time in various situations, which often necessitates the use of powerful ground station machines. In order to make image-based geo-localization online and pragmatic for lightweight embedded systems on UAVs, we propose a framework that is reliable in changing scenes, flexible about computing resource allocation and adaptable to common camera placements. The framework is comprised of two stages: offline database preparation and online inference. At the first stage, color images and depth maps are rendered as seen from potential vehicle poses quantized over the satellite and topography maps of anticipated flying areas. A database is then populated with the global and local descriptors of the rendered images. At the second stage, for each captured real-world query image, top global matches are retrieved from the database and the vehicle pose is further refined via local descriptor matching. We present field experiments of image-based localization on two different UAV platforms to validate our results. Shuxiao Chen, Mark W. Mueller, Koushil Sreenath |
IROS | 4 |
| 2020 | Staging energy sources to extend flight time of a multirotor UAVabstractEnergy sources such as batteries do not decrease in mass after consumption, unlike combustion-based fuels. We present the concept of staging energy sources, i.e. consuming energy in stages and ejecting used stages, to progressively reduce the mass of aerial vehicles in-flight which reduces power consumption, and consequently increases flight time. A flight time vs. energy storage mass analysis is presented to show the endurance benefit of staging to multirotors. We consider two specific problems in discrete staging - optimal order of staging given a certain number of energy sources, and optimal partitioning of a given energy storage mass budget into a given number of stages. We then derive results for a continuously staged case of an internal combustion engine driving propellers. Notably, we show that a multirotor powered by internal combustion has an upper limit on achievable flight time independent of the available fuel mass. Lastly, we validate the analysis with flight experiments on a custom two-stage battery-powered quadcopter. This quadcopter can eject a battery stage after consumption in-flight using a custom-designed mechanism, and continue hovering using the next stage. The experimental flight times match well with those predicted from the analysis for our vehicle. We achieve a 19% increase in flight time using the batteries in two stages as compared to a single stage. Karan P. Jain, Jerry Tang, Koushil Sreenath, Mark W. Mueller |
IROS | 3 |
| 2020 | Animated Cassie: A Dynamic Relatable Robotic CharacterabstractCreating robots with emotional personalities will transform the usability of robots in the real-world. As previous emotive social robots are mostly based on statically stable robots whose mobility is limited, this paper develops an animation to real-world pipeline that enables dynamic bipedal robots that can twist, wiggle, and walk to behave with emotions. First, an animation method is introduced to design emotive motions for the virtual robot's character. Second, a dynamics optimizer is used to convert the animated motion to dynamically feasible motion. Third, real-time standing and walking controllers and an automaton are developed to bring the virtual character to life. This framework is deployed on a bipedal robot Cassie and validated in experiments. To the best of our knowledge, this paper is one of the first to present an animatronic dynamic legged robot that is able to perform motions with desired emotional attributes. We term robots that use dynamic motions to convey emotions as Dynamic Relatable Robotic Characters. Zhongyu Li 0003, Christine Cummings, Koushil Sreenath |
IROS | 3 |
| 2020 | Dynamic Legged Manipulation of a Ball Through Multi-Contact OptimizationabstractThe feet of robots are typically used to design locomotion strategies, such as balancing, walking, and running. However, they also have great potential to perform manipulation tasks. In this paper, we propose a model predictive control (MPC) framework for a quadrupedal robot to dynamically balance on a ball and simultaneously manipulate it to follow various trajectories such as straight lines, sinusoids, circles and in-place turning. We numerically validate our controller on the Mini Cheetah robot using different gaits including trotting, bounding, and pronking on the ball. Bike Zhang, Jun Zeng 0002, Koushil Sreenath |
IROS | 5 |
| 2018 | Safe Teleoperation of Dynamic UAVs Through Control Barrier FunctionsabstractThis paper presents a method for assisting human operators to teleoperate highly dynamic systems such as quadrotors inside a constrained environment with safety guarantees. Our method enables human operators to focus on manually operating and flying quadrotor systems without the need to focus on avoiding potential obstacles. This is achieved with the presented supervisory controller overriding human input to enforce safety constraints when necessary. This method can be used as an assistive training solution for novice pilots to begin flying quadrotors without crashing them. Our supervisory controller uses an Exponential control barrier function based quadratic program to achieve safe human teleoperated flight. We demonstrate and validate our control approach through several experiments with multiple users with varying skill levels for three different scenarios of a quadrotor flying in a motion capture environment with virtual and physical constraints. Koushil Sreenath |
ICRA | 2 |
| 2017 | A framework for efficient teleoperation via online adaptationabstractWe propose a task-independent adaptive teleoperation methodology that seeks to improve operator performance and efficiency by concurrently modeling user intent and adapting the set of available actions according to the predicted intent. User input selects a robot motion from a finite set of dynamically feasible and safe motions, represented as a motion primitive library. User intent is modeled as a probabilistic distribution with respect to future actions that represents the likelihood of action selection given recent user input, which can be formulated independent of task, environment, or user. As the intent model becomes increasingly confident, the action set is adapted in order to reduce the error between the intended and actual performance. Experimental evaluation of teleoperating a quadrotor for nonaggressive, single-intent maneuvers such as following a racetrack and conducting a free-hand helix motion shows improved performance, validating that the approach provides efficient adaptation towards achieving the user intent. Xuning Yang, Koushil Sreenath, Nathan Michael |
ICRA | 2 |
| 2017 | Multi-robot Trajectory Generation for an Aerial Payload Transport System
Sarah Y. Tang, Koushil Sreenath, Vijay Kumar 0001 |
ISRR | 2 |
| 2016 | Optimal control for geometric motion planning of a robot diverabstractInertial reorientation of airborne articulated bodies has been an active area of research in the robotics community, as this behavior can help guide dynamic robots to a safe landing with minimal damage. The main objective of this work is emulating the aggressive and large angle correction maneuvers, like somersaults, that are performed by human divers. To this end, a planar three link robot, called DiverBot, is proposed. By considering a gravity-free scenario, a local connection is obtained between joint angles and the body orientation, resulting in a reduction in the system dynamics. An optimal control policy applied on this reduced configuration space yielded diving maneuvers that are dynamically feasible. Numerical results show that the DiverBot can execute one somersault without drift and multiple somersaults with minimal drift. Roberto Shu, Avinash Siravuru, Akshara Rai, Tony Dear, Koushil Sreenath, Howie Choset |
IROS | 5 |
| 2016 | Dynamic Walking on Stepping Stones with Gait Library and Control Barrier Functions
Quan Nguyen 0004, Xingye Da, Jessy W. Grizzle, Koushil Sreenath |
WAFR | 4 |
| 2016 | Symbolic Computation of Dynamics on Smooth Manifolds
Brian Bittner, Koushil Sreenath |
WAFR | 2 |
| 2015 | The Reaction Mass Biped: Equations of motion, hybrid model for walking and trajectory tracking controlabstractPendulum models have been studied as benchmark problems for development of nonlinear control schemes, as well as reduced-order models for the dynamics analysis of locomotion of humanoid robots. This work provides a generalization of the previously introduced Reaction Mass Pendulum (RMP), which is a multibody inverted pendulum model, to a bipedal model that can better model bipedal locomotion. The RMP consists of an extensible “leg” and a “body” with moving proof masses that give rise to a variable rotational inertia. The Reaction Mass Biped (RMB) introduced here has two legs, one of which takes the role of a stance leg and the other performs as a swing leg during bipedal locomotion. The bipedal walking dynamics model of the RMB is therefore hybrid, with the roles of stance leg and swing leg interchanged after each cycle. The dynamics model is developed using a variational mechanics approach, without using generalized coordinates for the rotational degrees of freedom. This dynamics model has thirteen degrees of freedom, all of which are considered to be actuated in the control design. A set of desired state trajectories that can enable bipedal walking in straight and curved lines are generated. A control scheme is then designed for asymptotically stable tracking of this set of trajectories with an almost global domain of attraction. Numerical simulation results confirm the stability of this tracking control scheme for different walking trajectories of the RMB. Koushil Sreenath, Amit K. Sanyal |
ICRA | 1 |
| 2014 | Toward image based visual servoing for aerial grasping and perchingabstractThis paper addresses the dynamics, control, planning, and visual servoing for micro aerial vehicles to perform high-speed aerial grasping tasks. We draw inspiration from agile, fast-moving birds, such as raptors, that detect, locate, and execute high-speed swoop maneuvers to capture prey. Since these grasping maneuvers are predominantly in the sagittal plane, we consider the planar system and present mathematical models and algorithms for motion planning and control, required to incorporate similar capabilities in quadrotors equipped with a monocular camera. In particular, we develop a dynamical model directly in the image space, show that this is a differentially-flat system with the image features serving as flat outputs, outline a method for generating trajectories directly in the image feature space, develop a geometric visual controller that considers the second order dynamics (in contrast to most visual servoing controllers that assume first order dynamics), and present validation of our methods through both simulations and experiments. Justin Thomas, Giuseppe Loianno, Koushil Sreenath, Vijay Kumar 0001 |
ICRA | 3 |
| 2013 | A partially observable hybrid system model for bipedal locomotion for adapting to terrain variationsabstractWe propose a methodology of applying PoMDPs at a sufficiently high abstraction of a high-dimensional continuous-time partially observable hybrid system. In particular, we develop a two-layer hybrid controller, where the higher-level PoMDP-based hybrid controller learns the boundaries between various modes and appropriately switches between them. The modes partition the state-space and represent a closed-loop hybrid system with a lower-level hybrid controller. We apply this methodology onto the problem of bipedal walking on varying terrain, where the gradient change in the terrain is only partially observable (due to poor and noisy sensors.) We develop three lower-level hybrid controllers that result in robust walking on level ground, up and down ramps. The higher-level PoMDP-based hybrid controller then learns the boundary between these controllers and is used to perform appropriate controller switching. With only a coarse, discrete estimate of walking speed, the controller enables traversing terrain both with long sustained constant slopes, and with rapid changes in slope. Simulation results are presented on a 26-dimensional planar bipedal robot model that incorporates contact forces and friction. Koushil Sreenath, Connie R. Hill Jr., Vijay Kumar 0001 |
HSCC | 1 |
| 2013 | Trajectory generation and control of a quadrotor with a cable-suspended load - A differentially-flat hybrid systemabstractA quadrotor with a cable-suspended load with eight degrees of freedom and four degrees underactuation is considered and the system is established to be a differentially-flat hybrid system. Using the flatness property, a trajectory generation method is presented that enables finding nominal trajectories with various constraints that not only result in minimal load swing if required, but can also cause a large swing in the load for dynamically agile motions. A control design is presented for the system specialized to the planar case, that enables tracking of either the quadrotor attitude, the load attitude or the position of the load. Stability proofs for the controller design and experimental validation of the proposed controller are presented. Koushil Sreenath, Nathan Michael, Vijay Kumar 0001 |
ICRA | 1 |
| 2012 | Switching control design for accommodating large step-down disturbances in bipedal robot walkingabstractThis paper presents a feedback controller that allows MABEL, a kneed, planar bipedal robot, with 1 m-long legs, to accommodate an abrupt 20 cm decrease in ground height. The robot is provided information on neither where the step down occurs, nor by how much. After the robot has stepped off a raised platform, however, the height of the platform can be estimated from the lengths of the legs and the angles of the robot's joints. A real-time control strategy is implemented that uses this on-line estimate of step-down height to switch from a baseline controller, that is designed for flat-ground walking, to a second controller, that is designed to attenuate torso oscillation resulting from the step-down disturbance. After one step, the baseline controller is re-applied. The control strategy is developed on a simplified-design model of the robot and then verified on a more realistic model before being evaluated experimentally. The paper concludes with experimental results showing MABEL (blindly) stepping off a 20 cm high platform. Hae-Won Park 0002, Koushil Sreenath, Alireza Ramezani, Jessy W. Grizzle |
ICRA | 2 |
| 2012 | Design and experimental implementation of a compliant hybrid zero dynamics controller with active force control for running on MABELabstractThis paper presents a control design based on the method of virtual constraints and hybrid zero dynamics to achieve stable running on MABEL, a planar biped with compliance. In particular, a time-invariant feedback controller is designed such that the closed-loop system not only respects the natural compliance of the open-loop system, but also enables active force control within the compliant hybrid zero dynamics and results in exponentially stable running gaits. The compliant-hybrid-zero-dynamics-based controller with active force control is implemented experimentally and shown to realize stable running gaits on MABEL at an average speed of 1.95 m/s (4.4 mph) and a peak speed of 3.06 m/s (6.8 mph). The obtained gait has flight phases upto 39% of the gait, and an estimated ground clearance of 7.5 – 10 cm. Koushil Sreenath, Hae-Won Park 0002, Jessy W. Grizzle |
ICRA | 1 |