Christian Pek

dblp:144/3352 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
13since 2021 · last 2024
0000-0001-7461-920XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 11 since 2021Systems, architecture and hardware · 8 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2024 SEQUEL: Semi-Supervised Preference-based RL with Query Synthesis via Latent Interpolation
abstract
Preference-based reinforcement learning (RL) poses as a recent research direction in robot learning, by allowing humans to teach robots through preferences on pairs of desired behaviours. Nonetheless, to obtain realistic robot policies, an arbitrarily large number of queries is required to be answered by humans. In this work, we approach the sample-efficiency challenge by presenting a technique which synthesizes queries, in a semi-supervised learning perspective. To achieve this, we leverage latent variational autoencoder (VAE) representations of trajectory segments (sequences of state-action pairs). Our approach manages to produce queries which are closely aligned with those labeled by humans, while avoiding excessive uncertainty according to the human preference predictions as determined by reward estimations. Additionally, by introducing variation without deviating from the original human’s intents, more robust reward function representations are achieved. We compare our approach to recent state-of-the-art preference-based RL semi-supervised learning techniques. Our experimental findings reveal that we can enhance the generalization of the estimated reward function without requiring additional human intervention. Lastly, to confirm the practical applicability of our approach, we conduct experiments involving actual human users in a simulated social navigation setting. Videos of the experiments can be found at https://sites.google.com/view/rl-sequel
Daniel Marta, Simon Holk, Christian Pek, Iolanda Leite
ICRA3
2023 Increasing Perceived Safety in Motion Planning for Human-Drone Interaction
abstract
Safety is crucial for autonomous drones to operate close to humans. Besides avoiding unwanted or harmful contact, people should also perceive the drone as safe. Existing safe motion planning approaches for autonomous robots, such as drones, have primarily focused on ensuring physical safety, e.g., by imposing constraints on motion planners. However, studies indicate that ensuring physical safety does not necessarily lead to perceived safety. Prior work in Human-Drone Interaction (HDI) shows that factors such as the drone's speed and distance to the human are important for perceived safety. Building on these works, we propose a parameterized control barrier function (CBF) that constrains the drone's maximum deceleration and minimum distance to the human and update its parameters on people's ratings of perceived safety. We describe an implementation and evaluation of our approach. Results of a within-subject user study (N=15) show that we can improve perceived safety of a drone by adjusting to people individually.
Sanne van Waveren, Rasmus Rudling, Iolanda Leite, Patric Jensfelt, Christian Pek
HRI5
2023 Aligning Human Preferences with Baseline Objectives in Reinforcement Learning
abstract
Practical implementations of deep reinforcement learning (deep RL) have been challenging due to an amplitude of factors, such as designing reward functions that cover every possible interaction. To address the heavy burden of robot reward engineering, we aim to leverage subjective human preferences gathered in the context of human-robot interaction, while taking advantage of a baseline reward function when available. By considering baseline objectives to be designed beforehand, we are able to narrow down the policy space, solely requesting human attention when their input matters the most. To allow for control over the optimization of different objectives, our approach contemplates a multi-objective setting. We achieve human-compliant policies by sequentially training an optimal policy from a baseline specification and collecting queries on pairs of trajectories. These policies are obtained by training a reward estimator to generate Pareto optimal policies that include human preferred behaviours. Our approach ensures sample efficiency and we conducted a user study to collect real human preferences, which we utilized to obtain a policy on a social navigation environment.
Daniel Marta, Simon Holk, Christian Pek, Jana Tumova, Iolanda Leite
ICRA3
2023 Risk-aware Spatio-temporal Logic Planning in Gaussian Belief Spaces
abstract
In many real-world robotic scenarios, we cannot assume exact knowledge about a robot's state due to unmodeled dynamics or noisy sensors. Planning in belief space addresses this problem by tightly coupling perception and planning modules to obtain trajectories that take into account the environment's stochasticity. However, existing works are often limited to tasks such as the classic reach-avoid problem and do not provide risk awareness. We propose a risk-aware planning strategy in belief space that minimizes the risk of violating a given specification and enables a robot to actively gather information about its state. We use Risk Signal Temporal Logic (RiSTL) as a specification language in belief space to express complex spatio-temporal missions including predicates over Gaussian beliefs. We synthesize trajectories for challenging scenarios that cannot be expressed through classical reach-avoid properties and show that risk-aware objectives improve the uncertainty reduction in a robot's belief.
Matti Vahs, Christian Pek, Jana Tumova
ICRA2
2023 VARIQuery: VAE Segment-Based Active Learning for Query Selection in Preference-Based Reinforcement Learning
abstract
Human-in-the-loop reinforcement learning (RL) methods actively integrate human knowledge to create reward functions for various robotic tasks. Learning from preferences shows promise as alleviates the requirement of demonstrations by querying humans on state-action sequences. However, the limited granularity of sequence-based approaches complicates temporal credit assignment. The amount of human querying is contingent on query quality, as redundant queries result in excessive human involvement. This paper addresses the often-overlooked aspect of query selection, which is closely related to active learning (AL). We propose a novel query selection approach that leverages variational autoencoder (VAE) representations of state sequences. In this manner, we formulate queries that are diverse in nature while simultaneously taking into account reward model estimations. We compare our approach to the current state-of-the-art query selection methods in preference-based RL, and find ours to be either on-par or more sample efficient through extensive benchmarking on simulated environments relevant to robotics. Lastly, we conduct an online study to verify the effectiveness of our query selection approach with real human feedback and examine several metrics related to human effort.
Daniel Marta, Simon Holk, Christian Pek, Jana Tumova, Iolanda Leite
IROS3
2023 Generating Scenarios from High-Level Specifications for Object Rearrangement Tasks
abstract
Rearranging objects is an essential skill for robots. To quickly teach robots new rearrangements tasks, we would like to generate training scenarios from high-level specifications that define the relative placement of objects for the task at hand. Ideally, to guide the robot's learning we also want to be able to rank these scenarios according to their difficulty. Prior work has shown how generating diverse scenario from specifications and providing the robot with easy-to-difficult samples can improve the learning. Yet, existing scenario generation methods typically cannot generate diverse scenarios while controlling their difficulty. We address this challenge by conditioning generative models on spatial logic specifications to generate spatially-structured scenarios that meet the specification and desired difficulty level. Our experiments showed that generative models are more effective and data-efficient than rejection sam-pling and that the spatially-structured scenarios can drastically improve training of downstream tasks by orders of magnitude.
Sanne van Waveren, Christian Pek, Iolanda Leite, Jana Tumova, Danica Kragic
IROS2
2023 Safe Data-Driven Model Predictive Control of Systems With Complex Dynamics
abstract
In this article, we address the task and safety performance of data-driven model predictive controllers (DD-MPC) for systems with complex dynamics, i.e., temporally or spatially varying dynamics that may also be discontinuous. The three challenges we focus on are the accuracy of learned models, the receding horizon-induced myopic predictions of DD-MPC, and the active encouragement of safety. To learn accurate models for DD-MPC, we cautiously, yet effectively, explore the dynamical system with rapidly exploring random trees (RRT) to collect a uniform distribution of samples in the state-input space and overcome the common distribution shift in model learning. The learned model is further used to construct an RRT tree that estimates how close the model's predictions are to the desired target. This information is used in the cost function of the DD-MPC to minimize the short-sighted effect of its receding horizon nature. To promote safety, we approximate sets of safe states using demonstrations of exclusively safe trajectories, i.e., without unsafe examples, and encourage the controller to generate trajectories close to the sets. As a running example, we use abrokenversion of an inverted pendulum where the friction abruptly changes in certain regions. Furthermore, we showcase the adaptation of our method to a real-world robotic application with complex dynamics: robotic food-cutting. Our results show that our proposed control framework effectively avoids unsafe states with higher success rates than baseline controllers that employ models from controlled demonstrations and even random actions.
Ioanna Mitsioni, Pouria Tajvar, Danica Kragic, Jana Tumova, Christian Pek
IEEE Trans. Robotics5
2022 Correct Me If I'm Wrong: Using Non-Experts to Repair Reinforcement Learning Policies
abstract
Reinforcement learning has shown great potential for learning sequential decision-making tasks. Yet, it is difficult to anticipate all possible real-world scenarios during training, causing robots to inevitably fail in the long run. Many of these failures are due to variations in the robot's environment. Usually experts are called to correct the robot's behavior; however, some of these failures do not necessarily require an expert to solve them. In this work, we query non-experts online for help and explore 1) if/how non-experts can provide feedback to the robot after a failure and 2) how the robot can use this feedback to avoid such failures in the future by generating shields that restrict or correct its high-level actions. We demonstrate our approach on common daily scenarios of a simulated kitchen robot. The results indicate that non-experts can indeed understand and repair robot failures. Our generated shields accelerate learning and improve data-efficiency during retraining.
Sanne van Waveren, Christian Pek, Jana Tumova, Iolanda Leite
HRI2
2022 Foresee the Unseen: Sequential Reasoning about Hidden Obstacles for Safe Driving
abstract
Safe driving requires autonomous vehicles to anticipate potential hidden traffic participants and other unseen objects, such as a cyclist hidden behind a large vehicle, or an object on the road hidden behind a building. Existing methods are usually unable to consider all possible shapes and orientations of such obstacles. They also typically do not reason about observations of hidden obstacles over time, leading to conservative anticipations. We overcome these limitations by (1) modeling possible hidden obstacles as a set of states of a point mass model and (2) sequential reasoning based on reachability analysis and previous observations. Based on (1), our method is safer, since we anticipate obstacles of arbitrary unknown shapes and orientations. In addition, (2) increases the available drivable space when planning trajectories for autonomous vehicles. In our experiments, we demonstrate that our method, at no expense of safety, gives rise to significant reductions in time to traverse various intersection scenarios from the CommonRoad Benchmark Suite.
José Manuel Gaspar Sánchez, Truls Nyberg, Christian Pek, Jana Tumova, Martin Törngren
IV3
2021 Encoding Human Driving Styles in Motion Planning for Autonomous Vehicles
abstract
Driving styles play a major role in the acceptance and use of autonomous vehicles. Yet, existing motion planning techniques can often only incorporate simple driving styles that are modeled by the developers of the planner and not tailored to the passenger. We present a new approach to encode human driving styles through the use of signal temporal logic and its robustness metrics. Specifically, we use a penalty structure that can be used in many motion planning frameworks, and calibrate its parameters to model different automated driving styles. We combine this penalty structure with a set of signal temporal logic formula, based on the Responsibility-Sensitive Safety model, to generate trajectories that we expected to correlate with three different driving styles: aggressive, neutral, and defensive. An online study showed that people perceived different parameterizations of the motion planner as unique driving styles, and that most people tend to prefer a more defensive automated driving style, which correlated to their self-reported driving style.
Jesper Karlsson, Sanne van Waveren, Christian Pek, Ilaria Torre 0002, Iolanda Leite, Jana Tumova
ICRA3
2021 Risk-aware Motion Planning for Autonomous Vehicles with Safety Specifications
abstract
Ensuring the safety of autonomous vehicles (AV s) in uncertain traffic scenarios is a major challenge. In this paper, we address the problem of computing the risk that AV s violate a given safety specification in uncertain traffic scenarios, where state estimates are not perfect. We propose a risk measure that captures the probability of violating the specification and determines the average expected severity of violation. Using highway scenarios of the US101 dataset and Responsible Sensitive Safety (RSS) as an example specification, we demonstrate the effectiveness and benefits of our proposed risk measure. By incorporating the risk measure into a trajectory planner, we enable AVs to plan minimal-risk trajectories and to quantify trade-offs between risk and progress in traffic scenarios.
Truls Nyberg, Christian Pek, Laura Dal Col, Christoffer Norén, Jana Tumova
IV2
2021 Learning Task Constraints in Visual-Action Planning from Demonstrations
abstract
Visual planning approaches have shown great success for decision making tasks with no explicit model of the state space. Learning a suitable representation and constructing a latent space where planning can be performed allows non-experts to setup and plan motions by just providing images. However, learned latent spaces are usually not semantically-interpretable, and thus it is difficult to integrate task constraints. We propose a novel framework to determine whether plans satisfy constraints given demonstrations of policies that satisfy or violate the constraints. The demonstrations are realizations of Linear Temporal Logic formulas which are employed to train Long Short-Term Memory (LSTM) networks directly in the latent space representation. We demonstrate that our architecture enables designers to easily specify, compose and integrate task constraints and achieves high performance in terms of accuracy. Furthermore, this visual planning framework enables human interaction, coping the environment changes that a human worker may involve. We show the flexibility of the method on a box pushing task in a simulated warehouse setting with different task constraints.
Francesco Esposito, Christian Pek, Michael C. Welle, Danica Kragic
RO-MAN2
2021 Fail-Safe Motion Planning for Online Verification of Autonomous Vehicles Using Convex Optimization
abstract
Safe motion planning for autonomous vehicles is a challenging task, since the exact future motion of other traffic participant is usually unknown. In this article, we present a verification technique ensuring that autonomous vehicles do not cause collisions by using fail-safe trajectories. Fail-safe trajectories are executed if the intended motion of the autonomous vehicle causes a safety-critical situation. Our verification technique is real-time capable and operates under the premise that intended trajectories are only executed if they have been verified as safe. The benefits of our proposed approach are demonstrated in different scenarios on an actual vehicle. Moreover, we present the first in-depth analysis of our verification technique used in dense urban traffic. Our results indicate that fail-safe motion planning has the potential to drastically reduce accidents while not resulting in overly conservative behaviors of the autonomous vehicle.
Christian Pek, Matthias Althoff
IEEE Trans. Robotics1
2020 Provably-Safe Cooperative Driving via Invariably Safe Sets
abstract
We address the problem of provably-safe cooperative driving for a group of vehicles that operate in mixed traffic scenarios, where both autonomous and human-driven vehicles are present. Our method is based on Invariably Safe Sets (ISSs), which are sets of states that let each of the cooperative vehicles remain safe for an infinite time horizon. The potential conflicts between the ISSs of a group of cooperative vehicles are resolved by examining and negotiating their Safe Maneuver Corridors. As a result, each vehicle obtains its negotiated ISS, which is used as target sets for motion planning. We demonstrate the applicability and benefits of our method on various traffic scenarios from the CommonRoad benchmark suite.
Edmond Irani Liu, Christian Pek, Matthias Althoff
IV2
2020 CommonRoad Drivability Checker: Simplifying the Development and Validation of Motion Planning Algorithms
abstract
Collision avoidance, kinematic feasibility, and road-compliance must be validated to ensure the drivability of planned motions for autonomous vehicles. Although these tasks are highly repetitive, computationally efficient toolboxes are still unavailable. The CommonRoad Drivability Checker-an open-source toolbox-unifies these mentioned checks. It is compatible with the CommonRoad benchmark suite, which additionally facilitates the development of motion planners. Our toolbox drastically reduces the effort of developing and validating motion planning algorithms. Numerical experiments show that our toolbox is real-time capable and can be used in real test vehicles.
Christian Pek, Vitaliy Rusinov, Stefanie Manzinger, Murat Can Üste, Matthias Althoff
IV1
2018 Efficient Computation of Invariably Safe States for Motion Planning of Self-Driving Vehicles
abstract
Safe motion planning requires that a vehicle reaches a set of safe states at the end of the planning horizon. However, safe states of vehicles have not yet been systematically defined in the literature, nor does a computationally efficient way to obtain them for online motion planning exist. To tackle the aforementioned issues, we introduce invariably safe sets. These are regions that allow vehicles to remain safe for an infinite time horizon. We show how invariably safe sets can be computed and propose a tight under-approximation which can be obtained efficiently in linear time with respect to the number of traffic participants. We use invariably safe sets to lift safety verification from finite to infinite time horizons. In addition, our sets can be used to determine the existence of feasible evasive maneuvers and the criticality of scenarios by computing the time-to-react metric.
Christian Pek, Matthias Althoff
IROS1
2018 Efficient Mixed-Integer Programming for Longitudinal and Lateral Motion Planning of Autonomous Vehicles
abstract
The application of continuous optimization to motion planning of autonomous vehicles has enjoyed increasing popularity in recent years. In order to maintain low computation times, it is advantageous to have a convex formulation, in general requiring the planning problem to be separated into a longitudinal and lateral component. However, this decoupling of the motion often results in infeasible trajectories in situations in which both components need to be heavily linked, e.g., when planning swerving maneuvers to avoid a collision with obstacles. In this work, we propose an approach which extends the convex optimization problem of the longitudinal component to incorporate changing constraints, allowing us to guarantee feasibility of the resulting combined trajectory. Furthermore, we provide additional safety guarantees for the planned motion by integrating formal safety distances assuming infinite precision arithmetic. Our approach is demonstrated using simulated lane change maneuvers.
Christina Miller, Christian Pek, Matthias Althoff
Intelligent Vehicles Symposium2
2017 Verifying the safety of lane change maneuvers of self-driving vehicles based on formalized traffic rules
abstract
Validating the safety of self-driving vehicles requires an enormous amount of testing. By applying formal verification methods, we can prove the correctness of the vehicles' behavior, which at the same time reduces remaining risks and the need for extensive testing. However, current safety approaches do not consider liabilities of traffic participants if a collision occurs. Utilizing formalized traffic rules to verify motion plans allows this problem to be solved. We present a novel approach for verifying the safety of lane change maneuvers, using formalized traffic rules according to the Vienna Convention on Road Traffic. This allows us to provide additional guarantees that if a collision occurs, the self-driving vehicle is not responsible. Furthermore, we consider misbehavior of other traffic participants during lane changes and propose feasible solutions to avoid or mitigate a potential collision. The approach has been evaluated using real traffic data provided by the NGSIM project as well as simulated lane changes.
Christian Pek, Peter Zahn, Matthias Althoff
Intelligent Vehicles Symposium1
2016 Simplifying synchronization in cooperative robot tasks - an enhancement of the Manipulation Primitive paradigm
abstract
Complex handling and assembly tasks can be gradually decomposed into simpler sub-tasks until a level of elementary 'primitive' tasks is reached, which can be expressed generically by so-called Manipulation Primitives. Manipulation Primitives can hence be utilized as the foundation of a unified paradigm to specify and execute complex sensor-based robot tasks. By embedding these Manipulation Primitives into a fully hierarchical structure, the specification and reuse of common sub-tasks can be simplified immediately. However, without appropriate synchronization mechanisms the user-friendly and effective specification of tasks for multi-robot systems based on Manipulation Primitives still remains an issue. The underlying place-transition net formalism allows for synchronization by including the notions of forks and joins into the net but this approach is neither user-friendly nor does it allow for immediate support of more sophisticated synchronization mechanisms such as range synchronization. To eliminate this drawback of the Manipulation Primitive Concept, this contribution proposes powerful but at the same time easy-to-use synchronization primitives and introduces a revised Manipulation Primitive paradigm, which takes full advantage of the developed primitives. The advantages of the proposed concepts regarding the specification of tasks for industrial multi-robot systems are illustrated with several examples.
Christian Pek, Arne Muxfeldt, Daniel Kubus
ETFA1
2014 Any Suggestions? Active Schema Support for Structuring Web Information
Silviu Homoceanu, Felix Geilert, Christian Pek, Wolf-Tilo Balke
DASFAA (2)3