Kai Oliver Arras

dblp:a/KOArras · also Kai O. Arras · DBLP profile ↗
← Back
70ranked-venue papers
7as first author
14since 2021 · last 2026
0009-0001-3003-8857ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 68 · 7 first-author · 14 since 2021Systems, architecture and hardware · 50 · 7 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 11 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2026 Context-Aware Generation and Modulation of Expressive Motion Behavior using Multimodal Foundation Models
abstract
Expressive robot motion positively impacts human-robot interaction by improving user engagement, likability, or task performance. We present a novel approach to automatically generate and modulate complex, expressive full-body gesture sequences from a multimodal (text, audio, video) context description. Based on a unified mathematical implementation of the Principles of Animation using Dynamic Movement Primitives, we use multimodal foundation models to generate such sequences along with parametric motion variations that are highly context-aware. This method extends the state of the art in terms of flexibility and generality by being interpretable, composable, and working across different robot morphologies. Moreover, we integrate the system into a continuous control framework and leverage knowledge distillation to learn a much smaller model, significantly improving token efficiency and system latency. Results from a user study with a human-like platform indicate that participants judged our system’s motions to be better aligned with the interaction context than motions produced under a non-modulated or API-level condition. We also demonstrate how the system generates expressive motion for robots with different kinematics, showcasing its versatility. Paper webpage: https://gen-mod-expressive-motion-behavior.github.io/.
Till Hielscher, Fabio Scaparro, Kai Oliver Arras
HRI3
2025 Multi-Heuristic Robotic Bin Packing of Regular and Irregular Objects
abstract
The increasing demand in e-commerce, combined with labor shortages and rising wages, is driving the rapid automation of warehouse operations. A critical aspect of this shift is bin packing, where diverse unknown items of varying sizes and shapes must be optimally arranged within a bin or container. Robot bin packing is receiving growing attention and presents unique challenges due to the broad range of objects, packing rules, and task-specific requirements. In response, we propose So-Pack, a generalist packing heuristic for irregularly shaped objects integrated into a flexible, weighted multi-heuristic planning system. The system demonstrates robust performance across general packing scenarios and flexibility to adapt to changing packing rules and specific end-user requirements. Experimental results show that the system outperforms state-of-the-art approaches in key metrics on a new challenging dataset of retail objects in real-world applications.
Tim Nickel, Richard Bormann, Kai Oliver Arras
ICRA3
2025 Interactive Expressive Motion Generation Using Dynamic Movement Primitives
abstract
Our goal is to enable social robots to interact autonomously with humans in a realistic, engaging, and expressive manner. The 12 Principles of Animation are a well-established framework animators use to create movements that make characters appear convincing, dynamic, and emotionally expressive. This paper proposes a novel approach that leverages Dynamic Movement Primitives (DMPs) to implement key animation principles, providing a learnable, explainable, modulable, online adaptable and composable model for automatic expressive motion generation. DMPs, originally developed for general imitation learning in robotics and grounded in a spring-damper system design, offer mathematical properties that make them particularly suitable for this task. Specifically, they enable modulation of the intensities of individual principles and facilitate the decomposition of complex, expressive motion sequences into learnable and parametrizable primitives. We present the mathematical formulation of the parameterized animation principles and demonstrate the effectiveness of our framework through experiments and application on three robotic platforms with different kinematic configurations, in simulation, on actual robots and in a user study. Our results show that the approach allows for creating diverse and nuanced expressions using a single base model.
Till Hielscher, Andreas Bulling, Kai Oliver Arras
IROS3
2025 Snuggle-Pack: Speeding Up Multi-Heuristic Packing Planning of Complex Objects
abstract
Efficient object packing is a fundamental challenge in logistics and industrial automation. This work introduces Snuggle-Pack, a novel 3D packing algorithm that integrates Fast Fourier Transform (FFT)-based spatial analysis with a multi-heuristic optimization framework to achieve real-time, high-density packing. Unlike traditional heuristic-based approaches that rely on 2D simplifications, our method operates in a fully 3D volumetric space, ensuring collision-free, stable, and physically feasible placements. At its core, our approach employs a proximity-aware and support-sensitive placement strategy, which encourages objects to fit snugly within their surroundings —hence the name—, optimizing space utilization through ne-grained collision metrics. We evaluate our method on the YCB and IPA-3D1K datasets in both previewed and ad-hoc packing scenarios. Our experiments show that Snuggle-Pack significantly outperforms the state of the art, achieving up to 25% higher packing densities or, alternatively, accelerating computation by up to 10×. Moreover, our framework allows for dynamic adaptation to custom constraints, such as balanced center of mass, weight limitations on fragile items, and safety proximity constraints. These results highlight Snuggle-Pack as an efficient, flexible, and scalable solution for industrial robotic packing tasks.
Tim Nickel, Richard Bormann, Kai Oliver Arras
IROS3
2025 FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction
abstract
The concept of 3D scene graphs is increasingly recognized as a powerful semantic and hierarchical representation of the environment. Current approaches often address this at a coarse, object-level resolution. In contrast, our goal is to develop a representation that enables robots to directly interact with their environment by identifying both the location of functional interactive elements and how these can be used. To achieve this, we focus on detecting and storing objects at a finer resolution, focusing on affordance-relevant parts. The primary challenge lies in the scarcity of data that extends beyond instance-level detection and the inherent difficulty of capturing detailed object features using robotic sensors. We leverage currently available 3D resources to generate 2D data and train a detector, which is then used to augment the standard 3D scene graph generation pipeline. Through our experiments, we demonstrate that our approach achieves functional element segmentation comparable to state-of-the-art 3D models and that our augmentation enables task-driven affordance grounding with higher accuracy than the current solutions. See our project page at https://fungraph.github.io.
Dennis Rotondi, Fabio Scaparro, Hermann Blum, Kai Oliver Arras
IROS4
2025 Incremental and Interactive Exploration of Robot Appearance Designs Using GenAI
abstract
A robot’s physical design impacts user acceptance, engagement, and trust while also influencing social and functional expectations about the robot’s capabilities. Robot design, and industrial design in general, can also be a driver of differentiation in a competitive market. In this paper, we leverage generative AI to incrementally explore design spaces of robot appearances. With the goal of overcoming training data bias of text-to-image models that favor stereotypical morphologies and designs, we propose a set of generation methods that enable a designer to guide the exploration process through text-, style-, and structure-based specifications from user or client feedback. Further, using Low-Rank Adaptation for model fine-tuning, the method allows to define the aesthetic direction by an image collection that conveys a particular style or theme ("mood boards"). The experiments demonstrate that our extensions retain image quality in terms of statistical and structural features and allow for both diversity and specificity in the design process. In a case study, we apply this method in the user-centered design process and discuss its opportunities and limitations.
Till Hielscher, Arne Bartenbach, Wolf Leonhardt, Kai Oliver Arras
RO-MAN4
2024 STARK: A Unified Framework for Strongly Coupled Simulation of Rigid and Deformable Bodies with Frictional Contact
abstract
The use of simulation in robotics is increasingly widespread for the purpose of testing, synthetic data generation and skill learning. A relevant aspect of simulation for a variety of robot applications is physics-based simulation of robot-object interactions. This involves the challenge of accurately modeling and implementing different mechanical systems such as rigid and deformable bodies as well as their interactions via constraints, contact or friction. Most state-of-the-art physics engines commonly used in robotics either cannot couple deformable and rigid bodies in the same framework, lack important systems such as cloth or shells, have stability issues in complex friction-dominated setups or cannot robustly prevent penetrations. In this paper, we propose a framework for strongly coupled simulation of rigid and deformable bodies with focus on usability, stability, robustness and easy access to state-of-the-art deformation and frictional contact models. Our system uses the Finite Element Method (FEM) to model deformable solids, the Incremental Potential Contact (IPC) approach for frictional contact and a robust second order optimizer to ensure stable and penetration-free solutions to tight tolerances. It is a general purpose framework, not tied to a particular use case such as grasping or learning, it is written in C++ and comes with a Python interface. We demonstrate our system’s ability to reproduce complex real-world experiments where a mobile vacuum robot interacts with a towel on different floor types and towel geometries. Our system is able to reproduce 100% of the qualitative outcomes observed in the laboratory environment. The simulation pipeline, named Stark (the German word for strong, as in strong coupling) is made open-source.
José Antonio Fernández-Fernández, Ralph Lange, Stefan Laible, Kai Oliver Arras, Jan Bender
ICRA4
2023 Semantically Informed MPC for Context-Aware Robot Exploration
abstract
We investigate the task of object goal navigation in unknown environments where a target object is given as a semantic label (e.g. find a couch). This task is challenging as it requires the robot to consider the semantic context in diverse settings (e.g. TVs are often nearby couches). Most of the prior work tackles this problem under the assumption of a discrete action policy whereas we present an approach with continuous control which brings it closer to real world applications. In this paper, we use information-theoretic model predictive control on dense cost maps to bring object goal navigation closer to real robots with kinodynamic constraints. We propose a deep neural network framework to learn cost maps that encode semantic context and guide the robot towards the target object. We also present a novel way of fusing mid-level visual representations in our architecture to provide additional semantic cues for cost map prediction. The experiments show that our method leads to more efficient and accurate goal navigation with higher quality paths than the reported baselines. The results also indicate the importance of mid-level representations for navigation by improving the success rate by 8 percentage points.
Yash Goel, Narunas Vaskevicius, Luigi Palmieri, Nived Chebrolu, Kai Oliver Arras, Cyrill Stachniss
IROS5
2023 Proactive Model Predictive Control with Multi-Modal Human Motion Prediction in Cluttered Dynamic Environments
abstract
For robots navigating in dynamic environments, exploiting and understanding uncertain human motion prediction is key to generate efficient, safe and legible actions. The robot may perform poorly and cause hindrances if it does not reason over possible, multi-modal future social interactions. With the goal of enhancing autonomous navigation in cluttered environments, we propose a novel formulation for nonlinear model predictive control including multi-modal predictions of human motion. As a result, our approach leads to less conservative, smooth and intuitive human-aware navigation with reduced risk of collisions, and shows a good balance between task efficiency, collision avoidance and human comfort. To show its effectiveness, we compare our approach against the state of the art in crowded simulated environments, and with real-world human motion data from the THOR dataset. This comparison shows that we are able to improve task efficiency, keep a larger distance to humans and significantly reduce the collision time, when navigating in cluttered dynamic environ-ments. Furthermore, the method is shown to work robustly with different state-of-the-art human motion predictors.
Lukas Heuer, Luigi Palmieri, Andrey Rudenko, Anna Mannucci, Martin Magnusson 0002, Kai Oliver Arras
IROS6
2023 CLiFF-LHMP: Using Spatial Dynamics Patterns for Long- Term Human Motion Prediction
abstract
Human motion prediction is important for mobile service robots and intelligent vehicles to operate safely and smoothly around people. The more accurate predictions are, particularly over extended periods of time, the better a system can, e.g., assess collision risks and plan ahead. In this paper, we propose to exploit maps of dynamics (MoDs, a class of general representations of place-dependent spatial motion patterns, learned from prior observations) for long-term human motion prediction (LHMP). We present a new MoD-informed human motion prediction approach, named CLiFF-LHMP, which is data efficient, explainable, and insensitive to errors from an upstream tracking system. Our approach uses CLiFF -map, a specific MoD trained with human motion data recorded in the same environment. We bias a constant velocity prediction with samples from the CLiFF-map to generate multi-modal trajectory predictions. In two public datasets we show that this algorithm outperforms the state of the art for predictions over very extended periods of time, achieving 45 % more accurate prediction performance at 50s compared to the baseline.
Andrey Rudenko, Tomasz Kucner, Luigi Palmieri, Kai Oliver Arras, Achim J. Lilienthal, Martin Magnusson 0002
IROS5
2023 Advantages of Multimodal versus Verbal-Only Robot-to-Human Communication with an Anthropomorphic Robotic Mock Driver
abstract
Robots are increasingly used in shared environments with humans, making effective communication a necessity for successful human-robot interaction. In our work, we study a crucial component: active communication of robot intent. Here, we present an anthropomorphic solution where a humanoid robot communicates the intent of its host robot acting as an “Anthropomorphic Robotic Mock Driver” (ARMoD). We evaluate this approach in two experiments in which participants work alongside a mobile robot on various tasks, while the ARMoD communicates a need for human attention, when required, or gives instructions to collaborate on a joint task. The experiments feature two interaction styles of the ARMoD: a verbal-only mode using only speech and a multimodal mode, additionally including robotic gaze and pointing gestures to support communication and register intent in space. Our results show that the multimodal interaction style, including head movements and eye gaze as well as pointing gestures, leads to more natural fixation behavior. Participants naturally identified and fixated longer on the areas relevant for intent communication, and reacted faster to instructions in collaborative tasks. Our research further indicates that the ARMoD intent communication improves engagement and social interaction with mobile robots in workplace settings.
Tim Schreiter, Lucas Morillo-Méndez, Ravi Chadalavada, Andrey Rudenko, Erik Billing, Martin Magnusson 0002, Kai Oliver Arras, Achim J. Lilienthal
RO-MAN7
2022 The Atlas Benchmark: an Automated Evaluation Framework for Human Motion Prediction
abstract
Human motion trajectory prediction, an essential task for autonomous systems in many domains, has been on the rise in recent years. With a multitude of new methods proposed by different communities, the lack of standardized benchmarks and objective comparisons is increasingly becoming a major limitation to assess progress and guide further research. Existing benchmarks are limited in their scope and flexibility to conduct relevant experiments and to account for contextual cues of agents and environments. In this paper we present Atlas, a benchmark to systematically evaluate human motion trajectory prediction algorithms in a unified framework. Atlas offers data preprocessing functions, hyperparameter optimization, comes with popular datasets and has the flexibility to setup and conduct underexplored yet relevant experiments to analyze a method’s accuracy and robustness. In an example application of Atlas, we compare five popular model- and learning-based predictors and find that, when properly applied, early physics-based approaches are still remarkably competitive. Such results confirm the necessity of benchmarks like Atlas.
Andrey Rudenko, Luigi Palmieri, Wanting Huang, Achim J. Lilienthal, Kai Oliver Arras
RO-MAN5
2022 How a Social Robot's Vocalization Affects Children's Speech, Learning, and Interaction
abstract
A wider incorporation of robots into classrooms is hampered by current technological limitations on full autonomy in social robots. Automated speech recognition, for example, a key enabler for vocal communication, is still unable to perform with sufficient accuracy. Past studies have shown that humans adjust their speech patterns to accommodate less skilled interlocutors. If such a response holds in human-robot interactions as well, we may be able to exploit it to lessen the burden on social robots and enable rich, autonomous vocal communication. In this paper we explore whether a robot’s speaking ability could have an impact on children’s speech patterns, learning, and engagement by designing an interaction where a child and a robot collaborate on a Tower of Hanoi puzzle. Sixteen children aged 7-14 completed this collaborative task partnered with a social robot that communicated with either high verbal (full sentences), low verbal (short phrases or single words), or nonverbal (sound-based utterances) vocalization. While we found no significant impact on children’s speech patterns or learning due to the robot’s method of vocalization, children in the non-verbal condition had a significantly lower perception of the robot’s intelligence along with higher rates of providing feedback and more instances of undoing its moves. This suggests that a link may exist between a robot’s perceived speaking ability and children’s confidence in that robot’s overall intelligence and capability in a collaborative task, as well as their empathy towards a peer they perceive as less skilled in the task.
Lauren L. Wright, Aditi Kothiyal, Kai Oliver Arras, Barbara Bruno
RO-MAN3
2021 Cross-Modal Analysis of Human Detection for Robotics: An Industrial Case Study
abstract
Advances in sensing and learning algorithms have led to increasingly mature solutions for human detection by robots, particularly in selected use-cases such as pedestrian detection for self-driving cars or close-range person detection in consumer settings. Despite this progress, the simple question which sensor-algorithm combination is best suited for a person detection task at handƒ remains hard to answer. In this paper, we tackle this issue by conducting a systematic cross-modal analysis of sensor-algorithm combinations typically used in robotics. We compare the performance of state-of-the-art person detectors for 2D range data, 3D lidar, and RGB-D data as well as selected combinations thereof in a challenging industrial use-case.We further address the related problems of data scarcity in the industrial target domain, and that recent research on human detection in 3D point clouds has mostly focused on autonomous driving scenarios. To leverage these methodological advances for robotics applications, we utilize a simple, yet effective multi-sensor transfer learning strategy by extending a strong image-based RGB-D detector to provide cross-modal supervision for lidar detectors in the form of weak 3D bounding box labels.Our results show a large variance among the different approaches in terms of detection performance, generalization, frame rates and computational requirements. As our use-case contains difficulties representative for a wide range of service robot applications, we believe that these results point to relevant open challenges for further research and provide valuable support to practitioners for the design of their robot system.
Timm Linder, Narunas Vaskevicius, Robert Schirmer, Kai Oliver Arras
IROS4
2020 Multi-Path Learning for Object Pose Estimation Across Domains
abstract
We introduce a scalable approach for object pose estimation trained on simulated RGB views of multiple 3D models together. We learn an encoding of object views that does not only describe an implicit orientation of all objects seen during training, but can also relate views of untrained objects. Our single-encoder-multi-decoder network is trained using a technique we denote ”multi-path learning”: While the encoder is shared by all objects, each decoder only reconstructs views of a single object. Consequently, views of different instances do not have to be separated in the latent space and can share common features. The resulting encoder generalizes well from synthetic to real data and across various instances, categories, model types and datasets. We systematically investigate the learned encodings, their generalization, and iterative refinement strategies on the ModelNet40 and T-LESS dataset. Despite training jointly on multiple objects, our 6D Object Detection pipeline achieves state-of-the-art results on T-LESS at much lower runtimes than competing approaches.
Martin Sundermeyer, Maximilian Durner, En Yen Puang, Zoltan-Csaba Marton, Narunas Vaskevicius, Kai Oliver Arras, Rudolph Triebel
CVPR6
2020 Metric-Scale Truncation-Robust Heatmaps for 3D Human Pose Estimation
abstract
Heatmap representations have formed the basis of 2D human pose estimation systems for many years, but their generalizations for 3D pose have only recently been considered. This includes 2.5D volumetric heatmaps, whose X and Y axes correspond to image space and the Z axis to metric depth around the subject. To obtain metric-scale predictions, these methods must include a separate, explicit post-processing step to resolve scale ambiguity. Further, they cannot encode body joint positions outside of the image boundaries, leading to incomplete pose estimates in case of image truncation. We address these limitations by proposing metric-scale truncation-robust (MeTRo) volumetric heatmaps, whose dimensions are defined in metric 3D space near the subject, instead of being aligned with image space. We train a fully-convolutional network to estimate such heatmaps from monocular RGB in an end-to-end manner. This reinterpretation of the heatmap dimensions allows us to estimate complete metric-scale poses without test-time knowledge of the focal length or person distance and without relying on anthropometric heuristics in post-processing. Furthermore, as the image space is decoupled from the heatmap space, the network can learn to reason about joints beyond the image boundary. Using ResNet-50 without any additional learned layers, we obtain state-of-the-art results on the Human3.6M and MPI-INF-3DHP benchmarks. As our method is simple and fast, it can become a useful component for real-time top-down multi-person pose estimation systems. We make our code publicly available to facilitate further research.
István Sárándi, Timm Linder, Kai Oliver Arras, Bastian Leibe
FG3
2020 Accurate detection and 3D localization of humans using a novel YOLO-based RGB-D fusion approach and synthetic training data
abstract
While 2D object detection has made significant progress, robustly localizing objects in 3D space under presence of occlusion is still an unresolved issue. Our focus in this work is on real-time detection of human 3D centroids in RGB-D data. We propose an image-based detection approach which extends the YOLO v3 architecture with a 3D centroid loss and mid-level feature fusion to exploit complementary information from both modalities. We employ a transfer learning scheme which can benefit from existing large-scale 2D object detection datasets, while at the same time learning end-to-end 3D localization from our highly randomized, diverse synthetic RGB-D dataset with precise 3D groundtruth. We further propose a geometrically more accurate depth-aware crop augmentation for training on RGB-D data, which helps to improve 3D localization accuracy. In experiments on our challenging intralogistics dataset, we achieve state-of-the-art performance even when learning 3D localization just from synthetic data.
Timm Linder, Kilian Y. Pfeiffer, Narunas Vaskevicius, Robert Schirmer, Kai Oliver Arras
ICRA5
2020 An NMPC Approach using Convex Inner Approximations for Online Motion Planning with Guaranteed Collision Avoidance
abstract
Even though mobile robots have been around for decades, trajectory optimization and continuous time collision avoidance remain subject of active research. Existing methods trade off between path quality, computational complexity, and kinodynamic feasibility. This work approaches the problem using a nonlinear model predictive control (NMPC) framework, that is based on a novel convex inner approximation of the collision avoidance constraint. The proposed Convex Inner ApprOximation (CIAO) method finds kinodynamically feasible and continuous time collision free trajectories, in few iterations, typically one. For a feasible initialization, the approach is guaranteed to find a feasible solution, i.e. it preserves feasibility. Our experimental evaluation shows that CIAO outperforms state of the art baselines in terms of planning efficiency and path quality. Experiments show that it also efficiently scales to high-dimensional systems. Furthermore real-world experiments demonstrate its capability of unifying trajectory optimization and tracking for safe motion planning in dynamic environments.
Tobias Schoels, Luigi Palmieri, Kai Oliver Arras, Moritz Diehl
ICRA3
2020 Plug-and-Play SLAM: A Unified SLAM Architecture for Modularity and Ease of Use
abstract
Simultaneous Localization and Mapping (SLAM) is considered a mature research field with numerous applications and publicly available open-source systems. Despite this maturity, existing SLAM systems often rely on ad-hoc implementations or are tailored to predefined sensor setups. In this work, we tackle these issues, proposing a novel unified SLAM architecture specifically designed to standardize the SLAM problem and to address heterogeneous sensor configurations. Thanks to its modularity and design patterns, the presented framework is easy to extend, maximizes code reuse and improves computational efficiency. We show in our experiments with a variety of typical sensor configurations that these advantages come without compromising state-of-the-art SLAM performance. The result demonstrates the architecture's relevance for facilitating further research in (multi-sensor) SLAM and its transfer into practical applications.
Mirco Colosi, Irvin Aloise, Tiziano Guadagnino, Dominik Schlegel, Bartolomeo Della Corte, Kai Oliver Arras, Giorgio Grisetti
IROS6
2019 Informed Information Theoretic Model Predictive Control
abstract
The problem of minimizing cost in nonlinear control systems with uncertainties or disturbances remains a major challenge. Model predictive control (MPC), and in particular sampling-based MPC has recently shown great success in complex domains such as aggressive driving with highly nonlinear dynamics. Sampling-based methods rely on a prior distribution to generate samples in the first place. Obviously, the choice of this distribution highly influences efficiency of the controller. Existing approaches such as sampling around the control trajectory of the previous time step perform suboptimally, especially in multi-modal or highly dynamic settings. In this work, we therefore propose to learn models that generate samples in low-cost areas of the state-space, conditioned on the environment and on contextual information of the task to solve. By using generative models as an informed sampling distribution, our approach exploits guidance from the learned models and at the same time maintains robustness properties of the MPC methods. We use Conditional Variational Autoencoders (CVAE) to learn distributions that imitate samples from a training dataset containing optimized controls. An extensive evaluation in the autonomous navigation domain suggests that replacing previous sampling schemes with our learned models considerably improves performance in terms of path quality and planning efficiency.
Raphael Kusumoto, Luigi Palmieri, Markus Spies, Akos Csiszar, Kai Oliver Arras
ICRA5
2019 Better Lost in Transition Than Lost in Space: SLAM State Machine
abstract
A Simultaneous Localization and Mapping (SLAM) system is a complex program consisting of several interconnected components with different functionalities such as optimization, tracking or loop detection. Whereas the literature addresses in detail how enhancing the algorithmic aspects of the individual components improves SLAM performance, the modal aspects, such as when to localize, relocalize or close a loop, are usually left aside. In this paper, we address the modal aspects of a SLAM system and show that the design of the modal controller has a strong impact on SLAM performance in particular in terms of robustness against unforeseen events such as sensor failures, perceptual aliasing or kidnapping. We preset a novel taxonomy for the components of a modern SLAM system, investigate their interplay and propose a highly modular architecture of a generic SLAM system using the Unified Modeling LanguageTM(UML) state machine formalism. The result, called SLAM state machine, is compared to the modal controller of several state-of-the-art SLAM systems and evaluated in two experiments. We demonstrate that our state machine handles unforeseen events much more robustly than the state-of-the-art systems.
Mirco Colosi, Sebastian Haug, Peter Biber, Kai Oliver Arras, Giorgio Grisetti
IROS4
2018 Semantic Labeling of Indoor Environments from 3D RGB Maps
abstract
We present an approach to automatically assign semantic labels to rooms reconstructed from 3D RGB maps of apartments. Evidence for the room types is generated using state-of-the-art deep-learning techniques for scene classification and object detection based on automatically generated virtual RGB views, as well as from a geometric analysis of the map's 3D structure. The evidence is merged in a conditional random field, using statistics mined from different datasets of indoor environments. We evaluate our approach qualitatively and quantitatively and compare it to related methods.
Manuel Brucker, Maximilian Durner, Rares Ambrus, Zoltan-Csaba Marton, Axel Wendt, Patric Jensfelt, Kai Oliver Arras, Rudolph Triebel
ICRA7
2018 Gradient-Informed Path Smoothing for Wheeled Mobile Robots
abstract
Planning smooth trajectories is important for the safe, efficient and comfortable operation of mobile robots, such as wheeled robots moving in crowded environments or cars moving at high speed. Asymptotically optimal sampling-based motion planners can be used to generate such trajectories. However, to achieve the necessary efficiency for the realtime operation of robots, one often uses their initial feasible trajectories or the trajectories of non-optimal motion planners instead, typically after a post-smoothing step. We propose a gradient-informed post-smoothing algorithm, called GRIPS, that deforms given trajectories by locally optimizing the placement of vertices while satisfying the system's kinodynamic constraints. We show experimentally that GRIPS typically produces trajectories of significantly smaller length and higher smoothness than several existing post-smoothing algorithms.
Eric Heiden, Luigi Palmieri, Sven Koenig, Kai Oliver Arras, Gaurav S. Sukhatme
ICRA4
2018 Joint Long-Term Prediction of Human Motion Using a Planning-Based Social Force Approach
abstract
The ability to perceive and predict future positions of dynamic objects is essential for mobile robots and intelligent vehicles in dynamic environments. In this paper, we present a novel planning-based approach for long-term human motion prediction that accounts for local interactions and can accurately predict joint motion of multiple agents. Long-term predictions are handled using an MDP formulation that computes a set of stochastic motion policies. To obtain distributions over future motion trajectories, we sample the policies with a weighted random walk algorithm in which each person is locally influenced by social forces from other nearby agents. Unlike related work, the algorithm is environment-aware, can account for individual agent velocities, requires no training phase and makes joint predictions for multiple agents. Experiments in simulation and with real data show that our method makes more accurate predictions than two state-of-the-art methods in terms of probabilistic and geometrical performance measures.
Andrey Rudenko, Luigi Palmieri, Kai Oliver Arras
ICRA3
2018 Human Motion Prediction Under Social Grouping Constraints
abstract
Accurate long-term prediction of human motion in populated spaces is an important but difficult task for mobile robots and intelligent vehicles. What makes this task challenging is that human motion is influenced by a large variety of factors including the person's intention, the presence, attributes, actions, social relations and social norms of other surrounding agents, and the geometry and semantics of the environment. In this paper, we consider the problem of computing human motion predictions that account for such factors. We formulate the task as an MDP planning problem with stochastic policies and propose a weighted random walk algorithm in which each agent is locally influenced by social forces from other nearby agents. The novelty of this paper is that we incorporate social grouping information into the prediction process reflecting the soft formation constraints that groups typically impose to their members' motion. We show that our method makes more accurate predictions than three state-of-the-art methods in terms of probabilistic and geometrical performance metrics.
Andrey Rudenko, Luigi Palmieri, Achim J. Lilienthal, Kai Oliver Arras
IROS4
2017 Kinodynamic motion planning on Gaussian mixture fields
abstract
We present a mobile robot motion planning approach under kinodynamic constraints that exploits learned perception priors in the form of continuous Gaussian mixture fields. Our Gaussian mixture fields are statistical multi-modal motion models of discrete objects or continuous media in the environment that encode e.g. the dynamics of air or pedestrian flows. We approach this task using a recently proposed circular linear flow field map based on semi-wrapped GMMs whose mixture components guide sampling and rewiring in an RRT* algorithm using a steer function for non-holonomic mobile robots. In our experiments with three alternative baselines, we show that this combination allows the planner to very efficiently generate high-quality solutions in terms of path smoothness, path length as well as natural yet minimum control effort motions through multi-modal representations of Gaussian mixture fields.
Luigi Palmieri, Tomasz Kucner, Martin Magnusson 0002, Achim J. Lilienthal, Kai Oliver Arras
ICRA5
2016 On multi-modal people tracking from mobile platforms in very crowded and dynamic environments
abstract
Tracking people is a key technology for robots and intelligent systems in human environments. Many person detectors, filtering methods and data association algorithms for people tracking have been proposed in the past 15+ years in both the robotics and computer vision communities, achieving decent tracking performances from static and mobile platforms in real-world scenarios. However, little effort has been made to compare these methods, analyze their performance using different sensory modalities and study their impact on different performance metrics. In this paper, we propose a fully integrated real-time multi-modal laser/RGB-D people tracking framework for moving platforms in environments like a busy airport terminal. We conduct experiments on two challenging new datasets collected from a first-person perspective, one of them containing very dense crowds of people with up to 30 individuals within close range at the same time. We consider four different, recently proposed tracking methods and study their impact on seven different performance metrics, in both single and multi-modal settings. We extensively discuss our findings, which indicate that more complex data association methods may not always be the better choice, and derive possible future research directions.
Timm Linder, Stefan Breuers, Bastian Leibe, Kai Oliver Arras
ICRA4
2016 Learning socially normative robot navigation behaviors with Bayesian inverse reinforcement learning
abstract
Mobile robots that navigate in populated environments require the capacity to move efficiently, safely and in human-friendly ways. In this paper, we address this task using a learning approach that enables a mobile robot to acquire navigation behaviors from demonstrations of socially normative human behavior. In the past, such approaches have been typically used to learn only simple behaviors under relatively controlled conditions using rigid representations or with methods that scale poorly to large domains. We thus develop a flexible graph-based representation able to capture relevant task structure and extend Bayesian inverse reinforcement learning to use sampled trajectories from this representation. In experiments with a real robot and a large-scale pedestrian simulator, we are able to show that the approach enables a robot to learn complex navigation behaviors of varying degrees of social normativeness using the same set of simple features.
Billy Okal, Kai Oliver Arras
ICRA2
2016 RRT-based nonholonomic motion planning using any-angle path biasing
abstract
RRT and RRT* have become popular planning techniques, in particular for high-dimensional systems such as wheeled robots with complex nonholonomic constraints. Their planning times, however, can scale poorly for such robots, which has motivated researchers to study hierarchical techniques that grow the RRT trees in more focused ways. Along this line, we introduce Theta*-RRT that hierarchically combines (discrete) any-angle search with (continuous) RRT motion planning for nonholonomic wheeled robots. Theta*-RRT is a variant of RRT that generates a trajectory by expanding a tree of geodesics toward sampled states whose distribution summarizes geometric information of the any-angle path. We show experimentally, for both a differential drive system and a high-dimensional truck-and-trailer system, that Theta*-RRT finds shorter trajectories significantly faster than four baseline planners (RRT, A*-RRT, RRT*, A*-RRT*) without loss of smoothness, while A*-RRT* and RRT* (and thus also Informed RRT*) fail to generate a first trajectory sufficiently fast in environments with complex nonholonomic constraints. We also prove that Theta*-RRT retains the probabilistic completeness of RRT for all small-time controllable systems that use an analytical steer function.
Luigi Palmieri, Sven Koenig, Kai Oliver Arras
ICRA3
2016 Practical Bayesian Inverse Reinforcement Learning for Robot Navigation
Billy Okal, Kai Oliver Arras
ECML/PKDD (3)2
2016 Errare humanum est: Erroneous robots in human-robot interaction
abstract
Perfect memory, strong reasoning abilities and flawless performance are typical cognitive traits associated with robots. In contrast, forgetting and erroneous reasoning are typical cognitive patterns of humans. This discrepancy may fundamentally affect the way how robots and humans interact and collaborate together and is today still little explored. In this paper, we investigate the effect of differences between erroneous and perfect robots in a competitive scenario in which humans and robots solve reasoning tasks and memorize numbers. Participants are randomly assigned to one of two groups: in the first group they interact with a perfect, flawless robot, while in the second, they interact with a human-like robot with occasional errors and imperfect memorizing abilities. Participants rate attitude, sympathy, and attributes of the robot in a questionnaire and we measure their task performance. The results show that the erroneous robot triggered more positive emotions but lead to a lower human performance than the perfect one. Effects of both conditions on the group of students with and without technical background are reported.
Marco Ragni, Andrey Rudenko, Barbara Kuhnert, Kai Oliver Arras
RO-MAN4
2015 Real-time full-body human gender recognition in (RGB)-D data
abstract
Understanding social context is an important skill for robots that share a space with humans. In this paper, we address the problem of recognizing gender, a key piece of information when interacting with people and understanding human social relations and rules. Unlike previous work which typically considered faces or frontal body views in image data, we address the problem of recognizing gender in RGB-D data from side and back views as well. We present a large, gender-balanced, annotated, multi-perspective RGB-D dataset with full-body views of over a hundred different persons captured with both the Kinect v1 and Kinect v2 sensor. We then learn and compare several classifiers on the Kinect v2 data using a HOG baseline, two state-of-the-art deep-learning methods, and a recent tessellation-based learning approach. Originally developed for person detection in 3D data, the latter is able to learn the best selection, location and scale of a set of simple point cloud features. We show that for gender recognition, it outperforms the other approaches for both standing and walking people while being very efficient to compute with classification rates up to 150 Hz.
Timm Linder, Sven Wehner, Kai Oliver Arras
ICRA3
2015 Distance metric learning for RRT-based motion planning with constant-time inference
abstract
The distance metric is a key component in RRT-based motion planning that deeply affects coverage of the state space, path quality and planning time. With the goal to speed up planning time, we introduce a learning approach to approximate the distance metric for RRT-based planners. By exploiting a novel steer function which solves the two-point boundary value problem for wheeled mobile robots, we train a simple nonlinear parametric model with constant-time inference that is shown to predict distances accurately in terms of regression and ranking performance. In an extensive analysis we compare our approach to an Euclidean distance baseline, consider four alternative regression models and study the impact of domain-specific feature expansion. The learning approach is shown to be faster in planning time by several factors at negligible loss of path quality.
Luigi Palmieri, Kai Oliver Arras
ICRA2
2015 Real-time full-body human attribute classification in RGB-D using a tessellation boosting approach
abstract
Robots that cooperate and interact with humans require the capacity to detect and track people, analyze their behavior and understand human social relations and rules. A key piece of information for such tasks are human attributes like gender, age, hair or clothing. In this paper, we address the problem of recognizing such attributes in RGB-D data from varying full-body views. To this end, we extend a recent tessellation boosting approach which learns the best selection, location and scale of a set of simple RGB-D features. The approach outperforms the original approach and a HOG baseline for five human attributes including gender, has long hair, has long trousers, has long sleeves and has jacket. Experiments on a multi-perspective RGB-D dataset with full-body views of over a hundred different persons show that the method is able to robustly recognize multiple attributes across different view directions and distances to the sensor with accuracies up to 90%. Our methods runs in real-time, achieving a classification rate of around 300 Hz for a single attribute.
Timm Linder, Kai Oliver Arras
IROS2
2014 Schedule-Based Robotic Search for Multiple Residents in a Retirement Home Environment
abstract
In this paper we address the planning problem of a robot searching for multiple residents in a retirement home in order to remind them of an upcoming multi-person recreational activity before a given deadline. We introduce a novel Multi-User Schedule Based (M-USB) Search approach which generates a high-level-plan to maximize the number of residents that are found within the given time frame. From the schedules of the residents, the layout of the retirement home environment as well as direct observations by the robot, we obtain spatio-temporal likelihood functions for the individual residents. The main contribution of our work is the development of a novel approach to compute a reward to find a search plan for the robot using: 1) the likelihood functions, 2) the availabilities of the residents, and 3) the order in which the residents should be found. Simulations were conducted on a floor of a real retirement home to compare our proposed M-USB Search approach to a Weighted Informed Walk and a Random Walk. Our results show that the proposed M-USB Search finds residents in a shorter amount of time by visiting fewer rooms when compared to the other approaches.
Markus Sebastian Schwenk, Tiago Stegun Vaquero, Goldie Nejat, Kai Oliver Arras
AAAI4
2014 Multi-model hypothesis tracking of groups of people in RGB-D data
Timm Linder, Kai Oliver Arras
FUSION2
2014 Robotic tele-presence with DARYL in the wild
abstract
This paper describes the results of a qualitative analysis of questionnaire data collected during a public exhibition of our robotic tele-presence system. In Summer 2013 the mildly humanized robot DARYL could be tried out by the general public during our University's science fair in the city center. People were given the chance to communicate through the robot with their peers and to perceive the world through the "eyes" and "ears" of the robot by means of a head-mounted display with attached headphones. An operator's voice was instantaneously transmitted to the robot's location and his or her head movements were tracked to enable direct, intuitive control of the robot's head movements. Twenty-seven people were interviewed in a structured way about their impressions and opinions after having either operated or interacted with the tele-operated robot. A careful analysis of the acquired data reveals a rather positive evaluation of the tele-presence system and interesting opinions about suitable application areas. These findings may guide designers of robotic tele-presence systems, a research area of increasing popularity.
Christian Becker-Asano, Kai Oliver Arras, Bernhard Nebel
HAI2
2014 A novel RRT extend function for efficient and smooth mobile robot motion planning
abstract
In this paper we introduce a novel RRT extend function for wheeled mobile robots. The approach computes closed-loop forward simulations based on the kinematic model of the robot and enables the planner to efficiently generate smooth and feasible paths that connect any pairs of states. We extend the control law of an existing discontinuous state feedback controller to make it usable as an RRT extend function and prove that all relevant stability properties are retained. We study the properties of the new approach as extender for RRT and RRT* and compare it systematically to a spline-based approach and a large and small set of motion primitives. The results show that our approach generally produces smoother paths to the goal in less time with smaller trees. For RRT*, the approach produces also the shortest paths and achieves the lowest cost solutions when given more planning time.
Luigi Palmieri, Kai Oliver Arras
IROS2
2014 Inverse Reinforcement Learning algorithms and features for robot navigation in crowds: An experimental comparison
abstract
For mobile robots which operate in human populated environments, modeling social interactions is key to understand and reproduce people's behavior. A promising approach to this end is Inverse Reinforcement Learning (IRL) as it allows to model the factors that motivate people's actions instead of the actions themselves. A crucial design choice in IRL is the selection of features that encode the agent's context. In related work, features are typically chosen ad hoc without systematic evaluation of the alternatives and their actual impact on the robot's task. In this paper, we introduce a new software framework to systematically investigate the effect features and learning algorithms used in the literature. We also present results for the task of socially compliant robot navigation in crowds, evaluating two different IRL approaches and several feature sets in large-scale simulations. The results are benchmarked according to a proposed set of objective and subjective performance metrics.
Dizan Vasquez, Billy Okal, Kai Oliver Arras
IROS3
2014 R2-D2 Reloaded: A flexible sound synthesis system for sonic human-robot interaction design
abstract
A key skill for social robots is the ability to communicate their inner state to humans. In this paper, we explore abstracted robot-specific ways of interaction as an alternative to human-like or animal-like social cues. In particular, we present a sound system as a novel modality that extends a robot's ability for non-verbal communication. Unlike prior work which used pre-recorded audio samples to this end, we propose a flexible architecture with a generalized sound synthesizer that uses the principle of modulation to shape the sound in real-time by external and internal stimuli from the robot or the interaction. This allows for almost unlimited possibilities in the design of an expressive auditory social cue for human-robot interaction. We instantiate the architecture and report on example design choices for the sound synthesis principle, the real-time synthesizer, the sound modulation routings, and a sound sequence composer. We then demonstrate the system's ability for affect communication of primary and secondary emotions on a social robot.
Markus Sebastian Schwenk, Kai Oliver Arras
RO-MAN2
2013 Robot embodiment, operator modality, and social interaction in tele-existence: a project outline
Christian Becker-Asano, Severin Gustorff, Kai Oliver Arras, Kohei Ogawa, Shuichi Nishio, Hiroshi Ishiguro, Bernhard Nebel
HRI3
2012 Leveraging RGB-D Data: Adaptive fusion and domain adaptation for object detection
abstract
Vision and range sensing belong to the richest sensory modalities for perception in robotics and related fields. This paper addresses the problem of how to best combine image and range data for the task of object detection. In particular, we propose a novel adaptive fusion approach, hierarchical Gaussian Process mixtures of experts, able to account for missing information and cross-cue data consistency. The hierarchy is a two-tier architecture that for each modality, each frame and each detection computes a weight function using Gaussian Processes that reflects the confidence of the respective information. We further propose a method called cross-cue domain adaptation that makes use of large image data sets to improve the depth-based object detector for which only few training samples exist. In the experiments that include a comparison with alternative sensor fusion schemes, we demonstrate the viability of the proposed methods and achieve significant improvements in classification accuracy.
Luciano Spinello, Kai Oliver Arras
ICRA2
2012 Socially-aware robot navigation: A learning approach
abstract
The ability to act in a socially-aware way is a key skill for robots that share a space with humans. In this paper we address the problem of socially-aware navigation among people that meets objective criteria such as travel time or path length as well as subjective criteria such as social comfort. Opposed to model-based approaches typically taken in related work, we pose the problem as an unsupervised learning problem. We learn a set of dynamic motion prototypes from observations of relative motion behavior of humans found in publicly available surveillance data sets. The learned motion prototypes are then used to compute dynamic cost maps for path planning using an any-angle A* algorithm. In the evaluation we demonstrate that the learned behaviors are better in reproducing human relative motion in both criteria than a Proxemics-based baseline method.
Matthias Luber, Luciano Spinello, Jens Silva, Kai Oliver Arras
IROS4
2012 Robot-specific social cues in emotional body language
abstract
Humans use very sophisticated ways of bodily emotion expression combining facial expressions, sound, gestures and full body posture. Like others, we want to apply these aspects of human communication to ease the interaction between robots and users. In doing so we believe there is a need to consider what abstraction of human social communicative behaviors is appropriate for robots. The study reported in this paper is a pilot study to not offer simulated emotion but to offer an abstracted robot version of emotion expressions and an evaluation to what extent users interpret these robot expressions as the intended emotional states. To this end, we present the mobile, mildly humanized robot Daryl, for which we created six motion sequences that combine human-like, animal-like, and robot-specific social cues. The results of a user study (N=29) show that despite the absence of facial expressions and articulated extremities, subjects' interpretation of Daryl's emotional states were congruent with the abstracted emotion display. These results demonstrate that abstract displays of emotion that combine human-like, animal-like, and robot-specific modalities could in fact be an alternative to complex facial expressions and will feed into ongoing work identifying robot-specific social cues.
Stephanie Embgen, Matthias Luber, Christian Becker-Asano, Marco Ragni, Vanessa Evers, Kai Oliver Arras
RO-MAN6
2012 Audio-based human activity recognition using Non-Markovian Ensemble Voting
abstract
Human activity recognition is a key component for socially enabled robots to effectively and naturally interact with humans. In this paper we exploit the fact that many human activities produce characteristic sounds from which a robot can infer the corresponding actions. We propose a novel recognition approach called Non-Markovian Ensemble Voting (NEV) able to classify multiple human activities in an online fashion without the need for silence detection or audio stream segmentation. Moreover, the method can deal with activities that are extended over undefined periods in time. In a series of experiments in real reverberant environments, we are able to robustly recognize 22 different sounds that correspond to a number of human activities in a bathroom and kitchen context. Our method outperforms several established classification techniques.
Johannes A. Stork, Luciano Spinello, Jens Silva, Kai Oliver Arras
RO-MAN4
2011 Better models for people tracking
abstract
People tracking is a key component for robots operating in populated environments. Previous works have employed different filtering and data association techniques for this purpose that typically rely on a set of generic assumptions on target behavior and detector characteristics. In this paper, we focus on these assumptions rather than the tracking approach itself and show that with informed models, people tracking can be made substantially more accurate without compromising efficiency. Concretely, we present better, human-specific models for the occurrence of new tracks, false alarms, track occlusions, and track deletions. In the experiments with a large-scale outdoor data set collected with a laser range finder, the models and combinations thereof are experimentally compared using a multi-hypothesis baseline tracker and the CLEAR MOT metrics. The results show how some models selectively improve tracking performance at the expense of other measures. The final combination is then able to resolve the trade-offs, leading to a reduction of data association errors by more than a factor of two at the same cost.
Matthias Luber, Gian Diego Tipaldi, Kai Oliver Arras
ICRA3
2011 Tracking people in 3D using a bottom-up top-down detector
abstract
People detection and tracking is a key component for robots and autonomous vehicles in human environments. While prior work mainly employed image or 2D range data for this task, in this paper, we address the problem using 3D range data. In our approach, a top-down classifier selects hypotheses from a bottom-up detector, both based on sets of boosted features. The bottom-up detector learns a layered person model from a bank of specialized classifiers for different height levels of people that collectively vote into a continuous space. Modes in this space represent detection candidates that each postulate a segmentation hypothesis of the data. In the top-down step, the candidates are classified using features that are computed in voxels of a boosted volume tessellation. We learn the optimal volume tessellation as it enables the method to stably deal with sparsely sampled and articulated objects. We then combine the detector with tracking in 3D for which we take a multi-target multi-hypothesis tracking approach. The method neither needs a ground plane assumption nor relies on background learning. The results from experiments in populated urban environments demonstrate 3D tracking and highly robust people detection up to 20 m with equal error rates of at least 93%.
Luciano Spinello, Matthias Luber, Kai Oliver Arras
ICRA3
2011 I want my coffee hot! Learning to find people under spatio-temporal constraints
abstract
In this paper we present a probabilistic model for spatio-temporal patterns of human activities that enable robots to blend themselves into the workflows and daily routines of people. The model, called spatial affordance map, is a non-homogeneous spatial Poisson process that relates space, time and occurrence probability of activity events. We describe how learning and inference is made and present a novel planning algorithm that produces paths which maximize the probability to encounter a person. We show that the problem is a special class of the orienteering problem that can be solved as a finite horizon Markov decision process. We develop a simulator of populated office environments to validate the model and the planning algorithm. The simulated agents follow activity patterns learned by administering a questionnaire to 27 colleagues over two weeks. The experiments shows that the model is statistically valid with respect to both the Anderson-Darling test and the expected waiting time estimation. They further show that the proposed algorithm is able to find optimal paths.
Gian Diego Tipaldi, Kai Oliver Arras
ICRA2
2011 People tracking in RGB-D Data with on-line boosted target models
abstract
People tracking is a key component for robots that are deployed in populated environments. Previous works have used cameras and 2D and 3D range finders for this task. In this paper, we present a 3D people detection and tracking approach using RGB-D data. We combine a novel multi-cue person detector for RGB-D data with an on-line detector that learns individual target models. The two detectors are integrated into a decisional framework with a multi-hypothesis tracker that controls on-line learning through a track interpretation feedback. For on-line learning, we take a boosting approach using three types of RGB-D features and a confidence maximization search in 3D space. The approach is general in that it neither relies on background learning nor a ground plane assumption. For the evaluation, we collect data in a populated indoor environment using a setup of three Microsoft Kinect sensors with a joint field of view. The results demonstrate reliable 3D tracking of people in RGB-D data and show how the framework is able to avoid drift of the on-line detector and increase the overall tracking performance.
Matthias Luber, Luciano Spinello, Kai Oliver Arras
IROS3
2011 People detection in RGB-D Data
abstract
People detection is a key issue for robots and intelligent systems sharing a space with people. Previous works have used cameras and 2D or 3D range finders for this task. In this paper, we present a novel people detection approach for RGB-D data. We take inspiration from the Histogram of Oriented Gradients (HOG) detector to design a robust method to detect people in dense depth data, called Histogram of Oriented Depths (HOD). HOD locally encodes the direction of depth changes and relies on an depth-informed scale-space search that leads to a 3-fold acceleration of the detection process. We then propose Combo-HOD, a RGB-D detector that probabilistically combines HOD and HOG. The experiments include a comprehensive comparison with several alternative detection approaches including visual HOG, several variants of HOD, a geometric person detector for 3D point clouds, and an Haar-based AdaBoost detector. With an equal error rate of 85% in a range up to 8m, the results demonstrate the robustness of HOD and Combo-HOD on a real-world data set collected with a Kinect sensor in a populated indoor environment.
Luciano Spinello, Kai Oliver Arras
IROS2
2011 Please do not disturb! Minimum interference coverage for social robots
abstract
In this paper we address the problem of human-aware coverage planning. We first present an approach to learn and model human activity events in a probabilistic spatio-temporal map using spatial Poisson processes. We then propose a coverage planner for paths that minimize the interference probability with people. To this end, we pose the coverage problem as an asymmetric traveling salesman problem with time-dependent costs (ATDTSP) derived from the information in the map. The approach enables a noisy robotic vacuum in a home scenario, for instance, to learn to avoid busy places at certain times of the day such as the kitchen at lunch time. We evaluate the planner using a simulator of people in a home environment to generate typical weekday activity patterns. In the experiments with a regular TSP planner and two modified TSP heuristics, the proposed coverage planner significantly reduces interference with people in terms of number of disturbed persons and overall disturbance time.
Gian Diego Tipaldi, Kai Oliver Arras
IROS2
2010 A Layered Approach to People Detection in 3D Range Data
abstract
People tracking is a key technology for autonomous systems, intelligent cars and social robots operating in populated environments. What makes the task difficult is that the appearance of humans in range data can change drastically as a function of body pose, distance to the sensor, self-occlusion and occlusion by other objects. In this paper we propose a novel approach to pedestrian detection in 3D range data based on supervised learning techniques to create a bank of classifiers for different height levels of the human body. In particular, our approach applies AdaBoost to train a strong classifier from geometrical and statistical features of groups of neighboring points at the same height. In a second step, the AdaBoost classifiers mutually enforce their evidence across different heights by voting into a continuous space. Pedestrians are finally found efficiently by mean-shift search for local maxima in the voting space. Experimental results carried out with 3D laser range data illustrate the robustness and efficiency of our approach even in cluttered urban environments. The learned people detector reaches a classification rate up to 96% from a single 3D scan.
Luciano Spinello, Kai Oliver Arras, Rudolph Triebel, Roland Siegwart
AAAI2
2010 Exploiting Repetitive Object Patterns for Model Compression and Completion
Luciano Spinello, Rudolph Triebel, Dizan Vasquez, Kai Oliver Arras, Roland Siegwart
ECCV (5)4
2010 People tracking with human motion predictions from social forces
abstract
For many tasks in populated environments, robots need to keep track of current and future motion states of people. Most approaches to people tracking make weak assumptions on human motion such as constant velocity or acceleration. But even over a short period, human behavior is more complex and influenced by factors such as the intended goal, other people, objects in the environment, and social rules. This motivates the use of more sophisticated motion models for people tracking especially since humans frequently undergo lengthy occlusion events. In this paper, we consider computational models developed in the cognitive and social science communities that describe individual and collective pedestrian dynamics for tasks such as crowd behavior analysis. In particular, we integrate a model based on a social force concept into a multi-hypothesis target tracker. We show how the refined motion predictions translate into more informed probability distributions over hypotheses and finally into a more robust tracking behavior and better occlusion handling. In experiments in indoor and outdoor environments with data from a laser range finder, the social force model leads to more accurate tracking with up to two times fewer data association errors.
Matthias Luber, Johannes A. Stork, Gian Diego Tipaldi, Kai Oliver Arras
ICRA4
2010 FLIRT - Interest regions for 2D range data
abstract
Local image features are used for a wide range of applications in computer vision and range imaging. While there is a great variety of detector-descriptor combinations for image data and 3D point clouds, there is no general method readily available for 2D range data. For this reason, the paper first proposes a set of benchmark experiments on detector repeatability and descriptor matching performance using known indoor and outdoor data sets for robot navigation. Secondly, the paper introduces FLIRT that stands for Fast Laser Interest Region Transform, a multi-scale interest region operator for 2D range data. FLIRT combines the best detector with the best descriptor, experimentally found in a comprehensive analysis of alternative detector and descriptor approaches. The analysis yields repeatability and matching performance results similar to the values found for features in the computer vision literature, encouraging a wide range of applications of FLIRT on 2D range data. We finally show how FLIRT can be used in conjunction with RANSAC to address the loop closing/global localization problem in SLAM in indoor as well as outdoor environments. The results demonstrate that FLIRT features have a great potential for robot navigation in terms of precision-recall performance, efficiency and generality.
Gian Diego Tipaldi, Kai Oliver Arras
ICRA2
2009 Tracking groups of people with a multi-model hypothesis tracker
abstract
People in densely populated environments typically form groups that split and merge. In this paper we track groups of people so as to reflect this formation process and gain efficiency in situations where maintaining the state of individual people would be intractable. We pose the group tracking problem as a recursive multi-hypothesis model selection problem in which we hypothesize over both, the partitioning of tracks into groups (models) and the association of observations to tracks (assignments). Model hypotheses that include split, merge, and continuation events are first generated in a data-driven manner and then validated by means of the assignment probabilities conditioned on the respective model. Observations are found by clustering points from a laser range finder given a background model and associated to existing group tracks using the minimum average Hausdorff distance. Experiments with a stationary and a moving platform show that, in populated environments, tracking groups is clearly more efficient than tracking people separately. Our system runs in real-time on a typical desktop computer.
Boris Lau, Kai Oliver Arras, Wolfram Burgard
ICRA2
2009 Place-Dependent People Tracking
Matthias Luber, Gian Diego Tipaldi, Kai Oliver Arras
ISRR3
2008 Efficient people tracking in laser range data using a multi-hypothesis leg-tracker with adaptive occlusion probabilities
abstract
We present an approach to laser-based people tracking using a multi-hypothesis tracker that detects and tracks legs separately with Kalman filters, constant velocity motion models, and a multi-hypothesis data association strategy. People are defined as high-level tracks consisting of two legs that are found with little model knowledge. We extend the data association so that it explicitly handles track occlusions in addition to detections and deletions. Additionally, we adapt the corresponding probabilities in a situation-dependent fashion so as to reflect the fact that legs frequently occlude each other. Experimental results carried out with a mobile robot illustrate that our approach can robustly and efficiently track multiple people even in situations of high levels of occlusion.
Kai Oliver Arras, Slawomir Grzonka, Matthias Luber, Wolfram Burgard
ICRA1
2007 Using Boosted Features for the Detection of People in 2D Range Data
abstract
This paper addresses the problem of detecting people in two dimensional range scans. Previous approaches have mostly used pre-defined features for the detection and tracking of people. We propose an approach that utilizes a supervised learning technique to create a classifier that facilitates the detection of people. In particular, our approach applies AdaBoost to train a strong classifier from simple features of groups of neighboring beams corresponding to legs in range data. Experimental results carried out with laser range data illustrate the robustness of our approach even in cluttered office environments
Kai Oliver Arras, Óscar Martínez Mozos, Wolfram Burgard
ICRA1
2004 2D Mapping of Cluttered Indoor Environments by Means of 3D Perception
abstract
This paper presents a combination of a 3D laser sensor and a line-base SLAM algorithm which together produce 2D line maps of highly cluttered indoor environments. The key of the described method is the replacement of commonly used 2D laser range sensors by 3D perception. A straightforward algorithm extracts a virtual 2D scan that also contains partially occluded walls. These virtual scans are used as input for SLAM using line segments as features. The paper presents the used algorithms and experimental results that were made in a former industrial bakery. The focus lies on scenes that are known to be problematic for pure 2D systems. The results demonstrate that mapping indoor environments can be made robust with respect to both, poor odometry and clutter.
Oliver Wulf, Kai Oliver Arras, Henrik I. Christensen, Bernardo Wagner
ICRA2
2003 A navigation framework for multiple mobile robots and its application at the Expo.02 exhibition
abstract
This paper presents a navigation framework which enables multiple mobile robots to attain individual goals, coordinate their actions and work safely and reliably in a highly dynamic environment. We give an overview of the framework architecture, its layering and the subsystems reactive obstacle avoidance, local path planning, global path planning, multi-robot planning and localization. The latter receives particular attention as the localization problem is a key issue for navigation in unmodified and difficult environments. The framework permits a lightweight implementation on a fully autonomous robot. This is the result of a design effort striving for compact representations and computational efficiency. The experimental testbed was the "Robotics" pavilion at the Swiss National Exhibition Expo.02 where ten fully autonomous robots were interacting with more than half a million visitors during a five-month period on 3316 km.
Kai Oliver Arras, Roland Philippsen, Nicola Tomatis, Marc De Battista, Martin Schilt, Roland Siegwart
ICRA1
2003 Designing a secure and robust mobile interacting robot for the long term
abstract
This paper presents the genesis of RoboX. This tour guide robot has been built from the scratch based on the experience of the Autonomous Systems Lab. The production of 11 of those machines has been realized by a spin-off of the lab: BlueBotics SA. The goal was to maximize the autonomy and interactivity of the mobile platform while ensuring high robustness, security and performance. The result is an interactive moving machine which can operate in human environments and interacts by seeing humans, talking to and looking at them, showing icons and asking them to answer its questions. The complete design of mechanics, electronics and software is presented in the first part. Then, as extraordinary test bed, the Robotics exhibition at Expo.02 (Swiss National Exhibition) permits to establish meaningful statistics over 5 months (from May 15 to October 20, 2002) with up to 11 robots operating at the same time.
Nicola Tomatis, Gregoire Terrien, Ralph Piguet, Daniel Burnier, Samir Bouabdallah, Kai Oliver Arras, Roland Siegwart
ICRA6
2003 Multi-resolution SLAM for Real World Navigation
Agostino Martinelli, Adriana Tapus, Kai Oliver Arras, Roland Siegwart
ISRR3
2002 Feature-Based Multi-Hypothesis Localization and Tracking for Mobile Robots using Geometric Constraints
abstract
In this paper we present a new probabilistic feature-based approach to multi-hypothesis global localization and pose tracking. Hypotheses are generated using a constraint-based search in the interpretation tree of possible local-to-global pairings. This results in a set of robot location hypotheses of unbounded accuracy. For tracking, the same constraint-based technique is used. It performs track splitting as soon as location ambiguities arise from uncertainties and sensing. This yields a very robust localization technique which can deal with significant errors from odometry, collisions and kidnapping. Simulation experiments and first tests with a real robot demonstrate these properties at very low computational cost. The presented approach is theoretically sound which makes that the only parameter is the significance level on which all statistical decisions are taken.
Kai Oliver Arras, José A. Castellanos 0001, Roland Siegwart
ICRA1
2002 Real-Time Obstacle Avoidance for Polygonal Robots with a Reduced Dynamic Window
abstract
In this paper we present an approach to obstacle avoidance and local path planning for polygonal robots. It decomposes the task into a model stage and a planning stage. The model stage accounts for robot shape and dynamics using a reduced dynamic window. The planning stage produces collision-free local paths with a velocity profile. We present an analytical solution to the distance to collision problem for polygonal robots, avoiding thus the use of look-up tables. The approach has been tested in simulation and on two non-holonomic rectangular robots where a cycle time of 10 Hz was reached under full CPU load. During a long-term experiment over 5 km travel distance, the method demonstrated its practicability.
Kai Oliver Arras, Jan Persson, Nicola Tomatis, Roland Siegwart
ICRA1
2001 A Hybrid Approach for Robust and Precise Mobile Robot Navigation with Compact Environment Modeling
abstract
In this paper a new localization approach combining the metric and topological paradigm is presented. The main idea is to connect local metric maps by means of a global topological map. This allows a compact environment model which does not require global metric consistency and permits both precision and robustness. The method uses a 360 degree laser scanner in order to extract lines for the metric localization and doors, discontinuities and hallways for the topological approach. The approach has been widely tested in a 50/spl times/25 m portion of the institute building with the new fully autonomous robot Donald Duck. 25 randomly generated test missions were performed with a success ratio of 96% and a mean error at the goal point of 9 mm for an overall trajectory length of 1.15 km.
Nicola Tomatis, Illah R. Nourbakhsh, Kai Oliver Arras, Roland Siegwart
ICRA3
2000 Multisensor on-the-fly localization using laser and vision
abstract
In this paper a multisensor setup for localization consisting of a 360 degree laser range finder and a monocular vision system is presented. Its practicability under conditions of continuous localization during motion in real-time (referred to as on-the-fly localization) is investigated in large-scale experiments. The features in use are infinite horizontal lines for the laser and vertical lines for the camera providing an extremely compact environment representation. They are extracted using physically well-grounded models for all sensors and passed to the Kalman filter for fusion and position estimation. Very high localization precision is obtained in general. The vision information has been found to further increase this precision, particular in the orientation, already with a moderate number of matched features. The results were obtained with a fully autonomous system where extensive tests with an overall length of more than 1.4 km and 9,500 localization cycles have been conducted. Furthermore, general aspects of multisensor on-the-fly localization are discussed.
Kai Oliver Arras, Nicola Tomatis, Roland Siegwart
IROS1
2000 The need for autonomy and real-time in mobile robotics: a case study of XO/2 and Pygmalion
abstract
Starting from a user point of view the paper discusses the requirements of a development environment (operating system and programming language) for mechatronic systems, especially mobile robots. We argue that user require ments from research, education, ergonomics and applications impose a certain functionality on the embedded operating system and programming language, and that a deadline-driven real-time operating system helps to fulfil these requirements. A case study of the operating system XO/2, its programming language Oberon-2 and the mobile robot Pygmalion is presented. XO/2 explicitly addresses issues like scalabilty, safety and abstraction, previously found to be relevant for many user scenarios.
Roberto Brega, Nicola Tomatis, Kai Oliver Arras
IROS3
2000 The autonomous miniature robot Alice: from prototypes to applications
abstract
We present an overview of the prototype family of Alice miniature mobile robots and the improvements achieved so far. Applications are often the final objective but also an incentive to correct and enhance the robot abilities. The research carried out with Alice and various real-world applications, which exceed the robot's use as a research prototype, is presented. They include local and global localization, map building, control strategies for semi-autonomous operation via Internet and Matlab, its use for robot soccer tournaments and as a research platform for studies of collective behaviors.
Gilles Caprari, Kai Oliver Arras, Roland Siegwart
IROS2
1998 Hybrid, High-Precision Localisation for the Mail Distributing Mobile Robot System MOPS
abstract
Describes the new localisation algorithms under implementation for the mail distributing mobile robot, MOPS, of the Institute of Robotics, Swiss Federal Institute of Technology Zurich. Using geometric primitives as features, we employ consistent probabilistic feature extraction, clustering, matching and estimation of the vehicle position and orientation. The extracted features and their first-order covariance estimates are used, together with a world model, by an extended Kalman filter so as to get an optimal estimate of MOPS' current pose vector and the associated uncertainty. The line extraction consists of an initial segmentation, based on a feature-independent compactness measure in the model space, and a subsequent probabilistic clustering step. This yields a highly accurate and efficient localisation.
Kai Oliver Arras, Sjur J. Vestli
ICRA1