VLDB 2026 Research / reviewers in the wild / expert
Kai Oliver Arras
dblp:a/KOArras · also Kai O. Arras
· DBLP profile ↗
70ranked-venue papers
7as first author
14since 2021 · last 2026
0009-0001-3003-8857ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 68 · 7 first-author · 14 since 2021Systems, architecture and hardware · 50 · 7 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 11 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Context-Aware Generation and Modulation of Expressive Motion Behavior using Multimodal Foundation ModelsabstractExpressive robot motion positively impacts human-robot interaction by improving user engagement, likability, or task performance. We present a novel approach to automatically generate and modulate complex, expressive full-body gesture sequences from a multimodal (text, audio, video) context description. Based on a unified mathematical implementation of the Principles of Animation using Dynamic Movement Primitives, we use multimodal foundation models to generate such sequences along with parametric motion variations that are highly context-aware. This method extends the state of the art in terms of flexibility and generality by being interpretable, composable, and working across different robot morphologies. Moreover, we integrate the system into a continuous control framework and leverage knowledge distillation to learn a much smaller model, significantly improving token efficiency and system latency. Results from a user study with a human-like platform indicate that participants judged our system’s motions to be better aligned with the interaction context than motions produced under a non-modulated or API-level condition. We also demonstrate how the system generates expressive motion for robots with different kinematics, showcasing its versatility. Paper webpage: https://gen-mod-expressive-motion-behavior.github.io/. Till Hielscher, Fabio Scaparro, Kai Oliver Arras |
HRI | 3 |
| 2025 | Multi-Heuristic Robotic Bin Packing of Regular and Irregular ObjectsabstractThe increasing demand in e-commerce, combined with labor shortages and rising wages, is driving the rapid automation of warehouse operations. A critical aspect of this shift is bin packing, where diverse unknown items of varying sizes and shapes must be optimally arranged within a bin or container. Robot bin packing is receiving growing attention and presents unique challenges due to the broad range of objects, packing rules, and task-specific requirements. In response, we propose So-Pack, a generalist packing heuristic for irregularly shaped objects integrated into a flexible, weighted multi-heuristic planning system. The system demonstrates robust performance across general packing scenarios and flexibility to adapt to changing packing rules and specific end-user requirements. Experimental results show that the system outperforms state-of-the-art approaches in key metrics on a new challenging dataset of retail objects in real-world applications. Tim Nickel, Richard Bormann, Kai Oliver Arras |
ICRA | 3 |
| 2025 | Interactive Expressive Motion Generation Using Dynamic Movement PrimitivesabstractOur goal is to enable social robots to interact autonomously with humans in a realistic, engaging, and expressive manner. The 12 Principles of Animation are a well-established framework animators use to create movements that make characters appear convincing, dynamic, and emotionally expressive. This paper proposes a novel approach that leverages Dynamic Movement Primitives (DMPs) to implement key animation principles, providing a learnable, explainable, modulable, online adaptable and composable model for automatic expressive motion generation. DMPs, originally developed for general imitation learning in robotics and grounded in a spring-damper system design, offer mathematical properties that make them particularly suitable for this task. Specifically, they enable modulation of the intensities of individual principles and facilitate the decomposition of complex, expressive motion sequences into learnable and parametrizable primitives. We present the mathematical formulation of the parameterized animation principles and demonstrate the effectiveness of our framework through experiments and application on three robotic platforms with different kinematic configurations, in simulation, on actual robots and in a user study. Our results show that the approach allows for creating diverse and nuanced expressions using a single base model. Till Hielscher, Andreas Bulling, Kai Oliver Arras |
IROS | 3 |
| 2025 | Snuggle-Pack: Speeding Up Multi-Heuristic Packing Planning of Complex ObjectsabstractEfficient object packing is a fundamental challenge in logistics and industrial automation. This work introduces Snuggle-Pack, a novel 3D packing algorithm that integrates Fast Fourier Transform (FFT)-based spatial analysis with a multi-heuristic optimization framework to achieve real-time, high-density packing. Unlike traditional heuristic-based approaches that rely on 2D simplifications, our method operates in a fully 3D volumetric space, ensuring collision-free, stable, and physically feasible placements. At its core, our approach employs a proximity-aware and support-sensitive placement strategy, which encourages objects to fit snugly within their surroundings —hence the name—, optimizing space utilization through ne-grained collision metrics. We evaluate our method on the YCB and IPA-3D1K datasets in both previewed and ad-hoc packing scenarios. Our experiments show that Snuggle-Pack significantly outperforms the state of the art, achieving up to 25% higher packing densities or, alternatively, accelerating computation by up to 10×. Moreover, our framework allows for dynamic adaptation to custom constraints, such as balanced center of mass, weight limitations on fragile items, and safety proximity constraints. These results highlight Snuggle-Pack as an efficient, flexible, and scalable solution for industrial robotic packing tasks. Tim Nickel, Richard Bormann, Kai Oliver Arras |
IROS | 3 |
| 2025 | FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene InteractionabstractThe concept of 3D scene graphs is increasingly recognized as a powerful semantic and hierarchical representation of the environment. Current approaches often address this at a coarse, object-level resolution. In contrast, our goal is to develop a representation that enables robots to directly interact with their environment by identifying both the location of functional interactive elements and how these can be used. To achieve this, we focus on detecting and storing objects at a finer resolution, focusing on affordance-relevant parts. The primary challenge lies in the scarcity of data that extends beyond instance-level detection and the inherent difficulty of capturing detailed object features using robotic sensors. We leverage currently available 3D resources to generate 2D data and train a detector, which is then used to augment the standard 3D scene graph generation pipeline. Through our experiments, we demonstrate that our approach achieves functional element segmentation comparable to state-of-the-art 3D models and that our augmentation enables task-driven affordance grounding with higher accuracy than the current solutions. See our project page at https://fungraph.github.io. Dennis Rotondi, Fabio Scaparro, Hermann Blum, Kai Oliver Arras |
IROS | 4 |
| 2025 | Incremental and Interactive Exploration of Robot Appearance Designs Using GenAIabstractA robot’s physical design impacts user acceptance, engagement, and trust while also influencing social and functional expectations about the robot’s capabilities. Robot design, and industrial design in general, can also be a driver of differentiation in a competitive market. In this paper, we leverage generative AI to incrementally explore design spaces of robot appearances. With the goal of overcoming training data bias of text-to-image models that favor stereotypical morphologies and designs, we propose a set of generation methods that enable a designer to guide the exploration process through text-, style-, and structure-based specifications from user or client feedback. Further, using Low-Rank Adaptation for model fine-tuning, the method allows to define the aesthetic direction by an image collection that conveys a particular style or theme ("mood boards"). The experiments demonstrate that our extensions retain image quality in terms of statistical and structural features and allow for both diversity and specificity in the design process. In a case study, we apply this method in the user-centered design process and discuss its opportunities and limitations. Till Hielscher, Arne Bartenbach, Wolf Leonhardt, Kai Oliver Arras |
RO-MAN | 4 |
| 2024 | STARK: A Unified Framework for Strongly Coupled Simulation of Rigid and Deformable Bodies with Frictional ContactabstractThe use of simulation in robotics is increasingly widespread for the purpose of testing, synthetic data generation and skill learning. A relevant aspect of simulation for a variety of robot applications is physics-based simulation of robot-object interactions. This involves the challenge of accurately modeling and implementing different mechanical systems such as rigid and deformable bodies as well as their interactions via constraints, contact or friction. Most state-of-the-art physics engines commonly used in robotics either cannot couple deformable and rigid bodies in the same framework, lack important systems such as cloth or shells, have stability issues in complex friction-dominated setups or cannot robustly prevent penetrations. In this paper, we propose a framework for strongly coupled simulation of rigid and deformable bodies with focus on usability, stability, robustness and easy access to state-of-the-art deformation and frictional contact models. Our system uses the Finite Element Method (FEM) to model deformable solids, the Incremental Potential Contact (IPC) approach for frictional contact and a robust second order optimizer to ensure stable and penetration-free solutions to tight tolerances. It is a general purpose framework, not tied to a particular use case such as grasping or learning, it is written in C++ and comes with a Python interface. We demonstrate our system’s ability to reproduce complex real-world experiments where a mobile vacuum robot interacts with a towel on different floor types and towel geometries. Our system is able to reproduce 100% of the qualitative outcomes observed in the laboratory environment. The simulation pipeline, named Stark (the German word for strong, as in strong coupling) is made open-source. José Antonio Fernández-Fernández, Ralph Lange, Stefan Laible, Kai Oliver Arras, Jan Bender |
ICRA | 4 |
| 2023 | Semantically Informed MPC for Context-Aware Robot ExplorationabstractWe investigate the task of object goal navigation in unknown environments where a target object is given as a semantic label (e.g. find a couch). This task is challenging as it requires the robot to consider the semantic context in diverse settings (e.g. TVs are often nearby couches). Most of the prior work tackles this problem under the assumption of a discrete action policy whereas we present an approach with continuous control which brings it closer to real world applications. In this paper, we use information-theoretic model predictive control on dense cost maps to bring object goal navigation closer to real robots with kinodynamic constraints. We propose a deep neural network framework to learn cost maps that encode semantic context and guide the robot towards the target object. We also present a novel way of fusing mid-level visual representations in our architecture to provide additional semantic cues for cost map prediction. The experiments show that our method leads to more efficient and accurate goal navigation with higher quality paths than the reported baselines. The results also indicate the importance of mid-level representations for navigation by improving the success rate by 8 percentage points. Yash Goel, Narunas Vaskevicius, Luigi Palmieri, Nived Chebrolu, Kai Oliver Arras, Cyrill Stachniss |
IROS | 5 |
| 2023 | Proactive Model Predictive Control with Multi-Modal Human Motion Prediction in Cluttered Dynamic EnvironmentsabstractFor robots navigating in dynamic environments, exploiting and understanding uncertain human motion prediction is key to generate efficient, safe and legible actions. The robot may perform poorly and cause hindrances if it does not reason over possible, multi-modal future social interactions. With the goal of enhancing autonomous navigation in cluttered environments, we propose a novel formulation for nonlinear model predictive control including multi-modal predictions of human motion. As a result, our approach leads to less conservative, smooth and intuitive human-aware navigation with reduced risk of collisions, and shows a good balance between task efficiency, collision avoidance and human comfort. To show its effectiveness, we compare our approach against the state of the art in crowded simulated environments, and with real-world human motion data from the THOR dataset. This comparison shows that we are able to improve task efficiency, keep a larger distance to humans and significantly reduce the collision time, when navigating in cluttered dynamic environ-ments. Furthermore, the method is shown to work robustly with different state-of-the-art human motion predictors. Lukas Heuer, Luigi Palmieri, Andrey Rudenko, Anna Mannucci, Martin Magnusson 0002, Kai Oliver Arras |
IROS | 6 |
| 2023 | CLiFF-LHMP: Using Spatial Dynamics Patterns for Long- Term Human Motion PredictionabstractHuman motion prediction is important for mobile service robots and intelligent vehicles to operate safely and smoothly around people. The more accurate predictions are, particularly over extended periods of time, the better a system can, e.g., assess collision risks and plan ahead. In this paper, we propose to exploit maps of dynamics (MoDs, a class of general representations of place-dependent spatial motion patterns, learned from prior observations) for long-term human motion prediction (LHMP). We present a new MoD-informed human motion prediction approach, named CLiFF-LHMP, which is data efficient, explainable, and insensitive to errors from an upstream tracking system. Our approach uses CLiFF -map, a specific MoD trained with human motion data recorded in the same environment. We bias a constant velocity prediction with samples from the CLiFF-map to generate multi-modal trajectory predictions. In two public datasets we show that this algorithm outperforms the state of the art for predictions over very extended periods of time, achieving 45 % more accurate prediction performance at 50s compared to the baseline. Andrey Rudenko, Tomasz Kucner, Luigi Palmieri, Kai Oliver Arras, Achim J. Lilienthal, Martin Magnusson 0002 |
IROS | 5 |
| 2023 | Advantages of Multimodal versus Verbal-Only Robot-to-Human Communication with an Anthropomorphic Robotic Mock DriverabstractRobots are increasingly used in shared environments with humans, making effective communication a necessity for successful human-robot interaction. In our work, we study a crucial component: active communication of robot intent. Here, we present an anthropomorphic solution where a humanoid robot communicates the intent of its host robot acting as an “Anthropomorphic Robotic Mock Driver” (ARMoD). We evaluate this approach in two experiments in which participants work alongside a mobile robot on various tasks, while the ARMoD communicates a need for human attention, when required, or gives instructions to collaborate on a joint task. The experiments feature two interaction styles of the ARMoD: a verbal-only mode using only speech and a multimodal mode, additionally including robotic gaze and pointing gestures to support communication and register intent in space. Our results show that the multimodal interaction style, including head movements and eye gaze as well as pointing gestures, leads to more natural fixation behavior. Participants naturally identified and fixated longer on the areas relevant for intent communication, and reacted faster to instructions in collaborative tasks. Our research further indicates that the ARMoD intent communication improves engagement and social interaction with mobile robots in workplace settings. Tim Schreiter, Lucas Morillo-Méndez, Ravi Chadalavada, Andrey Rudenko, Erik Billing, Martin Magnusson 0002, Kai Oliver Arras, Achim J. Lilienthal |
RO-MAN | 7 |
| 2022 | The Atlas Benchmark: an Automated Evaluation Framework for Human Motion PredictionabstractHuman motion trajectory prediction, an essential task for autonomous systems in many domains, has been on the rise in recent years. With a multitude of new methods proposed by different communities, the lack of standardized benchmarks and objective comparisons is increasingly becoming a major limitation to assess progress and guide further research. Existing benchmarks are limited in their scope and flexibility to conduct relevant experiments and to account for contextual cues of agents and environments. In this paper we present Atlas, a benchmark to systematically evaluate human motion trajectory prediction algorithms in a unified framework. Atlas offers data preprocessing functions, hyperparameter optimization, comes with popular datasets and has the flexibility to setup and conduct underexplored yet relevant experiments to analyze a method’s accuracy and robustness. In an example application of Atlas, we compare five popular model- and learning-based predictors and find that, when properly applied, early physics-based approaches are still remarkably competitive. Such results confirm the necessity of benchmarks like Atlas. Andrey Rudenko, Luigi Palmieri, Wanting Huang, Achim J. Lilienthal, Kai Oliver Arras |
RO-MAN | 5 |
| 2022 | How a Social Robot's Vocalization Affects Children's Speech, Learning, and InteractionabstractA wider incorporation of robots into classrooms is hampered by current technological limitations on full autonomy in social robots. Automated speech recognition, for example, a key enabler for vocal communication, is still unable to perform with sufficient accuracy. Past studies have shown that humans adjust their speech patterns to accommodate less skilled interlocutors. If such a response holds in human-robot interactions as well, we may be able to exploit it to lessen the burden on social robots and enable rich, autonomous vocal communication. In this paper we explore whether a robot’s speaking ability could have an impact on children’s speech patterns, learning, and engagement by designing an interaction where a child and a robot collaborate on a Tower of Hanoi puzzle. Sixteen children aged 7-14 completed this collaborative task partnered with a social robot that communicated with either high verbal (full sentences), low verbal (short phrases or single words), or nonverbal (sound-based utterances) vocalization. While we found no significant impact on children’s speech patterns or learning due to the robot’s method of vocalization, children in the non-verbal condition had a significantly lower perception of the robot’s intelligence along with higher rates of providing feedback and more instances of undoing its moves. This suggests that a link may exist between a robot’s perceived speaking ability and children’s confidence in that robot’s overall intelligence and capability in a collaborative task, as well as their empathy towards a peer they perceive as less skilled in the task. Lauren L. Wright, Aditi Kothiyal, Kai Oliver Arras, Barbara Bruno |
RO-MAN | 3 |
| 2021 | Cross-Modal Analysis of Human Detection for Robotics: An Industrial Case StudyabstractAdvances in sensing and learning algorithms have led to increasingly mature solutions for human detection by robots, particularly in selected use-cases such as pedestrian detection for self-driving cars or close-range person detection in consumer settings. Despite this progress, the simple question which sensor-algorithm combination is best suited for a person detection task at handƒ remains hard to answer. In this paper, we tackle this issue by conducting a systematic cross-modal analysis of sensor-algorithm combinations typically used in robotics. We compare the performance of state-of-the-art person detectors for 2D range data, 3D lidar, and RGB-D data as well as selected combinations thereof in a challenging industrial use-case.We further address the related problems of data scarcity in the industrial target domain, and that recent research on human detection in 3D point clouds has mostly focused on autonomous driving scenarios. To leverage these methodological advances for robotics applications, we utilize a simple, yet effective multi-sensor transfer learning strategy by extending a strong image-based RGB-D detector to provide cross-modal supervision for lidar detectors in the form of weak 3D bounding box labels.Our results show a large variance among the different approaches in terms of detection performance, generalization, frame rates and computational requirements. As our use-case contains difficulties representative for a wide range of service robot applications, we believe that these results point to relevant open challenges for further research and provide valuable support to practitioners for the design of their robot system. Timm Linder, Narunas Vaskevicius, Robert Schirmer, Kai Oliver Arras |
IROS | 4 |
| 2020 | Multi-Path Learning for Object Pose Estimation Across DomainsabstractWe introduce a scalable approach for object pose estimation trained on simulated RGB views of multiple 3D models together. We learn an encoding of object views that does not only describe an implicit orientation of all objects seen during training, but can also relate views of untrained objects. Our single-encoder-multi-decoder network is trained using a technique we denote ”multi-path learning”: While the encoder is shared by all objects, each decoder only reconstructs views of a single object. Consequently, views of different instances do not have to be separated in the latent space and can share common features. The resulting encoder generalizes well from synthetic to real data and across various instances, categories, model types and datasets. We systematically investigate the learned encodings, their generalization, and iterative refinement strategies on the ModelNet40 and T-LESS dataset. Despite training jointly on multiple objects, our 6D Object Detection pipeline achieves state-of-the-art results on T-LESS at much lower runtimes than competing approaches. Martin Sundermeyer, Maximilian Durner, En Yen Puang, Zoltan-Csaba Marton, Narunas Vaskevicius, Kai Oliver Arras, Rudolph Triebel |
CVPR | 6 |
| 2020 | Metric-Scale Truncation-Robust Heatmaps for 3D Human Pose EstimationabstractHeatmap representations have formed the basis of 2D human pose estimation systems for many years, but their generalizations for 3D pose have only recently been considered. This includes 2.5D volumetric heatmaps, whose X and Y axes correspond to image space and the Z axis to metric depth around the subject. To obtain metric-scale predictions, these methods must include a separate, explicit post-processing step to resolve scale ambiguity. Further, they cannot encode body joint positions outside of the image boundaries, leading to incomplete pose estimates in case of image truncation. We address these limitations by proposing metric-scale truncation-robust (MeTRo) volumetric heatmaps, whose dimensions are defined in metric 3D space near the subject, instead of being aligned with image space. We train a fully-convolutional network to estimate such heatmaps from monocular RGB in an end-to-end manner. This reinterpretation of the heatmap dimensions allows us to estimate complete metric-scale poses without test-time knowledge of the focal length or person distance and without relying on anthropometric heuristics in post-processing. Furthermore, as the image space is decoupled from the heatmap space, the network can learn to reason about joints beyond the image boundary. Using ResNet-50 without any additional learned layers, we obtain state-of-the-art results on the Human3.6M and MPI-INF-3DHP benchmarks. As our method is simple and fast, it can become a useful component for real-time top-down multi-person pose estimation systems. We make our code publicly available to facilitate further research. István Sárándi, Timm Linder, Kai Oliver Arras, Bastian Leibe |
FG | 3 |
| 2020 | Accurate detection and 3D localization of humans using a novel YOLO-based RGB-D fusion approach and synthetic training dataabstractWhile 2D object detection has made significant progress, robustly localizing objects in 3D space under presence of occlusion is still an unresolved issue. Our focus in this work is on real-time detection of human 3D centroids in RGB-D data. We propose an image-based detection approach which extends the YOLO v3 architecture with a 3D centroid loss and mid-level feature fusion to exploit complementary information from both modalities. We employ a transfer learning scheme which can benefit from existing large-scale 2D object detection datasets, while at the same time learning end-to-end 3D localization from our highly randomized, diverse synthetic RGB-D dataset with precise 3D groundtruth. We further propose a geometrically more accurate depth-aware crop augmentation for training on RGB-D data, which helps to improve 3D localization accuracy. In experiments on our challenging intralogistics dataset, we achieve state-of-the-art performance even when learning 3D localization just from synthetic data. Timm Linder, Kilian Y. Pfeiffer, Narunas Vaskevicius, Robert Schirmer, Kai Oliver Arras |
ICRA | 5 |
| 2020 | An NMPC Approach using Convex Inner Approximations for Online Motion Planning with Guaranteed Collision AvoidanceabstractEven though mobile robots have been around for decades, trajectory optimization and continuous time collision avoidance remain subject of active research. Existing methods trade off between path quality, computational complexity, and kinodynamic feasibility. This work approaches the problem using a nonlinear model predictive control (NMPC) framework, that is based on a novel convex inner approximation of the collision avoidance constraint. The proposed Convex Inner ApprOximation (CIAO) method finds kinodynamically feasible and continuous time collision free trajectories, in few iterations, typically one. For a feasible initialization, the approach is guaranteed to find a feasible solution, i.e. it preserves feasibility. Our experimental evaluation shows that CIAO outperforms state of the art baselines in terms of planning efficiency and path quality. Experiments show that it also efficiently scales to high-dimensional systems. Furthermore real-world experiments demonstrate its capability of unifying trajectory optimization and tracking for safe motion planning in dynamic environments. Tobias Schoels, Luigi Palmieri, Kai Oliver Arras, Moritz Diehl |
ICRA | 3 |
| 2020 | Plug-and-Play SLAM: A Unified SLAM Architecture for Modularity and Ease of UseabstractSimultaneous Localization and Mapping (SLAM) is considered a mature research field with numerous applications and publicly available open-source systems. Despite this maturity, existing SLAM systems often rely on ad-hoc implementations or are tailored to predefined sensor setups. In this work, we tackle these issues, proposing a novel unified SLAM architecture specifically designed to standardize the SLAM problem and to address heterogeneous sensor configurations. Thanks to its modularity and design patterns, the presented framework is easy to extend, maximizes code reuse and improves computational efficiency. We show in our experiments with a variety of typical sensor configurations that these advantages come without compromising state-of-the-art SLAM performance. The result demonstrates the architecture's relevance for facilitating further research in (multi-sensor) SLAM and its transfer into practical applications. Mirco Colosi, Irvin Aloise, Tiziano Guadagnino, Dominik Schlegel, Bartolomeo Della Corte, Kai Oliver Arras, Giorgio Grisetti |
IROS | 6 |
| 2019 | Informed Information Theoretic Model Predictive ControlabstractThe problem of minimizing cost in nonlinear control systems with uncertainties or disturbances remains a major challenge. Model predictive control (MPC), and in particular sampling-based MPC has recently shown great success in complex domains such as aggressive driving with highly nonlinear dynamics. Sampling-based methods rely on a prior distribution to generate samples in the first place. Obviously, the choice of this distribution highly influences efficiency of the controller. Existing approaches such as sampling around the control trajectory of the previous time step perform suboptimally, especially in multi-modal or highly dynamic settings. In this work, we therefore propose to learn models that generate samples in low-cost areas of the state-space, conditioned on the environment and on contextual information of the task to solve. By using generative models as an informed sampling distribution, our approach exploits guidance from the learned models and at the same time maintains robustness properties of the MPC methods. We use Conditional Variational Autoencoders (CVAE) to learn distributions that imitate samples from a training dataset containing optimized controls. An extensive evaluation in the autonomous navigation domain suggests that replacing previous sampling schemes with our learned models considerably improves performance in terms of path quality and planning efficiency. Raphael Kusumoto, Luigi Palmieri, Markus Spies, Akos Csiszar, Kai Oliver Arras |
ICRA | 5 |
| 2019 | Better Lost in Transition Than Lost in Space: SLAM State MachineabstractA Simultaneous Localization and Mapping (SLAM) system is a complex program consisting of several interconnected components with different functionalities such as optimization, tracking or loop detection. Whereas the literature addresses in detail how enhancing the algorithmic aspects of the individual components improves SLAM performance, the modal aspects, such as when to localize, relocalize or close a loop, are usually left aside. In this paper, we address the modal aspects of a SLAM system and show that the design of the modal controller has a strong impact on SLAM performance in particular in terms of robustness against unforeseen events such as sensor failures, perceptual aliasing or kidnapping. We preset a novel taxonomy for the components of a modern SLAM system, investigate their interplay and propose a highly modular architecture of a generic SLAM system using the Unified Modeling LanguageTM(UML) state machine formalism. The result, called SLAM state machine, is compared to the modal controller of several state-of-the-art SLAM systems and evaluated in two experiments. We demonstrate that our state machine handles unforeseen events much more robustly than the state-of-the-art systems. Mirco Colosi, Sebastian Haug, Peter Biber, Kai Oliver Arras, Giorgio Grisetti |
IROS | 4 |
| 2018 | Semantic Labeling of Indoor Environments from 3D RGB MapsabstractWe present an approach to automatically assign semantic labels to rooms reconstructed from 3D RGB maps of apartments. Evidence for the room types is generated using state-of-the-art deep-learning techniques for scene classification and object detection based on automatically generated virtual RGB views, as well as from a geometric analysis of the map's 3D structure. The evidence is merged in a conditional random field, using statistics mined from different datasets of indoor environments. We evaluate our approach qualitatively and quantitatively and compare it to related methods. Manuel Brucker, Maximilian Durner, Rares Ambrus, Zoltan-Csaba Marton, Axel Wendt, Patric Jensfelt, Kai Oliver Arras, Rudolph Triebel |
ICRA | 7 |
| 2018 | Gradient-Informed Path Smoothing for Wheeled Mobile RobotsabstractPlanning smooth trajectories is important for the safe, efficient and comfortable operation of mobile robots, such as wheeled robots moving in crowded environments or cars moving at high speed. Asymptotically optimal sampling-based motion planners can be used to generate such trajectories. However, to achieve the necessary efficiency for the realtime operation of robots, one often uses their initial feasible trajectories or the trajectories of non-optimal motion planners instead, typically after a post-smoothing step. We propose a gradient-informed post-smoothing algorithm, called GRIPS, that deforms given trajectories by locally optimizing the placement of vertices while satisfying the system's kinodynamic constraints. We show experimentally that GRIPS typically produces trajectories of significantly smaller length and higher smoothness than several existing post-smoothing algorithms. Eric Heiden, Luigi Palmieri, Sven Koenig, Kai Oliver Arras, Gaurav S. Sukhatme |
ICRA | 4 |
| 2018 | Joint Long-Term Prediction of Human Motion Using a Planning-Based Social Force ApproachabstractThe ability to perceive and predict future positions of dynamic objects is essential for mobile robots and intelligent vehicles in dynamic environments. In this paper, we present a novel planning-based approach for long-term human motion prediction that accounts for local interactions and can accurately predict joint motion of multiple agents. Long-term predictions are handled using an MDP formulation that computes a set of stochastic motion policies. To obtain distributions over future motion trajectories, we sample the policies with a weighted random walk algorithm in which each person is locally influenced by social forces from other nearby agents. Unlike related work, the algorithm is environment-aware, can account for individual agent velocities, requires no training phase and makes joint predictions for multiple agents. Experiments in simulation and with real data show that our method makes more accurate predictions than two state-of-the-art methods in terms of probabilistic and geometrical performance measures. Andrey Rudenko, Luigi Palmieri, Kai Oliver Arras |
ICRA | 3 |
| 2018 | Human Motion Prediction Under Social Grouping ConstraintsabstractAccurate long-term prediction of human motion in populated spaces is an important but difficult task for mobile robots and intelligent vehicles. What makes this task challenging is that human motion is influenced by a large variety of factors including the person's intention, the presence, attributes, actions, social relations and social norms of other surrounding agents, and the geometry and semantics of the environment. In this paper, we consider the problem of computing human motion predictions that account for such factors. We formulate the task as an MDP planning problem with stochastic policies and propose a weighted random walk algorithm in which each agent is locally influenced by social forces from other nearby agents. The novelty of this paper is that we incorporate social grouping information into the prediction process reflecting the soft formation constraints that groups typically impose to their members' motion. We show that our method makes more accurate predictions than three state-of-the-art methods in terms of probabilistic and geometrical performance metrics. Andrey Rudenko, Luigi Palmieri, Achim J. Lilienthal, Kai Oliver Arras |
IROS | 4 |
| 2017 | Kinodynamic motion planning on Gaussian mixture fieldsabstractWe present a mobile robot motion planning approach under kinodynamic constraints that exploits learned perception priors in the form of continuous Gaussian mixture fields. Our Gaussian mixture fields are statistical multi-modal motion models of discrete objects or continuous media in the environment that encode e.g. the dynamics of air or pedestrian flows. We approach this task using a recently proposed circular linear flow field map based on semi-wrapped GMMs whose mixture components guide sampling and rewiring in an RRT* algorithm using a steer function for non-holonomic mobile robots. In our experiments with three alternative baselines, we show that this combination allows the planner to very efficiently generate high-quality solutions in terms of path smoothness, path length as well as natural yet minimum control effort motions through multi-modal representations of Gaussian mixture fields. Luigi Palmieri, Tomasz Kucner, Martin Magnusson 0002, Achim J. Lilienthal, Kai Oliver Arras |
ICRA | 5 |
| 2016 | On multi-modal people tracking from mobile platforms in very crowded and dynamic environmentsabstractTracking people is a key technology for robots and intelligent systems in human environments. Many person detectors, filtering methods and data association algorithms for people tracking have been proposed in the past 15+ years in both the robotics and computer vision communities, achieving decent tracking performances from static and mobile platforms in real-world scenarios. However, little effort has been made to compare these methods, analyze their performance using different sensory modalities and study their impact on different performance metrics. In this paper, we propose a fully integrated real-time multi-modal laser/RGB-D people tracking framework for moving platforms in environments like a busy airport terminal. We conduct experiments on two challenging new datasets collected from a first-person perspective, one of them containing very dense crowds of people with up to 30 individuals within close range at the same time. We consider four different, recently proposed tracking methods and study their impact on seven different performance metrics, in both single and multi-modal settings. We extensively discuss our findings, which indicate that more complex data association methods may not always be the better choice, and derive possible future research directions. Timm Linder, Stefan Breuers, Bastian Leibe, Kai Oliver Arras |
ICRA | 4 |
| 2016 | Learning socially normative robot navigation behaviors with Bayesian inverse reinforcement learningabstractMobile robots that navigate in populated environments require the capacity to move efficiently, safely and in human-friendly ways. In this paper, we address this task using a learning approach that enables a mobile robot to acquire navigation behaviors from demonstrations of socially normative human behavior. In the past, such approaches have been typically used to learn only simple behaviors under relatively controlled conditions using rigid representations or with methods that scale poorly to large domains. We thus develop a flexible graph-based representation able to capture relevant task structure and extend Bayesian inverse reinforcement learning to use sampled trajectories from this representation. In experiments with a real robot and a large-scale pedestrian simulator, we are able to show that the approach enables a robot to learn complex navigation behaviors of varying degrees of social normativeness using the same set of simple features. Billy Okal, Kai Oliver Arras |
ICRA | 2 |
| 2016 | RRT-based nonholonomic motion planning using any-angle path biasingabstractRRT and RRT* have become popular planning techniques, in particular for high-dimensional systems such as wheeled robots with complex nonholonomic constraints. Their planning times, however, can scale poorly for such robots, which has motivated researchers to study hierarchical techniques that grow the RRT trees in more focused ways. Along this line, we introduce Theta*-RRT that hierarchically combines (discrete) any-angle search with (continuous) RRT motion planning for nonholonomic wheeled robots. Theta*-RRT is a variant of RRT that generates a trajectory by expanding a tree of geodesics toward sampled states whose distribution summarizes geometric information of the any-angle path. We show experimentally, for both a differential drive system and a high-dimensional truck-and-trailer system, that Theta*-RRT finds shorter trajectories significantly faster than four baseline planners (RRT, A*-RRT, RRT*, A*-RRT*) without loss of smoothness, while A*-RRT* and RRT* (and thus also Informed RRT*) fail to generate a first trajectory sufficiently fast in environments with complex nonholonomic constraints. We also prove that Theta*-RRT retains the probabilistic completeness of RRT for all small-time controllable systems that use an analytical steer function. Luigi Palmieri, Sven Koenig, Kai Oliver Arras |
ICRA | 3 |
| 2016 | Practical Bayesian Inverse Reinforcement Learning for Robot Navigation
Billy Okal, Kai Oliver Arras |
ECML/PKDD (3) | 2 |
| 2016 | Errare humanum est: Erroneous robots in human-robot interactionabstractPerfect memory, strong reasoning abilities and flawless performance are typical cognitive traits associated with robots. In contrast, forgetting and erroneous reasoning are typical cognitive patterns of humans. This discrepancy may fundamentally affect the way how robots and humans interact and collaborate together and is today still little explored. In this paper, we investigate the effect of differences between erroneous and perfect robots in a competitive scenario in which humans and robots solve reasoning tasks and memorize numbers. Participants are randomly assigned to one of two groups: in the first group they interact with a perfect, flawless robot, while in the second, they interact with a human-like robot with occasional errors and imperfect memorizing abilities. Participants rate attitude, sympathy, and attributes of the robot in a questionnaire and we measure their task performance. The results show that the erroneous robot triggered more positive emotions but lead to a lower human performance than the perfect one. Effects of both conditions on the group of students with and without technical background are reported. Marco Ragni, Andrey Rudenko, Barbara Kuhnert, Kai Oliver Arras |
RO-MAN | 4 |
| 2015 | Real-time full-body human gender recognition in (RGB)-D dataabstractUnderstanding social context is an important skill for robots that share a space with humans. In this paper, we address the problem of recognizing gender, a key piece of information when interacting with people and understanding human social relations and rules. Unlike previous work which typically considered faces or frontal body views in image data, we address the problem of recognizing gender in RGB-D data from side and back views as well. We present a large, gender-balanced, annotated, multi-perspective RGB-D dataset with full-body views of over a hundred different persons captured with both the Kinect v1 and Kinect v2 sensor. We then learn and compare several classifiers on the Kinect v2 data using a HOG baseline, two state-of-the-art deep-learning methods, and a recent tessellation-based learning approach. Originally developed for person detection in 3D data, the latter is able to learn the best selection, location and scale of a set of simple point cloud features. We show that for gender recognition, it outperforms the other approaches for both standing and walking people while being very efficient to compute with classification rates up to 150 Hz. Timm Linder, Sven Wehner, Kai Oliver Arras |
ICRA | 3 |
| 2015 | Distance metric learning for RRT-based motion planning with constant-time inferenceabstractThe distance metric is a key component in RRT-based motion planning that deeply affects coverage of the state space, path quality and planning time. With the goal to speed up planning time, we introduce a learning approach to approximate the distance metric for RRT-based planners. By exploiting a novel steer function which solves the two-point boundary value problem for wheeled mobile robots, we train a simple nonlinear parametric model with constant-time inference that is shown to predict distances accurately in terms of regression and ranking performance. In an extensive analysis we compare our approach to an Euclidean distance baseline, consider four alternative regression models and study the impact of domain-specific feature expansion. The learning approach is shown to be faster in planning time by several factors at negligible loss of path quality. Luigi Palmieri, Kai Oliver Arras |
ICRA | 2 |
| 2015 | Real-time full-body human attribute classification in RGB-D using a tessellation boosting approachabstractRobots that cooperate and interact with humans require the capacity to detect and track people, analyze their behavior and understand human social relations and rules. A key piece of information for such tasks are human attributes like gender, age, hair or clothing. In this paper, we address the problem of recognizing such attributes in RGB-D data from varying full-body views. To this end, we extend a recent tessellation boosting approach which learns the best selection, location and scale of a set of simple RGB-D features. The approach outperforms the original approach and a HOG baseline for five human attributes including gender, has long hair, has long trousers, has long sleeves and has jacket. Experiments on a multi-perspective RGB-D dataset with full-body views of over a hundred different persons show that the method is able to robustly recognize multiple attributes across different view directions and distances to the sensor with accuracies up to 90%. Our methods runs in real-time, achieving a classification rate of around 300 Hz for a single attribute. Timm Linder, Kai Oliver Arras |
IROS | 2 |
| 2014 | Schedule-Based Robotic Search for Multiple Residents in a Retirement Home EnvironmentabstractIn this paper we address the planning problem of a robot searching for multiple residents in a retirement home in order to remind them of an upcoming multi-person recreational activity before a given deadline. We introduce a novel Multi-User Schedule Based (M-USB) Search approach which generates a high-level-plan to maximize the number of residents that are found within the given time frame. From the schedules of the residents, the layout of the retirement home environment as well as direct observations by the robot, we obtain spatio-temporal likelihood functions for the individual residents. The main contribution of our work is the development of a novel approach to compute a reward to find a search plan for the robot using: 1) the likelihood functions, 2) the availabilities of the residents, and 3) the order in which the residents should be found. Simulations were conducted on a floor of a real retirement home to compare our proposed M-USB Search approach to a Weighted Informed Walk and a Random Walk. Our results show that the proposed M-USB Search finds residents in a shorter amount of time by visiting fewer rooms when compared to the other approaches. Markus Sebastian Schwenk, Tiago Stegun Vaquero, Goldie Nejat, Kai Oliver Arras |
AAAI | 4 |
| 2014 | Multi-model hypothesis tracking of groups of people in RGB-D data
Timm Linder, Kai Oliver Arras |
FUSION | 2 |
| 2014 | Robotic tele-presence with DARYL in the wildabstractThis paper describes the results of a qualitative analysis of questionnaire data collected during a public exhibition of our robotic tele-presence system. In Summer 2013 the mildly humanized robot DARYL could be tried out by the general public during our University's science fair in the city center. People were given the chance to communicate through the robot with their peers and to perceive the world through the "eyes" and "ears" of the robot by means of a head-mounted display with attached headphones. An operator's voice was instantaneously transmitted to the robot's location and his or her head movements were tracked to enable direct, intuitive control of the robot's head movements. Twenty-seven people were interviewed in a structured way about their impressions and opinions after having either operated or interacted with the tele-operated robot. A careful analysis of the acquired data reveals a rather positive evaluation of the tele-presence system and interesting opinions about suitable application areas. These findings may guide designers of robotic tele-presence systems, a research area of increasing popularity. Christian Becker-Asano, Kai Oliver Arras, Bernhard Nebel |
HAI | 2 |
| 2014 | A novel RRT extend function for efficient and smooth mobile robot motion planningabstractIn this paper we introduce a novel RRT extend function for wheeled mobile robots. The approach computes closed-loop forward simulations based on the kinematic model of the robot and enables the planner to efficiently generate smooth and feasible paths that connect any pairs of states. We extend the control law of an existing discontinuous state feedback controller to make it usable as an RRT extend function and prove that all relevant stability properties are retained. We study the properties of the new approach as extender for RRT and RRT* and compare it systematically to a spline-based approach and a large and small set of motion primitives. The results show that our approach generally produces smoother paths to the goal in less time with smaller trees. For RRT*, the approach produces also the shortest paths and achieves the lowest cost solutions when given more planning time. Luigi Palmieri, Kai Oliver Arras |
IROS | 2 |
| 2014 | Inverse Reinforcement Learning algorithms and features for robot navigation in crowds: An experimental comparisonabstractFor mobile robots which operate in human populated environments, modeling social interactions is key to understand and reproduce people's behavior. A promising approach to this end is Inverse Reinforcement Learning (IRL) as it allows to model the factors that motivate people's actions instead of the actions themselves. A crucial design choice in IRL is the selection of features that encode the agent's context. In related work, features are typically chosen ad hoc without systematic evaluation of the alternatives and their actual impact on the robot's task. In this paper, we introduce a new software framework to systematically investigate the effect features and learning algorithms used in the literature. We also present results for the task of socially compliant robot navigation in crowds, evaluating two different IRL approaches and several feature sets in large-scale simulations. The results are benchmarked according to a proposed set of objective and subjective performance metrics. Dizan Vasquez, Billy Okal, Kai Oliver Arras |
IROS | 3 |
| 2014 | R2-D2 Reloaded: A flexible sound synthesis system for sonic human-robot interaction designabstractA key skill for social robots is the ability to communicate their inner state to humans. In this paper, we explore abstracted robot-specific ways of interaction as an alternative to human-like or animal-like social cues. In particular, we present a sound system as a novel modality that extends a robot's ability for non-verbal communication. Unlike prior work which used pre-recorded audio samples to this end, we propose a flexible architecture with a generalized sound synthesizer that uses the principle of modulation to shape the sound in real-time by external and internal stimuli from the robot or the interaction. This allows for almost unlimited possibilities in the design of an expressive auditory social cue for human-robot interaction. We instantiate the architecture and report on example design choices for the sound synthesis principle, the real-time synthesizer, the sound modulation routings, and a sound sequence composer. We then demonstrate the system's ability for affect communication of primary and secondary emotions on a social robot. Markus Sebastian Schwenk, Kai Oliver Arras |
RO-MAN | 2 |
| 2013 | Robot embodiment, operator modality, and social interaction in tele-existence: a project outline
Christian Becker-Asano, Severin Gustorff, Kai Oliver Arras, Kohei Ogawa, Shuichi Nishio, Hiroshi Ishiguro, Bernhard Nebel |
HRI | 3 |
| 2012 | Leveraging RGB-D Data: Adaptive fusion and domain adaptation for object detectionabstractVision and range sensing belong to the richest sensory modalities for perception in robotics and related fields. This paper addresses the problem of how to best combine image and range data for the task of object detection. In particular, we propose a novel adaptive fusion approach, hierarchical Gaussian Process mixtures of experts, able to account for missing information and cross-cue data consistency. The hierarchy is a two-tier architecture that for each modality, each frame and each detection computes a weight function using Gaussian Processes that reflects the confidence of the respective information. We further propose a method called cross-cue domain adaptation that makes use of large image data sets to improve the depth-based object detector for which only few training samples exist. In the experiments that include a comparison with alternative sensor fusion schemes, we demonstrate the viability of the proposed methods and achieve significant improvements in classification accuracy. Luciano Spinello, Kai Oliver Arras |
ICRA | 2 |
| 2012 | Socially-aware robot navigation: A learning approachabstractThe ability to act in a socially-aware way is a key skill for robots that share a space with humans. In this paper we address the problem of socially-aware navigation among people that meets objective criteria such as travel time or path length as well as subjective criteria such as social comfort. Opposed to model-based approaches typically taken in related work, we pose the problem as an unsupervised learning problem. We learn a set of dynamic motion prototypes from observations of relative motion behavior of humans found in publicly available surveillance data sets. The learned motion prototypes are then used to compute dynamic cost maps for path planning using an any-angle A* algorithm. In the evaluation we demonstrate that the learned behaviors are better in reproducing human relative motion in both criteria than a Proxemics-based baseline method. Matthias Luber, Luciano Spinello, Jens Silva, Kai Oliver Arras |
IROS | 4 |
| 2012 | Robot-specific social cues in emotional body languageabstractHumans use very sophisticated ways of bodily emotion expression combining facial expressions, sound, gestures and full body posture. Like others, we want to apply these aspects of human communication to ease the interaction between robots and users. In doing so we believe there is a need to consider what abstraction of human social communicative behaviors is appropriate for robots. The study reported in this paper is a pilot study to not offer simulated emotion but to offer an abstracted robot version of emotion expressions and an evaluation to what extent users interpret these robot expressions as the intended emotional states. To this end, we present the mobile, mildly humanized robot Daryl, for which we created six motion sequences that combine human-like, animal-like, and robot-specific social cues. The results of a user study (N=29) show that despite the absence of facial expressions and articulated extremities, subjects' interpretation of Daryl's emotional states were congruent with the abstracted emotion display. These results demonstrate that abstract displays of emotion that combine human-like, animal-like, and robot-specific modalities could in fact be an alternative to complex facial expressions and will feed into ongoing work identifying robot-specific social cues. Stephanie Embgen, Matthias Luber, Christian Becker-Asano, Marco Ragni, Vanessa Evers, Kai Oliver Arras |
RO-MAN | 6 |
| 2012 | Audio-based human activity recognition using Non-Markovian Ensemble VotingabstractHuman activity recognition is a key component for socially enabled robots to effectively and naturally interact with humans. In this paper we exploit the fact that many human activities produce characteristic sounds from which a robot can infer the corresponding actions. We propose a novel recognition approach called Non-Markovian Ensemble Voting (NEV) able to classify multiple human activities in an online fashion without the need for silence detection or audio stream segmentation. Moreover, the method can deal with activities that are extended over undefined periods in time. In a series of experiments in real reverberant environments, we are able to robustly recognize 22 different sounds that correspond to a number of human activities in a bathroom and kitchen context. Our method outperforms several established classification techniques. Johannes A. Stork, Luciano Spinello, Jens Silva, Kai Oliver Arras |
RO-MAN | 4 |
| 2011 | Better models for people trackingabstractPeople tracking is a key component for robots operating in populated environments. Previous works have employed different filtering and data association techniques for this purpose that typically rely on a set of generic assumptions on target behavior and detector characteristics. In this paper, we focus on these assumptions rather than the tracking approach itself and show that with informed models, people tracking can be made substantially more accurate without compromising efficiency. Concretely, we present better, human-specific models for the occurrence of new tracks, false alarms, track occlusions, and track deletions. In the experiments with a large-scale outdoor data set collected with a laser range finder, the models and combinations thereof are experimentally compared using a multi-hypothesis baseline tracker and the CLEAR MOT metrics. The results show how some models selectively improve tracking performance at the expense of other measures. The final combination is then able to resolve the trade-offs, leading to a reduction of data association errors by more than a factor of two at the same cost. Matthias Luber, Gian Diego Tipaldi, Kai Oliver Arras |
ICRA | 3 |
| 2011 | Tracking people in 3D using a bottom-up top-down detectorabstractPeople detection and tracking is a key component for robots and autonomous vehicles in human environments. While prior work mainly employed image or 2D range data for this task, in this paper, we address the problem using 3D range data. In our approach, a top-down classifier selects hypotheses from a bottom-up detector, both based on sets of boosted features. The bottom-up detector learns a layered person model from a bank of specialized classifiers for different height levels of people that collectively vote into a continuous space. Modes in this space represent detection candidates that each postulate a segmentation hypothesis of the data. In the top-down step, the candidates are classified using features that are computed in voxels of a boosted volume tessellation. We learn the optimal volume tessellation as it enables the method to stably deal with sparsely sampled and articulated objects. We then combine the detector with tracking in 3D for which we take a multi-target multi-hypothesis tracking approach. The method neither needs a ground plane assumption nor relies on background learning. The results from experiments in populated urban environments demonstrate 3D tracking and highly robust people detection up to 20 m with equal error rates of at least 93%. Luciano Spinello, Matthias Luber, Kai Oliver Arras |
ICRA | 3 |
| 2011 | I want my coffee hot! Learning to find people under spatio-temporal constraintsabstractIn this paper we present a probabilistic model for spatio-temporal patterns of human activities that enable robots to blend themselves into the workflows and daily routines of people. The model, called spatial affordance map, is a non-homogeneous spatial Poisson process that relates space, time and occurrence probability of activity events. We describe how learning and inference is made and present a novel planning algorithm that produces paths which maximize the probability to encounter a person. We show that the problem is a special class of the orienteering problem that can be solved as a finite horizon Markov decision process. We develop a simulator of populated office environments to validate the model and the planning algorithm. The simulated agents follow activity patterns learned by administering a questionnaire to 27 colleagues over two weeks. The experiments shows that the model is statistically valid with respect to both the Anderson-Darling test and the expected waiting time estimation. They further show that the proposed algorithm is able to find optimal paths. Gian Diego Tipaldi, Kai Oliver Arras |
ICRA | 2 |
| 2011 | People tracking in RGB-D Data with on-line boosted target modelsabstractPeople tracking is a key component for robots that are deployed in populated environments. Previous works have used cameras and 2D and 3D range finders for this task. In this paper, we present a 3D people detection and tracking approach using RGB-D data. We combine a novel multi-cue person detector for RGB-D data with an on-line detector that learns individual target models. The two detectors are integrated into a decisional framework with a multi-hypothesis tracker that controls on-line learning through a track interpretation feedback. For on-line learning, we take a boosting approach using three types of RGB-D features and a confidence maximization search in 3D space. The approach is general in that it neither relies on background learning nor a ground plane assumption. For the evaluation, we collect data in a populated indoor environment using a setup of three Microsoft Kinect sensors with a joint field of view. The results demonstrate reliable 3D tracking of people in RGB-D data and show how the framework is able to avoid drift of the on-line detector and increase the overall tracking performance. Matthias Luber, Luciano Spinello, Kai Oliver Arras |
IROS | 3 |
| 2011 | People detection in RGB-D DataabstractPeople detection is a key issue for robots and intelligent systems sharing a space with people. Previous works have used cameras and 2D or 3D range finders for this task. In this paper, we present a novel people detection approach for RGB-D data. We take inspiration from the Histogram of Oriented Gradients (HOG) detector to design a robust method to detect people in dense depth data, called Histogram of Oriented Depths (HOD). HOD locally encodes the direction of depth changes and relies on an depth-informed scale-space search that leads to a 3-fold acceleration of the detection process. We then propose Combo-HOD, a RGB-D detector that probabilistically combines HOD and HOG. The experiments include a comprehensive comparison with several alternative detection approaches including visual HOG, several variants of HOD, a geometric person detector for 3D point clouds, and an Haar-based AdaBoost detector. With an equal error rate of 85% in a range up to 8m, the results demonstrate the robustness of HOD and Combo-HOD on a real-world data set collected with a Kinect sensor in a populated indoor environment. Luciano Spinello, Kai Oliver Arras |
IROS | 2 |
| 2011 | Please do not disturb! Minimum interference coverage for social robotsabstractIn this paper we address the problem of human-aware coverage planning. We first present an approach to learn and model human activity events in a probabilistic spatio-temporal map using spatial Poisson processes. We then propose a coverage planner for paths that minimize the interference probability with people. To this end, we pose the coverage problem as an asymmetric traveling salesman problem with time-dependent costs (ATDTSP) derived from the information in the map. The approach enables a noisy robotic vacuum in a home scenario, for instance, to learn to avoid busy places at certain times of the day such as the kitchen at lunch time. We evaluate the planner using a simulator of people in a home environment to generate typical weekday activity patterns. In the experiments with a regular TSP planner and two modified TSP heuristics, the proposed coverage planner significantly reduces interference with people in terms of number of disturbed persons and overall disturbance time. Gian Diego Tipaldi, Kai Oliver Arras |
IROS | 2 |
| 2010 | A Layered Approach to People Detection in 3D Range DataabstractPeople tracking is a key technology for autonomous systems, intelligent cars and social robots operating in populated environments. What makes the task difficult is that the appearance of humans in range data can change drastically as a function of body pose, distance to the sensor, self-occlusion and occlusion by other objects. In this paper we propose a novel approach to pedestrian detection in 3D range data based on supervised learning techniques to create a bank of classifiers for different height levels of the human body. In particular, our approach applies AdaBoost to train a strong classifier from geometrical and statistical features of groups of neighboring points at the same height. In a second step, the AdaBoost classifiers mutually enforce their evidence across different heights by voting into a continuous space. Pedestrians are finally found efficiently by mean-shift search for local maxima in the voting space. Experimental results carried out with 3D laser range data illustrate the robustness and efficiency of our approach even in cluttered urban environments. The learned people detector reaches a classification rate up to 96% from a single 3D scan. Luciano Spinello, Kai Oliver Arras, Rudolph Triebel, Roland Siegwart |
AAAI | 2 |
| 2010 | Exploiting Repetitive Object Patterns for Model Compression and Completion
Luciano Spinello, Rudolph Triebel, Dizan Vasquez, Kai Oliver Arras, Roland Siegwart |
ECCV (5) | 4 |
| 2010 | People tracking with human motion predictions from social forcesabstractFor many tasks in populated environments, robots need to keep track of current and future motion states of people. Most approaches to people tracking make weak assumptions on human motion such as constant velocity or acceleration. But even over a short period, human behavior is more complex and influenced by factors such as the intended goal, other people, objects in the environment, and social rules. This motivates the use of more sophisticated motion models for people tracking especially since humans frequently undergo lengthy occlusion events. In this paper, we consider computational models developed in the cognitive and social science communities that describe individual and collective pedestrian dynamics for tasks such as crowd behavior analysis. In particular, we integrate a model based on a social force concept into a multi-hypothesis target tracker. We show how the refined motion predictions translate into more informed probability distributions over hypotheses and finally into a more robust tracking behavior and better occlusion handling. In experiments in indoor and outdoor environments with data from a laser range finder, the social force model leads to more accurate tracking with up to two times fewer data association errors. Matthias Luber, Johannes A. Stork, Gian Diego Tipaldi, Kai Oliver Arras |
ICRA | 4 |
| 2010 | FLIRT - Interest regions for 2D range dataabstractLocal image features are used for a wide range of applications in computer vision and range imaging. While there is a great variety of detector-descriptor combinations for image data and 3D point clouds, there is no general method readily available for 2D range data. For this reason, the paper first proposes a set of benchmark experiments on detector repeatability and descriptor matching performance using known indoor and outdoor data sets for robot navigation. Secondly, the paper introduces FLIRT that stands for Fast Laser Interest Region Transform, a multi-scale interest region operator for 2D range data. FLIRT combines the best detector with the best descriptor, experimentally found in a comprehensive analysis of alternative detector and descriptor approaches. The analysis yields repeatability and matching performance results similar to the values found for features in the computer vision literature, encouraging a wide range of applications of FLIRT on 2D range data. We finally show how FLIRT can be used in conjunction with RANSAC to address the loop closing/global localization problem in SLAM in indoor as well as outdoor environments. The results demonstrate that FLIRT features have a great potential for robot navigation in terms of precision-recall performance, efficiency and generality. Gian Diego Tipaldi, Kai Oliver Arras |
ICRA | 2 |
| 2009 | Tracking groups of people with a multi-model hypothesis trackerabstractPeople in densely populated environments typically form groups that split and merge. In this paper we track groups of people so as to reflect this formation process and gain efficiency in situations where maintaining the state of individual people would be intractable. We pose the group tracking problem as a recursive multi-hypothesis model selection problem in which we hypothesize over both, the partitioning of tracks into groups (models) and the association of observations to tracks (assignments). Model hypotheses that include split, merge, and continuation events are first generated in a data-driven manner and then validated by means of the assignment probabilities conditioned on the respective model. Observations are found by clustering points from a laser range finder given a background model and associated to existing group tracks using the minimum average Hausdorff distance. Experiments with a stationary and a moving platform show that, in populated environments, tracking groups is clearly more efficient than tracking people separately. Our system runs in real-time on a typical desktop computer. Boris Lau, Kai Oliver Arras, Wolfram Burgard |
ICRA | 2 |
| 2009 | Place-Dependent People Tracking
Matthias Luber, Gian Diego Tipaldi, Kai Oliver Arras |
ISRR | 3 |
| 2008 | Efficient people tracking in laser range data using a multi-hypothesis leg-tracker with adaptive occlusion probabilitiesabstractWe present an approach to laser-based people tracking using a multi-hypothesis tracker that detects and tracks legs separately with Kalman filters, constant velocity motion models, and a multi-hypothesis data association strategy. People are defined as high-level tracks consisting of two legs that are found with little model knowledge. We extend the data association so that it explicitly handles track occlusions in addition to detections and deletions. Additionally, we adapt the corresponding probabilities in a situation-dependent fashion so as to reflect the fact that legs frequently occlude each other. Experimental results carried out with a mobile robot illustrate that our approach can robustly and efficiently track multiple people even in situations of high levels of occlusion. Kai Oliver Arras, Slawomir Grzonka, Matthias Luber, Wolfram Burgard |
ICRA | 1 |
| 2007 | Using Boosted Features for the Detection of People in 2D Range DataabstractThis paper addresses the problem of detecting people in two dimensional range scans. Previous approaches have mostly used pre-defined features for the detection and tracking of people. We propose an approach that utilizes a supervised learning technique to create a classifier that facilitates the detection of people. In particular, our approach applies AdaBoost to train a strong classifier from simple features of groups of neighboring beams corresponding to legs in range data. Experimental results carried out with laser range data illustrate the robustness of our approach even in cluttered office environments Kai Oliver Arras, Óscar Martínez Mozos, Wolfram Burgard |
ICRA | 1 |
| 2004 | 2D Mapping of Cluttered Indoor Environments by Means of 3D PerceptionabstractThis paper presents a combination of a 3D laser sensor and a line-base SLAM algorithm which together produce 2D line maps of highly cluttered indoor environments. The key of the described method is the replacement of commonly used 2D laser range sensors by 3D perception. A straightforward algorithm extracts a virtual 2D scan that also contains partially occluded walls. These virtual scans are used as input for SLAM using line segments as features. The paper presents the used algorithms and experimental results that were made in a former industrial bakery. The focus lies on scenes that are known to be problematic for pure 2D systems. The results demonstrate that mapping indoor environments can be made robust with respect to both, poor odometry and clutter. Oliver Wulf, Kai Oliver Arras, Henrik I. Christensen, Bernardo Wagner |
ICRA | 2 |
| 2003 | A navigation framework for multiple mobile robots and its application at the Expo.02 exhibitionabstractThis paper presents a navigation framework which enables multiple mobile robots to attain individual goals, coordinate their actions and work safely and reliably in a highly dynamic environment. We give an overview of the framework architecture, its layering and the subsystems reactive obstacle avoidance, local path planning, global path planning, multi-robot planning and localization. The latter receives particular attention as the localization problem is a key issue for navigation in unmodified and difficult environments. The framework permits a lightweight implementation on a fully autonomous robot. This is the result of a design effort striving for compact representations and computational efficiency. The experimental testbed was the "Robotics" pavilion at the Swiss National Exhibition Expo.02 where ten fully autonomous robots were interacting with more than half a million visitors during a five-month period on 3316 km. Kai Oliver Arras, Roland Philippsen, Nicola Tomatis, Marc De Battista, Martin Schilt, Roland Siegwart |
ICRA | 1 |
| 2003 | Designing a secure and robust mobile interacting robot for the long termabstractThis paper presents the genesis of RoboX. This tour guide robot has been built from the scratch based on the experience of the Autonomous Systems Lab. The production of 11 of those machines has been realized by a spin-off of the lab: BlueBotics SA. The goal was to maximize the autonomy and interactivity of the mobile platform while ensuring high robustness, security and performance. The result is an interactive moving machine which can operate in human environments and interacts by seeing humans, talking to and looking at them, showing icons and asking them to answer its questions. The complete design of mechanics, electronics and software is presented in the first part. Then, as extraordinary test bed, the Robotics exhibition at Expo.02 (Swiss National Exhibition) permits to establish meaningful statistics over 5 months (from May 15 to October 20, 2002) with up to 11 robots operating at the same time. Nicola Tomatis, Gregoire Terrien, Ralph Piguet, Daniel Burnier, Samir Bouabdallah, Kai Oliver Arras, Roland Siegwart |
ICRA | 6 |
| 2003 | Multi-resolution SLAM for Real World Navigation
Agostino Martinelli, Adriana Tapus, Kai Oliver Arras, Roland Siegwart |
ISRR | 3 |
| 2002 | Feature-Based Multi-Hypothesis Localization and Tracking for Mobile Robots using Geometric ConstraintsabstractIn this paper we present a new probabilistic feature-based approach to multi-hypothesis global localization and pose tracking. Hypotheses are generated using a constraint-based search in the interpretation tree of possible local-to-global pairings. This results in a set of robot location hypotheses of unbounded accuracy. For tracking, the same constraint-based technique is used. It performs track splitting as soon as location ambiguities arise from uncertainties and sensing. This yields a very robust localization technique which can deal with significant errors from odometry, collisions and kidnapping. Simulation experiments and first tests with a real robot demonstrate these properties at very low computational cost. The presented approach is theoretically sound which makes that the only parameter is the significance level on which all statistical decisions are taken. Kai Oliver Arras, José A. Castellanos 0001, Roland Siegwart |
ICRA | 1 |
| 2002 | Real-Time Obstacle Avoidance for Polygonal Robots with a Reduced Dynamic WindowabstractIn this paper we present an approach to obstacle avoidance and local path planning for polygonal robots. It decomposes the task into a model stage and a planning stage. The model stage accounts for robot shape and dynamics using a reduced dynamic window. The planning stage produces collision-free local paths with a velocity profile. We present an analytical solution to the distance to collision problem for polygonal robots, avoiding thus the use of look-up tables. The approach has been tested in simulation and on two non-holonomic rectangular robots where a cycle time of 10 Hz was reached under full CPU load. During a long-term experiment over 5 km travel distance, the method demonstrated its practicability. Kai Oliver Arras, Jan Persson, Nicola Tomatis, Roland Siegwart |
ICRA | 1 |
| 2001 | A Hybrid Approach for Robust and Precise Mobile Robot Navigation with Compact Environment ModelingabstractIn this paper a new localization approach combining the metric and topological paradigm is presented. The main idea is to connect local metric maps by means of a global topological map. This allows a compact environment model which does not require global metric consistency and permits both precision and robustness. The method uses a 360 degree laser scanner in order to extract lines for the metric localization and doors, discontinuities and hallways for the topological approach. The approach has been widely tested in a 50/spl times/25 m portion of the institute building with the new fully autonomous robot Donald Duck. 25 randomly generated test missions were performed with a success ratio of 96% and a mean error at the goal point of 9 mm for an overall trajectory length of 1.15 km. Nicola Tomatis, Illah R. Nourbakhsh, Kai Oliver Arras, Roland Siegwart |
ICRA | 3 |
| 2000 | Multisensor on-the-fly localization using laser and visionabstractIn this paper a multisensor setup for localization consisting of a 360 degree laser range finder and a monocular vision system is presented. Its practicability under conditions of continuous localization during motion in real-time (referred to as on-the-fly localization) is investigated in large-scale experiments. The features in use are infinite horizontal lines for the laser and vertical lines for the camera providing an extremely compact environment representation. They are extracted using physically well-grounded models for all sensors and passed to the Kalman filter for fusion and position estimation. Very high localization precision is obtained in general. The vision information has been found to further increase this precision, particular in the orientation, already with a moderate number of matched features. The results were obtained with a fully autonomous system where extensive tests with an overall length of more than 1.4 km and 9,500 localization cycles have been conducted. Furthermore, general aspects of multisensor on-the-fly localization are discussed. Kai Oliver Arras, Nicola Tomatis, Roland Siegwart |
IROS | 1 |
| 2000 | The need for autonomy and real-time in mobile robotics: a case study of XO/2 and PygmalionabstractStarting from a user point of view the paper discusses the requirements of a development environment (operating system and programming language) for mechatronic systems, especially mobile robots. We argue that user require ments from research, education, ergonomics and applications impose a certain functionality on the embedded operating system and programming language, and that a deadline-driven real-time operating system helps to fulfil these requirements. A case study of the operating system XO/2, its programming language Oberon-2 and the mobile robot Pygmalion is presented. XO/2 explicitly addresses issues like scalabilty, safety and abstraction, previously found to be relevant for many user scenarios. Roberto Brega, Nicola Tomatis, Kai Oliver Arras |
IROS | 3 |
| 2000 | The autonomous miniature robot Alice: from prototypes to applicationsabstractWe present an overview of the prototype family of Alice miniature mobile robots and the improvements achieved so far. Applications are often the final objective but also an incentive to correct and enhance the robot abilities. The research carried out with Alice and various real-world applications, which exceed the robot's use as a research prototype, is presented. They include local and global localization, map building, control strategies for semi-autonomous operation via Internet and Matlab, its use for robot soccer tournaments and as a research platform for studies of collective behaviors. Gilles Caprari, Kai Oliver Arras, Roland Siegwart |
IROS | 2 |
| 1998 | Hybrid, High-Precision Localisation for the Mail Distributing Mobile Robot System MOPSabstractDescribes the new localisation algorithms under implementation for the mail distributing mobile robot, MOPS, of the Institute of Robotics, Swiss Federal Institute of Technology Zurich. Using geometric primitives as features, we employ consistent probabilistic feature extraction, clustering, matching and estimation of the vehicle position and orientation. The extracted features and their first-order covariance estimates are used, together with a world model, by an extended Kalman filter so as to get an optimal estimate of MOPS' current pose vector and the associated uncertainty. The line extraction consists of an initial segmentation, based on a feature-independent compactness measure in the model space, and a subsequent probabilistic clustering step. This yields a highly accurate and efficient localisation. Kai Oliver Arras, Sjur J. Vestli |
ICRA | 1 |