Martin Magnusson 0002

dblp:m/MartinMagnusson2 · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0001-8658-2985ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 5 first-author · 17 since 2021Systems, architecture and hardware · 31 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Evaluating Efficiency and Engagement in Scripted and LLM-Enhanced Human-Robot Interactions
abstract
To achieve natural and intuitive interaction with people, HRI frameworks combine a wide array of methods for human perception, intention communication, human-aware navigation and collaborative action. In practice, when encountering unpredictable behavior of people or unexpected states of the environment, these frameworks may lack the ability to dynamically recognize such states, adapt and recover to resume the interaction. Large Language Models (LLMs), owing to their advanced reasoning capabilities and context retention, present a promising solution for enhancing robot adaptability. This potential, however, may not directly translate to improved interaction metrics. This paper considers a representative interaction with an industrial robot involving approach, instruction, and object manipulation, implemented in two conditions: (1) fully scripted and (2) including LLM-enhanced responses. We use gaze tracking and questionnaires to measure the participants' task efficiency, engagement, and robot perception. The results indicate higher SUbjective ratings for the LLM condition, but objective metrics show that the scripted condition performs comparably, particularly in efficiency and focus during simple tasks. We also note that the scripted condition may have an edge over LLM-enhanced responses in terms of response latency and energy consumption, especially for trivial and repetitive interactions.
Tim Schreiter, Jens Rüppel, Rishi Hazra, Andrey Rudenko, Martin Magnusson 0002, Achim J. Lilienthal
HRI5
2025 KEA: Keeping Exploration Alive by Proactively Coordinating Exploration Strategies
abstract
Soft Actor-Critic (SAC) has achieved notable success in continuous control tasks but struggles in sparse reward settings, where infrequent rewards make efficient exploration challenging. While novelty-based exploration methods address this issue by encouraging the agent to explore novel states, they are not trivial to apply to SAC. In particular, managing the interaction between novelty-based exploration and SAC’s stochastic policy can lead to inefficient exploration and redundant sample collection. In this paper, we propose KEA (Keeping Exploration Alive) which tackles the inefficiencies in balancing exploration strategies when combining SAC with novelty-based exploration. KEA integrates a novelty-augmented SAC with a standard SAC agent, proactively coordinated via a switching mechanism. This coordination allows the agent to maintain stochasticity in high-novelty regions, enhancing exploration efficiency and reducing repeated sample collection. We first analyze this potential issue in a 2D navigation task, and then evaluate KEA on the DeepSea hard-exploration benchmark as well as sparse reward control tasks from the DeepMind Control Suite. Compared to state-of-the-art novelty-based exploration baselines, our experiments show that KEA significantly improves learning efficiency and robustness in sparse reward setups.
Shih-Min Yang, Martin Magnusson 0002, Johannes A. Stork, Todor Stoyanov
ICML2
2025 Fast Online Learning of CLiFF-Maps in Changing Environments
abstract
Maps of dynamics are effective representations of motion patterns learned from prior observations, with recent research demonstrating their ability to enhance various downstream tasks such as human-aware robot navigation, long-term human motion prediction, and robot localization. Current advancements have primarily concentrated on methods for learning maps of human flow in environments where the flow is static, i.e., not assumed to change over time. In this paper we propose an online update method of the CLiFF-map (an advanced map of dynamics type that models motion patterns as velocity and orientation mixtures) to actively detect and adapt to human flow changes. As new observations are collected, our goal is to update a CLiFF-map to effectively and accurately integrate them, while retaining relevant historic motion patterns. The proposed online update method maintains a probabilistic representation in each observed location, updating parameters by continuously tracking sufficient statistics. In experiments using both synthetic and real-world datasets, we show that our method is able to maintain accurate representations of human motion dynamics, contributing to high performance flow-compliant planning downstream tasks, while being orders of magnitude faster than the comparable baselines.
Andrey Rudenko, Luigi Palmieri, Lukas Heuer, Achim J. Lilienthal, Martin Magnusson 0002
ICRA6
2025 Gaze-supported Large Language Model Framework for Bi-directional Human-Robot Interaction
abstract
The rapid development of Large Language Models (LLMs) creates an exciting potential for flexible, general knowledge-driven Human-Robot Interaction (HRI) systems for assistive robots. Existing HRI systems demonstrate great progress in interpreting and following user instructions, action generation, and robot task solving. On the other hand, bi-directional, multimodal, and context-aware support of the user in collaborative tasks still remains an open challenge. In this paper, we present a gaze- and speech-informed interface to the assistive robot, which is able to perceive the working environment from multiple vision inputs and support the dynamic user in their tasks. Our system is designed to be modular and transferable to adapt to diverse tasks and robots, and it is real-time capable due to the language-based interaction state representation and fast on-board perception modules. Its development was supported by multiple public dissemination events, contributing important considerations for improved robustness and user experience. Furthermore, in a lab study, we compare the performance and user ratings of our system with those of a traditional scripted HRI pipeline. Our findings indicate that an LLM-based approach enhances adaptability and marginally improves user engagement and task execution metrics but may produce redundant output, while a scripted pipeline is well suited for more straightforward tasks.
Jens Rüppel, Andrey Rudenko, Tim Schreiter, Martin Magnusson 0002, Achim J. Lilienthal
RO-MAN4
2024 Doppler-only Single-scan 3D Vehicle Odometry
abstract
We present a novel 3D odometry method that recovers the full motion of a vehicle only from a Doppler-capable range sensor. It leverages the radial velocities measured from the scene, estimating the sensor’s velocity from a single scan. The vehicle’s 3D motion, defined by its linear and angular velocities, is calculated taking into consideration its kinematic model which provides a constraint between the velocity measured at the sensor frame and the vehicle frame.Experiments carried out prove the viability of our single-sensor method compared to mounting an additional IMU. Our method provides a more reliable translation of the sensor, compared to the errors linked to IMUs due to noise and biases. Its short-term accuracy and fast operation (∼5ms) make it a proper candidate to supply the initialization to more complex localization algorithms or mapping pipelines. Not only does it reduce the error of the mapper, but it does so at a comparable level of accuracy as an IMU would. All without the need to mount and calibrate an extra sensor on the vehicle.
Andres Galeote-Luque, Vladimir Kubelka, Martin Magnusson 0002, José-Raúl Ruiz-Sarmiento, Javier González 0001
ICRA3
2024 Benchmarking Multi-Robot Coordination in Realistic, Unstructured Human-Shared Environments
abstract
Coordinating a fleet of robots in unstructured, human-shared environments is challenging. Human behavior is hard to predict, and its uncertainty impacts the performance of the robotic fleet. Various multi-robot planning and coordination algorithms have been proposed, including Multi-Agent Path Finding (MAPF) methods to precedence-based algorithms. However, it is still unclear how human presence impacts different coordination strategies in both simulated environments and the real world. With the goal of studying and further improving multi-robot planning capabilities in those settings, we propose a method to develop and benchmark different multi-robot coordination algorithms in realistic, unstructured and human-shared environments. To this end, we introduce a multi-robot benchmark framework that is based on state-of-the-art open-source navigation and simulation frameworks and can use different types of robots, environments and human motion models. We show a possible application of the benchmark framework with two different environments and three centralized coordination methods (two MAPF algorithms and a loosely-coupled coordination method based on precedence constraints). We evaluate each environment for different human densities to investigate its impact on each coordination method. We also present preliminary results that show how informing each coordination method about human presence can help the coordination method to find faster paths for the robots.
Lukas Heuer, Luigi Palmieri, Anna Mannucci, Sven Koenig, Martin Magnusson 0002
ICRA5
2024 Do we need scan-matching in radar odometry?
abstract
There is a current increase in the development of "4D" Doppler-capable radar and lidar range sensors that produce 3D point clouds where all points also have information about the radial velocity relative to the sensor. 4D radars in particular are interesting for object perception and navigation in low-visibility conditions (dust, smoke) where lidars and cameras typically fail. With the advent of high-resolution Doppler-capable radars comes the possibility of estimating odometry from single point clouds, foregoing the need for scan registration which is error-prone in feature-sparse field environments. We compare several odometry estimation methods, from direct integration of Doppler/IMU data and Kalman filter sensor fusion to 3D scan-to-scan and scan-to-map registration, on three datasets with data from two recent 4D radars and two IMUs. Surprisingly, our results show that the odometry from Doppler and IMU data alone give similar or better results than 3D point cloud registration. In our experiments, the position drift can be as low as 0.9% over 1.8 and 4.5km trajectories. That allows accurate estimation of 6-DOF ego-motion over long distances also in feature-sparse mine environments. These results are useful not least for applications of navigation with resource-constrained robot platforms in feature-sparse and low-visibility conditions such as mining, construction, and search & rescue operations.
Vladimir Kubelka, Emil Fritz, Martin Magnusson 0002
ICRA3
2024 3QFP: Efficient neural implicit surface reconstruction using Tri-Quadtrees and Fourier feature Positional encoding
abstract
Neural implicit surface representations are currently receiving a lot of interest as a means to achieve high-fidelity surface reconstruction at a low memory cost, compared to traditional explicit representations. However, state-of-the-art methods still struggle with excessive memory usage and non-smooth surfaces. This is particularly problematic in large-scale applications with sparse inputs, as is common in robotics use cases. To address these issues, we first introduce a sparse structure, tri-quadtrees, which represents the environment using learnable features stored in three planar quadtree projections. Secondly, we concatenate the learnable features with a Fourier feature positional encoding. The combined features are then decoded into signed distance values through a small multilayer perceptron. We demonstrate that this approach facilitates smoother reconstruction with a higher completion ratio with fewer holes. Compared to two recent baselines, one implicit and one explicit, our approach requires only 10%–50% as much memory, while achieving competitive quality. The code is released on https://github.com/ljjTYJR/3QFP.
Malcolm Mielle, Achim J. Lilienthal, Martin Magnusson 0002
ICRA4
2024 Learning Extrinsic Dexterity with Parameterized Manipulation Primitives
abstract
Many practically relevant robot grasping problems feature a target object for which all grasps are occluded, e.g., by the environment. Single-shot grasp planning invariably fails in such scenarios. Instead, it is necessary to first manipulate the object into a configuration that affords a grasp. We solve this problem by learning a sequence of actions that utilize the environment to change the object’s pose. Concretely, we employ hierarchical reinforcement learning to combine a sequence of learned parameterized manipulation primitives. By learning the low-level manipulation policies, our approach can control the object’s state through exploiting interactions between the object, the gripper, and the environment. Designing such a complex behavior analytically would be infeasible under uncontrolled conditions, as an analytic approach requires accurate physical modeling of the interaction and contact dynamics. In contrast, we learn a hierarchical policy model that operates directly on depth perception data, without the need for object detection, pose estimation, or manual design of controllers. We evaluate our approach on picking box-shaped objects of various weight, shape, and friction properties from a constrained table-top workspace. Our method transfers to a real robot and is able to successfully complete the object picking task in 98% of experimental trials.
Shih-Min Yang, Martin Magnusson 0002, Johannes A. Stork, Todor Stoyanov
ICRA2
2024 LaCE-LHMP: Airflow Modelling-Inspired Long-Term Human Motion Prediction By Enhancing Laminar Characteristics in Human Flow
abstract
Long-term human motion prediction (LHMP) is essential for safely operating autonomous robots and vehicles in populated environments. It is fundamental for various applications, including motion planning, tracking, human-robot interaction and safety monitoring. However, accurate prediction of human trajectories is challenging due to complex factors, including, for example, social norms and environmental conditions. The influence of such factors can be captured through Maps of Dynamics (MoDs), which encode spatial motion patterns learned from (possibly scattered and partial) past observations of motion in the environment and which can be used for data-efficient, interpretable motion prediction (MoD-LHMP). To address the limitations of prior work, especially regarding accuracy and sensitivity to anomalies in long-term prediction, we propose the Laminar Component Enhanced LHMP approach (LaCE-LHMP). Our approach is inspired by data-driven airflow modelling, which estimates laminar and turbulent flow components and uses predominantly the laminar components to make flow predictions. Based on the hypothesis that human trajectory patterns also manifest laminar flow (that represents predictable motion) and turbulent flow components (that reflect more unpredictable and arbitrary motion), LaCE-LHMP extracts the laminar patterns in human dynamics and uses them for human motion prediction. We demonstrate the superior prediction performance of LaCE-LHMP through benchmark comparisons with state-of-the-art LHMP methods, offering an unconventional perspective and a more intuitive understanding of human movement patterns.
Han Fan, Andrey Rudenko, Martin Magnusson 0002, Erik Schaffernicht, Achim J. Lilienthal
ICRA4
2024 High-Fidelity SLAM Using Gaussian Splatting with Rendering-Guided Densification and Regularized Optimization
abstract
We propose a dense RGBD SLAM system based on 3D Gaussian Splatting that provides metrically accurate pose tracking and visually realistic reconstruction. To this end, we first propose a Gaussian densification strategy based on the rendering loss to map unobserved areas and refine reobserved areas. Second, we introduce extra regularization parameters to alleviate the "forgetting" problem during contiunous mapping, where parameters tend to overfit the latest frame and result in decreasing rendering quality for previous frames. Both mapping and tracking are performed with Gaussian parameters by minimizing re-rendering loss in a differentiable way. Compared to recent neural and concurrently developed Gaussian splatting RGBD SLAM baselines, our method achieves state-of-the-art results on the synthetic dataset Replica and competitive results on the real-world dataset TUM. The code is released on https://github.com/ljjTYJR/HF-SLAM.
Malcolm Mielle, Achim J. Lilienthal, Martin Magnusson 0002
IROS4
2024 Human Gaze and Head Rotation during Navigation, Exploration and Object Manipulation in Shared Environments with Robots
abstract
The human gaze is an important cue to signal intention, attention, distraction, and the regions of interest in the immediate surroundings. Gaze tracking can transform how robots perceive, understand, and react to people, enabling new modes of robot control, interaction, and collaboration. In this paper, we use gaze tracking data from a rich dataset of human motion (THÖR-MAGNI) to investigate the coordination between gaze direction and head rotation of humans engaged in various indoor activities involving navigation, interaction with objects, and collaboration with a mobile robot. In particular, we study the spread and central bias of fixations in diverse activities and examine the correlation between gaze direction and head rotation. We introduce various human motion metrics to enhance the understanding of gaze behavior in dynamic interactions. Finally, we apply semantic object labeling to decompose the gaze distribution into activity-relevant regions.
Tim Schreiter, Andrey Rudenko, Martin Magnusson 0002, Achim J. Lilienthal
RO-MAN3
2023 Proactive Model Predictive Control with Multi-Modal Human Motion Prediction in Cluttered Dynamic Environments
abstract
For robots navigating in dynamic environments, exploiting and understanding uncertain human motion prediction is key to generate efficient, safe and legible actions. The robot may perform poorly and cause hindrances if it does not reason over possible, multi-modal future social interactions. With the goal of enhancing autonomous navigation in cluttered environments, we propose a novel formulation for nonlinear model predictive control including multi-modal predictions of human motion. As a result, our approach leads to less conservative, smooth and intuitive human-aware navigation with reduced risk of collisions, and shows a good balance between task efficiency, collision avoidance and human comfort. To show its effectiveness, we compare our approach against the state of the art in crowded simulated environments, and with real-world human motion data from the THOR dataset. This comparison shows that we are able to improve task efficiency, keep a larger distance to humans and significantly reduce the collision time, when navigating in cluttered dynamic environ-ments. Furthermore, the method is shown to work robustly with different state-of-the-art human motion predictors.
Lukas Heuer, Luigi Palmieri, Andrey Rudenko, Anna Mannucci, Martin Magnusson 0002, Kai Oliver Arras
IROS5
2023 CLiFF-LHMP: Using Spatial Dynamics Patterns for Long- Term Human Motion Prediction
abstract
Human motion prediction is important for mobile service robots and intelligent vehicles to operate safely and smoothly around people. The more accurate predictions are, particularly over extended periods of time, the better a system can, e.g., assess collision risks and plan ahead. In this paper, we propose to exploit maps of dynamics (MoDs, a class of general representations of place-dependent spatial motion patterns, learned from prior observations) for long-term human motion prediction (LHMP). We present a new MoD-informed human motion prediction approach, named CLiFF-LHMP, which is data efficient, explainable, and insensitive to errors from an upstream tracking system. Our approach uses CLiFF -map, a specific MoD trained with human motion data recorded in the same environment. We bias a constant velocity prediction with samples from the CLiFF-map to generate multi-modal trajectory predictions. In two public datasets we show that this algorithm outperforms the state of the art for predictions over very extended periods of time, achieving 45 % more accurate prediction performance at 50s compared to the baseline.
Andrey Rudenko, Tomasz Kucner, Luigi Palmieri, Kai Oliver Arras, Achim J. Lilienthal, Martin Magnusson 0002
IROS7
2023 Advantages of Multimodal versus Verbal-Only Robot-to-Human Communication with an Anthropomorphic Robotic Mock Driver
abstract
Robots are increasingly used in shared environments with humans, making effective communication a necessity for successful human-robot interaction. In our work, we study a crucial component: active communication of robot intent. Here, we present an anthropomorphic solution where a humanoid robot communicates the intent of its host robot acting as an “Anthropomorphic Robotic Mock Driver” (ARMoD). We evaluate this approach in two experiments in which participants work alongside a mobile robot on various tasks, while the ARMoD communicates a need for human attention, when required, or gives instructions to collaborate on a joint task. The experiments feature two interaction styles of the ARMoD: a verbal-only mode using only speech and a multimodal mode, additionally including robotic gaze and pointing gestures to support communication and register intent in space. Our results show that the multimodal interaction style, including head movements and eye gaze as well as pointing gestures, leads to more natural fixation behavior. Participants naturally identified and fixated longer on the areas relevant for intent communication, and reacted faster to instructions in collaborative tasks. Our research further indicates that the ARMoD intent communication improves engagement and social interaction with mobile robots in workplace settings.
Tim Schreiter, Lucas Morillo-Méndez, Ravi Chadalavada, Andrey Rudenko, Erik Billing, Martin Magnusson 0002, Kai Oliver Arras, Achim J. Lilienthal
RO-MAN6
2023 Lidar-Level Localization With Radar? The CFEAR Approach to Accurate, Fast, and Robust Large-Scale Radar Odometry in Diverse Environments
abstract
This article presents an accurate, highly efficient, and learning-free method for large-scale odometry estimation using spinning radar, empirically found to generalize well across very diverse environments—outdoors, from urban to woodland, and indoors in warehouses and mines—without changing parameters. Our method integrates motion compensation within a sweep with one-to-many scan registration that minimizes distances between nearby oriented surface points and mitigates outliers with a robust loss function. Extending our previous approach conservative filtering for efficient and accurate radar odometry (CFEAR), we present an in-depth investigation on a wider range of datasets, quantifying the importance of filtering, resolution, registration cost and loss functions, keyframe history, and motion compensation. We present a new solving strategy and configuration that overcomes previous issues with sparsity and bias, and improves our state-of-the-art by 38%, thus, surprisingly, outperforming radar simultaneous localization and mapping (SLAM) and approaching lidar SLAM. The most accurate configuration achieves 1.09% error at 5 Hz on the Oxford benchmark, and the fastest achieves 1.79% error at 160 Hz.
Daniel Adolfsson, Martin Magnusson 0002, Anas W. Alhashimi, Achim J. Lilienthal, Henrik Andreasson
IEEE Trans. Robotics2
2021 Robust Frequency-Based Structure Extraction
abstract
State of the art mapping algorithms can produce high-quality maps. However, they are still vulnerable to clutter and outliers which can affect map quality and in consequence hinder the performance of a robot, and further map processing for semantic understanding of the environment. This paper presents ROSE, a method for building-level structure detection in robotic maps. ROSE exploits the fact that indoor environments usually contain walls and straight-line elements along a limited set of orientations. Therefore metric maps often have a set of dominant directions. ROSE extracts these directions and uses this information to segment the map into structure and clutter through filtering the map in the frequency domain (an approach substantially underutilised in the mapping applications). Removing the clutter in this way makes wall detection (e.g. using the Hough transform) more robust. Our experiments demonstrate that (1) the application of ROSE for decluttering can substantially improve structural feature retrieval (e.g., walls) in cluttered environments, (2) ROSE can successfully distinguish between clutter and structure in the map even with substantial amount of noise and (3) ROSE can numerically assess the amount of structure in the map.
Tomasz Kucner, Matteo Luperto, Stephanie Lowry, Martin Magnusson 0002, Achim J. Lilienthal
ICRA4
2021 CFEAR Radarodometry - Conservative Filtering for Efficient and Accurate Radar Odometry
abstract
This paper presents an accurate, highly efficient and learning free method for large-scale radar odometry estimation. By using a simple filtering technique that keeps the strongest returns, we produce a clean radar data representation and reconstruct surface normals for efficient and accurate scan matching. Registration is carried out by minimizing a point-to-line metric and robustness to outliers is achieved using a Huber loss. Drift is additionally reduced by jointly registering the latest scan to a history of keyframes. We found that our odometry pipeline generalize well to different sensor models and datasets without changing a single parameter. We evaluate our method in three widely different environments and demonstrate an improvement over spatially cross validated state-of-the-art with an overall translation error of 1.76% in a public urban radar odometry benchmark, running merely on a single laptop CPU thread at 55 Hz.
Daniel Adolfsson, Martin Magnusson 0002, Anas W. Alhashimi, Achim J. Lilienthal, Henrik Andreasson
IROS2
2020 Localising Faster: Efficient and precise lidar-based robot localisation in large-scale environments
abstract
This paper proposes a novel approach for global localisation of mobile robots in large-scale environments. Our method leverages learning-based localisation and filtering-based localisation, to localise the robot efficiently and precisely through seeding Monte Carlo Localisation (MCL) with a deeplearned distribution. In particular, a fast localisation system rapidly estimates the 6-DOF pose through a deep-probabilistic model (Gaussian Process Regression with a deep kernel), then a precise recursive estimator refines the estimated robot pose according to the geometric alignment. More importantly, the Gaussian method (i.e. deep probabilistic localisation) and nonGaussian method (i.e. MCL) can be integrated naturally via importance sampling. Consequently, the two systems can be integrated seamlessly and mutually benefit from each other. To verify the proposed framework, we provide a case study in large-scale localisation with a 3D lidar sensor. Our experiments on the Michigan NCLT long-term dataset show that the proposed method is able to localise the robot in 1.94 s on average (median of 0.8 s) with precision 0.75 m in a largescale environment of approximately 0.5 km2.
Li Sun 0005, Daniel Adolfsson, Martin Magnusson 0002, Henrik Andreasson, Ingmar Posner, Tom Duckett
ICRA3
2020 Natural Criteria for Comparison of Pedestrian Flow Forecasting Models
abstract
Models of human behaviour, such as pedestrian flows, are beneficial for safe and efficient operation of mobile robots. We present a new methodology for benchmarking of pedestrian flow models based on the afforded safety of robot navigation in human-populated environments. While previous evaluations of pedestrian flow models focused on their predictive capabilities, we assess their ability to support safe path planning and scheduling. Using real-world datasets gathered continuously over several weeks, we benchmark state-of-the-art pedestrian flow models, including both time-averaged and time-sensitive models. In the evaluation, we use the learned models to plan robot trajectories and then observe the number of times when the robot gets too close to humans, using a predefined social distance threshold. The experiments show that while traditional evaluation criteria based on model fidelity differ only marginally, the introduced criteria vary significantly depending on the model used, providing a natural interpretation of the expected safety of the system. For the time-averaged flow models, the number of encounters increases linearly with the percentage operating time of the robot, as might be reasonably expected. By contrast, for the time-sensitive models, the number of encounters grows sublinearly with the percentage operating time, by planning to avoid congested areas and times.
Tomas Vintr, Zhi Yan 0001, Kerem Eyisoy, Filip Kubis, Jan Blaha, Jirí Ulrich, Chittaranjan Srinivas Swaminathan, Sergi Molina Mellado, Tomasz Kucner, Martin Magnusson 0002, Grzegorz Cielniak, Jan Faigl, Tom Duckett, Achim J. Lilienthal, Tomás Krajník
IROS10
2018 2D Spatial Keystone Transform for Sub-Pixel Motion Extraction from Noisy Occupancy Grid Map
abstract
In this paper, we propose a novel sub-pixel motion extraction method, called as Two Dimensional Spatial Keystone Transform (2DS-KST), for the motion detection and estimation from successive noisy Occupancy Grid Maps (OGMs). It extends the KST in radar imaging or motion compensation to 2D real spatial case, based on multiple hypotheses about possible directions of moving obstacles. Simulation results show that 2DS-KST has a good performance on the extraction of sub-pixel motions in very noisy environment, especially for those slowly moving obstacles.
Hongqi Fan, Tomasz Kucner, Martin Magnusson 0002, Achim J. Lilienthal
FUSION4
2018 A Method to Segment Maps from Different Modalities Using Free Space Layout MAORIS: Map of Ripples Segmentation
abstract
How to divide floor plans or navigation maps into semantic representations, such as rooms and corridors, is an important research question in fields such as human-robot interaction, place categorization, or semantic mapping. While most works focus on segmenting robot built maps, those are not the only types of map a robot, or its user, can use. We present a method for segmenting maps from different modalities, focusing on robot built maps and hand-drawn sketch maps, and show better results than state of the art for both types. Our method segments the map by doing a convolution between the distance image of the map and a circular kernel, and grouping pixels of the same value. Segmentation is done by detecting ripple-like patterns where pixel values vary quickly, and merging neighboring regions with similar values. We identify a flaw in the segmentation evaluation metric used in recent works and propose a metric based on Matthews correlation coefficient (MCC). We compare our results to ground-truth segmentations of maps from a publicly available dataset, on which we obtain a better MCC than the state of the art with 0.98 compared to 0.65 for a recent Voronoi-based segmentation method and 0.70 for the DuDe segmentation method. We also provide a dataset of sketches of an indoor environment, with two possible sets of ground truth segmentations, on which our method obtains an MCC of 0.56 against 0.28 for the Voronoi-based segmentation method and 0.30 for DuDe.
Malcolm Mielle, Martin Magnusson 0002, Achim J. Lilienthal
ICRA2
2018 Down the CLiFF: Flow-Aware Tralatory Planning Under Motion Pattern Uncertainty
abstract
In this paper we address the problem of flow-aware trajectory planning in dynamic environments considering flow model uncertainty. Flow-aware planning aims to plan trajectories that adhere to existing flow motion patterns in the environment, with the goal to make robots more efficient, less intrusive and safer. We use a statistical model called CLiFF-map that can map flow patterns for both continuous media and discrete objects. We propose novel cost and biasing functions for an RRT* planning algorithm, which exploits all the information available in the CLiFF-map model, including uncertainties due to flow variability or partial observability. Qualitatively, a benefit of our approach is that it can also be tuned to yield trajectories with different qualities such as exploratory or cautious, depending on application requirements. Quantitatively, we demonstrate that our approach produces more flow-compliant trajectories, compared to two baselines.
Chittaranjan Srinivas Swaminathan, Tomasz Kucner, Martin Magnusson 0002, Luigi Palmieri, Achim J. Lilienthal
IROS3
2018 A Dual PHD Filter for Effective Occupancy Filtering in a Highly Dynamic Environment
abstract
Environment monitoring remains a major challenge for mobile robots, especially in densely cluttered or highly populated dynamic environments, where uncertainties originated from environment and sensor significantly challenge the robot's perception. This paper proposes an effective occupancy filtering method called the dual probability hypothesis density (DPHD) filter, which models uncertain phenomena, such as births, deaths, occlusions, false alarms, and miss detections, by using random finite sets. The key insight of our method lies in the connection of the idea of dynamic occupancy with the concepts of the phase space density in gas kinetic and the PHD in multiple target tracking. By modeling the environment as a mixture of static and dynamic parts, the DPHD filter separates the dynamic part from the static one with a unified filtering process, but has a higher computational efficiency than existing Bayesian Occupancy Filters (BOFs). Moreover, an adaptive newborn function and a detection model considering occlusions are proposed to improve the filtering efficiency further. Finally, a hybrid particle implementation of the DPHD filter is proposed, which uses a box particle filter with constant discrete states and an ordinary particle filter with a time-varying number of particles in a continuous state space to process the static part and the dynamic part, respectively. This filter has a linear complexity with respect to the number of grid cells occupied by dynamic obstacles. Real-world experiments on data collected by a lidar at a busy roundabout demonstrate that our approach can handle monitoring of a highly dynamic environment in real time.
Hongqi Fan, Tomasz Kucner, Martin Magnusson 0002, Tiancheng Li 0002, Achim J. Lilienthal
IEEE Trans. Intell. Transp. Syst.3
2017 Kinodynamic motion planning on Gaussian mixture fields
abstract
We present a mobile robot motion planning approach under kinodynamic constraints that exploits learned perception priors in the form of continuous Gaussian mixture fields. Our Gaussian mixture fields are statistical multi-modal motion models of discrete objects or continuous media in the environment that encode e.g. the dynamics of air or pedestrian flows. We approach this task using a recently proposed circular linear flow field map based on semi-wrapped GMMs whose mixture components guide sampling and rewiring in an RRT* algorithm using a steer function for non-holonomic mobile robots. In our experiments with three alternative baselines, we show that this combination allows the planner to very efficiently generate high-quality solutions in terms of path smoothness, path length as well as natural yet minimum control effort motions through multi-modal representations of Gaussian mixture fields.
Luigi Palmieri, Tomasz Kucner, Martin Magnusson 0002, Achim J. Lilienthal, Kai Oliver Arras
ICRA3
2017 Semi-supervised 3D place categorisation by descriptor clustering
abstract
Place categorisation; i.e., learning to group perception data into categories based on appearance; typically uses supervised learning and either visual or 2D range data. This paper shows place categorisation from 3D data without any training phase. We show that, by leveraging the NDT histogram descriptor to compactly encode 3D point cloud appearance, in combination with standard clustering techniques, it is possible to classify public indoor data sets with accuracy comparable to, and sometimes better than, previous supervised training methods. We also demonstrate the effectiveness of this approach to outdoor data, with an added benefit of being able to hierarchically categorise places into sub-categories based on a user-selected threshold. This technique relieves users of providing relevant training data, and only requires them to adjust the sensitivity to the number of place categories, and provide a semantic label to each category after the process is completed.
Martin Magnusson 0002, Tomasz Kucner, Saeed Gholami Shahbandi, Henrik Andreasson, Achim J. Lilienthal
IROS1
2017 Incorporating ego-motion uncertainty estimates in range data registration
abstract
Local scan registration approaches commonly only utilize ego-motion estimates (e.g. odometry) as an initial pose guess in an iterative alignment procedure. This paper describes a new method to incorporate ego-motion estimates, including uncertainty, into the objective function of a registration algorithm. The proposed approach is particularly suited for feature-poor and self-similar environments, which typically present challenges to current state of the art registration algorithms. Experimental evaluation shows significant improvements in accuracy when using data acquired by Automatic Guided Vehicles (AGVs) in industrial production and warehouse environments.
Henrik Andreasson, Daniel Adolfsson, Todor Stoyanov, Martin Magnusson 0002, Achim J. Lilienthal
IROS4
2017 Semantic-assisted 3D normal distributions transform for scan registration in environments with limited structure
abstract
Point cloud registration is a core problem of many robotic applications, including simultaneous localization and mapping. The Normal Distributions Transform (NDT) is a method that fits a number of Gaussian distributions to the data points, and then uses this transform as an approximation of the real data, registering a relatively small number of distributions as opposed to the full point cloud. This approach contributes to NDT's registration robustness and speed but leaves room for improvement in environments of limited structure. To address this limitation we propose a method for the introduction of semantic information extracted from the point clouds into the registration process. The paper presents a large scale experimental evaluation of the algorithm against NDT on two publicly available benchmark data sets. For the purpose of this test a measure of smoothness is used for the semantic partitioning of the point clouds. The results indicate that the proposed method improves the accuracy, robustness and speed of NDT registration, especially in unstructured environments, making NDT suitable for a wider range of applications.
Anestis Zaganidis, Martin Magnusson 0002, Tom Duckett, Grzegorz Cielniak
IROS2
2015 Beyond points: Evaluating recent 3D scan-matching algorithms
abstract
Given that 3D scan matching is such a central part of the perception pipeline for robots, thorough and large-scale investigations of scan matching performance are still surprisingly few. A crucial part of the scientific method is to perform experiments that can be replicated by other researchers in order to compare different results. In light of this fact, this paper presents a thorough comparison of 3D scan registration algorithms using a recently published benchmark protocol which makes use of a publicly available challenging data set that covers a wide range of environments. In particular, we evaluate two types of recent 3D registration algorithms - one local and one global. Both approaches take local surface structure into account, rather than matching individual points. After well over 100 000 individual tests, we conclude that algorithms using the normal distributions transform (NDT) provides accurate results compared to a modern implementation of the iterative closest point (ICP) method, when faced with scan data that has little overlap and weak geometric structure. We also demonstrate that the minimally uncertain maximum consensus (MUMC) algorithm provides accurate results in structured environments without needing an initial guess, and that it provides useful measures to detect whether it has succeeded or not. We also propose two amendments to the experimental protocol, in order to provide more valuable results in future implementations.
Martin Magnusson 0002, Narunas Vaskevicius, Todor Stoyanov, Kaustubh Pathak, Andreas Birk 0002
ICRA1
2013 Improving point-cloud accuracy from a moving platform in field operations
abstract
This paper presents a method for improving the quality of distorted 3D point clouds made from a vehicle equipped with a laser scanner moving over uneven terrain. Existing methods that use 3D point-cloud data (for tasks such as mapping, localisation, and object detection) typically assume that each point cloud is accurate. For autonomous robots moving in rough terrain, it is often the case that the vehicle moves a substantial amount during the acquisition of one point cloud, in which case the data will be distorted. The method proposed in this paper is capable of increasing the accuracy of 3D point clouds, without assuming any specific features of the environment (such as planar walls), without resorting to a “stop-scan-go” approach, and without relying on specialised and expensive hardware. Each new point cloud is matched to the previous using normal-distribution-transform (NDT) registration, after which a mini-loop closure is performed with a local, per-scan, graph-based SLAM method. The proposed method increases the accuracy of both the measured platform trajectory and the point cloud. The method is validated on both real-world and simulated data.
Hakan Almqvist, Martin Magnusson 0002, Todor Stoyanov, Achim J. Lilienthal
ICRA2
2013 Conditional transition maps: Learning motion patterns in dynamic environments
abstract
In this paper we introduce a method for learning motion patterns in dynamic environments. Representations of dynamic environments have recently received an increasing amount of attention in the research community. Understanding dynamic environments is seen as one of the key challenges in order to enable autonomous navigation in real-world scenarios. However, representing the temporal dimension is a challenge yet to be solved. In this paper we introduce a spatial representation, which encapsulates the statistical dynamic behavior observed in the environment. The proposed Conditional Transition Map (CTMap) is a grid-based representation that associates a probability distribution for an object exiting the cell, given its entry direction. The transition parameters are learned from a temporal signal of occupancy on cells by using a local-neighborhood cross-correlation method. In this paper, we introduce the CTMap, the learning approach and present a proof-of-concept method for estimating future paths of dynamic objects, called Conditional Probability Propagation Tree (CPPTree). The evaluation is done using a real-world dataset collected at a busy roundabout.
Tomasz Kucner, Jari Saarinen, Martin Magnusson 0002, Achim J. Lilienthal
IROS3
2012 Point set registration through minimization of the L2 distance between 3D-NDT models
abstract
Point set registration-the task of finding the best fitting alignment between two sets of point samples, is an important problem in mobile robotics. This article proposes a novel registration algorithm, based on the distance between Three-Dimensional Normal Distributions Transforms. 3D-NDT models - a sub-class of Gaussian Mixture Models with uniformly weighted, largely disjoint components, can be quickly computed from range point data. The proposed algorithm constructs 3D-NDT representations of the input point sets and then formulates an objective function based on the L2distance between the considered models. Analytic first and second order derivatives of the objective function are computed and used in a standard Newton method optimization scheme, to obtain the best-fitting transformation. The proposed algorithm is evaluated and shown to be more accurate and faster, compared to a state of the art implementation of the Iterative Closest Point and 3D-NDT Point-to-Distribution algorithms.
Todor Stoyanov, Martin Magnusson 0002, Achim J. Lilienthal
ICRA2
2011 On the accuracy of the 3D Normal Distributions Transform as a tool for spatial representation
abstract
The Three-Dimensional Normal Distributions Transform (3D-NDT) is a spatial modeling technique with applications in point set registration, scan similarity comparison, change detection and path planning. This work concentrates on evaluating three common variations of the 3D-NDT in terms of accuracy of representing sampled semi-structured environments. In a novel approach to spatial representation quality measurement, the 3D geometrical modeling task is formulated as a classification problem and its accuracy is evaluated with standard machine learning performance metrics. In this manner the accuracy of the 3D-NDT variations is shown to be comparable to, and in some cases to outperform that of the standard occupancy grid mapping model.
Todor Stoyanov, Martin Magnusson 0002, Hakan Almqvist, Achim J. Lilienthal
ICRA2
2011 Consistent pile-shape quantification for autonomous wheel loaders
abstract
This paper presents a study of approaches for selecting an efficient attack pose when loading piled materials with industrial construction vehicles. Automated handling of piled materials is a highly desired goal in many construction and mining applications. The main contributions of the paper are an experimental study of two novel approaches for selecting an attack pose from 3D data, compared to previously published approaches and extensions thereof. The outcome is based on quantitative validation, both with simulated data and data from a real-world scenario with nontrivial ground geometry.
Martin Magnusson 0002, Hakan Almqvist
IROS1
2010 Path planning in 3D environments using the Normal Distributions Transform
abstract
Planning feasible paths in fully three-dimensional environments is a challenging problem. Application of existing algorithms typically requires the use of limited 3D representations that discard potentially useful information. This article proposes a novel approach to path planning that utilizes a full 3D representation directly: the Three-Dimensional Normal Distributions Transform (3D-NDT). The well known wavefront planner is modified to use 3D-NDT as a basis for map representation and evaluated using both indoor and outdoor data sets. The use of 3D-NDT for path planning is thus demonstrated to be a viable choice with good expressive capabilities.
Todor Stoyanov, Martin Magnusson 0002, Henrik Andreasson, Achim J. Lilienthal
IROS2
2009 Appearance-based loop detection from 3D laser data using the normal distributions transform
abstract
We propose a new approach to appearance based loop detection from metric 3D maps, exploiting the NDT surface representation. Locations are described with feature histograms based on surface orientation and smoothness, and loop closure can be detected by matching feature histograms. We also present a quantitative performance evaluation using two real-world data sets, showing that the proposed method works well in different environments.
Martin Magnusson 0002, Henrik Andreasson, Andreas Nüchter, Achim J. Lilienthal
ICRA1
2009 Evaluation of 3D registration reliability and speed - A comparison of ICP and NDT
abstract
To advance robotic science it is important to perform experiments that can be replicated by other researchers to compare different methods. However, these comparisons tend to be biased, since re-implementations of reference methods often lack thoroughness and do not include the hands-on experience obtained during the original development process. This paper presents a thorough comparison of 3D scan registration algorithms based on a 3D mapping field experiment, carried out by two research groups that are leading in the field of 3D robotic mapping. The iterative closest points algorithm (ICP) is compared to the normal distributions transform (NDT). We also present an improved version of NDT with a substantially larger valley of convergence than previously published versions.
Martin Magnusson 0002, Andreas Nüchter, Christopher Lörken, Achim J. Lilienthal, Joachim Hertzberg
ICRA1
2008 Registration of colored 3D point clouds with a Kernel-based extension to the normal distributions transform
abstract
We present a new algorithm for scan registration of colored 3D point data which is an extension to the Normal Distributions Transform (NDT). The probabilistic approach of NDT is extended to a color-aware registration algorithm by modeling the point distributions as Gaussian mixture-models in color space. We discuss different point cloud registration techniques, as well as alternative variants of the proposed algorithm. Results showing improved robustness of the proposed method using real-world data acquired with a mobile robot and a time-of-flight camera are presented.
Benjamin Huhle, Martin Magnusson 0002, Wolfgang Straßer, Achim J. Lilienthal
ICRA2
2007 Has somethong changed here? Autonomous difference detection for security patrol robots
abstract
This paper presents a system for autonomous change detection with a security patrol robot. In an initial step a reference model of the environment is created and changes are then detected with respect to the reference model as differences in coloured 3D point clouds, which are obtained from a 3D laser range scanner and a CCD camera. The suggested approach introduces several novel aspects, including a registration method that utilizes local visual features to determine point correspondences (thus essentially working without an initial pose estimate) and the 3D-NDT representation with adaptive cell size to efficiently represent both the spatial and colour aspects of the reference model. Apart from a detailed description of the individual parts of the difference detection system, a qualitative experimental evaluation in an indoor lab environment is presented, which demonstrates that the suggested system is able register and detect changes in spatial 3D data and also to detect changes that occur in colour space and are not observable using range values only.
Henrik Andreasson, Martin Magnusson 0002, Achim J. Lilienthal
IROS2