EDBT 2026 Demo / reviewers in the wild / expert
Patric Jensfelt
dblp:60/4015
· DBLP profile ↗
91ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-1170-7162ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 80 · 5 first-author · 18 since 2021Systems, architecture and hardware · 54 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 13 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiMo: High-Speed Objects Motion Compensation in Point Clouds (Abstract Reprint)abstractLiDAR point cloud is essential for autonomous vehicles, but motion distortions from dynamic objects degrade the data quality. While previous work has considered distortions caused by ego motion, distortions caused by other moving objects remain largely overlooked, leading to errors in object shape and position. This distortion is particularly pronounced in high-speed environments such as highways and in multi-LiDAR configurations, a common setup for heavy vehicles. To address this challenge, we introduce HiMo, a pipeline that repurposes scene flow estimation for non-ego motion compensation, correcting the representation of dynamic objects in point clouds. We further propose SeFlow++, a real-time scene flow estimator that achieves state-of-the-art performance on both scene flow and motion compensation. We validate HiMo through extensive experiments on Argoverse 2, ZOD and a newly collected real-world dataset featuring highway driving and multi-LiDAR-equipped heavy vehicles. Qingwen Zhang, Ajinkya Khoche, Yi Yang 0095, Sina Sharif Mansouri, Olov Andersson, Patric Jensfelt |
AAAI | 7 |
| 2026 | BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal PretrainingabstractZero-shot 3D object classification is crucial for real-world applications like autonomous driving, however it is often hindered by a significant domain gap between the synthetic data used for training and the sparse, noisy LiDAR scans encountered in the real-world. Current methods trained solely on synthetic data fail to generalize to outdoor scenes, while those trained only on real data lack the semantic diversity to recognize rare or unseen objects. We introduce BlendCLIP, a multimodal pretraining framework that bridges this synthetic-to-real gap by strategically combining the strengths of both domains. We first propose a pipeline to generate a large-scale dataset of object-level triplets—consisting of a point cloud, image, and text description—mined directly from real-world driving data and human annotated 3D boxes. Our core contribution is a curriculum-based data mixing strategy that first grounds the model in the semantically rich synthetic CAD data before progressively adapting it to the specific characteristics of real-world scans. Our experiments show that our approach is highly label-efficient: introducing as few as 1.5% real-world samples per batch into training boosts zero-shot accuracy on the nuScenes benchmark by 27%. Consequently, our final model achieves state-of-the-art performance on challenging outdoor datasets like nuScenes and TruckScenes, improving over the best prior method by 19.3% on nuScenes, while maintaining strong generalization on diverse synthetic benchmarks. Our findings demonstrate that effective domain adaptation, not full-scale real-world annotation, is the key to unlocking robust open-vocabulary 3D perception. Our code and dataset will be released upon acceptance on https://github.com/kesu1/BlendCLIP. Ajinkya Khoche, Gergo László Nagy, Maciej Wozniak 0001, Thomas Gustafsson, Patric Jensfelt |
WACV | 5 |
| 2025 | Take a Chance on Me: How Robot Performance and Risk Behaviour Affects Trust and Risk-TakingabstractReal-world human-robot interactions often encompass uncertainty. This uncertainty can be handled in different ways, for example by designing robot planners to be more or less risk-tolerant. However, how users actually perceive different risk-taking behaviours in robots has yet to be described. Additionally, in the absence of guarantees on optimal robot performance, the interaction between risk and performance on user perceptions is also unclear. To address this gap, we conducted a user study with 84 participants investigating how robot performance and risk behaviour affects users' trust and risk-taking decisions. Participants collaborated with a Franka robot arm to perform a block-stacking task. We compared a robot which displays consistent but sub-optimal behaviours to a robot displaying risky but occasionally optimal behaviour. Risky robot behaviour led to higher trust than consistent behaviour when the robot was on average good at stacking blocks (high expectation), but lower trust when the robot was on average bad at stacking blocks (low expectation). Individual risk-willingness also predicted likelihood of selecting the risky robot over the consistent robot for future interactions, but only when the average expectation was low. These findings have implications for risk-aware planning and decision-making in mixed human-robot systems. Rebecca Stower, Anna Gautier, Maciej Wozniak 0001, Patric Jensfelt, Jana Tumova, Iolanda Leite |
HRI | 4 |
| 2025 | ArgoTweak: Towards Self-Updating HD Maps Through Structured PriorsabstractReliable integration of prior information is crucial for self-verifying and self-updating HD maps. However, no public dataset includes the required triplet of prior maps, current maps, and sensor data. As a result, existing methods must rely on synthetic priors, which create inconsistencies and lead to a significant sim2real gap. To address this, we introduce ArgoTweak, the first dataset to complete the triplet with realistic map priors. At its core, ArgoTweak employs a bijective mapping framework, breaking down large-scale modifications into fine-grained atomic changes at the map element level, thus ensuring interpretability. This paradigm shift enables accurate change detection and integration while preserving unchanged elements with high fidelity. Experiments show that training models on ArgoTweak significantly reduces the sim2real gap compared to synthetic priors. Extensive ablations further highlight the impact of structured priors and detailed change annotations. By establishing a benchmark for explainable, prior-aided HD mapping, ArgoTweak advances scalable, self-improving mapping solutions. The dataset, baselines, map modification toolbox, and further resources are available at https://kth-rpl.github.io/ArgoTweak/. Lena Wild, Rafael Valencia, Patric Jensfelt |
ICCV | 3 |
| 2025 | SSF: Sparse Long-Range Scene Flow for Autonomous DrivingabstractScene flow enables an understanding of the motion characteristics of the environment in the 3D world. It gains particular significance in the long-range, where object-based perception methods might fail due to sparse observations far away. Although significant advancements have been made in scene flow pipelines to handle large-scale point clouds, a gap remains in scalability with respect to long-range. We attribute this limitation to the common design choice of using dense feature grids, which scale quadratically with range. In this paper, we propose Sparse Scene Flow (SSF), a general pipeline for long-range scene flow, adopting a sparse convolution based backbone for feature extraction. This approach introduces a new challenge: a mismatch in size and ordering of sparse feature maps between time-sequential point scans. To address this, we propose a sparse feature fusion scheme, that augments the feature maps with virtual voxels at missing locations. Additionally, we propose a range-wise metric that implicitly gives greater importance to faraway points. Our method, SSF, achieves state-of-the-art results on the Argoverse2 dataset, demonstrating strong performance in long-range scene flow estimation. Our code is open-sourced at https://github.com/KTH-RPL/SSF.git. Ajinkya Khoche, Qingwen Zhang, Laura Pereira Sánchez, Aron Asefaw, Sina Sharif Mansouri, Patric Jensfelt |
ICRA | 6 |
| 2025 | DeltaFlow: An Efficient Multi-frame Scene Flow Estimation MethodabstractPrevious dominant methods for scene flow estimation focus mainly on input from two consecutive frames, neglecting valuable information in the temporal domain. While recent trends shift towards multi-frame reasoning, they suffer from rapidly escalating computational costs as the number of frames grows. To leverage temporal information more efficiently, we propose DeltaFlow ($\Delta$Flow), a lightweight 3D framework that captures motion cues via a $\Delta$ scheme, extracting temporal features with minimal computational cost, regardless of the number of frames. Additionally, scene flow estimation faces challenges such as imbalanced object class distributions and motion inconsistency. To tackle these issues, we introduce a Category-Balanced Loss to enhance learning across underrepresented classes and an Instance Consistency Loss to enforce coherent object motion, improving flow accuracy. Extensive evaluations on the Argoverse 2, Waymo and nuScenes datasets show that $\Delta$Flow achieves state-of-the-art performance with up to 22\% lower error and $2\times$ faster inference compared to the next-best multi-frame supervised method, while also demonstrating a strong cross-domain generalization ability. The code is open-sourced at https://github.com/Kin-Zhang/DeltaFlow along with trained model weights. Qingwen Zhang, Yushan Zhang, Yixi Cai, Olov Andersson, Patric Jensfelt |
NeurIPS | 6 |
| 2025 | Fusion in Context: A Multimodal Approach to Affective State RecognitionabstractAccurate recognition of human emotions is a crucial challenge in affective computing and human-robot interaction (HRI). Emotional states play a vital role in shaping behaviors, decisions, and social interactions. However, emotional expressions can be influenced by contextual factors, leading to misinterpretations if context is not considered. Multimodal fusion, combining modalities like facial expressions, speech, and physiological signals, has shown promise in improving affect recognition. This paper proposes a transformer-based multimodal fusion approach that leverages facial thermal data, facial action units, and textual context information for context-aware emotion recognition. We explore modality-specific encoders to learn tailored representations, which are then fused and processed by a shared transformer encoder to capture temporal dependencies and interactions. The proposed method is evaluated on a dataset collected from participants engaged in a tangible tabletop Pacman game designed to induce various affective states. Our results demonstrate improvements from incorporating contextual information and multimodal fusion, achieving 89% F1 score with our full model compared to 65% for action units alone and 30% for thermal data alone. Youssef Mohamed, Séverin Lemaignan, Arzu Güneysu, Patric Jensfelt, Christian Smith |
RO-MAN | 4 |
| 2025 | Are You an Expert? Instruction Adaptation Using Multi-Modal Affect Detections with Thermal Imaging and ContextabstractHuman-robot interactions increasingly require adaptive instruction delivery, yet robots struggle to calibrate instruction detail levels without explicit user input. We present a system that automatically modulates instruction granularity using real-time affect detection through multi-modal fusion of thermal imaging, facial expressions, and contextual information. Our transformer-based architecture integrates these signals to enable decisions about instruction delivery based on detected user states. In a between-subjects study (N=40), participants completed assembly tasks under either manual adjustment or automatic adaptation conditions. Results showed significantly fewer manual adjustments in the adaptive condition (0.7 vs 2.0 per session), with comparable user satisfaction across conditions. This work shows the effectiveness of affect-driven adaptive instruction in human-robot interaction, contributing to more responsive robotic interfaces while providing guidelines for balancing automation with user control. Youssef Mohamed, Séverin Lemaignan, Arzu Güneysu, Patric Jensfelt, Christian Smith |
RO-MAN | 4 |
| 2025 | Neural Graph Map: Dense Mapping with Efficient Loop Closure IntegrationabstractNeural field-based SLAM methods typically employ a single, monolithic field as their scene representation. This prevents efficient incorporation of loop closure constraints and limits scalability. To address these shortcomings, we propose a novel RGB-D neural mapping framework in which the scene is represented by a collection of lightweight neural fields which are dynamically anchored to the pose graph of a sparse visual SLAM system. Our approach shows the ability to integrate large-scale loop closures, while re-quiring only minimal reintegration. Furthermore, we verify the scalability of our approach by demonstrating success-ful building-scale mapping taking multiple loop closures into account during the optimization, and show that our method outperforms existing state-of-the-art approaches on large scenes in terms of quality and runtime. Our code is available open-source at https://github.com/KTH-RPL/neural_graph_mapping. Leonard Bruns, Jun Zhang 0102, Patric Jensfelt |
WACV | 3 |
| 2025 | HiMo: High-Speed Objects Motion Compensation in Point CloudsabstractLiDAR point cloud is essential for autonomous vehicles, but motion distortions from dynamic objects degrade the data quality. While previous work has considered distortions caused by ego motion, distortions caused by other moving objects remain largely overlooked, leading to errors in object shape and position. This distortion is particularly pronounced in high-speed environments such as highways and in multi-LiDAR configurations, a common setup for heavy vehicles. To address this challenge, we introduce HiMo, a pipeline that repurposes scene flow estimation for non-ego motion compensation, correcting the representation of dynamic objects in point clouds. During the development of HiMo, we observed that existing self-supervised scene flow estimators often produce degenerate or inconsistent estimates under high-speed distortion. We further propose SeFlow++, a real-time scene flow estimator that achieves state-of-the-art performance on both scene flow and motion compensation. Since well-established motion distortion metrics are absent in the literature, we introduce two evaluation metrics: compensation accuracy at a point level and shape similarity of objects. We validate HiMo through extensive experiments on Argoverse 2, ZOD and a newly collected real-world dataset featuring highway driving and multi-LiDAR-equipped heavy vehicles. Our findings show that HiMo improves the geometric consistency and visual fidelity of dynamic objects in LiDAR point clouds, benefiting downstream tasks such as semantic segmentation and 3D detection. See https://kin-zhang.github.io/HiMo for more details. Qingwen Zhang, Ajinkya Khoche, Yi Yang 0095, Sina Sharif Mansouri, Olov Andersson, Patric Jensfelt |
IEEE Trans. Robotics | 7 |
| 2024 | MCD: Diverse Large-Scale Multi-Campus Dataset for Robot PerceptionabstractPerception plays a crucial role in various robot applications. However, existing well-annotated datasets are biased towards autonomous driving scenarios, while unlabelled SLAM datasets are quickly over-fitted, and often lack environment and domain variations. To expand the frontier of these fields, we introduce a comprehensive dataset named MCD (Multi-Campus Dataset), featuring a wide range of sensing modalities, high-accuracy ground truth, and diverse challenging environments across three Eurasian university campuses. MCD comprises both CCS (Classical Cylindrical Spinning) and NRE (Non-Repetitive Epicyclic) lidars, high-quality IMUs (Inertial Measurement Units), cameras, and UWB (Ultra-WideBand) sensors. Further-more, in a pioneering effort, we introduce semantic annotations of 29 classes over 59k sparse NRE lidar scans across three domains, thus providing a novel challenge to existing semantic segmentation research upon this largely unexplored modality. Finally, we propose, for the first time to the best of our knowledge, continuous-time ground truth based on optimization-based registration of lidar-inertial data on three survey-grade prior maps, each several times larger than the next largest publicly available ones. We conduct a rigorous evaluation of numerous state-of-the-art algorithms on MCD, report their performance, and highlight the challenges awaiting solutions from the research community. Thien-Minh Nguyen, Shenghai Yuan 0001, Thien Hoang Nguyen, Pengyu Yin, Haozhi Cao, Lihua Xie 0001, Maciej Wozniak 0001, Patric Jensfelt, Marko Thiel 0002, Justin Ziegenbein, Noel Blunder |
CVPR | 8 |
| 2024 | SeFlow: A Self-supervised Scene Flow Method in Autonomous Driving
Qingwen Zhang, Yi Yang 0095, Peizheng Li, Olov Andersson, Patric Jensfelt |
ECCV (1) | 5 |
| 2024 | DeFlow: Decoder of Scene Flow Network in Autonomous DrivingabstractScene flow estimation determines a scene’s 3D motion field, by predicting the motion of points in the scene, especially for aiding tasks in autonomous driving. Many networks with large-scale point clouds as input use voxelization to create a pseudo-image for real-time running. However, the voxelization process often results in the loss of point-specific features. This gives rise to a challenge in recovering those features for scene flow tasks. Our paper introduces DeFlow which enables a transition from voxel-based features to point features using Gated Recurrent Unit (GRU) refinement. To further enhance scene flow estimation performance, we formulate a novel loss function that accounts for the data imbalance between static and dynamic points. Evaluations on the Argoverse 2 scene flow task reveal that DeFlow achieves state-of-the-art results on large-scale point cloud data, demonstrating that our network has better performance and efficiency compared to others. The code is available at https://github.com/KTH-RPL/deflow. Qingwen Zhang, Yi Yang 0095, Ruoyu Geng, Patric Jensfelt |
ICRA | 5 |
| 2024 | Conditional Variational Autoencoders for Probabilistic Pose RegressionabstractRobots rely on visual relocalization to estimate their pose from camera images when they lose track. One of the challenges in visual relocalization is repetitive structures in the operation environment of the robot. This calls for probabilistic methods that support multiple hypotheses for robot’s pose. We propose such a probabilistic method to predict the posterior distribution of camera poses given an observed image. Our proposed training strategy results in a generative model of camera poses given an image, which can be used to draw samples from the pose posterior distribution. Our method is streamlined and well-founded in theory and outperforms existing methods on localization in presence of ambiguities. Fereidoon Zangeneh, Leonard Bruns, Amit Dekel, Alessandro Pieropan, Patric Jensfelt |
IROS | 5 |
| 2024 | Towards Long-Range 3D Object Detection for Autonomous Vehiclesabstract3D object detection at long-range is crucial for ensuring the safety and efficiency of self-driving vehicles, allowing them to accurately perceive and react to objects, obstacles, and potential hazards from a distance. But most current state-of-the-art LiDAR based methods are range limited due to sparsity at long-range, which generates a form of domain gap between points closer to and farther away from the ego vehicle. Another related problem is the label imbalance for faraway objects, which inhibits the performance of Deep Neural Networks at long-range. To address the above limitations, we investigate two ways to improve long-range performance of current LiDAR-based 3D detectors. First, we combine two 3D detection networks, referred to as range experts, one specializing at near to mid-range objects, and one at long-range 3D detection. To train a detector at long-range under a scarce label regime, we further weigh the loss according to the labelled point’s distance from ego vehicle. Second, we augment LiDAR scans with virtual points generated using Multimodal Virtual Points (MVP), a readily available image-based depth completion algorithm. Our experiments on the long-range Argoverse2 (AV2) dataset indicate that MVP is more effective in improving long range performance, while maintaining a straightforward implementation. On the other hand, the range experts offer a computationally efficient and simpler alternative, avoiding dependency on image-based segmentation networks and perfect camera-LiDAR calibration. Ajinkya Khoche, Laura Pereira Sánchez, Nazre Batool, Sina Sharif Mansouri, Patric Jensfelt |
IV | 5 |
| 2023 | Increasing Perceived Safety in Motion Planning for Human-Drone InteractionabstractSafety is crucial for autonomous drones to operate close to humans. Besides avoiding unwanted or harmful contact, people should also perceive the drone as safe. Existing safe motion planning approaches for autonomous robots, such as drones, have primarily focused on ensuring physical safety, e.g., by imposing constraints on motion planners. However, studies indicate that ensuring physical safety does not necessarily lead to perceived safety. Prior work in Human-Drone Interaction (HDI) shows that factors such as the drone's speed and distance to the human are important for perceived safety. Building on these works, we propose a parameterized control barrier function (CBF) that constrains the drone's maximum deceleration and minimum distance to the human and update its parameters on people's ratings of perceived safety. We describe an implementation and evaluation of our approach. Results of a within-subject user study (N=15) show that we can improve perceived safety of a drone by adjusting to people individually. Sanne van Waveren, Rasmus Rudling, Iolanda Leite, Patric Jensfelt, Christian Pek |
HRI | 4 |
| 2023 | A Probabilistic Framework for Visual Localization in Ambiguous ScenesabstractVisual localization allows autonomous robots to relocalize when losing track of their pose by matching their current observation with past ones. However, ambiguous scenes pose a challenge for such systems, as repetitive structures can be viewed from many distinct, equally likely camera poses, which means it is not sufficient to produce a single best pose hypothesis. In this work, we propose a probabilistic framework that for a given image predicts the arbitrarily shaped posterior distribution of its camera pose. We do this via a novel formulation of camera pose regression using variational inference, which allows sampling from the predicted distribution. Our method outperforms existing methods on localization in ambiguous scenes. We open-source our approach and share our recorded data sequence at github.com/efreidun/vapor. Fereidoon Zangeneh, Leonard Bruns, Amit Dekel, Alessandro Pieropan, Patric Jensfelt |
ICRA | 5 |
| 2023 | Happily Error After: Framework Development and User Study for Correcting Robot Perception Errors in Virtual RealityabstractWhile we can see robots in more areas of our lives, they still make errors. One common cause of failure stems from the robot perception module when detecting objects. Allowing users to correct such errors can help improve the interaction and prevent the same errors in the future. Consequently, we investigate the effectiveness of a virtual reality (VR) framework for correcting perception errors of a Franka Panda robot. We conducted a user study with 56 participants who interacted with the robot using both VR and screen interfaces. Participants learned to collaborate with the robot faster in the VR interface compared to the screen interface. Additionally, participants found the VR interface more immersive, enjoyable, and expressed a preference for using it again. These findings suggest that VR interfaces may offer advantages over screen interfaces for human-robot interaction in erroneous environments. Maciej Wozniak 0001, Rebecca Stower, Patric Jensfelt, André Pereira 0001 |
RO-MAN | 3 |
| 2022 | FloorGenT: Generative Vector Graphic Model of Floor Plans for RoboticsabstractFloor plans are the basis of reasoning in and communicating about indoor environments. In this paper, we show that by modelling floor plans as sequences of line segments seen from a particular point of view, recent advances in autoregressive sequence modelling can be leveraged to model and predict floor plans. The line segments are canonicalized and translated to sequence of tokens and an attention-based neural network is used to fit a one-step distribution over next tokens. We fit the network to sequences derived from a set of large-scale floor plans, and demonstrate the capabilities of the model in four scenarios: novel floor plan generation, completion of partially observed floor plans, generation of floor plans from simulated sensor data, and finally, the applicability of a floor plan model in predicting the shortest distance with partial knowledge of the environment. Ludvig Ericson, Patric Jensfelt |
IROS | 2 |
| 2022 | In Memoriam: Jan-Olof Eklundh
Atsuto Maki, Danica Kragic, Hedvig Kjellström, Hossein Azizpour, Josephine Sullivan, Mårten Björkman, Patric Jensfelt, Stefan Carlsson, Tony Lindeberg, Yngve Sundblad |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2021 | Cross-layer Configuration Optimization for Localization on Resource-constrained DevicesabstractMobile devices are increasingly expected to sup-port high-performance cyber-physical applications in small form factors, e.g., drones and rovers. However, the gap between hardware limitations of these devices and application requirements is still prohibitive – conflicting goals such as robust, accurate, and efficient execution must be managed carefully to achieve acceptable operation. In this paper, we explore the tradeoff between performance and efficiency in such cyber-physical systems, specifically with respect to localization (a core task for any mobile autonomous device). We perform a design space exploration (DSE) given a number of configurable parameters for both localization algorithm and platform layers. Given the configuration space, we formulate a cross-layer multi-objective optimization problem to explore the tradeoff between localization accuracy and power consumption. We then propose a predictive model for robust execution that can be used to determine desirable configurations at runtime in the face of environmental changes. Sandra Hernández, José Araújo, Patric Jensfelt, Ananya Muddukrishna, Bryan Donyanavard |
IROS | 3 |
| 2020 | Adversarial Feature Training for Generalizable Robotic Visuomotor ControlabstractDeep reinforcement learning (RL) has enabled training action-selection policies, end-to-end, by learning a function which maps image pixels to action outputs. However, it's application to visuomotor robotic policy training has been limited because of the challenge of large-scale data collection when working with physical hardware. A suitable visuomotor policy should perform well not just for the task-setup it has been trained for, but also for all varieties of the task, including novel objects at different viewpoints surrounded by task-irrelevant objects. However, it is impractical for a robotic setup to sufficiently collect interactive samples in a RL framework to generalize well to novel aspects of a task. In this work, we demonstrate that by using adversarial training for domain transfer, it is possible to train visuomotor policies based on RL frameworks, and then transfer the acquired policy to other novel task domains. We propose to leverage the deep RL capabilities to learn complex visuomotor skills for uncomplicated task setups, and then exploit transfer learning to generalize to new task domains provided only still images of the task in the target domain. We evaluate our method on two real robotic tasks, picking and pouring, and compare it to a number of prior works, demonstrating its superiority. Xi Chen 0051, Ali Ghadirzadeh, Mårten Björkman, Patric Jensfelt |
ICRA | 4 |
| 2019 | Object Detection Approach for Robot Grasp DetectionabstractIn this paper, we focus on the robot grasping problem with parallel grippers using image data. For this task, we propose and implement an end-to-end approach. In order to detect the good grasping poses for a parallel gripper from RGB images, we have employed transfer learning for a Convolutional Neural Network (CNN) based object detection architecture. Our obtained results show that, the adapted network either outperforms or is on-par with the state-of-the art methods on a benchmark dataset. We also performed grasping experiments on a real robot platform to evaluate our method's real world performance. Hakan Karaoguz, Patric Jensfelt |
ICRA | 2 |
| 2019 | Knowledge is Never Enough: Towards Web Aided Deep Open World RecognitionabstractWhile today's robots are able to perform sophisticated tasks, they can only act on objects they have been trained to recognize. This is a severe limitation: any robot will inevitably see new objects in unconstrained settings, and thus will always have visual knowledge gaps. However, standard visual modules are usually built on a limited set of classes and are based on the strong prior that an object must belong to one of those classes. Identifying whether an instance does not belong to the set of known categories (i.e. open set recognition), only partially tackles this problem, as a truly autonomous agent should be able not only to detect what it does not know, but also to extend dynamically its knowledge about the world. We contribute to this challenge with a deep learning architecture that can dynamically update its known classes in an end-to-end fashion. The proposed deep network, based on a deep extension of a non-parametric model, detects whether a perceived object belongs to the set of categories known by the system and learns it without the need to retrain the whole system from scratch. Annotated images about the new category can be provided by an `oracle' (i.e. human supervision), or by autonomous mining of the Web. Experiments on two different databases and on a robot platform demonstrate the promise of our approach. Massimiliano Mancini, Hakan Karaoguz, Elisa Ricci 0001, Patric Jensfelt, Barbara Caputo |
ICRA | 4 |
| 2019 | Meta-Learning for Multi-objective Reinforcement LearningabstractMulti-objective reinforcement learning (MORL) is the generalization of standard reinforcement learning (RL) approaches to solve sequential decision making problems that consist of several, possibly conflicting, objectives. Generally, in such formulations, there is no single optimal policy which optimizes all the objectives simultaneously, and instead, a number of policies has to be found each optimizing a preference of the objectives. In this paper, we introduce a novel MORL approach by training a meta-policy, a policy simultaneously trained with multiple tasks sampled from a task distribution, for a number of randomly sampled Markov decision processes (MDPs). In other words, the MORL is framed as a meta-learning problem, with the task distribution given by a distribution over the preferences. We demonstrate that such a formulation results in a better approximation of the Pareto optimal solutions in terms of both the optimality and the computational efficiency. We evaluated our method on obtaining Pareto optimal policies using a number of continuous control problems with high degrees of freedom. Xi Chen 0051, Ali Ghadirzadeh, Mårten Björkman, Patric Jensfelt |
IROS | 4 |
| 2019 | Detection and Tracking of General Movable Objects in Large Three-Dimensional MapsabstractThis paper studies the problem of detection and tracking of general objects with semistatic dynamics observed by a mobile robot moving in a large environment. A key problem is that due to the environment scale, the robot can only observe a subset of the objects at any given time. Since some time passes between observations of objects in different places, the objects might be moved when the robot is not there. We propose a model for this movement in which the objects typically only move locally, but with some small probability they jump longer distances through what we call global motion. For filtering, we decompose the posterior over local and global movements into two linked processes. The posterior over the global movements and measurement associations is sampled, while we track the local movement analytically using Kalman filters. This novel filter is evaluated on point cloud data gathered autonomously by a mobile robot over an extended period of time. We show that tracking jumping objects is feasible, and that the proposed probabilistic treatment outperforms previous methods when applied to real world data. The key to efficient probabilistic tracking in this scenario is focused sampling of the object posteriors. Nils Bore, Johan Ekekrantz, Patric Jensfelt, John Folkesson |
IEEE Trans. Robotics | 3 |
| 2018 | The Obstacle-restriction Method for Tele-operation of Unmanned Aerial Vehicles with Restricted MotionabstractThis paper presents a collision avoidance method for tele-operated unmanned aerial vehicles (UAVs). The method is designed to assist the operator at all times, such that the operator can focus solely on the main objectives instead of avoiding obstacles. We restrict the altitude to be fixed in a three dimensional environment to simplify the control and operation of the UAV. The method contributes a number of desired properties not found in other collision avoidance systems for tele-operated UAVs. Our method i) can handle situations where there is no input from the user by actively stopping and proceeding to avoid obstacles, ii) allows the operator to slide between prioritizing staying away from objects and getting close to them in a safe way when so required, and iii) provides for intuitive control by not deviating too far from the control input of the operator. We demonstrate the effectiveness of the method in real world experiments with a physical hexacopter in different indoor scenarios. We also present simulation results where we compare controlling the UAV with and without our method activated. Daniel Duberg, Patric Jensfelt |
ICARCV | 2 |
| 2018 | Semantic Labeling of Indoor Environments from 3D RGB MapsabstractWe present an approach to automatically assign semantic labels to rooms reconstructed from 3D RGB maps of apartments. Evidence for the room types is generated using state-of-the-art deep-learning techniques for scene classification and object detection based on automatically generated virtual RGB views, as well as from a geometric analysis of the map's 3D structure. The evidence is merged in a conditional random field, using statistics mined from different datasets of indoor environments. We evaluate our approach qualitatively and quantitatively and compare it to related methods. Manuel Brucker, Maximilian Durner, Rares Ambrus, Zoltan-Csaba Marton, Axel Wendt, Patric Jensfelt, Kai Oliver Arras, Rudolph Triebel |
ICRA | 6 |
| 2018 | Interactive, Collaborative Robots: Challenges and OpportunitiesabstractRobotic technology has transformed manufacturing industry ever since the first industrial robot was put in use in the beginning of the 60s. The challenge of developing flexible solutions where production lines can be quickly re-planned, adapted and structured for new or slightly changed products is still an important open problem. Industrial robots today are still largely preprogrammed for their tasks, not able to detect errors in their own performance or to robustly interact with a complex environment and a human worker. The challenges are even more serious when it comes to various types of service robots. Full robot autonomy, including natural interaction, learning from and with human, safe and flexible performance for challenging tasks in unstructured environments will remain out of reach for the foreseeable future. In the envisioned future factory setups, home and office environments, humans and robots will share the same workspace and perform different object manipulation tasks in a collaborative manner. We discuss some of the major challenges of developing such systems and provide examples of the current state of the art. Danica Kragic, Joakim Gustafson, Hakan Karaoguz, Patric Jensfelt, Robert Krug 0002 |
IJCAI | 4 |
| 2018 | Deep Reinforcement Learning to Acquire Navigation Skills for Wheel-Legged Robots in Complex EnvironmentsabstractMobile robot navigation in complex and dynamic environments is a challenging but important problem. Reinforcement learning approaches fail to solve these tasks efficiently due to reward sparsities, temporal complexities and high-dimensionality of sensorimotor spaces which are inherent in such problems. We present a novel approach to train action policies to acquire navigation skills for wheel-legged robots using deep reinforcement learning. The policy maps height-map image observations to motor commands to navigate to a target position while avoiding obstacles. We propose to acquire the multifaceted navigation skill by learning and exploiting a number of manageable navigation behaviors. We also introduce a domain randomization technique to improve the versatility of the training samples. We demonstrate experimentally a significant improvement in terms of data-efficiency, success rate, robustness against irrelevant sensory data, and also the quality of the maneuver skills. Xi Chen 0051, Ali Ghadirzadeh, John Folkesson, Mårten Björkman, Patric Jensfelt |
IROS | 5 |
| 2018 | Kitting in the Wild through Online Domain AdaptationabstractTechnological developments call for increasing perception and action capabilities of robots. Among other skills, vision systems that can adapt to any possible change in the working conditions are needed. Since these conditions are unpredictable, we need benchmarks which allow to assess the generalization and robustness capabilities of our visual recognition algorithms. In this work we focus on robotic kitting in unconstrained scenarios. As a first contribution, we present a new visual dataset for the kitting task. Differently from standard object recognition datasets, we provide images of the same objects acquired under various conditions where camera, illumination and background are changed. This novel dataset allows for testing the robustness of robot visual recognition algorithms to a series of different domain shifts both in isolation and unified. Our second contribution is a novel online adaptation algorithm for deep models, based on batch-normalization layers, which allows to continuously adapt a model to the current working conditions. Differently from standard domain adaptation algorithms, it does not require any image from the target domain at training time. We benchmark the performance of the algorithm on the proposed dataset, showing its capability to fill the gap between the performances of a standard architecture and its counterpart adapted offline to the given target domain. Massimiliano Mancini, Hakan Karaoguz, Elisa Ricci 0001, Patric Jensfelt, Barbara Caputo |
IROS | 4 |
| 2017 | Autonomous meshing, texturing and recognition of object models with a mobile robotabstractWe present a system for creating object models from RGB-D views acquired autonomously by a mobile robot. We create high-quality textured meshes of the objects by approximating the underlying geometry with a Poisson surface. Our system employs two optimization steps, first registering the views spatially based on image features, and second aligning the RGB images to maximize photometric consistency with respect to the reconstructed mesh. We show that the resulting models can be used robustly for recognition by training a Convolutional Neural Network (CNN) on images rendered from the reconstructed meshes. We perform experiments on data collected autonomously by a mobile robot both in controlled and uncontrolled scenarios. We compare quantitatively and qualitatively to previous work to validate our approach. Rares Ambrus, Nils Bore, John Folkesson, Patric Jensfelt |
IROS | 4 |
| 2017 | Geometric and visual terrain classification for autonomous mobile navigationabstractIn this paper, we present a multi-sensory terrain classification algorithm with a generalized terrain representation using semantic and geometric features. We compute geometric features from lidar point clouds and extract pixel-wise semantic labels from a fully convolutional network that is trained using a dataset with a strong focus on urban navigation. We use data augmentation to overcome the biases of the original dataset and apply transfer learning to adapt the model to new semantic labels in off-road environments. Finally, we fuse the visual and geometric features using a random forest to classify the terrain traversability into three classes: safe, risky and obstacle. We implement the algorithm on our four-wheeled robot and test it in novel environments including both urban and off-road scenes which are distinct from the training environments and under summer and winter conditions. We provide experimental result to show that our algorithm can perform accurate and fast prediction of terrain traversability in a mixture of environments with a small set of training data. Fabian Schilling, Xi Chen 0051, John Folkesson, Patric Jensfelt |
IROS | 4 |
| 2017 | Human-centric partitioning of the environmentabstractIn this paper, we present an object based approach for human-centric partitioning of the environment. Our approach for determining the human-centric regions is to detect the objects that are commonly associated with frequent human presence. In order to detect these objects, we employ state of the art perception techniques. The detected objects are stored with their spatio-temporal information in the robot's memory to be later used for generating the regions. The advantages of our method is that it is autonomous, requires only a small set of perceptual data and does not even require people to be present while generating the regions. The generated regions are validated using a 1-month dataset collected in an indoor office environment. The experimental results show that although a small set of perceptual data is used, the regions are generated at densely occupied locations. Hakan Karaoguz, Nils Bore, John Folkesson, Patric Jensfelt |
RO-MAN | 4 |
| 2017 | Non-parametric spatial context structure learning for autonomous understanding of human environmentsabstractAutonomous scene understanding by object classification today, crucially depends on the accuracy of appearance based robotic perception. However, this is prone to difficulties in object detection arising from unfavourable lighting conditions and vision unfriendly object properties. In our work, we propose a spatial context based system which infers object classes utilising solely structural information captured from the scenes to aid traditional perception systems. Our system operates on novel spatial features (IFRC) that are robust to noisy object detections; It also caters to on-the-fly learned knowledge modification improving performance with practise. IFRC are aligned with human expression of 3D space, thereby facilitating easy HRI and hence simpler supervised learning. We tested our spatial context based system to successfully conclude that it can capture spatio structural information to do joint object classification to not only act as a vision aide, but sometimes even perform on par with appearance based robotic vision. Akshaya Thippur, Johannes A. Stork, Patric Jensfelt |
RO-MAN | 3 |
| 2017 | Robot task planning and explanation in open and uncertain worlds
Marc Hanheide, Moritz Göbelbecker, Graham S. Horn, Andrzej Pronobis, Kristoffer Sjöö, Alper Aydemir, Patric Jensfelt, Charles Gretton, Richard Dearden, Miroslav Janícek, Hendrik Zender, Geert-Jan M. Kruijff, Nick Hawes, Jeremy L. Wyatt |
Artif. Intell. | 7 |
| 2016 | Building a human behavior map from local observationsabstractThis paper presents a novel method for classifying regions from human movements in service robots' working environments. The entire space is segmented subject to the class type according to the functionality or affordance of each place which accommodates a typical human behavior. This is achieved based on a grid map in two steps. First a probabilistic model is developed to capture human movements for each grid cell by using a non-ergodic HMM. Then the learned transition probabilities corresponding to these movements are used to cluster all cells by using the K-means algorithm. The knowledge of typical human movements for each location, represented by the prototypes from K-means and summarized in a ‘behavior-based map’, enables a robot to adjust the strategy of interacting with people according to where they are located, and thus greatly enhances its capability to assist people. The performance of the proposed classification method is demonstrated by experimental results from 8 hours of data that are collected in a kitchen environment. Patric Jensfelt, John Folkesson |
RO-MAN | 2 |
| 2015 | A Comparison of Qualitative and Metric Spatial Relation Models for Scene UnderstandingabstractObject recognition systems can be unreliable when run in isolation depending on only image based features, but their performance can be improved when taking scene context into account. In this paper, we present techniques to model and infer object labels in real scenes based on a variety of spatial relations — geometric features which capture how objects co-occur — and compare their efficacy in the context of augmenting perception based object classification in real-world table-top scenes. We utilise a long-term dataset of office table-tops for qualitatively comparing the performances of these techniques. On this dataset, we show that more intricate techniques, have a superior performance but do not generalise well on small training data. We also show that techniques using coarser information perform crudely but sufficiently well in standalone scenarios and generalise well on small training data. We conclude the paper, expanding on the insights we have gained through these comparisons and comment on a few fundamental topics with respect to long-term autonomous robots. Akshaya Thippur, Christopher Burbridge, Lars Kunze, Marina Alberti, John Folkesson, Patric Jensfelt, Nick Hawes |
AAAI | 6 |
| 2015 | Querying 3D Data by Adjacency Graphs
Nils Bore, Patric Jensfelt, John Folkesson |
ICVS | 2 |
| 2015 | Unsupervised learning of spatial-temporal models of objects in a long-term autonomy scenarioabstractWe present a novel method for clustering segmented dynamic parts of indoor RGB-D scenes across repeated observations by performing an analysis of their spatial-temporal distributions. We segment areas of interest in the scene using scene differencing for change detection. We extend the Meta-Room method and evaluate the performance on a complex dataset acquired autonomously by a mobile robot over a period of 30 days. We use an initial clustering method to group the segmented parts based on appearance and shape, and we further combine the clusters we obtain by analyzing their spatial-temporal behaviors. We show that using the spatial-temporal information further increases the matching accuracy. Rares Ambrus, Johan Ekekrantz, John Folkesson, Patric Jensfelt |
IROS | 4 |
| 2015 | Multi-scale conditional transition map: Modeling spatial-temporal dynamics of human movements with local and long-term correlationsabstractThis paper presents a novel approach to modeling the dynamics of human movements with a grid-based representation. The model we propose, termed as Multi-scale Conditional Transition Map (MCTMap), is an inhomogeneous HMM process that describes transitions of human location state in spatial and temporal space. Unlike existing work, our method is able to capture both local correlations and long-term dependencies on faraway initiating events. This enables the learned model to incorporate more information and to generate an informative representation of human existence probabilities across the grid map and along the temporal axis for intelligent interaction of the robot, such as avoiding or meeting the human. Our model consists of two levels. For each grid cell, we formulate the local dynamics using a variant of the left-to-right HMM, and thus explicitly model the exiting direction from the current cell. The dependency of this process on the entry direction is captured by employing the Input-Output HMM (IOHMM). On the higher level, we introduce the place where the whole trajectory originated into the IOHMM framework forming a hierarchical input structure to capture long-term dependencies. The capabilities of our method are verified by experimental results from 10 hours of data collected in an office corridor environment. Patric Jensfelt, John Folkesson |
IROS | 2 |
| 2014 | KTH-3D-TOTAL: A 3D dataset for discovering spatial structures for long-term autonomous learningabstractLong-term autonomous learning of human environments entails modelling and generalizing over distinct variations in: object instances in different scenes, and different scenes with respect to space and time. It is crucial for the robot to recognize the structure and context in spatial arrangements and exploit these to learn models which capture the essence of these distinct variations. Table-tops posses a typical structure repeatedly seen in human environments and are identified by characteristics of being personal spaces of diverse functionalities and dynamically changing due to human interactions. In this paper, we present a 3D dataset of 20 office table-tops manually observed and scanned 3 times a day as regularly as possible over 19 days (461 scenes) and subsequently, manually annotated with 18 different object classes, including multiple instances. We analyse the dataset to discover spatial structures and patterns in their variations. The dataset can, for example, be used to study the spatial relations between objects and long-term environment models for applications such as activity recognition, context and functionality estimation and anomaly detection. Akshaya Thippur, Rares Ambrus, Gaurav Agrawal, Adria Gallart del Burgo, Janardhan Haryadi Ramesh, Mayank Kumar Jha, Malepati Bala Siva Sai Akhil, Nishan Bhavanishankar Shetty, John Folkesson, Patric Jensfelt |
ICARCV | 10 |
| 2014 | Meta-rooms: Building and maintaining long term spatial models in a dynamic worldabstractWe present a novel method for re-creating the static structure of cluttered office environments - which we define as the “meta-room” - from multiple observations collected by an autonomous robot equipped with an RGB-D depth camera over extended periods of time. Our method works directly with point clusters by identifying what has changed from one observation to the next, removing the dynamic elements and at the same time adding previously occluded objects to reconstruct the underlying static structure as accurately as possible. The process of constructing the meta-rooms is iterative and it is designed to incorporate new data as it becomes available, as well as to be robust to environment changes. The latest estimate of the meta-room is used to differentiate and extract clusters of dynamic objects from observations. In addition, we present a method for re-identifying the extracted dynamic objects across observations thus mapping their spatial behaviour over extended periods of time. Rares Ambrus, Nils Bore, John Folkesson, Patric Jensfelt |
IROS | 4 |
| 2014 | Combining top-down spatial reasoning and bottom-up object class recognition for scene understandingabstractMany robot perception systems are built to only consider intrinsic object features to recognise the class of an object. By integrating both top-down spatial relational reasoning and bottom-up object class recognition the overall performance of a perception system can be improved. In this paper we present a unified framework that combines a 3D object class recognition system with learned, spatial models of object relations. In robot experiments we show that our combined approach improves the classification results on real world office desks compared to pure bottom-up perception. Hence, by using spatial knowledge during object class recognition perception becomes more efficient and robust and robots can understand scenes more effectively. Lars Kunze, Christopher Burbridge, Marina Alberti, Akshaya Thippur, John Folkesson, Patric Jensfelt, Nick Hawes |
IROS | 6 |
| 2014 | Modeling motion patterns of dynamic objects by IOHMMabstractThis paper presents a novel approach to model motion patterns of dynamic objects, such as people and vehicles, in the environment with the occupancy grid map representation. Corresponding to the ever-changing nature of the motion pattern of dynamic objects, we model each occupancy grid cell by an IOHMM, which is an inhomogeneous variant of the HMM. This distinguishes our work from existing methods which use the conventional HMM, assuming motion evolving according to a stationary process. By introducing observations of neighbor cells in the previous time step as input of IOHMM, the transition probabilities in our model are dependent on the occurrence of events in the cell's neighborhood. This enables our method to model the spatial correlation of dynamics across cells. A sequence processing example is used to illustrate the advantage of our model over conventional HMM based methods. Results from the experiments in an office corridor environment demonstrate that our method is capable of capturing dynamics of such human living environments. Rares Ambrus, Patric Jensfelt, John Folkesson |
IROS | 3 |
| 2013 | Active Visual Object Search in Unknown Environments Using Uncertain SemanticsabstractIn this paper, we study the problem of active visual search (AVS) in large, unknown, or partially known environments. We argue that by making use of uncertain semantics of the environment, a robot tasked with finding an object can devise efficient search strategies that can locate everyday objects at the scale of an entire building floor, which is previously unknown to the robot. To realize this, we present a probabilistic model of the search environment, which allows for prioritizing the search effort to those parts of the environment that are most promising for a specific object type. Further, we describe a method for reasoning about the unexplored part of the environment for goal-directed exploration with the purpose of object search. We demonstrate the validity of our approach by comparing it with two other search systems in terms of search trajectory length and time. First, we implement a greedy coverage-based search strategy that is found in previous work. Second, we let human participants search for objects as an alternative comparison for our method. Our results show that AVS strategies that exploit uncertain semantics of the environment are a very promising idea, and our method pushes the state-of-the-art forward in AVS. Alper Aydemir, Andrzej Pronobis, Moritz Göbelbecker, Patric Jensfelt |
IEEE Trans. Robotics | 4 |
| 2012 | Large-scale semantic mapping and reasoning with heterogeneous modalitiesabstractThis paper presents a probabilistic framework combining heterogeneous, uncertain, information such as object observations, shape, size, appearance of rooms and human input for semantic mapping. It abstracts multi-modal sensory information and integrates it with conceptual common-sense knowledge in a fully probabilistic fashion. It relies on the concept of spatial properties which make the semantic map more descriptive, and the system more scalable and better adapted for human interaction. A probabilistic graphical model, a chaingraph, is used to represent the conceptual information and perform spatial reasoning. Experimental results from online system tests in a large unstructured office environment highlight the system's ability to infer semantic room categories, predict existence of objects and values of other spatial properties as well as reason about unexplored space. Andrzej Pronobis, Patric Jensfelt |
ICRA | 2 |
| 2012 | Exploiting and modeling local 3D structure for predicting object locationsabstractIn this paper, we argue that there is a strong correlation between local 3D structure and object placement in everyday scenes. We call this the 3D context of the object. In previous work, this is typically hand-coded and limited to flat horizontal surfaces. In contrast, we propose to use a more general model for 3D context and learn the relationship between 3D context and different object classes. This way, we can capture more complex 3D contexts without implementing specialized routines. We present extensive experiments with both qualitative and quantitative evaluations of our method for different object classes. We show that our method can be used in conjunction with an object detection algorithm to reduce the rate of false positives. Our results support that the 3D structure surrounding objects in everyday scenes is a strong indicator of their placement and that it can give significant improvements in the performance of, for example, an object detection system. For evaluation, we have collected a large dataset of Microsoft Kinect frames from five different locations, which we also make publicly available. Alper Aydemir, Patric Jensfelt |
IROS | 2 |
| 2012 | What can we learn from 38, 000 rooms? Reasoning about unexplored space in indoor environmentsabstractMany robotics tasks require the robot to predict what lies in the unexplored part of the environment. Although much work focuses on building autonomous robots that operate indoors, indoor environments are neither well understood nor analyzed enough in the literature. In this paper, we propose and compare two methods for predicting both the topology and the categories of rooms given a partial map. The methods are motivated by the analysis of two large annotated floor plan data sets corresponding to the buildings of the MIT and KTH campuses. In particular, utilizing graph theory, we discover that local complexity remains unchanged for growing global complexity in real-world indoor environments, a property which we exploit. In total, we analyze 197 buildings, 940 floors and over 38,000 real-world rooms. Such a large set of indoor places has not been investigated before in the previous work. We provide extensive experimental results and show the degree of transferability of spatial knowledge between two geographically distinct locations. We also contribute the KTH data set and the software tools to with it. Alper Aydemir, Patric Jensfelt, John Folkesson |
IROS | 2 |
| 2011 | Search in the real world: Active visual object search based on spatial relationsabstractObjects are integral to a robot's understanding of space. Various tasks such as semantic mapping, pick-and-carry missions or manipulation involve interaction with objects. Previous work in the field largely builds on the assumption that the object in question starts out within the ready sensory reach of the robot. In this work we aim to relax this assumption by providing the means to perform robust and large-scale active visual object search. Presenting spatial relations that describe topological relationships between objects, we then show how to use these to create potential search actions. We introduce a method for efficiently selecting search strategies given probabilities for those relations. Finally we perform experiments to verify the feasibility of our approach. Alper Aydemir, Kristoffer Sjöö, John Folkesson, Andrzej Pronobis, Patric Jensfelt |
ICRA | 5 |
| 2011 | Home alone: Autonomous extension and correction of spatial representationsabstractIn this paper we present an account of the problems faced by a mobile robot given an incomplete tour of an unknown environment, and introduce a collection of techniques which can generate successful behaviour even in the presence of such problems. Underlying our approach is the principle that an autonomous system must be motivated to act to gather new knowledge, and to validate and correct existing knowledge. This principle is embodied in Dora, a mobile robot which features the aforementioned techniques: shared representations, non-monotonic reasoning, and goal generation and management. To demonstrate how well this collection of techniques work in real-world situations we present a comprehensive analysis of the Dora system's performance over multiple tours in an in door environment. In this analysis Dora successfully completed 18 of 21 attempted runs, with all but 3 of these successes requiring one or more of the integrated techniques to recover from problems. Nick Hawes, Marc Hanheide, Jack Hargreaves, Ben Page, Hendrik Zender, Patric Jensfelt |
ICRA | 6 |
| 2011 | Learning spatial relations from functional simulationabstractRobots acting in complex environments need not only be aware of objects, but also of the relationships objects have with each other. This paper suggests a conceptualization of these relationships in terms of task-relevant functional distinctions, such as support, location control, protection and confinement. Being able to discern such relations in a scene will be important for robots in practical tasks; accordingly, it is demonstrated how predictive models can be trained using data from physics simulations. The resulting models are shown to be both highly predictive and intuitively reasonable. Kristoffer Sjöö, Patric Jensfelt |
IROS | 2 |
| 2010 | Global robot localization with random finite set statistics
Adrian N. Bishop, Patric Jensfelt |
FUSION | 2 |
| 2010 | Simultaneous object class and pose estimation for mobile robotic applications with minimalistic recognitionabstractIn this paper we address the problem of simultaneous object class and pose estimation using nothing more than object class label measurements from a generic object classifier. We detail a method for designing a likelihood function over the robot configuration space. This function provides a likelihood measure of an object being of a certain class given that the robot (from some position) sees and recognizes an object as being of some (possibly different) class. Using this likelihood function in a recursive Bayesian framework allows us to achieve a kind of spatial averaging and determine the object pose (up to certain ambiguities to be made precise). We show how inter-class confusion from certain robot viewpoints can actually increase the ability to determine the object pose. Our approach is motivated by the idea of minimalistic sensing since we use only class label measurements albeit we attempt to estimate the object pose in addition to the class. Alper Aydemir, Adrian N. Bishop, Patric Jensfelt |
ICRA | 3 |
| 2010 | Mechanical support as a spatial abstraction for mobile robotsabstractMotivated by functional interpretations of spatial language terms, and the need for cognitively plausible and practical abstractions for mobile service robots, we present a spatial representation based on the physical support of one object by another, corresponding to the preposition “on”. A perceptual model for evaluating this relation is suggested, and experiments - simulated as well as using a real robot - are presented. We indicate how this model can be used for important tasks such as communication of spatial knowledge, abstract reasoning and learning, taking as an example direct and indirect visual search. We also demonstrate the model experimentally and show that it produces intuitively feasible results from visual scene analysis as well as synthetic distributions that can be put to a number of uses. Kristoffer Sjöö, Alper Aydemir, Thomas Morwald, Patric Jensfelt |
IROS | 5 |
| 2009 | A stochastically stable solution to the problem of robocentric mappingabstractThis paper provides a novel solution for robocentric mapping using an autonomous mobile robot. The robot dynamic model is the standard unicycle model and the robot is assumed to measure both the range and relative bearing to the landmarks. The algorithm introduced in this paper relies on a coordinate transformation and an extended Kalman filter like algorithm. The coordinate transformation considered in this paper has not been previously considered for robocentric mapping applications. Moreover, we provide a rigorous stochastic stability analysis of the filter employed and we examine the conditions under which the mean-square estimation error converges to a steady-state value. Adrian N. Bishop, Patric Jensfelt |
ICRA | 2 |
| 2008 | Active gaze control for attentional visual SLAMabstractIn this paper, we introduce an approach to active camera control for visual SLAM. Features, detected by a biologically motivated attention system, are tracked over several frames to determine stable landmarks. Matching of features to database entries enables global loop closing. The focus of this paper is the active camera control module, which supports the system with three behaviours: (i) A tracking behaviour tracks promising landmarks and prevents them from leaving the field of view, (ii) A redetection behaviour directs the camera actively to regions where landmarks are expected and thus supports loop closing, (iii) Finally, an exploration behaviour investigates regions without landmarks and enables a more uniform distribution of landmarks. Several real-world experiments show that the active camera control outperforms the passive system considerably. Simone Frintrop, Patric Jensfelt |
ICRA | 2 |
| 2008 | Hybrid laser and vision based object search and localizationabstractWe describe a method for an autonomous robot to efficiently locate one or more distinct objects in a realistic environment using monocular vision. We demonstrate how to efficiently subdivide acquired images into interest regions for the robot to zoom in on, using receptive field cooccurrence histograms. Objects are recognized through SIFT feature matching and the positions of the objects are estimated. Assuming a 2D map of the robot's surroundings and a set of navigation nodes between which it is free to move, we show how to compute an efficient sensing plan that allows the robot's camera to cover the environment, while obeying restrictions on the different objects' maximum and minimum viewing distances. The approach has been implemented on a real robotic system and results are presented showing its practicability and the quality of the position estimates obtained. Dorian Gálvez-López, Kristoffer Sjöö, Chandana Paul, Patric Jensfelt |
ICRA | 4 |
| 2008 | Towards robust place recognition for robot localizationabstractLocalization and context interpretation are two key competences for mobile robot systems. Visual place recognition, as opposed to purely geometrical models, holds promise of higher flexibility and association of semantics to the model. Ideally, a place recognition algorithm should be robust to dynamic changes and it should perform consistently when recognizing a room (for instance a corridor) in different geographical locations. Also, it should be able to categorize places, a crucial capability for transfer of knowledge and continuous learning. In order to test the suitability of visual recognition algorithms for these tasks, this paper presents a new database, acquired in three different labs across Europe. It contains image sequences of several rooms under dynamic changes, acquired at the same time with a perspective and omnidirectional camera, mounted on a socket. We assess this new database with an appearance- based algorithm that combines local features with support vector machines through an ad-hoc kernel. Results show the effectiveness of the approach and the value of the database. Muhammad Muneeb Ullah, Andrzej Pronobis, Barbara Caputo, Jie Luo 0018, Patric Jensfelt, Henrik I. Christensen |
ICRA | 5 |
| 2008 | Attentional Landmarks and Active Gaze Control for Visual SLAMabstractThis paper is centered around landmark detection, tracking, and matching for visual simultaneous localization and mapping using a monocular vision system with active gaze control. We present a system that specializes in creating and maintaining a sparse set of landmarks based on a biologically motivated feature-selection strategy. A visual attention system detects salient features that are highly discriminative and ideal candidates for visual landmarks that are easy to redetect. Features are tracked over several frames to determine stable landmarks and to estimate their 3-D position in the environment. Matching of current landmarks to database entries enables loop closing. Active gaze control allows us to overcome some of the limitations of using a monocular vision system with a relatively small field of view. It supports 1) the tracking of landmarks that enable a better pose estimation, 2) the exploration of regions without landmarks to obtain a better distribution of landmarks in the environment, and 3) the active redetection of landmarks to enable loop closing in situations in which a fixed camera fails to close the loop. Several real-world experiments show that accurate pose estimation is obtained with the presented system and that active camera control outperforms the passive approach. Simone Frintrop, Patric Jensfelt |
IEEE Trans. Robotics | 2 |
| 2007 | An Integrated Robotic System for Spatial Understanding and Situated Interaction in Indoor Environments
Hendrik Zender, Patric Jensfelt, Óscar Martínez Mozos, Geert-Jan M. Kruijff, Wolfram Burgard |
AAAI | 2 |
| 2007 | EKF SLAM updates in O(n) with Divide and Conquer SLAMabstractIn this paper we describe divide and conquer SLAM (D&C SLAM), an algorithm for performing simultaneous localization and mapping using the extended Kalman filter. D&C SLAM overcomes the two fundamental limitations of standard EKF SLAM: 1.) the computational cost per step is reduced from O(n2) to O(n) (the cost full SLAM is reduced from O(n3) to O(n2)); 2.) the resulting vehicle and map estimates have better consistency properties than standard EKF SLAM in the sense that the computed state covariance adequately represents the real error in the estimation. Unlike many current large scale EKF SLAM techniques, this algorithm computes an exact solution, without relying on approximations or simplifications to reduce computational complexity. Also, estimates and covariances are available when needed by data association without any further computation. Empirical results show that, as a bi-product of reduced computations, and without losing precision because of approximations, D&C SLAM has better consistency properties than standard EKF SLAM. Both characteristics allow to extend the range of environments that can be mapped in real time using EKF. We describe the algorithm and study its computational cost and consistency properties. Lina María Paz, Patric Jensfelt, Juan D. Tardós, José Neira |
ICRA | 2 |
| 2007 | Incremental learning for place recognition in dynamic environmentsabstractVision-based place recognition is a desirable feature for an autonomous mobile system. In order to work in realistic scenarios, visual recognition algorithms should be adaptive, i.e. should be able to learn from experience and adapt continuously to changes in the environment. This paper presents a discriminative incremental learning approach to place recognition. We use a recently introduced version of the incremental SVM, which allows to control the memory requirements as the system updates its internal representation. At the same time, it preserves the recognition performance of the batch algorithm. In order to assess the method, we acquired a database capturing the intrinsic variability of places over time. Extensive experiments show the power and the potential of the approach. Jie Luo 0018, Andrzej Pronobis, Barbara Caputo, Patric Jensfelt |
IROS | 4 |
| 2007 | Human- and Situation-Aware People FollowingabstractThe paper presents an approach to intelligent, interactive people following for autonomous robots. The approach combines robust methods for simultaneous localization and mapping and for people tracking in order to yield a socially and environmentally sensitive people following behavior. Unlike current purely reactive approaches ("nearest point following") it enables the robot to follow a human in a socially acceptable way, providing verbal and non-verbal feedback to the user where necessary. At the same time, the robot makes use of information about the spatial and functional organization of its environment, so that it can anticipate likely actions performed by a human, and adjust its motion accordingly. As a result, the robot's behaviors become less reactive and more intuitive when following people around an indoor environment. The approach has been fully implemented and tested. Hendrik Zender, Patric Jensfelt, Geert-Jan M. Kruijff |
RO-MAN | 2 |
| 2007 | The M-Space Feature Representation for SLAMabstractIn this paper, a new feature representation for simultaneous localization and mapping (SLAM) is discussed. The representation addresses feature symmetries and constraints explicitly to make the basic model numerically robust. In previous SLAM work, complete initialization of features is typically performed prior to introduction of a new feature into the map. This results in delayed use of new data. To allow early use of sensory data, the new feature representation addresses the use of features that initially have been partially observed. This is achieved by explicitly modelling the subspace of a feature that has been observed. In addition to accounting for the special properties of each feature type, the commonalities can be exploited in the new representation to create a feature framework that allows for interchanging of SLAM algorithms, sensor and features. Experimental results are presented using a low-cost Web-cam, a laser range scanner, and combinations thereof. John Folkesson, Patric Jensfelt, Henrik I. Christensen |
IEEE Trans. Robotics | 2 |
| 2006 | Clarification dialogues in human-augmented mappingabstractAn approach to dialogue based interaction for resolution of ambiguities encountered as part of Human-Augmented Mapping (HAM) is presented. The paper focuses on issues related to spatial organisation and localisation. The dialogue pattern naturally arises as robots are introduced to novel environments. The paper discusses an approach based on the notion of Questions under Discussion (QUD). The presented approach has been implemented on a mobile platform that has dialogue capabilities and methods for metric SLAM. Experimental results from a pilot study clearly demonstrate that the system can resolve problematic situations. Geert-Jan M. Kruijff, Hendrik Zender, Patric Jensfelt, Henrik I. Christensen |
HRI | 3 |
| 2006 | A Framework for Vision Based bearing only 3D SLAMabstractThis paper presents a framework for 3D vision based bearing only SLAM using a single camera, an interesting setup for many real applications due to its low cost. The focus in is on the management of the features to achieve real-time performance in extraction, matching and loop detection. For matching image features to map landmarks a modified, rotationally variant SIFT descriptor is used in combination with a Harris-Laplace detector. To reduce the complexity in the map estimation while maintaining matching performance only a few, high quality, image features are used for map landmarks. The rest of the features are used for matching. The framework has been combined with an EKF implementation for SLAM. Experiments performed in indoor environments are presented. These experiments demonstrate the validity and effectiveness of the approach. In particular they show how the robot is able to successfully match current image features to the map when revisiting an area Patric Jensfelt, Danica Kragic, John Folkesson, Mårten Björkman |
ICRA | 1 |
| 2006 | Nonholonomic Epipolar Visual ServoingabstractA significant amount of work has been reported in the area of visual servoing during the last decade. However, most of the contributions are applied in cases of holonomic robots. More recently, the use of visual feedback for control of nonholonomic vehicles has been reported. Some of the examples are docking and parallel parking maneuvers of cars or vision-based stabilization of a mobile manipulator to a desired pose with respect to a target of interest. Still, many of the approaches are mostly interested in the control part of visual servoing loop considering very simple vision algorithms based on artificial markers. In this paper, we present an approach for nonholonomic visual servoing based on epipolar geometry. The method facilitates a classical teach-by-showing approach where a reference image is used to define the desired pose (position and orientation) of the robot. The major contribution of the paper is the design of the control law that considers nonholonomic constraints of the robot as well as the robust feature detection and matching process based on scale and rotation invariant image features. An extensive experimental evaluation has been performed in a realistic indoor setting and the results are summarized in the paper Gonzalo López-Nicolás, Carlos Sagüés, Josechu J. Guerrero, Danica Kragic, Patric Jensfelt |
ICRA | 5 |
| 2006 | Fault Detection for Mobile Robots using Redundant Positioning SystemsabstractReliable navigation is a very important part of an autonomous mobile robot system. This means for instance that the robot should not lose track of its position, even if unexpected events like wheel slip and collisions occur. The standard approach to this problem is to construct a navigation system that is robust in itself. This paper proposes that detecting faults can also be made outside the normal navigation system, as an additional fault detector. Besides increasing the robustness, a means for detecting deviations is obtained, which can be important for the rest of the robot system, for instance the top level planner. The method uses two or more sources of robot position estimates, and compares them to detect unexpected deviation without getting deceived by drift or different characteristics in the position systems it gets information from. Both relative and absolute position sources can be used, meaning that existing positioning systems already implemented can be used in the detector. For detection purposes, an extended Kalman filter is used in conjunction with a CUSUM test. The detector is able to not only detect faults, but also give an estimate of when the fault occurred, which is useful for doing fault recovery. The detector is easy to implement, as it requires no modification of existing systems. Also the computational demands are very low. The approach is implemented and demonstrated on a mobile robot, using odometry and a scan matcher as sources of position information. It is shown that the system is able to detect wheel slip in real-time Paul Sundvall, Patric Jensfelt |
ICRA | 2 |
| 2006 | SLAM using Visual Scan-Matching with Distinguishable 3D PointsabstractScan-matching based on data from a laser scanner is frequently used for mapping and localization. This paper presents an scan-matching approach based instead on visual information from a stereo system. The scale invariant feature transform (SIFT) is used together with epipolar constraints to get high matching precision between the stereo images. Calculating the 3D position of the corresponding points in the world results in a visual scan where each point has a descriptor attached to it. These descriptors can be used when matching scans acquired from different positions. Just like in the work with laser based scan matching a map can be defined as a set of reference scans and their corresponding acquisition point. In essence this reduces each visual scan that can consist of hundreds of points to a single entity for which only the corresponding robot pose has to be estimated in the map. This reduces the overall complexity of the map. The SIFT descriptor attached to each of the points in the reference allows for robust matching and detection of loop closing situations. The paper presents real-world experimental results from an indoor office environment Federico Bertolli, Patric Jensfelt, Henrik I. Christensen |
IROS | 2 |
| 2006 | Integrating Active Mobile Robot Object Recognition and SLAM in Natural EnvironmentsabstractLinking semantic and spatial information has become an important research area in robotics since, for robots interacting with humans and performing tasks in natural environments, it is of foremost importance to be able to reason beyond simple geometrical and spatial levels. In this paper, we consider this problem in a service robot scenario where a mobile robot autonomously navigates in a domestic environment, builds a map as it moves along, localizes its position in it, recognizes objects on its way and puts them in the map. The experimental evaluation is performed in a realistic setting where the main concentration is put on the synergy of object recognition and simultaneous localization and mapping systems Staffan Ekvall, Patric Jensfelt, Danica Kragic |
IROS | 2 |
| 2006 | Attentional Landmark Selection for Visual SLAMabstractIn this paper, we introduce a new method to automatically detect useful landmarks for visual SLAM. A biologically motivated attention system detects regions of interest which "pop-out" automatically due to strong contrasts and the uniqueness of features. This property makes the regions easily redetectable and thus they are useful candidates for visual landmarks. Matching based on scene prediction and feature similarity allows not only short-term tracking of the regions, but also redetection in loop closing situations. The paper demonstrates how regions are determined and how they are matched reliably. Various experimental results on real-world data show that the landmarks are useful with respect to be tracked in consecutive frames and to enable closing loops Simone Frintrop, Patric Jensfelt, Henrik I. Christensen |
IROS | 2 |
| 2006 | Design of an Office-Guide Robot for Social Interaction StudiesabstractIn this paper, the design of an office-guide robot for social interaction studies is presented. We are interested in studying the impact of passage behaviours in casual encounters. While the system offers assistance in locating the appropriate office that a visitor wants to reach, it is expected to engage in a passing behaviour to allow free passage for other persons that it may encounter. Through use of such an approach it is possible to study the effect of social interaction in a situation that is much more natural than out-of-context user studies. The system has been tested in an early evaluation phase when it worked for almost 7 hours. A total of 64 interactions with people were registered and 13 passage behaviors were performed to conclude that this framework can be successfully used for the evaluation of passing behaviors in natural contexts of operation Elena Pacchierotti, Henrik I. Christensen, Patric Jensfelt |
IROS | 3 |
| 2006 | A Discriminative Approach to Robust Visual Place RecognitionabstractAn important competence for a mobile robot system is the ability to localize and perform context interpretation. This is required to perform basic navigation and to facilitate local specific services. Usually localization is performed based on a purely geometric model. Through use of vision and place recognition a number of opportunities open up in terms of flexibility and association of semantics to the model. To achieve this we present an appearance based method for place recognition. The method is based on a large margin classifier in combination with a rich global image descriptor. The method is robust to variations in illumination and minor scene changes. The method is evaluated across several different cameras, changes in time-of-day and weather conditions. The results clearly demonstrate the value of the approach. Andrzej Pronobis, Barbara Caputo, Patric Jensfelt, Henrik I. Christensen |
IROS | 3 |
| 2006 | A Discriminative Approach to Robust Visual Place RecognitionabstractAn important competence for a mobile robot system is the ability to localize and perform context interpretation. This is required to perform basic navigation and to facilitate local specific services. Usually localization is performed based on a purely geometric model. Through use of vision and place recognition a number of opportunities open up in terms of flexibility and association of semantics to the model. To achieve this, the present paper presents an appearance based method for place recognition. The method is based on a large margin classifier in combination with a rich global image descriptor. The method is robust to variations in illumination and minor scene changes. The method is evaluated across several different cameras, changes in time-of-day and weather conditions. The results clearly demonstrate the value of the approach Andrzej Pronobis, Barbara Caputo, Patric Jensfelt, Henrik I. Christensen |
IROS | 3 |
| 2006 | Augmenting SLAM with Object Detection in a Service Robot FrameworkabstractIn a service robot scenario, we are interested in a task of building maps of the environment that include automatically recognized objects. Most systems for simultaneous localization and mapping (SLAM) build maps that are only used for localizing the robot. Such maps are typically based on grids or different types of features such as point and lines. Here, we augment the process with an object recognition system that detects objects in the environment and puts them in the map generated by the SLAM system. During task execution, the robot can use this information to reason about objects, places and their relationships. The metric map is also split into topological entities corresponding to rooms. In this way, the user can command the robot to retrieve an object from a particular room or get help from a robot when searching for a certain object Patric Jensfelt, Staffan Ekvall, Danica Kragic, Daniel Aarno |
RO-MAN | 1 |
| 2006 | Situated dialogue and understanding spatial organization: Knowing what is where and what you can do thereabstractThe paper presents an HRI architecture for human-augmented mapping. Through interaction with a human, the robot can augment its autonomously learnt metric map with qualitative information about locations and objects in the environment. The system implements various interaction strategies observed in independent Wizard-of-Oz studies. The paper discusses an ontology-based approach to representing and inferring 2.5D spatial organization, and presents how knowledge of spatial organization can be acquired autonomously or through spoken dialogue interaction Geert-Jan M. Kruijff, Hendrik Zender, Patric Jensfelt, Henrik I. Christensen |
RO-MAN | 3 |
| 2006 | Evaluation of Passing Distance for Social RobotsabstractCasual encounters with mobile robots for nonexperts can be a challenge due to lack of an interaction model. The present work is based on the rules from proxemics which are used to design a passing strategy. In narrow corridors the lateral distance of passage is a key parameter to consider. An implemented system has been used in a small study to verify the basic parametric design for such a system. In total 10 subjects evaluated variations in proxemics for encounters with a robot in a corridor setting. The user feedback indicates that entering the intimate sphere of people is less comfortable, however a too significant avoidance is also considered unnecessary. Adequate signaling of avoidance is a behaviour that must be carefully tuned Elena Pacchierotti, Henrik I. Christensen, Patric Jensfelt |
RO-MAN | 3 |
| 2005 | Vision SLAM in the Measurement SubspaceabstractIn this paper we describe an approach to feature representation for simultaneous localization and mapping, SLAM. It is a general representation for features that addresses symmetries and constraints in the feature coordinates. Furthermore, the representation allows for the features to be added to the map with partial initialization. This is an important property when using oriented vision features where angle information can be used before their full pose is known. The number of the dimensions for a feature can grow with time as more information is acquired. At the same time as the special properties of each type of feature are accounted for, the commonalities of all map features are also exploited to allow SLAM algorithms to be interchanged as well as choice of sensors and features. In other words the SLAM implementation need not be changed at all when changing sensors and features and vice versa. Experimental results both with vision and range data and combinations thereof are presented. John Folkesson, Patric Jensfelt, Henrik I. Christensen |
ICRA | 2 |
| 2005 | Graphical SLAM using vision and the measurement subspaceabstractIn this paper we combine a graphical approach for simultaneous localization and mapping, SLAM, with a feature representation that addresses symmetries and constraints in the feature coordinates, the measurement subspace, M-space. The graphical method has the advantages of delayed linearizations and soft commitment to feature measurement matching. It also allows large maps to be built up as a network of small local patches, star nodes. This local map net is then easier to work with. The formation of the star nodes is explicitly stable and invariant with all the symmetries of the original measurements. All linearization errors are kept small by using a local frame. The construction of this invariant star is made clearer by the M-space feature representation. The M-space allows the symmetries and constraints of the measurements to be explicitly represented. We present results using both vision and laser sensors. John Folkesson, Patric Jensfelt, Henrik I. Christensen |
IROS | 2 |
| 2004 | An Interactive Interface for Service RobotsabstractIn this paper, we present an initial design of an interactive interface for a service robot based on multisensor fusion. We show how the integration of speech, vision and laser range data can be performed using a high level of abstraction. Guided by a number of scenarios commonly used in a service robot framework, the experimental evaluation will show the benefit of sensory integration which allows the design of a robust and natural interaction system using a set of simple perceptual algorithms. Elin Anna Topp, Danica Kragic, Patric Jensfelt, Henrik I. Christensen |
ICRA | 3 |
| 2003 | An experimental comparison of localisation methods, the MHL sessionsabstractIn this paper we compare multi hypothesis localisation (MHL)-which is a mobile robot localisation method based on multi hypothesis tracking - with six other methods reported in the literature. The comparison is performed using a standard set of test data and corresponding evaluation tools, thus facilitating a direct comparison of the obtained results. The experiments show that MHL compares favourably to all other methods in terms of recovering when the robot has been kidnapped. When using a validation gate for filtering out noisy measurements, MHL and the standard extended Kalman filter both perform as well as all other reported methods in terms of accuracy while being faster to compute. Steen Kristensen, Patric Jensfelt |
IROS | 2 |
| 2002 | Systems Integration for Real-World Manipulation TasksabstractA system developed to demonstrate integration of a number of key research areas such as localization, recognition, visual tracking, visual servoing and grasping is presented together with the underlying methodology adopted to facilitate the integration. Through sequencing of basic skills, provided by the above mentioned competencies, the system has the potential to carry out flexible grasping for fetch and carry in realistic environments. Through careful fusion of reactive and deliberative control and use of multiple sensory modalities a significant flexibility is achieved. Experimental verification of the integrated system is presented. Lars Petersson, Patric Jensfelt, Dennis Tell, M. Strandberg, Danica Kragic, Henrik I. Christensen |
ICRA | 2 |
| 2001 | Pose tracking using laser scanning and minimalistic environmental modelsabstractKeeping track of the position and orientation over time using sensor data, i.e., pose tracking, is a central component in many mobile robot systems. In this paper, we present a Kalman filter-based approach utilizing a minimalistic environmental model. By continuously updating the pose, matching the sensor data to the model is straightforward and outliers can be filtered out effectively by validation gates. The minimalistic model paves the way for a low-complexity algorithm with a high degree of robustness and accuracy. Robustness here refers both to being able to track the pose for a long time, but also handling changes and clutter in the environment. This robustness is gained by the minimalistic model only capturing the stable and large scale features of the environment. The effectiveness of the pose tracking is demonstrated through a number of experiments, including a run of 90 min., which clearly establishes the robustness of the method. Patric Jensfelt, Henrik I. Christensen |
IEEE Trans. Robotics Autom. | 1 |
| 2001 | Active global localization for a mobile robot using multiple hypothesis trackingabstractWe present a probabilistic approach for mobile robot localization using an incomplete topological world model. The method, called the multi-hypothesis localization (MHL), uses multi-hypothesis Kalman filter based pose tracking combined with a probabilistic formulation of hypothesis correctness to generate and track Gaussian pose hypotheses online. Apart from a lower computational complexity, this approach has the advantage over traditional grid based methods that incomplete and topological world model information can be utilized. Furthermore, the method generates movement commands for the platform to enhance the gathering of information for the pose estimation process. Extensive experiments are presented from two different environments, a typical office environment and an old hospital building. Patric Jensfelt, Steen Kristensen |
IEEE Trans. Robotics Autom. | 1 |
| 2000 | Using Multiple Gaussian Hypotheses to Represent Probability Distributions for Mobile Robot LocalizationabstractA new mobile robot localization technique is presented which uses multiple Gaussian hypotheses to represent the probability distribution of the robot location in the environment. Sensor data is assumed to be provided in the form of a Gaussian distribution over the space of robot poses. A tree of hypotheses is built, representing the possible data association histories for the system. Covariance intersection is used for the fusion of the Gaussians whenever a data association decision is taken. However, such a tree can grow without bound and so rules are introduced for the elimination of the least likely hypotheses from the tree and for the proper re-distribution of their probabilities. This technique is applied to a feature-based mobile robot localization scheme and experimental results are given demonstrating the effectiveness of the scheme. David J. Austin, Patric Jensfelt |
ICRA | 2 |
| 2000 | Feature Based Condensation for Mobile Robot LocalizationabstractMuch attention has been given to CONDENSATION methods for mobile robot localization. This has resulted in somewhat of a breakthrough in representing uncertainty for mobile robots. In this paper we use CONDENSATION with planned sampling as a tool for doing feature based global localization in a large and semi-structured environment. This paper presents a comparison of four different feature types: sonar based triangulation points and point pairs, as well as lines and doors extracted using a laser scanner. We show experimental results that highlight the information content of the different features, and point to fruitful combinations. Accuracy, computation time and the ability to narrow down the search space are among the measures used to compare the features. From the comparison of the features, some general guidelines are drawn for determining good feature types. Patric Jensfelt, David J. Austin, Olle Wijk |
ICRA | 1 |
| 2000 | Experiments on Augmenting Condensation for Mobile Robot LocalizationabstractWe study some modifications of the CONDENSATION algorithm. The case studied is feature based mobile robot localization in a large scale environment. The required sample set size for making the CONDENSATION algorithm converge properly can in many cases require too much computation. To manage with a sample set size which in the normal case would cause the CONDENSATION algorithm to break down. We study two modifications. The first strategy, called "CONDENSATION with random sampling", takes part of the sample set and spreads it randomly over the environment the robot operates in. The second strategy, called "CONDENSATION with planned sampling", places part of the sample set at planned positions based on the detected features. From the experiments we conclude that the second strategy is the best and can reduce the sample set size by at feast a factor of 40. Patric Jensfelt, Olle Wijk, David J. Austin |
ICRA | 1 |
| 2000 | Active exploration for feature based global localizationabstractPresents an algorithm for active exploration of the environment by a mobile robot when performing global localization. During the localization process interesting regions for future exploration are selected based on already detected features and on the hypotheses generated by the localization algorithm. The localization process is improved by presenting it a richer set of features. The proposed algorithm provides highly robust global localization in real world environments with very low computational effort spent in finding exploration goal points. Experimental results are given, demonstrating the effectiveness of the algorithm in a number of different situations. Marco Seiz, Patric Jensfelt, Henrik I. Christensen |
IROS | 2 |
| 1999 | Laser Based Pose TrackingabstractThe trend in localization is towards using more and more detailed models of the world. Our aim is to deal with the question of how simple a model can be used to provide and maintain pose information in an in-door setting. In this paper a Kalman filter based method for continuous position updating using a laser scanner is presented. By updating the position at a high frequency the matching problem becomes tractable and outliers can effectively be filtered out by means of validation gates. The experimental results presented show that the method performs very well in an in-door environment. Patric Jensfelt, Henrik I. Christensen |
ICRA | 1 |
| 1998 | Triangulation based Fusion of Ultrasonic Sensor DataabstractUltrasonic sensors are still one of the most widely used sensors in mobile robotics. A notorious problem in the use of sonar data is the lack of good spatial resolution, which typically results in a high uncertainty in the resulting map of the environment. In the paper a triangulation technique is used for filtering of data so as to obtain an improved grid map of the environment. The basic technique is described and it is outlined how it can be used for identification of natural landmarks. Olle Wijk, Patric Jensfelt, Henrik I. Christensen |
ICRA | 2 |