EDBT 2026 Demo / reviewers in the wild / expert
John Folkesson
dblp:83/5784
· DBLP profile ↗
41ranked-venue papers
9as first author
7since 2021 · last 2025
0000-0002-7796-1438ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 7 first-author · 7 since 2021Systems, architecture and hardware · 25 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Non-Myopic Layered Bayesian Optimization for Large-Scale Bathymetric Informative Path PlanningabstractInformative path planning (IPP) applied to bathy-metric mapping allows AUVs to focus on feature-rich areas to quickly reduce uncertainty and increase mapping efficiency. Existing methods based on Bayesian optimization (BO) over Gaussian Process (GP) maps work well on small scenarios but they are short-sighted and computationally heavy when mapping larger areas, hindering deployment in real applications. To overcome this, we present a 2-layered BO IPP method that performs non-myopic, online planning in a tree search fashion over large Stochastic Variational GP maps, while respecting the AUV dynamical constraints and accounting for localization uncertainty. Our framework outperforms the standard industrial lawn-mowing pattern and a myopic baseline in a set of hardware in the loop (HIL) experiments in an embedded platform over real bathymetry areas. Alexander Kiessling, Ignacio Torroba, Chelsea Sidrane, Ivan Stenius, Jana Tumova, John Folkesson |
ICRA | 6 |
| 2025 | Side Scan Sonar-based SLAM for Autonomous Algae Farm MonitoringabstractThe transition of seaweed farming to an alternative food source on an industrial scale relies on automating its processes through smart farming, equivalent to land agriculture. Key to this process are autonomous underwater vehicles (AUVs) via their capacity to automate crop and structural inspections. However, the current bottleneck for their deployment is ensuring safe navigation within farms, which requires an accurate, online estimate of the AUV pose and map of the infrastructure. To enable this, we propose an efficient side scan sonar-based (SSS) simultaneous localization and mapping (SLAM) framework that exploits the geometry of kelp farms via modeling structural ropes in the back-end as sequences of individual landmarks from each SSS ping detection, instead of combining detections into elongated representations. Our method outperforms state of the art solutions in hardware in the loop (HIL) experiments on a real AUV survey in a kelp farm. The framework and dataset can be found at https://github.com/julRusVal/sss_farm_slam. Julian Valdez, Ignacio Torroba, John Folkesson, Ivan Stenius |
IROS | 3 |
| 2024 | Boundary Factors for Seamless State Estimation between Autonomous Underwater Docking PhasesabstractAutonomous underwater docking is of the utmost importance for expanding the capabilities of Autonomous Underwater Vehicles (AUVs). Due to a historical focus on underwater docking to only static targets, the research gap in underwater docking to dynamically active targets has been left relatively untouched. We address the state estimation problem that arises when trying to rendezvous a chaser AUV with a dynamic target by modeling the scenario as a factor graph optimization-based Simultaneous Localization and Mapping problem. We present a set of boundary factors that aid the inference process by seamlessly transitioning the target’s state between the different observability stages, intrinsic to any dynamic docking scenario. We benchmark the performance of our approach using the Stonefish simulated environment. Aldo Terán Espinoza, Antonio Terán Espinoza, John Folkesson, Peter Sigray, Jakob Kuttenkeuler |
ICRA | 3 |
| 2024 | Benchmarking Classical and Learning-Based Multibeam Point Cloud RegistrationabstractDeep learning has shown promising results for multiple 3D point cloud registration datasets. However, in the underwater domain, most registration of multibeam echo-sounder (MBES) point cloud data are still performed using classical methods in the iterative closest point (ICP) family. In this work, we curate and release DotsonEast Dataset, a semi-synthetic MBES registration dataset constructed from an autonomous underwater vehicle in West Antarctica. Using this dataset, we systematically benchmark the performance of 2 classical and 4 learning-based methods. The experimental results show that the learning-based methods work well for coarse alignment, and are better at recovering rough transforms consistently at high overlap (20-50%). In comparison, GICP (a variant of ICP) performs well for fine alignment and is better across all metrics at extremely low overlap (10%). To the best of our knowledge, this is the first work to benchmark both learning-based and classical registration methods on an AUV-based MBES dataset. To facilitate future research, both the code and data are made available online.1 Jun Zhang 0102, Nils Bore, John Folkesson, Anna Wåhlin |
ICRA | 4 |
| 2024 | Hard Cases Detection in Motion Prediction by Vision-Language Foundation ModelsabstractAddressing hard cases in autonomous driving, such as anomalous road users, extreme weather conditions, and complex traffic interactions, presents significant challenges. To ensure safety, it is crucial to detect and manage these scenarios effectively for autonomous driving systems. However, the rarity and high-risk nature of these cases demand extensive, diverse datasets for training robust models. Vision-Language Foundation Models (VLMs) have shown remarkable zero-shot capabilities as being trained on extensive datasets. This work explores the potential of VLMs in detecting hard cases in autonomous driving. We demonstrate the capability of VLMs such as GPT-4v in detecting hard cases in traffic participant motion prediction on both agent and scenario levels. We introduce a feasible pipeline where VLMs, fed with sequential image frames with designed prompts, effectively identify challenging agents or scenarios, which are verified by existing prediction models. Moreover, by taking advantage of this detection of hard cases by VLMs, we further improve the training efficiency of the existing motion prediction pipeline by performing data selection for the training samples suggested by GPT. We show the effectiveness and feasibility of our pipeline incorporating VLMs with state-of-the-art methods on NuScenes datasets. The code is accessible at https://github.com/KTH-RPL/Detect_VLM. Yi Yang 0095, Qingwen Zhang, Kei Ikemura, Nazre Batool, John Folkesson |
IV | 5 |
| 2023 | Data-driven Loop Closure Detection in Bathymetric Point Clouds for Underwater SLAMabstractSimultaneous localization and mapping (SLAM) frameworks for autonomous navigation rely on robust data association to identify loop closures for back-end trajectory optimization. In the case of autonomous underwater vehicles (AUVs) equipped with multibeam echosounders (MBES), data association is particularly challenging due to the scarcity of identifiable landmarks in the seabed, the large drift in deadreckoning navigation estimates to which AUVs are prone and the low resolution characteristic of MBES data. Deep learning solutions to loop closure detection have shown excellent performance on data from more structured environments. However, their transfer to the seabed domain is not immediate and efforts to port them are hindered by the lack of bathymetric datasets. Thus, in this paper we propose a neural network architecture aimed to showcase the potential of adapting such techniques to correspondence matching in bathymetric data. We train our framework on real bathymetry from an AUV mission and evaluate its performance on the tasks of loop closure detection and coarse point cloud alignment. Finally, we show its potential against a more traditional method and release both its implementation and the dataset used. Jiarui Tan, Ignacio Torroba, Yiping Xie 0002, John Folkesson |
ICRA | 4 |
| 2021 | Interpretability in Contact-Rich Manipulation via Kinodynamic ImagesabstractDeep Neural Networks (NNs) have been widely utilized in contact-rich manipulation tasks to model the complicated contact dynamics. However, NN-based models are often difficult to decipher which can lead to seemingly inexplicable behaviors and unidentifiable failure cases. In this work, we address the interpretability of NN-based models by introducing the kinodynamic images. We propose a methodology that creates images from kinematic and dynamic data of contact-rich manipulation tasks. By using images as the state representation, we enable the application of interpretability modules that were previously limited to vision-based tasks. We use this representation to train a Convolutional Neural Network (CNN) and we extract interpretations with Grad-CAM to produce visual explanations. Our method is versatile and can be applied to any classification problem in manipulation tasks to visually interpret which parts of the input drive the model’s decisions and distinguish its failure modes, regardless of the features used. Our experiments demonstrate that our method enables detailed visual inspections of sequences in a task, and high-level evaluations of a model’s behavior. Code for this work is available at [1]. Ioanna Mitsioni, Joonatan Mänttäri, Yiannis Karayiannidis, John Folkesson, Danica Kragic |
ICRA | 4 |
| 2020 | Interpreting Video Features: A Comparison of 3D Convolutional Networks and Convolutional LSTM Networks
Joonatan Mänttäri, Sofia Broomé, John Folkesson, Hedvig Kjellström |
ACCV (5) | 3 |
| 2019 | Towards Autonomous Industrial-Scale Bathymetric SurveyingabstractBoth higher efficiency and cost reduction can be gained from automating bathymetric surveying for offshore applications such as pipeline, telecommunication or power cables installation and inspection on the seabed. We present a SLAM system that optimizes the geo-referencing of bathymetry surveys by fusing the dead-reckoning sensor data from the surveying vehicle with constraints from the maximization of the geometric consistency of overlapping regions of the survey. The framework has been extensively tested on bathymetric maps from both simulation and several actual industrial surveys and has proved robustness over different types of terrain. We demonstrate that our system is able to maximize the consistency of the final map even when there are large sections of the survey with reduced topographic variation. The framework has been made publicly available together with the simulation environment used to test it and some of the datasets. Ignacio Torroba, Nils Bore, John Folkesson |
IROS | 3 |
| 2019 | Incorporating Uncertainty in Predicting Vehicle Maneuvers at Intersections With Complex InteractionsabstractHighly automated driving systems are required to make robust decisions in many complex driving environments, such as urban intersections with high traffic. In order to make as informed and safe decisions as possible, it is necessary for the system to be able to predict the future maneuvers and positions of other traffic agents, as well as to provide information about the uncertainty in the prediction to the decision making module. While Bayesian approaches are a natural way of modeling uncertainty, recently deep learning-based methods have emerged to address this need as well. However, balancing the computational and system complexity, while also taking into account agent interactions and uncertainties, remains a difficult task. The work presented in this paper proposes a method of producing predictions of other traffic agents' trajectories in intersections with a singular Deep Learning module, while incorporating uncertainty and the interactions between traffic participants. The accuracy of the generated predictions is tested on a simulated intersection with a high level of interaction between agents, and different methods of incorporating uncertainty are compared. Preliminary results show that the CVAE-based method produces qualitatively and quantitatively better measurements of uncertainty and manage to more accurately assign probability to the future occupied space of traffic agents. Joonatan Mänttäri, John Folkesson |
IV | 2 |
| 2019 | Detection and Tracking of General Movable Objects in Large Three-Dimensional MapsabstractThis paper studies the problem of detection and tracking of general objects with semistatic dynamics observed by a mobile robot moving in a large environment. A key problem is that due to the environment scale, the robot can only observe a subset of the objects at any given time. Since some time passes between observations of objects in different places, the objects might be moved when the robot is not there. We propose a model for this movement in which the objects typically only move locally, but with some small probability they jump longer distances through what we call global motion. For filtering, we decompose the posterior over local and global movements into two linked processes. The posterior over the global movements and measurement associations is sampled, while we track the local movement analytically using Kalman filters. This novel filter is evaluated on point cloud data gathered autonomously by a mobile robot over an extended period of time. We show that tracking jumping objects is feasible, and that the proposed probabilistic treatment outperforms previous methods when applied to real world data. The key to efficient probabilistic tracking in this scenario is focused sampling of the object posteriors. Nils Bore, Johan Ekekrantz, Patric Jensfelt, John Folkesson |
IEEE Trans. Robotics | 4 |
| 2018 | Deep Reinforcement Learning to Acquire Navigation Skills for Wheel-Legged Robots in Complex EnvironmentsabstractMobile robot navigation in complex and dynamic environments is a challenging but important problem. Reinforcement learning approaches fail to solve these tasks efficiently due to reward sparsities, temporal complexities and high-dimensionality of sensorimotor spaces which are inherent in such problems. We present a novel approach to train action policies to acquire navigation skills for wheel-legged robots using deep reinforcement learning. The policy maps height-map image observations to motor commands to navigate to a target position while avoiding obstacles. We propose to acquire the multifaceted navigation skill by learning and exploiting a number of manageable navigation behaviors. We also introduce a domain randomization technique to improve the versatility of the training samples. We demonstrate experimentally a significant improvement in terms of data-efficiency, success rate, robustness against irrelevant sensory data, and also the quality of the maneuver skills. Xi Chen 0051, Ali Ghadirzadeh, John Folkesson, Mårten Björkman, Patric Jensfelt |
IROS | 3 |
| 2018 | Learning to Predict Lane Changes in Highway Scenarios Using Dynamic Filters On a Generic Traffic RepresentationabstractIn highway driving scenarios it is important for highly automated driving systems to be able to recognize and predict the intended maneuvers of other drivers in order to make robust and informed decisions. Many methods utilize the current kinematics of vehicles to make these predictions, but it is possible to examine the relations between vehicles as well to gain more information about the traffic scene and make more accurate predictions. The work presented in this paper proposes a novel method of predicting lane change maneuvers in highway scenarios using deep learning and a generic visual representation of the traffic scene. Experimental results suggest that by operating on the visual representation, the spacial relations between arbitrary vehicles can be captured by our method and used for more informed predictions without the need for explicit dynamic or driver interaction models. The proposed method is evaluated on highway driving scenarios using the Interstate-80 dataset and compared to a kinematics based prediction model, with results showing that the proposed method produces more robust predictions across the prediction horizon than the comparison model. Joonatan Mänttäri, John Folkesson, Erik Ward |
Intelligent Vehicles Symposium | 2 |
| 2018 | Towards Risk Minimizing Trajectory Planning in On-Road ScenariosabstractTrajectory planning for autonomous vehicles should attempt to minimize expected risk given noisy sensor data and uncertain predictions of the near future. In this paper, we present a trajectory planning approach for on-road scenarios where we use a graph search approximation. Uncertain predictions of other vehicles are accounted for by a novel inference technique that allows efficient calculation of the probability of dangerous outcomes for set of modeled situation types. Erik Ward, John Folkesson |
Intelligent Vehicles Symposium | 2 |
| 2017 | Autonomous meshing, texturing and recognition of object models with a mobile robotabstractWe present a system for creating object models from RGB-D views acquired autonomously by a mobile robot. We create high-quality textured meshes of the objects by approximating the underlying geometry with a Poisson surface. Our system employs two optimization steps, first registering the views spatially based on image features, and second aligning the RGB images to maximize photometric consistency with respect to the reconstructed mesh. We show that the resulting models can be used robustly for recognition by training a Convolutional Neural Network (CNN) on images rendered from the reconstructed meshes. We perform experiments on data collected autonomously by a mobile robot both in controlled and uncontrolled scenarios. We compare quantitatively and qualitatively to previous work to validate our approach. Rares Ambrus, Nils Bore, John Folkesson, Patric Jensfelt |
IROS | 3 |
| 2017 | Geometric and visual terrain classification for autonomous mobile navigationabstractIn this paper, we present a multi-sensory terrain classification algorithm with a generalized terrain representation using semantic and geometric features. We compute geometric features from lidar point clouds and extract pixel-wise semantic labels from a fully convolutional network that is trained using a dataset with a strong focus on urban navigation. We use data augmentation to overcome the biases of the original dataset and apply transfer learning to adapt the model to new semantic labels in off-road environments. Finally, we fuse the visual and geometric features using a random forest to classify the terrain traversability into three classes: safe, risky and obstacle. We implement the algorithm on our four-wheeled robot and test it in novel environments including both urban and off-road scenes which are distinct from the training environments and under summer and winter conditions. We provide experimental result to show that our algorithm can perform accurate and fast prediction of terrain traversability in a mixture of environments with a small set of training data. Fabian Schilling, Xi Chen 0051, John Folkesson, Patric Jensfelt |
IROS | 3 |
| 2017 | Human-centric partitioning of the environmentabstractIn this paper, we present an object based approach for human-centric partitioning of the environment. Our approach for determining the human-centric regions is to detect the objects that are commonly associated with frequent human presence. In order to detect these objects, we employ state of the art perception techniques. The detected objects are stored with their spatio-temporal information in the robot's memory to be later used for generating the regions. The advantages of our method is that it is autonomous, requires only a small set of perceptual data and does not even require people to be present while generating the regions. The generated regions are validated using a 1-month dataset collected in an indoor office environment. The experimental results show that although a small set of perceptual data is used, the regions are generated at densely occupied locations. Hakan Karaoguz, Nils Bore, John Folkesson, Patric Jensfelt |
RO-MAN | 3 |
| 2016 | Vehicle localization with low cost radar sensorsabstractAutonomous vehicles rely on GPS aided by motion sensors to localize globally within the road network. However, not all driving surfaces have satellite visibility. Therefore, it is important to augment these systems with localization based on environmental sensing such as cameras, lidar and radar in order to increase reliability and robustness. In this work we look at using radar for localization. Radar sensors are available in compact format devices well suited to automotive applications. Past work on localization using radar in automotive applications has been based on careful sensor modeling and Sequential Monte Carlo, (Particle) filtering. In this work we investigate the use of the Iterative Closest Point, ICP, algorithm together with an Extended Kalman filter, EKF, for localizing a vehicle equipped with automotive grade radars. Experiments using data acquired on public roads shows that this computationally simpler approach yields sufficiently accurate results on par with more complex methods. Erik Ward, John Folkesson |
Intelligent Vehicles Symposium | 2 |
| 2016 | Building a human behavior map from local observationsabstractThis paper presents a novel method for classifying regions from human movements in service robots' working environments. The entire space is segmented subject to the class type according to the functionality or affordance of each place which accommodates a typical human behavior. This is achieved based on a grid map in two steps. First a probabilistic model is developed to capture human movements for each grid cell by using a non-ergodic HMM. Then the learned transition probabilities corresponding to these movements are used to cluster all cells by using the K-means algorithm. The knowledge of typical human movements for each location, represented by the prototypes from K-means and summarized in a ‘behavior-based map’, enables a robot to adjust the strategy of interacting with people according to where they are located, and thus greatly enhances its capability to assist people. The performance of the proposed classification method is demonstrated by experimental results from 8 hours of data that are collected in a kitchen environment. Patric Jensfelt, John Folkesson |
RO-MAN | 3 |
| 2015 | A Comparison of Qualitative and Metric Spatial Relation Models for Scene UnderstandingabstractObject recognition systems can be unreliable when run in isolation depending on only image based features, but their performance can be improved when taking scene context into account. In this paper, we present techniques to model and infer object labels in real scenes based on a variety of spatial relations — geometric features which capture how objects co-occur — and compare their efficacy in the context of augmenting perception based object classification in real-world table-top scenes. We utilise a long-term dataset of office table-tops for qualitatively comparing the performances of these techniques. On this dataset, we show that more intricate techniques, have a superior performance but do not generalise well on small training data. We also show that techniques using coarser information perform crudely but sufficiently well in standalone scenarios and generalise well on small training data. We conclude the paper, expanding on the insights we have gained through these comparisons and comment on a few fundamental topics with respect to long-term autonomous robots. Akshaya Thippur, Christopher Burbridge, Lars Kunze, Marina Alberti, John Folkesson, Patric Jensfelt, Nick Hawes |
AAAI | 5 |
| 2015 | Unsupervised robot learning to predict person motionabstractSocially interacting robots will need to understand the intentions and recognize the behaviors of people they come in contact with. In this paper we look at how a robot can learn to recognize and predict people's intended path based on its own observations of people over time. Our approach uses people tracking on the robot from either RGBD cameras or LIDAR. The tracks are separated into homogeneous motion classes using a pre-trained SVM. Then the individual classes are clustered and prototypes are extracted from each cluster. These are then used to predict a person's future motion based on matching to a partial prototype and using the rest of the prototype as the predicted motion. Results from experiments in a kitchen environment in our lab demonstrate the capabilities of the proposed method. Shuang Xiao, John Folkesson |
ICRA | 3 |
| 2015 | Querying 3D Data by Adjacency Graphs
Nils Bore, Patric Jensfelt, John Folkesson |
ICVS | 3 |
| 2015 | Unsupervised learning of spatial-temporal models of objects in a long-term autonomy scenarioabstractWe present a novel method for clustering segmented dynamic parts of indoor RGB-D scenes across repeated observations by performing an analysis of their spatial-temporal distributions. We segment areas of interest in the scene using scene differencing for change detection. We extend the Meta-Room method and evaluate the performance on a complex dataset acquired autonomously by a mobile robot over a period of 30 days. We use an initial clustering method to group the segmented parts based on appearance and shape, and we further combine the clusters we obtain by analyzing their spatial-temporal behaviors. We show that using the spatial-temporal information further increases the matching accuracy. Rares Ambrus, Johan Ekekrantz, John Folkesson, Patric Jensfelt |
IROS | 3 |
| 2015 | Multi-scale conditional transition map: Modeling spatial-temporal dynamics of human movements with local and long-term correlationsabstractThis paper presents a novel approach to modeling the dynamics of human movements with a grid-based representation. The model we propose, termed as Multi-scale Conditional Transition Map (MCTMap), is an inhomogeneous HMM process that describes transitions of human location state in spatial and temporal space. Unlike existing work, our method is able to capture both local correlations and long-term dependencies on faraway initiating events. This enables the learned model to incorporate more information and to generate an informative representation of human existence probabilities across the grid map and along the temporal axis for intelligent interaction of the robot, such as avoiding or meeting the human. Our model consists of two levels. For each grid cell, we formulate the local dynamics using a variant of the left-to-right HMM, and thus explicitly model the exiting direction from the current cell. The dependency of this process on the entry direction is captured by employing the Input-Output HMM (IOHMM). On the higher level, we introduce the place where the whole trajectory originated into the IOHMM framework forming a hierarchical input structure to capture long-term dependencies. The capabilities of our method are verified by experimental results from 10 hours of data collected in an office corridor environment. Patric Jensfelt, John Folkesson |
IROS | 3 |
| 2014 | KTH-3D-TOTAL: A 3D dataset for discovering spatial structures for long-term autonomous learningabstractLong-term autonomous learning of human environments entails modelling and generalizing over distinct variations in: object instances in different scenes, and different scenes with respect to space and time. It is crucial for the robot to recognize the structure and context in spatial arrangements and exploit these to learn models which capture the essence of these distinct variations. Table-tops posses a typical structure repeatedly seen in human environments and are identified by characteristics of being personal spaces of diverse functionalities and dynamically changing due to human interactions. In this paper, we present a 3D dataset of 20 office table-tops manually observed and scanned 3 times a day as regularly as possible over 19 days (461 scenes) and subsequently, manually annotated with 18 different object classes, including multiple instances. We analyse the dataset to discover spatial structures and patterns in their variations. The dataset can, for example, be used to study the spatial relations between objects and long-term environment models for applications such as activity recognition, context and functionality estimation and anomaly detection. Akshaya Thippur, Rares Ambrus, Gaurav Agrawal, Adria Gallart del Burgo, Janardhan Haryadi Ramesh, Mayank Kumar Jha, Malepati Bala Siva Sai Akhil, Nishan Bhavanishankar Shetty, John Folkesson, Patric Jensfelt |
ICARCV | 9 |
| 2014 | Meta-rooms: Building and maintaining long term spatial models in a dynamic worldabstractWe present a novel method for re-creating the static structure of cluttered office environments - which we define as the “meta-room” - from multiple observations collected by an autonomous robot equipped with an RGB-D depth camera over extended periods of time. Our method works directly with point clusters by identifying what has changed from one observation to the next, removing the dynamic elements and at the same time adding previously occluded objects to reconstruct the underlying static structure as accurately as possible. The process of constructing the meta-rooms is iterative and it is designed to incorporate new data as it becomes available, as well as to be robust to environment changes. The latest estimate of the meta-room is used to differentiate and extract clusters of dynamic objects from observations. In addition, we present a method for re-identifying the extracted dynamic objects across observations thus mapping their spatial behaviour over extended periods of time. Rares Ambrus, Nils Bore, John Folkesson, Patric Jensfelt |
IROS | 3 |
| 2014 | Combining top-down spatial reasoning and bottom-up object class recognition for scene understandingabstractMany robot perception systems are built to only consider intrinsic object features to recognise the class of an object. By integrating both top-down spatial relational reasoning and bottom-up object class recognition the overall performance of a perception system can be improved. In this paper we present a unified framework that combines a 3D object class recognition system with learned, spatial models of object relations. In robot experiments we show that our combined approach improves the classification results on real world office desks compared to pure bottom-up perception. Hence, by using spatial knowledge during object class recognition perception becomes more efficient and robust and robots can understand scenes more effectively. Lars Kunze, Christopher Burbridge, Marina Alberti, Akshaya Thippur, John Folkesson, Patric Jensfelt, Nick Hawes |
IROS | 5 |
| 2014 | Modeling motion patterns of dynamic objects by IOHMMabstractThis paper presents a novel approach to model motion patterns of dynamic objects, such as people and vehicles, in the environment with the occupancy grid map representation. Corresponding to the ever-changing nature of the motion pattern of dynamic objects, we model each occupancy grid cell by an IOHMM, which is an inhomogeneous variant of the HMM. This distinguishes our work from existing methods which use the conventional HMM, assuming motion evolving according to a stationary process. By introducing observations of neighbor cells in the previous time step as input of IOHMM, the transition probabilities in our model are dependent on the occurrence of events in the cell's neighborhood. This enables our method to model the spatial correlation of dynamics across cells. A sequence processing example is used to illustrate the advantage of our model over conventional HMM based methods. Results from the experiments in an office corridor environment demonstrate that our method is capable of capturing dynamics of such human living environments. Rares Ambrus, Patric Jensfelt, John Folkesson |
IROS | 4 |
| 2012 | What can we learn from 38, 000 rooms? Reasoning about unexplored space in indoor environmentsabstractMany robotics tasks require the robot to predict what lies in the unexplored part of the environment. Although much work focuses on building autonomous robots that operate indoors, indoor environments are neither well understood nor analyzed enough in the literature. In this paper, we propose and compare two methods for predicting both the topology and the categories of rooms given a partial map. The methods are motivated by the analysis of two large annotated floor plan data sets corresponding to the buildings of the MIT and KTH campuses. In particular, utilizing graph theory, we discover that local complexity remains unchanged for growing global complexity in real-world indoor environments, a property which we exploit. In total, we analyze 197 buildings, 940 floors and over 38,000 real-world rooms. Such a large set of indoor places has not been investigated before in the previous work. We provide extensive experimental results and show the degree of transferability of spatial knowledge between two geographically distinct locations. We also contribute the KTH data set and the software tools to with it. Alper Aydemir, Patric Jensfelt, John Folkesson |
IROS | 3 |
| 2011 | Search in the real world: Active visual object search based on spatial relationsabstractObjects are integral to a robot's understanding of space. Various tasks such as semantic mapping, pick-and-carry missions or manipulation involve interaction with objects. Previous work in the field largely builds on the assumption that the object in question starts out within the ready sensory reach of the robot. In this work we aim to relax this assumption by providing the means to perform robust and large-scale active visual object search. Presenting spatial relations that describe topological relationships between objects, we then show how to use these to create potential search actions. We introduce a method for efficiently selecting search strategies given probabilities for those relations. Finally we perform experiments to verify the feasibility of our approach. Alper Aydemir, Kristoffer Sjöö, John Folkesson, Andrzej Pronobis, Patric Jensfelt |
ICRA | 3 |
| 2011 | The Antiparticle Filter - An Adaptive Nonlinear Estimator
John Folkesson |
ISRR | 1 |
| 2009 | Autonomy through SLAM for an Underwater Robot
John Folkesson, John J. Leonard |
ISRR | 1 |
| 2007 | Feature tracking for underwater navigation using sonarabstractTracking sonar features in real time on an underwater robot is a challenging task. One reason is the low observability of the sonar in some directions. For example, using a blazed array sonar one observes range and the angle to the array axis with fair precision. The angle around the axis is poorly constrained. This situation is problematic for tracking features in world frame Cartesian coordinates as the error surfaces will not be ellipsoids. Thus Gaussian tracking of the features will not work properly. The situation is similar to the problem of tracking features in camera images. There the unconstrained direction is depth and its errors are highly non-Gaussian. We propose a solution to the sonar problem that is analogous to the successful inverse depth feature parameterization for vision tracking, introduced by [1]. We parameterize the features by the robot pose where it was first seen and the range/bearing from that pose. Thus the 3D features have 9 parameters that specify their world coordinates. We use a nonlinear transformation on the poorly observed bearing angle to give a more accurate Gaussian approximation to the uncertainty. These features are tracked in a SLAM framework until there is enough information to initialize world frame Cartesian coordinates for them. The more compact representation can then be used for a global SLAM or localization purposes. We present results for a system running real time underwater SLAM/localization. These results show that the parameterization leads to greater consistency in the feature location estimates. John Folkesson, John J. Leonard, Jacques Leederkerken, Rob Williams |
IROS | 1 |
| 2007 | Closing the Loop With Graphical SLAMabstractThe problem of simultaneous localization and mapping (SLAM) is addressed using a graphical method. The main contributions are a computational complexity that scales well with the size of the environment, the elimination of most of the linearization inaccuracies, and a more flexible and robust data association. We also present a detection criteria for closing loops. We show how multiple topological constraints can be imposed on the graphical solution by a process of coarse fitting followed by fine tuning. The coarse fitting is performed using an approximate system. This approximate system can be shown to possess all the local symmetries. Observations made during the SLAM process often contain symmetries, that is to say, directions of change to the state space that do not affect the observed quantities. It is important that these directions do not shift as we approximate the system by, for example, linearization. The approximate system is both linear and block diagonal. This makes it a very simple system to work with especially when imposing global topological constraints on the solution. These global constraints are nonlinear. We show how these constraints can be discovered automatically. We develop a method of testing multiple hypotheses for data matching using the graph. This method is derived from statistical theory and only requires simple counting of observations. The central insight is to examine the probability of not observing the same features on a return to a region. We present results with data from an outdoor scenario using a SICK laser scanner. John Folkesson, Henrik I. Christensen |
IEEE Trans. Robotics | 1 |
| 2007 | The M-Space Feature Representation for SLAMabstractIn this paper, a new feature representation for simultaneous localization and mapping (SLAM) is discussed. The representation addresses feature symmetries and constraints explicitly to make the basic model numerically robust. In previous SLAM work, complete initialization of features is typically performed prior to introduction of a new feature into the map. This results in delayed use of new data. To allow early use of sensory data, the new feature representation addresses the use of features that initially have been partially observed. This is achieved by explicitly modelling the subspace of a feature that has been observed. In addition to accounting for the special properties of each feature type, the commonalities can be exploited in the new representation to create a feature framework that allows for interchanging of SLAM algorithms, sensor and features. Experimental results are presented using a low-cost Web-cam, a laser range scanner, and combinations thereof. John Folkesson, Patric Jensfelt, Henrik I. Christensen |
IEEE Trans. Robotics | 1 |
| 2006 | A Framework for Vision Based bearing only 3D SLAMabstractThis paper presents a framework for 3D vision based bearing only SLAM using a single camera, an interesting setup for many real applications due to its low cost. The focus in is on the management of the features to achieve real-time performance in extraction, matching and loop detection. For matching image features to map landmarks a modified, rotationally variant SIFT descriptor is used in combination with a Harris-Laplace detector. To reduce the complexity in the map estimation while maintaining matching performance only a few, high quality, image features are used for map landmarks. The rest of the features are used for matching. The framework has been combined with an EKF implementation for SLAM. Experiments performed in indoor environments are presented. These experiments demonstrate the validity and effectiveness of the approach. In particular they show how the robot is able to successfully match current image features to the map when revisiting an area Patric Jensfelt, Danica Kragic, John Folkesson, Mårten Björkman |
ICRA | 3 |
| 2005 | Vision SLAM in the Measurement SubspaceabstractIn this paper we describe an approach to feature representation for simultaneous localization and mapping, SLAM. It is a general representation for features that addresses symmetries and constraints in the feature coordinates. Furthermore, the representation allows for the features to be added to the map with partial initialization. This is an important property when using oriented vision features where angle information can be used before their full pose is known. The number of the dimensions for a feature can grow with time as more information is acquired. At the same time as the special properties of each type of feature are accounted for, the commonalities of all map features are also exploited to allow SLAM algorithms to be interchanged as well as choice of sensors and features. In other words the SLAM implementation need not be changed at all when changing sensors and features and vice versa. Experimental results both with vision and range data and combinations thereof are presented. John Folkesson, Patric Jensfelt, Henrik I. Christensen |
ICRA | 1 |
| 2005 | Graphical SLAM using vision and the measurement subspaceabstractIn this paper we combine a graphical approach for simultaneous localization and mapping, SLAM, with a feature representation that addresses symmetries and constraints in the feature coordinates, the measurement subspace, M-space. The graphical method has the advantages of delayed linearizations and soft commitment to feature measurement matching. It also allows large maps to be built up as a network of small local patches, star nodes. This local map net is then easier to work with. The formation of the star nodes is explicitly stable and invariant with all the symmetries of the original measurements. All linearization errors are kept small by using a local frame. The construction of this invariant star is made clearer by the M-space feature representation. The M-space allows the symmetries and constraints of the measurements to be explicitly represented. We present results using both vision and laser sensors. John Folkesson, Patric Jensfelt, Henrik I. Christensen |
IROS | 1 |
| 2004 | Graphical SLAM - a Self-correcting MapabstractWe describe an approach to simultaneous localization and mapping, SLAM. This approach has the highly desirable property of robustness to data association errors. Another important advantage of our algorithm is that non-linearities are computed exactly, so that global constraints can be imposed even if they result in large shifts to the map. We represent the map as a graph and use the graph to find an efficient map update algorithm. We also show how topological consistency can be imposed on the map, such as, closing a loop. The algorithm has been implemented on an outdoor robot and we have experimental validation of our ideas. We also explain how the graph can be simplified leading to linear approximations of sections of the map. This reduction gives us a natural way to connect local map patches into a much larger global map. John Folkesson, Henrik I. Christensen |
ICRA | 1 |
| 2003 | Outdoor exploration and SLAM using a compressed filterabstractIn this paper we describe the use of automatic explorationfor autonomous mapping of outdoor scenes. We describe areal-time SLAM implementation along with an autonomous explorationalgorithm. We have ... John Folkesson, Henrik I. Christensen |
ICRA | 1 |
| 2003 | PDA interface for a field robotabstractOperating robots in an outdoor setting poses interesting problems in terms of interaction. To interact with the robot there is a need for a flexible computer interface. In this paper a PDA-based (personal digital assistant, i.e. a handheld computer) approach to robot interaction is presented. The system is designed to allow non-expert users to utilise the robot for operation in an urban exploration setup. The basic design is outlined and a first set of experiments are reported. Carl Lundberg, Carl Barck-Holst, John Folkesson, Henrik I. Christensen |
IROS | 3 |