Lars Hammarstrand

dblp:82/9456 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0001-5676-1392ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving
abstract
Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly, autonomous driving generates vast amounts of spatiotemporal data, alluding to the possibility of harnessing scale to learn the underlying geometric and semantic structure of the environment and its evolution over time. In this direction, we propose a geometric and semantic self-supervised pre-training method, GASP, that learns a unified representation by predicting, at any queried future point in spacetime, (1) general occupancy, capturing the evolving structure of the 3D scene; (2) ego occupancy, modeling the ego vehicle path through the environment; and (3) distilled high-level features from a vision foundation model. By modeling geometric and semantic 4D occupancy fields instead of raw sensor measurements, the model learns a structured, generalizable representation of the environment and its evolution through time. We validate GASP on multiple autonomous driving benchmarks, demonstrating significant improvements in semantic occupancy forecasting, online mapping, and ego trajectory prediction. Our results demonstrate that continuous 4D geometric and semantic occupancy prediction provides a scalable and effective pre-training paradigm for autonomous driving. For code and additional visualizations, see our project page.
William Ljungbergh, Adam Lilja, Adam Tonderski, Arvid Laveno Ling, Carl Lindström, Willem Verbeke, Junsheng Fu, Christoffer Petersson, Lars Hammarstrand, Michael Felsberg
WACV9
2026 Semi-Supervised Hierarchical Open-Set Classification
abstract
Hierarchical open-set classification handles previously unseen classes by assigning them to the most appropriate high-level category in a class taxonomy. We extend this paradigm to the semi-supervised setting, enabling the use of large-scale, uncurated datasets containing a mixture of known and unknown classes to improve the hierarchical open-set performance. To this end, we propose a teacher-student framework based on pseudo-labeling. Two key components are introduced: 1) subtree pseudo-labels, which provide reliable supervision in the presence of unknown data, and 2) age-gating, a mechanism that mitigates overconfidence in pseudo-labels. Experiments show that our framework outperforms self-supervised pretraining followed by supervised adaptation, and even matches the fully supervised counterpart when using only 20 labeled samples per class on the iNaturalist19 benchmark. Our code is available at https://github.com/walline/semihoc.
Erik Wallin, Fredrik Kahl, Lars Hammarstrand
WACV3
2025 ProHOC: Probabilistic Hierarchical Out-of-Distribution Classification via Multi-Depth Networks
abstract
Out-of-distribution (OOD) detection in deep learning has traditionally been framed as a binary task, where samples are either classified as belonging to the known classes or marked as OOD, with little attention given to the semantic relationships between OOD samples and the in-distribution (ID) classes. We propose a framework for detecting and classifying OOD samples in a given class hierarchy. Specifically, we aim to predict OOD data to their correct internal nodes of the class hierarchy, whereas the known ID classes should be predicted as their corresponding leaf nodes. Our approach leverages the class hierarchy to create a probabilistic model and we implement this model by using networks trained for ID classification at multiple hierarchy depths. We conduct experiments on three datasets with predefined class hierarchies and show the effectiveness of our method. Our code is available at https://github.com/walline/prohoc.
Erik Wallin, Fredrik Kahl, Lars Hammarstrand
CVPR3
2025 Bayesian Motion Estimation for Articulated Heavy Vehicles; A Damper-Based Model for Coupling Force
abstract
Accurately estimating articulation angle, coupling force, and lateral velocity in articulated heavy vehicles is critical for accident prevention and energy efficiency. However, despite their importance, these quantities - especially the coupling force - have received limited attention in terms of practical and computationally efficient estimation methods. To bridge this gap, we propose a novel modeling approach that conceptualizes the coupling as a rigid damper. This formulation significantly reduces computational complexity while maintaining high estimation accuracy. Within a Bayesian estimation framework, we employ an unscented Kalman filter (UKF) for real-time inference of the vehicle states. We validate our method on high-fidelity simulation data with realistic scenarios and sensor noise. The results demonstrate the effectiveness of our method, highlighting its potential for enhancing vehicle safety and performance in practical applications.
Axel Ceder, Lars Hammarstrand, Mats Jonasson, Murat Kumru, Leo Laine
FUSION2
2024 Localization is All You Evaluate: Data Leakage in Online Mapping Datasets and How to Fix it
abstract
The task of online mapping is to predict a local map using current sensor observations, e.g. from lidar and camera, without relying on a pre-built map. State-of-the-art methods are based on supervised learning and are trained predominantly using two datasets: nuScenes and Argoverse 2. However, these datasets revisit the same geographic locations across training, validation, and test sets. Specifically, over 80% of nuScenes and 40% of Argoverse 2 validation and test samples are less than 5 m from a training sample. At test time, the methods are thus evaluated more on how well they localize within a memorized implicit map built from the training data than on extrapolating to unseen locations. Naturally, this data leakage causes inflated performance numbers and we propose geographically disjoint data splits to reveal the true performance in unseen environments. Experimental results show that methods perform considerably worse, some dropping more than 45 mAp, when trained and evaluated on proper data splits. Additionally, a reassessment of prior design choices reveals diverging conclusions from those based on the original split. Notably, the impact of lifting methods and the support from auxiliary tasks (e.g., depth supervision) on performance appears less substantial or follows a different trajectory than previously perceived. https://github.com/LiljaAdam/geographical-splits
Adam Lilja, Junsheng Fu, Erik Stenborg, Lars Hammarstrand
CVPR4
2024 ProSub: Probabilistic Open-Set Semi-supervised Learning with Subspace-Based Out-of-Distribution Detection
Erik Wallin, Lennart Svensson, Fredrik Kahl, Lars Hammarstrand
ECCV (61)4
2024 Improving Open-Set Semi-Supervised Learning with Self-Supervision
abstract
Open-set semi-supervised learning (OSSL) embodies a practical scenario within semi-supervised learning, wherein the unlabeled training set encompasses classes absent from the labeled set. Many existing OSSL methods assume that these out-of-distribution data are harmful and put effort into excluding data belonging to unknown classes from the training objective. In contrast, we propose an OSSL framework that facilitates learning from all unlabeled data through self-supervision. Additionally, we utilize an energy-based score to accurately recognize data belonging to the known classes, making our method well-suited for handling uncurated data in deployment. We show through extensive experimental evaluations that our method yields state-of-the-art results on many of the evaluated benchmark problems in terms of closed-set accuracy and open-set recognition when compared with existing methods for OSSL. Our code is available at https://github.com/walline/ssl-tf2-sefoss.
Erik Wallin, Lennart Svensson, Fredrik Kahl, Lars Hammarstrand
WACV4
2022 DoubleMatch: Improving Semi-Supervised Learning with Self-Supervision
abstract
Following the success of supervised learning, semi-supervised learning (SSL) is now becoming increasingly popular. SSL is a family of methods, which in addition to a labeled training set, also use a sizable collection of unlabeled data for fitting a model. Most of the recent successful SSL methods are based on pseudo-labeling approaches: letting confident model predictions act as training labels. While these methods have shown impressive results on many benchmark datasets, a drawback of this approach is that not all unlabeled data are used during training. We propose a new SSL algorithm, DoubleMatch, which combines the pseudo-labeling technique with a self-supervised loss, enabling the model to utilize all unlabeled data in the training process. We show that this method achieves state-of-the-art accuracies on multiple benchmark datasets while also reducing training times compared to existing SSL methods. Code is available at https://github.com/walline/doublematch.
Erik Wallin, Lennart Svensson, Fredrik Kahl, Lars Hammarstrand
ICPR4
2022 Long-Term Visual Localization Revisited
abstract
Visual localization enables autonomous vehicles to navigate in their surroundings and augmented reality applications to link virtual to real worlds. Practical visual localization approaches need to be robust to a wide variety of viewing conditions, including day-night changes, as well as weather and seasonal variations, while providing highly accurate six degree-of-freedom (6DOF) camera pose estimates. In this paper, we extend three publicly available datasets containing images captured under a wide variety of viewing conditions, but lacking camera pose information, with ground truth pose information, making evaluation of the impact of various factors on 6DOF camera pose estimation accuracy possible. We also discuss the performance of state-of-the-art localization approaches on these datasets. Additionally, we release around half of the poses for all conditions, and keep the remaining half private as a test set, in the hopes that this will stimulate research on long-term visual localization, learned local image features, and related research areas. Our datasets are available at visuallocalization.net, where we are also hosting a benchmarking server for automatic evaluation of results on the test set. The presented state-of-the-art results are to a large degree based on submissions to our server.
Carl Toft, Will Maddern, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Tomás Pajdla, Fredrik Kahl, Torsten Sattler
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Back to the Feature: Learning Robust Camera Localization From Pixels To Pose
abstract
Camera pose estimation in known scenes is a 3D geometry task recently tackled by multiple learning algorithms. Many regress precise geometric quantities, like poses or 3D points, from an input image. This either fails to generalize to new viewpoints or ties the model parameters to a specific scene. In this paper, we go Back to the Feature: we argue that deep networks should focus on learning robust and invariant visual features, while the geometric estimation should be left to principled algorithms. We introduce PixLoc, a scene-agnostic neural network that estimates an accurate 6-DoF pose from an image and a 3D model. Our approach is based on the direct alignment of multiscale deep features, casting camera localization as metric learning. PixLoc learns strong data priors by end-to-end training from pixels to pose and exhibits exceptional generalization to new scenes by separating model parameters and scene geometry. The system can localize in large environments given coarse pose priors but also improve the accuracy of sparse feature matching by jointly refining keypoints and poses with little overhead. The code will be publicly available at github.com/cvg/pixloc.
Paul-Edouard Sarlin, Ajaykumar Unagar, Måns Larsson, Hugo Germain, Carl Toft, Viktor Larsson, Marc Pollefeys, Vincent Lepetit, Lars Hammarstrand, Fredrik Kahl, Torsten Sattler
CVPR9
2021 Extended Object Tracking Using Sets Of Trajectories with a PHD Filter
Jakob Sjudin, Martin Marcusson, Lennart Svensson, Lars Hammarstrand
FUSION4
2020 Using Image Sequences for Long-Term Visual Localization
abstract
Estimating the pose of a camera in a known scene, i.e., visual localization, is a core task for applications such as self-driving cars. In many scenarios, image sequences are available and existing work on combining single-image localization with odometry offers to unlock their potential for improving localization performance. Still, the largest part of the literature focuses on single-image localization and ignores the availability of sequence data. The goal of this paper is to demonstrate the potential of image sequences in challenging scenarios, e.g., under day-night or seasonal changes. Combining ideas from the literature, we describe a sequence-based localization pipeline that combines odometry with both a coarse and a fine localization module. Experiments on long-term localization datasets show that combining single-image global localization against a prebuilt map with a visual odometry/SLAM pipeline improves performance to a level where the extended CMU Seasons dataset can be considered solved. We show that SIFT features can perform on par with modern state-of-the-art features in our framework, despite being much weaker and a magnitude faster to compute. Our code is publicly available at github.com/rulllars.
Erik Stenborg, Torsten Sattler, Lars Hammarstrand
3DV3
2019 A Cross-Season Correspondence Dataset for Robust Semantic Segmentation
abstract
In this paper, we present a method to utilize 2D-2D point matches between images taken during different image conditions to train a convolutional neural network for semantic segmentation. Enforcing label consistency across the matches makes the final segmentation algorithm robust to seasonal changes. We describe how these 2D-2D matches can be generated with little human interaction by geometrically matching points from 3D models built from images. Two cross-season correspondence datasets are created providing 2D-2D matches across seasonal changes as well as from day to night. The datasets are made publicly available to facilitate further research. We show that adding the correspondences as extra supervision during training improves the segmentation performance of the convolutional neural network, making it more robust to seasonal changes and weather conditions.
Måns Larsson, Erik Stenborg, Lars Hammarstrand, Marc Pollefeys, Torsten Sattler, Fredrik Kahl
CVPR3
2019 Fine-Grained Segmentation Networks: Self-Supervised Segmentation for Improved Long-Term Visual Localization
abstract
Long-term visual localization is the problem of estimating the camera pose of a given query image in a scene whose appearance changes over time. It is an important problem in practice that is, for example, encountered in autonomous driving. In order to gain robustness to such changes, long-term localization approaches often use segmantic segmentations as an invariant scene representation, as the semantic meaning of each scene part should not be affected by seasonal and other changes. However, these representations are typically not very discriminative due to the very limited number of available classes. In this paper, we propose a novel neural network, the Fine-Grained Segmentation Network (FGSN), that can be used to provide image segmentations with a larger number of labels and can be trained in a self-supervised fashion. In addition, we show how FGSNs can be trained to output consistent labels across seasonal changes. We show through extensive experiments that integrating the fine-grained segmentations produced by our FGSNs into existing localization algorithms leads to substantial improvements in localization performance.
Måns Larsson, Erik Stenborg, Carl Toft, Lars Hammarstrand, Torsten Sattler, Fredrik Kahl
ICCV4
2018 Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions
abstract
Visual localization enables autonomous vehicles to navigate in their surroundings and augmented reality applications to link virtual to real worlds. Practical visual localization approaches need to be robust to a wide variety of viewing condition, including day-night changes, as well as weather and seasonal variations, while providing highly accurate 6 degree-of-freedom (6DOF) camera pose estimates. In this paper, we introduce the first benchmark datasets specifically designed for analyzing the impact of such factors on visual localization. Using carefully created ground truth poses for query images taken under a wide variety of conditions, we evaluate the impact of various factors on 6DOF camera pose estimation accuracy through extensive experiments with state-of-the-art localization approaches. Based on our results, we draw conclusions about the difficulty of different conditions, showing that long-term localization is far from solved, and propose promising avenues for future work, including sequence-based localization approaches and the need for better local features. Our benchmark is available at visuallocalization.net.
Torsten Sattler, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Fredrik Kahl, Tomás Pajdla
CVPR5
2018 Semantic Match Consistency for Long-Term Visual Localization
Carl Toft, Erik Stenborg, Lars Hammarstrand, Lucas Brynte, Marc Pollefeys, Torsten Sattler, Fredrik Kahl
ECCV (2)3
2018 Long-Term Visual Localization Using Semantically Segmented Images
abstract
Robust cross-seasonal localization is one of the major challenges in long-term visual navigation of autonomous vehicles. In this paper, we exploit recent advances in semantic segmentation of images, i.e., where each pixel is assigned a label related to the type of object it represents, to attack the problem of long-term visual localization. We show that semantically labeled 3D point maps of the environment, together with semantically segmented images, can be efficiently used for vehicle localization without the need for detailed feature descriptors (SIFT, SURF, etc.), Thus, instead of depending on hand-crafted feature descriptors, we rely on the training of an image segmenter. The resulting map takes up much less storage space compared to a traditional descriptor based map. A particle filter based semantic localization solution is compared to one based on SIFT-features, and even with large seasonal variations over the year we perform on par with the larger and more descriptive SIFT-features, and are able to localize with an error below 1 m most of the time.
Erik Stenborg, Carl Toft, Lars Hammarstrand
ICRA3
2016 Using a single band GNSS receiver to improve relative positioning in autonomous cars
abstract
We show how the combination of a single band global navigation satellite systems (GNSS) receiver, standard automotive level inertial measurement unit (IMU), and wheel speed sensors, can be used for relative positioning with accuracy on a decimeter scale. It is realized without the need for expensive dual band receivers, base stations or long initialization times. This is implemented and evaluated in a natural driving environment against a reference systems and against two simple base line systems; one using only IMU and wheel speed sensors, the other also adding basic GNSS. The proposed solution provides substantially slower error growth than either of the two base line systems.
Erik Stenborg, Lars Hammarstrand
Intelligent Vehicles Symposium2
2016 Long-Range Road Geometry Estimation Using Moving Vehicles and Roadside Observations
abstract
This paper presents an algorithm for estimating the shape of the road ahead of a host vehicle equipped with the following onboard sensors: a camera, a radar, and vehicle internal sensors. The aim is to accurately describe the road geometry up to 200 m ahead in highway scenarios. This purpose is accomplished by deriving a precise clothoid-based road model for which we design a Bayesian fusion framework. Using this framework, the road geometry is estimated using sensor observations on the shape of the lane markings, the heading of leading vehicles, and the position of roadside radar reflectors. The evaluation on sensor data shows that the proposed algorithm is capable of capturing the shape of the road well, even in challenging mountainous highways.
Lars Hammarstrand, Maryam Fatemi, Ángel F. García-Fernández, Lennart Svensson
IEEE Trans. Intell. Transp. Syst.1
2016 Driver-Gaze Zone Estimation Using Bayesian Filtering and Gaussian Processes
abstract
In this paper, we propose a Bayesian filtering approach that uses information from camera-based driver monitoring systems and filtering techniques to find the probability that the driver is looking in different zones. In particular, the focus is on a set of zones directly related either to active driving or to visual distraction, such as the road, the mirrors, the infotainment display, or control buttons. For systems that do not provide direct observations of the gaze direction or as a complement to noisy gaze data, we propose to use probabilistic functions that describe the gaze direction as a function of head pose and eye closure. It is further shown how these functions can be estimated from data with know visual focus points using Gaussian processes. Evaluation on data from two driver monitoring systems shows a significant improvement compared with the gaze zone estimates based on unprocessed data.
Malin Lundgren, Lars Hammarstrand, Tomas McKelvey
IEEE Trans. Intell. Transp. Syst.2
2014 Vehicle self-localization using off-the-shelf sensors and a detailed map
abstract
In the research on autonomous vehicles, self-localization is an important problem to solve. In this paper we present a localization algorithm based on a map and a set of off-the-shelf sensors, with the purpose of evaluating this low-cost solution with respect to localization performance. The used test vehicle is equipped with a Global Positioning System receiver, a gyroscope, wheel speed sensors, a camera providing information about lane markings, and a radar detecting landmarks along the road. Evaluation shows that the localization result is within or close to the requirements for autonomous driving when lane markers and good radar landmarks are present. However, it also indicates that the solution is not robust enough to handle situations when one of these information sources is absent.
Malin Lundgren, Erik Stenborg, Lennart Svensson, Lars Hammarstrand
Intelligent Vehicles Symposium4
2014 Bayesian Road Estimation Using Onboard Sensors
abstract
This paper describes an algorithm for estimating the road ahead of a host vehicle based on the measurements from several onboard sensors: a camera, a radar, wheel speed sensors, and an inertial measurement unit. We propose a novel road model that is able to describe the road ahead with higher accuracy than the usual polynomial model. We also develop a Bayesian fusion system that uses the following information from the surroundings: lane marking measurements obtained by the camera and leading vehicle and stationary object measurements obtained by a radar-camera fusion system. The performance of our fusion algorithm is evaluated in several drive tests. As expected, the more information we use, the better the performance is.
Ángel F. García-Fernández, Lars Hammarstrand, Maryam Fatemi, Lennart Svensson
IEEE Trans. Intell. Transp. Syst.2
2013 A Probabilistic Framework for Decision-Making in Collision Avoidance Systems
abstract
This paper is concerned with the problem of decision-making in systems that assist drivers in avoiding collisions. An important aspect of these systems is not only assisting the driver when needed but also not disturbing the driver with unnecessary interventions. Aimed at improving both of these properties, a probabilistic framework is presented for jointly evaluating the driver acceptance of an intervention and the necessity thereof to automatically avoid a collision. The intervention acceptance is modeled as high if it estimated that the driver judges the situation as critical, based on the driver's observations and predictions of the traffic situation. One advantage with the proposed framework is that interventions can be initiated at an earlier stage when the estimated driver acceptance is high. Using a simplified driver model, the framework is applied to a few different types of collision scenarios. The results show that the framework has appealing properties, both with respect to increasing the system benefit and to decreasing the risk of unnecessary interventions.
Mattias Brännström 0001, Fredrik Sandblom, Lars Hammarstrand
IEEE Trans. Intell. Transp. Syst.3
2012 A study of MAP estimation techniques for nonlinear filtering
Maryam Fatemi, Lennart Svensson, Lars Hammarstrand, Mark R. Morelande
FUSION3
2012 A cardinality preserving multitarget multi-Bernoulli RFS tracker
Vishal Cholapadi Ravindra, Lennart Svensson, Lars Hammarstrand, Mark R. Morelande
FUSION3
2011 A New Vehicle Motion Model for Improved Predictions and Situation Assessment
abstract
Reliable and accurate vehicle motion models are of vital importance for automotive active safety systems for a number of reasons. First of all, these models are necessary in tracking algorithms that provide the safety system with information. Second, the motion model is often used by the safety application to make long-term predictions about the future traffic situation. These predictions are then part of the basic data used by the system to determine if, when, and how to intervene. In this paper, we suggest a framework for designing accurate vehicle motion models. The resulting models differ from conventional models in that the expected control input from the driver is included. By also providing a methodology for a formal treatment of the uncertainties, a model structure well suited, e.g., in a tracking algorithm, is obtained. To utilize the framework in an application will require careful design and validation of submodels to calculate the expected driver control input. We illustrate the potential of the framework by examining the performance for a specific model example using real measurements. The properties are compared with those of a constant acceleration model. Evaluations indicate that the proposed model yields better predictions and that it has an ability to estimate the prediction uncertainties.
Joakim Sorstedt, Lennart Svensson, Fredrik Sandblom, Lars Hammarstrand
IEEE Trans. Intell. Transp. Syst.4