Kalle Åström

dblp:45/2097 · also Karl Åström 0002 · DBLP profile ↗
← Back
126ranked-venue papers
14as first author
24since 2021 · last 2025
0000-0002-8689-7810ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 94 · 8 first-author · 17 since 2021Artificial intelligence and machine learning · 93 · 14 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Algebraic analysis of Doppler-based positioning with application to LEO satellites
abstract
In this paper we revisit positioning from Doppler measurements, using techniques from algebraic geometry to produce new theoretical insights and efficient localization algorithms. First, we give a full characterization of the problem, identifying for each problem their minimal configurations, that is the minimum number of receivers and transmitters to have a finite number of solutions. We also study the degree of each configuration. Finally, we show how these algebraic techniques can be applied to obtain an accurate position estimation in challenging environments, such as in the presence of noise and outliers. We also describe how the derived optimized algebraic solver can be integrated in a positioning engine for Low Earth Orbit position, navigation and timing (LEO PNT) systems.
Luca Ferranti, Kalle Åström, Magnus Oskarsson, Fabricio dos Santos Prol
IPIN2
2024 NeuroNCAP: Photorealistic Closed-Loop Safety Testing for Autonomous Driving
William Ljungbergh, Adam Tonderski, Joakim Johnander, Holger Caesar, Kalle Åström, Michael Felsberg, Christoffer Petersson
ECCV (30)5
2024 Geometry-Biased Transformer for Robust Multi-View 3D Human Pose Reconstruction
abstract
We address the challenges in estimating 3D human poses from multiple views under occlusion and with limited overlapping views. We approach multi-view, single-person 3D human pose reconstruction as a regression problem and propose a novel encoder-decoder Transformer architecture to estimate 3D poses from multi-view 2D pose sequences. The encoder refines 2D skeleton joints detected across different views and times, fusing multi-view and temporal information through global self-attention. We enhance the encoder by incorporating a geometry-biased attention mechanism, effectively leveraging geometric relationships between views. Additionally, we use detection scores provided by the 2D pose detector to further guide the encoder's attention based on the reliability of the 2D detections. The decoder subsequently regresses the 3D pose sequence from these refined tokens, using pre-defined queries for each joint. To enhance the generalization of our method to unseen scenes and improve resilience to missing joints, we implement strategies including scene centering, synthetic views, and token dropout. We conduct extensive experiments on three benchmark public datasets, Human3.6M, CMU Panoptic and Occlusion-Persons. Our results demonstrate the efficacy of our approach, particularly in occluded scenes and when few views are available, which are traditionally challenging scenarios for triangulation-based methods.
Olivier Moliner, Sangxia Huang, Kalle Åström
FG3
2024 GCC-PHAT Re-Imagined - A U-Net Filter for Audio TDOA Peak-Selection
abstract
Time-difference-of-arrival (TDOA) estimation from GCC-PHAT is not always as straight forward as finding the maximum peak. This work views the GCC output as an image, with time on the vertical axis and TDOA horizontally, to explore if image-to-image machine learning methods can make a more robust filter. The Structure from Sound Database provides audio recorded with a distributed microphone setup and a moving sound source. The audio was fed to GCC-PHAT without pre-processing, and images were produced for batch processing. The ground truth, the direct-path TDOA, shows a continuous curve through time. The GCC output image has a similar curve, but obscured by noise and not at all times texturally different from the multi-path components. The main approach tested is binary semantic segmentation with a U-Net. A challenge is the extreme class imbalance within the image. Preliminary results indicate that the method is valid to detect curves, yet more work is needed to single out the direct path TDOA with confidence.
Jens Gulin, Kalle Åström
ICASSP2
2024 SONNET: Enhancing Time Delay Estimation by Leveraging Simulated Audio
Erik Tegler, Magnus Oskarsson, Kalle Åström
ICPR (20)3
2024 The LuViRA Dataset: Synchronized Vision, Radio, and Audio Sensors for Indoor Localization
abstract
We present a synchronized multisensory dataset for accurate and robust indoor localization: the Lund University Vision, Radio, and Audio (LuViRA) Dataset. The dataset includes color images, corresponding depth maps, inertial measurement unit (IMU) readings, channel response between a 5G massive multiple-input and multiple-output (MIMO) testbed and user equipment, audio recorded by 12 microphones, and accurate six degrees of freedom (6DOF) pose ground truth of 0.5 mm. We synchronize these sensors to ensure that all data is recorded simultaneously. A camera, speaker, and transmit antenna are placed on top of a slowly moving service robot, and 89 trajectories are recorded. Each trajectory includes 20 to 50 seconds of recorded sensor data and ground truth labels. Data from different sensors can be used separately or jointly to perform localization tasks, and data from the motion capture (mocap) system is used to verify the results obtained by the localization algorithms. The main aim of this dataset is to enable research on sensor fusion with the most commonly used sensors for localization tasks. Moreover, the full dataset or some parts of it can also be used for other research areas such as channel estimation, image classification, etc. Our dataset is available at: https://github.com/ilaydayaman/LuViRA_Dataset
Ilayda Yaman, Guoda Tian, Martin Larsson, Patrik Persson, Michiel Sandra, Alexander Dürr, Erik Tegler, Nikhil Challa, Henrik Garde, Fredrik Tufvesson, Kalle Åström, Ove Edfors, Steffen Malkowsky, Liang Liu 0002
ICRA11
2024 wav2pos: Sound Source Localization using Masked Autoencoders
abstract
We present a novel approach to the 3D sound source localization task for distributed ad-hoc microphone arrays by formulating it as a set-to-set regression problem. By training a multi-modal masked autoencoder model that operates on audio recordings and microphone coordinates, we show that such a formulation allows for accurate localization of the sound source, by reconstructing coordinates masked in the input. Our approach is flexible in the sense that a single model can be used with an arbitrary number of microphones, even when a subset of audio recordings and microphone coordinates are missing. We test our method on simulated and real-world recordings of music and speech in indoor environments, and demonstrate competitive performance compared to both classical and other learning based localization methods.
Axel Berg, Jens Gulin, Mark O'Connor, Chuteng Zhou, Kalle Åström, Magnus Oskarsson
IPIN5
2024 LidarCLIP or: How I Learned to Talk to Point Clouds
abstract
Research connecting text and images has recently seen several breakthroughs, with models like CLIP, DALL•E 2, and Stable Diffusion. However, the connection between text and other visual modalities, such as lidar data, has received less attention, prohibited by the lack of text-lidar datasets. In this work, we propose LidarCLIP, a mapping from automotive point clouds to a pre-existing CLIP embedding space. Using image-lidar pairs, we supervise a point cloud encoder with the image CLIP embeddings, effectively relating text and lidar data with the image domain as an intermediary. We show the effectiveness of Lidar-CLIP by demonstrating that lidar-based retrieval is generally on par with image-based retrieval, but with complementary strengths and weaknesses. By combining image and lidar features, we improve upon both single-modality methods and enable a targeted search for challenging detection scenarios under adverse sensor conditions. We also explore zero-shot classification and show that LidarCLIP outperforms existing attempts to use CLIP for point clouds by a large margin. Finally, we leverage our compatibility with CLIP to explore a range of applications, such as point cloud captioning and lidar-to-image generation, without any additional training. Code and pre-trained models at github.com/atonderski/lidarclip.
Georg Hess, Adam Tonderski, Christoffer Petersson, Kalle Åström, Lennart Svensson
WACV4
2024 Estimates of Temporal Edge Detection Filters in Human Vision
abstract
Edge detection is an important process in human visual processing. However, as far as we know, few attempts have been made to map the temporal edge detection filters in human vision. To that end, we devised a user study and collected data from which we derived estimates of human temporal edge detection filters based on three different models, including the derivative of the infinite symmetric exponential function and temporal contrast sensitivity function. We analyze our findings using several different methods, including extending the filter to higher frequencies than were shown during the experiment. In addition, we show a proof of concept that our filter may be used in spatiotemporal image quality metrics by incorporating it into a flicker detection pipeline.
Pontus Ebelin, Gyorgy Denes, Tomas Akenine-Möller, Kalle Åström, Magnus Oskarsson, William McIlhagga
ACM Trans. Appl. Percept.4
2023 Revisiting the P3P Problem
abstract
One of the classical multi-view geometry problems is the so called P3P problem, where the absolute pose of a calibrated camera is determined from three 2D-to-3D correspondences. Since these solvers form a critical component of many vision systems (e.g. in localization and Structure-from-Motion), there have been significant effort in developing faster and more stable algorithms. While the current state-of-the-art solvers are both extremely fast and stable, there still exist configurations where they break down. In this paper we algebraically formulate the problem as finding the intersection of two conics. With this formulation we are able to analytically characterize the real roots of the polynomial system and employ a tailored solution strategy for each problem instance. The result is a fast and stable solver, that is able to correctly solve cases where competing methods might fail. Our experimental evaluation shows that we outperform the current state-of-the-art methods both in terms of speed and success rate.
Yaqing Ding 0001, Jian Yang 0003, Viktor Larsson, Carl Olsson, Kalle Åström
CVPR5
2023 Minimal Solutions to Generalized Three-View Relative Pose Problem
abstract
For a generalized (or non-central) camera model, the minimal problem for two views of six points has efficient solvers. However, minimal problems of three views with four points and three views of six lines have not yet been explored and solved, despite the efforts from the computer vision community. This paper develops the formulations of these two minimal problems and shows how state-of-the-art GPU implementations of Homotopy Continuation solver can be used effectively. The proposed methods are evaluated on both synthetic and real datasets, demonstrating that they are fast, accurate and that they improve on structure from motion estimations, when employed in an hypothesis and test setting.
Yaqing Ding 0001, Chiang-Heng Chien, Viktor Larsson, Kalle Åström, Benjamin B. Kimia
ICCV4
2022 Trilateration Using Motion Models
Martin Larsson, Erik Tegler, Kalle Åström, Magnus Oskarsson
FUSION3
2022 Multiple Offsets Multilateration: A New Paradigm for Sensor Network Calibration with Unsynchronized Reference Nodes
abstract
Positioning using wave signal measurements is used in several applications, such as GPS systems, structure from sound and Wifi based positioning. Mathematically, such problems require the computation of the positions of receivers and/or transmitters as well as time offsets if the devices are unsynchronized. In this paper, we expand the previous state-of-the-art on positioning formulations by introducing Multiple Offsets Multilateration (MOM), a new mathematical framework to compute the receivers positions with pseudoranges from unsynchronized reference transmitters at known positions. This could be applied in several scenarios, for example structure from sound and positioning with LEO satellites. We mathematically describe MOM, determining how many receivers and transmitters are needed for the network to be solvable, a study on the number of possible distinct solutions is presented and stable solvers based on homotopy continuation are derived. The solvers are shown to be efficient and robust to noise both for synthetic and real audio data.
Luca Ferranti, Kalle Åström, Magnus Oskarsson, Jani Boutellier, Juho Kannala
ICASSP2
2022 Joint Handwritten Text Recognition and Word Classification for Tabular Information Extraction
abstract
In this paper, we present a system for extracting tabular information from loosely structured handwritten documents. The system consists of three parts, (i) a u-net like CNN-based method for text detection and segmentation, (ii) a new attention-based method for simultaneous text recognition and classification of word-parts, and (iii) a method for matching the word parts into a tabular structure for each entry. A key contribution is the observation that the new attention-based recognition and classification module makes it possible for improved spatial analysis of the tabular information. The method is evaluated on a unique historical document: The Swedish Wealth Tax of 1571, consisting of 11,453 pages of hand-written tax records. The evaluation shows that the system provides a significant improvement to the state-of-the-art to the problem of tabular extraction from loosely structured historical documents.
Christopher Blomqvist, Kerstin Enflo, Andreas Jakobsson, Kalle Åström
ICPR4
2022 Minimal Solvers for Point Cloud Matching with Statistical Deformations
Gabrielle Flood, Erik Tegler, David Gillsjö, Anders Heyden, Kalle Åström
ICPR5
2022 Semantic Room Wireframe Detection from a Single View
abstract
Reconstruction of indoor surfaces with limited texture information or with repeated textures, a situation common in walls and ceilings, may be difficult with a monocular Structure from Motion system. We propose a Semantic Room Wireframe Detection task to predict a Semantic Wireframe from a single perspective image. Such predictions may be used with shape priors to estimate the Room Layout and aid reconstruction. To train and test the proposed algorithm we create a new set of annotations from the simulated Structured3D dataset. We show qualitatively that the SRW-Net handles complex room geometries better than previous Room Layout Estimation algorithms while quantitatively out-performing the baseline in non-semantic Wireframe Detection.
David Gillsjö, Gabrielle Flood, Kalle Åström
ICPR3
2022 Fast and efficient minimal solvers for quadric based camera pose estimation
abstract
In this paper we address absolute camera pose estimation. An efficient (and standard) way to solve this problem, is to use sparse keypoint correspondences. In many cases point features are not available, or are unstable over time and viewing conditions. We propose a framework based on silhouettes of quadric surfaces, with special emphasis on cylinders. We provide mathematical analysis of the problem of projected cylinders in particular, but also general quadrics. We develop a number of minimal solvers for estimating camera pose from silhouette lines of cylinders, given different calibration and cylinder properties. These solvers can be used efficiently in bootstrapping robust estimation schemes, such as RANSAC. Note that even though we have lines as image features, this is a different case than line based pose estimation, since we do not have 2D-line to 3D-line correspondences. We perform synthetic accuracy and robustness tests and evaluate on a number of real case scenarios.
Anna Gummeson, Johanna Engman, Kalle Åström, Magnus Oskarsson
ICPR3
2022 Extending GCC-PHAT using Shift Equivariant Neural Networks
abstract
Speaker localization using microphone arrays depends on accurate time delay estimation techniques. For decades, methods based on the generalized cross correlation with phase transform (GCC-PHAT) have been widely adopted for this purpose. Recently, the GCC-PHAT has also been used to provide input features to neural networks in order to remove the effects of noise and reverberation, but at the cost of losing theoretical guarantees in noise-free conditions. We propose a novel approach to extending the GCC-PHAT, where the received signals are filtered using a shift equivariant neural network that preserves the timing information contained in the signals. By extensive experiments we show that our model consistently reduces the error of the GCC-PHAT in adverse environments, with guarantees of exact time delay recovery in ideal conditions.
Axel Berg, Mark O'Connor, Kalle Åström, Magnus Oskarsson
INTERSPEECH3
2021 Parameterization of Ambiguity in Monocular Depth Prediction
abstract
Monocular depth estimation is a highly challenging problem that is often addressed with deep neural networks. While these use recognition of high level image features to predict reasonably looking depth maps, the result often has poor metric accuracy. Moreover, the standard feed forward architecture does not allow modification of the prediction based on cues other than the image.In this paper we relax the monocular depth estimation task by proposing a network that allows us to complement image features with a set of auxiliary variables. These allow disambiguation when image features are not enough to accurately pinpoint the exact depth map and can be thought of as a low dimensional parameterization of the surfaces that are reasonable monocular predictions. By searching the parameterization we can combine monocular estimation with traditional photoconsistency or geometry based methods to achieve both visually appealing and metrically accurate surface estimations. Since we relax the problem we are able to work with smaller networks than current architectures. In addition we design a self-supervised training scheme, eliminating the need for ground truth image depth-map pairs. Our experimental evaluation shows that our method generates more accurate depth maps and generalizes better than competing state-of-the-art approaches.
Patrik Persson, Linn Öström, Carl Olsson, Kalle Åström
3DV4
2021 Sensor Networks TDOA Self-Calibration: 2D Complexity Analysis and Solutions
abstract
Given a network of receivers and transmitters, the process of determining their positions from measured pseudoranges is known as network self-calibration. In this paper we consider 2D networks with synchronized receivers but unsynchronized transmitters and the corresponding calibration techniques, known as Time-Difference-Of-Arrival (TDOA) techniques. Despite previous work, TDOA self-calibration is computationally challenging. Iterative algorithms are very sensitive to the initialization, causing convergence issues. In this paper, we present a novel approach, which gives an algebraic solution to two previously unsolved scenarios. We also demonstrate that our solvers produce an excellent initial value for non-linear optimisation algorithms, leading to a full pipeline robust to noise.
Luca Ferranti, Kalle Åström, Magnus Oskarsson, Jani Boutellier, Juho Kannala
ICASSP2
2021 Fast and Robust Stratified Self-Calibration Using Time-Difference-Of-Arrival Measurements
abstract
In this paper we study the problem of estimating receiver and sender positions using time-difference-of-arrival measurements. For this, we use a stratified, two-tiered approach. In the first step the problem is converted to a low-rank matrix estimation problem. We present new, efficient solvers for the minimal problems of this low-rank problem. These solvers are used in a hypothesis and test manner to efficiently remove outliers and find an initial estimate which is used for the subsequent step. Once a promising solution is obtained for a sufficiently large subset of the receivers and senders, the solution can be extended to the remaining receivers and senders. These steps are then combined with robust local optimization using the initial inlier set and the initial estimate as a starting point. The proposed system is verified on both real and synthetic data.
Martin Larsson, Gabrielle Flood, Magnus Oskarsson, Kalle Åström
ICASSP4
2021 Accurate Indoor Positioning Based on Learned Absolute and Relative Models
abstract
To improve the accuracy of indoor positioning systems it can be useful to combine different types of sensor data. This paper describes deep learning methods both for estimating absolute positions and for performing pedestrian dead reckoning, and then how to combine the resulting estimates using weighted least squares optimization. The positioning model is based on a custom neural network which uses measurements of received signal strength indication from one instant of time as input. The model for estimating relative positions is on the other hand based on inertial sensors, the accelerometer, magnetometer and gyroscope. The position estimates are then combined using a least squares approach with weights based on the standard deviations of errors in predictions from the used models.
Christoffer Kjellson, Martin Larsson, Kalle Åström, Magnus Oskarsson
IPIN3
2021 Reconfigurable Multi-Access Pattern Vector Memory for Real-Time ORB Feature Extraction
abstract
This work presents an on-chip memory subsystem envisioned for real-time applications performing Oriented FAST and Rotated Brief (ORB) feature extraction for Simultaneous Localization and Mapping (SLAM) systems. For autonomous navigation of battery-powered devices, feature-based SLAM is a computationally frugal alternative to direct methods. This paper thoroughly analyses ORB multiple memory access patterns, exploring possible systematic parallelism and hardware-biased algorithmic enhancements, alleviating requirements on bandwidth and reducing redundant accesses. Enabling those, a suitable multi-bank parallel memory featuring run-time reconfigurable address generation, image allotment, and close-to-memory data-shuffling is proposed. As case study, a 30 Frames-Per-Second (FPS) VGA-resolution ORB-capable 8-bank memory is evaluated using 22 FDX technology, running at 909 MHz, with a negligible area overhead of 0.3%, reducing operand accesses between 54 - 160× relative to Sudoku-like and scalar memories.
Lucas Ferreira, Steffen Malkowsky, Patrik Persson, Kalle Åström, Liang Liu 0002
ISCAS4
2021 Efficient Real-Time Radial Distortion Correction for UAVs
abstract
In this paper we present a novel algorithm for onboard radial distortion correction for unmanned aerial vehicles (UAVs) equipped with an inertial measurement unit (IMU), that runs in real-time. This approach makes calibration procedures redundant, thus allowing for exchange of optics extemporaneously. By utilizing the IMU data, the cameras can be aligned with the gravity direction. This allows us to work with fewer degrees of freedom, and opens up for further intrinsic calibration. We propose a fast and robust minimal solver for simultaneously estimating the focal length, radial distortion profile and motion parameters from homographies. The proposed solver is tested on both synthetic and real data, and perform better or on par with state-of-the-art methods relying on pre-calibration procedures. Code available at: https://github.com/marcusvaltonen/HomLib.1
Marcus Valtonen Örnhag, Patrik Persson, Mårten Wadenbäck, Kalle Åström, Anders Heyden
WACV4
2020 Upgrade Methods for Stratified Sensor Network Self-Calibration
abstract
Estimating receiver and sender positions is often solved using a stratified, two-tiered approach. In the first step the problem is converted to a low-rank matrix estimation problem. The second step can be seen as an affine upgrade. This affine upgrade is the focus of this paper. In the paper new efficient algorithms for solving for the upgrade parameters using minimal data are presented. It is also shown how to combine such solvers as initial estimates, either directly or after a hypothesis and test step, in optimization of likelihood. The system is verified on both real and synthetic data.
Martin Larsson, Gabrielle Flood, Magnus Oskarsson, Kalle Åström
ICASSP4
2020 Calibration and Absolute Pose Estimation of Trinocular Linear Camera Array for Smart City Applications
abstract
A method for calibrating a Trinocular Linear Camera Array (TLCA) for traffic surveillance applications, such as towards smart cities, is presented. A TLCA-specific parametrization guarantees that the calibration finds a model where all the cameras are on a straight line. The method uses both a chequerboard close to the camera, as well as measured 3D points far from the camera: points measured in world coordinates, as well as their corresponding 2D points found manually in the images. Superior calibration accuracy can be obtained compared to standard methods using only a single data source, largely due to the use of chequerboards, while the line constraint in the parametrization allows for joint rectification. The improved triangulation accuracy, from 8-12 cm to around 6 cm when calibrating with 30-50 points in our experiment, allowing better road user analysis. The method is demonstrated by a proof-of-concept application where a point cloud is generated from multiple disparity maps, visualizing road user detections in 3D.
Martin Ahrnbom, Mikael G. Nilsson, Håkan Ardö, Kalle Åström, Oksana Yastremska-Kravchenko, Aliaksei Laureshyn
ICPR4
2020 Prediction of Obstructive Coronary Artery Disease from Myocardial Perfusion Scintigraphy using Deep Neural Networks
abstract
For diagnosis and risk assessment in patients with stable ischemic heart disease, myocardial perfusion scintigraphy is one of the most common cardiological examinations performed today. There are however many motivations for why an artificial intelligence algorithm would provide useful input to this task. For example to reduce the subjectiveness and save time for the nuclear medicine physicians working with this time consuming task. In this work we have developed a deep learning algorithm for multi-label classification based on a convolutional neural network to estimate the probability of obstructive coronary artery disease in the left anterior artery, left circumflex artery and right coronary artery. The prediction is based on data from myocardial perfusion scintigraphy studies conducted in a dedicated Cadmium-Zinc-Telluride cardio camera (D-SPECT Spectrum Dynamics). Data from 588 patients was available, with stress images in both upright and supine position, as well as a number of auxiliary parameters such as angina symptoms and age. The data was used to train and evaluate the algorithm using 5-fold cross-validation. We achieve state-of-the-art results for this task with an area under the receiver operating characteristics curve of 0.89 as average on per-vessel level and 0.95 on per-patient level.
Ida Arvidsson, Niels Chr. Overgaard, Kalle Åström, Anders Heyden, Miguel Ochoa Figueroa, Jeronimo Frias Rose, Anette Davidsson
ICPR3
2020 Generic Merging of Structure from Motion Maps with a Low Memory Footprint
abstract
With the development of cheap image sensors, the amount of available image data have increased enormously, and the possibility of using crowdsourced collection methods has emerged. This calls for development of ways to handle all these data. In this paper, we present new tools that will enable efficient, flexible and robust map merging. Assuming that separate optimisations have been performed for the individual maps, we show how only relevant data can be stored in a low memory footprint representation. We use these representations to perform map merging so that the algorithm is invariant to the merging order and independent of the choice of coordinate system. The result is a robust algorithm that can be applied to several maps simultaneously. The result of a merge can also be represented with the same type of low-memory footprint format, which enables further merging and updating of the map in a hierarchical way. Furthermore, the method can perform loop closing and also detect changes in the scene between the capture of the different image sequences. Using both simulated and real data - from both a hand held mobile phone and from a drone - we verify the performance of the proposed method.
Gabrielle Flood, David Gillsjö, Patrik Persson, Anders Heyden, Kalle Åström
ICPR5
2020 In Depth Bayesian Semantic Scene Completion
abstract
For autonomous agents moving around in our world, mapping of the environment is essential. This is their only perception of their surrounding, what is not measured is unknown. Humans have learned from experience what to expect in certain environments, for example in indoor offices or supermarkets. This work studies Semantic Scene Completion which aims to predict a 3D semantic segmentation of our surroundings, even though some areas are occluded. For this we construct a Bayesian Convolutional Neural Network (BCNN), which is not only able to perform the segmentation, but also predict model uncertainty. This is an important feature not present in standard CNNs. We show on the MNIST dataset that the Bayesian approach performs equal or better to the standard CNN when processing digits unseen in the training phase when looking at accuracy, precision and recall. With the added benefit of having better calibrated scores and the ability to express model uncertainty. We then show results for the Semantic Scene Completion task where a category is introduced at test time on the SUNCG dataset. In this more complex task the Bayesian approach outperforms the standard CNN. Showing better Intersection over Union score and excels in Average Precision and separation scores.
David Gillsjö, Kalle Åström
ICPR2
2020 Better Prior Knowledge Improves Human-Pose-Based Extrinsic Camera Calibration
abstract
Accurate extrinsic calibration of wide baseline multi-camera systems enables better understanding of 3D scenes for many applications and is of great practical importance. Classical Structure-from-Motion calibration methods require special calibration equipment so that accurate point correspondences can be detected between different views. In addition, an operator with some training is usually needed to ensure that data is collected in a way that leads to good calibration accuracy. This limits the ease of adoption of such technologies. Recently, methods have been proposed to use human pose estimation models to establish point correspondences, thus removing the need for any special equipment. The challenge with this approach is that human pose estimation algorithms typically produce much less accurate feature points compared to classical patch-based methods. Another problem is that ambient human motion might not be optimal for calibration. We build upon prior works and introduce several novel ideas to improve the accuracy of human-pose-based extrinsic calibration. Our first contribution is a robust reprojection loss based on a better understanding of the sources of pose estimation error. Our second contribution is a 3D human pose likelihood model learned from motion capture data. We demonstrate significant improvements in calibration accuracy by evaluating our method on four publicly available datasets.
Olivier Moliner, Sangxia Huang, Kalle Åström
ICPR3
2020 Minimal Solvers for Indoor UAV Positioning
abstract
In this paper we consider a collection of relative pose problems which arise naturally in applications for visual indoor UAV navigation. We focus on cases where additional information from an onboard IMU is available and thus provides a partial extrinsic calibration through the gravitational vector. The solvers are designed for a partially calibrated camera, for a variety of realistic indoor scenarios, which makes it possible to navigate using images of the ground floor. Current state-of-the-art solvers use more general assumptions, such as using arbitrary planar structures; however, these solvers do not yield adequate reconstructions for real scenes, nor do they perform fast enough to be incorporated in real-time systems. We show that the proposed solvers enjoy better numerical stability, are faster, and require fewer point correspondences, compared to state-of-the-art solvers. These properties are vital components for robust navigation in real-time systems, and we demonstrate on both synthetic and real data that our method outperforms other methods, and yields superior motion estimation.
Marcus Valtonen Örnhag, Patrik Persson, Mårten Wadenbäck, Kalle Åström, Anders Heyden
ICPR4
2019 Robust Self-calibration of Constant Offset Time-difference-of-arrival
abstract
In this paper we study the problem of estimating receiver and sender positions from time-difference-of-arrival measurements, assuming an unknown constant time-difference-of-arrival offset. This problem is relevant for example for repetitive sound events. In this paper it is shown that there are three minimal cases to the problem. One of these (the five receiver, five sender problem) is of particular importance. A fast solver (with run-time under 4 µs) is given. We show how this solver can be used in robust estimation algorithms, based on RANSAC, for obtaining an initial estimate followed by local optimization using a robust error norm. The system is verified on both real and synthetic data.
Kenneth Batstone, Gabrielle Flood, Thejasvi Beleyur, Viktor Larsson, Holger R. Goerlitz, Magnus Oskarsson, Kalle Åström
ICASSP7
2019 Optimal Trilateration Is an Eigenvalue Problem
abstract
The problem of estimating receiver or sender node positions from measured receiver-sender distances is a key issue in different applications such as microphone array calibration, radio antenna array calibration, mapping and positioning using UWB or using round-trip-time measurements between mobile phones and WiFi-units. In this paper we address the problem of optimally estimating a receiver position given a number of distance measurements to known sender positions, so called trilateration. We show that this problem can be rephrased as an eigenvalue problem. We also address different error models and the multilateration setting where an additional offset is also unknown, and show that these problems can be modeled using the same framework.
Martin Larsson, Viktor Larsson, Kalle Åström, Magnus Oskarsson
ICASSP3
2019 Collaborative Merging of Radio SLAM Maps in View of Crowd-sourced Data Acquisition and Big Data
abstract
Indoor localization and navigation is a much researched and difficult problem. The best solutions, usually use expensive specialized equipment and/or prior calibration of some form. To the average person with smart or Internet-Of-Things devices, these solutions are not feasible, particularly in large scales. With hardware advancements making Ultra-Wideband devices more accurate and low powered, this unlocks the potential of having such devices in commonplace around factories and homes, enabling an alternative method of navigation. Therefore, indoor anchor calibration becomes a key problem in order to implement these devices efficiently and effectively. In this paper, we present a method to fuse radio SLAM (also known as Time-Of-Arrival self-calibration) maps together in a linear way. In doing so we are then able to collaboratively calibrate the anchor positions in 3D to native precision of the devices. Furthermore, we introduce an automatic scheme to determine which of the maps are best to use to further improve the anchor calibration and its robustness but also show which maps could be discarded. Additionally, when a map is fused in a linear way, it is a very computationally cheap process and produces a reasonable map which is required to push for crowd-sourced data acquisition.
Kenneth Batstone, Magnus Oskarsson, Kalle Åström
ICPRAM3
2019 Massive MIMO-Based Localization and Mapping Exploiting Phase Information of Multipath Components
abstract
In this paper, we present a robust multipath-based localization and mapping framework that exploits the phases of specular multipath components (MPCs) using a massive multiple-input multiple-output (MIMO) array at the base station. Utilizing the phase information related to the propagation distances of the MPCs enables the possibility of localization with extraordinary accuracy even with limited bandwidth. The specular MPC parameters along with the parameters of the noise and the dense multipath component (DMC) are tracked using an extended Kalman filter (EKF), which enables to preserve the distance-related phase changes of the MPC complex amplitudes. The DMC comprises all non-resolvable MPCs, which occur due to finite measurement aperture. The estimation of the DMC parameters enhances the estimation quality of the specular MPCs and, therefore, also the quality of localization and mapping. The estimated MPC propagation distances are subsequently used as input to a distance-based localization and mapping algorithm. This algorithm does not need prior knowledge about the surrounding environment and base station position. The performance is demonstrated with real radio-channel measurements using an antenna array with 128 ports at the base station side and a standard cellular signal bandwidth of 40MHz. The results show that the high accuracy localization is possible even with such a low bandwidth.
Xuhong Li 0001, Erik Leitinger, Magnus Oskarsson, Kalle Åström, Fredrik Tufvesson
IEEE Trans. Wirel. Commun.4
2018 Beyond Grobner Bases: Basis Selection for Minimal Solvers
abstract
Many computer vision applications require robust estimation of the underlying geometry, in terms of camera motion and 3D structure of the scene. These robust methods often rely on running minimal solvers in a RANSAC framework. In this paper we show how we can make polynomial solvers based on the action matrix method faster, by careful selection of the monomial bases. These monomial bases have traditionally been based on a Grobner basis for the polynomial ideal. Here we describe how we can enumerate all such bases in an efficient way. We also show that going beyond Grobner bases leads to more efficient solvers in many cases. We present a novel basis sampling scheme that we evaluate on a number of problems.
Viktor Larsson, Magnus Oskarsson, Kalle Åström, Alge Wallis, Zuzana Kukelova, Tomás Pajdla
CVPR3
2018 Estimating Uncertainty in Time-difference and Doppler Estimates
abstract
Sound and radio can be used to estimate the distance between a transmitter and a sender by correlating the emitted and received signal. Alternatively by correlating two received signals it is possible to estimate distance difference. Such methods can be divided into methods that are robust to noise and reverberation, but give limited precision and sub-sample refinements that are sensitive to noise, but give higher precision when initialized close to the real translation. In this paper we develop stochastic models that can explain the limits in the precision of such sub-sample time-difference estimates. Using such models we provide new methods for precise estimates of time-differences as well as Doppler effects. The method is verified on both synthetic and real data.
Gabrielle Flood, Anders Heyden, Kalle Åström
ICPRAM3
2018 Registration and Merging Maps with Uncertainties
abstract
In this paper we address the problem of registering and merging two maps in two dimensions, given covariance estimates of the two maps. We show that if two maps are given in the same coordinate system, then the problem of merging them in a statistically optimal way can be formulated as a linear least squares problem, but if they are given in different coordinate systems as well the problem becomes highly non-linear and nonconvex. We show how we can relax the problem slightly in order to optimize over the registration (i.e. putting the two maps in the same coordinate system) and at the same time optimize over the merged map. The approach is based on finding all stationary points of the optimization problem and evaluating these to choose the global optimum. We show on synthetic data that in many cases the proposed approach gives better results than naively registering and merging the maps. We also show results on real data, where we merge maps given by time-of-arrival measurements, and in these cases simpler linear methods perform just a good as the proposed method.
Martin Larsson, Kalle Åström, Magnus Oskarsson
IPIN2
2017 Efficient Solvers for Minimal Problems by Syzygy-Based Reduction
abstract
In this paper we study the problem of automatically generating polynomial solvers for minimal problems. The main contribution is a new method for finding small elimination templates by making use of the syzygies (i.e. the polynomial relations) that exist between the original equations. Using these syzygies we can essentially parameterize the set of possible elimination templates. We evaluate our method on a wide variety of problems from geometric computer vision and show improvement compared to both handcrafted and automatically generated solvers. Furthermore we apply our method on two previously unsolved relative orientation problems.
Viktor Larsson, Kalle Åström, Magnus Oskarsson
CVPR2
2017 The Misty Three Point Algorithm for Relative Pose
abstract
There is a significant interest in scene reconstruction from underwater images given its utility for oceanic research and for recreational image manipulation. In this paper we propose a novel algorithm for two view camera motion estimation for underwater imagery. Our method leverages the constraints provided by the attenuation properties of water and its effects on the appearance of the color to determine the depth difference of a point with respect to the two observing views of the underwater cameras. Additionally, we propose an algorithm, leveraging the depth differences of three such observed points, to estimate the relative pose of the cameras. Given the unknown underwater attenuation coefficients, our method estimates the relative motion up to scale. The results are represented as a generalized camera. We evaluate our method on both real data and simulated data.
Tobias Palmér, Kalle Åström, Jan-Michael Frahm
CVPR2
2017 Polynomial Solvers for Saturated Ideals
abstract
In this paper we present a new method for creating polynomial solvers for problems where a (possibly infinite) subset of the solutions are undesirable or uninteresting. These solutions typically arise from simplifications made during modeling, but can also come from degeneracies which are inherent to the geometry of the original problem. The proposed approach extends the standard action matrix method to saturated ideals. This allows us to add constraints that some polynomials should be non-zero on the solutions. This does not only offer the possibility of improved performance by removing superfluous solutions, but makes a larger class of problems tractable. Previously, problems with infinitely many solutions could not be solved directly using the action matrix method as it requires a zero-dimensional ideal. In contrast we only require that after removing the unwanted solutions only finitely many remain. We evaluate our method on three applications, optimal triangulation, time-of-arrival self-calibration and optimal vanishing point estimation.
Viktor Larsson, Kalle Åström, Magnus Oskarsson
ICCV2
2017 Semantic segmentation of microscopic images of H&E stained prostatic tissue using CNN
abstract
There is a need for an automatic Gleason scoring system that can be used for prostate cancer diagnosis. Today the diagnoses are determined by pathologists manually, which is both a complex and a time-consuming task. To reduce the pathologists' workload, but also to reduce variations between different pathologists, an automatic classification system would be of great use. Some previous works have aimed for this, but still more work needs to be done. It is probable that such a tool would benefit from having access to individually segmented, pathologically relevant objects from the images. Therefore, we have developed an algorithm for semantic segmentation of the microscopic images of H&E stained prostate tissue into Background, Stroma, Epithelial Cytoplasm and Nuclei. This algorithm is based on deep learning, or more specifically a convolutional neural network. The network design is inspired by architectures that previously have been proved successful in different applications. It consists of a contracting and an expanding part, which are symmetrical. We have reached an accuracy of 80 %, as measured by the mean of the intersection over union, for segmentation into four classes. Previous works have only investigated nuclei segmentation, and our network performed similar but for the more challenging task of four class segmentation.
Johan Isaksson, Ida Arvidsson, Kalle Åström, Anders Heyden
IJCNN3
2017 Towards real-time time-of-arrival self-calibration using ultra-wideband anchors
abstract
Indoor localisation is a currently a key issue, from robotics to the Internet of Things. With hardware advancements making Ultra-Wideband devices more accurate and low powered (potentially even passive), this unlocks the potential of having such devices in common place around factories and homes, enabling an alternative method of navigation. Therefore, anchor calibration indoors becomes a key problem in order to implement these devices efficiently and effectively. In this paper, we study the possibility for sequentially gathering Ultra-Wideband Time-of-Arrival measurements and using previously studied robust solvers, merge solutions together in order to calculate anchor positions in 3D in real-time. Here it is assumed that there is no prior knowledge of the anchor positions. This is then validated using Ultra-Wideband Time-of-Arrival data gathered by a Bitcraze Crazyflie quadcopter in 2D motion, 3D motion and full flight.
Kenneth Batstone, Magnus Oskarsson, Kalle Åström
IPIN3
2017 Robust phase-based positioning using massive MIMO with limited bandwidth
abstract
This paper presents a phase-based positioning framework using a massive MIMO system. The phase-based distance estimates of MPCs together with other parameters are tracked with an EKF, the state dimension of which varies with the birth-death processes of paths. The RIMAX and the modeling of dense multipath components in the framework further enhance the quality of parameter tracking by providing an accurate initial state and the underlying noise covariance. The tracked MPCs are fed into a time-of-arrival self-calibration positioning algorithm for simultaneous trajectory and environment estimation. Throughout the positioning process, no prior knowledge of the surrounding environment and base station position is needed. The performance is demonstrated with the measurement of a 2D complex movement, which was performed in a sports hall with an antenna array with 128 ports as base station using a standard cellular bandwidth of 40 MHz. The positioning result shows that the mean deviation of the estimated user equipment trajectory from the ground truth is 13 cm. In summary, the proposed framework is a promising high-resolution radio-based positioning solution for current and next generation cellular systems.
Xuhong Li 0001, Kenneth Batstone, Kalle Åström, Magnus Oskarsson, Carl Gustafson, Fredrik Tufvesson
PIMRC3
2016 Trust No One: Low Rank Matrix Factorization Using Hierarchical RANSAC
abstract
In this paper we present a system for performing low rank matrix factorization. Low-rank matrix factorization is an essential problem in many areas, including computer vision with applications in affine structure-from-motion, photometric stereo, and non-rigid structure from motion. We specifically target structured data patterns, with outliers and large amounts of missing data. Using recently developed characterizations of minimal solutions to matrix factorization problems with missing data, we show how these can be used as building blocks in a hierarchical system that performs bootstrapping on all levels. This gives a robust and fast system, with state-of-the-art performance.
Magnus Oskarsson, Kenneth Batstone, Kalle Åström
CVPR3
2016 Uncovering Symmetries in Polynomial Systems
Viktor Larsson, Kalle Åström
ECCV (3)2
2016 Recovering planar motion from homographies obtained using a 2.5-point solver for a polynomial system
abstract
We present a minimal solver for a special kind of homography arising in applications with planar camera motion (e.g. mobile robotics applications). Since the camera motion we consider only has five degrees of freedom, an explicit parametrisation allows us to reduce the required number of point correspondences to 2.5. Using fewer point correspondences is beneficial when used together with RANSAC, but more importantly, the proposed special solver ensures that the estimated homography is of the correct type (in contrast to the DLT, which estimates a general homography). Our method works by enforcing eleven independent polynomial constraints on the elements of this kind of homography matrix, through the framework of the action matrix method for solving polynomial equations. Some analytical investigation using symbolic software has been conducted in order to understand the properties of the polynomial system, and these results have been used to help guide our design of the solver. Additionally, we provide a direct method to recover the sought motion parameters from the homography matrix. We demonstrate that it is possible to recover both the homography and its generating parameters efficiently and accurately.
Mårten Wadenbäck, Kalle Åström, Anders Heyden
ICIP2
2016 Calibration, positioning and tracking in a refractive and reflective scene
abstract
We propose a framework for calibration, positioning and tracking in a scene viewed by multiple cameras, through a flat refractive surface and one or several flat reflective walls. Refractions are explicitly modeled by Snell's law and reflections are handled using virtual points. A novel bundle adjustment framework is introduced for solving the nonlinear equations of refractions and the linear equations of reflections, which in addition enables optimization for calibration and positioning. The numerical accuracy of the solutions is investigated on synthetic data, and the influence of noise in image points for several settings of refractive and reflective planes is presented. The performance of the framework is evaluated on real data and confirms the validity of the physical model. Examples of how to use the framework to back-project image coordinates, forward-project scene points and estimate the refractive and reflective planes are presented. Lastly, an application of the system on real data from a biological experiment on small aquatic organisms is presented.
Tobias Palmér, Giuseppe Bianco, Mikael T. Ekvall, Lars-Anders Hansson, Kalle Åström
ICPR5
2016 Smartphone positioning in multi-floor environments without calibration or added infrastructure
abstract
Indoor positioning for smartphone users has received a lot of attention in recent years. While many solutions have been developed, most rely on a need for pre-deployment of infrastructure or collecting ground truth data to train on. In this paper we see what can be done using existing WiFi-infrastructure and Received Signal Strength from these to smartphones, not using any calibration of the signal environment or manually set WiFi positions. We expand on previous work by using a multi-floor model taking into account dampening between floors, and optimize a target function consisting of least squares residuals, to find positions for WiFis and the smartphone measurement locations simultaneously. Pressure sensors are used to do floor estimation. The method was tested inside two multi-story buildings, with 5 stories each, with median errors of smartphone positions of 12.5m and 16.4m and with WiFi median position errors of 7.16m and 19.4m respectively. Correct floor detection was achieved for 96% of all smartphone positions.
Simon Burgess 0002, Kalle Åström, Mikael Hogstrom, Bjorn Lindquist
IPIN2
2016 Sparse Localization of Harmonic Audio Sources
abstract
In this paper, we propose a novel method for estimating the locations of near- and/or far-field harmonic audio sources impinging on an arbitrary, but calibrated, sensor array. Using a joint pitch and location estimation formed in two steps, we first estimate the fundamental frequencies and complex amplitudes under a sinusoidal model assumption, whereafter the location of each source is found by utilizing both the difference in phase and the relative attenuation of the magnitude estimates. As audio recordings often consist of multi-pitch signals exhibiting some degree of reverberation, where both the number of pitches and the source locations are unknown, we propose to use sparse heuristics to avoid the necessity of detailed a priori assumptions on the spectral and spatial model orders. The method’s performance is evaluated using both simulated and measured audio data, with the former showing that the proposed method achieves near-optimal performance, whereas the latter confirms the method’s feasibility when used with real recordings.
Stefan Ingi Adalbjornsson, Ted Kronvall, Simon Burgess 0002, Kalle Åström, Andreas Jakobsson
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Linking Entities Across Images and Text
abstract
This paper describes a set of methods to link entities across images and text.As a corpus, we used a data set of images, where each image is commented by a short caption and where the regions in the images are manually segmented and labeled with a category.We extracted the entity mentions from the captions and we computed a semantic similarity between the mentions and the region labels.We also measured the statistical associations between these mentions and the labels and we combined them with the semantic similarity to produce mappings in the form of pairs consisting of a region label and a caption entity.In a second step, we used the syntactic relationships between the mentions and the spatial relationships between the regions to rerank the lists of candidate mappings.To evaluate our methods, we annotated a test set of 200 images, where we manually linked the image regions to their corresponding mentions in the captions.Eventually, we could match objects in pictures to their correct mentions for nearly 89 percent of the segments, when such a matching exists.
Rebecka Weegar, Kalle Åström, Pierre Nugues
CoNLL2
2015 Absolute pose for cameras under flat refractive interfaces
abstract
This paper studies the problem of determining the absolute pose of a perspective camera observing a scene through a known refractive plane, the flat boundary between transparent media with different refractive indices. Efficient minimal solvers are developed for the 2D, known orientation and known rotation axis cases, and near-minimal solvers for the general calibrated and unknown focal length cases. We show that ambiguities in the equations of Snell's law give rise to a large number of false solutions, increasing the complexity of the problem. Evaluation of the solvers on both synthetic and real data show excellent numerical performance, and the necessity of explicitly modelling refraction to obtain accurate pose estimates.
Sebastian Haner, Kalle Åström
CVPR2
2015 On the minimal problems of low-rank matrix factorization
abstract
Low-rank matrix factorization is an essential problem in many areas including computer vision, with applications in e.g. affine structure-from-motion, photometric stereo, and non-rigid structure from motion. However, very little attention has been drawn to minimal cases for this problem or to using the minimal configuration of observations to find the solution. Minimal problems are useful when either outliers are present or the observation matrix is sparse. In this paper, we first give some theoretical insights on how to generate all the minimal problems of a given size using Laman graph theory. We then propose a new parametrization and a building-block scheme to solve these minimal problems by extending the solution from a small sized minimal problem. We test our solvers on synthetic data as well as real data with outliers or a large portion of missing data and show that our method can handle the cases when other iterative methods, based on convex relaxation, fail.
Fangyuan Jiang, Magnus Oskarsson, Kalle Åström
CVPR3
2015 Tractable Algorithms for Robust Model Estimation
Olof Enqvist, Erik Ask, Fredrik Kahl, Kalle Åström
Int. J. Comput. Vis.4
2015 TOA sensor network self-calibration for receiver and transmitter spaces with difference in dimension
Simon Burgess 0002, Yubin Kuang, Kalle Åström
Signal Process.3
2014 A Minimal Solution to Relative Pose with Unknown Focal Length and Radial Distortion
Fangyuan Jiang, Yubin Kuang, Jan Erik Solem, Kalle Åström
ACCV (2)4
2014 Minimal Solvers for Relative Pose with a Single Unknown Radial Distortion
abstract
In this paper, we study the problems of estimating relative pose between two cameras in the presence of radial distortion. Specifically, we consider minimal problems where one of the cameras has no or known radial distortion. There are three useful cases for this setup with a single unknown distortion: (i) fundamental matrix estimation where the two cameras are uncalibrated, (ii) essential matrix estimation for a partially calibrated camera pair, (iii) essential matrix estimation for one calibrated camera and one camera with unknown focal length. We study the parameterization of these three problems and derive fast polynomial solvers based on Gröbner basis methods. We demonstrate the numerical stability of the solvers on synthetic data. The minimal solvers have also been applied to real imagery with convincing results.
Yubin Kuang, Jan Erik Solem, Fredrik Kahl, Kalle Åström
CVPR4
2014 Partial Symmetry in Polynomial Systems and Its Applications in Computer Vision
abstract
Algorithms for solving systems of polynomial equations are key components for solving geometry problems in computer vision. Fast and stable polynomial solvers are essential for numerous applications e.g. minimal problems or finding for all stationary points of certain algebraic errors. Recently, full symmetry in the polynomial systems has been utilized to simplify and speed up state-of-the-art polynomial solvers based on Gröbner basis method. In this paper, we further explore partial symmetry (i.e. where the symmetry lies in a subset of the variables) in the polynomial systems. We develop novel numerical schemes to utilize such partial symmetry. We then demonstrate the advantage of our schemes in several computer vision problems. In both synthetic and real experiments, we show that utilizing partial symmetry allow us to obtain faster and more accurate polynomial solvers than the general solvers.
Yubin Kuang, Yinqiang Zheng, Kalle Åström
CVPR3
2014 Revisiting Trifocal Tensor Estimation Using Lines
abstract
In this paper, we revisit the problem of estimating the trifocal tensor from image line measurements. With measurements of corresponding lines in three views, a linear method [1] requiring 13 lines was developed to estimate the trifocal tensor from which projective reconstruction of the scene is made possible. By further utilizing the nonlinear constraints on the trifocal tensor, we propose several new linear solvers that require fewer number of lines (10,11,12) than the previous linear method. We use methods based on algebraic geometry to incorporate the non-linear constraints in the estimation. We demonstrate the performance of the proposed solvers on synthetic data. We also test the solvers on real images and obtain promising results.
Yubin Kuang, Magnus Oskarsson, Kalle Åström
ICPR3
2014 Prime Rigid Graphs and Multidimensional Scaling with Missing Data
abstract
In this paper we investigate the problem of embedding a number of points given certain (but typically not all) inter-pair distance measurements. This problem is relevant for multi-dimensional scaling problems with missing data, and is applicable within anchor-free sensor network node calibration and anchor-free node localization using radio or sound TOA measurements. There are also applications within chemistry for deducing molecular 3D structure given inter-atom distance measurements and within machine learning and visualization of data, where only similarity measures between sample points are provided. The problem has been studied previously within the field of rigid graph theory. Our aim is here to construct numerically stable and efficient solvers for finding all embeddings of such minimal rigid graphs. The method is based on the observation that all graphs are either irreducibly rigid, here called prime rigid graphs, or contain smaller rigid graphs. By solving the embedding problem for the prime rigid graphs and for ways of assembling such graphs to other minimal rigid graphs, we show how to (i) calculate the number of embeddings and (ii) construct numerically stable and efficient algorithms for obtaining all embeddings given inter-node measurements. The solvers are verified with experiments on simulated data.
Magnus Oskarsson, Kalle Åström, Anna Torstensson
ICPR2
2014 Image Segmentation and Labeling Using Free-Form Semantic Annotation
abstract
In this paper we investigate the problem of segmenting images using the information in text annotations. In contrast to the general image understanding problem, this type of annotation guided segmentation is less ill-posed in the sense that for the output there is higher consensus among human annotations. In the paper we present a system based on a combined visual and semantic pipeline. In the visual pipeline, a list of tentative figure-ground segmentations is first proposed. Each such segmentation is classified into a set of visual categories. In the natural language processing pipeline, the text is parsed and chunked into objects. Each chunk is then compared with the visual categories and the relative distance is computed using the word-net structure. The final choice of segments and their correspondence to the chunked objects are then obtained using combinatorial optimization. The output is compared to manually annotated ground-truth images. The results are promising and there are several interesting avenues for continued research.
Agnes Tegen, Rebecka Weegar, Linus Hammarlund, Magnus Oskarsson, Fangyuan Jiang, Dennis Medved, Pierre Nugues, Kalle Åström
ICPR8
2014 Combining Text Semantics and Image Geometry to Improve Scene Interpretation
abstract
Inthispaper,wedescribeanovelsystemthatidentifiesrelationsbetweentheobjectsextractedfromanimage. We started from the idea that in addition to the geometric and visual properties of the image objects, we could exploit lexical and semantic information from the text accompanying the image. As experimental set up, we gathered a corpus of images from Wikipedia as well as their associated articles. We extracted two types of objects: human beings and horses and we considered three relations that could hold between them: Ride, Lead, or None. We used geometric features as a baseline to identify the relations between the entities and we describe the improvements brought by the addition of bag-of-wordf eatures and predicate–arguments tructures we derived from the text. The best semantic model resulted in a relative error reduction of more than 18% over the baseline.
Dennis Medved, Fangyuan Jiang, Peter Exner, Magnus Oskarsson, Pierre Nugues, Kalle Åström
ICPRAM6
2013 Minimal Solvers for Unsynchronized TDOA Sensor Network Calibration
Simon Burgess 0002, Yubin Kuang, Johannes Wendeberg, Kalle Åström, Christian Schindelhauer
ALGOSENSORS4
2013 Time delay estimation for TDOA self-calibration using truncated nuclear norm regularization
abstract
Measurements with unknown time delays are common in different applications such as microphone array, radio antenna array calibration, where the sources (e.g. sounds) are transmitted in unknown time instants. In this paper, we present a method for estimating unknown time delays from Time-Difference-of-Arrival (TDOA) measurements. We propose a novel rank constraint on a matrix depending on the measurements and the unknown time delays. The time delays are recovered by solving a truncated nuclear norm minimization problem using alternating direction method of multipliers (ADMM). We show in synthetic experiments that the proposed method recovers the time delays with good accuracy for noisy and missing data.
Fangyuan Jiang, Yubin Kuang, Kalle Åström
ICASSP3
2013 A complete characterization and solution to the microphone position self-calibration problem
abstract
This paper presents a complete characterization and solution to microphone position self-calibration problem for time-of-arrival (TOA) measurements. This is the problem of determining the positions of receivers and transmitters given all receiver-transmitter distances. Such calibration problems arise in application such as calibration of radio antenna networks, audio or ultra-sound arrays and WiFi transmitter arrays. We show for what cases such calibration problems are well-defined and derive efficient and numerically stable algorithms for the minimal TOA based self-calibration problems. The proposed algorithms are non-iterative and require no assumptions on the sensor positions. Experiments on synthetic data show that the minimal solvers are numerically stable and perform well on noisy data. The solvers are also tested on two real datasets with good results.
Yubin Kuang, Simon Burgess 0002, Anna Torstensson, Kalle Åström
ICASSP4
2013 Single antenna anchor-free UWB positioning based on multipath propagation
abstract
Radio based localization and tracking usually require multiple receivers/transmitters or a known floor plan. This paper presents a method for anchor free indoor positioning based on single antenna ultra wideband (UWB) measurements. By using time of arrival information from multipath propagation components stemming from scatterers with different, but unknown, positions we estimate the movement of the receiver as well as the angle of arrival of the considered multipath components. Experiments are shown for real indoor data measured in a lecture room with promising results. Simultaneous estimation of both receiver motion, transmitter and scatterer positions is performed using an factorization based approach followed by non-linear least squares optimization. A RANSAC approach to automatic matching of data has also been implemented and tested. The resulting reconstruction is compared to ground truth motion as given by the antenna positioner. The resulting accuracy is in the order of one cm.
Yubin Kuang, Kalle Åström, Fredrik Tufvesson
ICC2
2013 Pose Estimation with Unknown Focal Length Using Points, Directions and Lines
abstract
In this paper, we study the geometry problems of estimating camera pose with unknown focal length using combination of geometric primitives. We consider points, lines and also rich features such as quivers, i.e.\ points with one or more directions. We formulate the problems as polynomial systems where the constraints for different primitives are handled in a unified way. We develop efficient polynomial solvers for each of the derived cases with different combinations of primitives. The availability of these solvers enables robust pose estimation with unknown focal length for wider classes of features. Such rich features allow for fewer feature correspondences and generate larger inlier sets with higher probability. We demonstrate in synthetic experiments that our solvers are fast and numerically stable. For real images, we show that our solvers can be used in RANSAC loops to provide good initial solutions.
Yubin Kuang, Kalle Åström
ICCV2
2013 Revisiting the PnP Problem: A Fast, General and Optimal Solution
abstract
In this paper, we revisit the classical perspective-n-point (PnP) problem, and propose the first non-iterative O(n) solution that is fast, generally applicable and globally optimal. Our basic idea is to formulate the PnP problem into a functional minimization problem and retrieve all its stationary points by using the Gr"obner basis technique. The novelty lies in a non-unit quaternion representation to parameterize the rotation and a simple but elegant formulation of the PnP problem into an unconstrained optimization problem. Interestingly, the polynomial system arising from its first-order optimality condition assumes two-fold symmetry, a nice property that can be utilized to improve speed and numerical stability of a Grobner basis solver. Experiment results have demonstrated that, in terms of accuracy, our proposed solution is definitely better than the state-of-the-art O(n) methods, and even comparable with the reprojection error minimization method.
Yinqiang Zheng, Yubin Kuang, Shigeki Sugimoto, Kalle Åström, Masatoshi Okutomi
ICCV4
2013 Minimal Structure and Motion Problems for TOA and TDOA Measurements with Collinearity Constraints
Erik Ask, Simon Burgess 0002, Kalle Åström
ICPRAM3
2012 Robust Fitting for Multiple View Geometry
Olof Enqvist, Erik Ask, Fredrik Kahl, Kalle Åström
ECCV (1)4
2012 Numerically Stable Optimization of Polynomial Solvers for Minimal Problems
Yubin Kuang, Kalle Åström
ECCV (3)2
2012 Exploiting p-fold symmetries for faster polynomial equation solving
Erik Ask, Yubin Kuang, Kalle Åström
ICPR3
2012 Node localization in unsynchronized time of arrival sensor networks
Simon Burgess 0002, Yubin Kuang, Kalle Åström
ICPR3
2012 Pose estimation from minimal dual-receiver configurations
Simon Burgess 0002, Yubin Kuang, Kalle Åström
ICPR3
2012 Understanding TOA and TDOA Network Calibration using Far Field Approximation as Initial Estimate
Yubin Kuang, Erik Ask, Simon Burgess 0002, Kalle Åström
ICPRAM (2)4
2012 Measurement of bitumen coverage of stones for road building, based on digital image analysis
abstract
The top layer of a road is made up of a mixture of stones and bitumen and the durability is dependent on how well the bitumen adheres to the stones. The standard way of determining the bitumen coverage in the industry is the so called rolling bottle method, where a number of stones covered with bitumen are put in a rolling bottle and the bitumen coverage is estimated after different times. This paper describes a novel method for measuring the bitumen coverage of the stones by using advanced segmentation methods instead of manual inspection. The stones are put on a table and a number of images with different exposure times are taken. The images are normalized and the stones are segmented from the background based on a threshold obtained from an optimality criterion. Then the bitumen covered parts of the stones are segmented based on a graph-cut method. The results are compared to manual inspection and are well in agreement with these.
Hanna Källén, Anders Heyden, Kalle Åström, Per Lindh
WACV3
2010 Optimizing Visual Vocabularies Using Soft Assignment Entropies
Yubin Kuang, Kalle Åström, Lars Kopp, Magnus Oskarsson, Martin Byröd
ACCV (4)2
2010 Conjugate Gradient Bundle Adjustment
Martin Byröd, Kalle Åström
ECCV (2)2
2010 Multi-camera Platform Calibration Using Multi-linear Constraints
abstract
We present a novel calibration method for multi-camera platforms, based on multi-linear constraints. The calibration method can recover the relative orientation between the different cameras on the platform, even when there are no corresponding feature points between the cameras, i.e. there are no overlaps between the cameras. It is shown that two translational motions in different directions are sufficient to linearly recover the rotational part of the relative orientation. Then two general motions, including both translation and rotation, are sufficient to linearly recover the translational part of the relative orientation. However, as a consequence of the speed-scale ambiguity the absolute scale of the translational part can not be determined if no prior information about the motions are known, e.g. from dead reckoning. It is shown that in case of planar motion, the vertical component of the translational part can not be determined. However, if at least one feature point can be seen in two different cameras, this vertical component can also be estimated. Finally, the performance of the proposed method is shown in simulated experiments.
Patrik Nyman, Anders Heyden, Kalle Åström
ICPR3
2010 Object tracking with measurements from single or multiple cameras
abstract
To be able to determine the position of a static object in 3D space by means of computer vision, it has to be seen by cameras from at least two different view points. The same applies for measuring the position of a moving object based on images captured at one single time instant. However, if the cameras are not synchronized in time, or if a moving object is not visible in all images, one can not rely on using matching pictures for making accurate position estimates of dynamical objects. This paper presents a strategy to track an object with known dynamical model, using a series of images where no pair has to be captured simultaneously. It even allows tracking of a point object in 3D space using a single static camera.
Magnus Linderoth, Anders Robertsson, Kalle Åström, Rolf Johansson 0001
ICRA3
2010 Fast and robust numerical solutions to minimal problems for cameras with radial distortion
Zuzana Kukelova, Martin Byröd, Klas Josephson, Tomás Pajdla, Kalle Åström
Comput. Vis. Image Underst.5
2010 Global Optimization for One-Dimensional Structure and Motion Problems
abstract
We study geometric reconstruction problems in one-dimensional retina vision. In such problems, the scene is modeled as a two-dimensional plane, and the camera sensor produces one-dimensional images of the scene. Our main contribution is an efficient method for computing the global optimum to the structure and motion problem with respect to the $L_{\infty}$ norm of the reprojection errors. One-dimensional cameras have proven useful in several applications, most prominently for autonomous vehicles, where they are used to provide inexpensive and reliable navigational systems. Previous results on one-dimensional vision are limited to the classification and solving of minimal cases, bundle adjustment for finding local optima, and linear algorithms for algebraic cost functions. In contrast, we present an approach for finding globally optimal solutions with respect to the $L_{\infty}$ norm of the angular reprojection errors. We show how to solve intersection and resection problems as well as the problem of simultaneous localization and mapping (SLAM). The algorithm is robust to use when there are missing data, which means that all points are not necessarily seen in all images. Our approach has been tested on a variety of different scenarios, both real and synthetic. The algorithm shows good performance for intersection and resection and for SLAM with up to five views. For more views the high dimension of the search space tends to give long running times. The experimental section also gives interesting examples showing that for one-dimensional cameras with limited field of view the SLAM problem is often inherently ill-conditioned.
Olof Enqvist, Fredrik Kahl, Carl Olsson, Kalle Åström
SIAM J. Imaging Sci.4
2009 Bundle Adjustment using Conjugate Gradients with Multiscale Preconditioning
abstract
Bundle adjustment is a key component of almost any feature based 3D reconstruction system, used to compute accurate estimates of calibration parameters and structure and motion configurations. These problems tend to be very large, often involving thousands of variables. Thus, efficient optimization methods are crucial. The traditional Levenberg Marquardt algorithm with a direct sparse solver can be efficiently adapted to the special structure of the problem and works well for small to medium size setups. However, for larger scale configurations the cubic computational complexity makes this approach pro- hibitively expensive. The natural step here is to turn to iterative methods for solving the normal equations such as conjugate gradients. So far, there has been little progress in this direction. This is probably due to the lack of suitable pre-conditioners, which are con- sidered essential for the success of any iterative linear solver. In this paper, we show how multi scale representations, derived from the underlying geometric layout of the problem, can be used to dramatically increase the power of straight forward preconditioners such as Gauss-Seidel.
Martin Byröd, Kalle Åström
BMVC2
2009 Minimal Solutions for Panoramic Stitching with Radial Distortion
abstract
This paper presents a solution to panoramic image stitching of two images with co- inciding optical centers, but unknown focal length and radial distortion. The algorithm operates with a minimal set of corresponding points (three) which means that it is well suited for use in any RANSAC style algorithm for simultaneous estimation of geometry and outlier rejection. Compared to a previous method for this problem, we are able to guarantee that the right solution is found in all cases. The solution is obtained by solving a small system of polynomial equations. The proposed algorithm has been integrated in a complete multi image stitching system and we evaluate its performance on real images with lens distortion. We demonstrate both quantitative and qualitative improvements compared to state of the art methods.
Martin Byröd, Matthew A. Brown, Kalle Åström
BMVC3
2009 Fast and Stable Polynomial Equation Solving and Its Application to Computer Vision
Martin Byröd, Klas Josephson, Kalle Åström
Int. J. Comput. Vis.3
2008 Fast and robust numerical solutions to minimal problems for cameras with radial distortion
abstract
A number of minimal problems of structure from motion for cameras with radial distortion have recently been studied and solved in some cases. These problems are known to be numerically very challenging and in several cases there exist no known practical algorithm yielding solutions in floating point arithmetic. We make some crucial observations concerning the floating point implementation of Gröbner basis computations and use these new insights to formulate fast and stable algorithms for two minimal problems with radial distortion previously solved in exact rational arithmetic only: (i) simultaneous estimation of essential matrix and a common radial distortion parameter for two partially calibrated views and six image point correspondences and (ii) estimation of fundamental matrix and two different radial distortion parameters for two uncalibrated views and nine image point correspondences. We demonstrate on simulated and real experiments that these two problems can be efficiently solved in floating point arithmetic.
Martin Byröd, Zuzana Kukelova, Klas Josephson, Tomás Pajdla, Kalle Åström
CVPR5
2008 A Column-Pivoting Based Strategy for Monomial Ordering in Numerical Gröbner Basis Calculations
Martin Byröd, Klas Josephson, Kalle Åström
ECCV (4)3
2008 MDL patch correspondences on unlabeled images
abstract
Automatic construction of shape and appearance models from examples via establishing correspondences across the training set has been successful in the last decades. One successful measure for establishing correspondences of high quality is minimum description length (MDL). In other approaches it has been shown that parts+geometry models which model the appearance of parts of the object and the geometric relation between the parts have been successful for automatic model building. In this paper it is shown how to fuse the above approaches and use MDL to fully automatically build optimal parts+geometry models from unlabeled images.
Johan Karlsson 0002, Kalle Åström
ICPR2
2007 Fast Optimal Three View Triangulation
Martin Byröd, Klas Josephson, Kalle Åström
ACCV (2)3
2007 Image-Based Localization Using Hybrid Feature Correspondences
abstract
Where am I and what am I seeing? This is a classical vision problem and this paper presents a solution based on efficient use of a combination of 2D and 3D features. Given a model of a scene, the objective is to find the relative camera location of a new input image. Unlike traditional hypothesize-and-test methods that try to estimate the unknown camera position based on 3D model features only, or alternatively, based on 2D model features only, we show that using a mixture of such features, that is, a hybrid correspondence set, may improve performance. We use minimal cases of structure-from-motion for hypothesis generation in a RANSAC engine. For this purpose, several new and useful minimal cases are derived for calibrated, semi-calibrated and uncalibrated settings. Based on algebraic geometry methods, we show how these minimal hybrid cases can be solved efficiently. The whole approach has been validated on both synthetic and real data, and we demonstrate improvements compared to previous work.
Klas Josephson, Martin Byröd, Fredrik Kahl, Kalle Åström
CVPR4
2007 An L Approach to Structure and Motion Problems in 1D-Vision
abstract
The structure and motion problem of multiple one-dimensional projections of a two-dimensional environment is studied. One-dimensional cameras have proven useful in several different applications, most prominently for autonomous guided vehicles, but also in ordinary vision for analysing planar motion and the projection of lines. Previous results on one-dimensional vision are limited to classifying and solving minimal cases, bundle adjustment for finding local minima to the structure and motion problem and linear algorithms based on algebraic cost functions. In this paper, we present a method for finding the global minimum to the structure and motion problem using the max norm of reprojection errors. We show how the optimal solution can be computed efficiently using simple linear programming techniques. The algorithms have been tested on a variety of different scenarios, both real and synthetic, with good performance. In addition, we show how to solve the multiview triangulation problem, the camera pose problem and how to dualize the algorithm in the Carlsson duality sense, all within the same framework.
Kalle Åström, Olof Enqvist, Carl Olsson, Fredrik Kahl, Richard I. Hartley
ICCV1
2007 Improving Numerical Accuracy of Gröbner Basis Polynomial Equation Solvers
abstract
This paper presents techniques for improving the numerical stability of Grobner basis solvers for polynomial equations. Recently Grobner basis methods have been used successfully to solve polynomial equations arising in global optimization e.g. three view triangulation and in many important minimal cases of structure from motion. Such methods work extremely well for problems of reasonably low degree, involving a few variables. Currently, the limiting factor in using these methods for larger and more demanding problems is numerical difficulties. In the paper we (i) show how to change basis in the quotient space R[x]/I and propose a strategy for selecting a basis which improves the conditioning of a crucial elimination step, (ii) use this technique to devise a Grobner basis with improved precision and (iii) show how solving for the eigenvalues instead of eigenvectors can be used to improve precision further while retaining the same speed. We study these methods on some of the latest reported uses of Grobner basis methods and demonstrate dramatically improved numerical precision using these new techniques making it possible to solve a larger class of problems than previously.
Martin Byröd, Klas Josephson, Kalle Åström
ICCV3
2004 Structure and Motion Problems for Multiple Rigidly Moving Cameras
Henrik Stewénius, Kalle Åström
ECCV (3)2
2004 Minimal projective reconstruction for combinations of points and lines in three views
Magnus Oskarsson, Andrew Zisserman, Kalle Åström
Image Vis. Comput.3
2003 Minimizing the description length using steepest descent
abstract
Recently there has been much attention to MDL and its effectiveness in automatic shape modelling. One problem of this technique has been the slow convergence of the optimization step. In this paper the Jacobian of the objective function is derived. Being able to calculate the Jacobian, a variety of optimisation techniques can be considered. In this paper we apply steepest descent and show that it is more efcient than the previously proposed Nelder-Mead Simplex optimisation.
Anders Ericsson, Kalle Åström
BMVC2
2003 An affine invariant deformable shape representation for general curves
abstract
Automatic construction of shape models from examples has been the focus of intense research during the last couple of years. These methods have proved to be useful for shape segmentation, tracking and shape understanding. In this paper novel theory to automate shape modelling is described. The theory is intrinsically defined for curves although curves are infinite dimensional objects. The theory is independent of parameterisation and affine transformations. We suggest a method for implementing the ideas and compare it to minimising the description length of the model (MDL). It turns out that the accuracy of the two methods is comparable. Both the MDL and our approach can get stuck at local minima. Our algorithm is less computational expensive and relatively good solutions are obtained after a few iterations. The MDL is, however, better suited at fine-tuning the parameters given good initial estimates to the problem. It is shown that a combination of the two methods outperforms either on its own.
Anders Ericsson, Kalle Åström
ICCV2
2002 Minimal Projective Reconstruction for Combinations of Points and Lines in Three Views
Magnus Oskarsson, Andrew Zisserman, Kalle Åström
BMVC3
2002 Robust Factorization
abstract
Factorization algorithms for recovering structure and motion from an image stream have many advantages, but they usually require a set of well-tracked features. Such a set is in generally not available in practical applications. There is thus a need for making factorization algorithms deal effectively with errors in the tracked features. We propose a new and computationally efficient algorithm for applying an arbitrary error function in the factorization scheme. This algorithm enables the use of robust statistical techniques and arbitrary noise models for the individual features. These techniques and models enable the factorization scheme to deal effectively with mismatched features, missing features, and noise on the individual features. The proposed approach further includes a new method for Euclidean reconstruction that significantly improves convergence of the factorization algorithms. The proposed algorithm has been implemented as a modification of the Christy-Horaud factorization scheme, which yields a perspective reconstruction. Based on this implementation, a considerable increase in error tolerance is demonstrated on real and synthetic data. The proposed scheme can, however, be applied to most other factorization algorithms.
Henrik Aanæs, Rune Fisker, Kalle Åström, Jens Michael Carstensen
IEEE Trans. Pattern Anal. Mach. Intell.3
2001 Critical Configurations for N-view Projective Reconstruction
abstract
In this paper we give a characterization of critical configurations for projective reconstruction with any number of points and views. A set of cameras and points is said to be critical if the projected image points are insufficient to determine the placement of the points and the cameras uniquely, up to a projective transformation. For two views, the critical configurations are well-known. In this paper it is shown that a configuration of n 3 cameras and in points all lying on the intersection of two distinct ruled quadrics is critical. In distinction to the two-view case, which in general allows two alternative solutions, there is a family of ambiguous reconstructions for the n-view case. As a partial converse, it Is shown that for any critical configuration, all the points lie on the intersection of two ruled quadrics.
Fredrik Kahl, Richard I. Hartley, Kalle Åström
CVPR (2)3
2001 Ambiguous Configurations for the 1D Structure and Motion Problem
Fredrik Kahl, Kalle Åström
ICCV2
2001 Classifying and Solving Minimal Structure and Motion Problems with Missing Data
Magnus Oskarsson, Kalle Åström, Niels Chr. Overgaard
ICCV2
2001 Reconstruction of General Curves, Using Factorization and Bundle Adjustment
Rikard Berthilsson, Kalle Åström, Anders Heyden
Int. J. Comput. Vis.2
2001 Preface - Selected Papers from Statistical Methods for Image Processing
Kalle Åström, Bjarne K. Ersbøll
Pattern Recognit. Lett.1
2000 Multiple View Vision
abstract
This paper presents a short historical overview of multiple view vision and, in particular, the estimation of both camera geometry and scene models using only images as input. This problem (the structure and motion problem) can be seen as the mathematical inverse of the computer graphics problem. The basic structure and motion problems for different feature types are discussed as well as some recent methods that estimate not only the scene geometry, but also radiance, irradiance and illumination properties.
Kalle Åström
ICPR1
2000 Automatic geometric reasoning in structure and motion estimatio
Magnus Oskarsson, Kalle Åström
Pattern Recognit. Lett.2
1999 Structure and Motion from Lines under Affine Projections
Kalle Åström, Anders Heyden, Fredrik Kahl, Magnus Oskarsson
ICCV1
1999 Reconstruction of Curves in R3, using Factorization and Bundle Adjustment
abstract
In this paper we extend the notion of affine shape, introduced by Sparr (1995, 1996), from finite point sets to curves. The extension makes it possible to reconstruct 3D-curves up to projective transformations, from a number of their 2D-projections. We also extend the bundle adjustment technique from point features to curves. The first step of the curve reconstruction algorithm is based on affine shape, is independent of choice of coordinates, robust, does not rely on any preselected parameters and works for an arbitrary number of images. In particular this means that a solution is given to the aperture problem of finding point correspondences between curves. The second step takes advantage of any knowledge of measurement errors in the images. This is possible by extending the bundle adjustment technique to curves. Finally, experiments are performed on both synthetic and real data to show the performance and applicability of the algorithm.
Rikard Berthilsson, Kalle Åström, Anders Heyden
ICCV2
1999 Flexible Calibration: Minimal Cases for Auto-Calibration
abstract
This paper deals with the concept of auto-calibration, i.e. methods to calibrate a camera on-line. In particular we deal with minimal conditions on the intrinsic parameters needed to make a Euclidean reconstruction, called flexible calibration. The main theoretical results are that it is only needed to know that one intrinsic parameter is constant. The method is based on an initial projective reconstruction, which is upgraded to a Euclidean one. The number of images needed increases with the complexity of the constraints, but the number of points needed is only the number needed in order to obtain a projective reconstruction. The theoretical results are exemplified in a number of experiments. An algorithm, based on bundle adjustments and a linear initialization method are presented and experiments are performed on both synthetic and real data.
Anders Heyden, Kalle Åström
ICCV2
1999 Generalised Epipolar Constraints
Kalle Åström, Roberto Cipolla, Peter J. Giblin
Int. J. Comput. Vis.1
1999 Motion Estimation in Image Sequences Using the Deformation of Apparent Contours
abstract
The problem of determining the camera motion from apparent contours or silhouettes of a priori unknown curved 3D surfaces is considered. In a sequence of images, it is shown how to use the generalized epipolar constraint on apparent contours. One such constraint is obtained for each epipolar tangency point in each image pair. An accurate algorithm for computing the motion is presented based on a maximum likelihood estimate. It is shown how to generate initial estimates on the camera motion using only the tracked contours. It is also shown that in theory the motion can be calculated from the deformation of a single contour. The algorithm has been tested on several real image sequences, for both Euclidean and projective reconstruction. The resulting motion estimate is compared to motion estimates calculated independently using standard feature-based methods. The motion estimate is also used to classify the silhouettes as curves or apparent contours. The statistical evaluation shows that the technique gives accurate and stable results.
Kalle Åström, Fredrik Kahl
IEEE Trans. Pattern Anal. Mach. Intell.1
1998 Minimal Conditions on Intrinsic Parameters for Euclidean Reconstruction
Anders Heyden, Kalle Åström
ACCV (2)2
1998 Motion Estimation in Image Sequences Using the Deformation of Apparent Contours
abstract
The problem of determining the camera motion from apparent contours or silhouettes of curved three-dimensional surfaces is considered. In a sequence of images is shown how to use the generalized epipolar constraint on apparent contours. One such constraint is obtained for each epipolar tangency point in each image pair. Thus in theory the motion can be calculated from the deformation of a single contour. A robust algorithm for computing the motion is presented based on the maximum likelihood estimate. It is shown how to generate initial estimates on the camera motion using only the tracked contours. It is also shown how to improve this estimate by maximizing the likelihood function. The algorithm has been tested on real image sequences. The result is compared to that of using only point features. The statistical evaluation shows that the technique gives accurate and stable results.
Fredrik Kahl, Kalle Åström
ICCV2
1998 Continuous Time Matching Constraints for Image Streams
Kalle Åström, Anders Heyden
Int. J. Comput. Vis.1
1997 Reconstruction of 3D-Curves from 2D-Images Using Affine Shape Methods for Curves
abstract
In this paper, we propose an algorithm for doing reconstruction of general 3D-curves from a number of 2D-images taken by uncalibrated cameras. No point correspondences between the images are assumed. The curve and the view points are uniquely reconstructed, modulo common projective transformations and the point correspondence problem is solved. Furthermore, the algorithm is independent of the choice of coordinates, as it is based on orthogonal projections and aligning subspaces. The algorithm is based on an extension of affine shape of finite point configurations to more general objects.
Rikard Berthilsson, Kalle Åström
CVPR2
1997 Euclidean Reconstruction from Image Sequences with Varying and Unknown Focal Length and Principal Point
abstract
The special case of reconstruction from image sequences taken by cameras with skew equal to 0 and aspect ratio equal to 1 has been treated. These type of cameras, here called cameras with Euclidean image planes, represent rigid projections where neither the principal point nor the focal length is known, it is shown that it is possible to reconstruct an unknown object from images taken by a camera with Euclidean image plane up to similarity transformations, i.e., Euclidean transformations plus changes in the global scale. An algorithm, using bundle adjustment techniques, has been implemented. The performance of the algorithm is shown on simulated data.
Anders Heyden, Kalle Åström
CVPR2
1997 Reply to Pizlo, Rosenfeld, and Weiss
Kalle Åström
Comput. Vis. Image Underst.1
1997 Simplifications of multilinear forms for sequences of images
Anders Heyden, Kalle Åström
Image Vis. Comput.2
1996 Multilinear Constraints in the Infinitesimal-time Cas
abstract
In this paper we study the infinitesimal-time case of the so called multilinear constraints that exist for each subsequence in a sequence of images. These constraints link the infinitesimal motion of the image points with the infinitesimal viewer motion. The analysis is done both for calibrated and uncalibrated cameras. Two simplifications are also presented for the uncalibrated camera case. One simplification is made using affine reduction and kinetic depth. The second simplification is based upon a projective reduction with respect to the image of a planar patch.
Kalle Åström, Anders Heyden
CVPR1
1996 Generalised Epipolar Constraints
Kalle Åström, Roberto Cipolla, Peter J. Giblin
ECCV (2)1
1996 Algebraic Varieties in Multiple View Geometry
Anders Heyden, Kalle Åström
ECCV (2)2
1996 Stochastic modelling and analysis of sub-pixel edge detection
abstract
Stochastic analysis of edge detectors can be made either by theoretical modeling of the image formation process and the edge detectors or by empirical stochastic analysis of the edge locations. In this paper we study and model the image formation process in detail. In particular the much neglected discretisation process is modelled and taken into account. This makes it possible to define and analyse sub-pixel edge detection. The theoretical results are verified through stochastic analysis of both simulated and real image data.
Kalle Åström, Anders Heyden
ICPR1
1996 Stochastic analysis of scale-space smoothing
abstract
In the high-level operations of computer vision it is taken for granted that image features have been reliably detected. This paper addresses the problem of feature extraction by scale-space methods. This paper is based on two key ideas: to investigate the stochastic properties of scale-space representations, and to investigate the interplay between discrete and continuous images. These investigations are then used to predict the stochastic properties of sub-pixel feature detectors.
Kalle Åström, Anders Heyden
ICPR1
1996 Euclidean reconstruction from constant intrinsic parameters
abstract
A new method for Euclidean reconstruction from sequences of images taken by uncalibrated cameras, with constant intrinsic parameters, is described. Our approach leads to a variant of the so called Kruppa equations. It is shown that it is possible to calculate the intrinsic parameters as well as the Euclidean reconstruction from at least three images. The novelty of our approach is that we build our calculation on a projective reconstruction obtained without the assumption on constant intrinsic parameters. This assumption simplifies the analysis, because a projective reconstruction is already obtained and we need "only" to find the correct Euclidean reconstruction among all possible projective reconstructions.
Anders Heyden, Kalle Åström
ICPR2
1995 Motion from the Frontier of Curved Surfaces
abstract
The frontier of a curved surface is the envelope of contour generators showing the boundary, at least locally, of the visible region swept out under viewer motion. In general, the outlines of curved surfaces (apparent contours) from different viewpoints are generated by different contour generators on the surface and hence do not provide a constraint on viewer motion. We show that frontier points, however, have projections which correspond to a real point on the surface and can be used to constrain viewer motion by the epipolar constraint. We show how to recover viewer motion from frontier points for both continuous and discrete motion, calibrated and uncalibrated cameras. We present preliminary results of an iterative scheme to recover the epipolar line structure from real image sequences using only the outlines of curved surfaces. A statistical evaluation as also performed to estimate the stability of the solution.>
Roberto Cipolla, Kalle Åström, Peter J. Giblin
ICCV2
1995 Fundamental Limitations on Projective Invariants of Planar Curves
abstract
In this paper, some fundamental limitations of projective invariants of non-algebraic planar curves are discussed. It is shown that all curves within a large class can be mapped arbitrarily close to a circle by projective transformations. It is also shown that arbitrarily close to each of a finite number of closed planar curves there is one member of a set of projectively equivalent curves. Thus a continuous projective invariant on closed curves is constant. This also limits the possibility of finding so called projective normalisation schemes for closed planar curves.>
Kalle Åström
IEEE Trans. Pattern Anal. Mach. Intell.1
1994 Affine and Projective Normalization of Planar Curves and Regions
Kalle Åström
ECCV (2)1