Thomas B. Schön

dblp:85/4891 · DBLP profile ↗
← Back
63ranked-venue papers
2as first author
18since 2021 · last 2025
0000-0001-5183-234XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2025 Conditioning diffusion models by explicit forward-backward bridging
abstract
Given an unconditional diffusion model targeting a joint model $\pi(x, y)$, using it to perform conditional simulation $\pi(x \mid y)$ is still largely an open question and is typically achieved by learning conditional drifts to the denoising SDE after the fact. In this work, we express \emph{exact} conditional simulation within the \emph{approximate} diffusion model as an inference problem on an augmented space corresponding to a partial SDE bridge. This perspective allows us to implement efficient and principled particle Gibbs and pseudo-marginal samplers marginally targeting the conditional distribution $\pi(x \mid y)$. Contrary to existing methodology, our methods do not introduce any additional approximation to the unconditional diffusion model aside from the Monte Carlo error. We showcase the benefits and drawbacks of our approach on a series of synthetic and real data examples.
Adrien Corenflos, Zheng Zhao 0004, Thomas B. Schön, Simo Särkkä, Jens Sjölund
AISTATS3
2025 Efficient Optimization Algorithms for Linear Adversarial Training
abstract
Adversarial training can be used to learn models that are robust against perturbations. For linear models, it can be formulated as a convex optimization problem. Compared to methods proposed in the context of deep learning, leveraging the optimization structure allows significantly faster convergence rates. Still, the use of generic convex solvers can be inefficient for large-scale problems. Here, we propose tailored optimization algorithms for the adversarial training of linear models, which render large-scale regression and classification problems more tractable. For regression problems, we propose a family of solvers based on iterative ridge regression and, for classification, a family of solvers based on projected gradient descent. The methods are based on extended variable reformulations of the original problem. We illustrate their efficiency in numerical examples.
Antônio H. Ribeiro, Thomas B. Schön, Dave Zachariah, Francis R. Bach
AISTATS2
2025 Safe exploration in reproducing kernel Hilbert spaces
abstract
Popular safe Bayesian optimization (BO) algorithms learn control policies for safety-critical systems in unknown environments. However, most algorithms make a smoothness assumption, which is encoded by a known bounded norm in a reproducing kernel Hilbert space (RKHS). The RKHS is a potentially infinite-dimensional space, and it remains unclear how to reliably obtain the RKHS norm of an unknown function. In this work, we propose a safe BO algorithm capable of estimating the RKHS norm from data. We provide statistical guarantees on the RKHS norm estimation, integrate the estimated RKHS norm into existing confidence intervals and show that we retain theoretical guarantees, and prove safety of the resulting safe BO algorithm. We apply our algorithm to safely optimize reinforcement learning policies on physics simulators and on a real inverted pendulum, demonstrating improved performance, safety, and scalability compared to the state-of-the-art.
Abdullah Tokmak, Kiran G. Krishnan, Thomas B. Schön, Dominik Baumann
AISTATS3
2025 Kernel Learning with Adversarial Features: Numerical Efficiency and Adaptive Regularization
abstract
Adversarial training has emerged as a key technique to enhance model robustness against adversarial input perturbations. Many of the existing methods rely on computationally expensive min-max problems that limit their application in practice. We propose a novel formulation of adversarial training in reproducing kernel Hilbert spaces, shifting from input to feature-space perturbations. This reformulation enables the exact solution of inner maximization and efficient optimization. It also provides a regularized estimator that naturally adapts to the noise level and the smoothness of the underlying function. We establish conditions under which the feature-perturbed formulation is a relaxation of the original problem and propose an efficient optimization algorithm based on iterative kernel ridge regression. We provide generalization bounds that help to understand the properties of the method. We also extend the formulation to multiple kernel learning. Empirical evaluation shows good performance in both clean and adversarial settings.
Antônio H. Ribeiro, David Vävinggren, Dave Zachariah, Thomas B. Schön, Francis R. Bach
NeurIPS4
2024 On Feynman-Kac training of partial Bayesian neural networks
abstract
Recently, partial Bayesian neural networks (pBNNs), which only consider a subset of the parameters to be stochastic, were shown to perform competitively with full Bayesian neural networks. However, pBNNs are often multi-modal in the latent variable space and thus challenging to approximate with parametric models. To address this problem, we propose an efficient sampling-based training strategy, wherein the training of a pBNN is formulated as simulating a Feynman-Kac model. We then describe variations of sequential Monte Carlo samplers that allow us to simultaneously estimate the parameters and the latent posterior distribution of this model at a tractable computational cost. Using various synthetic and real-world datasets we show that our proposed training scheme outperforms the state of the art in terms of predictive performance.
Zheng Zhao 0004, Sebastian Mair 0001, Thomas B. Schön, Jens Sjölund
AISTATS3
2024 Controlling Vision-Language Models for Multi-Task Image Restoration
abstract
Vision-language models such as CLIP have shown great impact on diverse downstream tasks for zero-shot or label-free predictions. However, when it comes to low-level vision such as image restoration their performance deteriorates dramatically due to corrupted inputs. In this paper, we present a degradation-aware vision-language model (DA-CLIP) to better transfer pretrained vision-language models to low-level vision tasks as a multi-task framework for image restoration. More specifically, DA-CLIP trains an additional controller that adapts the fixed CLIP image encoder to predict high-quality feature embeddings. By integrating the embedding into an image restoration network via cross-attention, we are able to pilot the model to learn a high-fidelity image reconstruction. The controller itself will also output a degradation feature that matches the real corruptions of the input, yielding a natural classifier for different degradation types. In addition, we construct a mixed degradation dataset with synthetic captions for DA-CLIP training. Our approach advances state-of-the-art performance on both degradation-specific and unified image restoration tasks, showing a promising direction of prompting image restoration with large-scale pretrained vision-language models. Our code is available at https://github.com/Algolzw/daclip-uir.
Ziwei Luo 0002, Fredrik K. Gustafsson, Zheng Zhao 0004, Jens Sjölund, Thomas B. Schön
ICLR5
2024 No Double Descent in Principal Component Regression: A High-Dimensional Analysis
abstract
Understanding the generalization properties of large-scale models necessitates incorporating realistic data assumptions into the analysis. Therefore, we consider Principal Component Regression (PCR)—combining principal component analysis and linear regression—on data from a low-dimensional manifold. We present an analysis of PCR when the data is sampled from a spiked covariance model, obtaining fundamental asymptotic guarantees for the generalization risk of this model. Our analysis is based on random matrix theory and allows us to provide guarantees for high-dimensional data. We additionally present an analysis of the distribution shift between training and test data. The results allow us to disentangle the effects of (1) the number of parameters, (2) the data-generating model and, (3) model misspecification on the generalization risk. The use of PCR effectively regularizes the model and prevents the interpolation peak of the double descent. Our theoretical findings are empirically validated in simulation, demonstrating their practical relevance.
Daniel Gedon, Antônio H. Ribeiro, Thomas B. Schön
ICML3
2024 Online Learning in Motion Modeling for Intra-interventional Image Sequences
Niklas Gunnarsson, Jens Sjölund, Peter Kimstrand, Thomas B. Schön
MICCAI (2)4
2024 Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning
abstract
Diffusion policy has shown a strong ability to express complex action distributions in offline reinforcement learning (RL). However, it suffers from overestimating Q-value functions on out-of-distribution (OOD) data points due to the offline dataset limitation. To address it, this paper proposes a novel entropy-regularized diffusion policy and takes into account the confidence of the Q-value prediction with Q-ensembles. At the core of our diffusion policy is a mean-reverting stochastic differential equation (SDE) that transfers the action distribution into a standard Gaussian form and then samples actions conditioned on the environment state with a corresponding reverse-time process. We show that the entropy of such a policy is tractable and that can be used to increase the exploration of OOD samples in offline RL training. Moreover, we propose using the lower confidence bound of Q-ensembles for pessimistic Q-value function estimation. The proposed approach demonstrates state-of-the-art performance across a range of tasks in the D4RL benchmarks, significantly improving upon existing diffusion-based policies. The code is available at https://github.com/ruoqizzz/entropy-offlineRL.
Ruoqi Zhang, Ziwei Luo 0002, Jens Sjölund, Thomas B. Schön, Per Mattsson
NeurIPS4
2024 Uncertainty Estimation with Recursive Feature Machines
abstract
In conventional regression analysis, predictions are typically represented as point estimates derived from covariates. The Gaussian Process (GP) offer a kernel-based framework that predicts and quantifies associated uncertainties. However, kernel-based methods often underperform ensemble-based decision tree approaches in regression tasks involving tabular and categorical data. Recently, Recursive Feature Machines (RFMs) were proposed as a novel feature-learning kernel which strengthens the capabilities of kernel machines. In this study, we harness the power of these RFMs in a probabilistic GP-based approach to enhance uncertainty estimation through feature extraction within kernel methods. We employ this learned kernel for in-depth uncertainty analysis. On tabular datasets, our RFM-based method surpasses other leading uncertainty estimation techniques, including NGBoost and CatBoost-ensemble. Additionally, when assessing out-of-distribution performance, we found that boosting-based methods are surpassed by our RFM-based approach.
Daniel Gedon, Amirhesam Abedsoltan, Thomas B. Schön, Mikhail Belkin
UAI3
2024 Safe Reinforcement Learning in Uncertain Contexts
abstract
When deploying machine learning algorithms in the real world, guaranteeing safety is an essential asset. Existing safe learning approaches typically consider continuous variables, i.e., regression tasks. However, in practice, robotic systems are also subject to discrete, external environmental changes, e.g., having to carry objects of certain weights or operating on frozen, wet, or dry surfaces. Such influences can be modeled as discretecontextvariables. In the existing literature, such contexts are, if considered, mostly assumed to be known. In this work, we drop this assumption and show how we can perform safe learning when we cannot directly measure the context variables. To achieve this, we derive frequentist guarantees for multiclass classification, allowing us to estimate the current context from measurements. Furthermore, we propose an approach for identifying contexts through experiments. We discuss under which conditions we can retain theoretical guarantees and demonstrate the applicability of our algorithm on a Furuta pendulum with camera measurements of different weights that serve as contexts.
Dominik Baumann, Thomas B. Schön
IEEE Trans. Robotics2
2023 Image Restoration with Mean-Reverting Stochastic Differential Equations
abstract
This paper presents a stochastic differential equation (SDE) approach for general-purpose image restoration. The key construction consists in a mean-reverting SDE that transforms a high-quality image into a degraded counterpart as a mean state with fixed Gaussian noise. Then, by simulating the corresponding reverse-time SDE, we are able to restore the origin of the low-quality image without relying on any task-specific prior knowledge. Crucially, the proposed mean-reverting SDE has a closed-form solution, allowing us to compute the ground truth time-dependent score and learn it with a neural network. Moreover, we propose a maximum likelihood objective to learn an optimal reverse trajectory that stabilizes the training and improves the restoration results. The experiments show that our proposed method achieves highly competitive performance in quantitative comparisons on image deraining, deblurring, and denoising, setting a new state-of-the-art on two deraining datasets. Finally, the general applicability of our approach is further demonstrated via qualitative results on image super-resolution, inpainting, and dehazing. Code is available at https://github.com/Algolzw/image-restoration-sde.
Ziwei Luo 0002, Fredrik K. Gustafsson, Zheng Zhao 0004, Jens Sjölund, Thomas B. Schön
ICML5
2023 Regularization properties of adversarially-trained linear regression
abstract
State-of-the-art machine learning models can be vulnerable to very small input perturbations that are adversarially constructed. Adversarial training is an effective approach to defend against it. Formulated as a min-max problem, it searches for the best solution when the training data were corrupted by the worst-case attacks. Linear models are among the simple models where vulnerabilities can be observed and are the focus of our study. In this case, adversarial training leads to a convex optimization problem which can be formulated as the minimization of a finite sum. We provide a comparative analysis between the solution of adversarial training in linear regression and other regularization methods. Our main findings are that: (A) Adversarial training yields the minimum-norm interpolating solution in the overparameterized regime (more parameters than data), as long as the maximum disturbance radius is smaller than a threshold. And, conversely, the minimum-norm interpolator is the solution to adversarial training with a given radius. (B) Adversarial training can be equivalent to parameter shrinking methods (ridge regression and Lasso). This happens in the underparametrized region, for an appropriate choice of adversarial radius and zero-mean symmetrically distributed covariates. (C) For $\ell_\infty$-adversarial training---as in square-root Lasso---the choice of adversarial radius for optimal bounds does not depend on the additive noise variance. We confirm our theoretical findings with numerical examples.
Antônio H. Ribeiro, Dave Zachariah, Francis R. Bach, Thomas B. Schön
NeurIPS4
2023 Invertible Kernel PCA With Random Fourier Features
abstract
Kernel principal component analysis (kPCA) is a widely studied method to construct a low-dimensional data representation after a nonlinear transformation. The prevailing method to reconstruct the original input signal from kPCA—an important task for denoising—requires us to solve a supervised learning problem. In this paper, we present an alternative method where the reconstruction follows naturally from the compression step. We first approximate the kernel with random Fourier features. Then, we exploit the fact that the nonlinear transformation is invertible in a certain subdomain. Hence, the nameinvertible kernel PCA (ikPCA). We experiment with different data modalities and show that ikPCA performs similarly to kPCA with supervised reconstruction on denoising tasks, making it a strong alternative.
Daniel Gedon, Antônio H. Ribeiro, Niklas Wahlstrom, Thomas B. Schön
IEEE Signal Process. Lett.4
2022 Learning Proposals for Practical Energy-Based Regression
abstract
Energy-based models (EBMs) have experienced a resurgence within machine learning in recent years, including as a promising alternative for probabilistic regression. However, energy-based regression requires a proposal distribution to be manually designed for training, and an initial estimate has to be provided at test-time. We address both of these issues by introducing a conceptually simple method to automatically learn an effective proposal distribution, which is parameterized by a separate network head. To this end, we derive a surprising result, leading to a unified training objective that jointly minimizes the KL divergence from the proposal to the EBM, and the negative log-likelihood of the EBM. At test-time, we can then employ importance sampling with the trained proposal to efficiently evaluate the learned EBM and produce stand-alone predictions. Furthermore, we utilize our derived training objective to learn mixture density networks (MDNs) with a jointly trained energy-based teacher, consistently outperforming conventional MDN training on four real-world regression tasks within computer vision. Code is available at https://github.com/fregu856/ebms_proposals.
Fredrik K. Gustafsson, Martin Danelljan, Thomas B. Schön
AISTATS3
2022 Unsupervised dynamic modeling of medical image transformations
Niklas Gunnarsson, Jens Sjölund, Peter Kimstrand, Thomas B. Schön
FUSION4
2022 Direct Transmittance Estimation in Heterogeneous Participating Media Using Approximated Taylor Expansions
abstract
Evaluating the transmittance between two points along a ray is a key component in solving the light transport through heterogeneous participating media and entails computing an intractable exponential of the integrated medium's extinction coefficient. While algorithms for estimating this transmittance exist, there is a lack of theoretical knowledge about their behaviour, which also prevent new theoretically sound algorithms from being developed. For this purpose, we introduce a new class of unbiased transmittance estimators based on random sampling or truncation of a Taylor expansion of the exponential function. In contrast to classical tracking algorithms, these estimators are non-analogous to the physical light transport process and directly sample the underlying extinction function without performing incremental advancement. We present several versions of the new class of estimators, based on either importance sampling or Russian roulette to provide finite unbiased estimators of the infinite Taylor series expansion. We also show that the well known ratio tracking algorithm can be seen as a special case of the new class of estimators. Lastly, we conduct performance evaluations on both the central processing unit (CPU) and the graphics processing unit (GPU), and the results demonstrate that the new algorithms outperform traditional algorithms for heterogeneous mediums.
Daniel Jönsson, Joel Kronander, Jonas Unger, Thomas B. Schön, Magnus Wrenninge
IEEE Trans. Vis. Comput. Graph.4
2021 How Convolutional Neural Networks Deal with Aliasing
abstract
The convolutional neural network (CNN) remains an essential tool in solving computer vision problems. Standard convolutional architectures consist of stacked layers of operations that progressively downscale the image. Aliasing is a well-known side-effect of downsampling that may take place: it causes high-frequency components of the original signal to become indistinguishable from its low-frequency components. While downsampling takes place in the max-pooling layers or in the strided-convolutions in these models, there is no explicit mechanism that prevents aliasing from taking place in these layers. Due to the impressive performance of these models, it is natural to suspect that they, somehow, implicitly deal with this distortion. The question we aim to answer in this paper is simply: "how and to what extent do CNNs counteract aliasing?" We explore the question by means of two examples: In the first, we assess the CNNs capability of distinguishing oscillations at the input, showing that the redundancies in the intermediate channels play an important role in succeeding at the task; In the second, we show that an image classifier CNN while, in principle, capable of implementing anti-aliasing filters, does not prevent aliasing from taking place in the intermediate layers.
Antônio H. Ribeiro, Thomas B. Schön
ICASSP2
2020 Beyond exploding and vanishing gradients: analysing RNN training using attractors and smoothness
abstract
The exploding and vanishing gradient problem has been the major conceptual principle behind most architecture and training improvements in recurrent neural networks (RNNs) during the last decade. In this paper, we argue that this principle, while powerful, might need some refinement to explain recent developments. We refine the concept of exploding gradients by reformulating the problem in terms of the cost function smoothness, which gives insight into higher-order derivatives and the existence of regions with many close local minima. We also clarify the distinction between vanishing gradients and the need for the RNN to learn attractors to fully use its expressive power. Through the lens of these refinements, we shed new light on recent developments in the RNN field, namely stable RNN and unitary (or orthogonal) RNNs.
Antônio H. Ribeiro, Koen Tiels, Luis Antonio Aguirre, Thomas B. Schön
AISTATS4
2020 How to Train Your Energy-Based Model for Regression
Fredrik Gustafsson, Martin Danelljan, Radu Timofte, Thomas B. Schön
BMVC4
2020 Energy-Based Models for Deep Probabilistic Regression
Fredrik K. Gustafsson, Martin Danelljan, Goutam Bhat, Thomas B. Schön
ECCV (20)4
2020 Particle Filter with Rejection Control and Unbiased Estimator of the Marginal Likelihood
abstract
We consider the combined use of resampling and partial rejection control in sequential Monte Carlo methods, also known as particle filters. While the variance reducing properties of rejection control are known, there has not been (to the best of our knowledge) any work on unbiased estimation of the marginal likelihood (also known as the model evidence or the normalizing constant) in this type of particle filter. Being able to estimate the marginal likelihood without bias is highly relevant for model comparison, computation of interpretable and reliable confidence intervals, and in exact approximation methods, such as particle Markov chain Monte Carlo. In the paper we present a particle filter with rejection control that enables unbiased estimation of the marginal likelihood.
Jan Kudlicka, Lawrence Murray, Thomas B. Schön, Fredrik Lindsten
ICASSP3
2019 Conditionally Independent Multiresolution Gaussian Processes
abstract
The multiresolution Gaussian process (GP) has gained increasing attention as a viable approach towards improving the quality of approximations in GPs that scale well to large-scale data. Most of the current constructions assume full independence across resolutions. This assumption simplifies the inference, but it underestimates the uncertainties in transitioning from one resolution to another. This in turn results in models which are prone to overfitting in the sense of excessive sensitivity to the chosen resolution, and predictions which are non-smooth at the boundaries. Our contribution is a new construction which instead assumes conditional independence among GPs across resolutions. We show that relaxing the full independence assumption enables robustness against overfitting, and that it delivers predictions that are smooth at the boundaries. Our new model is compared against current state of the art on 2 synthetic and 9 real-world datasets. In most cases, our new conditionally independent construction performed favorably when compared against models based on the full independence assumption. In particular, it exhibits little to no signs of overfitting.
Jalil Taghia, Thomas B. Schön
AISTATS2
2019 Evaluating model calibration in classification
abstract
Probabilistic classifiers output a probability distribution on target classes rather than just a class prediction. Besides providing a clear separation of prediction and decision making, the main advantage of probabilistic models is their ability to represent uncertainty about predictions. In safety-critical applications, it is pivotal for a model to possess an adequate sense of uncertainty, which for probabilistic classifiers translates into outputting probability distributions that are consistent with the empirical frequencies observed from realized outcomes. A classifier with such a property is called calibrated. In this work, we develop a general theoretical calibration evaluation framework grounded in probability theory, and point out subtleties present in model calibration evaluation that lead to refined interpretations of existing evaluation techniques. Lastly, we propose new ways to quantify and visualize miscalibration in probabilistic classification, including novel multidimensional reliability diagrams.
Juozas Vaicenavicius, David Widmann, Carl R. Andersson, Fredrik Lindsten, Jacob Roll, Thomas B. Schön
AISTATS6
2019 Inferring Heterogeneous Causal Effects in Presence of Spatial Confounding
abstract
We address the problem of inferring the causal effect of an exposure on an outcome across space, using observational data. The data is possibly subject to unmeasured confounding variables which, in a standard approach, must be adjusted for by estimating a nuisance function. Here we develop a method that eliminates the nuisance function, while mitigating the resulting errors-in-variables. The result is a robust and accurate inference method for spatially varying heterogeneous causal effects. The properties of the method are demonstrated on synthetic as well as real data from Germany and the US.
Muhammad Osama 0001, Dave Zachariah, Thomas B. Schön
ICML3
2019 Robust exploration in linear quadratic reinforcement learning
abstract
Learning to make decisions in an uncertain and dynamic environment is a task of fundamental performance in a number of domains. This paper concerns the problem of learning control policies for an unknown linear dynamical system so as to minimize a quadratic cost function. We present a method, based on convex optimization, that accomplishes this task ‘robustly’, i.e., the worst-case cost, accounting for system uncertainty given the observed data, is minimized. The method balances exploitation and exploration, exciting the system in such a way so as to reduce uncertainty in the model parameters to which the worst-case cost is most sensitive. Numerical simulations and application to a hardware-in-the-loop servo-mechanism are used to demonstrate the approach, with appreciable performance and robustness gains over alternative methods observed in both.
Jack Umenberger, Mina Ferizbegovic, Thomas B. Schön, Håkan Hjalmarsson
NeurIPS3
2019 Probabilistic Programming for Birth-Death Models of Evolution Using an Alive Particle Filter with Delayed Sampling
Jan Kudlicka, Lawrence Murray, Fredrik Ronquist, Thomas B. Schön
UAI4
2019 A Fast and Robust Algorithm for Orientation Estimation Using Inertial Sensors
abstract
We present a novel algorithm for online, real-time orientation estimation. Our algorithm integrates gyroscope data and corrects the resulting orientation estimate for integration drift using accelerometer and magnetometer data. This correction is computed, at each time instance, using a single gradient descent step with fixed step length. This fixed step length results in robustness against model errors, e.g., caused by large accelerations or by short-term magnetic field disturbances, which we numerically illustrate using Monte Carlo simulations. Our algorithm estimates a three-dimensional update to the orientation rather than the entire orientation itself. This reduces the computational complexity by approximately 1/3 with respect to the state of the art. It also improves the quality of the resulting estimates, specifically when the orientation corrections are large. We illustrate the efficacy of the algorithm using experimental data.
Manon Kok, Thomas B. Schön
IEEE Signal Process. Lett.2
2018 Delayed Sampling and Automatic Rao-Blackwellization of Probabilistic Programs
abstract
We introduce a dynamic mechanism for the solution of analytically-tractable substructure in probabilistic programs, using conjugate priors and affine transformations to reduce variance in Monte Carlo estimators. For inference with Sequential Monte Carlo, this automatically yields improvements such as locally-optimal proposals and Rao–Blackwellization. The mechanism maintains a directed graph alongside the running program that evolves dynamically as operations are triggered upon it. Nodes of the graph represent random variables, edges the analytically-tractable relationships between them. Random variables remain in the graph for as long as possible, to be sampled only when they are used by the program in a way that cannot be resolved analytically. In the meantime, they are conditioned on as many observations as possible. We demonstrate the mechanism with a few pedagogical examples, as well as a linear-nonlinear state-space model with simulated data, and an epidemiological model with real data of a dengue outbreak in Micronesia. In all cases one or more variables are automatically marginalized out to significantly reduce variance in estimates of the marginal likelihood, in the final case facilitating a random-weight or pseudo-marginal-type importance sampler for parameter estimation. We have implemented the approach in Anglican and a new probabilistic programming language called Birch.
Lawrence Murray, Daniel Lundén, Jan Kudlicka, David Broman, Thomas B. Schön
AISTATS5
2018 Auxiliary-Particle-Filter-Based Two-Filter Smoothing for Wiener State-Space Models
abstract
In this paper, we propose an auxiliary-particle-filter-based two-filter smoother for Wiener state-space models. The proposed smoother exploits the model structure in order to obtain an analytical solution for the backward dynamics, which is introduced artificially in other two-filter smoothers. Furthermore, Gaussian approximations to the optimal proposal density and the adjustment multipliers are derived for both the forward and backward filters. The proposed algorithm is evaluated and compared to existing smoothing algorithms in a numerical example where it is shown that it performs similarly to the state of the art in terms of the root mean squared error at lower computational cost for large numbers of particles.
Roland Hostettler, Thomas B. Schön
FUSION2
2018 Learning Localized Spatio-Temporal Models From Streaming Data
abstract
We address the problem of predicting spatio-temporal processes with temporal patterns that vary across spatial regions, when data is obtained as a stream. That is, when the training dataset is augmented sequentially. Specifically, we develop a localized spatio-temporal covariance model of the process that can capture spatially varying temporal periodicities in the data. We then apply a covariance-fitting methodology to learn the model parameters which yields a predictor that can be updated sequentially with each new data point. The proposed method is evaluated using both synthetic and real climate data which demonstrate its ability to accurately predict data missing in spatial regions over time.
Muhammad Osama 0001, Dave Zachariah, Thomas B. Schön
ICML3
2018 Learning convex bounds for linear quadratic control policy synthesis
abstract
Learning to make decisions from observed data in dynamic environments remains a problem of fundamental importance in a numbers of fields, from artificial intelligence and robotics, to medicine and finance. This paper concerns the problem of learning control policies for unknown linear dynamical systems so as to maximize a quadratic reward function. We present a method to optimize the expected value of the reward over the posterior distribution of the unknown system parameters, given data. The algorithm involves sequential convex programing, and enjoys reliable local convergence and robust stability guarantees. Numerical simulations and stabilization of a real-world inverted pendulum are used to demonstrate the approach, with strong performance and robustness properties observed in both.
Jack Umenberger, Thomas B. Schön
NeurIPS2
2018 Modeling and Interpolation of the Ambient Magnetic Field by Gaussian Processes
abstract
Anomalies in the ambient magnetic field can be used as features in indoor positioning and navigation. By using Maxwell's equations, we derive and present a Bayesian nonparametric probabilistic modeling approach for interpolation and extrapolation of the magnetic field. We model the magnetic field components jointly by imposing a Gaussian process (GP) prior to the latent scalar potential of the magnetic field. By rewriting the GP model in terms of a Hilbert space representation, we circumvent the computational pitfalls associated with GP modeling and provide a computationally efficient and physically justified modeling tool for the ambient magnetic field. The model allows for sequential updating of the estimate and time-dependent changes in the magnetic field. The model is shown to work well in practice in different applications. We demonstrate mapping of the magnetic field both with an inexpensive Raspberry Pi powered robot and on foot using a standard smartphone.
Arno Solin, Manon Kok, Niklas Wahlstrom, Thomas B. Schön, Simo Särkkä
IEEE Trans. Robotics4
2017 Prediction Performance After Learning in Gaussian Process Regression
abstract
This paper considers the quantification of the prediction performance in Gaussian process regression. The standard approach is to base the prediction error bars on the theoretical predictive variance, which is a lower bound on the mean square-error (MSE). This approach, however, does not take into account that the statistical model is learned from the data. We show that this omission leads to a systematic underestimation of the prediction errors. Starting from a generalization of the Cramér-Rao bound, we derive a more accurate MSE bound which provides a measure of uncertainty for prediction of Gaussian processes. The improved bound is easily computed and we illustrate it using synthetic and real data examples.
Johan Wågberg, Dave Zachariah, Thomas B. Schön, Petre Stoica
AISTATS3
2017 Linearly constrained Gaussian processes
abstract
We consider a modification of the covariance function in Gaussian processes to correctly account for known linear constraints. By modelling the target function as a transformation of an underlying function, the constraints are explicitly incorporated in the model such that they are guaranteed to be fulfilled by any sample drawn or prediction made. We also propose a constructive procedure for designing the transformation operator and illustrate the result on both simulated and real-data examples.
Carl Jidling, Niklas Wahlstrom, Adrian Wills, Thomas B. Schön
NIPS4
2016 Computationally Efficient Bayesian Learning of Gaussian Process State Space Models
abstract
Gaussian processes allow for flexible specification of prior assumptions of unknown dynamics in state space models. We present a procedure for efficient Bayesian learning in Gaussian process state space models, where the representation is formed by projecting the problem onto a set of approximate eigenfunctions derived from the prior covariance structure. Learning under this family of models can be conducted using a carefully crafted particle MCMC algorithm. This scheme is computationally efficient and yet allows for a fully Bayesian treatment of the problem. Compared to conventional system identification tools or existing learning methods, we show competitive performance and reliable quantification of uncertainties in the model.
Andreas Svensson, Arno Solin, Simo Särkkä, Thomas B. Schön
AISTATS4
2016 Using Convolution to Estimate the Score Function for Intractable State-Transition Models
abstract
This letter revisits a recent result for the score function estimation by making use of a basic property of the convolution operator. With a probabilistic view of the convolution property, the score function can be estimated as a function of the mean of the parameter posterior distribution, when a pseudoprior distribution for the parameter is introduced. In this letter, the pseudoprior for the parameter is assumed to be Gaussian, which is a common choice in practice. In the end, a toy numerical experiment is implemented to show the efficacy of the estimator, where a particle Markov chain Monte Carlo (MCMC) method is used to sample from the parameter posterior distribution.
Liang Dai 0002, Thomas B. Schön
IEEE Signal Process. Lett.2
2015 Nested Sequential Monte Carlo Methods
abstract
We propose nested sequential Monte Carlo (NSMC), a methodology to sample from sequences of probability distributions, even where the random variables are high-dimensional. NSMC generalises the SMC framework by requiring only approximate, properly weighted, samples from the SMC proposal distribution, while still resulting in a correct SMC algorithm. Furthermore, NSMC can in itself be used to produce such properly weighted samples. Consequently, one NSMC sampler can be used to construct an efficient high-dimensional proposal distribution for another NSMC sampler, and this nesting of the algorithm can be done to an arbitrary degree. This allows us to consider complex and high-dimensional models using SMC. We show results that motivate the efficacy of our approach on several filtering problems with dimensions in the order of 100 to 1000.
Christian A. Naesseth, Fredrik Lindsten, Thomas B. Schön
ICML3
2015 On the Exponential Convergence of the Kaczmarz Algorithm
abstract
The Kaczmarz algorithm (KA) is a popular method for solving a system of linear equations. In this note we derive a new exponential convergence result for the KA. The key allowing us to establish the new result is to rewrite the KA in such a way that its solution path can be interpreted as the output from a particular dynamical system. The asymptotic stability results of the corresponding dynamical system can then be leveraged to prove exponential convergence of the KA. The new bound is also compared to existing bounds.
Liang Dai 0002, Thomas B. Schön
IEEE Signal Process. Lett.2
2014 Capacity estimation of two-dimensional channels using Sequential Monte Carlo
abstract
We derive a new Sequential-Monte-Carlo-based algorithm to estimate the capacity of two-dimensional channel models. The focus is on computing the noiseless capacity of the 2-D (1, ∞) run-length limited constrained channel, but the underlying idea is generally applicable. The proposed algorithm is profiled against a state-of-the-art method, yielding more than an order of magnitude improvement in estimation accuracy for a given computation time.
Christian A. Naesseth, Fredrik Lindsten, Thomas B. Schön
ITW3
2014 Detecting and positioning overtaking vehicles using 1D optical flow
abstract
We are concerned with the problem of detecting an overtaking vehicle using a single camera mounted behind the ego-vehicle windscreen. The proposed solution makes use of 1D optical flow evaluated along lines parallel to the motion of the overtaking vehicles. The 1D optical flow is computed by tracking features along these lines. Based on these features, the position of the overtaking vehicle can also be estimated. The proposed solution has been implemented and tested in real time with promising results. The video data was recorded during test drives in normal traffic conditions in Sweden and Germany.
Daniel Hultqvist, Jacob Roll, Fredrik Svensson, Johan Dahlin, Thomas B. Schön
Intelligent Vehicles Symposium5
2014 Sequential Monte Carlo for Graphical Models
Christian A. Naesseth, Fredrik Lindsten, Thomas B. Schön
NIPS3
2014 Particle gibbs with ancestor sampling
Fredrik Lindsten, Michael I. Jordan, Thomas B. Schön
J. Mach. Learn. Res.3
2013 Particle metropolis hastings using Langevin dynamics
abstract
Particle Markov Chain Monte Carlo (PMCMC) samplers allow for routine inference of parameters and states in challenging nonlinear problems. A common choice for the parameter proposal is a simple random walk sampler, which can scale poorly with the number of parameters. In this paper, we propose to use log-likelihood gradients, i.e. the score, in the construction of the proposal, akin to the Langevin Monte Carlo method, but adapted to the PMCMC framework. This can be thought of as a way to guide a random walk proposal by using drift terms that are proportional to the score function. The method is successfully applied to a stochastic volatility model and the drift term exhibits intuitive behaviour.
Johan Dahlin, Fredrik Lindsten, Thomas B. Schön
ICASSP3
2013 MEMS-based inertial navigation based on a magnetic field map
abstract
This paper presents an approach for 6D pose estimation where MEMS inertial measurements are complemented with magnetometer measurements assuming that a model (map) of the magnetic field is known. The resulting estimation problem is solved using a Rao-Blackwellized particle filter. In our experimental study the magnetic field is generated by a magnetic coil giving rise to a magnetic field that we can model using analytical expressions. The experimental results show that accurate position estimates can be obtained in the vicinity of the coil, where the magnetic field is strong.
Manon Kok, Niklas Wahlstrom, Thomas B. Schön, Fredrik Gustafsson
ICASSP3
2013 Rao-Blackwellized particle smoothers for mixed linear/nonlinear state-space models
abstract
We consider the smoothing problem for a class of conditionally linear Gaussian state-space (CLGSS) models, referred to as mixed linear/nonlinear models. In contrast to the better studied hierarchical CLGSS models, these allow for an intricate cross dependence between the linear and the nonlinear parts of the state vector. We derive a Rao-Blackwellized particle smoother (RBPS) for this model class by exploiting its tractable substructure. The smoother is of the forward filtering/backward simulation type. A key feature of the proposed method is that, unlike existing RBPS for this model class, the linear part of the state vector is marginalized out in both the forward direction and in the backward direction.
Fredrik Lindsten, Pete Bunch, Simon J. Godsill, Thomas B. Schön
ICASSP4
2013 Adaptive stopping for fast particle smoothing
abstract
Particle smoothing is useful for offline state inference and parameter learning in nonlinear/non-Gaussian state-space models. However, many particle smoothers, such as the popular forward filter/backward simulator (FFBS), are plagued by a quadratic computational complexity in the number of particles. One approach to tackle this issue is to use rejection-sampling-based FFBS (RS-FFBS), which asymptotically reaches linear complexity. In practice, however, the constants can be quite large and the actual gain in computational time limited. In this contribution, we develop a hybrid method, governed by an adaptive stopping rule, in order to exploit the benefits, but avoid the drawbacks, of RS-FFBS. The resulting particle smoother is shown in a simulation study to be considerably more computationally efficient than both FFBS and RS-FFBS.
Ehsan Taghavi, Fredrik Lindsten, Lennart Svensson, Thomas B. Schön
ICASSP4
2013 Modeling magnetic fields using Gaussian processes
abstract
Starting from the electromagnetic theory, we derive a Bayesian non-parametric model allowing for joint estimation of the magnetic field and the magnetic sources in complex environments. The model is a Gaussian process which exploits the divergence- and curl-free properties of the magnetic field by combining well-known model components in a novel manner. The model is estimated using magnetometer measurements and spatial information implicitly provided by the sensor. The model and the associated estimator are validated on both simulated and real world experimental data producing Bayesian nonparametric maps of magnetized objects.
Niklas Wahlstrom, Manon Kok, Thomas B. Schön, Fredrik Gustafsson
ICASSP3
2013 Bayesian Inference and Learning in Gaussian Process State-Space Models with Particle MCMC
abstract
State-space models are successfully used in many areas of science, engineering and economics to model time series and dynamical systems. We present a fully Bayesian approach to inference and learning in nonlinear nonparametric state-space models. We place a Gaussian process prior over the transition dynamics, resulting in a flexible model able to capture complex dynamical phenomena. However, to enable efficient inference, we marginalize over the dynamics of the model and instead infer directly the joint smoothing distribution through the use of specially tailored Particle Markov Chain Monte Carlo samplers. Once an approximation of the smoothing distribution is computed, the state transition predictive distribution can be formulated analytically. We make use of sparse Gaussian process models to greatly reduce the computational complexity of the approach.
Roger Frigola, Fredrik Lindsten, Thomas B. Schön, Carl E. Rasmussen
NIPS3
2012 On mixture reduction for multiple target tracking
Tohid Ardeshiri, Umut Orguner, Christian Lundquist, Thomas B. Schön
FUSION4
2012 Calibration of a magnetometer in combination with inertial sensors
Manon Kok, Jeroen D. Hol, Thomas B. Schön, Fredrik Gustafsson, Henk Luinge
FUSION3
2012 On the use of backward simulation in the particle Gibbs sampler
abstract
The particle Gibbs (PG) sampler was introduced in [1] as a way to incorporate a particle filter (PF) in a Markov chain Monte Carlo (MCMC) sampler. The resulting method was shown to be an efficient tool for joint Bayesian parameter and state inference in nonlinear, non-Gaussian state-space models. However, the mixing of the PG kernel can be very poor when there is severe degeneracy in the PF. Hence, the success of the PG sampler heavily relies on the, often unrealistic, assumption that we can implement a PF without suffering from any considerate degeneracy. However, as pointed out by Whiteley [2] in the discussion following [1], the mixing can be improved by adding a backward simulation step to the PG sampler. Here, we investigate this further, derive an explicit PG sampler with backward simulation (denoted PG-BSi) and show that this indeed is a valid MCMC method. Furthermore, we show in a numerical example that backward simulation can lead to a considerable increase in performance over the standard PG sampler.
Fredrik Lindsten, Thomas B. Schön
ICASSP2
2012 Ancestor Sampling for Particle Gibbs
abstract
We present a novel method in the family of particle MCMC methods that we refer to as particle Gibbs with ancestor sampling (PG-AS). Similarly to the existing PG with backward simulation (PG-BS) procedure, we use backward sampling to (considerably) improve the mixing of the PG kernel. Instead of using separate forward and backward sweeps as in PG-BS, however, we achieve the same effect in a single forward sweep. We apply the PG-AS framework to the challenging class of non-Markovian state-space models. We develop a truncation strategy of these models that is applicable in principle to any backward-simulation-based method, but which is particularly well suited to the PG-AS framework. In particular, as we show in a simulation study, PG-AS can yield an order-of-magnitude improved accuracy relative to PG-BS due to its robustness to the truncation error. Several application examples are discussed, including Rao-Blackwellized particle smoothing and inference in degenerate state-space models.
Fredrik Lindsten, Michael I. Jordan, Thomas B. Schön
NIPS3
2011 Bicycle tracking using ellipse extraction
Tohid Ardeshiri, Fredrik Larsson, Fredrik Gustafsson, Thomas B. Schön, Michael Felsberg
FUSION4
2010 Torchlight Navigation
abstract
A common computer vision task is navigation and mapping. Many indoor navigation tasks require depth knowledge of flat, unstructured surfaces (walls, floor, ceiling). With passive illumination only, this is an ill-posed problem. Inspired by small children using a torchlight, we use a spotlight for active illumination. Using our torchlight approach, depth and orientation estimation of unstructured, flat surfaces boils down to estimation of ellipse parameters. The extraction of ellipses is very robust and requires little computational effort.
Michael Felsberg, Fredrik Larsson, Han Wang 0001, Anders Ynnerman, Thomas B. Schön
ICPR5
2010 Geo-referencing for UAV navigation using environmental classification
abstract
A UAV navigation system relying on GPS is vulnerable to signal failure, making a drift free backup system necessary. We introduce a vision based geo-referencing system that uses pre-existing maps to reduce the long term drift. The system classifies an image according to its environmental content and thereafter matches it to an environmentally classified map over the operational area. This map matching provides a measurement of the absolute location of the UAV, that can easily be incorporated into a sensor fusion framework. Experiments show that the geo-referencing system reduces the long term drift in UAV navigation, enhancing the ability of the UAV to navigate accurately over large areas without the use of GPS.
Fredrik Lindsten, Jonas Callmer, Henrik Ohlsson, David Törnqvist, Thomas B. Schön, Fredrik Gustafsson
ICRA5
2010 Learning to close the loop from 3D point clouds
abstract
This paper presents a new solution to the loop closing problem for 3D point clouds. Loop closing is the problem of detecting the return to a previously visited location, and constitutes an important part of the solution to the Simultaneous Localisation and Mapping (SLAM) problem. It is important to achieve a low level of false alarms, since closing a false loop can have disastrous effects in a SLAM algorithm. In this work, the point clouds are described using features, which efficiently reduces the dimension of the data by a factor of 300 or more. The machine learning algorithm AdaBoost is used to learn a classifier from the features. All features are invariant to rotation, resulting in a classifier that is invariant to rotation. The presented method does neither rely on the discretisation of 3D space, nor on the extraction of lines, corners or planes. The classifier is extensively evaluated on publicly available outdoor and indoor data, and is shown to be able to robustly and accurately determine whether a pair of point clouds is from the same location or not. Experiments show detection rates of 63% for outdoor and 53% for indoor data at a false alarm rate of 0%. Furthermore, the classifier is shown to generalise well when trained on outdoor data and tested on indoor data in a SLAM experiment.
Karl Granström, Thomas B. Schön
IROS2
2008 A new algorithm for calibrating a combined camera and IMU sensor unit
abstract
This paper is concerned with the problem of estimating the relative translation and orientation between an inertial measurement unit and a camera which are rigidly connected. The key is to realise that this problem is in fact an instance of a standard problem within the area of system identification, referred to as a gray-box problem. We propose a new algorithm for estimating the relative translation and orientation, which does not require any additional hardware, except a piece of paper with a checkerboard pattern on it. Furthermore, covariance expressions are provided for all involved estimates. The experimental results shows that the method works well in practice.
Jeroen D. Hol, Thomas B. Schön, Fredrik Gustafsson
ICARCV2
2008 Detecting spurious features using parity space
abstract
Detection of spurious features is instrumental in many computer vision applications. The standard approach is feature based, where extracted features are matched between the image frames. This approach requires only vision, but is computer intensive and not yet suitable for real-time applications. We propose an alternative based on algorithms from the statistical fault detection literature. It is based on image data and an inertial measurement unit (IMU). The principle of analytical redundancy is applied to batches of measurements from a sliding time window. The resulting algorithm is fast and scalable, and requires only feature positions as inputs from the computer vision system. It is also pointed out that the algorithm can be extended to also detect non-stationary features (moving targets for instance). The algorithm is applied to real data from an unmanned aerial vehicle in a navigation application.
David Törnqvist, Thomas B. Schön, Fredrik Gustafsson
ICARCV2
2008 Relative pose calibration of a spherical camera and an IMU
abstract
This paper is concerned with the problem of estimating the relative translation and orientation of an inertial measurement unit and a spherical camera, which are rigidly connected. The key is to realize that this problem is in fact an instance of a standard problem within the area of system identification, referred to as a gray-box problem. We propose a new algorithm for estimating the relative translation and orientation, which does not require any additional hardware, except a piece of paper with a checkerboard pattern on it. The experimental results show that the method works well in practice.
Jeroen D. Hol, Thomas B. Schön, Fredrik Gustafsson
ISMAR2
2007 A framework for simultaneous localization and mapping utilizing model structure
abstract
This contribution aims at unifying two trends in applied particle filtering (PF). The first trend is the major impact in simultaneous localization and mapping (slam) applications, utilizing the FastSLAM algorithm. The second one is the implications of the marginalized particle filter (MPF) or the Rao-Blackwellized particle filter (RBPF) in positioning and tracking applications. An algorithm is introduced, which merges FastSLAM and MPF, and the result is an MPF algorithm for slam applications, where state vectors of higher dimensions can be used. Results using experimental data from a 3D slam development environment, fusing measurements from inertial sensors (accelerometer and gyro) and vision are presented.
Thomas B. Schön, Rickard Karlsson, David Törnqvist, Fredrik Gustafsson
FUSION1
2006 Sensor Fusion for Augmented Reality
abstract
In augmented reality (AR), the position and orientation of the camera have to be estimated with high accuracy and low latency. This nonlinear estimation problem is studied in the present paper. The proposed solution makes use of measurements from inertial sensors and computer vision. These measurements are fused using a Kalman filtering framework, incorporating a rather detailed model for the dynamics of the camera. Experiments show that the resulting filter provides good estimates of the camera motion, even during fast movements
Jeroen D. Hol, Thomas B. Schön, Fredrik Gustafsson, Per J. Slycke
FUSION2
2003 A note on state estimation as a convex optimization problem
abstract
The Kalman filter computes the maximum a posteriori (MAP) estimate of the states for linear state space models with Gaussian noise. We interpret the Kalman filter as the solution to a convex optimization problem, and show that we can generalize the MAP state estimator to any noise with a log-concave density function and any combination of linear equality and convex inequality constraints on the states. We illustrate the principle on a hidden Markov model, where the state vector contains probabilities that are positive and sum to one.
Thomas B. Schön, Fredrik Gustafsson, Anders Hansson
ICASSP (6)1