EDBT 2026 Demo / reviewers in the wild / expert
Matthias W. Seeger
dblp:43/5832
· DBLP profile ↗
43ranked-venue papers
16as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 13 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Theory of computation · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
35 papers |
Optimization for machine learning · 28% Probabilistic and Bayesian machine learning · 26% Trustworthy machine learning · 22% | |
| Theoretical computer science
5 papers |
Algorithmic game theory and mechanism design · 59% Mathematical optimization · 28% Computational geometry · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 50% Distributed systems · 33% Cloud and datacenter computing · 17% |
Topics — the 30 heaviest of 84, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning
hyperparameter optimization |
1.7 | 5 | 2023 | Optimizing Hyperparameters with Conformal Quantile Regression · ICML 2023 Amazon SageMaker Automatic Model Tuning: Scalable Gradient-Free Optimization · KDD 2021 Scalable Hyperparameter Transfer Learning · NeurIPS 2018 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
1.4 | 2 | 2024 | Fortuna: A Library for Uncertainty Quantification in Deep Learning · J. Mach. Learn. Res. 2024 Optimizing Hyperparameters with Conformal Quantile Regression · ICML 2023 |
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
1.2 | 4 | 2021 | BORE: Bayesian Optimization by Density-Ratio Estimation · ICML 2021 Scalable Hyperparameter Transfer Learning · NeurIPS 2018 Bayesian Optimization with Tree-structured Dependencies · ICML 2017 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
1.2 | 4 | 2024 | Fortuna: A Library for Uncertainty Quantification in Deep Learning · J. Mach. Learn. Res. 2024 Bayesian Intermittent Demand Forecasting for Large Inventories · NIPS 2016 Bayesian Inference and Optimal Design for the Sparse Linear Model · J. Mach. Learn. Res. 2008 |
Machine learning › Trustworthy machine learning
calibration |
0.8 | 1 | 2024 | Fortuna: A Library for Uncertainty Quantification in Deep Learning · J. Mach. Learn. Res. 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Explaining Probabilistic Models with Distributional Values · ICML 2024 |
Algorithmic game theory and mechanism design
cooperative game theory |
0.8 | 1 | 2024 | Explaining Probabilistic Models with Distributional Values · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.7 | 6 | 2013 | Fast Dual Variational Inference for Non-Conjugate Latent Gaussian Models · ICML (3) 2013 Large Scale Variational Bayesian Inference for Structured Scale Mixture Models · ICML 2012 Gaussian Covariance and Scalable Variational Inference · ICML 2010 |
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction |
0.7 | 1 | 2023 | Optimizing Hyperparameters with Conformal Quantile Regression · ICML 2023 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.5 | 7 | 2017 | Bayesian Optimization with Tree-structured Dependencies · ICML 2017 Information Consistency of Nonparametric Gaussian Process Methods · IEEE Trans. Inf. Theory 2008 Worst-Case Bounds for Gaussian Process Models · NIPS 2005 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
acquisition function |
0.5 | 1 | 2021 | BORE: Bayesian Optimization by Density-Ratio Estimation · ICML 2021 |
Machine learning › Optimization for machine learning
black-box optimization |
0.5 | 1 | 2021 | BORE: Bayesian Optimization by Density-Ratio Estimation · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
density ratio estimation |
0.5 | 1 | 2021 | BORE: Bayesian Optimization by Density-Ratio Estimation · ICML 2021 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
expected improvement |
0.5 | 1 | 2021 | BORE: Bayesian Optimization by Density-Ratio Estimation · ICML 2021 |
Machine learning › Optimization for machine learning › black-box optimization
zeroth-order optimization |
0.5 | 1 | 2021 | Amazon SageMaker Automatic Model Tuning: Scalable Gradient-Free Optimization · KDD 2021 |
Machine learning › Transfer learning and domain adaptation › transferability estimation
feature transferability |
0.4 | 1 | 2020 | LEEP: A New Measure to Evaluate Transferability of Learned Representations · ICML 2020 |
Machine learning › Representation and self-supervised learning › representation analysis
representation evaluation |
0.4 | 1 | 2020 | LEEP: A New Measure to Evaluate Transferability of Learned Representations · ICML 2020 |
Machine learning › Transfer learning and domain adaptation
transferability estimation |
0.4 | 1 | 2020 | LEEP: A New Measure to Evaluate Transferability of Learned Representations · ICML 2020 |
Machine learning › Learning paradigms › multi-task learning
multi-task transfer learning |
0.3 | 1 | 2018 | Scalable Hyperparameter Transfer Learning · NeurIPS 2018 |
Machine learning › Deep learning architectures and training
state space model |
0.3 | 1 | 2018 | Deep State Space Models for Time Series Forecasting · NeurIPS 2018 |
Data mining › predictive modeling › forecasting
demand prediction |
0.3 | 1 | 2017 | Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017 |
Machine learning and data management
machine learning pipeline |
0.3 | 1 | 2017 | Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017 |
High-performance computing
cluster computing |
0.3 | 1 | 2017 | Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017 |
Distributed systems
distributed machine learning |
0.3 | 1 | 2017 | Probabilistic Demand Forecasting at Scale · Proc. VLDB Endow. 2017 |
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design |
0.3 | 3 | 2009 | Speeding up Magnetic Resonance Image Acquisition by Bayesian Multi-Slice Adaptive Compressed Sensing · NIPS 2009 Bayesian Experimental Design of Magnetic Resonance Imaging Sequences · NIPS 2008 Compressed sensing and Bayesian experimental design · ICML 2008 |
Machine learning › Time series and sequential data › time series modeling
demand forecasting |
0.2 | 1 | 2016 | Bayesian Intermittent Demand Forecasting for Large Inventories · NIPS 2016 |
Machine learning › Trustworthy machine learning › interpretability › explainable AI
probabilistic explanation |
0.2 | 1 | 2024 | Explaining Probabilistic Models with Distributional Values · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
scalable variational inference |
0.2 | 2 | 2010 | Gaussian Covariance and Scalable Variational Inference · ICML 2010 Bayesian Experimental Design of Magnetic Resonance Imaging Sequences · NIPS 2008 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.2 | 5 | 2008 | Cross-Validation Optimization for Large Scale Structured Classification Kernel Methods · J. Mach. Learn. Res. 2008 Cross-Validation Optimization for Large Scale Hierarchical Classification Kernel Methods · NIPS 2006 The Effect of the Input Density Distribution on Kernel-based Classifiers · ICML 2000 |
Medical and health informatics › medical imaging
magnetic resonance imaging |
0.2 | 2 | 2009 | Speeding up Magnetic Resonance Image Acquisition by Bayesian Multi-Slice Adaptive Compressed Sensing · NIPS 2009 Bayesian Experimental Design of Magnetic Resonance Imaging Sequences · NIPS 2008 |
Methods — techniques the papers use, named apart from their topics
gaussian process · 1.5shapley value · 1.5distributional values · 1.5analytic expressions for gaussian/bernoulli/categorical payoffs · 1.5bayesian inference · 1.0bayesian optimization · 1.0conformal prediction · 0.8conformalized quantile regression · 0.7feature engineering · 0.6ensembling · 0.6warm-starting · 0.5random search · 0.5early stopping · 0.5density ratio estimation · 0.5binary classification · 0.5newton-raphson · 0.2kalman smoothing · 0.2operator spectrum · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Explaining Probabilistic Models with Distributional ValuesabstractA large branch of explainable machine learning is grounded in cooperative game theory. However, research indicates that game-theoretic explanations may mislead or be hard to interpret. We argue that often there is a critical mismatch between what one wishes to explain (e.g. the output of a classifier) and what current methods such as SHAP explain (e.g. the scalar probability of a class). This paper addresses such gap for probabilistic models by generalising cooperative games and value operators. We introduce the *distributional values*, random variables that track changes in the model output (e.g. flipping of the predicted class) and derive their analytic expressions for games with Gaussian, Bernoulli and Categorical payoffs. We further establish several characterising properties, and show that our framework provides fine-grained and insightful explanations with case studies on vision and language models. Luca Franceschi 0001, Michele Donini, Cédric Archambeau, Matthias W. Seeger |
ICML | 4 |
| 2024 | Fortuna: A Library for Uncertainty Quantification in Deep LearningabstractWe present Fortuna, an open-source library for uncertainty quantification in deep learning. Fortuna supports a range of calibration techniques, such as conformal prediction that can be applied to any trained neural network to generate reliable uncertainty estimates, and scalable Bayesian inference methods that can be applied to deep neural networks trained from scratch for improved uncertainty quantification and accuracy. By providing a coherent framework for advanced uncertainty quantification methods, Fortuna simplifies the process of benchmarking and helps practitioners build robust AI systems. Gianluca Detommaso, Alberto Gasparin, Michele Donini, Matthias W. Seeger, Andrew Gordon Wilson, Cédric Archambeau |
J. Mach. Learn. Res. | 4 |
| 2023 | Optimizing Hyperparameters with Conformal Quantile RegressionabstractMany state-of-the-art hyperparameter optimization (HPO) algorithms rely on model-based optimizers that learn surrogate models of the target function to guide the search. Gaussian processes are the de facto surrogate model due to their ability to capture uncertainty. However, they make strong assumptions about the observation noise, which might not be warranted in practice. In this work, we propose to leverage conformalized quantile regression which makes minimal assumptions about the observation noise and, as a result, models the target function in a more realistic and robust fashion which translates to quicker HPO convergence on empirical benchmarks. To apply our method in a multi-fidelity setting, we propose a simple, yet effective, technique that aggregates observed results across different resource levels and outperforms conventional methods across many empirical tasks. David Salinas, Jacek Golebiowski, Aaron Klein, Matthias W. Seeger, Cédric Archambeau |
ICML | 4 |
| 2021 | BORE: Bayesian Optimization by Density-Ratio EstimationabstractBayesian optimization (BO) is among the most effective and widely-used blackbox optimization methods. BO proposes solutions according to an explore-exploit trade-off criterion encoded in an acquisition function, many of which are computed from the posterior predictive of a probabilistic surrogate model. Prevalent among these is the expected improvement (EI). The need to ensure analytical tractability of the predictive often poses limitations that can hinder the efficiency and applicability of BO. In this paper, we cast the computation of EI as a binary classification problem, building on the link between class-probability estimation and density-ratio estimation, and the lesser-known link between density-ratios and EI. By circumventing the tractability constraints, this reformulation provides numerous advantages, not least in terms of expressiveness, versatility, and scalability. Louis C. Tiao, Aaron Klein, Matthias W. Seeger, Edwin V. Bonilla, Cédric Archambeau, Fabio Ramos 0001 |
ICML | 3 |
| 2021 | Amazon SageMaker Automatic Model Tuning: Scalable Gradient-Free OptimizationabstractTuning complex machine learning systems is challenging. Machine learning typically requires to set hyperparameters, be it regularization, architecture, or optimization parameters, whose tuning is critical to achieve good predictive performance. To democratize access to machine learning systems, it is essential to automate the tuning. This paper presents Amazon SageMaker Automatic Model Tuning (AMT), a fully managed system for gradient-free optimization at scale. AMT finds the best version of a trained machine learning model by repeatedly evaluating it with different hyperparameter configurations. It leverages either random search or Bayesian optimization to choose the hyperparameter values resulting in the best model, as measured by the metric chosen by the user. AMT can be used with built-in algorithms, custom algorithms, and Amazon SageMaker pre-built containers for machine learning frameworks. We discuss the core functionality, system architecture, our design principles, and lessons learned. We also describe more advanced features of AMT, such as automated early stopping and warm-starting, showing in experiments their benefits to users. Valerio Perrone, Huibin Shen, Aida Zolic, Iaroslav Shcherbatyi, Amr Ahmed 0004, Tanya Bansal, Michele Donini, Fela Winkelmolen, Rodolphe Jenatton, Jean Baptiste Faddoul, Barbara Pogorzelska, Miroslav Miladinovic, Krishnaram Kenthapadi, Matthias W. Seeger, Cédric Archambeau |
KDD | 14 |
| 2021 | A Nonmyopic Approach to Cost-Constrained Bayesian OptimizationabstractBayesian optimization (BO) is a popular method for optimizing expensive-to-evaluate black-box functions. BO budgets are typically given in iterations, which implicitly assumes each evaluation has the same cost. In fact, in many BO applications, evaluation costs vary significantly in different regions of the search space. In hyperparameter optimization, the time spent on neural network training increases with layer size; in clinical trials, the monetary cost of drug compounds vary; and in optimal control, control actions have differing complexities. Cost-constrained BO measures convergence with alternative cost metrics such as time, money, or energy, for which the sample efficiency of standard BO methods is ill-suited. For cost-constrained BO, cost efficiency is far more important than sample efficiency. In this paper, we formulate cost-constrained BO as a constrained Markov decision process (CMDP), and develop an efficient rollout approximation to the optimal CMDP policy that takes both the cost and future iterations into account. We validate our method on a collection of hyperparameter optimization problems as well as a sensor set selection application. Eric Hans Lee, David Eriksson, Valerio Perrone, Matthias W. Seeger |
UAI | 4 |
| 2020 | LEEP: A New Measure to Evaluate Transferability of Learned RepresentationsabstractWe introduce a new measure to evaluate the transferability of representations learned by classifiers. Our measure, the Log Expected Empirical Prediction (LEEP), is simple and easy to compute: when given a classifier trained on a source data set, it only requires running the target data set through this classifier once. We analyze the properties of LEEP theoretically and demonstrate its effectiveness empirically. Our analysis shows that LEEP can predict the performance and convergence speed of both transfer and meta-transfer learning methods, even for small or imbalanced data. Moreover, LEEP outperforms recently proposed transferability measures such as negative conditional entropy and H scores. Notably, when transferring from ImageNet to CIFAR100, LEEP can achieve up to 30% improvement compared to the best competing method in terms of the correlations with actual transfer accuracy. Viet Cuong Nguyen, Tal Hassner, Matthias W. Seeger, Cédric Archambeau |
ICML | 3 |
| 2018 | Scalable Hyperparameter Transfer LearningabstractBayesian optimization (BO) is a model-based approach for gradient-free black-box function optimization, such as hyperparameter optimization. Typically, BO relies on conventional Gaussian process (GP) regression, whose algorithmic complexity is cubic in the number of evaluations. As a result, GP-based BO cannot leverage large numbers of past function evaluations, for example, to warm-start related BO runs. We propose a multi-task adaptive Bayesian linear regression model for transfer learning in BO, whose complexity is linear in the function evaluations: one Bayesian linear regression model is associated to each black-box function optimization problem (or task), while transfer learning is achieved by coupling the models through a shared deep neural net. Experiments show that the neural net learns a representation suitable for warm-starting the black-box optimization problems and that BO runs can be accelerated when the target black-box function (e.g., validation loss) is learned together with other related signals (e.g., training loss). The proposed method was found to be at least one order of magnitude faster that methods recently published in the literature. Valerio Perrone, Rodolphe Jenatton, Matthias W. Seeger, Cédric Archambeau |
NeurIPS | 3 |
| 2018 | Deep State Space Models for Time Series ForecastingabstractWe present a novel approach to probabilistic time series forecasting that combines state space models with deep learning. By parametrizing a per-time-series linear state space model with a jointly-learned recurrent neural network, our method retains desired properties of state space models such as data efficiency and interpretability, while making use of the ability to learn complex patterns from raw data offered by deep learning approaches. Our method scales gracefully from regimes where little training data is available to regimes where data from millions of time series can be leveraged to learn accurate models. We provide qualitative as well as quantitative results with the proposed method, showing that it compares favorably to the state-of-the-art. Syama Sundar Rangapuram, Matthias W. Seeger, Jan Gasthaus, Lorenzo Stella, Yuyang Wang 0001, Tim Januschowski |
NeurIPS | 2 |
| 2017 | Bayesian Optimization with Tree-structured DependenciesabstractBayesian optimization has been successfully used to optimize complex black-box functions whose evaluations are expensive. In many applications, like in deep learning and predictive analytics, the optimization domain is itself complex and structured. In this work, we focus on use cases where this domain exhibits a known dependency structure. The benefit of leveraging this structure is twofold: we explore the search space more efficiently and posterior inference scales more favorably with the number of observations than Gaussian Process-based approaches published in the literature. We introduce a novel surrogate model for Bayesian optimization which combines independent Gaussian Processes with a linear model that encodes a tree-based dependency structure and can transfer information between overlapping decision sequences. We also design a specialized two-step acquisition function that explores the search space more effectively. Our experiments on synthetic tree-structured functions and the tuning of feedforward neural networks trained on a range of binary classification datasets show that our method compares favorably with competing approaches. Rodolphe Jenatton, Cédric Archambeau, Javier González 0002, Matthias W. Seeger |
ICML | 4 |
| 2017 | Probabilistic Demand Forecasting at ScaleabstractWe present a platform built on large-scale, data-centric machine learning (ML) approaches, whose particular focus is demand forecasting in retail. At its core, this platform enables the training and application of probabilistic demand forecasting models, and provides convenient abstractions and support functionality for forecasting problems. The platform comprises of a complex end-to-end machine learning system built on Apache Spark, which includes data preprocessing, feature engineering, distributed learning, as well as evaluation, experimentation and ensembling. Furthermore, it meets the demands of a production system and scales to large catalogues containing millions of items. We describe the challenges of building such a platform and discuss our design decisions. We detail aspects on several levels of the system, such as a set of general distributed learning schemes, our machinery for ensembling predictions, and a high-level dataflow abstraction for modeling complex ML pipelines. To the best of our knowledge, we are not aware of prior work on real-world demand forecasting systems which rivals our approach in terms of scalability. Joos-Hendrik Böse, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, Dustin Lange, David Salinas, Sebastian Schelter, Matthias W. Seeger, Yuyang Wang 0001 |
Proc. VLDB Endow. | 8 |
| 2016 | Bayesian Intermittent Demand Forecasting for Large InventoriesabstractWe present a scalable and robust Bayesian method for demand forecasting in the context of a large e-commerce platform, paying special attention to intermittent and bursty target statistics. Inference is approximated by the Newton-Raphson algorithm, reduced to linear-time Kalman smoothing, which allows us to operate on several orders of magnitude larger problems than previous related work. In a study on large real-world sales datasets, our method outperforms competing approaches on fast and medium moving items. Matthias W. Seeger, David Salinas, Valentin Flunkert |
NIPS | 1 |
| 2015 | Expectation Propagation for Rectified Linear Poisson Regression
Young-Jun Ko, Matthias W. Seeger |
ACML | 2 |
| 2014 | Scalable Collaborative Bayesian Preference LearningabstractLearning about users’ utilities from preference, discrete choice or implicit feedback data is of integral importance in e-commerce, targeted advertising and web search. Due to the sparsity and diffuse nature of data, Bayesian approaches hold much promise, yet most prior work does not scale up to realistic data sizes. We shed light on why inference for such settings is computationally difficult for standard machine learning methods, most of which focus on predicting explicit ratings only. To simplify the difficulty, we present a novel expectation maximization algorithm, driven by expectation propagation approximate inference, which scales to very large datasets without requiring strong factorization assumptions. Our utility model uses both latent bilinear collaborative filtering and non-parametric Gaussian process (GP) regression. In experiments on large real-world datasets, our method gives substantially better results than either matrix factorization or GPs in isolation, and converges significantly faster. Mohammad Emtiyaz Khan, Young-Jun Ko, Matthias W. Seeger |
AISTATS | 3 |
| 2013 | Fast Dual Variational Inference for Non-Conjugate Latent Gaussian ModelsabstractLatent Gaussian models (LGMs) are widely used in statistics and machine learning. Bayesian inference in non-conjugate LGM is difficult due to intractable integrals involving the Gaussian prior and non-conjugate likelihoods. Algorithms based on Variational Gaussian (VG) approximations are widely employed since they strike a favorable balance between accuracy, generality, speed, and ease of use. However, the structure of optimization problems associated with them remains poorly understood, and standard solvers take too long to converge. In this paper, we derive a novel dual variational inference approach, which exploits the convexity property of the VG approximations. The implications of our approach is that we obtain an algorithm that solves a convex optimization problem, reduces the number of variational parameters, and converges much faster than previous methods. Using real world data, we demonstrate these advantages on a variety of LGMs including Gaussian process classification and latent Gaussian Markov random fields. Mohammad Emtiyaz Khan, Aleksandr Y. Aravkin, Michael P. Friedlander, Matthias W. Seeger |
ICML (3) | 4 |
| 2012 | Large Scale Variational Bayesian Inference for Structured Scale Mixture Models
Young-Jun Ko, Matthias W. Seeger |
ICML | 2 |
| 2012 | Information-Theoretic Regret Bounds for Gaussian Process Optimization in the Bandit SettingabstractMany applications require optimizing an unknown, noisy function that is expensive to evaluate. We formalize this task as a multiarmed bandit problem, where the payoff function is either sampled from a Gaussian process (GP) or has low norm in a reproducing kernel Hilbert space. We resolve the important open problem of deriving regret bounds for this setting, which imply novel convergence rates for GP optimization. We analyze an intuitive Gaussian process upper confidence bound (GP-UCB) algorithm, and bound its cumulative regret in terms of maximal in- formation gain, establishing a novel connection between GP optimization and experimental design. Moreover, by bounding the latter in terms of operator spectra, we obtain explicit sublinear regret bounds for many commonly used covariance functions. In some important cases, our bounds have surprisingly weak dependence on the dimensionality. In our experiments on real sensor data, GP-UCB compares favorably with other heuristical GP optimization approaches. Niranjan Srinivas, Andreas Krause 0001, Sham M. Kakade, Matthias W. Seeger |
IEEE Trans. Inf. Theory | 4 |
| 2011 | Large Scale Bayesian Inference and Experimental Design for Sparse Linear ModelsabstractMany problems of low-level computer vision and image processing, such as denoising, deconvolution, tomographic reconstruction or superresolution, can be addressed by maximizing the posterior distribution of a sparse linear model (SLM). We show how higher-order Bayesian decision-making problems, such as optimizing image acquisition in magnetic resonance scanners, can be addressed by querying the SLM posterior covariance, unrelated to the density's mode. We propose a scalable algorithmic framework, with which SLM posteriors over full, high-resolution images can be approximated for the first time, solving a variational optimization problem which is convex if and only if posterior mode finding is convex. These methods successfully drive the optimization of sampling trajectories for real-world magnetic resonance imaging through Bayesian experimental design, which has not been attempted before. Our methodology provides new insight into similarities and differences between sparse reconstruction and approximate Bayesian inference, and has important implications for compressive sensing of real-world images. Parts of this work have been presented at conferences [M. Seeger, H. Nickisch, R. Pohmann, and B. Schölkopf, in Advances in Neural Information Processing Systems 21, D. Koller, D. Schuurmans, Y. Bengio, and L. Bottou, eds., Curran Associates, Red Hook, NY, 2009, pp. 1441–1448; H. Nickisch and M. Seeger, in Proceedings of the 26th International Conference on Machine Learning, L. Bottou and M. Littman, eds., Omni Press, Madison, WI, 2009, pp. 761–768]. Matthias W. Seeger, Hannes Nickisch |
SIAM J. Imaging Sci. | 1 |
| 2010 | Gaussian Covariance and Scalable Variational Inference
Matthias W. Seeger |
ICML | 1 |
| 2010 | Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design
Niranjan Srinivas, Andreas Krause 0001, Sham M. Kakade, Matthias W. Seeger |
ICML | 4 |
| 2009 | Convex variational Bayesian inference for large scale generalized linear modelsabstractWe show how variational Bayesian inference can be implemented for very large generalized linear models. Our relaxation is proven to be a convex problem for any log-concave model. We provide a generic double loop algorithm for solving this relaxation on models with arbitrary super-Gaussian potentials. By iteratively decoupling the criterion, most of the work can be done by solving large linear systems, rendering our algorithm orders of magnitude faster than previously proposed solvers for the same problem. We evaluate our method on problems of Bayesian active learning for large binary classification models, and show how to address settings with many candidates and sequential inclusion steps. Hannes Nickisch, Matthias W. Seeger |
ICML | 2 |
| 2009 | Workshop summary: Numerical mathematics in machine learning
Matthias W. Seeger, Suvrit Sra, John P. Cunningham |
ICML | 1 |
| 2009 | Speeding up Magnetic Resonance Image Acquisition by Bayesian Multi-Slice Adaptive Compressed SensingabstractWe show how to sequentially optimize magnetic resonance imaging measurement designs over stacks of neighbouring image slices, by performing convex variational inference on a large scale non-Gaussian linear dynamical system, tracking dominating directions of posterior covariance without imposing any factorization constraints. Our approach can be scaled up to high-resolution images by reductions to numerical mathematics primitives and parallelization on several levels. In a first study, designs are found that improve significantly on others chosen independently for each slice or drawn at random. Matthias W. Seeger |
NIPS | 1 |
| 2008 | Learning Inverse Dynamics: a Comparison
Duy Nguyen-Tuong, Jan Peters 0001, Matthias W. Seeger, Bernhard Schölkopf |
ESANN | 3 |
| 2008 | Compressed sensing and Bayesian experimental designabstractWe relate compressed sensing (CS) with Bayesian experimental design and provide a novel efficient approximate method for the latter, based on expectation propagation. In a large comparative study about linearly measuring natural images, we show that the simple standard heuristic of measuring wavelet coefficients top-down systematically outperforms CS methods using random measurements; the sequential projection optimisation approach of (Ji & Carin, 2007) performs even worse. We also show that our own approximate Bayesian method is able to learn measurement filters on full images efficiently which outperform the wavelet heuristic. To our knowledge, ours is the first successful attempt at "learning compressed sensing" for images of realistic size. In contrast to common CS methods, our framework is not restricted to sparse signals, but can readily be applied to other notions of signal complexity or noise models. We give concrete ideas how our method can be scaled up to large signal representations. Matthias W. Seeger, Hannes Nickisch |
ICML | 1 |
| 2008 | Local Gaussian Process Regression for Real Time Online Model LearningabstractLearning in real-time applications, e.g., online approximation of the inverse dynamics model for model-based robot control, requires fast online regression techniques. Inspired by local learning, we propose a method to speed up standard Gaussian Process regression (GPR) with local GP models (LGP). The training data is partitioned in local regions, for each an individual GP model is trained. The prediction for a query point is performed by weighted estimation using nearby local models. Unlike other GP approximations, such as mixtures of experts, we use a distance based measure for partitioning of the data and weighted prediction. The proposed method achieves online learning and prediction in real-time. Comparisons with other nonparametric regression methods show that LGP has higher accuracy than LWPR and close to the performance of standard GPR and nu-SVR. Duy Nguyen-Tuong, Matthias W. Seeger, Jan Peters 0001 |
NIPS | 2 |
| 2008 | Bayesian Experimental Design of Magnetic Resonance Imaging SequencesabstractWe show how improved sequences for magnetic resonance imaging can be found through automated optimization of Bayesian design scores. Combining recent advances in approximate Bayesian inference and natural image statistics with high-performance numerical computation, we propose the first scalable Bayesian experimental design framework for this problem of high relevance to clinical and brain research. Our solution requires approximate inference for dense, non-Gaussian models on a scale seldom addressed before. We propose a novel scalable variational inference algorithm, and show how powerful methods of numerical mathematics can be modified to compute primitives in our framework. Our approach is evaluated on a realistic setup with raw data from a 3T MR scanner. Matthias W. Seeger, Hannes Nickisch, Rolf Pohmann, Bernhard Schölkopf |
NIPS | 1 |
| 2008 | Bayesian Inference and Optimal Design for the Sparse Linear Model
Matthias W. Seeger |
J. Mach. Learn. Res. | 1 |
| 2008 | Cross-Validation Optimization for Large Scale Structured Classification Kernel Methods
Matthias W. Seeger |
J. Mach. Learn. Res. | 1 |
| 2008 | Information Consistency of Nonparametric Gaussian Process MethodsabstractBayesian nonparametric models are widely and successfully used for statistical prediction. While posterior consistency properties are well studied in quite general settings, results have been proved using abstract concepts such as metric entropy, and they come with subtle conditions which are hard to validate and not intuitive when applied to concrete models. Furthermore, convergence rates are difficult to obtain. By focussing on the concept of information consistency for Bayesian Gaussian process (GP)models, consistency results and convergence rates are obtained via a regret bound on cumulative log loss. These results depend strongly on the covariance function of the prior process, thereby giving a novel interpretation to penalization with reproducing kernel Hilbert space norms and to commonly used covariance function classes and their parameters. The proof of the main result employs elementary convexity arguments only. A theorem of Widom is used in order to obtain precise convergence rates for several covariance functions widely used in practice. Matthias W. Seeger, Sham M. Kakade, Dean P. Foster |
IEEE Trans. Inf. Theory | 1 |
| 2007 | Bayesian Inference for Sparse Generalized Linear Models
Matthias W. Seeger, Sebastian Gerwinn, Matthias Bethge |
ECML | 1 |
| 2007 | Bayesian Inference for Spiking Neuron Models with a Sparsity PriorabstractGeneralized linear models are the most commonly used tools to describe the stim- ulus selectivity of sensory neurons. Here we present a Bayesian treatment of such models. Using the expectation propagation algorithm, we are able to approximate the full posterior distribution over all weights. In addition, we use a Laplacian prior to favor sparse solutions. Therefore, stimulus features that do not critically influence neural activity will be assigned zero weights and thus be effectively excluded by the model. This feature selection mechanism facilitates both the in- terpretation of the neuron model as well as its predictive abilities. The posterior distribution can be used to obtain confidence intervals which makes it possible to assess the statistical significance of the solution. In neural data analysis, the available amount of experimental measurements is often limited whereas the pa- rameter space is large. In such a situation, both regularization by a sparsity prior and uncertainty estimates for the model parameters are essential. We apply our method to multi-electrode recordings of retinal ganglion cells and use our uncer- tainty estimate to test the statistical significance of functional couplings between neurons. Furthermore we used the sparsity of the Laplace prior to select those filters from a spike-triggered covariance analysis that are most informative about the neural response. Sebastian Gerwinn, Jakob H. Macke, Matthias W. Seeger, Matthias Bethge |
NIPS | 3 |
| 2006 | Cross-Validation Optimization for Large Scale Hierarchical Classification Kernel MethodsabstractWe propose a highly efficient framework for kernel multi-class models with a large and structured set of classes. Kernel parameters are learned automatically by maximizing the cross-validation log likelihood, and predictive probabilities are estimated. We demonstrate our approach on large scale text classification tasks with hierarchical class structure, achieving state-of-the-art results in an order of magnitude less time than previous work. Matthias W. Seeger |
NIPS | 1 |
| 2005 | Worst-Case Bounds for Gaussian Process ModelsabstractWe present a competitive analysis of some non-parametric Bayesian al- gorithms in a worst-case online learning setting, where no probabilistic assumptions about the generation of the data are made. We consider models which use a Gaussian process prior (over the space of all func- tions) and provide bounds on the regret (under the log loss) for com- monly used non-parametric Bayesian algorithms — including Gaussian regression and logistic regression — which show how these algorithms can perform favorably under rather general conditions. These bounds ex- plicitly handle the infinite dimensionality of these non-parametric classes in a natural way. We also make formal connections to the minimax and minimum description length (MDL) framework. Here, we show precisely how Bayesian Gaussian regression is a minimax strategy. Sham M. Kakade, Matthias W. Seeger, Dean P. Foster |
NIPS | 2 |
| 2005 | Fast Gaussian Process Regression using KD-Treesabstract1 Introduction We consider (regression) estimation of a function x u(x ) from noisy observations. If the data-generating process is not well understood, simple parametric learning algorithms, for example ones from the generalized linear model (GLM) family, may be hard to apply because of the difficulty of choosing good features. In contrast, the nonparametric Gaussian process (GP) model [19] offers a flexible and powerful alternative. However, a major drawback of GP models is that the computational cost of learning is about O(n 3 ), and the cost of making a single prediction is O(n), where n is the number of training examples. This high computational complexity severely limits its scalability to large problems, and we believe has proved a significant barrier to the wider adoption of the GP model. In this paper, we address the scaling issue by recognizing that learning and predictions with a GP regression (GPR) model can be implemented using the matrix-vector multiplication (MVM) primitive z K z . Here, K Rn,n is the kernel matrix, and z Rn is an arbitrary vector. For the wide class of so-called isotropic kernels, MVM can be approximated efficiently by arranging the dataset in a tree-type multiresolution data structure such as kd-trees [13], ball trees [11], or cover trees [1]. This approximation can sometimes be made orders of magnitude faster than the direct computation, without sacrificing much in terms of accuracy. Further, the storage requirements for the tree is O(n), while a direct storage of the kernel matrix would require O(n2 ) spare. We demonstrate the efficiency of the tree approach on several large datasets. In the sequel, for the sake of simplicity we will focus on kd-trees (even though it is known that kd-trees do not scale well to high dimensional data). However, it is also completely straightforward to apply the ideas in this paper to other tree-type data structures, for example ball trees and cover trees, which typically scale significantly better to high dimensional data. 2 The Gaussian Process Regression Model Suppose that we observe some data D = {(xi , yi ) | i = 1, . . . , n}, xi X , yi R, sampled independently and identically distributed (i.i.d.) from some unknown distribution. Our goal is to predict the response y on future test points x with small mean-squared error under the data distribution. Our model consists of a latent (unobserved) function x u so that yi = ui + i , where ui = u(xi ), and the i are independent Gaussian noise variables with zero mean and variance 2 > 0. Following the Bayesian paradigm, we place a prior distribution P (u()) on the function u() and use the posterior distribution P (u()|D) N (y |u , 2 I )P (u()) in order to predict y on new points x . Here, y = [y1 , . . . , yn ]T and u = [u1 , . . . , un ]T are vectors in Rn , and N (|, ) is the density of a Gaussian with mean and covariance . For a GPR model, the prior distribution is a (zero-mean) Gaussian process defined in terms of a positive definite kernel (or covariance) function K : X 2 R. For the purposes of this paper, a GP can be thought of as a mapping from arbitrary finite subsets ~ {x i } X of points, to corresponding zero-mean Gaussian distributions with covariance ~ ~ ~~ matrix K = (K (x i , x j ))i,j . (This notation indicates that K is a matrix whose (i, j )~~ element is K (x i , x j ).) In this paper, we focus on the problem of speeding up GPR under the assumption that the kernel is monotonic isotropic. A kernel function K (x , x ) is called isotropic if it depends only on the Euclidean distance r = x - x 2 between the points, and it is monotonic isotropic if it can be written as a monotonic function of r. 3 Fast GPR predictions Since u(x1 ), u(x2 ), . . . , u(xn ) and u(x ) are jointly Gaussian, it is easy to see that the predictive (posterior) distribution P (u |D), u = u(x ) is given by , u (1) P (u |D) = N | kT M -1 y , K (x , x ) - kT M -1 k where k = [K (x , x1 ), . . . , K (x , xn )]T Rn , and M = K + 2 I , K = (K (xi , xj ))i,j . Therefore, if p = M -1 y , the optimal prediction under the model is u = kT p , and the predictive variance (of P (u |D)) can be used to quantify our uncer^ tainty in the prediction. Details can be found in [19]. ([16] also provides a tutorial on GPs.) Once p is determined, making a prediction now requires that we compute in in kT p = K (x , xi )pi = wi p i (2) which is O(n) since it requires scanning through the entire training set and computing K (x , xi ) for each xi in the training set. When the training set is very large, this becomes prohibitively slow. In such situations, it is desirable to use a fast approximation instead of the exact direct implementation. 3.1 Weighted Sum Approximation The computations in Equation 2 can be thought of as a weighted sum, where w i = K (x , xi ) is the weight on the i-th summand pi . We observe that if the dataset is divided into groups where all data points in a group have similar weights, then it is possible to compute a fast approximation to the above weighted sum. For example, let G be a set of data points that all have weights near some value w. The contribution to the weighted sum by points in G is i i i i i wi p i = w pi + (wi - w)pi = w pi + ipi i i i i i i i Where i = wi - w. Assuming that :xi G pi is known in advance, w i :xi G pi can tihen be computed in constant time and used as an approximation to :xi G wi pi if ipi is small. :xi G Yirong Shen, Andrew Y. Ng, Matthias W. Seeger |
NIPS | 3 |
| 2004 | Gaussian Processes For Machine LearningabstractGaussian processes (GPs) are natural generalisations of multivariate Gaussian random variables to infinite (countably or continuous) index sets. GPs have been applied in a large number of fields to a diverse range of ends, and very many deep theoretical analyses of various properties are available. This paper gives an introduction to Gaussian processes on a fairly elementary level with special emphasis on characteristics relevant in machine learning. It draws explicit connections to branches such as spline smoothing models and support vector machines in which similar ideas have been investigated. Gaussian process models are routinely used to solve hard machine learning problems. They are attractive because of their flexible non-parametric nature and computational simplicity. Treated within a Bayesian framework, very powerful statistical methods can be implemented which offer valid estimates of uncertainties in our predictions and generic model selection procedures cast as nonlinear optimization problems. Their main drawback of heavy computational scaling has recently been alleviated by the introduction of generic sparse approximations.13,78,31 The mathematical literature on GPs is large and often uses deep concepts which are not required to fully understand most machine learning applications. In this tutorial paper, we aim to present characteristics of GPs relevant to machine learning and to show up precise connections to other "kernel machines" popular in the community. Our focus is on a simple presentation, but references to more detailed sources are provided. Matthias W. Seeger |
Int. J. Neural Syst. | 1 |
| 2002 | Fast Sparse Gaussian Process Methods: The Informative Vector MachineabstractWe present a framework for sparse Gaussian process (GP) methods which uses forward selection with criteria based on information- theoretic principles, previously suggested for active learning. Our goal is not only to learn d{sparse predictors (which can be evalu- ated in O(d) rather than O(n), d (cid:28) n, n the number of training points), but also to perform training under strong restrictions on time and memory requirements. The scaling of our method is at most O(n (cid:1) d2), and in large real-world classi(cid:12)cation experiments we show that it can match prediction performance of the popular support vector machine (SVM), yet can be signi(cid:12)cantly faster in training. In contrast to the SVM, our approximation produces esti- mates of predictive probabilities (‘error bars’), allows for Bayesian model selection and is less complex in implementation. Neil D. Lawrence, Matthias W. Seeger, Ralf Herbrich |
NIPS | 2 |
| 2002 | PAC-Bayesian Generalisation Error Bounds for Gaussian Process Classification
Matthias W. Seeger |
J. Mach. Learn. Res. | 1 |
| 2001 | An Improved Predictive Accuracy Bound for Averaging Classifiers
John Langford 0001, Matthias W. Seeger, Nimrod Megiddo |
ICML | 2 |
| 2001 | Covariance Kernels from Bayesian Generative ModelsabstractWe propose the framework of mutual information kernels for learning covariance kernels, as used in Support Vector machines and Gaussian process classifiers, from unlabeled task data using Bayesian techniques. We describe an implementation of this frame(cid:173) work which uses variational Bayesian mixtures of factor analyzers in order to attack classification problems in high-dimensional spaces where labeled data is sparse, but unlabeled data is abundant. Matthias W. Seeger |
NIPS | 1 |
| 2000 | The Effect of the Input Density Distribution on Kernel-based Classifiers
Christopher K. I. Williams, Matthias W. Seeger |
ICML | 2 |
| 2000 | Using the Nyström Method to Speed Up Kernel Machines
Christopher K. I. Williams, Matthias W. Seeger |
NIPS | 2 |
| 1999 | Bayesian Model Selection for Support Vector Machines, Gaussian Processes and Other Kernel Classifiers
Matthias W. Seeger |
NIPS | 1 |