VLDB 2026 Research / reviewers in the wild / expert
Stephen J. Roberts
dblp:64/1485 · also Stephen Roberts 0001
· DBLP profile ↗
117ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0002-9305-9268ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 86 · 9 first-author · 20 since 2021Databases, data management, data science and information retrieval · 17 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 since 2021Computer networks · 4Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training instabilities favor flatter solutions in gradient descentabstractClassical analyses of gradient descent (GD) define a stability threshold based on the largest eigenvalue of the loss Hessian, often termed sharpness. When the learning rate lies below this threshold, training is stable and the loss decreases monotonically. Yet, modern deep networks often achieve their best performance beyond this regime. We demonstrate that such instabilities induce an implicit preference in GD, driving parameters toward flatter regions of the loss landscape and thereby improving generalization. The key mechanism is the Rotational Polarity of Eigenvectors (RPE), a geometric phenomenon in which the leading eigenvectors of the Hessian rotate during training instabilities. These rotations, which increase with learning rates, promote exploration and provably lead to flatter minima. This theoretical framework extends to stochastic GD, where instability-driven flattening persists and its empirical effects outweigh minibatch noise. Finally, we show that restoring instabilities in Adam further improves generalization. Together, these results establish and understand the constructive role of training instabilities in deep learning. Lawrence Wang, Stephen J. Roberts |
Neural Networks | 2 |
| 2024 | Iterate Averaging in the Quest for Best Test ErrorabstractWe analyse and explain the increased generalisation performance of iterate averaging using a Gaussian process perturbation model between the true and batch risk surface on the high dimensional quadratic. We derive three phenomena from our theoretical results: (1) The importance of combining iterate averaging (IA) with large learning rates and regularisation for improved generalisation. (2) Justification for less frequent averaging. (3) That we expect adaptive gradient methods to work equally well, or better, with iterate averaging than their non-adaptive counterparts. Inspired by these results, together with empirical investigations of the importance of appropriate regularisation for the solution diversity of the iterates, we propose two adaptive algorithms with iterate averaging. These give significantly better results compared to stochastic gradient descent (SGD), require less tuning and do not require early stopping or validation set monitoring. We showcase the efficacy of our approach on the CIFAR-10/100, ImageNet and Penn Treebank datasets on a variety of modern and classical network architectures. Diego Granziol, Nicholas P. Baskerville, Xingchen Wan, Samuel Albanie, Stephen J. Roberts |
J. Mach. Learn. Res. | 5 |
| 2023 | Nonparametric Boundary Geometry in Physics Informed Deep LearningabstractEngineering design problems frequently require solving systems of
partial differential equations with boundary conditions specified on
object geometries in the form of a triangular mesh. These boundary
geometries are provided by a designer and are problem dependent.
The efficiency of the design process greatly benefits from fast turnaround
times when repeatedly solving PDEs on various geometries. However,
most current work that uses machine learning to speed up the solution
process relies heavily on a fixed parameterization of the geometry, which
cannot be changed after training. This severely limits the possibility of
reusing a trained model across a variety of design problems.
In this work, we propose a novel neural operator architecture which accepts
boundary geometry, in the form of triangular meshes, as input and produces an
approximate solution to a given PDE as output. Once trained, the model can be
used to rapidly estimate the PDE solution over a new geometry, without the need for
retraining or representation of the geometry to a pre-specified parameterization. Scott Alexander Cameron, Arnu Pretorius, Stephen J. Roberts |
NeurIPS | 3 |
| 2022 | Same State, Different Task: Continual Reinforcement Learning without InterferenceabstractContinual Learning (CL) considers the problem of training an agent sequentially on a set of tasks while seeking to retain performance on all previous tasks. A key challenge in CL is catastrophic forgetting, which arises when performance on a previously mastered task is reduced when learning a new task. While a variety of methods exist to combat forgetting, in some cases tasks are fundamentally incompatible with each other and thus cannot be learnt by a single policy. This can occur, in reinforcement learning (RL) when an agent may be rewarded for achieving different goals from the same observation. In this paper we formalize this "interference" as distinct from the problem of forgetting. We show that existing CL methods based on single neural network predictors with shared replay buffers fail in the presence of interference. Instead, we propose a simple method, OWL, to address this challenge. OWL learns a factorized policy, using shared feature extraction layers, but separate heads, each specializing on a new task. The separate heads in OWL are used to prevent interference. At test time, we formulate policy selection as a multi-armed bandit problem, and show it is possible to select the best policy for an unknown task using feedback from the environment. The use of bandit algorithms allows the OWL agent to constructively re-use different continually learnt policies at different times during an episode. We show in multiple RL environments that existing replay based CL methods fail, while OWL is able to achieve close to optimal performance when training sequentially. Samuel Kessler, Jack Parker-Holder, Philip J. Ball, Stefan Zohren, Stephen J. Roberts |
AAAI | 5 |
| 2022 | Marginalising over Stationary Kernels with Bayesian QuadratureabstractMarginalising over families of Gaussian Process kernels produces flexible model classes with well-calibrated uncertainty estimates. Existing approaches require likelihood evaluations of many kernels, rendering them prohibitively expensive for larger datasets. We propose a Bayesian Quadrature scheme to make this marginalisation more efficient and thereby more practical. Through use of maximum mean discrepancies between distributions, we define a kernel over kernels that captures invariances between Spectral Mixture (SM) Kernels. Kernel samples are selected by generalising an information-theoretic acquisition function for warped Bayesian Quadrature. We show that our framework achieves more accurate predictions with better calibrated uncertainty than state-of-the-art baselines, especially when given limited (wall-clock) time budgets. Saad Hamid, Sebastian Schulze, Michael A. Osborne, Stephen J. Roberts |
AISTATS | 4 |
| 2022 | Uncertainty Estimation with a VAE-Classifier Hybrid ModelabstractWe propose a hybrid model that combines a generative unit and a discriminative classifier to quantify uncertainty in a classification task. The representation learning capability in the VAE module allows our method to learn more useful and generalizable features and outperform other purely discriminative classifiers when training labels are limited. With proper statistical treatment, the probabilistic encoder in our VAE module offers a convenient mechanism to express uncertainty for out-of-distribution (OOD) data. As a result, our method gives better calibrated uncertainty prediction. We demonstrate the effectiveness of our method on MNIST and a challenging medical image dataset for skin lesion diagnosis. Shuyu Lin, Ronald Clark, Agathoniki Trigoni, Stephen J. Roberts |
ICASSP | 4 |
| 2022 | Robust and Scalable SDE Learning: A Functional Perspective
Scott Alexander Cameron, Tyron Luke Cameron, Arnu Pretorius, Stephen J. Roberts |
ICLR | 4 |
| 2022 | Revisiting Design Choices in Offline Model Based Reinforcement Learning
Cong Lu, Philip J. Ball, Jack Parker-Holder, Michael A. Osborne, Stephen J. Roberts |
ICLR | 5 |
| 2022 | Stabilizing Off-Policy Deep Reinforcement Learning from PixelsabstractOff-policy reinforcement learning (RL) from pixel observations is notoriously unstable. As a result, many successful algorithms must combine different domain-specific practices and auxiliary losses to learn meaningful behaviors in complex environments. In this work, we provide novel analysis demonstrating that these instabilities arise from performing temporal-difference learning with a convolutional encoder and low-magnitude rewards. We show that this new visual deadly triad causes unstable training and premature convergence to degenerate solutions, a phenomenon we name catastrophic self-overfitting. Based on our analysis, we propose A-LIX, a method providing adaptive regularization to the encoder’s gradients that explicitly prevents the occurrence of catastrophic self-overfitting using a dual objective. By applying A-LIX, we significantly outperform the prior state-of-the-art on the DeepMind Control and Atari benchmarks without any data augmentation or auxiliary losses. Edoardo Cetin, Philip J. Ball, Stephen J. Roberts, Oya Çeliktutan |
ICML | 3 |
| 2022 | The ACM Multimedia 2022 Computational Paralinguistics Challenge: Vocalisations, Stuttering, Activity, & MosquitoesabstractThe ACM Multimedia 2022 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the Vocalisations and Stuttering Sub-Challenges, a classification on human non-verbal vocalisations and speech has to be made; the Activity Sub-Challenge aims at beyond-audio human activity recognition from smartwatch sensor data; and in the Mosquitoes Sub-Challenge, mosquitoes need to be detected. We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' ComParE and BoAW features, the auDeep toolkit, and deep feature extraction from pre-trained CNNs using the DeepSpectrum toolkit; in addition, we add end-to-end sequential modelling, and a log-mel-128-BNN. Björn W. Schuller, Anton Batliner, Shahin Amiriparian, Christian Bergler, Maurice Gerczuk, Natalie Holz, Pauline Larrouy-Maestri, Sebastian P. Bayerl, Korbinian Riedhammer, Adria Mallol-Ragolta, Maria Pateraki, Harry Coppock, Ivan Kiskin, Marianne Sinka, Stephen J. Roberts |
ACM Multimedia | 15 |
| 2022 | Learning General World Models in a Handful of Reward-Free DeploymentsabstractBuilding generally capable agents is a grand challenge for deep reinforcement learning (RL). To approach this challenge practically, we outline two key desiderata: 1) to facilitate generalization, exploration should be task agnostic; 2) to facilitate scalability, exploration policies should collect large quantities of data without costly centralized retraining. Combining these two properties, we introduce the reward-free deployment efficiency setting, a new paradigm for RL research. We then present CASCADE, a novel approach for self-supervised exploration in this new setting. CASCADE seeks to learn a world model by collecting data with a population of agents, using an information theoretic objective inspired by Bayesian Active Learning. CASCADE achieves this by specifically maximizing the diversity of trajectories sampled by the population through a novel cascading objective. We provide theoretical intuition for CASCADE which we show in a tabular setting improves upon naïve approaches that do not account for population diversity. We then demonstrate that CASCADE collects diverse task-agnostic datasets and learns agents that generalize zero-shot to novel, unseen downstream tasks on Atari, MiniGrid, Crafter and the DM Control Suite. Code and videos are available at https://ycxuyingchen.github.io/cascade/ Yingchen Xu, Jack Parker-Holder, Aldo Pacchiano, Philip J. Ball, Oleh Rybkin, Stephen J. Roberts, Tim Rocktäschel, Edward Grefenstette |
NeurIPS | 6 |
| 2022 | Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network TrainingabstractWe study the effect of mini-batching on the loss landscape of deep neural networks using spiked, field-dependent random matrix theory. We demonstrate that the magnitude of the extremal values of the batch Hessian are larger than those of the empirical Hessian. We also derive similar results for the Generalised Gauss-Newton matrix approximation of the Hessian. As a consequence of our theorems we derive an analytical expressions for the maximal learning rates as a function of batch size, informing practical training regimens for both stochastic gradient descent (linear scaling) and adaptive algorithms, such as Adam (square root scaling), for smooth, non-convex deep neural networks. Whilst the linear scaling for stochastic gradient descent has been derived under more restrictive conditions, which we generalise, the square root scaling rule for adaptive optimisers is, to our knowledge, completely novel. We validate our claims on the VGG/WideResNet architectures on the CIFAR-100 and ImageNet data sets. Based on our investigations of the sub-sampled Hessian we develop a stochastic Lanczos quadrature based on the fly learning rate and momentum learner, which avoids the need for expensive multiple evaluations for these key hyper-parameters and shows good preliminary results on the Pre-Residual Architecture for CIFAR-100. We further investigate the similarity between the Hessian spectrum of a multi-layer perceptron, trained on Gaussian mixture data, compared to that of deep neural networks trained on natural images. We find striking similarities, with both exhibiting rank degeneracy, a bulk spectrum and outliers to that spectrum. Furthermore, we show that ZCA whitening can remove such outliers early on in training before class separation occurs, but that outliers persist in later training. Diego Granziol, Stefan Zohren, Stephen J. Roberts |
J. Mach. Learn. Res. | 3 |
| 2022 | Adversarial Robustness Guarantees for Gaussian ProcessesabstractGaussian processes (GPs) enable principled computation of model uncertainty, making them attractive for safety-critical applications. Such scenarios demand that GP decisions are not only accurate, but also robust to perturbations. In this paper we present a framework to analyse adversarial robustness of GPs, defined as invariance of the model's decision to bounded perturbations. Given a compact subset of the input space $T\subseteq \mathbb{R}^d$, a point $x^*$ and a GP, we provide provable guarantees of adversarial robustness of the GP by computing lower and upper bounds on its prediction range in $T$. We develop a branch-and-bound scheme to refine the bounds and show, for any $\epsilon > 0$, that our algorithm is guaranteed to converge to values $\epsilon$-close to the actual values in finitely many iterations. The algorithm is anytime and can handle both regression and classification tasks, with analytical formulation for most kernels used in practice. We evaluate our methods on a collection of synthetic and standard benchmark data sets, including SPAM, MNIST and FashionMNIST. We study the effect of approximate inference techniques on robustness and demonstrate how our method can be used for interpretability. Our empirical results suggest that the adversarial robustness of GPs increases with accurate posterior estimation. Andrea Patanè, Arno Blaas, Luca Laurenti, Luca Cardelli, Stephen J. Roberts, Marta Z. Kwiatkowska |
J. Mach. Learn. Res. | 5 |
| 2021 | Learning Bijective Feature Maps for Linear ICAabstractSeparating high-dimensional data like images into independent latent factors, i.e independent component analysis (ICA), remains an open research problem. As we show, existing probabilistic deep generative models (DGMs), which are tailor-made for image data, underperform on non-linear ICA tasks. To address this, we propose a DGM which combines bijective feature maps with a linear ICA model to learn interpretable latent structures for high-dimensional data. Given the complexities of jointly training such a hybrid model, we introduce novel theory that constrains linear ICA to lie close to the manifold of orthogonal rectangular matrices, the Stiefel manifold. By doing so we create models that converge quickly, are easy to train, and achieve better unsupervised latent factor discovery than flow-based models, linear ICA, and Variational Autoencoders on images. Alexander Camuto, Matthew Willetts, Christopher C. Holmes, Brooks Paige, Stephen J. Roberts |
AISTATS | 5 |
| 2021 | Towards a Theoretical Understanding of the Robustness of Variational AutoencodersabstractWe make inroads into understanding the robustness of Variational Autoencoders (VAEs) to adversarial attacks and other input perturbations. While previous work has developed algorithmic approaches to attacking and defending VAEs, there remains a lack of formalization for what it means for a VAE to be robust. To address this, we develop a novel criterion for robustness in probabilistic models: $r$-robustness. We then use this to construct the first theoretical results for the robustness of VAEs, deriving margins in the input space for which we can provide guarantees about the resulting reconstruction. Informally, we are able to define a region within which any perturbation will produce a reconstruction that is similar to the original reconstruction. To support our analysis, we show that VAEs trained using disentangling methods not only score well under our robustness metrics, but that the reasons for this can be interpreted through our theoretical results. Alexander Camuto, Matthew Willetts, Stephen J. Roberts, Christopher C. Holmes, Tom Rainforth |
AISTATS | 3 |
| 2021 | Improving VAEs' Robustness to Adversarial Attack
Matthew Willetts, Alexander Camuto, Tom Rainforth, Stephen J. Roberts, Christopher C. Holmes |
ICLR | 4 |
| 2021 | Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline EnvironmentabstractReinforcement learning from large-scale offline datasets provides us with the ability to learn policies without potentially unsafe or impractical exploration. Significant progress has been made in the past few years in dealing with the challenge of correcting for differing behavior between the data collection and learned policies. However, little attention has been paid to potentially changing dynamics when transferring a policy to the online setting, where performance can be up to 90% reduced for existing methods. In this paper we address this problem with Augmented World Models (AugWM). We augment a learned dynamics model with simple transformations that seek to capture potential changes in physical properties of the robot, leading to more robust policies. We not only train our policy in this new setting, but also provide it with the sampled augmentation as a context, allowing it to adapt to changes in the environment. At test time we learn the context in a self-supervised fashion by approximating the augmentation which corresponds to the new environment. We rigorously evaluate our approach on over 100 different changed dynamics settings, and show that this simple approach can significantly improve the zero-shot generalization of a recent state-of-the-art baseline, often achieving successful policies where the baseline fails. Philip J. Ball, Cong Lu, Jack Parker-Holder, Stephen J. Roberts |
ICML | 4 |
| 2021 | Tuning Mixed Input Hyperparameters on the Fly for Efficient Population Based AutoRLabstractDespite a series of recent successes in reinforcement learning (RL), many RL algorithms remain sensitive to hyperparameters. As such, there has recently been interest in the field of AutoRL, which seeks to automate design decisions to create more general algorithms. Recent work suggests that population based approaches may be effective AutoRL algorithms, by learning hyperparameter schedules on the fly. In particular, the PB2 algorithm is able to achieve strong performance in RL tasks by formulating online hyperparameter optimization as time varying GP-bandit problem, while also providing theoretical guarantees. However, PB2 is only designed to work for \emph{continuous} hyperparameters, which severely limits its utility in practice. In this paper we introduce a new (provably) efficient hierarchical approach for optimizing \emph{both continuous and categorical} variables, using a new time-varying bandit algorithm specifically designed for the population based training regime. We evaluate our approach on the challenging Procgen benchmark, where we show that explicitly modelling dependence between data augmentation and other hyperparameters improves generalization. Jack Parker-Holder, Shaan Desai, Stephen J. Roberts |
NeurIPS | 4 |
| 2021 | Automatic Acoustic Mosquito Tagging with Bayesian Neural Networks
Ivan Kiskin, Adam D. Cobb, Marianne Sinka, Kathy Willis, Stephen J. Roberts |
ECML/PKDD (4) | 5 |
| 2021 | Hierarchical Indian buffet neural networks for Bayesian continual learningabstractWe place an Indian Buffet process (IBP) prior over the structure of a Bayesian Neural Network (BNN), thus allowing the complexity of the BNN to increase and decrease automatically. We further extend this model such that the prior on the structure of each hidden layer is shared globally across all layers, using a Hierarchical-IBP (H-IBP). We apply this model to the problem of resource allocation in Continual Learning (CL) where new tasks occur and the network requires extra resources. Our model uses online variational inference with reparameterisation of the Bernoulli and Beta distributions, which constitute the IBP and H-IBP priors. As we automatically learn the number of weights in each layer of the BNN, overfitting and underfitting problems are largely overcome. We show empirically that our approach offers a competitive edge over existing methods in CL. Samuel Kessler, Stefan Zohren, Stephen J. Roberts |
UAI | 4 |
| 2021 | Towards tractable optimism in model-based reinforcement learningabstractThe principle of optimism in the face of uncertainty is prevalent throughout sequential decision making problems such as multi-armed bandits and reinforcement learning (RL). To be successful, an optimistic RL algorithm must over-estimate the true value function (optimism) but not by so much that it is inaccurate (estimation error). In the tabular setting, many state-of-the-art methods produce the required optimism through approaches which are intractable when scaling to deep RL. We re-interpret these scalable optimistic model-based algorithms as solving a tractable noise augmented MDP. This formulation achieves a competitive regret bound: $\tilde{\mathcal{O}}( |\mathcal{S}|H\sqrt{|\mathcal{A}| T } )$ when augmenting using Gaussian noise, where $T$ is the total number of environment steps. We also explore how this trade-off changes in the deep RL setting, where we show empirically that estimation error is significantly more troublesome. However, we also show that if this error is reduced, optimistic model-based RL algorithms can match state-of-the-art performance in continuous control problems. Aldo Pacchiano, Philip J. Ball, Jack Parker-Holder, Krzysztof Choromanski, Stephen J. Roberts |
UAI | 5 |
| 2021 | A Bayesian Optimization Approach to Compute Nash Equilibrium of Potential Games Using Bandit FeedbackabstractAbstract Computing a Nash equilibrium for strategic multi-agent systems is challenging for black box systems. Motivated by the ubiquity of games involving exploitation of common resources, this paper considers the above problem for potential games. We use a Bayesian optimization framework to obtain novel algorithms to solve finite (discrete action spaces) and infinite (real interval action spaces) potential games, utilizing the structure of potential games. Numerical results illustrate the efficiency of the approach in computing a Nash equilibrium of static potential games and linear Nash equilibrium of dynamic potential games. Anup Aprem, Stephen J. Roberts |
Comput. J. | 2 |
| 2021 | Optimal pricing in black box producer-consumer Stackelberg games using revealed preference feedback
Anup Aprem, Stephen J. Roberts |
Neurocomputing | 2 |
| 2020 | Adversarial Robustness Guarantees for Classification with Gaussian ProcessesabstractWe investigate adversarial robustness of Gaussian Process classification (GPC) models. Specifically, given a compact subset of the input space $T\subseteq \mathbb{R}^d$ enclosing a test point $x^*$ and a GPC trained on a dataset $\mathcal{D}$, we aim to compute the minimum and the maximum classification probability for the GPC over all the points in $T$.In order to do so, we show how functions lower- and upper-bounding the GPC output in $T$ can be derived, and implement those in a branch and bound optimisation algorithm. For any error threshold $\epsilon > 0$ selected \emph{a priori}, we show that our algorithm is guaranteed to reach values $\epsilon$-close to the actual values in finitely many iterations.We apply our method to investigate the robustness of GPC models on a 2D synthetic dataset, the SPAM dataset and a subset of the MNIST dataset, providing comparisons of different GPC training techniques, and show how our method can be used for interpretability analysis. Our empirical analysis suggests that GPC robustness increases with more accurate posterior estimation. Arno Blaas, Andrea Patanè, Luca Laurenti, Luca Cardelli, Marta Z. Kwiatkowska, Stephen J. Roberts |
AISTATS | 6 |
| 2020 | Semi-Unsupervised Learning: Clustering and Classifying using Ultra-Sparse LabelsabstractIn semi-supervised learning for classification, i t is assumed that every ground truth class of data is present in the small labelled dataset. In many real-world sparsely-labelled datasets, it is possible that not all ground-truth classes are captured in the labelled dataset: a biased data collection process could result in some classes of data to be found only in the unlabelled dataset. We call this regime `semi-unsupervised learning', an extreme case of semi-supervised learning, where some classes have no labelled exemplars. First, we outline the pitfalls associated with trying to apply deep generative model (DGM)-based semi-supervised learning algorithms to datasets of this type. We then show how a combination of clustering and semi-supervised learning, using DGMs, can be brought to bear on this problem. We study several different datasets, showing how one can still learn effectively when half of the ground truth classes are entirely unlabelled and the other half are sparsely labelled. Matthew Willetts, Stephen J. Roberts, Christopher C. Holmes |
IEEE BigData | 2 |
| 2020 | Zero-shot and few-shot time series forecasting with ordinal regression recurrent neural networks
Bernardo Pérez Orozco, Stephen J. Roberts |
ESANN | 2 |
| 2020 | Humbug Zooniverse: A Crowd-Sourced Acoustic Mosquito DatasetabstractMosquitoes are the only known vector of malaria, which leads to hundreds of thousands of deaths each year. Understanding the number and location of potential mosquito vectors is of paramount importance to aid the reduction of malaria transmission cases. In recent years, deep learning has become widely used for bioacoustic classification tasks. In order to enable further research applications in this field, we release a new dataset of mosquito audio recordings. With over a thousand contributors, we obtained 195,434 labels of two second duration, of which approximately 10 percent signify mosquito events. We present an example use of the dataset, in which we train a convolutional neural network on log-Mel features, showcasing the information content of the labels. We hope this will become a vital resource for those researching all aspects of malaria, and add to the existing audio datasets for bioacoustic detection and signal processing. Ivan Kiskin, Adam D. Cobb, Lawrence Wang, Stephen J. Roberts |
ICASSP | 4 |
| 2020 | Anomaly Detection for Time Series Using VAE-LSTM Hybrid ModelabstractIn this work, we propose a VAE-LSTM hybrid model as an unsupervised approach for anomaly detection in time series. Our model utilizes both a VAE module for forming robust local features over short windows and a LSTM module for estimating the long term correlation in the series on top of the features inferred from the VAE module. As a result, our detection algorithm is capable of identifying anomalies that span over multiple time scales. We demonstrate the effectiveness of our detection algorithm on five real world problems and find our method outperforms three other commonly used detection methods. Shuyu Lin, Ronald Clark, Robert Birke, Sandro Schönborn, Agathoniki Trigoni, Stephen J. Roberts |
ICASSP | 6 |
| 2020 | Ready Policy One: World Building Through Active LearningabstractModel-Based Reinforcement Learning (MBRL) offers a promising direction for sample efficient learning, often achieving state of the art results for continuous control tasks. However many existing MBRL methods rely on combining greedy policies with exploration heuristics, and even those which utilize principled exploration bonuses construct dual objectives in an ad hoc fashion. In this paper we introduce Ready Policy One (RP1), a framework that views MBRL as an active learning problem, where we aim to improve the world model in the fewest samples possible. RP1 achieves this by utilizing a hybrid objective function, which crucially adapts during optimization, allowing the algorithm to trade off reward v.s. exploration at different stages of learning. In addition, we introduce a principled mechanism to terminate sample collection once we have a rich enough trajectory batch to improve the model. We rigorously evaluate our method on a variety of continuous control tasks, and demonstrate statistically significant gains over existing approaches. Philip J. Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski, Stephen J. Roberts |
ICML | 5 |
| 2020 | Bayesian Optimisation over Multiple Continuous and Categorical InputsabstractEfficient optimisation of black-box problems that comprise both continuous and categorical inputs is important, yet poses significant challenges. Current approaches, like one-hot encoding, severely increase the dimension of the search space, while separate modelling of category-specific data is sample-inefficient. Both frameworks are not scalable to practical applications involving multiple categorical variables, each with multiple possible values. We propose a new approach, Continuous and Categorical Bayesian Optimisation (CoCaBO), which combines the strengths of multi-armed bandits and Bayesian optimisation to select values for both categorical and continuous inputs. We model this mixed-type space using a Gaussian Process kernel, designed to allow sharing of information across multiple categorical variables; this allows CoCaBO to leverage all available data efficiently. We extend our method to the batch setting and propose an efficient selection procedure that dynamically balances exploration and exploitation whilst encouraging batch diversity. We demonstrate empirically that our method outperforms existing approaches on both synthetic and real-world optimisation tasks with continuous and categorical inputs. Bin Xin Ru, Ahsan S. Alvi, Michael A. Osborne, Stephen J. Roberts |
ICML | 5 |
| 2020 | Recurrent Neural Filters: Learning Independent Bayesian Filtering Steps for Time Series PredictionabstractDespite the recent popularity of deep generative state space models, few comparisons have been made between network architectures and the inference steps of the Bayesian filtering framework - with most models simultaneously approximating both state transition and update steps with a single recurrent neural network (RNN). In this paper, we introduce the Recurrent Neural Filter (RNF), a novel recurrent autoencoder architecture that learns distinct representations for each Bayesian filtering step, captured by a series of encoders and decoders. Testing this on three real-world time series datasets, we demonstrate that the decoupled representations learnt improve the accuracy of one-step-ahead forecasts while providing realistic uncertainty estimates, and also facilitate multistep prediction through the separation of encoder stages. Bryan Lim, Stefan Zohren, Stephen J. Roberts |
IJCNN | 3 |
| 2020 | Explicit Regularisation in Gaussian Noise InjectionsabstractWe study the regularisation induced in neural networks by Gaussian noise injections (GNIs). Though such injections have been extensively studied when applied to data, there have been few studies on understanding the regularising effect they induce when applied to network activations. Here we derive the explicit regulariser of GNIs, obtained by marginalising out the injected noise, and show that it penalises functions with high-frequency components in the Fourier domain; particularly in layers closer to a neural network's output. We show analytically and empirically that such regularisation produces calibrated classifiers with large classification margins. Alexander Camuto, Matthew Willetts, Umut Simsekli, Stephen J. Roberts, Christopher C. Holmes |
NeurIPS | 4 |
| 2020 | Provably Efficient Online Hyperparameter Optimization with Population-Based BanditsabstractMany of the recent triumphs in machine learning are dependent on well-tuned hyperparameters. This is particularly prominent in reinforcement learning (RL) where a small change in the configuration can lead to failure. Despite the importance of tuning hyperparameters, it remains expensive and is often done in a naive and laborious way. A recent solution to this problem is Population Based Training (PBT) which updates both weights and hyperparameters in a \emph{single training run} of a population of agents. PBT has been shown to be particularly effective in RL, leading to widespread use in the field. However, PBT lacks theoretical guarantees since it relies on random heuristics to explore the hyperparameter space. This inefficiency means it typically requires vast computational resources, which is prohibitive for many small and medium sized labs. In this work, we introduce the first provably efficient PBT-style algorithm, Population-Based Bandits (PB2). PB2 uses a probabilistic model to guide the search in an efficient way, making it possible to discover high performing hyperparameter configurations with far fewer agents than typically required by PBT. We show in a series of RL experiments that PB2 is able to achieve high performance with a modest computational budget. Jack Parker-Holder, Stephen J. Roberts |
NeurIPS | 3 |
| 2020 | Effective Diversity in Population Based Reinforcement LearningabstractExploration is a key problem in reinforcement learning, since agents can only learn from data they acquire in the environment. With that in mind, maintaining a population of agents is an attractive method, as it allows data be collected with a diverse set of behaviors. This behavioral diversity is often boosted via multi-objective loss functions. However, those approaches typically leverage mean field updates based on pairwise distances, which makes them susceptible to cycling behaviors and increased redundancy. In addition, explicitly boosting diversity often has a detrimental impact on optimizing already fruitful behaviors for rewards. As such, the reward-diversity trade off typically relies on heuristics. Finally, such methods require behavioral representations, often handcrafted and domain specific. In this paper, we introduce an approach to optimize all members of a population simultaneously. Rather than using pairwise distance, we measure the volume of the entire population in a behavioral manifold, defined by task-agnostic behavioral embeddings. In addition, our algorithm Diversity via Determinants (DvD), adapts the degree of diversity during training using online learning techniques. We introduce both evolutionary and gradient-based instantiations of DvD and show they effectively improve exploration without reducing performance when better exploration is not required. Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski, Stephen J. Roberts |
NeurIPS | 4 |
| 2020 | Thresholded ConvNet ensembles: neural networks for technical forecastingabstractAbstract Much of modern practice in financial forecasting relies on technicals, an umbrella term for several heuristics applying visual pattern recognition to price charts. Despite its ubiquity in financial media, the reliability of its signals remains a contentious and highly subjective form of ‘domain knowledge’. We investigate the predictive value of patterns in financial time series, applying machine learning and signal processing techniques to 22 years of US equity data. By reframing technical analysis as a poorly specified, arbitrarily preset feature-extractive layer in a deep neural network, we show that better convolutional filters can be learned directly from the data, and provide visual representations of the features being identified. We find that an ensemble of shallow, thresholded convolutional neural networks optimised over different resolutions achieves state-of-the-art performance on this domain, outperforming technical methods while retaining some of their interpretability. Sid Ghoshal, Stephen J. Roberts |
Neural Comput. Appl. | 2 |
| 2020 | Bioacoustic detection with wavelet-conditioned convolutional neural networksabstractMany real-world time series analysis problems are characterized by low signal-to-noise ratios and compounded by scarce data. Solutions to these types of problems often rely on handcrafted features extracted in the time or frequency domain. Recent high-profile advances in deep learning have improved performance across many application domains; however, they typically rely on large data sets that may not always be available. This paper presents an application of deep learning for acoustic event detection in a challenging, data-scarce, real-world problem. We show that convolutional neural networks (CNNs), operating on wavelet transformations of audio recordings, demonstrate superior performance over conventional classifiers that utilize handcrafted features. Our key result is that wavelet transformations offer a clear benefit over the more commonly used short-time Fourier transform. Furthermore, we show that features, handcrafted for a particular dataset, do not generalize well to other datasets. Conversely, CNNs trained on generic features are able to achieve comparable results across multiple datasets, along with outperforming human labellers. We present our results on the application of both detecting the presence of mosquitoes and the classification of bird species. Ivan Kiskin, Davide Zilli, Yunpeng Li 0001, Marianne Sinka, Kathy Willis, Stephen J. Roberts |
Neural Comput. Appl. | 6 |
| 2019 | Asynchronous Batch Bayesian Optimisation with Improved Local PenalisationabstractBatch Bayesian optimisation (BO) has been successfully applied to hyperparameter tuning using parallel computing, but it is wasteful of resources: workers that complete jobs ahead of others are left idle. We address this problem by developing an approach, Penalising Locally for Asynchronous Bayesian Optimisation on K Workers (PLAyBOOK), for asynchronous parallel BO. We demonstrate empirically the efficacy of PLAyBOOK and its variants on synthetic tasks and a real-world problem. We undertake a comparison between synchronous and asynchronous BO, and show that asynchronous BO often outperforms synchronous batch BO in both wall-clock time and sample efficiency. Ahsan S. Alvi, Bin Xin Ru, Jan-P. Calliess, Stephen J. Roberts, Michael A. Osborne |
ICML | 4 |
| 2019 | Gaussian Processes for Personalized Interpretable Volatility Metrics in the Step-Down WardabstractPatients in a hospital step-down unit require a level of care that is between that of the intensive care unit (ICU) and that of the general ward. While many patients remain physiologically stabilized, others will suffer clinical emergencies and be readmitted to the ICU, with a subsequent high risk of mortality. Had the associated physiological deterioration been detected early, the emergency may have been less severe or avoided entirely. Current clinical monitoring is largely heuristic, requiring manual calculation of risk scores and the use of heuristic decision criteria. Technical drawbacks include ignoring the time-series dynamics of physiological measurements, and lacking patient-specificity (i.e., personalization of models to the individual patient). In this paper, we demonstrate how Gaussian process regression models can supplement current monitoring practice by providing interpretable and intuitive illustrations of erratic vital-sign volatility. These personalized volatility metrics may provide significantly advanced warning of deterioration, while minimizing the false alarms that induce so-called alarm fatigue. While many AI-based approaches to healthcare are criticized for being uninterpretable "black-box" methods, the cause of alarms generated from the proposed methods are explicitly interpretable and intuitive. We conclude that intelligent computational inference using methods such as those proposed can enhance current clinical decision making and potentially save lives. Glen Wright Colopy, Stephen J. Roberts, David A. Clifton |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Novel Exploration Techniques (NETs) for Malaria Policy InterventionsabstractThe task of decision-making under uncertainty is daunting, especially for problems which have significant complexity. Healthcare policy makers across the globe are facing problems under challenging constraints, with limited tools to help them make data driven decisions. In this work we frame the process of finding an optimal malaria policy as a stochastic multi-armed bandit problem, and implement three agent based strategies to explore the policy space. We apply a Gaussian Process regression to the findings of each agent, both for comparison and to account for stochastic results from simulating the spread of malaria in a fixed population. The generated policy spaces are compared with published results to give a direct reference with human expert decisions for the same simulated population. Our novel approach provides a powerful resource for policy makers, and a platform which can be readily extended to capture future more nuanced policy spaces. Oliver Bent, Sekou L. Remy, Stephen J. Roberts, Aisha Walcott-Bryant |
AAAI | 3 |
| 2018 | Optimization, Fast and Slow: Optimally Switching between Local and Bayesian OptimizationabstractWe develop the first Bayesian Optimization algorithm, BLOSSOM, which selects between multiple alternative acquisition functions and traditional local optimization at each step. This is combined with a novel stopping condition based on expected regret. This pairing allows us to obtain the best characteristics of both local and Bayesian optimization, making efficient use of function evaluations while yielding superior convergence to the global minimum on a selection of optimization problems, and also halting optimization once a principled and intuitive stopping condition has been fulfilled. Mark McLeod, Stephen J. Roberts, Michael A. Osborne |
ICML | 2 |
| 2018 | Identifying Sources and Sinks in the Presence of Multiple Agents with Gaussian Process Vector CalculusabstractIn systems of multiple agents, identifying the cause of observed agent dynamics is challenging. Often, these agents operate in diverse, non-stationary environments, where models rely on hand-crafted environment-specific features to infer influential regions in the system's surroundings. To overcome the limitations of these inflexible models, we present GP-LAPLACE, a technique for locating sources and sinks from trajectories in time-varying fields. Using Gaussian processes, we jointly infer a spatio-temporal vector field, as well as canonical vector calculus operations on that field. Notably, we do this from only agent trajectories without requiring knowledge of the environment, and also obtain a metric for denoting the significance of inferred causal features in the environment by exploiting our probabilistic method. To evaluate our approach, we apply it to both synthetic and real-world GPS data, demonstrating the applicability of our technique in the presence of multiple agents, as well as its superiority over existing methods. Adam D. Cobb, Richard Everett 0001, Andrew Markham, Stephen J. Roberts |
KDD | 4 |
| 2018 | Improved Stochastic Trace Estimation using Mutually Unbiased Bases
Jack K. Fitzsimons, Michael A. Osborne, Stephen J. Roberts, Joseph F. Fitzsimons |
UAI | 3 |
| 2018 | Provenance Network Analytics - An approach to data analytics using data provenanceabstractProvenance network analytics is a novel data analytics approach that helps infer properties of data, such as quality or importance, from their provenance. Instead of analysing application data, which are typically domain-dependent, it analyses the data’s provenance as represented using the World Wide Web Consortium’s domain-agnostic PROV data model. Specifically, the approach proposes a number of network metrics for provenance data and applies established machine learning techniques over such metrics to build predictive models for some key properties of data. Applying this method to the provenance of real-world data from three different applications, we show that it can successfully identify the owners of provenance documents, assess the quality of crowdsourced data, and identify instructions from chat messages in an alternate-reality game with high levels of accuracy. By so doing, we demonstrate the different ways the proposed provenance network metrics can be used in analysing data, providing the foundation for provenance-based data analytics. Trung Dong Huynh, Mark Ebden, Joel E. Fischer, Stephen J. Roberts, Luc Moreau 0001 |
Data Min. Knowl. Discov. | 4 |
| 2018 | Bayesian Optimization of Personalized Models for Patient Vital-Sign MonitoringabstractGaussian process regression (GPR) provides a means to generate flexible personalized models of time series of patient vital signs. These models can perform useful clinical inference in ways that population-based models cannot. A challenge for the use of personalized models is that they must be amenable to a wide range of parameterizations, to accommodate the plausible physiology of any patient in the population. Additionally, optimal performance is typically achieved when models are regularized in light of the knowledge of the physiology of the individual patient. In this paper, we describe a method to build GP models with varying complexity (via covariance kernels) and regularization (via fixed priors over hyperparameters) on a patient-specific level, for the purpose of robust vital-sign forecasting. To this end, our results present evidence in support of two main hypotheses: 1) the use of patient-specific models can outperform population-based models for useful clinical tasks, such as vital-sign forecasting; and 2) the optimal values of (hyper)parameters of these models are best determined by sophisticated methods of optimization, due to high correlation between dimensions of the search space. The resulting models are sufficiently robust to inform clinicians of a patient's vital-sign trajectory and warn of imminent deterioration. Glen Wright Colopy, Stephen J. Roberts, David A. Clifton |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Distribution of Gaussian Process Arc LengthsabstractWe present the first treatment of the arc length of the GP with more than a single output dimension. GPs are commonly used for tasks such as trajectory modelling, where path length is a crucial quantity of interest. Previously, only paths in one dimension have been considered, with no theoretical consideration of higher dimensional problems. We fill the gap in the existing literature by deriving the moments of the arc length for a stationary GP with multiple output dimensions. A new method is used to derive the mean of a one-dimensional GP over a finite interval, by considering the distribution of the arc length integrand. This technique is used to derive an approximate distribution over the arc length of a vector valued GP in $\mathbbR^n$ by moment matching the distribution. Numerical simulations confirm our theoretical derivations. Justin Bewsher, Alessandra Tosi, Michael A. Osborne, Stephen J. Roberts |
AISTATS | 4 |
| 2017 | Entropic determinants of massive matricesabstractThe ability of many powerful machine learning algorithms to deal with large data sets without compromise is often hampered by computationally expensive linear algebra tasks, of which calculating the log determinant is a canonical example. In this paper we demonstrate the optimality of Maximum Entropy methods in approximating such calculations. We prove the equivalence between mean value constraints and sample expectations in the big data limit, that Covariance matrix eigenvalue distributions can be completely defined by moment information and that the reduction of the self entropy of a maximum entropy proposal distribution, achieved by adding more moments reduces the KL divergence between the proposal and true eigenvalue distribution. We empirically verify our results on a variety of SparseSuite matrices and establish best practices. Diego Granziol, Stephen J. Roberts |
IEEE BigData | 2 |
| 2017 | Entropic Trace Estimates for Log Determinants
Jack K. Fitzsimons, Diego Granziol, Kurt Cutajar, Michael A. Osborne, Maurizio Filippone, Stephen J. Roberts |
ECML/PKDD (1) | 6 |
| 2017 | Optimal Client Recommendation for Market Makers in Illiquid Financial Products
Dieter Hendricks, Stephen J. Roberts |
ECML/PKDD (3) | 2 |
| 2017 | Bayesian Heatmaps: Probabilistic Classification with Multiple Unreliable Information Sources
Edwin Simpson, Steven Reece, Stephen J. Roberts |
ECML/PKDD (2) | 3 |
| 2017 | Bayesian Inference of Log Determinants
Jack K. Fitzsimons, Kurt Cutajar, Maurizio Filippone, Michael A. Osborne, Stephen J. Roberts |
UAI | 5 |
| 2016 | Latent Point Process AllocationabstractWe introduce a probabilistic model for the factorisation of continuous Poisson process rate functions. Our model can be thought of as a topic model for Poisson point processes in which each point is assigned to one of a set of latent rate functions that are shared across multiple outputs. We show that the model brings a means of incorporating structure in point process inference beyond the state-of-the-art. We derive an efficient variational inference scheme for the model based on sparse Gaussian processes that scales linearly in the number of data points. Finally, we demonstrate, using examples from spatial and temporal statistics, how the model can be used for discovering hidden structure with greater precision than standard frequentist approaches. Chris M. Lloyd, Tom Gunter, Michael A. Osborne, Stephen J. Roberts, Tom Nickson |
AISTATS | 4 |
| 2016 | Human-agent collaboration for disaster response
Sarvapali D. Ramchurn, Feng Wu 0001, Wenchao Jiang, Joel E. Fischer, Steven Reece, Stephen J. Roberts, Tom Rodden, Christopher Greenhalgh, Nicholas R. Jennings |
Auton. Agents Multi Agent Syst. | 6 |
| 2016 | A Disaster Response System based on Human-Agent Collectives
Sarvapali D. Ramchurn, Trung Dong Huynh, Feng Wu 0001, Yuki Ikuno, Jack Flann, Luc Moreau 0001, Joel E. Fischer, Wenchao Jiang, Tom Rodden, Edwin Simpson, Steven Reece, Stephen J. Roberts, Nicholas R. Jennings |
J. Artif. Intell. Res. | 12 |
| 2016 | String and Membrane Gaussian ProcessesabstractIn this paper we introduce a novel framework for making exact nonparametric Bayesian inference on latent functions that is particularly suitable for Big Data tasks. Firstly, we introduce a class of stochastic processes we refer to as string Gaussian processes (string GPs which are not to be mistaken for Gaussian processes operating on text). We construct string GPs so that their finite- dimensional marginals exhibit suitable local conditional independence structures, which allow for scalable, distributed, and flexible nonparametric Bayesian inference, without resorting to approximations, and while ensuring some mild global regularity constraints. Furthermore, string GP priors naturally cope with heterogeneous input data, and the gradient of the learned latent function is readily available for explanatory analysis. Secondly, we provide some theoretical results relating our approach to the standard GP paradigm. In particular, we prove that some string GPs are Gaussian processes, which provides a complementary global perspective on our framework. Finally, we derive a scalable and distributed MCMC scheme for supervised learning tasks under string GP priors. The proposed MCMC scheme has computational time complexity $\mathcal{O}(N)$ and memory requirement $\mathcal{O}(dN)$, where $N$ is the data size and $d$ the dimension of the input space. We illustrate the efficacy of the proposed approach on several synthetic and real-world data sets, including a data set with $6$ millions input points and $8$ attributes. Yves-Laurent Kom Samo, Stephen J. Roberts |
J. Mach. Learn. Res. | 2 |
| 2015 | Variational Inference for Gaussian Process Modulated Poisson ProcessesabstractWe present the first fully variational Bayesian inference scheme for continuous Gaussian-process-modulated Poisson processes. Such point processes are used in a variety of domains, including neuroscience, geo-statistics and astronomy, but their use is hindered by the computational cost of existing inference schemes. Our scheme: requires no discretisation of the domain; scales linearly in the number of observed events; and is many orders of magnitude faster than previous sampling based approaches. The resulting algorithm is shown to outperform standard methods on synthetic examples, coal mining disaster data and in the prediction of Malaria incidences in Kenya. Chris M. Lloyd, Tom Gunter, Michael A. Osborne, Stephen J. Roberts |
ICML | 4 |
| 2015 | Scalable Nonparametric Bayesian Inference on Point Processes with Gaussian ProcessesabstractIn this paper we propose an efficient, scalable non-parametric Gaussian process model for inference on Poisson point processes. Our model does not resort to gridding the domain or to introducing latent thinning points. Unlike competing models that scale as O(n^3) over n data points, our model has a complexity O(nk^2) where k << n. We propose a MCMC sampler and show that the model obtained is faster, more accurate and generates less correlated samples than competing approaches on both synthetic and real-life data. Finally, we show that our model easily handles data sizes not considered thus far by alternate approaches. Yves-Laurent Kom Samo, Stephen J. Roberts |
ICML | 2 |
| 2015 | Language Understanding in the Wild: Combining Crowdsourcing and Machine LearningabstractSocial media has led to the democratisation of opinion sharing. A wealth of information about public opinions, current events, and authors' insights into specific topics can be gained by understanding the text written by users. However, there is a wide variation in the language used by different authors in different contexts on the web. This diversity in language makes interpretation an extremely challenging task. Crowdsourcing presents an opportunity to interpret the sentiment, or topic, of free-text. However, the subjectivity and bias of human interpreters raise challenges in inferring the semantics expressed by the text. To overcome this problem, we present a novel Bayesian approach to language understanding that relies on aggregated crowdsourced judgements. Our model encodes the relationships between labels and text features in documents, such as tweets, web articles, and blog posts, accounting for the varying reliability of human labellers. It allows inference of annotations that scales to arbitrarily large pools of documents. Our evaluation using two challenging crowdsourcing datasets shows that by efficiently exploiting language models learnt from aggregated crowdsourced labels, we can provide up to 25% improved classifications when only a small portion, less than 4% of documents has been labelled. Compared to the six state-of-the-art methods, we reduce by up to 67% the number of crowd responses required to achieve comparable accuracy. Our method was a joint winner of the CrowdFlower - CrowdScale 2013 Shared Task challenge at the conference on Human Computation and Crowdsourcing (HCOMP 2013). Edwin Simpson, Matteo Venanzi, Steven Reece, Pushmeet Kohli, John Guiver, Stephen J. Roberts, Nicholas R. Jennings |
WWW | 6 |
| 2015 | Modeling the Thermal Dynamics of Buildings: A Latent-Force- Model-Based ApproachabstractMinimizing the energy consumed by heating, ventilation, and air conditioning (HVAC) systems of residential buildings without impacting occupants’ comfort has been highlighted as an important artificial intelligence (AI) challenge. Typically, approaches that seek to address this challenge use a model that captures the thermal dynamics within a building, also referred to as a thermal model. Among thermal models, gray-box models are a popular choice for modeling the thermal dynamics of buildings. They combine knowledge of the physical structure of a building with various data-driven inputs and are accurate estimators of the state (internal temperature). However, existing gray-box models require a detailed specification of all the physical elements that can affect the thermal dynamics of a building a priori. This limits their applicability, particularly in residential buildings, where additional dynamics can be induced by human activities such as cooking, which contributes additional heat, or opening of windows, which leads to additional leakage of heat. Since the incidence of these additional dynamics is rarely known, their combined effects cannot readily be accommodated within existing models. To overcome this limitation and improve the general applicability of gray-box models, we introduce a novel model, which we refer to as a latent force thermal model of the thermal dynamics of a building, or LFM-TM. Our model is derived from an existing gray-box thermal model, which is augmented with an extra term referred to as the learned residual. This term is capable of modeling the effect of any a priori unknown additional dynamic, which, if not captured, appears as a structure in a thermal model’s residual (the error induced by the model). More importantly, the learned residual can also capture the effects of physical elements such as a building’s envelope or the lags in a heating system, leading to a significant reduction in complexity compared to existing models. To evaluate the performance of LFM-TM, we apply it to two independent data sources. The first is an established dataset, referred to as the FlexHouse data, which was previously used for evaluating the efficacy of existing gray-box models [Bacher and Madsen 2011]. The second dataset consists of heating data logged within homes located on the University of Southampton campus, which were specifically instrumented to collect data for our thermal modeling experiments. On both datasets, we show that LFM-TM outperforms existing models in its ability to accurately fit the observed data, generate accurate day-ahead internal temperature predictions, and explain a large amount of the variability in the future observations. This, along with the fact that we also use a corresponding efficient sequential inference scheme for LFM-TM, makes it an ideal candidate for model-based predictive control, where having accurate online predictions of internal temperatures is essential for high-quality solutions. Siddhartha Ghosh, Steven Reece, Alex Rogers, Stephen J. Roberts, Areej Malibari, Nicholas R. Jennings |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2014 | Sampling for Inference in Probabilistic Models with Fast Bayesian Quadrature
Tom Gunter, Michael A. Osborne, Roman Garnett, Philipp Hennig, Stephen J. Roberts |
NIPS | 5 |
| 2014 | Efficient Bayesian Nonparametric Modelling of Structured Point Processes
Tom Gunter, Chris M. Lloyd, Michael A. Osborne, Stephen J. Roberts |
UAI | 4 |
| 2014 | Efficient state-space inference of periodic latent force models
Steven Reece, Siddhartha Ghosh, Alex Rogers, Stephen J. Roberts, Nicholas R. Jennings |
J. Mach. Learn. Res. | 4 |
| 2014 | Maritime abnormality detection using Gaussian processes
Steven Reece, Stephen J. Roberts, Ioannis Psorakis, Iead Rezek |
Knowl. Inf. Syst. | 3 |
| 2013 | Interpretation of Crowdsourced Activities Using Provenance Network AnalysisabstractUnderstanding the dynamics of a crowdsourcing application and controlling the quality of the data it generates is challenging, partly due to the lack of tools to do so. Provenance is a domain-independent means to represent what happened in an application, which can help verify data and infer their quality. It can also reveal the processes that led to a data item and the interactions of contributors with it. Provenance patterns can manifest real-world phenomena such as a significant interest in a piece of content, providing an indication of its quality, or even issues such as undesirable interactions within a group of contributors. This paper presents an application-independent methodology for analyzing provenance graphs, constructed from provenance records, to learn about such patterns and to use them for assessing some key properties of crowdsourced data, such as their quality, in an automated manner. Validating this method on the provenance records of CollabMap, an online crowdsourcing mapping application, we demonstrated an accuracy level of over 95% for the trust classification of data generated by the crowd therein. Trung Dong Huynh, Mark Ebden, Matteo Venanzi, Sarvapali D. Ramchurn, Stephen J. Roberts, Luc Moreau 0001 |
HCOMP | 5 |
| 2013 | Multi-Agent Planning with Mixed-Integer Programming and Adaptive Interaction Constraint Generation (Extended Abstract)abstractWe consider multi-agent planning in which the agents' optimal plans are solutions to mixed-integer programs (MIP) that are coupled via integer constraints. While in principle, one could find the joint solution by combining the separate problems into one large joint centralized MIP, this approach rapidly becomes intractable for growing numbers of agents and large problem domains. To address this issue, we propose an iterative approach that combines conflict detection with constraint-generation whereby the agents plan repeatedly until all conflicts are resolved. In each planning iteration, the agents plan with as few other agents and interaction-constraints as possible. This yields an optimal method that can reduce computation markedly. We test our approach in the context of multi-agent collision avoidance in graphs with indivisible flows. Our initial simulations on randomized graph routing problems confirm predicted optimality and reduced computational effort. Jan-P. Calliess, Stephen J. Roberts |
SOCS | 2 |
| 2012 | Online Maritime Abnormality Detection Using Gaussian Processes and Extreme Value TheoryabstractNovelty, or abnormality, detection aims to identify patterns within data streams that do not conform to expected behaviour. This paper introduces a novelty detection technique using a combination of Gaussian Processes and extreme value theory to identify anomalous behaviour in streaming data. The proposed combination of continuous and count stochastic processes is a principled approach towards dynamic extreme value modeling that accounts for the dynamics in the time series, the streaming nature of its observation as well as its sampling process. The approach is tested on both synthetic and real data, showing itself to be effective in our primary application of maritime vessel track analysis. Steven Reece, Stephen J. Roberts, Iead Rezek |
ICDM | 3 |
| 2012 | Active Learning of Model Evidence Using Bayesian QuadratureabstractNumerical integration is an key component of many problems in scientific computing, statistical modelling, and machine learning. Bayesian Quadrature is a model-based method for numerical integration which, relative to standard Monte Carlo methods, offers increased sample efficiency and a more robust estimate of the uncertainty in the estimated integral. We propose a novel Bayesian Quadrature approach for numerical integration when the integrand is non-negative, such as the case of computing the marginal likelihood, predictive distribution, or normalising constant of a probabilistic model. Our approach approximately marginalises the quadrature model's hyperparameters in closed form, and introduces an active learning scheme to optimally select function evaluations, as opposed to using Monte Carlo samples. We demonstrate our method on both a number of synthetic benchmarks and a real scientific problem from astronomy. Michael A. Osborne, David Duvenaud, Roman Garnett, Carl E. Rasmussen, Stephen J. Roberts, Zoubin Ghahramani |
NIPS | 5 |
| 2012 | Real-time information processing of environmental sensor network data using bayesian gaussian processesabstractIn this article, we consider the problem faced by a sensor network operator who must infer, in real time, the value of some environmental parameter that is being monitored at discrete points in space and time by a sensor network. We describe a powerful and generic approach built upon an efficient multi-output Gaussian process that facilitates this information acquisition and processing. Our algorithm allows effective inference even with minimal domain knowledge, and we further introduce a formulation of Bayesian Monte Carlo to permit the principled management of the hyperparameters introduced by our flexible models. We demonstrate how our methods can be applied in cases where the data is delayed, intermittently missing, censored, and/or correlated. We validate our approach using data collected from three networks of weather sensors and show that it yields better inference performance than both conventional independent Gaussian processes and the Kalman filter. Finally, we show that our formalism efficiently reuses previous computations by following an online update procedure as new data sequentially arrives, and that this results in a four-fold increase in computational speed in the largest cases considered. Michael A. Osborne, Stephen J. Roberts, Alex Rogers, Nicholas R. Jennings |
ACM Trans. Sens. Networks | 2 |
| 2011 | Determining intent using hard/soft data and Gaussian process classifiers
Steven Reece, Stephen J. Roberts, David Nicholson, Chris M. Lloyd |
FUSION | 2 |
| 2011 | Graph marginalization for rapid assignment in wide-area surveillance
Mark Ebden, Stephen J. Roberts |
Ad Hoc Networks | 2 |
| 2011 | Bayesian inference for an adaptive Ordered Probit model: An application to Brain Computer Interfacing
Jiwon Yoon 0001, Stephen J. Roberts, Matthew Dyson, John Q. Gan |
Neural Networks | 2 |
| 2010 | Active Data Selection for Sensor Networks with Faults and ChangepointsabstractWe describe a Bayesian formalism for the intelligent selection of observations from sensor networks that may intermittently undergo faults or changepoints. Such active data selection is performed with the goal of taking as few observations as necessary in order to maintain a reasonable level of uncertainty about the variables of interest. The presence of faults/changepoints is not always obvious and therefore our algorithm must first detect their occurrence. Having done so, our selection of observations must be appropriately altered. Faults corrupt our observations, reducing their impact; changepoints (abrupt changes in the characteristics of data) may require the transition to an entirely different sampling schedule. Our solution is to employ a Gaussian process formalism that allows for sequential time-series prediction about variables of interest along with a decision theoretic approach to the problem of selecting observations. Michael A. Osborne, Roman Garnett, Stephen J. Roberts |
AINA | 3 |
| 2010 | An introduction to Gaussian processes for the Kalman filter expert
Steven Reece, Stephen J. Roberts |
FUSION | 2 |
| 2010 | Bayesian optimization for sensor set selectionabstractWe consider the problem of selecting an optimal set of sensors, as determined, for example, by the predictive accuracy of the resulting sensor network. Given an underlying metric between pairs of set elements, we introduce a natural metric between sets of sensors for this task. Using this metric, we can construct covariance functions over sets, and thereby perform Gaussian process inference over a function whose domain is a power set. If the function has additional inputs, our covariances can be readily extended to incorporate them---allowing us to consider, for example, functions over both sets and time. These functions can then be optimized using Gaussian process global optimization (GPGO). We use the root mean squared error (RMSE) of the predictions made using a set of sensors at a particular time as an example of such a function to be optimized; the optimal point specifies the best choice of sensor locations. We demonstrate the resulting method by dynamically selecting the best subset of a given set of weather sensors for the prediction of the air temperature across the United Kingdom. Roman Garnett, Michael A. Osborne, Stephen J. Roberts |
IPSN | 3 |
| 2010 | Sequential Bayesian Prediction in the Presence of Changepoints and FaultsabstractWe introduce a new sequential algorithm for making robust predictions in the presence of changepoints. Unlike previous approaches, which focus on the problem of detecting and locating changepoints, our algorithm focuses on the problem of making predictions even when such changes might be present. We introduce nonstationary covariance functions to be used in Gaussian process prediction that model such changes, and then proceed to demonstrate how to effectively manage the hyperparameters associated with those covariance functions. We further introduce covariance functions to be used in situations where our observation model undergoes changes, as is the case for sensor faults. By using Bayesian quadrature, we can integrate out the hyperparameters, allowing us to calculate the full marginal predictive distribution. Furthermore, if desired, the posterior distribution over putative changepoint locations can be calculated as a natural byproduct of our prediction algorithm. Roman Garnett, Michael A. Osborne, Steven Reece, Alex Rogers, Stephen J. Roberts |
Comput. J. | 5 |
| 2010 | Sequential Dynamic Classification Using Latent Variable ModelsabstractAdaptive classification is an important online problem in data analysis. The nonlinear and nonstationary nature of much data makes standard static approaches unsuitable. In this paper, we propose a set of sequential dynamic classification algorithms based on extension of nonlinear variants of Bayesian Kalman processes and dynamic generalized linear models. The approaches are shown to work well not only in their ability to track changes in the underlying decision surfaces but also in their ability to handle in a principled manner missing data. We investigate both situations in which target labels are unobserved and also where incoming sensor data are unavailable. We extend the models to allow for active label requesting for use in situations in which there is a cost associated with such information and hence a fully labelled target set is prohibitive. Seung Min Lee, Stephen J. Roberts |
Comput. J. | 2 |
| 2010 | Sequential non-stationary dynamic classification with sparse feedback
D. R. Lowne, Stephen J. Roberts, Roman Garnett |
Pattern Recognit. | 2 |
| 2010 | The Near Constant Acceleration Gaussian Process Kernel for TrackingabstractTime series prediction is traditionally the domain of the state-based Kalman filter and very general Kalman filter process models, such as the near constant acceleration model (NCAM), have been developed to successfully track moving targets. However, the standard Kalman filter uses Markov process models and, consequently, it is difficult to track processes which include a complex periodic component. Gaussian processes are a generalisation of the Kalman filter and are able to model periodic behaviour efficiently and succinctly. However, no equivalent Gaussian process model for near constant acceleration has been formulated. We develop an equivalent Gaussian process kernel for NCAM to be used for time-series prediction. Steven Reece, Stephen J. Roberts |
IEEE Signal Process. Lett. | 2 |
| 2010 | Robust Measurement Validation in Target Tracking Using Geometric StructureabstractSelection schemes for forming data validation regions for target tracking are discussed in this paper. We develop a novel algorithm, less sensitive to gate size than conventional approaches. This new gate selection method combines a conventional threshold based algorithm with a geometric metric measure based on theVoronoi diagram. An adaptive search based on the Voronoi measure is then used to select valid data points for target tracking. Jiwon Yoon 0001, Stephen J. Roberts |
IEEE Signal Process. Lett. | 2 |
| 2009 | Multi-sensor fault recovery in the presence of known and unknown fault types
Steven Reece, Stephen J. Roberts, Christopher Claxton, David Nicholson |
FUSION | 2 |
| 2009 | Sequential Bayesian prediction in the presence of changepointsabstractWe introduce a new sequential algorithm for making robust predictions in the presence of changepoints. Unlike previous approaches, which focus on the problem of detecting and locating changepoints, our algorithm focuses on the problem of making predictions even when such changes might be present. We introduce nonstationary covariance functions to be used in Gaussian process prediction that model such changes, then proceed to demonstrate how to effectively manage the hyperparameters associated with those covariance functions. By using Bayesian quadrature, we can integrate out the hyperparameters, allowing us to calculate the marginal predictive distribution. Furthermore, if desired, the posterior distribution over putative changepoint locations can be calculated as a natural byproduct of our prediction algorithm. Roman Garnett, Michael A. Osborne, Stephen J. Roberts |
ICML | 3 |
| 2009 | Bayesian Methods for Image Super-ResolutionabstractWe present a novel method of Bayesian image super-resolution in which marginalization is carried out over latent parameters such as geometric and photometric registration and the image point-spread function. Related Bayesian super-resolution approaches marginalize over the high-resolution image, necessitating the use of an unfavourable image prior, whereas our method allows for more realistic image prior distributions, and reduces the dimension of the integral considerably, removing the main computational bottleneck of algorithms such as Tipping and Bishop's Bayesian image super-resolution. We show results on real and synthetic datasets to illustrate the efficacy of our method. Lyndsey C. Pickup, David P. Capel, Stephen J. Roberts, Andrew Zisserman |
Comput. J. | 3 |
| 2009 | Adaptive classification for Brain Computer Interface systems using Sequential Monte Carlo sampling
Jiwon Yoon 0001, Stephen J. Roberts, Matthew Dyson, John Q. Gan |
Neural Networks | 2 |
| 2008 | On-line novelty detection using the Kalman filter and extreme value theoryabstractNovelty detection is concerned with identifying abnormal system behaviours and abrupt changes from one regime to another. This paper proposes an on-line (causal) novelty detection method capable of detecting both outliers and regime change points in sequential time-series data. Our approach is based on a Kalman filter in order to model time-series data and extreme value theory is used to compute a novelty measure in a principled manner. The proposed approach is shown to be effective via experiments on several real-world data sets. Hyoungjoo Lee, Stephen J. Roberts |
ICPR | 2 |
| 2008 | Adaptive Classification by Hybrid EKF with Truncated Filtering: Brain Computer Interfacing
Jiwon Yoon 0001, Stephen J. Roberts, Matthew Dyson, John Q. Gan |
IDEAL | 2 |
| 2008 | Towards Real-Time Information Processing of Sensor Network Data Using Computationally Efficient Multi-output Gaussian ProcessesabstractIn this paper, we describe a novel, computationally efficient algorithm that facilitates the autonomous acquisition of readings from sensor networks (deciding when and which sensor to acquire readings from at any time), and which can, with minimal domain knowledge, perform a range of information processing tasks including modelling the accuracy of the sensor readings, predicting the value of missing sensor readings, and predicting how the monitored environmental variables will evolve into the future. Our motivating scenario is the need to provide situational awareness support to first responders at the scene of a large scale incident, and to this end, we describe a novel iterative formulation of a multi-output Gaussian process that can build and exploit a probabilistic model of the environmental variables being measured (including the correlations and delays that exist between them). We validate our approach using data collected from a network of weather sensors located on the south coast of England. Michael A. Osborne, Stephen J. Roberts, Alex Rogers, Sarvapali D. Ramchurn, Nicholas R. Jennings |
IPSN | 2 |
| 2008 | Information Agents for Pervasive Sensor NetworksabstractIn this paper, we describe an information agent, that resides on a mobile computer or personal digital assistant (PDA), that can autonomously acquire sensor readings from pervasive sensor networks (deciding when and which sensor to acquire readings from at any time). Moreover, it can perform a range of information processing tasks including modelling the accuracy of the sensor readings, predicting the value of missing sensor readings, and predicting how the monitored environmental parameters will evolve into the future. Our motivating scenario is the need to provide situational awareness support to first responders at the scene of a large scale incident, and we describe how we use an iterative formulation of a multi-output Gaussian process to build a probabilistic model of the environmental parameters being measured by local sensors, and the correlations and delays that exist between them. We validate our approach using data collected from a network of weather sensors located on the south coast of England. Alex Rogers, Mike Osborne, Sarvapali D. Ramchurn, Stephen J. Roberts, Nicholas R. Jennings |
PerCom | 4 |
| 2008 | On Similarities between Inference in Game Theory and Machine LearningabstractIn this paper, we elucidate the equivalence between inference in game theory and machine learning. Our aim in so doing is to establish an equivalent vocabulary between the two domains so as to facilitate developments at the intersection of both fields, and as proof of the usefulness of this approach, we use recent developments in each field to make useful improvements to the other. More specifically, we consider the analogies between smooth best responses in fictitious play and Bayesian inference methods. Initially, we use these insights to develop and demonstrate an improved algorithm for learning in games based on probabilistic moderation. That is, by integrating over the distribution of opponent strategies (a Bayesian approach within machine learning) rather than taking a simple empirical average (the approach used in standard fictitious play) we derive a novel moderated fictitious play algorithm and show that it is more likely than standard fictitious play to converge to a payoff-dominant but risk-dominated Nash equilibrium in a simple coordination game. Furthermore we consider the converse case, and show how insights from game theory can be used to derive two improved mean field variational learning algorithms. We first show that the standard update rule of mean field variational learning is analogous to a Cournot adjustment within game theory. By analogy with fictitious play, we then suggest an improved update rule, and show that this results in fictitious variational play, an improved mean field variational learning algorithm that exhibits better convergence in highly or strongly connected graphical models. Second, we use a recent advance in fictitious play, namely dynamic fictitious play, to derive a derivative action variational learning algorithm, that exhibits superior convergence properties on a canonical machine learning problem (clustering a mixture distribution). Iead Rezek, David S. Leslie, Steven Reece, Stephen J. Roberts, Alex Rogers, Rajdeep K. Dash, Nicholas R. Jennings |
J. Artif. Intell. Res. | 4 |
| 2007 | A Multi-Dimensional Trust Model for Heterogeneous Contract Observations
Steven Reece, Stephen J. Roberts, Alex Rogers, Nicholas R. Jennings |
AAAI | 2 |
| 2006 | Optimizing and Learning for Super-resolutionabstractIn multiple-image super-resolution, a high resolution image is estimated from a number of lower-resolution images. This involves computing the parameters of a generative imaging model (such as geometric and photometric registration, and blur) and obtaining a MAP estimate by minimizing a cost function including an appropriate prior. We consider the quite general geometric registration situation modelled by a plane projective transformation, and make two novel contributions: (i) in previous approaches the MAP estimate has been obtained by first computing and fixing the registration, and then computing the super-resolution image with this registration. We demonstrate that superior estimates are obtained by optimizing over both the registration and image; (ii) the parameters of the edge preserving prior are learnt automatically from the data, rather than being set by trial and error. We show examples on a number of real sequences including multiple stills, digital video, and DVDs of movies. 1 Lyndsey C. Pickup, Stephen J. Roberts, Andrew Zisserman |
BMVC | 2 |
| 2006 | Computational Mechanism Design for Information Fusion within Sensor NetworksabstractConventional centralised information fusion and control architectures will be challenged by developments in sensor networks that allow sophisticated autonomous sensors, owned by different stakeholders with individual goals, to interact and share information. Given this, we advocate the use of tools and techniques from computational mechanism design (CMD), a field at the intersection of computer science, game theory and economics, to address the challenges posed by these networks. In particular, CMD allows us to engineer networks with desirable system-wide properties, in which sensors act as rational selfish agents, each attempting to fulfil their own individuals goals through the exchange of observations and information. In this paper, we present our work developing such networks. Specifically, we discuss our development of a generic and principled information valuation metric for sensor networks and we report our experiences applying it within a real world information fusion sensor network scenario Alex Rogers, Rajdeep K. Dash, Nicholas R. Jennings, Steven Reece, Stephen J. Roberts |
FUSION | 5 |
| 2006 | Nonlinear, Biophysically-Informed Speech Pathology DetectionabstractThis paper reports a simple nonlinear approach to online acoustic speech pathology detection for automatic screening purposes. Straightforward linear preprocessing followed by two nonlinear measures, based parsimoniously upon the biophysics of speech production, combined with subsequent linear classification, achieves an overall normal/pathological detection performance of 91.4%, and over 99% with rejection of 15% ambiguous cases. This compares favourably with more complex, computationally intensive methods based on a large number of linear and other measures. This demonstrates that nonlinear approaches to speech pathology detection, informed by biophysics, can be both simple and robust, and are amenable to implementation as online algorithms Max A. Little, Patrick E. McSharry, Irene M. Moroz, Stephen J. Roberts |
ICASSP (2) | 4 |
| 2006 | Bayesian Image Super-resolution, ContinuedabstractThis paper develops a multi-frame image super-resolution approach from a Bayesian view-point by marginalizing over the unknown registration parameters relating the set of input low-resolution views. In Tipping and Bishop’s Bayesian image super-resolution approach [16], the marginalization was over the super- resolution image, necessitating the use of an unfavorable image prior. By inte- grating over the registration parameters rather than the high-resolution image, our method allows for more realistic prior distributions, and also reduces the dimen- sion of the integral considerably, removing the main computational bottleneck of the other algorithm. In addition to the motion model used by Tipping and Bishop, illumination components are introduced into the generative model, allowing us to handle changes in lighting as well as motion. We show results on real and synthetic datasets to illustrate the efficacy of this approach. Lyndsey C. Pickup, David P. Capel, Stephen J. Roberts, Andrew Zisserman |
NIPS | 3 |
| 2005 | Depth of anaesthesia assessment with generative polyspectral modelsabstractThe application of anaesthetic agents is known to have significant effects on the EEG waveform. Information extraction now routinely goes beyond second order spectral analysis, as obtained via power spectral methods, and uses higher order spectral methods. In this paper we present a model which generalises the autoregressive class of polyspectral models by having a semi-parametric description of the residual probability density. We estimate the model in the variational Bayesian framework and extract higher order spectral features. Testing their importance for depth of anaesthesia classification is done on three different EEG data sets collected under exposure to different agents. The results show that significant improvements can be made over standard methods of estimating higher order spectra. The results also indicate that in two out of three anaesthetic agents, better classification can be achieved with higher order spectral features. Iead Rezek, Stephen J. Roberts, Ellini Siva, R. Conradt |
ICMLA | 2 |
| 2004 | Hierarchy, priors and wavelets: structure and signal modelling using ICA
Stephen J. Roberts, Evangelos Roussos, Rizwan Choudrey |
Signal Process. | 1 |
| 2003 | Markov Models for Automated ECG Interval AnalysisabstractWe examine the use of hidden Markov and hidden semi-Markov mod- els for automatically segmenting an electrocardiogram waveform into its constituent waveform features. An undecimated wavelet transform is used to generate an overcomplete representation of the signal that is more appropriate for subsequent modelling. We show that the state dura- tions implicit in a standard hidden Markov model are ill-suited to those of real ECG features, and we investigate the use of hidden semi-Markov models for improved state duration modelling. Nicholas P. Hughes, Lionel Tarassenko, Stephen J. Roberts |
NIPS | 3 |
| 2003 | A Sampled Texture Prior for Image Super-ResolutionabstractSuper-resolution aims to produce a high-resolution image from a set of one or more low-resolution images by recovering or inventing plausible high-frequency image content. Typical approaches try to reconstruct a high-resolution image using the sub-pixel displacements of several low- resolution images, usually regularized by a generic smoothness prior over the high-resolution image space. Other methods use training data to learn low-to-high-resolution matches, and have been highly successful even in the single-input-image case. Here we present a domain-specific im- age prior in the form of a p.d.f. based upon sampled images, and show that for certain types of super-resolution problems, this sample-based prior gives a significant improvement over other common multiple-image super-resolution techniques. Lyndsey C. Pickup, Stephen J. Roberts, Andrew Zisserman |
NIPS | 2 |
| 2003 | Variational Mixture of Bayesian Independent Component AnalyzersabstractThere has been growing interest in subspace data modeling over the past few years. Methods such as principal component analysis, factor analysis, and independent component analysis have gained in popularity and have found many applications in image modeling, signal processing, and data compression, to name just a few. As applications and computing power grow, more and more sophisticated analyses and meaningful representations are sought. Mixture modeling methods have been proposed for principal and factor analyzers that exploit local gaussian features in the subspace manifolds. Meaningful representations may be lost, however, if these local features are nongaussian or discontinuous. In this article, we propose extending the gaussian analyzers mixture model to an independent component analyzers mixture model. We employ recent developments in variational Bayesian inference and structure determination to construct a novel approach for modeling nongaussian, discontinuous manifolds. We automatically determine the local dimensionality of each manifold and use variational inference to calculate the optimum number of ICA components needed in our mixture model. We demonstrate our framework on complex synthetic data and illustrate its application to real data by decomposing functional magnetic resonance images into meaningful-and medically useful-features. Rizwan Choudrey, Stephen J. Roberts |
Neural Comput. | 2 |
| 2003 | Data decomposition using independent component analysis with prior constraints
Stephen J. Roberts, Rizwan Choudrey |
Pattern Recognit. | 1 |
| 2002 | Adaptive Classification by Variational Kalman FilteringabstractWe propose in this paper a probabilistic approach for adaptive inference of generalized nonlinear classification that combines the computational advantage of a parametric solution with the flexibility of sequential sam- pling techniques. We regard the parameters of the classifier as latent states in a first order Markov process and propose an algorithm which can be regarded as variational generalization of standard Kalman filter- ing. The variational Kalman filter is based on two novel lower bounds that enable us to use a non-degenerate distribution over the adaptation rate. An extensive empirical evaluation demonstrates that the proposed method is capable of infering competitive classifiers both in stationary and non-stationary environments. Although we focus on classification, the algorithm is easily extended to other generalized nonlinear models. Peter Sykacek, Stephen J. Roberts |
NIPS | 2 |
| 2002 | Towards the automatic analysis of complex human body motions
Jens Rittscher, Andrew Blake 0001, Stephen J. Roberts |
Image Vis. Comput. | 3 |
| 2001 | Minimum-Entropy Data Clustering Using Reversible Jump Markov Chain Monte Carlo
Stephen J. Roberts, Christopher C. Holmes, Dave Denison |
ICANN | 1 |
| 2001 | Mixtures of Independent Component Analysers
Stephen J. Roberts, William D. Penny |
ICANN | 1 |
| 2001 | A Probabilistic Approach to High-Resolution Sleep Analysis
Peter Sykacek, Stephen J. Roberts, Iead Rezek, Arthur Flexer, Georg Dorffner |
ICANN | 2 |
| 2001 | Bayesian time series classificationabstractThis paper proposes an approach to classification of adjacent segments of a time series as being either of classes. We use a hierarchical model that consists of a feature extraction stage and a generative classifier which is built on top of these features. Such two stage approaches are often used in signal and image processing. The novel part of our work is that we link these stages probabilistically by using a latent feature space. To use one joint model is a Bayesian requirement, which has the advantage to fuse information according to its certainty. The classifier is implemented as hidden Markov model with Gaussian and Multinomial observation distributions defined on a suitably chosen representation of autoregressive models. The Markov dependency is mo- tivated by the assumption that successive classifications will be corre- lated. Inference is done with Markov chain Monte Carlo (MCMC) tech- niques. We apply the proposed approach to synthetic data and to classi- fication of EEG that was recorded while the subjects performed different cognitive tasks. All experiments show that using a latent feature space results in a significant improvement in generalization accuracy. Hence we expect that this idea generalizes well to other hierarchical models. Peter Sykacek, Stephen J. Roberts |
NIPS | 2 |
| 2001 | Minimum-Entropy Data Partitioning Using Reversible Jump Markov Chain Monte CarloabstractProblems in data analysis often require the unsupervised partitioning of a data set into classes. Several methods exist for such partitioning but many have the weakness of being formulated via strict parametric models (e.g., each class is modeled by a single Gaussian) or being computationally intensive in high-dimensional data spaces. We reconsider the notion of such cluster analysis in information-theoretic terms and show that an efficient partitioning may be given via a minimization of partition entropy. A reversible-jump sampling is introduced to explore the variable-dimension space of partition models. Stephen J. Roberts, Christopher C. Holmes, Dave Denison |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Maximum certainty data partitioning
Stephen J. Roberts, Richard M. Everson, Iead Rezek |
Pattern Recognit. | 1 |
| 1999 | Neural networks for predicting Kaposi's sarcomaabstractThis paper demonstrates a medical application of Bayesian neural networks, whose parameters and hyper-parameters are sampled from the posterior distribution by means of Monte Carlo Markov chain. The main objective is the determination of the relevance of various input variables. The paper focuses on typical difficulties one has to face when dealing with sparse data sets. Dirk Husmeier, Gillian S. Patton, Myra O. McClure, John R. W. Harris, Stephen J. Roberts |
IJCNN | 5 |
| 1999 | Dynamic logistic regressionabstractWe propose an online learning algorithm for training a logistic regression model on nonstationary classification problems. The nonstationarity is captured by modelling the weights in a logistic regression classifier as evolving according to a first order Markov process. The weights are updated using the extended Kalman filter formalism and nonstationarities are tracked by inferring a time-varying state noise variance parameter. We describe an algorithm for doing this based on maximising the evidence of updated predictions. The algorithm is illustrated on a number of synthetic problems. William D. Penny, Stephen J. Roberts |
IJCNN | 2 |
| 1999 | EEG-based communication via dynamic neural network modelsabstractThe overall aim of this research is to develop an EEG-based computer interface. We report on an offline analysis of EEG data recorded from 7 subjects performing two different pairs of cognitive tasks; motor imagery versus a baseline task and motor imagery versus a maths task. For the imagery versus baseline pairing, discrimination was good in three subjects, marginal in two and not possible in the other two. For the imagery versus maths pairing, discrimination was very good in two subjects, good in 4 and marginal in one. The data was analysed using lagged-AR feature vectors and a Bayesian logistic regression classifier with temporal smoothing. Enhanced spectra are shown highlighting differential spectral activity for each task pairing. The results suggest that combinations of different task pairings and dynamic neural network models have the potential to drastically reduce the time it takes for a new user to learn to use an EEG-based computer interface. William D. Penny, Stephen J. Roberts |
IJCNN | 2 |
| 1999 | Dynamic Models for Nonstationary Signal Segmentation
William D. Penny, Stephen J. Roberts |
Comput. Biomed. Res. | 2 |
| 1999 | Independent Component Analysis: A Flexible Nonlinearity and Decorrelating Manifold ApproachabstractIndependent component analysis (ICA) finds a linear transformation to variables that are maximally statistically independent. We examine ICA and algorithms for finding the best transformation from the point of view of maximizing the likelihood of the data. In particular, we discuss the way in which scaling of the unmixing matrix permits a "static" nonlinearity to adapt to various marginal densities. We demonstrate a new algorithm that uses generalized exponential functions to model the marginal densities and is able to separate densities with light tails. We characterize the manifold of decorrelating matrices and show that it lies along the ridges of high-likelihood unmixing matrices in the space of all unmixing matrices. We show how to find the optimum ICA matrix on the manifold of decorrelating matrices, and as an example we use the algorithm to find independent component basis vectors for an ensemble of portraits. Richard M. Everson, Stephen J. Roberts |
Neural Comput. | 2 |
| 1999 | An empirical evaluation of Bayesian sampling with hybrid Monte Carlo for training neural network classifiers
Dirk Husmeier, William D. Penny, Stephen J. Roberts |
Neural Networks | 3 |
| 1999 | Bayesian neural networks for classification: how useful is the evidence framework?
William D. Penny, Stephen J. Roberts |
Neural Networks | 2 |
| 1998 | Bayesian Approaches to Gaussian Mixture ModelingabstractA Bayesian-based methodology is presented which automatically penalizes overcomplex models being fitted to unknown data. We show that, with a Gaussian mixture model, the approach is able to select an "optimal" number of components in the model and so partition data sets. The performance of the Bayesian method is compared to other methods of optimal model selection and found to give good results. The methods are tested on synthetic and real data sets. Stephen J. Roberts, Dirk Husmeier, Iead Rezek, William D. Penny |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | Parametric and non-parametric unsupervised cluster analysis
Stephen J. Roberts |
Pattern Recognit. | 1 |
| 1996 | Scale-space unsupervised cluster analysisabstractMost scientific disciplines generate experimental data from an observed system about which we have may have little understanding of the data generating function. It is attractive, therefore, for an analysis system to break a complex data set into a series of piecewise similar groups or structures, each of which may then be regarded as a separate data state, for example, thus reducing overall data complexity. Cluster analysis has a long and rich history and excellent reviews of many methods may be found in Jain-Dubes (1988), Jain (1982), Hartigan (1975) and Everitt (1974). This paper presents a scale-space method of unsupervised clustering (the 'optimal' number of partitions is unknown a priori). Its performance is compared to that of a Gaussian-mixture model (GMM) approach using both maximum-likelihood and K-means algorithms. The multi-scale method may be seen as falling within the hierarchical clustering genre or as a method of scale-space (multiresolution) parameter estimation. We show that the GMM fails for data sets which are not multivariate Gaussian whilst the scale-space method is considerably more robust. Stephen J. Roberts |
ICPR | 1 |
| 1994 | A Probabilistic Resource Allocating Network for Novelty DetectionabstractThe detection of novel or abnormal input vectors is of importance in many monitoring tasks, such as fault detection in complex systems and detection of abnormal patterns in medical diagnostics. We have developed a robust method for novelty detection, which aims to minimize the number of heuristically chosen thresholds in the novelty decision process. We achieve this by growing a gaussian mixture model to form a representation of a training set of “normal” system states. When previously unseen data are to be screened for novelty we use the same threshold as was used during training to define a novelty decision boundary. We show on a sample problem of medical signal processing that this method is capable of providing robust novelty decision boundaries and apply the technique to the detection of epileptic seizures within a data record. Stephen J. Roberts, Lionel Tarassenko |
Neural Comput. | 1 |