VLDB 2026 Research / reviewers in the wild / expert
David Barber
dblp:47/4760
· DBLP profile ↗
80ranked-venue papers
20as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 70 · 18 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Training Neural Samplers with Reverse Diffusive KL DivergenceabstractTraining generative models to sample from unnormalized density functions is an important and challenging task in machine learning. Traditional training methods often rely on the reverse Kullback-Leibler (KL) divergence due to its tractability. However, the mode-seeking behavior of reverse KL hinders effective approximation of multi-modal target distributions. To address this, we propose to minimize the reverse KL along diffusion trajectories of both model and target densities. We refer to this objective as the reverse diffusive KL divergence, which allows the model to capture multiple modes. Leveraging this objective, we train neural samplers that can efficiently generate samples from the target distribution in one step. We demonstrate that our method enhances sampling performance across various Boltzmann distributions, including both synthetic multi-modal densities and n-body particle systems. Jiajun He 0003, Mingtian Zhang, David Barber, José Miguel Hernández-Lobato |
AISTATS | 4 |
| 2025 | Automatic Diagnosis of Hip and Knee Osteoarthritis from Medical Text RecordsabstractOsteoarthritis (OA) is a highly frequent musculoskeletal condition defined by the progressive degradation of joint cartilage and the underlying bone, resulting in the manifestation of pain and functional limitations. Identification of symptoms as early as possible for timely intervention is critical for effective pain management and treatment. Knee and hip OA are common in older patients which can greatly affect their mobility, lifestyle, and lead to other health complications. Electronic Medical Records (EMR) in primary care settings contain patients’ structured historical data including unstructured text data in the encounter chart notes. The unstructured notes are often very long, compiled from multiple patient-physician encounters, and contain medical jargon including personal data. The data offers a variety of computational challenges but it contains valuable information for disease diagnosis especially for detecting hip or knee OA. We demonstrate multiple keyword-based strategies to detect the OA-affected bone joints, including a simple rule-based approach and a machine learning based approach. We also provide an ablation study to show the effects of the different natural language text processing methods and validate our results against gold standard data labelled by a human expert. Our Random Forest (RF) model achieved the best result of 74.89% F1-score with OA related paragraph extraction. Jiahao Cai, Vidhi Kokel, Farhana Zulkernine, John A. Queenan, David Barber |
COMPSAC | 5 |
| 2025 | Improving Probabilistic Diffusion Models With Optimal Diagonal Covariance MatchingabstractThe probabilistic diffusion model has become highly effective across various domains. Typically, sampling from a diffusion model involves using a denoising distribution characterized by a Gaussian with a learned mean and either fixed or learned covariances. In this paper, we leverage the recently proposed covariance moment matching technique and introduce a novel method for learning the diagonal covariances. Unlike traditional data-driven covariance approximation approaches, our method involves directly regressing the optimal analytic covariance using a new, unbiased objective named Optimal Covariance Matching (OCM). This approach can significantly reduce the approximation error in covariance prediction. We demonstrate how our method can substantially enhance the sampling efficiency, recall rate and likelihood of both diffusion models and latent diffusion models. Zijing Ou, Mingtian Zhang, Andi Zhang 0001, Tim Z. Xiao, Yingzhen Li, David Barber |
ICLR | 6 |
| 2025 | Incremental Sequence Classification with Temporal ConsistencyabstractWe address the problem of incremental sequence classification, where predictions are updated as new elements in the sequence are revealed. Drawing on temporal-difference learning from reinforcement learning, we identify a temporal-consistency condition that successive predictions should satisfy. We leverage this condition to develop a novel loss function for training incremental sequence classifiers. Through a concrete example, we demonstrate that optimizing this loss can offer substantial gains in data efficiency. We apply our method to text classification tasks and show that it improves predictive accuracy over competing approaches on several benchmark datasets. We further evaluate our approach on the task of verifying large language model generations for correctness in grade-school math problems. Our results show that models trained with our method are better able to distinguish promising generations from unpromising ones after observing only a few tokens. Lucas Maystre, Gabriel Barello, Tudor Berariu, Aleix Cambray, Rares Dolga, Alvaro Ortega Gonzalez, Andrei Cristian Nica, David Barber |
NeurIPS | 8 |
| 2024 | Diffusive Gibbs SamplingabstractThe inadequate mixing of conventional Markov Chain Monte Carlo (MCMC) methods for multi-modal distributions presents a significant challenge in practical applications such as Bayesian inference and molecular dynamics. Addressing this, we propose Diffusive Gibbs Sampling (DiGS), an innovative family of sampling methods designed for effective sampling from distributions characterized by distant and disconnected modes. DiGS integrates recent developments in diffusion models, leveraging Gaussian convolution to create an auxiliary noisy distribution that bridges isolated modes in the original space and applying Gibbs sampling to alternately draw samples from both spaces. A novel Metropolis-within-Gibbs scheme is proposed to enhance mixing in the denoising sampling step. DiGS exhibits a better mixing property for sampling multi-modal distributions than state-of-the-art methods such as parallel tempering, attaining substantially improved performance across various tasks, including mixtures of Gaussians, Bayesian neural networks and molecular dynamics. Mingtian Zhang, Brooks Paige, José Miguel Hernández-Lobato, David Barber |
ICML | 5 |
| 2024 | Active Preference Learning for Large Language ModelsabstractAs large language models (LLMs) become more capable, fine-tuning techniques for aligning with human intent are increasingly important. A key consideration for aligning these models is how to most effectively use human resources, or model resources in the case where LLMs themselves are used as oracles. Reinforcement learning from Human or AI preferences (RLHF/RLAIF) is the most prominent example of such a technique, but is complex and often unstable. Direct Preference Optimization (DPO) has recently been proposed as a simpler and more stable alternative. In this work, we develop an active learning strategy for DPO to make better use of preference labels. We propose a practical acquisition function for prompt/completion pairs based on the predictive entropy of the language model and a measure of certainty of the implicit preference model optimized by DPO. We demonstrate how our approach improves both the rate of learning and final performance of fine-tuning on pairwise preference data. William Muldrew, Peter Hayes, Mingtian Zhang, David Barber |
ICML | 4 |
| 2024 | CenTime: Event-conditional modelling of censoring in survival analysisabstractSurvival analysis is a valuable tool for estimating the time until specific events, such as death or cancer recurrence, based on baseline observations. This is particularly useful in healthcare to prognostically predict clinically important events based on patient data. However, existing approaches often have limitations; some focus only on ranking patients by survivability, neglecting to estimate the actual event time, while others treat the problem as a classification task, ignoring the inherent time-ordered structure of the events. Additionally, the effective utilisation of censored samples-data points where the event time is unknown- is essential for enhancing the model's predictive accuracy. In this paper, we introduce CenTime, a novel approach to survival analysis that directly estimates the time to event. Our method features an innovative event-conditional censoring mechanism that performs robustly even when uncensored data is scarce. We demonstrate that our approach forms a consistent estimator for the event model parameters, even in the absence of uncensored data. Furthermore, CenTime is easily integrated with deep learning models with no restrictions on batch size or the number of uncensored samples. We compare our approach to standard survival analysis methods, including the Cox proportional-hazard model and DeepHit. Our results indicate that CenTime offers state-of-the-art performance in predicting time-to-death while maintaining comparable ranking performance. Our implementation is publicly available at https://github.com/ahmedhshahin/CenTime. Ahmed H. Shahin, Alexander C. Whitehead, Daniel C. Alexander, Joseph Jacob, David Barber |
Medical Image Anal. | 6 |
| 2023 | Moment Matching Denoising Gibbs SamplingabstractEnergy-Based Models (EBMs) offer a versatile framework for modelling complex data distributions. However, training and sampling from EBMs continue to pose significant challenges. The widely-used Denoising Score Matching (DSM) method for scalable EBM training suffers from inconsistency issues, causing the energy model to learn a noisy data distribution. In this work, we propose an efficient sampling framework: (pseudo)-Gibbs sampling with moment matching, which enables effective sampling from the underlying clean model when given a noisy model that has been well-trained via DSM. We explore the benefits of our approach compared to related methods and demonstrate how to scale the method to high-dimensional datasets. Mingtian Zhang, Alex Hawkins-Hooker, Brooks Paige, David Barber |
NeurIPS | 4 |
| 2022 | Prognostic Imaging Biomarker Discovery in Survival Analysis for Idiopathic Pulmonary Fibrosis
Ahmed H. Shahin, Eyjolfur Gudmundsson, Adam Szmul, Nesrin Mogulkoc, Frouke Van Beek, Christopher Brereton, Hendrik W. Van Es, Katarina Pontoppidan, Recep Savas, Timothy Wallis, Omer Unat, Marcel Veltkamp, Mark G. Jones, Coline H. M. Van Moorsel, David Barber, Joseph Jacob, Daniel C. Alexander |
MICCAI (8) | 17 |
| 2022 | Generalization Gap in Amortized InferenceabstractThe ability of likelihood-based probabilistic models to generalize to unseen data is central to many machine learning applications such as lossless compression. In this work, we study the generalization of a popular class of probabilistic model - the Variational Auto-Encoder (VAE). We discuss the two generalization gaps that affect VAEs and show that overfitting is usually dominated by amortized inference. Based on this observation, we propose a new training objective that improves the generalization of amortized inference. We demonstrate how our method can improve performance in the context of image modeling and lossless compression. Mingtian Zhang, Peter Hayes, David Barber |
NeurIPS | 3 |
| 2021 | Improving Gaussian mixture latent variable model convergence with Optimal TransportabstractGenerative models with both discrete and continuous latent variables are highly motivated by the structure of many real-world data sets. They present, however, subtleties in training often manifesting in the discrete latent variable not being leveraged. In this paper, we show why such models struggle to train using traditional log-likelihood maximization, and that they are amenable to training using the Optimal Transport framework of Wasserstein Autoencoders. We find our discrete latent variable to be fully leveraged by the model when trained, without any modifications to the objective function or significant fine tuning. Our model generates comparable samples to other approaches while using relatively simple neural networks, since the discrete latent variable carries much of the descriptive burden. Furthermore, the discrete latent provides significant control over generation. Benoit Gaujac, Ilya Feige, David Barber |
ACML | 3 |
| 2021 | Reducing the Computational Cost of Deep Generative Models with Binary Neural Networks
Thomas Bird, Friso H. Kingma, David Barber |
ICLR | 3 |
| 2021 | Addressing Catastrophic Forgetting in Few-Shot ProblemsabstractNeural networks are known to suffer from catastrophic forgetting when trained on sequential datasets. While there have been numerous attempts to solve this problem in large-scale supervised classification, little has been done to overcome catastrophic forgetting in few-shot classification problems. We demonstrate that the popular gradient-based model-agnostic meta-learning algorithm (MAML) indeed suffers from catastrophic forgetting and introduce a Bayesian online meta-learning framework that tackles this problem. Our framework utilises Bayesian online learning and meta-learning along with Laplace approximation and variational inference to overcome catastrophic forgetting in few-shot classification problems. The experimental evaluations demonstrate that our framework can effectively achieve this goal in comparison with various baselines. As an additional utility, we also demonstrate empirically that our framework is capable of meta-learning on sequentially arriving few-shot tasks from a stationary task distribution. Pau Ching Yap, Hippolyt Ritter, David Barber |
ICML | 3 |
| 2021 | Learning Disentangled Representations with the Wasserstein Autoencoder
Benoit Gaujac, Ilya Feige, David Barber |
ECML/PKDD (3) | 3 |
| 2020 | HiLLoC: lossless image compression with hierarchical latent variable models
James Townsend, Thomas Bird, Julius Kunze, David Barber |
ICLR | 4 |
| 2020 | Spread DivergenceabstractFor distributions $\mathbb{P}$ and $\mathbb{Q}$ with different supports or undefined densities, the divergence $\textrm{D}(\mathbb{P}||\mathbb{Q})$ may not exist. We define a Spread Divergence $\tilde{\textrm{D}}(\mathbb{P}||\mathbb{Q})$ on modified $\mathbb{P}$ and $\mathbb{Q}$ and describe sufficient conditions for the existence of such a divergence. We demonstrate how to maximize the discriminatory power of a given divergence by parameterizing and learning the spread. We also give examples of using a Spread Divergence to train implicit generative models, including linear models (Independent Components Analysis) and non-linear models (Deep Generative Networks). Mingtian Zhang, Peter Hayes, Thomas Bird, Raza Habib, David Barber |
ICML | 5 |
| 2019 | Tracking by Animation: Unsupervised Learning of Multi-Object Attentive TrackersabstractOnline Multi-Object Tracking (MOT) from videos is a challenging computer vision task which has been extensively studied for decades. Most of the existing MOT algorithms are based on the Tracking-by-Detection (TBD) paradigm combined with popular machine learning approaches which largely reduce the human effort to tune algorithm parameters. However, the commonly used supervised learning approaches require the labeled data (e.g., bounding boxes), which is expensive for videos. Also, the TBD framework is usually suboptimal since it is not end-to-end, i.e., it considers the task as detection and tracking, but not jointly. To achieve both label-free and end-to-end learning of MOT, we propose a Tracking-by-Animation framework, where a differentiable neural model first tracks objects from input frames and then animates these objects into reconstructed frames. Learning is then driven by the reconstruction error through backpropagation. We further propose a Reprioritized Attentive Tracking to improve the robustness of data association. Experiments conducted on both synthetic and real video datasets show the potential of the proposed model. Our project page is publicly available at: https://github.com/zhen-he/tracking-by-animation Jian Li 0003, Daxue Liu, Hangen He, David Barber |
CVPR | 5 |
| 2019 | Auxiliary Variational MCMC
Raza Habib, David Barber |
ICLR (Poster) | 2 |
| 2019 | Practical lossless compression with latent variables using bits back coding
James Townsend, Thomas Bird, David Barber |
ICLR (Poster) | 3 |
| 2019 | Improving latent variable descriptiveness by modelling rather than ad-hoc factorsabstractPowerful generative models, particularly in natural language modelling, are commonly trained by maximizing a variational lower bound on the data log likelihood. These models often suffer from poor use of their latent variable, with ad-hoc annealing factors used to encourage retention of information in the latent variable. We discuss an alternative and general approach to latent variable modelling, based on an objective that encourages a perfect reconstruction by tying a stochastic autoencoder with a variational autoencoder (VAE). This ensures by design that the latent variable captures information about the observations, whilst retaining the ability to generate well. Interestingly, although our model is fundamentally different to a VAE, the lower bound attained is identical to the standard VAE bound but with the addition of a simple pre-factor; thus, providing a formal interpretation of the commonly used, ad-hoc pre-factors in training VAEs. Alex Mansbridge, Roberto Fierimonte, Ilya Feige, David Barber |
Mach. Learn. | 4 |
| 2018 | Generating Sentences Using a Dynamic CanvasabstractWe introduce the Attentive Unsupervised Text (W)riter (AUTR), which is a word level generative model for natural language. It uses a recurrent neural network with a dynamic attention and canvas memory mechanism to iteratively construct sentences. By viewing the state of the memory at intermediate stages and where the model is placing its attention, we gain insight into how it constructs sentences. We demonstrate that AUTR learns a meaningful latent representation for each sentence, and achieves competitive log-likelihood lower bounds whilst being computationally efficient. It is effective at generating and reconstructing sentences, as well as imputing missing words. Harshil Shah, Bowen Zheng 0004, David Barber |
AAAI | 3 |
| 2018 | A Scalable Laplace Approximation for Neural Networks
Hippolyt Ritter, Aleksandar Botev, David Barber |
ICLR (Poster) | 3 |
| 2018 | Modular Networks: Learning to Decompose Neural ComputationabstractScaling model capacity has been vital in the success of deep learning. For a typical network, necessary compute resources and training time grow dramatically with model size. Conditional computation is a promising way to increase the number of parameters with a relatively small increase in resources. We propose a training algorithm that flexibly chooses neural modules based on the data to be processed. Both the decomposition and modules are learned end-to-end. In contrast to existing approaches, training does not rely on regularization to enforce diversity in module use. We apply modular networks both to image recognition and language modeling tasks, where we achieve superior performance compared to several baselines. Introspection reveals that modules specialize in interpretable contexts. Louis Kirsch, Julius Kunze, David Barber |
NeurIPS | 3 |
| 2018 | Online Structured Laplace Approximations for Overcoming Catastrophic ForgettingabstractWe introduce the Kronecker factored online Laplace approximation for overcoming catastrophic forgetting in neural networks. The method is grounded in a Bayesian online learning framework, where we recursively approximate the posterior after every task with a Gaussian, leading to a quadratic penalty on changes to the weights. The Laplace approximation requires calculating the Hessian around a mode, which is typically intractable for modern architectures. In order to make our method scalable, we leverage recent block-diagonal Kronecker factored approximations to the curvature. Our algorithm achieves over 90% test accuracy across a sequence of 50 instantiations of the permuted MNIST dataset, substantially outperforming related methods for overcoming catastrophic forgetting. Hippolyt Ritter, Aleksandar Botev, David Barber |
NeurIPS | 3 |
| 2018 | Generative Neural Machine TranslationabstractWe introduce Generative Neural Machine Translation (GNMT), a latent variable architecture which is designed to model the semantics of the source and target sentences. We modify an encoder-decoder translation model by adding a latent variable as a language agnostic representation which is encouraged to learn the meaning of the sentence. GNMT achieves competitive BLEU scores on pure translation tasks, and is superior when there are missing words in the source sentence. We augment the model to facilitate multilingual translation and semi-supervised learning without adding parameters. This framework significantly reduces overfitting when there is limited paired data available, and is effective for translating between pairs of languages not seen during training. Harshil Shah, David Barber |
NeurIPS | 2 |
| 2017 | Complementary Sum Sampling for Likelihood Approximation in Large Scale ClassificationabstractWe consider training probabilistic classifiers in the case that the number of classes is too large to perform exact normalisation over all classes. We show that the source of high variance in standard sampling approximations is due to simply not including the correct class of the datapoint into the approximation. To account for this we explicitly sum over a subset of classes and sample the remaining. We show that this simple approach is competitive with recently introduced non likelihood-based approximations. Aleksandar Botev, Bowen Zheng 0004, David Barber |
AISTATS | 3 |
| 2017 | Practical Gauss-Newton Optimisation for Deep LearningabstractWe present an efficient block-diagonal approximation to the Gauss-Newton matrix for feedforward neural networks. Our resulting algorithm is competitive against state-of-the-art first-order optimisation methods, with sometimes significant improvement in optimisation performance. Unlike first-order methods, for which hyperparameter tuning of the optimisation parameters is often a laborious process, our approach can provide good performance even when used with default settings. A side result of our work is that for piecewise linear transfer functions, the network objective function can have no differentiable local maxima, which may partially explain why such transfer functions facilitate effective optimisation. Aleksandar Botev, Hippolyt Ritter, David Barber |
ICML | 3 |
| 2017 | Nesterov's accelerated gradient and momentum as approximations to regularised update descentabstractWe present a unifying framework for adapting the update direction in gradient-based iterative optimization methods. As natural special cases we re-derive classical momentum and Nesterov's accelerated gradient method, lending a new intuitive interpretation to the latter algorithm. We show that a new algorithm, which we term Regularised Gradient Descent, can converge more quickly than either Nesterov's algorithm or the classical momentum algorithm. Aleksandar Botev, Guy Lever, David Barber |
IJCNN | 3 |
| 2017 | Overdispersed variational autoencodersabstractThe ability to fit complex generative probabilistic models to data is a key challenge in AI. Currently, variational methods are popular, but remain difficult to train due to high variance of the sampling methods employed. We introduce the overdispersed variational autoencoder and overdispersed importance weighted autoencoder, which combine overdispersed black box variational inference with the variational autoencoder and importance weighted autoencoder respectively. We use the log likelihood lower bounds and reparametrisation trick from the variational and importance weighted autoencoders, but rather than drawing samples from the variational distribution itself, we use importance sampling to draw samples from an overdispersed (i.e. heavier-tailed) proposal in the same family as the variational distribution. We run experiments on two different datasets, and show that this technique produces a lower variance estimate of the gradients, and reaches a higher bound on the log likelihood of the observed data. Harshil Shah, David Barber, Aleksandar Botev |
IJCNN | 2 |
| 2017 | Thinking Fast and Slow with Deep Learning and Tree SearchabstractSequential decision making problems, such as structured prediction, robotic control, and game playing, require a combination of planning policies and generalisation of those plans. In this paper, we present Expert Iteration (ExIt), a novel reinforcement learning algorithm which decomposes the problem into separate planning and generalisation tasks. Planning new policies is performed by tree search, while a deep neural network generalises those plans. Subsequently, tree search is improved by using the neural network policy to guide search, increasing the strength of new plans. In contrast, standard deep Reinforcement Learning algorithms rely on a neural network not only to generalise plans, but to discover them too. We show that ExIt outperforms REINFORCE for training a neural network to play the board game Hex, and our final tree search agent, trained tabula rasa, defeats MoHex1.0, the most recent Olympiad Champion player to be publicly released. Thomas W. Anthony 0001, Zheng Tian 0002, David Barber |
NIPS | 3 |
| 2017 | Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence LearningabstractLong Short-Term Memory (LSTM) is a popular approach to boosting the ability of Recurrent Neural Networks to store longer term temporal information. The capacity of an LSTM network can be increased by widening and adding layers. However, usually the former introduces additional parameters, while the latter increases the runtime. As an alternative we propose the Tensorized LSTM in which the hidden states are represented by tensors and updated via a cross-layer convolution. By increasing the tensor size, the network can be widened efficiently without additional parameters since the parameters are shared across different locations in the tensor; by delaying the output, the network can be deepened implicitly with little additional runtime since deep computations for each timestep are merged into temporal computations of the sequence. Experiments conducted on five challenging sequence learning tasks show the potential of the proposed model. Shaobing Gao, Liang Xiao 0007, Daxue Liu, Hangen He, David Barber |
NIPS | 6 |
| 2016 | Approximate Newton Methods for Policy Search in Markov Decision ProcessesabstractApproximate Newton methods are standard optimization tools which aim to maintain the benefits of Newton's method, such as a fast rate of convergence, while alleviating its drawbacks, such as computationally expensive calculation or estimation of the inverse Hessian. In this work we investigate approximate Newton methods for policy optimization in Markov decision processes (MDPs). We first analyse the structure of the Hessian of the total expected reward, which is a standard objective function for MDPs. We show that, like the gradient, the Hessian exhibits useful structure in the context of MDPs and we use this analysis to motivate two Gauss-Newton methods for MDPs. Like the Gauss- Newton method for non-linear least squares, these methods drop certain terms in the Hessian. The approximate Hessians possess desirable properties, such as negative definiteness, and we demonstrate several important performance guarantees including guaranteed ascent directions, invariance to affine transformation of the parameter space and convergence guarantees. We finally provide a unifying perspective of key policy search algorithms, demonstrating that our second Gauss- Newton algorithm is closely related to both the EM-algorithm and natural gradient ascent applied to MDPs, but performs significantly better in practice on a range of challenging domains. Thomas Furmston, Guy Lever, David Barber |
J. Mach. Learn. Res. | 3 |
| 2015 | Topic factor models: Uncovering thematic structure in equity market dataabstractWe examine the task of finding thematic structure in a data corpus comprising text and time series. To achieve this we introduce topic factor modelling (TFM). We develop a novel, joint generative model for both data types which resembles supervised latent Dirichlet allocation. TFM allows the decomp osition of time series into factors which also reflect the thematic content of the text. We describe a variational method for inference and demonstrate its effectiveness on a synthetic corpus. For a corpus of publicly available equity data, we show that a TFM can simultaneously and robustly model both stock price time series and text data describing the corresponding companies. We also discuss how topic modelling could assist with external tasks such as robust covariance estimation. Joe Staines, David Barber |
Intell. Data Anal. | 2 |
| 2014 | Gaussian Processes for Bayesian Estimation in Ordinary Differential EquationsabstractBayesian parameter estimation in coupled ordinary differential equations (ODEs) is challenging due to the high computational cost of numerical integration. In gradient matching a separate data model is introduced with the property that its gradient can be calculated easily. Parameter estimation is achieved by requiring consistency between the gradients computed from the data model and those specified by the ODE. We propose a Gaussian process model that directly links state derivative information with system observations, simplifying previous approaches and providing a natural generative model. David Barber |
ICML | 1 |
| 2014 | An Image Reconstruction Algorithm for 3-D Electrical Impedance MammographyabstractThe Sussex MK4 electrical impedance mammography system is especially designed for 3-D breast screening. It aims to diagnose breast cancer at an early stage when it is most treatable. Planar electrodes are employed in this system. The challenge with planar electrodes is the inaccuracy and poor sensitivity in the vertical direction for 3-D imaging. An enhanced image reconstruction algorithm using a duo-mesh method is proposed to improve the vertical accuracy and sensitivity. The novel part of the enhanced image reconstruction algorithm is the correction term. To evaluate the new algorithm, an image processing based error analysis method is presented, which not only can precisely assess the error of the reconstructed image but also locate the center and outline the center and outline the shape of the objects of interest. Although the enhanced image reconstruction algorithm and the image processing based error analysis method are designed for the Sussex MK4 system, they are applicable to all electrical impedance tomography systems, regardless of the hardware design. To validate the enhanced algorithm, performance results from simulations, phantoms and patients are presented. Gerald Sze, David Barber, Chris R. Chatwin |
IEEE Trans. Medical Imaging | 4 |
| 2013 | Optimization by Variational Bounding
Joe Staines, David Barber |
ESANN | 2 |
| 2013 | Gaussian Kullback-Leibler approximate inference
Edward Challis, David Barber |
J. Mach. Learn. Res. | 2 |
| 2012 | Bayesian Conditional Cointegration
Chris Bracegirdle, David Barber |
ICML | 2 |
| 2012 | Affine Independent Variational InferenceabstractWe present a method for approximate inference for a broad class of non-conjugate probabilistic models. In particular, for the family of generalized linear model target densities we describe a rich class of variational approximating densities which can be best fit to the target by minimizing the Kullback-Leibler divergence. Our approach is based on using the Fourier representation which we show results in efficient and scalable inference. Edward Challis, David Barber |
NIPS | 2 |
| 2012 | A Unifying Perspective of Parametric Policy Search Methods for Markov Decision ProcessesabstractParametric policy search algorithms are one of the methods of choice for the optimisation of Markov Decision Processes, with Expectation Maximisation and natural gradient ascent being considered the current state of the art in the field. In this article we provide a unifying perspective of these two algorithms by showing that their step-directions in the parameter space are closely related to the search direction of an approximate Newton method. This analysis leads naturally to the consideration of this approximate Newton method as an alternative gradient-based method for Markov Decision Processes. We are able show that the algorithm has numerous desirable properties, absent in the naive application of Newton's method, that make it a viable alternative to either Expectation Maximisation or natural gradient ascent. Empirical results suggest that the algorithm has excellent convergence and robustness properties, performing strongly in comparison to both Expectation Maximisation and natural gradient ascent. Thomas Furmston, David Barber |
NIPS | 2 |
| 2011 | Lagrange Dual Decomposition for Finite Horizon Markov Decision Processes
Thomas Furmston, David Barber |
ECML/PKDD (1) | 2 |
| 2011 | Efficient Inference in Markov Control Problems
Thomas Furmston, David Barber |
UAI | 2 |
| 2009 | A Simple Alternative Derivation of the Expectation Correction AlgorithmabstractThe switching linear dynamical system (SLDS) is a popular model in time-series analysis. However, the complexity of inferring the state of the latent variables scales exponentially with the length of the time-series, resulting in many approximation strategies in the literature. We focus on the recently devised expectation correction (EC) approximation which can be considered a form of Gaussian sum smoother. The algorithm has excellent numerical performance compared to a wide range of competing techniques, exploiting more fully the available information than, for example, generalised pseudo Bayes. We show that EC can be seen as an extension to the SLDS of the Rauch, Tung, Striebel inference algorithm for the linear dynamical system. This yields a simpler derivation of the EC algorithm and facilitates comparison with existing, similar approaches. Bertrand Mesot, David Barber |
IEEE Signal Process. Lett. | 2 |
| 2008 | Clique Matrices for Statistical Graph Decomposition and Parameterising Restricted Positive Definite Matrices
David Barber |
UAI | 1 |
| 2007 | Stable Belief Propagation in Gaussian DagsabstractWe consider approximate inference in the important class of Gaussian distributions corresponding to multiply-connected directed acylic networks (DAGs). We show how directed belief propagation can be implemented in a numerically stable manner by associating backward (λ) messages with an auxiliary variable, enabling intermediate computations to be carried out in moment form. We apply our method to the fast Fourier transform network with missing data, and show that the results are more accurate than those obtained using undirected belief propagation on the equivalent Markov network. David Barber, Peter Sollich |
ICASSP (2) | 1 |
| 2007 | A Bayesian Alternative to Gain Adaptation in Autoregressive Hidden Markov ModelsabstractModels dealing directly with the raw acoustic speech signal are an alternative to conventional feature-based HMMs. A popular way to model the raw speech signal is by means of an autoregressive (AR) process. Being too simple to cope with the nonlinearity of the speech signal, the AR process is generally embedded into a more elaborate model, such as the switching autoregressive HMM (SAR-HMM). A fundamental issue faced by models based on AR processes is that they are very sensitive to variations in the amplitude of the signal. One way to overcome this limitation is to use gain adaptation to adjust the amplitude by maximising the likelihood of the observed signal. However, adjusting model parameters by maximising test likelihoods is fundamentally outside the framework of standard statistical approaches to machine learning, since this may lead to overfitting when the models are sufficiently flexible. We propose a statistically principled alternative based on an exact Bayesian procedure in which priors are explicitly defined on the parameters of the AR process. Explicitly, we present the Bayesian SAR-HMM and compare the performance of this model against the standard gain-adapted SAR-HMM on a single digit recognition task, showing the effectiveness of the approach and suggesting thereby a principled and straightforward solution to the issue of gain adaptation. Bertrand Mesot, David Barber |
ICASSP (2) | 2 |
| 2007 | Bayesian Factorial Linear Gaussian State-Space Models for Biosignal DecompositionabstractWe discuss a method to extract independent dynamical systems underlying a single or multiple channels of observation. In particular, we search for one-dimensional subsignals to aid the interpretability of the decomposition. The method uses an approximate Bayesian analysis to determine automatically the number and appropriate complexity of the underlying dynamics, with a preference for the simplest solution. We apply this method to unfiltered EEG signals to discover low-complexity sources with preferential spectral properties, demonstrating improved interpretability of the extracted sources over related methods Silvia Chiappa, David Barber |
IEEE Signal Process. Lett. | 2 |
| 2007 | Switching Linear Dynamical Systems for Noise Robust Speech RecognitionabstractReal world applications such as hands-free dialling in cars may have to deal with potentially very noisy environments. Existing state-of-the-art solutions to this problem use feature-based HMMs, with a preprocessing stage to clean the noisy signal. However, the effect that raw signal noise has on the induced HMM features is poorly understood, and limits the performance of the HMM system. An alternative to feature-based HMMs is to model the raw signal, which has the potential advantage that including an explicit noise model is straightforward. Here we jointly model the dynamics of both the raw speech signal and the noise, using a switching linear dynamical system (SLDS). The new model was tested on isolated digit utterances corrupted by Gaussian noise. Contrary to the autoregressive HMM and its derivatives, which provides a model of uncorrupted raw speech, the SLDS is comparatively noise robust and also significantly outperforms a state-of-the-art feature-based HMM. The computational complexity of the SLDS scales exponentially with the length of the time series. To counter this we use expectation correction which provides a stable and accurate linear-time approximation for this important class of models, aiding their further application in acoustic modeling. Bertrand Mesot, David Barber |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Efficient Kalman Smoothing for Harmonic State-Space ModelsabstractHarmonic probabilistic models are common in signal analysis. Framed as a linear-Gaussian state-space model, smoothed inference scales as O(TH2) where H is twice the number of frequencies in the model and T is the length of the time-series. Due to their central role in acoustic modelling, fast effective inference in this model is of some considerable interest. We present a form of 'rotation-corrected' low-rank approximation for the backward pass of the Rauch-Tung-Striebel smoother. This provides an effective approximation with computation complexity Q(TSH) where S is the rank of the approximation David Barber |
ICASSP (3) | 1 |
| 2006 | Unified Inference for Variational Bayesian Linear Gaussian State-Space ModelsabstractLinear Gaussian State-Space Models are widely used and a Bayesian treatment of parameters is therefore of considerable interest. The approximate Variational Bayesian method applied to these models is an attractive approach, used successfully in applications ranging from acoustics to bioinformatics. The most challenging aspect of implementing the method is in performing inference on the hidden state sequence of the model. We show how to convert the inference problem so that standard Kalman Filtering/Smoothing recursions from the literature may be applied. This is in contrast to previously published approaches based on Belief Propagation. Our framework both simplifies and unifies the inference problem, so that future applications may be more easily developed. We demonstrate the elegance of the approach on Bayesian temporal ICA, with an application to finding independent dynamical processes underlying noisy EEG signals. 1 Linear Gaussian State-Space Models Linear Gaussian State-Space Models (LGSSMs)1 are fundamental in time-series analysis [1, 2, 3]. In these models the observations v1:T 2 are generated from an underlying dynamical system on h1:T according to: v v vt = B ht + t , t N (0V , V ), h h ht = Aht-1 + t , t N (0H , H ) , where N (, ) denotes a Gaussian with mean and covariance , and 0X denotes an X dimensional zero vector. The observation vt has dimension V and the hidden state ht has dimension H . Probabilistically, the LGSSM is defined by: p(v1:T , h1:T |) = p(v1 |h1 )p(h1 ) tT p(vt |ht )p(ht |ht-1 ), =2 with p(vt |ht ) = N (B ht , V ), p(ht |ht-1 ) = N (Aht-1 , H ), p(h1 ) = N (, ) and where = {A, B , H , V , , } denotes the model parameters. Because of the widespread use of these models, a Bayesian treatment of parameters is of considerable interest [4, 5, 6, 7, 8]. An exact implementation of the Bayesian LGSSM is formally intractable [8], and recently a Variational Bayesian (VB) approximation has been studied [4, 5, 6, 7, 9]. The most challenging part of implementing the VB method is performing inference over h1:T , and previous authors have developed their own specialized routines, based on Belief Propagation, since standard LGSSM inference routines appear, at first sight, not to be applicable. 1 2 Also called Kalman Filters/Smoothers, Linear Dynamical Systems. v1:T denotes v1 , . . . , vT . A key contribution of this paper is to show how the Variational Bayesian treatment of the LGSSM can be implemented using standard LGSSM inference routines. Based on the insight we provide, any standard inference method may be applied, including those specifically addressed to improve numerical stability [2, 10, 11]. In this article, we decided to describe the predictor-corrector and Rauch-Tung-Striebel recursions [2], and also suggest a small modification that reduces computational cost. The Bayesian LGSSM is particularly of interest when strong prior constraints are needed to find adequate solutions. One such case is in EEG signal analysis, whereby we wish to extract sources that evolve independently through time. Since EEG is particularly noisy [12], a prior that encourages sources to have preferential dynamics is advantageous. This application is discussed in Section 4, and demonstrates the ease of applying our VB framework. 2 Bayesian Linear Gaussian State-Space Models In the Bayesian treatment of the LGSSM, instead of considering the model parameters as fixed, ^ ^ we define a prior distribution p(|), where is a set of hyperparameters. Then: ^ ^ p(v1:T |) = p(v1:T |)p(|) . (1) In a full Bayesian treatment we would define additional prior distributions over the hyperparameters ^ . Here we take instead the ML-II (`evidence') framework, in which the optimal set of hyperpa^ ^ rameters is found by maximizing p(v1:T |) with respect to [6, 7, 9]. For the parameter priors, here we define Gaussians on the columns of A and B 3 : p(A|, H ) jH e- j 2 David Barber, Silvia Chiappa |
NIPS | 1 |
| 2006 | A Novel Gaussian Sum Smoother for Approximate Inference in Switching Linear Dynamical SystemsabstractWe introduce a method for approximate smoothed inference in a class of switching linear dynamical systems, based on a novel form of Gaussian Sum smoother. This class includes the switching Kalman Filter and the more general case of switch transitions dependent on the continuous latent state. The method improves on the standard Kim smoothing approach by dispensing with one of the key approximations, thus making fuller use of the available future information. Whilst the only central assumption required is projection to a mixture of Gaussians, we show that an additional conditional independence assumption results in a simpler but stable and accurate alternative. Unlike the alternative unstable Expectation Propagation procedure, our method consists only of a single forward and backward pass and is reminiscent of the standard smoothing `correction' recursions in the simpler linear dynamical system. The algorithm performs well on both toy experiments and in a large scale application to noise robust speech recognition. 1 Switching Linear Dynamical System The Linear Dynamical System (LDS) [1] is a key temporal model in which a latent linear process generates the observed series. For complex time-series which are not well described globally by a single LDS, we may break the time-series into segments, each modeled by a potentially different LDS. This is the basis for the Switching LDS (SLDS) [2, 3, 4, 5] where, for each time t, a switch variable st 1, . . . , S describes which of the LDSs is to be used. The observation (or `visible') vt RV is linearly related to the hidden state ht RH with additive noise by vt = B (st )ht + v (st ) p(vt |ht , st ) = N (B (st )ht , v (st )) (1 ) where N (, ) denotes a Gaussian distribution with mean and covariance . The transition dynamics of the continuous hidden state ht is linear, A ( ht = A(st )ht-1 + h (st ), p(ht |ht-1 , st ) = N (st )ht-1 , h (st ) 2) The switch st may depend on both the previous st-1 and ht-1 . This is an augmented SLDS (aSLDS), and defines the model p(v1:T , h1:T , s1:T ) = tT p(vt |ht , st )p(ht |ht-1 , st )p(st |ht-1 , st-1 ) =1 The standard SLDS[4] considers only switch transitions p(st |st-1 ). At time t = 1, p(s1 |h0 , s0 ) simply denotes the prior p(s1 ), and p(h1 |h0 , s1 ) denotes p(h1 |s1 ). The aim of this article is to address how to perform inference in the aSLDS. In particular we desire the filtered estimate p(ht , st |v1:t ) and the smoothed estimate p(ht , st |v1:T ), for any 1 t T . Both filtered and smoothed inference in the SLDS is intractable, scaling exponentially with time [4]. David Barber, Bertrand Mesot |
NIPS | 1 |
| 2006 | EEG classification using generative independent component analysis
Silvia Chiappa, David Barber |
Neurocomputing | 2 |
| 2006 | Expectation Correction for Smoothed Inference in Switching Linear Dynamical SystemsabstractWe introduce a method for approximate smoothed inference in a class of switching linear dynamical systems, based on a novel form of Gaussian Sum smoother. This class includes the switching Kalman 'Filter' and the more general case of switch transitions dependent on the continuous latent state. The method improves on the standard Kim smoothing approach by dispensing with one of the key approximations, thus making fuller use of the available future information. Whilst the central assumption required is projection to a mixture of Gaussians, we show that an additional conditional independence assumption results in a simpler but accurate alternative. Our method consists of a single Forward and Backward Pass and is reminiscent of the standard smoothing 'correction' recursions in the simpler linear dynamical system. The method is numerically stable and compares favourably against alternative approximations, both in cases where a single mixture component provides a good posterior approximation, and where a multimodal approximation is required. David Barber |
J. Mach. Learn. Res. | 1 |
| 2006 | Optimal Spike-Timing-Dependent Plasticity for Precise Action Potential Firing in Supervised LearningabstractIn timing-based neural codes, neurons have to emit action potentials at precise moments in time. We use a supervised learning paradigm to derive a synaptic update rule that optimizes by gradient ascent the likelihood of postsynaptic firing at one or several desired firing times. We find that the optimal strategy of up- and downregulating synaptic efficacies depends on the relative timing between presynaptic spike arrival and desired postsynaptic firing. If the presynaptic spike arrives before the desired postsynaptic spike timing, our optimal learning rule predicts that the synapse should become potentiated. The dependence of the potentiation on spike timing directly reflects the time course of an excitatory postsynaptic potential. However, our approach gives no unique reason for synaptic depression under reversed spike timing. In fact, the presence and amplitude of depression of synaptic efficacies for reversed spike timing depend on how constraints are implemented in the optimization problem. Two different constraints, control of postsynaptic rates and control of temporal locality, are studied. The relation of our results to spike-timing-dependent plasticity and reinforcement learning is discussed. Jean-Pascal Pfister, Taro Toyoizumi, David Barber, Wulfram Gerstner |
Neural Comput. | 3 |
| 2006 | A generative model for music transcriptionabstractIn this paper, we present a graphical model for polyphonic music transcription. Our model, formulated as a dynamical Bayesian network, embodies a transparent and computationally tractable approach to this acoustic analysis problem. An advantage of our approach is that it places emphasis on explicitly modeling the sound generation procedure. It provides a clear framework in which both high level (cognitive) prior information on music structure can be coupled with low level (acoustic physical) information in a principled manner to perform the analysis. The model is a special case of the, generally intractable, switching Kalman filter model. Where possible, we derive, exact polynomial time inference procedures, and otherwise efficient approximations. We argue that our generative model based approach is computationally feasible for many music applications and is readily extensible to more general auditory scene analysis scenarios. A. Taylan Cemgil, Hilbert J. Kappen, David Barber |
IEEE Trans. Speech Audio Process. | 3 |
| 2005 | generative independent component analysis for EEG classification
Silvia Chiappa, David Barber |
ESANN | 2 |
| 2005 | A graphical model for chord progressions embedded in a psychoacoustic spaceabstractChord progressions are the building blocks from which tonal music is constructed. Inferring chord progressions is thus an essential step towards modeling long term dependencies in music. In this paper, a distributed representation for chords is designed such that Euclidean distances roughly correspond to psychoacoustic dissimilarities. Parameters in the graphical models are learnt with the EM algorithm and the classical Junction Tree algorithm. Various model architectures are compared in terms of conditional out-of-sample likelihood. Both perceptual and statistical evidence show that binary trees related to meter are well suited to capture chord dependencies. Jean-François Paiement, Douglas Eck, Samy Bengio, David Barber |
ICML | 4 |
| 2005 | Kernelized Infomax ClusteringabstractWe propose a simple information-theoretic approach to soft clus- tering based on maximizing the mutual information I(x, y) between the unknown cluster labels y and the training patterns x with re- spect to parameters of specifically constrained encoding distribu- tions. The constraints are chosen such that patterns are likely to be clustered similarly if they lie close to specific unknown vectors in the feature space. The method may be conveniently applied to learning the optimal affinity matrix, which corresponds to learn- ing parameters of the kernelized encoder. The procedure does not require computations of eigenvalues of the Gram matrices, which makes it potentially attractive for clustering large data sets. Felix V. Agakov, David Barber |
NIPS | 2 |
| 2004 | Variational Information Maximization for Neural Coding
Felix V. Agakov, David Barber |
ICONIP | 2 |
| 2004 | An Auxiliary Variational Method
Felix V. Agakov, David Barber |
ICONIP | 2 |
| 2003 | Approximate Learning in Temporal Hidden Hopfield Models
Felix V. Agakov, David Barber |
ICANN | 2 |
| 2003 | Optimal Hebbian Learning: A Probabilistic Point of View
Jean-Pascal Pfister, David Barber, Wulfram Gerstner |
ICANN | 2 |
| 2003 | Information Maximization in Noisy Channels : A Variational ApproachabstractThe maximisation of information transmission over noisy channels is a common, albeit generally computationally difficult problem. We approach the difficulty of computing the mutual information for noisy channels by using a variational approximation. The re- sulting IM algorithm is analagous to the EM algorithm, yet max- imises mutual information, as opposed to likelihood. We apply the method to several practical examples, including linear compression, population encoding and CDMA. David Barber, Felix V. Agakov |
NIPS | 1 |
| 2002 | Learning in Spiking Neural AssembliesabstractWe consider a statistical framework for learning in a class of net- works of spiking neurons. Our aim is to show how optimal local learning rules can be readily derived once the neural dynamics and desired functionality of the neural assembly have been specifled, in contrast to other models which assume (sub-optimal) learning rules. Within this framework we derive local rules for learning tem- poral sequences in a model of spiking neurons and demonstrate its superior performance to correlation (Hebbian) based approaches. We further show how to include mechanisms such as synaptic de- pression and outline how the framework is readily extensible to learning in networks of highly complex spiking neurons. A stochas- tic quantal vesicle release mechanism is considered and implications on the complexity of learning discussed. David Barber |
NIPS | 1 |
| 2002 | Dynamic Bayesian Networks with Deterministic Latent TablesabstractThe application of latent/hidden variable Dynamic Bayesian Net- works is constrained by the complexity of marginalising over latent variables. For this reason either small latent dimensions or Gaus- sian latent conditional tables linearly dependent on past states are typically considered in order that inference is tractable. We suggest an alternative approach in which the latent variables are modelled using deterministic conditional probability tables. This specialisa- tion has the advantage of tractable inference even for highly com- plex non-linear/non-Gaussian visible conditional probability tables. This approach enables the consideration of highly complex latent dynamics whilst retaining the bene(cid:12)ts of a tractable probabilistic model. David Barber |
NIPS | 1 |
| 2001 | Deterministic Generative Models for Fast Feature Discovery
Machiel Westerdijk, David Barber, Wim Wiegerinck |
Data Min. Knowl. Discov. | 2 |
| 1999 | Gaussian Fields for Approximate Inference in Layered Sigmoid Belief Networks
David Barber, Peter Sollich |
NIPS | 1 |
| 1999 | Variational Cumulant Expansions for Intractable DistributionsabstractIntractable distributions present a common difficulty in inference within the probabilistic knowledge representation framework and variational methods have recently been popular in providing an approximate solution. In this article, we describe a perturbational approach in the form of a cumulant expansion which, to lowest order, recovers the standard Kullback-Leibler variational bound. Higher-order terms describe corrections on the variational approach without incurring much further computational cost. The relationship to other perturbational approaches such as TAP is also elucidated. We demonstrate the method on a particular class of undirected graphical models, Boltzmann machines, for which our simulation results confirm improved accuracy and enhanced stability during learning. David Barber, Piërre van de Laar |
J. Artif. Intell. Res. | 1 |
| 1998 | Tractable Variational Structures for Approximating Graphical Models
David Barber, Wim Wiegerinck |
NIPS | 1 |
| 1998 | Online Learning from Finite Training Sets and Robustness to Input BiasabstractWe analyze online gradient descent learning from finite training sets at noninfinitesimal learning rates eta. Exact results are obtained for the time-dependent generalization error of a simple model system: a linear network with a large number of weights N, trained on p = alphaN examples. This allows us to study in detail the effects of finite training set size alpha on, for example, the optimal choice of learning rate eta. We also compare online and offline learning, for respective optimal settings of eta at given final learning time. Online learning turns out to be much more robust to input bias and actually outperforms offline learning when such bias is present; for unbiased inputs, online and offline learning perform almost equally well. Peter Sollich, David Barber |
Neural Comput. | 2 |
| 1998 | Bayesian Classification With Gaussian ProcessesabstractWe consider the problem of assigning an input vector to one of m classes by predicting P(c|x) for c=1,...,m. For a two-class problem, the probability of class one given x is estimated by /spl sigma/(y(x)), where /spl sigma/(y)=1/(1+e/sup -y/). A Gaussian process prior is placed on y(x), and is combined with the training data to obtain predictions for new x points. We provide a Bayesian treatment, integrating over uncertainty in y and in the parameters that control the Gaussian process prior the necessary integration over y is carried out using Laplace's approximation. The method is generalized to multiclass problems (m>2) using the softmax function. We demonstrate the effectiveness of the method on a number of datasets. Christopher K. I. Williams, David Barber |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | Ensemble Learning for Multi-Layer Networks
David Barber, Christopher M. Bishop |
NIPS | 1 |
| 1997 | Radial Basis Functions: A Bayesian Treatment
David Barber, Bernhard Schottky |
NIPS | 1 |
| 1997 | On-line Learning from Finite Training Sets in Nonlinear Networks
Peter Sollich, David Barber |
NIPS | 2 |
| 1996 | Bayesian Model Comparison by Monte Carlo Chaining
David Barber, Christopher M. Bishop |
NIPS | 1 |
| 1996 | Gaussian Processes for Bayesian Classification via Hybrid Monte Carlo
David Barber, Christopher K. I. Williams |
NIPS | 1 |
| 1996 | Online Learning from Finite Training Sets: An Analytical Case Study
Peter Sollich, David Barber |
NIPS | 2 |
| 1996 | Does Extra Knowledge Necessarily Improve Generalization?abstractThe generalization error is a widely used performance measure employed in the analysis of adaptive learning systems. This measure is generally critically dependent on the knowledge that the system is given about the problem it is trying to learn. In this paper we examine to what extent it is necessarily the case that an increase in the knowledge that the system has about the problem will reduce the generalization error. Using the standard definition of the generalization error, we present simple cases for which the intuitive idea of “reducivity”—that more knowledge will improve generalization—does not hold. Under a simple approximation, however, we find conditions to satisfy “reducivity.” Finally, we calculate the effect of a specific constraint on the generalization error of the linear perceptron, in which the signs of the weight components are fixed. This particular restriction results in a significant improvement in generalization performance. David Barber, David Saad |
Neural Comput. | 1 |
| 1995 | Knowledge and generalisation in simple learning systems
David Barber, David Saad |
ESANN | 1 |
| 1995 | Test Error Fluctuations in Finite Linear PerceptronsabstractWe examine the fluctuations in the test error induced by random, finite, training and test sets for the linear perceptron of input dimension n with a spherically constrained weight vector. This variance enables us to address such issues as the partitioning of a data set into a test and training set. We find that the optimal assignment of the test set size scales with n2/3. David Barber, David Saad, Peter Sollich |
Neural Comput. | 1 |