Christopher M. Bishop

dblp:b/ChristopherMBishop · DBLP profile ↗
← Back
42ranked-venue papers
26as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 26 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
17 papers
Probabilistic and Bayesian machine learning · 70% Image recognition and object detection · 19% Learning paradigms · 5%
Computer graphics and multimedia
3 papers
Image and video processing · 74% Visualization and visual analytics · 26%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 72% Embedded and real-time systems · 14% Hardware accelerators and domain-specific architectures · 14%

Topics — the 30 heaviest of 46, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning
generative and discriminative models
0.122006
Principled Hybrids of Generative and Discriminative Models · CVPR (1) 2006
Generative versus Discriminative Methods for Object Recognition · CVPR (2) 2005
Computer vision › Image recognition and object detection
object recognition
0.122006
Principled Hybrids of Generative and Discriminative Models · CVPR (1) 2006
Generative versus Discriminative Methods for Object Recognition · CVPR (2) 2005
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian network
0.132005
Variational Message Passing · J. Mach. Learn. Res. 2005
VIBES: A Variational Inference Engine for Bayesian Networks · NIPS 2002
Approximating Posterior Distributions in Belief Networks Using Mixtures · NIPS 1997
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.122005
Variational Message Passing · J. Mach. Learn. Res. 2005
VIBES: A Variational Inference Engine for Bayesian Networks · NIPS 2002
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.122005
Variational Message Passing · J. Mach. Learn. Res. 2005
VIBES: A Variational Inference Engine for Bayesian Networks · NIPS 2002
Machine learning › Learning paradigms
semi-supervised learning
0.112006
Principled Hybrids of Generative and Discriminative Models · CVPR (1) 2006
Computer vision › Image recognition and object detection
object detection
0.112005
Generative versus Discriminative Methods for Object Recognition · CVPR (2) 2005
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
variational message passing
0.112005
Variational Message Passing · J. Mach. Learn. Res. 2005
Computer vision › Image recognition and object detection › object detection
weakly supervised object detection
0.112005
Generative versus Discriminative Methods for Object Recognition · CVPR (2) 2005
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.122002
VIBES: A Variational Inference Engine for Bayesian Networks · NIPS 2002
Approximating Posterior Distributions in Belief Networks Using Mixtures · NIPS 1997
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.031998
A Hierarchical Latent Variable Model for Data Visualization · IEEE Trans. Pattern Anal. Mach. Intell. 1998
GTM: A Principled Alternative to the Self-Organizing Map · NIPS 1996
EM Optimization of Latent-Variables Density Models · NIPS 1995
Image and video processing › super-resolution
bayesian super-resolution
0.012002
Bayesian Image Super-Resolution · NIPS 2002
Image and video processing
image registration
0.012002
Bayesian Image Super-Resolution · NIPS 2002
Image and video processing
super-resolution
0.012002
Bayesian Image Super-Resolution · NIPS 2002
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.021997
Approximating Posterior Distributions in Belief Networks Using Mixtures · NIPS 1997
Regression with Input-Dependent Noise: A Bayesian Treatment · NIPS 1996
Distributed systems › replication › replica consistency
replica synchronization
0.012001
Optimising Synchronisation Times for Mobile Devices · NIPS 2001
Distributed systems
replication
0.012001
Optimising Synchronisation Times for Mobile Devices · NIPS 2001
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.022002
Regression with Input-dependent Noise: A Gaussian Process Treatment · NIPS 1997
Bayesian Image Super-Resolution · NIPS 2002
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian image modeling
0.012000
Non-linear Bayesian Image Modelling · ECCV (1) 2000
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › latent gaussian model
Bayesian PCA
0.011998
Bayesian PCA · NIPS 1998
Visualization and visual analytics
data visualization
0.011998
A Hierarchical Latent Variable Model for Data Visualization · IEEE Trans. Pattern Anal. Mach. Intell. 1998
Visualization and visual analytics
dimensionality reduction
0.011998
A Hierarchical Latent Variable Model for Data Visualization · IEEE Trans. Pattern Anal. Mach. Intell. 1998
Algorithms and data structures › numerical linear algebra
dimensionality reduction
0.011998
Bayesian PCA · NIPS 1998
Algorithms and data structures › numerical linear algebra › dimensionality reduction
principal component analysis
0.011998
Bayesian PCA · NIPS 1998
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization
0.021998
EM Optimization of Latent-Variables Density Models · NIPS 1995
A Hierarchical Latent Variable Model for Data Visualization · IEEE Trans. Pattern Anal. Mach. Intell. 1998
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference
0.011997
Approximating Posterior Distributions in Belief Networks Using Mixtures · NIPS 1997
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.011997
Ensemble Learning for Multi-Layer Networks · NIPS 1997
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer neural network
0.011997
Ensemble Learning for Multi-Layer Networks · NIPS 1997
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian model selection
0.011996
Bayesian Model Comparison by Monte Carlo Chaining · NIPS 1996
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › probabilistic regression
bayesian regression
0.011996
Regression with Input-Dependent Noise: A Bayesian Treatment · NIPS 1996

Methods — techniques the papers use, named apart from their topics

bayesian inference · 0.1variational inference · 0.1expectation-maximization · 0.1probabilistic modeling · 0.1prior over parameters · 0.1objective function minimization · 0.1convex combination of objective functions · 0.1generative modeling · 0.1factorized variational approximation · 0.1discriminative modeling · 0.1belief propagation · 0.1message passing · 0.0marginalization · 0.0gaussian process prior · 0.0bayesian modeling · 0.0real-time feedback control · 0.0multilayer perceptron · 0.0hierarchical mixture model · 0.0
YearPublicationVenuePosition
2014 Students, Teachers, Exams and MOOCs: Predicting and Optimizing Attainment in Web-Based Education Using a Probabilistic Graphical Model
Bar Shalem, Yoram Bachrach, John Guiver, Christopher M. Bishop
ECML/PKDD (3)4
2013 Structural Expectation Propagation (SEP): Bayesian structure learning for networks with latent variables
abstract
Learning the structure of discrete Bayesian networks has been the subject of extensive research in machine learning, with most Bayesian approaches focusing on fully observed networks. One of few the methods that can handle networks with latent variables is the "structural EM algorithm" which interleaves greedy structure search with the estimation of latent variables and parameters, maintaining a single best network at each step. We introduce Structural Expectation Propagation (SEP), an extension of EP which can infer the structure of Bayesian networks having latent variables and missing data. SEP performs variational inference in a joint model of structure, latent variables, and parameters, offering two advantages: (i) it accounts for uncertainty in structure and parameter values when making local distribution updates (ii) it returns a variational distribution over network structures rather than a single network. We demonstrate the performance of SEP both on synthetic problems and on real-world clinical data.
Nevena Lazic, Christopher M. Bishop, John M. Winn
AISTATS2
2011 Embracing Uncertainty: Applied Machine Learning Comes of Age
Christopher M. Bishop
ECML/PKDD (1)1
2010 Embracing Uncertainty: The New Machine Intelligence
Christopher M. Bishop
KES (1)1
2006 Principled Hybrids of Generative and Discriminative Models
abstract
When labelled training data is plentiful, discriminative techniques are widely used since they give excellent generalization performance. However, for large-scale applications such as object recognition, hand labelling of data is expensive, and there is much interest in semi-supervised techniques based on generative models in which the majority of the training data is unlabelled. Although the generalization performance of generative models can often be improved by ‘training them discriminatively’, they can then no longer make use of unlabelled data. In an attempt to gain the benefit of both generative and discriminative approaches, heuristic procedure have been proposed [2, 3] which interpolate between these two extremes by taking a convex combination of the generative and discriminative objective functions. In this paper we adopt a new perspective which says that there is only one correct way to train a given model, and that a ‘discriminatively trained’ generative model is fundamentally a new model [7]. From this viewpoint, generative and discriminative models correspond to specific choices for the prior over parameters. As well as giving a principled interpretation of ‘discriminative training’, this approach opens door to very general ways of interpolating between generative and discriminative extremes through alternative choices of prior. We illustrate this framework using both synthetic data and a practical example in the domain of multi-class object recognition. Our results show that, when the supply of labelled training data is limited, the optimum performance corresponds to a balance between the purely generative and the purely discriminative.
Julia A. Lasserre, Christopher M. Bishop, Tom Minka
CVPR (1)2
2005 Generative versus Discriminative Methods for Object Recognition
abstract
Many approaches to object recognition are founded on probability theory, and can be broadly characterized as either generative or discriminative according to whether or not the distribution of the image features is modelled. Generative and discriminative methods have very different characteristics, as well as complementary strengths and weaknesses. In this paper we introduce new generative and discriminative models for object detection and classification based on weakly labelled training data. We use these models to illustrate the relative merits of the two approaches in the context of a data set of widely varying images of non-rigid objects (animals). Our results support the assertion that neither approach alone will be sufficient for large scale object recognition, and we discuss techniques for combining them.
Ilkay Ulusoy, Christopher M. Bishop
CVPR (2)2
2005 Robust Bayesian mixture modelling
Markus Svensén, Christopher M. Bishop
Neurocomputing2
2005 Variational Message Passing
abstract
Bayesian inference is now widely established as one of the principal foundations for machine learning. In practice, exact inference is rarely possible, and so a variety of approximation techniques have been developed, one of the most widely used being a deterministic framework called variational inference. In this paper we introduce Variational Message Passing (VMP), a general purpose algorithm for applying variational inference to Bayesian Networks. Like belief propagation, VMP proceeds by sending messages between nodes in the network and updating posterior beliefs using local operations at each node. Each such update increases a lower bound on the log evidence (unless already at a local maximum). In contrast to belief propagation, VMP can be applied to a very general class of conjugate-exponential models because it uses a factorised variational approximation. Furthermore, by introducing additional variational parameters, VMP can be applied to models containing non-conjugate distributions. The VMP framework also allows the lower bound to be evaluated, and this can be used both for model comparison and for detection of convergence. Variational message passing has been implemented in the form of a general purpose inference engine called VIBES ('Variational Inference for BayEsian networkS') which allows models to be specified graphically and then solved variationally without recourse to coding.
John M. Winn, Christopher M. Bishop
J. Mach. Learn. Res.2
2004 Robust Bayesian Mixture Modelling
Christopher M. Bishop, Markus Svensén
ESANN1
2003 Bayesian Hierarchical Mixtures of Experts
Christopher M. Bishop, Markus Svensén
UAI1
2002 VIBES: A Variational Inference Engine for Bayesian Networks
abstract
In recent years variational methods have become a popular tool for approximate inference and learning in a wide variety of proba- bilistic models. For each new application, however, it is currently necessary (cid:12)rst to derive the variational update equations, and then to implement them in application-speci(cid:12)c code. Each of these steps is both time consuming and error prone. In this paper we describe a general purpose inference engine called VIBES (‘Variational Infer- ence for Bayesian Networks’) which allows a wide variety of proba- bilistic models to be implemented and solved variationally without recourse to coding. New models are speci(cid:12)ed either through a simple script or via a graphical interface analogous to a drawing package. VIBES then automatically generates and solves the vari- ational equations. We illustrate the power and (cid:13)exibility of VIBES using examples from Bayesian mixture modelling.
Christopher M. Bishop, David J. Spiegelhalter, John M. Winn
NIPS1
2002 Bayesian Image Super-Resolution
abstract
The extraction of a single high-quality image from a set of low(cid:173) resolution images is an important problem which arises in fields such as remote sensing, surveillance, medical imaging and the ex(cid:173) traction of still images from video. Typical approaches are based on the use of cross-correlation to register the images followed by the inversion of the transformation from the unknown high reso(cid:173) lution image to the observed low resolution images, using regular(cid:173) ization to resolve the ill-posed nature of the inversion process. In this paper we develop a Bayesian treatment of the super-resolution problem in which the likelihood function for the image registra(cid:173) tion parameters is based on a marginalization over the unknown high-resolution image. This approach allows us to estimate the unknown point spread function, and is rendered tractable through the introduction of a Gaussian process prior over images. Results indicate a significant improvement over techniques based on MAP (maximum a-posteriori) point optimization of the high resolution image and associated registration parameters.
Michael E. Tipping, Christopher M. Bishop
NIPS2
2001 Probabilistic Modelling of Replica Divergence
abstract
It is common in distributed systems to replicate data. In many cases this data evolves in a consistent fashion, and this evolution can be modelled. A probabilistic model of the evolution allows us to estimate the divergence of the replicas and can be used by the application to alter its behaviour, for example to control synchronisation times, to determine the propagation of writes, and to convey to the user information about how much the data may have evolved. In this paper, we describe how the evolution of the data may be modelled and outline how the probabilistic model may be utilised in various applications, concentrating on a news database example.
Antony I. T. Rowstron, Neil D. Lawrence, Christopher M. Bishop
HotOS3
2001 Optimising Synchronisation Times for Mobile Devices
abstract
With the increasing number of users of mobile computing devices (e.g. personal digital assistants) and the advent of third generation mobile phones, wireless communications are becoming increasingly important. Many applications rely on the device maintaining a replica of a data-structure which is stored on a server, for exam(cid:173) ple news databases, calendars and e-mail. ill this paper we explore the question of the optimal strategy for synchronising such replicas. We utilise probabilistic models to represent how the data-structures evolve and to model user behaviour. We then formulate objective functions which can be minimised with respect to the synchronisa(cid:173) tion timings. We demonstrate, using two real world data-sets, that a user can obtain more up-to-date information using our approach.
Neil D. Lawrence, Antony I. T. Rowstron, Christopher M. Bishop, M. J. Taylor
NIPS3
2001 Feature representation and signal classification in fluorescence in-situ hybridization image analysis
abstract
Fast and accurate analysis of fluorescence in-situ hybridization images for signal counting will depend mainly upon two components: a classifier to discriminate between artifacts and valid signals of several fluorophores (colors), and well discriminating features to represent the signals. Our previous work (2001) has focused on the first component. To investigate the second component, we evaluate candidate feature sets by illustrating the probability density functions and scatter plots for the features. The analysis provides first insight into dependencies between features, indicates the relative importance of members of a feature set, and helps in identifying sources of potential classification errors. Class separability yielded by different feature subsets is evaluated using the accuracy of several neural network-based classification strategies, some of them hierarchical, as well as using a feature selection technique making use of a scatter criterion. Although applied to cytogenetics, the paper presents a comprehensive, unifying methodology of qualitative and quantitative evaluation of pattern feature representation essential for accurate image classification. This methodology is applicable to many other real-world pattern recognition problems.
Boaz Lerner, W. F. Clocksin, Seema Dhanjal, Maj A. Hultén, Christopher M. Bishop
IEEE Trans. Syst. Man Cybern. Part A5
2000 Non-linear Bayesian Image Modelling
Christopher M. Bishop, John M. Winn
ECCV (1)1
2000 Variational Relevance Vector Machines
Christopher M. Bishop, Michael E. Tipping
UAI1
1999 Neural Network-Based Wind Vector Retrieval from Satellite Scatterometer Data
Dan Cornford, Ian T. Nabney, Christopher M. Bishop
Neural Comput. Appl.3
1999 Mixtures of Probabilistic Principal Component Analysers
abstract
Principal component analysis (PCA) is one of the most popular techniques for processing, compressing, and visualizing data, although its effectiveness is limited by its global linearity. While nonlinear variants of PCA have been proposed, an alternative paradigm is to capture data complexity by a combination of local linear PCA projections. However, conventional PCA does not correspond to a probability density, and so there is no unique way to combine PCA models. Therefore, previous attempts to formulate mixture models for PCA have been ad hoc to some extent. In this article, PCA is formulated within a maximum likelihood framework, based on a specific form of gaussian latent variable model. This leads to a well-defined mixture model for probabilistic principal component analyzers, whose parameters can be determined using an expectation-maximization algorithm. We discuss the advantages of this model in the context of clustering, density modeling, and local dimensionality reduction, and we demonstrate its application to image compression and handwritten digit recognition.
Michael E. Tipping, Christopher M. Bishop
Neural Comput.2
1998 Bayesian PCA
Christopher M. Bishop
NIPS1
1998 Mixture Representations for Inference and Learning in Boltzmann Machines
Neil D. Lawrence, Christopher M. Bishop, Michael I. Jordan
UAI2
1998 Developments of the generative topographic mapping
Christopher M. Bishop, Markus Svensén, Christopher K. I. Williams
Neurocomputing1
1998 GTM: The Generative Topographic Mapping
abstract
Latent variable models represent the probability density of data in a space of several dimensions in terms of a smaller number of latent, or hidden, variables. A familiar example is factor analysis, which is based on a linear transformation between the latent space and the data space. In this article, we introduce a form of nonlinear latent variable model called the generative topographic mapping, for which the parameters of the model can be determined using the expectation-maximization algorithm. GTM provides a principled alternative to the widely used self-organizing map (SOM) of Kohonen (1982) and overcomes most of the significant limitations of the SOM. We demonstrate the performance of the GTM algorithm on a toy problem and on simulated data from flow diagnostics for a multiphase oil pipeline.
Christopher M. Bishop, Markus Svensén, Christopher K. I. Williams
Neural Comput.1
1998 A Hierarchical Latent Variable Model for Data Visualization
abstract
Visualization has proven to be a powerful and widely-applicable tool for the analysis and interpretation of multivariate data. Most visualization algorithms aim to find a projection from the data space down to a two-dimensional visualization space. However, for complex data sets living in a high-dimensional space, it is unlikely that a single two-dimensional projection can reveal all of the interesting structure. We therefore introduce a hierarchical visualization algorithm which allows the complete data set to be visualized at the top level, with clusters and subclusters of data points visualized at deeper levels. The algorithm is based on a hierarchical mixture of latent variable models, whose parameters are estimated using the expectation-maximization algorithm. We demonstrate the principle of the approach on a toy data set, and we then apply the algorithm to the visualization of a synthetic data set in 12 dimensions obtained from a simulation of multiphase flows in oil pipelines, and to data in 36 dimensions derived from satellite images.
Christopher M. Bishop, Michael E. Tipping
IEEE Trans. Pattern Anal. Mach. Intell.1
1997 Ensemble Learning for Multi-Layer Networks
David Barber, Christopher M. Bishop
NIPS2
1997 Approximating Posterior Distributions in Belief Networks Using Mixtures
Christopher M. Bishop, Neil D. Lawrence, Tommi S. Jaakkola, Michael I. Jordan
NIPS1
1997 Regression with Input-dependent Noise: A Gaussian Process Treatment
Paul W. Goldberg, Christopher K. I. Williams, Christopher M. Bishop
NIPS3
1996 Bayesian Inference of Noise Levels in Regression
Christopher M. Bishop, Cazhaow S. Quazaz
ICANN1
1996 GTM: A Principled Alternative to the Self-Organizing Map
Christopher M. Bishop, Markus Svensén, Christopher K. I. Williams
ICANN1
1996 Bayesian Model Comparison by Monte Carlo Chaining
David Barber, Christopher M. Bishop
NIPS2
1996 Regression with Input-Dependent Noise: A Bayesian Treatment
Christopher M. Bishop, Cazhaow S. Quazaz
NIPS1
1996 GTM: A Principled Alternative to the Self-Organizing Map
Christopher M. Bishop, Markus Svensén, Christopher K. I. Williams
NIPS1
1996 Modeling Conditional Probability Distributions for Periodic Variables
abstract
Most conventional techniques for estimating conditional probability densities are inappropriate for applications involving periodic variables. In this paper we introduce three related techniques for tackling such problems, and investigate their performance using synthetic data. We then apply these techniques to the problem of extracting the distribution of wind vector directions from radar scatterometer data gathered by a remote-sensing satellite.
Christopher M. Bishop, Ian T. Nabney
Neural Comput.1
1995 EM Optimization of Latent-Variables Density Models
Christopher M. Bishop, Markus Svensén, Christopher K. I. Williams
NIPS1
1995 Training with Noise is Equivalent to Tikhonov Regularization
abstract
It is well known that the addition of noise to the input data of a neural network during training can, in some circumstances, lead to significant improvements in generalization performance. Previous work has shown that such training with noise is equivalent to a form of regularization in which an extra term is added to the error function. However, the regularization term, which involves second derivatives of the error function, is not bounded below, and so can lead to difficulties if used directly in a learning algorithm based on error minimization. In this paper we show that for the purposes of network training, the regularization term can be reduced to a positive semi-definite form that involves only first derivatives of the network mapping. For a sum-of-squares error function, the regularization term belongs to the class of generalized Tikhonov regularizers. Direct minimization of the regularized error function provides a practical alternative to training with noise.
Christopher M. Bishop
Neural Comput.1
1995 Real-Time Control of a Tokamak Plasma Using Neural Networks
abstract
In this paper we present results from the first use of neural networks for real-time control of the high-temperature plasma in a tokamak fusion experiment. The tokamak is currently the principal experimental device for research into the magnetic confinement approach to controlled fusion. In an effort to improve the energy confinement properties of the high-temperature plasma inside tokamaks, recent experiments have focused on the use of noncircular cross-sectional plasma shapes. However, the accurate generation of such plasmas represents a demanding problem involving simultaneous control of several parameters on a time scale as short as a few tens of microseconds. Application of neural networks to this problem requires fast hardware, for which we have developed a fully parallel custom implementation of a multilayer perceptron, based on a hybrid of digital and analogue techniques.
Christopher M. Bishop, Paul S. Haynes, Mike E. U. Smith, Tom N. Todd, David L. Trotman
Neural Comput.1
1994 Real-Time Control of a Tokamak Plasma Using Neural Networks
abstract
This paper presents results from the first use of neural networks for the real-time feedback control of high temperature plasmas in a tokamak fusion experiment. The tokamak is currently the prin(cid:173) cipal experimental device for research into the magnetic confine(cid:173) ment approach to controlled fusion. In the tokamak, hydrogen plasmas, at temperatures of up to 100 Million K, are confined by strong magnetic fields. Accurate control of the position and shape of the plasma boundary requires real-time feedback control of the magnetic field structure on a time-scale of a few tens of mi(cid:173) croseconds. Software simulations have demonstrated that a neural network approach can give significantly better performance than the linear technique currently used on most tokamak experiments. The practical application of the neural network approach requires high-speed hardware, for which a fully parallel implementation of the multilayer perceptron, using a hybrid of digital and analogue technology, has been developed.
Christopher M. Bishop
NIPS1
1994 Estimating Conditional Probability Densities for Periodic Variables
abstract
Most of the common techniques for estimating conditional prob(cid:173) ability densities are inappropriate for applications involving peri(cid:173) odic variables. In this paper we introduce three novel techniques for tackling such problems, and investigate their performance us(cid:173) ing synthetic data. We then apply these techniques to the problem of extracting the distribution of wind vector directions from radar scatterometer data gathered by a remote-sensing satellite.
Christopher M. Bishop, Claire Legleye
NIPS1
1994 Fast Feedback Control of a High Temperature Fusion Plasma
Christopher M. Bishop, Paul S. Haynes, Mike E. U. Smith, Tom N. Todd, David L. Trotman
Neural Comput. Appl.1
1993 Reconstruction of Tokamak Density Profiles Using Feedforward Networks
Christopher M. Bishop, Iain Strachan, John O'Rourke, Geoff Maddison
Neural Comput. Appl.1
1993 Curvature-driven smoothing: a learning algorithm for feedforward networks
abstract
The performance of feedforward neural networks in real applications can often be improved significantly if use is made of a priori information. For interpolation problems this prior knowledge frequently includes smoothness requirements on the network mapping, and can be imposed by the addition to the error function of suitable regularization terms. The new error function, however, now depends on the derivatives of the network mapping, and so the standard backpropagation algorithm cannot be applied. In this letter, we derive a computationally efficient learning algorithm, for a feedforward network of arbitrary topology, which can be used to minimize such error functions. Networks having a single hidden layer, for which the learning algorithm simplifies, are treated as a special case.
Christopher M. Bishop
IEEE Trans. Neural Networks1
1991 A Fast Procedure for Retraining the Multilayer Perceptron
abstract
In this paper we describe a fast procedure for retraining a feedforward network, previously trained by error backpropagation, following a small change in the training data. This technique would permit fine calibration of individual neural network based control systems in a mass-production environment. We also derive a generalised error backpropagation algorithm which allows an exact evaluation of all of the terms in the Hessian matrix. The fast retraining procedure is illustrated using a simple example.
Christopher M. Bishop
Int. J. Neural Syst.1