VLDB 2026 Research / reviewers in the wild / expert
Mark A. Girolami
dblp:g/MarkAGirolami · also Mark Girolami
· DBLP profile ↗
105ranked-venue papers
25as first author
14since 2021 · last 2026
0000-0003-3008-253XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 19 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 1 since 2021Computer networks · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Embedding of Graphs With Missing Data by Soft ManifoldsabstractEmbedding graphs in continuous spaces is a key factor for automatic information extraction in diverse tasks (e.g., learning, inferring, predicting). The reliability of graph embeddings directly depends on how much the geometry of the manifold in continuous space matches the graph structure. State-of-the-art of manifold-based graph embedding algorithms assume that the projection on a tangential space of each point in the manifold (corresponding to a node in the graph) would locally resemble a Euclidean space. Although this condition helps in achieving efficient analytical solutions to the embedding problem, it is not an adequate set-up to work with modern real life graphs, that are characterized by weighted connections across nodes often computed over sparse datasets with missing records. In this work, we introduce a new class of manifold, named soft manifold, that can solve this situation. Soft manifolds are mathematical structures with spherical symmetry where the tangent spaces to each point are hypocycloids whose shape is defined according to the velocity of information propagation across the data points. Experimental results on reconstruction tasks on synthetic and real datasets show how the proposed approach enable more accurate and reliable characterization of graphs in continuous spaces with respect to the state-of-the-art. Andrea Marinoni, Pietro Liò, Alessandro Barp, Mark A. Girolami |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Autoencoders in Function SpaceabstractAutoencoders have found widespread application in both their original deterministic form and in their variational formulation (VAEs). In scientific applications and in image processing it is often of interest to consider data that are viewed as functions; while discretisation (of differential equations arising in the sciences) or pixellation (of images) renders problems finite dimensional in practice, conceiving first of algorithms that operate on functions, and only then discretising or pixellating, leads to better algorithms that smoothly operate between resolutions. In this paper function-space versions of the autoencoder (FAE) and variational autoencoder (FVAE) are introduced, analysed, and deployed. Well-definedness of the objective governing VAEs is a subtle issue, particularly in function space, limiting applicability. For the FVAE objective to be well defined requires compatibility of the data distribution with the chosen generative model; this can be achieved, for example, when the data arise from a stochastic differential equation, but is generally restrictive. The FAE objective, on the other hand, is well defined in many situations where FVAE fails to be. Pairing the FVAE and FAE objectives with neural operator architectures that can be evaluated on any mesh enables new applications of autoencoders to inpainting, superresolution, and generative modelling of scientific data. Justin Bunker, Mark A. Girolami, Hefin Lambley, Andrew M. Stuart, Timothy John Sullivan |
J. Mach. Learn. Res. | 2 |
| 2024 | Riemannian Laplace Approximation with the Fisher MetricabstractLaplace’s method approximates a target density with a Gaussian distribution at its mode. It is computationally efficient and asymptotically exact for Bayesian inference due to the Bernstein-von Mises theorem, but for complex targets and finite-data posteriors it is often too crude an approximation. A recent generalization of the Laplace Approximation transforms the Gaussian approximation according to a chosen Riemannian geometry providing a richer approximation family, while still retaining computational efficiency. However, as shown here, its properties depend heavily on the chosen metric, indeed the metric adopted in previous work results in approximations that are overly narrow as well as being biased even at the limit of infinite data. We correct this shortcoming by developing the approximation family further, deriving two alternative variants that are exact at the limit of infinite data, extending the theoretical analysis of the method, and demonstrating practical improvements in a range of experiments. Hanlin Yu, Marcelo Hartmann, Bernardo Williams, Mark A. Girolami, Arto Klami |
AISTATS | 4 |
| 2024 | Generating Origin-Destination Matrices in Neural Spatial Interaction ModelsabstractAgent-based models (ABMs) are proliferating as decision-making tools across policy areas in transportation, economics, and epidemiology. In these models, a central object of interest is the discrete origin-destination matrix which captures spatial interactions and agent trip counts between locations. Existing approaches resort to continuous approximations of this matrix and subsequent ad-hoc discretisations in order to perform ABM simulation and calibration. This impedes conditioning on partially observed summary statistics, fails to explore the multimodal matrix distribution over a discrete combinatorial support, and incurs discretisation errors. To address these challenges, we introduce a computationally efficient framework that scales linearly with the number of origin-destination pairs, operates directly on the discrete combinatorial space, and learns the agents' trip intensity through a neural differential equation that embeds spatial interactions. Our approach outperforms the prior art in terms of reconstruction error and ground truth matrix coverage, at a fraction of the computational cost. We demonstrate these benefits in two large-scale spatial mobility ABMs in Washington, DC and Cambridge, UK. Ioannis Zachos, Mark A. Girolami, Theodoros Damoulas |
NeurIPS | 2 |
| 2024 | Near Real-Time Social Distance Estimation In LondonabstractAbstract To mitigate the current COVID-19 pandemic, policy makers at the Greater London Authority, the regional governance body of London, UK, are reliant upon prompt, accurate and actionable estimations of lockdown and social distancing policy adherence. Transport for London, the local transportation department, reports they implemented over 700 interventions such as greater signage and expansion of pedestrian zoning at the height of the pandemic’s first wave with our platform providing key data for those decisions. Large well-defined heterogeneous compositions of pedestrian footfall and physical proximity are difficult to acquire, yet necessary to monitor city-wide activity (busyness) and consequently discern actionable policy decisions. To meet this challenge, we leverage our existing large-scale data processing urban air quality machine learning infrastructure to process over 900 camera feeds in near real-time to generate new estimates of social distancing adherence, group detection and camera stability. In this work, we describe our development and deployment of a computer vision and machine learning pipeline. It provides near immediate sampling and contextualization of activity and physical distancing on the streets of London via live traffic camera feeds. We introduce a platform for inspecting, calibrating and improving upon existing methods, describe the active deployment on real-time feeds and provide analysis over an 18 month period. James Walsh 0004, Oluwafunmilola Kesa, Andrew Wang 0004, Mihai Ilas, Patrick O'Hara, Oscar Giles, Neil Dhir, Mark A. Girolami, Theodoros Damoulas |
Comput. J. | 8 |
| 2024 | Bayesian dynamic modelling for probabilistic prediction of pavement conditionabstractSignificant funds have been allocated to maintain road networks each year in developed countries. Performance prediction is crucial for pavement management systems to adjust working plans and budget allocation. As a Bayesian nonparametric method, Gaussian process regression (GPR) is powerful in predicting nonlinear time series and quantifying uncertainty. However, it remains computationally intensive and fails to adapt to the time-varying characteristics. To address such issues, a dynamic GPR model is proposed for probabilistic prediction of the International Roughness Index (IRI) for flexible pavements. A moving window strategy is developed to substantially shrink the size of training data, which effectively alleviates computational cost and thus leads to a dynamic GPR. A genetic algorithm is then adopted to determine the optimal window size by considering the trade-off between computational efficiency and accuracy. A dataset acquired from Long-Term Pavement Performance (LTPP) is used to demonstrate the feasibility of the dynamic GPR. Its performance is compared to traditional GPR as well as dynamic and static Bayesian linear regression (BLR) models. The comparison results indicate that the proposed dynamic GPR can increase the accuracy by 0.86, 1.52, and 2.27 times for dynamic BLR, static GPR, and static BLR, respectively. It exhibits the best results in terms of accuracy and uncertainty metrics due to its nonlinear modelling and time-varying ability. Alix Marie d'Avigneau, Georgios M. Hadjidemetriou, Lavindra de Silva, Mark A. Girolami, Ioannis K. Brilakis |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Targeted Separation and Convergence with Kernel DiscrepanciesabstractMaximum mean discrepancies (MMDs) like the kernel Stein discrepancy (KSD) have grown central to a wide range of applications, including hypothesis testing, sampler selection, distribution approximation, and variational inference. In each setting, these kernel-based discrepancy measures are required to $(i)$ separate a target $\mathrm{P}$ from other probability measures or even $(ii)$ control weak convergence to $\mathrm{P}$. In this article we derive new sufficient and necessary conditions to ensure $(i)$ and $(ii)$. For MMDs on separable metric spaces, we characterize those kernels that separate Bochner embeddable measures and introduce simple conditions for separating all measures with unbounded kernels and for controlling convergence with bounded kernels. We use these results on $\mathbb{R}^d$ to substantially broaden the known conditions for KSD separation and convergence control and to develop the first KSDs known to exactly metrize weak convergence to $\mathrm{P}$. Along the way, we highlight the implications of our results for hypothesis testing, measuring and improving sample quality, and sampling with Stein variational gradient descent. Alessandro Barp, Carl-Johann Simon-Gabriel, Mark A. Girolami, Lester Mackey |
J. Mach. Learn. Res. | 3 |
| 2024 | Generative broad Bayesian (GBB) imputer for missing data imputation with uncertainty quantification
Sin-Chi Kuok, Ka-Veng Yuen, Tim J. Dodwell, Mark A. Girolami |
Knowl. Based Syst. | 4 |
| 2023 | Incorporating Reliability in Graph Information Propagation by Fluid Dynamics Diffusion: A case of Multimodal Semisupervised Deep LearningabstractClassic graph neural networks show some limitations in information extraction performance when applied to multimodal datasets. This is primarily due to such datasets having high volume, variety, and variability. In this paper, we propose structuring graph neural networks on a new graph representation based on fluid dynamics diffusion that allows us to incorporate the reliability of the features used to characterise each sample within the graph structure itself. This approach aims to address some of the major limitations of the classic graph-based learning structures, so to improve accuracy and robustness of the estimates. We show how this approach can help to strongly improve the quality of the analysis of classic graph neural networks. Experimental results are reported to support this point. Andrea Marinoni, Marine Mercier, Qian Shi 0001, Sivasakthy Selvakumaran, Mark A. Girolami |
ICASSP | 5 |
| 2023 | Random Grid Neural Processes for Parametric Partial Differential EquationsabstractWe introduce a new class of spatially stochastic physics and data informed deep latent models for parametric partial differential equations (PDEs) which operate through scalable variational neural processes. We achieve this by assigning probability measures to the spatial domain, which allows us to treat collocation grids probabilistically as random variables to be marginalised out. Adapting this spatial statistics view, we solve forward and inverse problems for parametric PDEs in a way that leads to the construction of Gaussian process models of solution fields. The implementation of these random grids poses a unique set of challenges for inverse physics informed deep learning frameworks and we propose a new architecture called Grid Invariant Convolutional Networks (GICNets) to overcome these challenges. We further show how to incorporate noisy data in a principled manner into our physics informed model to improve predictions for problems where data may be available but whose measurement location does not coincide with any fixed mesh or grid. The proposed method is tested on a nonlinear Poisson problem, Burgers equation, and Navier-Stokes equations, and we provide extensive numerical comparisons. We demonstrate significant computational advantages over current physics informed neural learning methods for parametric PDEs while improving the predictive capabilities and flexibility of these models. Arnaud Vadeboncoeur, Ieva Kazlauskaite, Yanni Papandreou, Fehmi Cirak, Mark A. Girolami, Ömer Deniz Akyildiz |
ICML | 5 |
| 2022 | Lagrangian manifold Monte Carlo on Monge patchesabstractThe efficiency of Markov Chain Monte Carlo (MCMC) depends on how the underlying geometry of the problem is taken into account. For distributions with strongly varying curvature, Riemannian metrics help in efficient exploration of the target distribution. Unfortunately, they have significant computational overhead due to e.g. repeated inversion of the metric tensor, and current geometric MCMC methods using the Fisher information matrix to induce the manifold are in practice slow. We propose a new alternative Riemannian metric for MCMC, by embedding the target distribution into a higher-dimensional Euclidean space as a Monge patch, thus using the induced metric determined by direct geometric reasoning. Our metric only requires first-order gradient information and has fast inverse and determinants, and allows reducing the computational complexity of individual iterations from cubic to quadratic in the problem dimensionality. We demonstrate how Lagrangian Monte Carlo in this metric efficiently explores the target distributions. Marcelo Hartmann, Mark A. Girolami, Arto Klami |
AISTATS | 2 |
| 2022 | A mixture modeling approach for clustering log files with coreset and user feedbackabstractMachine-generated log data can provide valuable insights into many critical areas such as system failures, network security, and performance optimization. The increasing prominence of this data in both volume and complexity requires data mining approaches that are both scalable and flexible. In this paper, we propose a new approach for clustering machine-generated logs which contains a novel combination of the use of the coreset with user feedback. The coreset allows us to efficiently summarize the data in a principled manner such that performance after fitting model parameters on the coreset is similar to the performance that would have been achieved with the full dataset. Furthermore, the formal approach we propose allows users to incorporate two different types of feedback, in the forms of labels and pairwise constraints, to further improve results and better deal with the increasing complexity and variety of log datasets. Justin Bunker, Kristal Curtis, Mark A. Girolami, Ram Sriharsha |
Pattern Recognit. Lett. | 3 |
| 2022 | Anomaly detection in streaming data with gaussian process based stochastic differential equations
Alex Glyn-Davies, Mark A. Girolami |
Pattern Recognit. Lett. | 2 |
| 2021 | Convergence Guarantees for Gaussian Process Means With Misspecified Likelihoods and SmoothnessabstractGaussian processes are ubiquitous in machine learning, statistics, and applied mathematics. They provide a flexible modelling framework for approximating functions, whilst simultaneously quantifying uncertainty. However, this is only true when the model is well-specified, which is often not the case in practice. In this paper, we study the properties of Gaussian process means when the smoothness of the model and the likelihood function are misspecified. In this setting, an important theoretical question of practical relevance is how accurate the Gaussian process approximations will be given the chosen model and the extent of the misspecification. The answer to this problem is particularly useful since it can inform our choice of model and experimental design. In particular, we describe how the experimental design and choice of kernel and kernel hyperparameters can be adapted to alleviate model misspecification. George Wynne, François-Xavier Briol, Mark A. Girolami |
J. Mach. Learn. Res. | 3 |
| 2020 | Dynamic content based rankingabstractWe introduce a novel state space model for a set of sequentially time-stamped partial rankings of items and textual descriptions for the items. Based on the data, the model infers text-based themes that are predictive of the rankings enabling forecasting tasks and performing trend analysis. We propose a scaled Gamma process based prior for capturing the underlying dynamics. Based on two challenging and contemporary real data collections, we show the model infers meaningful and useful textual themes as well as performs better than existing related dynamic models. Seppo Virtanen, Mark A. Girolami |
AISTATS | 2 |
| 2019 | Stein Point Markov Chain Monte CarloabstractAn important task in machine learning and statistics is the approximation of a probability measure by an empirical measure supported on a discrete point set. Stein Points are a class of algorithms for this task, which proceed by sequentially minimising a Stein discrepancy between the empirical measure and the target and, hence, require the solution of a non-convex optimisation problem to obtain each new point. This paper removes the need to solve this optimisation problem by, instead, selecting each new point based on a Markov chain sample path. This significantly reduces the computational cost of Stein Points and leads to a suite of algorithms that are straightforward to implement. The new algorithms are illustrated on a set of challenging Bayesian inference problems, and rigorous theoretical guarantees of consistency are established. Wilson Ye Chen, Alessandro Barp, François-Xavier Briol, Jackson Gorham, Mark A. Girolami, Lester Mackey, Chris J. Oates |
ICML | 5 |
| 2019 | Minimum Stein Discrepancy EstimatorsabstractWhen maximum likelihood estimation is infeasible, one often turns to score matching, contrastive divergence, or minimum probability flow to obtain tractable parameter estimates. We provide a unifying perspective of these techniques as minimum Stein discrepancy estimators, and use this lens to design new diffusion kernel Stein discrepancy (DKSD) and diffusion score matching (DSM) estimators with complementary strengths. We establish the consistency, asymptotic normality, and robustness of DKSD and DSM estimators, then derive stochastic Riemannian gradient descent algorithms for their efficient optimisation. The main strength of our methodology is its flexibility, which allows us to design estimators with desirable properties for specific models at hand by carefully selecting a Stein discrepancy. We illustrate this advantage for several challenging problems for score matching, such as non-smooth, heavy-tailed or light-tailed densities. Alessandro Barp, François-Xavier Briol, Andrew B. Duncan, Mark A. Girolami, Lester Mackey |
NeurIPS | 4 |
| 2019 | Multi-resolution Multi-task Gaussian ProcessesabstractWe consider evidence integration from potentially dependent observation processes under varying spatio-temporal sampling resolutions and noise levels. We offer a multi-resolution multi-task (MRGP) framework that allows for both inter-task and intra-task multi-resolution and multi-fidelity. We develop shallow Gaussian Process (GP) mixtures that approximate the difficult to estimate joint likelihood with a composite one and deep GP constructions that naturally handle biases. In doing so, we generalize existing approaches and offer information-theoretic corrections and efficient variational approximations. We demonstrate the competitiveness of MRGPs on synthetic settings and on the challenging problem of hyper-local estimation of air pollution levels across London from multiple sensing modalities operating at disparate spatio-temporal resolutions. Oliver Hamelijnck, Theodoros Damoulas, Kangrui Wang, Mark A. Girolami |
NeurIPS | 4 |
| 2019 | Precision-Recall Balanced Topic ModellingabstractTopic models are becoming increasingly relevant probabilistic models for dimensionality reduction of text data, inferring topics that capture meaningful themes of frequently co-occurring terms. We formulate topic modelling as an information retrieval task, where the goal is, based on the latent topic representation, to capture relevant term co-occurrence patterns. We evaluate performance for this task rigorously with regard to two types of errors, false negatives and positives, based on the well-known precision-recall trade-off and provide a statistical model that allows the user to balance between the contributions of the different error types. When the user focuses solely on the contribution of false negatives ignoring false positives altogether our proposed model reduces to a standard topic model. Extensive experiments demonstrate the proposed approach is effective and infers more coherent topics than existing related approaches. Seppo Virtanen, Mark A. Girolami |
NeurIPS | 2 |
| 2018 | Bayesian Quadrature for Multiple Related IntegralsabstractBayesian probabilistic numerical methods are a set of tools providing posterior distributions on the output of numerical methods. The use of these methods is usually motivated by the fact that they can represent our uncertainty due to incomplete/finite information about the continuous mathematical problem being approximated. In this paper, we demonstrate that this paradigm can provide additional advantages, such as the possibility of transferring information between several numerical methods. This allows users to represent uncertainty in a more faithful manner and, as a by-product, provide increased numerical efficiency. We propose the first such numerical method by extending the well-known Bayesian quadrature algorithm to the case where we are interested in computing the integral of several related functions. We then prove convergence rates for the method in the well-specified and misspecified cases, and demonstrate its efficiency in the context of multi-fidelity models for complex engineering systems and a problem of global illumination in computer graphics. Xiaoyue Xi, François-Xavier Briol, Mark A. Girolami |
ICML | 3 |
| 2018 | How Deep Are Deep Gaussian Processes?abstractRecent research has shown the potential utility of deep Gaussian processes. These deep structures are probability distributions, designed through hierarchical construction, which are conditionally Gaussian. In this paper, the current published body of work is placed in a common framework and, through recursion, several classes of deep Gaussian processes are defined. The resulting samples generated from a deep Gaussian process have a Markovian structure with respect to the depth parameter, and the effective depth of the resulting process is interpreted in terms of the ergodicity, or non-ergodicity, of the resulting Markov chain. For the classes of deep Gaussian processes introduced, we provide results concerning their ergodicity and hence their effective depth. We also demonstrate how these processes may be used for inference; in particular we show how a Metropolis-within-Gibbs construction across the levels of the hierarchy can be used to derive sampling tools which are robust to the level of resolution used to represent the functions on a computer. For illustration, we consider the effect of ergodicity in some simple numerical examples. Matthew M. Dunlop, Mark A. Girolami, Andrew M. Stuart, Aretha L. Teckentrup |
J. Mach. Learn. Res. | 2 |
| 2018 | Bat detective - Deep learning tools for bat acoustic signal detectionabstractPassive acoustic sensing has emerged as a powerful tool for quantifying anthropogenic impacts on biodiversity, especially for echolocating bat species. To better assess bat population trends there is a critical need for accurate, reliable, and open source tools that allow the detection and classification of bat calls in large collections of audio recordings. The majority of existing tools are commercial or have focused on the species classification task, neglecting the important problem of first localizing echolocation calls in audio which is particularly problematic in noisy recordings. We developed a convolutional neural network based open-source pipeline for detecting ultrasonic, full-spectrum, search-phase calls produced by echolocating bats. Our deep learning algorithms were trained on full-spectrum ultrasonic audio collected along road-transects across Europe and labelled by citizen scientists from www.batdetective.org. When compared to other existing algorithms and commercial systems, we show significantly higher detection performance of search-phase echolocation calls with our test sets. As an example application, we ran our detection pipeline on bat monitoring data collected over five years from Jersey (UK), and compared results to a widely-used commercial system. Our detection pipeline can be used for the automatic detection and monitoring of bat populations, and further facilitates their use as indicator species on a large scale. Our proposed pipeline makes only a small number of bat specific design decisions, and with appropriate training data it could be applied to detecting other species in audio. A crucial novelty of our work is showing that with careful, non-trivial, design and implementation considerations, state-of-the-art deep learning methods can be used for accurate and efficient monitoring in audio. Oisin Mac Aodha, Rory Gibb, Kate E. Barlow, Ella Browning, Michael Firman, Robin Freeman, Briana Harder, Libby Kinsey, Gary R. Mead, Stuart E. Newson, Ivan Pandourski, Stuart Parsons, Jon Russ, Abigel Szodoray-Paradi, Farkas Szodoray-Paradi, Elena Tilova, Mark A. Girolami, Gabriel J. Brostow, Kate E. Jones |
PLoS Comput. Biol. | 17 |
| 2017 | On the Sampling Problem for Kernel QuadratureabstractThe standard Kernel Quadrature method for numerical integration with random point sets (also called Bayesian Monte Carlo) is known to converge in root mean square error at a rate determined by the ratio s/d, where s and d encode the smoothness and dimension of the integrand. However, an empirical investigation reveals that the rate constant C is highly sensitive to the distribution of the random points. In contrast to standard Monte Carlo integration, for which optimal importance sampling is well-understood, the sampling distribution that minimises C for Kernel Quadrature does not admit a closed form. This paper argues that the practical choice of sampling distribution is an important open problem. One solution is considered; a novel automatic approach based on adaptive tempering and sequential Monte Carlo. Empirical results demonstrate a dramatic reduction in integration error of up to 4 orders of magnitude can be achieved with the proposed method. François-Xavier Briol, Chris J. Oates, Jon Cockayne, Wilson Ye Chen, Mark A. Girolami |
ICML | 5 |
| 2017 | Probabilistic Models for Integration Error in the Assessment of Functional Cardiac ModelsabstractThis paper studies the numerical computation of integrals, representing estimates or predictions, over the output $f(x)$ of a computational model with respect to a distribution $p(\mathrm{d}x)$ over uncertain inputs $x$ to the model. For the functional cardiac models that motivate this work, neither $f$ nor $p$ possess a closed-form expression and evaluation of either requires $\approx$ 100 CPU hours, precluding standard numerical integration methods. Our proposal is to treat integration as an estimation problem, with a joint model for both the a priori unknown function $f$ and the a priori unknown distribution $p$. The result is a posterior distribution over the integral that explicitly accounts for dual sources of numerical approximation error due to a severely limited computational budget. This construction is applied to account, in a statistically principled manner, for the impact of numerical errors that (at present) are confounding factors in functional cardiac model assessment. Chris J. Oates, Steven A. Niederer, Angela W. C. Lee, François-Xavier Briol, Mark A. Girolami |
NIPS | 5 |
| 2016 | Control Functionals for Quasi-Monte Carlo IntegrationabstractQuasi-Monte Carlo (QMC) methods are being adopted in statistical applications due to the increasingly challenging nature of numerical integrals that are now routinely encountered. For integrands with d-dimensions and derivatives of order α, an optimal QMC rule converges at a best-possible rate O(N^-α/d). However, in applications the value of αcan be unknown and/or a rate-optimal QMC rule can be unavailable. Standard practice is to employ \alpha_L-optimal QMC where the lower bound \alpha_L ≤αis known, but in general this does not exploit the full power of QMC. One solution is to trade-off numerical integration with functional approximation. This strategy is explored herein and shown to be well-suited to modern statistical computation. A challenging application to robotic arm data demonstrates a substantial variance reduction in predictions for mechanical torques. Chris J. Oates, Mark A. Girolami |
AISTATS | 2 |
| 2015 | Ordinal Mixed Membership ModelsabstractWe present a novel class of mixed membership models for joint distributions of groups of observations that co-occur with ordinal response variables for each group for learning statistical associations between the ordinal response variables and the observation groups. The class of proposed models addresses a requirement for predictive and diagnostic methods in a wide range of practical contemporary applications. In this work, by way of illustration, we apply the models to a collection of consumer-generated reviews of mobile software applications, where each review contains unstructured text data accompanied with an ordinal rating, and demonstrate that the models infer useful and meaningful recurring patterns of consumer feedback. We also compare the developed models to relevant existing works, which rely on improper statistical assumptions for ordinal variables, showing significant improvements both in predictive ability and knowledge extraction. Seppo Virtanen, Mark A. Girolami |
ICML | 2 |
| 2015 | Frank-Wolfe Bayesian Quadrature: Probabilistic Integration with Theoretical GuaranteesabstractThere is renewed interest in formulating integration as an inference problem, motivated by obtaining a full distribution over numerical error that can be propagated through subsequent computation. Current methods, such as Bayesian Quadrature, demonstrate impressive empirical performance but lack theoretical analysis. An important challenge is to reconcile these probabilistic integrators with rigorous convergence guarantees. In this paper, we present the first probabilistic integrator that admits such theoretical treatment, called Frank-Wolfe Bayesian Quadrature (FWBQ). Under FWBQ, convergence to the true value of the integral is shown to be exponential and posterior contraction rates are proven to be superexponential. In simulations, FWBQ is competitive with state-of-the-art methods and out-performs alternatives based on Frank-Wolfe optimisation. Our approach is applied to successfully quantify numerical error in the solution to a challenging model choice problem in cellular biology. François-Xavier Briol, Chris J. Oates, Mark A. Girolami, Michael A. Osborne |
NIPS | 3 |
| 2014 | Bat Call Identification with Gaussian Process Multinomial Probit Regression and a Dynamic Time Warping KernelabstractWe study the problem of identifying bat species from echolocation calls in order to build automated bioacoustic monitoring algorithms. We employ the Dynamic Time Warping algorithm which has been successfully applied for bird flight calls identification and show that classification performance is superior to hand crafted call shape parameters used in previous research. This highlights that generic bioacoustic software with good classification rates can be constructed with little domain knowledge. We conduct a study with field data of 21 bat species from the north and central Mexico using a multinomial probit regression model with Gaussian process prior and a full EP approximation of the posterior of latent function values. Results indicate high classification accuracy across almost all classes while misclassification rate across families of species is low highlighting the common evolutionary path of echolocation in bats. Vassilios Stathopoulos 0003, Veronica Zamora-Gutierrez, Kate E. Jones, Mark A. Girolami |
AISTATS | 4 |
| 2014 | Putting the Scientist in the Loop - Accelerating Scientific Progress with Interactive Machine LearningabstractTechnology drives advances in science. Giving scientists access to more powerful tools for collecting and understanding data enables them to both ask and answer new kinds questions that were previously beyond their reach. Of these new tools at their disposal, machine learning offers the opportunity to understand and analyze data at unprecedented scales and levels of detail. The standard machine learning pipeline consists of data labeling, feature extraction, training, and evaluation. However, without expert machine learning knowledge, it is difficult for scientists to optimally construct this pipeline to fully leverage machine learning in their work. Using ecology as a motivating example, we analyze a typical scientist's data collection and processing workflow and highlight many problems facing practitioners when attempting to capitalize on advances in machine learning and pattern recognition. Understanding these shortcomings allows us to outline several novel and underexplored research directions. We end with recommendations to motivate progress in future cross-disciplinary work. Oisin Mac Aodha, Vassilios Stathopoulos 0003, Gabriel J. Brostow, Michael Terry, Mark A. Girolami, Kate E. Jones |
ICPR | 5 |
| 2014 | mcmc_clib-an advanced MCMC sampling package for ode modelsabstractSUMMARY: We present a new C implementation of an advanced Markov chain Monte Carlo (MCMC) method for the sampling of ordinary differential equation (ode) model parameters. The software mcmc_clib uses the simplified manifold Metropolis-adjusted Langevin algorithm (SMMALA), which is locally adaptive; it uses the parameter manifold's geometry (the Fisher information) to make efficient moves. This adaptation does not diminish with MC length, which is highly advantageous compared with adaptive Metropolis techniques when the parameters have large correlations and/or posteriors substantially differ from multivariate Gaussians. The software is standalone (not a toolbox), though dependencies include the GNU scientific library and sundials libraries for ode integration and sensitivity analysis. AVAILABILITY AND IMPLEMENTATION: The source code and binary files are freely available for download at http://a-kramer.github.io/mcmc_clib/. This also includes example files and data. A detailed documentation, an example model and user manual are provided with the software. CONTACT: [email protected]. Andrei Kramer, Vassilios Stathopoulos 0003, Mark A. Girolami, Nicole Radde |
Bioinform. | 3 |
| 2014 | Pseudo-Marginal Bayesian Inference for Gaussian ProcessesabstractThe main challenges that arise when adopting Gaussian process priors in probabilistic modeling are how to carry out exact Bayesian inference and how to account for uncertainty on model parameters when making model-based predictions on out-of-sample data. Using probit regression as an illustrative working example, this paper presents a general and effective methodology based on the pseudo-marginal approach to Markov chain Monte Carlo that efficiently addresses both of these issues. The results presented in this paper show improvements over existing sampling methods to simulate from the posterior distribution over the parameters defining the covariance function of the Gaussian Process prior. This is particularly important as it offers a powerful tool to carry out full Bayesian inference of Gaussian Process based hierarchic statistical models in general. The results also demonstrate that Monte Carlo based integration of all model parameters is actually feasible in this class of models providing a superior quantification of uncertainty in predictions. Extensive comparisons with respect to state-of-the-art probabilistic classifiers confirm this assertion. Maurizio Filippone, Mark A. Girolami |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | A comparative evaluation of stochastic-based inference methods for Gaussian process models
Maurizio Filippone, Mingjun Zhong, Mark A. Girolami |
Mach. Learn. | 3 |
| 2013 | Online Learning with (Multiple) Kernels: A ReviewabstractThis review examines kernel methods for online learning, in particular, multiclass classification. We examine margin-based approaches, stemming from Rosenblatt's original perceptron algorithm, as well as nonparametric probabilistic approaches that are based on the popular gaussian process framework. We also examine approaches to online learning that use combinations of kernels--online multiple kernel learning. We present empirical validation of a wide range of methods on a protein fold recognition data set, where different biological feature types are available, and two object recognition data sets, Caltech101 and Caltech256, where multiple feature spaces are available in terms of different image feature extraction methods. Tom Diethe, Mark A. Girolami |
Neural Comput. | 2 |
| 2012 | A Bayesian Approach to Approximate Joint Diagonalization of Square Matrices
Mingjun Zhong, Mark A. Girolami |
ICML | 2 |
| 2012 | Markov Chain Monte Carlo Methods for State-Space Models with Point Process ObservationsabstractThis letter considers how a number of modern Markov chain Monte Carlo (MCMC) methods can be applied for parameter estimation and inference in state-space models with point process observations. We quantified the efficiencies of these MCMC methods on synthetic data, and our results suggest that the Reimannian manifold Hamiltonian Monte Carlo method offers the best performance. We further compared such a method with a previously tested variational Bayes method on two experimental data sets. Results indicate similar performance on the large data sets and superior performance on small ones. The work offers an extensive suite of MCMC algorithms evaluated on an important class of models for physiological signal analysis. Mark A. Girolami, Mahesan Niranjan |
Neural Comput. | 2 |
| 2010 | Addressing the Challenge of Defining Valid Proteomic Biomarkers and ClassifiersabstractBACKGROUND: The purpose of this manuscript is to provide, based on an extensive analysis of a proteomic data set, suggestions for proper statistical analysis for the discovery of sets of clinically relevant biomarkers. As tractable example we define the measurable proteomic differences between apparently healthy adult males and females. We choose urine as body-fluid of interest and CE-MS, a thoroughly validated platform technology, allowing for routine analysis of a large number of samples. The second urine of the morning was collected from apparently healthy male and female volunteers (aged 21-40) in the course of the routine medical check-up before recruitment at the Hannover Medical School. RESULTS: We found that the Wilcoxon-test is best suited for the definition of potential biomarkers. Adjustment for multiple testing is necessary. Sample size estimation can be performed based on a small number of observations via resampling from pilot data. Machine learning algorithms appear ideally suited to generate classifiers. Assessment of any results in an independent test-set is essential. CONCLUSIONS: Valid proteomic biomarkers for diagnosis and prognosis only can be defined by applying proper statistical data mining procedures. In particular, a justification of the sample size should be part of the study design. Mohammed Dakna, Keith Harris, Alexandros Kalousis, Sebastien Carpentier, Walter Kolch, Joost Schanstra, Marion Haubitz, Antonia Vlahou, Harald Mischak, Mark A. Girolami |
BMC Bioinform. | 10 |
| 2010 | Infinite factorization of multiple non-parametric views
Simon Rogers, Arto Klami, Janne Sinkkonen, Mark A. Girolami, Samuel Kaski |
Mach. Learn. | 4 |
| 2010 | Multiclass relevance vector machines: sparsity and accuracyabstractIn this paper, we investigate the sparsity and recognition capabilities of two approximate Bayesian classification algorithms, the multiclass multi-kernel relevance vector machines (mRVMs) that have been recently proposed. We provide an insight into the behavior of the mRVM models by performing a wide experimentation on a large range of real-world datasets. Furthermore, we monitor various model fitting characteristics that identify the predictive nature of the proposed methods and compare against existing classification techniques. By introducing novel convergence measures, sample selection strategies and model improvements, it is demonstrated that mRVMs can produce state-of-the-art results on multiclass discrimination problems. In addition, this is achieved by utilizing only a very small fraction of the available observation data. Ioannis Psorakis, Theodoros Damoulas, Mark A. Girolami |
IEEE Trans. Neural Networks | 3 |
| 2009 | Analysis of SVM with Indefinite KernelsabstractThe recent introduction of indefinite SVM by Luss and dAspremont [15] has effectively demonstrated SVM classification with a non-positive semi-definite kernel (indefinite kernel). This paper studies the properties of the objective function introduced there. In particular, we show that the objective function is continuously differentiable and its gradient can be explicitly computed. Indeed, we further show that its gradient is Lipschitz continuous. The main idea behind our analysis is that the objective function is smoothed by the penalty term, in its saddle (min-max) representation, measuring the distance between the indefinite kernel matrix and the proxy positive semi-definite one. Our elementary result greatly facilitates the application of gradient-based algorithms. Based on our analysis, we further develop Nesterovs smooth optimization approach [16,17] for indefinite SVM which has an optimal convergence rate for smooth problems. Experiments on various benchmark datasets validate our analysis and demonstrate the efficiency of our proposed algorithms. Yiming Ying, Colin Campbell, Mark A. Girolami |
NIPS | 3 |
| 2009 | Probabilistic assignment of formulas to mass peaks in metabolomics experimentsabstractMOTIVATION: High-accuracy mass spectrometry is a popular technology for high-throughput measurements of cellular metabolites (metabolomics). One of the major challenges is the correct identification of the observed mass peaks, including the assignment of their empirical formula, based on the measured mass. RESULTS: We propose a novel probabilistic method for the assignment of empirical formulas to mass peaks in high-throughput metabolomics mass spectrometry measurements. The method incorporates information about possible biochemical transformations between the empirical formulas to assign higher probability to formulas that could be created from other metabolites in the sample. In a series of experiments, we show that the method performs well and provides greater insight than assignments based on mass alone. In addition, we extend the model to incorporate isotope information to achieve even more reliable formula identification. AVAILABILITY: A supplementary document, Matlab code, data and further information are available from http://www.dcs.gla.ac.uk/inference/metsamp. Simon Rogers, Richard A. Scheltema, Mark A. Girolami, Rainer Breitling |
Bioinform. | 3 |
| 2009 | Combining feature spaces for classification
Theodoros Damoulas, Mark A. Girolami |
Pattern Recognit. | 2 |
| 2009 | Pattern recognition with a Bayesian kernel combination machine
Theodoros Damoulas, Mark A. Girolami |
Pattern Recognit. Lett. | 2 |
| 2008 | Inferring Sparse Kernel Combinations and Relevance Vectors: An Application to Subcellular Localization of ProteinsabstractIn this paper, we introduce two new formulations for multi-class multi-kernel relevance vector machines (m-RVMs) that explicitly lead to sparse solutions, both in samples and in number of kernels. This enables their application to large-scale multi-feature multinomial classification problems where there is an abundance of training samples, classes and feature spaces. The proposed methods are based on an expectation-maximization (EM) framework employing a multinomial probit likelihood and explicit pruning of non-relevant training samples. We demonstrate the methods on a low-dimensional artificial dataset. We then demonstrate the accuracy and sparsity of the method when applied to the challenging bioinformatics task of predicting protein subcellular localization. Theodoros Damoulas, Yiming Ying, Mark A. Girolami, Colin Campbell |
ICMLA | 3 |
| 2008 | Accelerating Bayesian Inference over Nonlinear Differential Equations with Gaussian ProcessesabstractIdentification and comparison of nonlinear dynamical systems using noisy and sparse experimental data is a vital task in many fields, however current methods are computationally expensive and prone to error due in part to the nonlinear nature of the likelihood surfaces induced. We present an accelerated sampling procedure which enables Bayesian inference of parameters in nonlinear ordinary and delay differential equations via the novel use of Gaussian processes (GP). Our method involves GP regression over time-series data, and the resulting derivative and time delay estimates make parameter inference possible without solving the dynamical system explicitly, resulting in dramatic savings of computational time. We demonstrate the speed and statistical accuracy of our approach using examples of both ordinary and delay differential equations, and provide a comprehensive comparison with current state of the art methods. Ben Calderhead, Mark A. Girolami, Neil D. Lawrence |
NIPS | 2 |
| 2008 | Probabilistic multi-class multi-kernel learning: on protein fold recognition and remote homology detectionabstractMOTIVATION: The problems of protein fold recognition and remote homology detection have recently attracted a great deal of interest as they represent challenging multi-feature multi-class problems for which modern pattern recognition methods achieve only modest levels of performance. As with many pattern recognition problems, there are multiple feature spaces or groups of attributes available, such as global characteristics like the amino-acid composition (C), predicted secondary structure (S), hydrophobicity (H), van der Waals volume (V), polarity (P), polarizability (Z), as well as attributes derived from local sequence alignment such as the Smith-Waterman scores. This raises the need for a classification method that is able to assess the contribution of these potentially heterogeneous object descriptors while utilizing such information to improve predictive performance. To that end, we offer a single multi-class kernel machine that informatively combines the available feature groups and, as is demonstrated in this article, is able to provide the state-of-the-art in performance accuracy on the fold recognition problem. Furthermore, the proposed approach provides some insight by assessing the significance of recently introduced protein features and string kernels. The proposed method is well-founded within a Bayesian hierarchical framework and a variational Bayes approximation is derived which allows for efficient CPU processing times. RESULTS: The best performance which we report on the SCOP PDB-40D benchmark data-set is a 70% accuracy by combining all the available feature groups from global protein characteristics but also including sequence-alignment features. We offer an 8% improvement on the best reported performance that combines multi-class k-nn classifiers while at the same time reducing computational costs and assessing the predictive power of the various available features. Furthermore, we examine the performance of our methodology on the SCOP 1.53 benchmark data-set that simulates remote homology detection and examine the combination of various state-of-the-art string kernels that have recently been proposed. Theodoros Damoulas, Mark A. Girolami |
Bioinform. | 2 |
| 2008 | vbmp: Variational Bayesian Multinomial Probit Regression for multi-class classification in RabstractSUMMARY: Vbmp is an R package for Gaussian Process classification of data over multiple classes. It features multinomial probit regression with Gaussian Process priors and estimates class posterior probabilities employing fast variational approximations to the full posterior. This software also incorporates feature weighting by means of Automatic Relevance Determination. Being equipped with only one main function and reasonable default values for optional parameters, vbmp combines flexibility with ease of usage as is demonstrated on a breast cancer microarray study. AVAILABILITY: The R library vbmp implementing this method is part of Bioconductor and can be downloaded from http://www.dcs.gla.ac.uk/~girolami Nicola Lama, Mark A. Girolami |
Bioinform. | 2 |
| 2008 | ParCrys: a Parzen window density estimation approach to protein crystallization propensity predictionabstractThe ability to rank proteins by their likely success in crystallization is useful in current Structural Biology efforts and in particular in high-throughput Structural Genomics initiatives. We present ParCrys, a Parzen Window approach to estimate a protein's propensity to produce diffraction-quality crystals. The Protein Data Bank (PDB) provided training data whilst the databases TargetDB and PepcDB were used to define feature selection data as well as test data independent of feature selection and training. ParCrys outperforms the OB-Score, SECRET and CRYSTALP on the data examined, with accuracy and Matthews correlation coefficient values of 79.1% and 0.582, respectively (74.0% and 0.227, respectively, on data with a 'real-world' ratio of positive:negative examples). ParCrys predictions and associated data are available from www.compbio.dundee.ac.uk/parcrys. Ian M. Overton, Gianandrea Padovani, Mark A. Girolami, Geoffrey J. Barton |
Bioinform. | 3 |
| 2008 | Investigating the correspondence between transcriptomic and proteomic expression profiles using coupled cluster modelsabstractMOTIVATION: Modern transcriptomics and proteomics enable us to survey the expression of RNAs and proteins at large scales. While these data are usually generated and analyzed separately, there is an increasing interest in comparing and co-analyzing transcriptome and proteome expression data. A major open question is whether transcriptome and proteome expression is linked and how it is coordinated. RESULTS: Here we have developed a probabilistic clustering model that permits analysis of the links between transcriptomic and proteomic profiles in a sensible and flexible manner. Our coupled mixture model defines a prior probability distribution over the component to which a protein profile should be assigned conditioned on which component the associated mRNA profile belongs to. We apply this approach to a large dataset of quantitative transcriptomic and proteomic expression data obtained from a human breast epithelial cell line (HMEC). The results reveal a complex relationship between transcriptome and proteome with most mRNA clusters linked to at least two protein clusters, and vice versa. A more detailed analysis incorporating information on gene function from the Gene Ontology database shows that a high correlation of mRNA and protein expression is limited to the components of some molecular machines, such as the ribosome, cell adhesion complexes and the TCP-1 chaperonin involved in protein folding. AVAILABILITY: Matlab code is available from the authors on request. Simon Rogers, Mark A. Girolami, Walter Kolch, Katrina M. Waters, Tao Liu 0003, Brian Thrall, H. Steven Wiley |
Bioinform. | 2 |
| 2008 | Bayesian ranking of biochemical system modelsabstractMOTIVATION: There often are many alternative models of a biochemical system. Distinguishing models and finding the most suitable ones is an important challenge in Systems Biology, as such model ranking, by experimental evidence, will help to judge the support of the working hypotheses forming each model. Bayes factors are employed as a measure of evidential preference for one model over another. Marginal likelihood is a key component of Bayes factors, however computing the marginal likelihood is a difficult problem, as it involves integration of nonlinear functions in multidimensional space. There are a number of methods available to compute the marginal likelihood approximately. A detailed investigation of such methods is required to find ones that perform appropriately for biochemical modelling. RESULTS: We assess four methods for estimation of the marginal likelihoods required for computing Bayes factors. The Prior Arithmetic Mean estimator, the Posterior Harmonic Mean estimator, the Annealed Importance Sampling and the Annealing-Melting Integration methods are investigated and compared on a typical case study in Systems Biology. This allows us to understand the stability of the analysis results and make reliable judgements in uncertain context. We investigate the variance of Bayes factor estimates, and highlight the stability of the Annealed Importance Sampling and the Annealing-Melting Integration methods for the purposes of comparing nonlinear models. AVAILABILITY: Models used in this study are available in SBML format as the supplementary material to this article. Vladislav Vyshemirsky, Mark A. Girolami |
Bioinform. | 2 |
| 2008 | BioBayes: A software package for Bayesian inference in systems biologyabstractMOTIVATION: There are several levels of uncertainty involved in the mathematical modelling of biochemical systems. There often may be a degree of uncertainty about the values of kinetic parameters, about the general structure of the model and about the behaviour of biochemical species which cannot be observed directly. The methods of Bayesian inference provide a consistent framework for modelling and predicting in these uncertain conditions. We present a software package for applying the Bayesian inferential methodology to problems in systems biology. RESULTS: Described herein is a software package, BioBayes, which provides a framework for Bayesian parameter estimation and evidential model ranking over models of biochemical systems defined using ordinary differential equations. The package is extensible allowing additional modules to be included by developers. There are no other such packages available which provide this functionality. Vladislav Vyshemirsky, Mark A. Girolami |
Bioinform. | 2 |
| 2008 | Bayesian ranking of biochemical system modelsabstractBioinformatics (2008), 24(6), 833–839 An error occurred in the text of the above article. Figure 1(a) and Figure 1(d) in the article and in the supplementary material must be depicted as in Figure 1 here. The differential equation for in Models 1 and 4 in the supplementary material should be . The SBML models supplied with the original article are correct and can be used without any changes. These changes do not undermine the validity of the methods discussed in the article, nor do they change the justification of the approach used. Schematic diagrams of biochemical models used in this study Vladislav Vyshemirsky, Mark A. Girolami |
Bioinform. | 2 |
| 2008 | Classifying EEG for brain computer interfaces using Gaussian processes
Mingjun Zhong, Fabien Lotte, Mark A. Girolami, Anatole Lécuyer |
Pattern Recognit. Lett. | 3 |
| 2008 | Bayesian inference for differential equations
Mark A. Girolami |
Theor. Comput. Sci. | 1 |
| 2007 | Detecting worm variants using machine learningabstractNetwork intrusion detection systems typically detect worms by examining packet or flow logs for known signatures. Not only does this approach mean worms cannot be detected until the signatures are created, but that variants of known worms will remain undetected since they will have different signatures. The intuitive solution is to write more generic signatures. This solution, however, would increase the false alarm rate and is therefore practically not feasible. This paper reports on the feasibility of using a machine learning technique to detect variants of known worms in real-time. Oliver Sharma, Mark A. Girolami, Joseph S. Sventek |
CoNEXT | 2 |
| 2007 | Bayesian model-based inference of transcription factor activityabstractBACKGROUND: In many approaches to the inference and modeling of regulatory interactions using microarray data, the expression of the gene coding for the transcription factor is considered to be an accurate surrogate for the true activity of the protein it produces. There are many instances where this is inaccurate due to post-translational modifications of the transcription factor protein. Inference of the activity of the transcription factor from the expression of its targets has predominantly involved linear models that do not reflect the nonlinear nature of transcription. We extend a recent approach to inferring the transcription factor activity based on nonlinear Michaelis-Menten kinetics of transcription from maximum likelihood to fully Bayesian inference and give an example of how the model can be further developed. RESULTS: We present results on synthetic and real microarray data. Additionally, we illustrate how gene and replicate specific delays can be incorporated into the model. CONCLUSION: We demonstrate that full Bayesian inference is appropriate in this application and has several benefits over the maximum likelihood approach, especially when the volume of data is limited. We also show the benefits of using a non-linear model over a linear model, particularly in the case of repression. Simon Rogers, Raya Khanin, Mark A. Girolami |
BMC Bioinform. | 3 |
| 2007 | An empirical analysis of the probabilistic K-nearest neighbour classifier
S. Manocha, Mark A. Girolami |
Pattern Recognit. Lett. | 2 |
| 2007 | Employing Latent Dirichlet Allocation for fraud detection in telecommunications
Dongshan Xing, Mark A. Girolami |
Pattern Recognit. Lett. | 2 |
| 2006 | Sparse Multinomial Logistic Regression via Bayesian L1 RegularisationabstractMultinomial logistic regression provides the standard penalised maximum- likelihood solution to multi-class pattern recognition problems. More recently, the development of sparse multinomial logistic regression models has found ap- plication in text processing and microarray classification, where explicit identifi- cation of the most informative features is of value. In this paper, we propose a sparse multinomial logistic regression method, in which the sparsity arises from the use of a Laplace prior, but where the usual regularisation parameter is inte- grated out analytically. Evaluation over a range of benchmark datasets reveals this approach results in similar generalisation performance to that obtained using cross-validation, but at greatly reduced computational expense. Gavin C. Cawley, Nicola L. C. Talbot, Mark A. Girolami |
NIPS | 3 |
| 2006 | Data Integration for Classification Problems Employing Gaussian Process PriorsabstractBy adopting Gaussian process priors a fully Bayesian solution to the problem of integrating possibly heterogeneous data sets within a classification setting is presented. Approximate inference schemes employing Variational & Expectation Propagation based methods are developed and rigorously assessed. We demonstrate our approach to integrating multiple data sets on a large scale protein fold prediction problem where we infer the optimal combinations of covariance functions and achieve state-of-the-art performance without resorting to any ad hoc parameter tuning and classifier combination. Mark A. Girolami, Mingjun Zhong |
NIPS | 1 |
| 2006 | Kernel Maximum Entropy Data Transformation and an Enhanced Spectral Clustering AlgorithmabstractWe propose a new kernel-based data transformation technique. It is founded on the principle of maximum entropy (MaxEnt) preservation, hence named kernel MaxEnt. The key measure is Renyi's entropy estimated via Parzen windowing. We show that kernel MaxEnt is based on eigenvectors, and is in that sense similar to kernel PCA, but may produce strikingly different transformed data sets. An enhanced spectral clustering algorithm is proposed, by replacing kernel PCA by kernel MaxEnt as an intermediate step. This has a major impact on performance. Robert Jenssen, Torbjørn Eltoft, Mark A. Girolami, Deniz Erdogmus |
NIPS | 3 |
| 2006 | Variational Bayesian Multinomial Probit Regression with Gaussian Process PriorsabstractIt is well known in the statistics literature that augmenting binary and polychotomous response models with gaussian latent variables enables exact Bayesian analysis via Gibbs sampling from the parameter posterior. By adopting such a data augmentation strategy, dispensing with priors over regression coefficients in favor of gaussian process (GP) priors over functions, and employing variational approximations to the full posterior, we obtain efficient computational methods for GP classification in the multiclass setting.1 The model augmentation with additional latent variables ensures full a posteriori class coupling while retaining the simple a priori independent GP covariance structure from which sparse approximations, such as multiclass informative vector machines (IVM), emerge in a natural and straightforward manner. This is the first time that a fully variational Bayesian treatment for multiclass GP classification has been developed without having to resort to additional explicit approximations to the nongaussian likelihood term. Empirical comparisons with exact analysis use Markov Chain Monte Carlo (MCMC) and Laplace approximations illustrate the utility of the variational approximation as a computationally economic alternative to full MCMC and it is shown to be more accurate than the Laplace approximation. Mark A. Girolami, Simon Rogers |
Neural Comput. | 1 |
| 2006 | Clustering via kernel decompositionabstractSpectral clustering methods were proposed recently which rely on the eigenvalue decomposition of an affinity matrix. In this letter, the affinity matrix is created from the elements of a nonparametric density estimator and then decomposed to obtain posterior probabilities of class membership. Hyperparameters are selected using standard cross-validation methods. Anna Szymkowiak-Have, Mark A. Girolami, Jan Larsen |
IEEE Trans. Neural Networks | 2 |
| 2005 | Hierarchic Bayesian models for kernel learningabstractThe integration of diverse forms of informative data by learning an optimal combination of base kernels in classification or regression problems can provide enhanced performance when compared to that obtained from any single data source. We present a Bayesian hierarchical model which enables kernel learning and present effective variational Bayes estimators for regression and classification. Illustrative experiments demonstrate the utility of the proposed method. Matlab code replicating results reported is available at http://www.dcs.gla.ac.uk/~srogers/kernel_comb.html. Mark A. Girolami, Simon Rogers |
ICML | 1 |
| 2005 | Probabilistic hyperspace analogue to languageabstractSong and Bruza [6] introduce a framework for Information Retrieval(IR) based on Gardenfor's three tiered cognitive model; Conceptual Spaces[4]. They instantiate a conceptual space using Hyperspace Analogue to Language (HAL[3] to generate higher order concepts which are later used for ad-hoc retrieval. In this poster, we propose an alternative implementation of the conceptual space by using a probabilistic HAL space (pHAL). To evaluate whether converting to such an implementation is beneficial we have performed an initial investigation comparing the concept combination of HAL against pHAL for the task of query expansion. Our experiments indicate that pHAL outperforms the original HAL method and that better query term selection methods can improve performance on both HAL and pHAL. Leif Azzopardi, Mark A. Girolami, Malcolm K. Crowe |
SIGIR | 2 |
| 2005 | A Bayesian regression approach to the inference of regulatory networks from gene expression dataabstractMotivation: There is currently much interest in reverse-engineering regulatory relationships between genes from microarray expression data. We propose a new algorithmic method for inferring such interactions between genes using data from gene knockout experiments. The algorithm we use is the Sparse Bayesian regression algorithm of Tipping and Faul. This method is highly suited to this problem as it does not require the data to be discretized, overcomes the need for an explicit topology search and, most importantly, requires no heuristic thresholding of the discovered connections. Results: Using simulated expression data, we are able to show that this algorithm outperforms a recently published correlation-based approach. Crucially, it does this without the need to set any ad hoc threshold on possible connections. Availability: Matlab code which allows all experimental results to be reproduced is available at http://www.dcs.gla.ac.uk/~srogers/reg_nets.html Contact: [email protected] Supplementary information: Appendices and supplementary figures mentioned in the text can be found at http://www.dcs.gla.ac.uk/~srogers/reg_nets.html Simon Rogers, Mark A. Girolami |
Bioinform. | 2 |
| 2005 | Sequential Activity Profiling: Latent Dirichlet Allocation of Markov Chains
Mark A. Girolami, Ata Kabán |
Data Min. Knowl. Discov. | 1 |
| 2005 | The Latent Process Decomposition of cDNA Microarray Data SetsabstractWe present a new computational technique (a software implementation, data sets, and supplementary information are available at http://www.enm.bris.ac.uk/lpd/) which enables the probabilistic analysis of cDNA microarray data and we demonstrate its effectiveness in identifying features of biomedical importance. A hierarchical Bayesian model, called Latent Process Decomposition (LPD), is introduced in which each sample in the data set is represented as a combinatorial mixture over a finite set of latent processes, which are expected to correspond to biological processes. Parameters in the model are estimated using efficient variational methods. This type of probabilistic model is most appropriate for the interpretation of measurement data generated by cDNA microarray technology. For determining informative substructure in such data sets, the proposed model has several important advantages over the standard use of dendrograms. First, the ability to objectively assess the optimal number of sample clusters. Second, the ability to represent samples and gene expression levels using a common set of latent variables (dendrograms cluster samples and gene expression values separately which amounts to two distinct reduced space representations). Third, in constrast to standard cluster models, observations are not assigned to a single cluster and, thus, for example, gene expression levels are modeled via combinations of the latent processes identified by the algorithm. We show this new method compares favorably with alternative cluster analysis methods. To illustrate its potential, we apply the proposed technique to several microarray data sets for cancer. For these data sets it successfully decomposes the data into known subtypes and indicates possible further taxonomic subdivision in addition to highlighting, in a wholly unsupervised manner, the importance of certain genes which are known to be medically significant. To illustrate its wider applicability, we also illustrate its performance on a microarray data set for yeast. Simon Rogers, Mark A. Girolami, Colin Campbell, Rainer Breitling |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2004 | An Assessment of Feature Relevance in Predicting Protein Function from Sequence
Ali Al-Shahib, Aik Choon Tan, Mark A. Girolami, David R. Gilbert |
IDEAL | 4 |
| 2004 | Topic based language models for ad hoc information retrievalabstractWe propose a topic based approach to language modelling for ad-hoc information retrieval (IR). Many smoothed estimators used for the multinomial query model in IR rely upon the estimated background collection probabilities. In this paper, we propose a topic based language modelling approach, that uses a more informative prior based on the topical content of a document. In our experiments, the proposed model provides comparable IR performance to the standard models, but when combined in a two stage language model, it outperforms all other estimated models. Leif Azzopardi, Mark A. Girolami, C. J. van Rijsbergen |
IJCNN | 2 |
| 2004 | User biased document language modellingabstractCapitalizing on the intuitive underlying assumptions of Language Modelling for Ad-Hoc Retrieval we present a novel approach that is capable of injecting the user's context of the document collection into the retrieval process. The preliminary findings from the evaluation undertaken suggest that improved IR performance is possible under certain circumstances. This motivates further investigation to determine the extent and significance of this improved performance. Leif Azzopardi, Mark A. Girolami, C. J. van Rijsbergen |
SIGIR | 2 |
| 2004 | Biologically valid linear factor models of gene expressionabstractAbstract Motivation: The identification of physiological processes underlying and generating the expression pattern observed in microarray experiments is a major challenge. Principal component analysis (PCA) is a linear multivariate statistical method that is regularly employed for that purpose as it provides a reduced-dimensional representation for subsequent study of possible biological processes responding to the particular experimental conditions. Making explicit the data assumptions underlying PCA highlights their lack of biological validity thus making biological interpretation of the principal components problematic. A microarray data representation which enables clear biological interpretation is a desirable analysis tool. Results: We address this issue by employing the probabilistic interpretation of PCA and proposing alternative linear factor models which are based on refined biological assumptions. A practical study on two well-understood microarray datasets highlights the weakness of PCA and the greater biological interpretability of the linear models we have developed. Availability: The model estimation routines are currently implemented as Matlab routines and these, as well as data and results reported, are available from the following URL: http://www.dcs.gla.ac.uk/~girolami/lfm/index.html Mark A. Girolami, Rainer Breitling |
Bioinform. | 1 |
| 2004 | Employing optimized combinations of one-class classifiers for automated currency validation
Mark A. Girolami, Gary Ross |
Pattern Recognit. | 2 |
| 2004 | Novelty detection employing an L2 optimal non-parametric density estimator
Mark A. Girolami |
Pattern Recognit. Lett. | 2 |
| 2003 | Simplicial Mixtures of Markov Chains: Distributed Modelling of Dynamic User ProfilesabstractTo provide a compact generative representation of the sequential activ- ity of a number of individuals within a group there is a tradeoff between the definition of individual specific and global models. This paper pro- poses a linear-time distributed model for finite state symbolic sequences representing traces of individual user activity by making the assump- tion that heterogeneous user behavior may be ‘explained’ by a relatively small number of common structurally simple behavioral patterns which may interleave randomly in a user-specific proportion. The results of an empirical study on three different sources of user traces indicates that this modelling approach provides an efficient representation scheme, re- flected by improved prediction performance as well as providing low- complexity and intuitively interpretable representations. Mark A. Girolami, Ata Kabán |
NIPS | 1 |
| 2003 | Investigating the relationship between language model perplexity and IR precision-recall measuresabstractAn empirical study has been conducted investigating the relationship between the performance of an aspect based language model in terms of perplexity and the corresponding information retrieval performance obtained. It is observed, on the corpora considered, that the perplexity of the language model has a systematic relationship with the achievable precision recall performance though it is not statistically significant. Leif Azzopardi, Mark A. Girolami, C. J. van Rijsbergen |
SIGIR | 2 |
| 2003 | On an equivalence between PLSI and LDAabstractLatent Dirichlet Allocation (LDA) is a fully generative approach to language modelling which overcomes the inconsistent generative semantics of Probabilistic Latent Semantic Indexing (PLSI). This paper shows that PLSI is a maximum a posteriori estimated LDA model under a uniform Dirichlet prior, therefore the perceived shortcomings of PLSI can be resolved and elucidated within the LDA framework. Mark A. Girolami, Ata Kabán |
SIGIR | 1 |
| 2003 | Topic Identification in Dynamical Text by Complexity Pursuit
Ella Bingham, Ata Kabán, Mark A. Girolami |
Neural Process. Lett. | 3 |
| 2003 | Probability Density Estimation from Optimally Condensed Data SamplesabstractThe requirement to reduce the computational cost of evaluating a point probability density estimate when employing a Parzen window estimator is a well-known problem. This paper presents the Reduced Set Density Estimator that provides a kernel-based density estimator which employs a small percentage of the available data sample and is optimal in the L/sub 2/ sense. While only requiring /spl Oscr/(N/sup 2/) optimization routines to estimate the required kernel weighting coefficients, the proposed method provides similar levels of performance accuracy and sparseness of representation as Support Vector Machine density estimation, which requires /spl Oscr/(N/sup 3/) optimization routines, and which has previously been shown to consistently outperform Gaussian Mixture Models. It is also demonstrated that the proposed density estimator consistently provides superior density estimates for similar levels of data reduction to that provided by the recently proposed Density-Based Multiscale Data Condensation algorithm and, in addition, has comparable computational scaling. The additional advantage of the proposed method is that no extra free parameters are introduced such as regularization, bin width, or condensation ratios, making this method a very simple and straightforward approach to providing a reduced set density estimator with comparable accuracy to that of the full sample Parzen density estimator. Mark A. Girolami |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | A General Framework for a Principled Hierarchical Visualization of Multivariate Data
Ata Kabán, Peter Tiño, Mark A. Girolami |
IDEAL | 3 |
| 2002 | Latent variable models for the topographic organisation of discrete and strictly positive data
Mark A. Girolami |
Neurocomputing | 1 |
| 2002 | A Dynamic Probabilistic Model to Visualise Topic Evolution in Text Streams
Ata Kabán, Mark A. Girolami |
J. Intell. Inf. Syst. | 2 |
| 2002 | A Probabilistic Framework for the Hierarchic Organisation and Classification of Document Collections
Alexei Vinokourov, Mark A. Girolami |
J. Intell. Inf. Syst. | 2 |
| 2002 | Orthogonal Series Density Estimation and the Kernel Eigenvalue ProblemabstractKernel principal component analysis has been introduced as a method of extracting a set of orthonormal nonlinear features from multivariate data, and many impressive applications are being reported within the literature. This article presents the view that the eigenvalue decomposition of a kernel matrix can also provide the discrete expansion coefficients required for a nonparametric orthogonal series density estimator. In addition to providing novel insights into nonparametric density estimation, this article provides an intuitively appealing interpretation for the nonlinear features extracted from data using kernel principal component analysis. Mark A. Girolami |
Neural Comput. | 1 |
| 2002 | Fast Extraction of Semantic Features from a Latent Semantic Indexed Text Corpus
Ata Kabán, Mark A. Girolami |
Neural Process. Lett. | 2 |
| 2002 | Mercer kernel-based clustering in feature spaceabstractThe article presents a method for both the unsupervised partitioning of a sample of data and the estimation of the possible number of inherent clusters which generate the data. This work exploits the notion that performing a nonlinear data transformation into some high dimensional feature space increases the probability of the linear separability of the patterns within the transformed space and therefore simplifies the associated data structure. It is shown that the eigenvectors of a kernel matrix which defines the implicit mapping provides a means to estimate the number of clusters inherent within the data and a computationally simple iterative procedure is presented for the subsequent feature space partitioning of the data. Mark A. Girolami |
IEEE Trans. Neural Networks | 1 |
| 2001 | Kernel PCA for Feature Extraction and De-Noising in Nonlinear Regression
Roman Rosipal, Mark A. Girolami, Leonard J. Trejo, Andrzej Cichocki |
Neural Comput. Appl. | 2 |
| 2001 | A Variational Method for Learning Sparse and Overcomplete RepresentationsabstractAn expectation-maximization algorithm for learning sparse and overcomplete data representations is presented. The proposed algorithm exploits a variational approximation to a range of heavy-tailed distributions whose limit is the Laplacian. A rigorous lower bound on the sparse prior distribution is derived, which enables the analytic marginalization of a lower bound on the data likelihood. This lower bound enables the development of an expectation-maximization algorithm for learning the overcomplete basis vectors and inferring the most probable basis coefficients. Mark A. Girolami |
Neural Comput. | 1 |
| 2001 | An Expectation-Maximization Approach to Nonlinear Component AnalysisabstractThe proposal of considering nonlinear principal component analysis as a kernel eigenvalue problem has provided an extremely powerful method of extracting nonlinear features for a number of classification and regression applications. Whereas the utilization of Mercer kernels makes the problem of computing principal components in, possibly, infinite-dimensional feature spaces tractable, there are still the attendant numerical problems of diagonalizing large matrices. In this contribution, we propose an expectation-maximization approach for performing kernel principal component analysis and show this to be a computationally efficient method, especially when the number of data points is large. Roman Rosipal, Mark A. Girolami |
Neural Comput. | 2 |
| 2001 | A Combined Latent Class and Trait Model for the Analysis and Visualization of Discrete DataabstractWe present a general framework for data analysis and visualization by means of topographic organization and clustering. Imposing distributional assumptions on the assumed underlying latent factors makes the proposed model suitable for both visualization and clustering. The system noise will be modeled in parametric form, as a member of the exponential family of distributions and this allows us to deal with different (continuous or discrete) types of observables in a unified framework. In this paper, we focus on discrete case formulations which, contrary to self organizing methods for continuous data, imply variants of Bregman divergencies as measures of dissimilarity between data and reference points and, also, define the matching nonlinear relation between latent and observable variables. Therefore, the trait variant of the model can be seen as a data-driven noisy nonlinear independent component analysis, which is capable of revealing meaningful structure in the multivariate observable data and visualizing it in two dimensions. The class variant (which performs the clustering) of our model performs data-driven parametric mixture modeling. The combined (trait and class) model along with the associated estimation procedures allows us to interpret the visualization result, in the sense of a topographic ordering. One important application of this work is the discovery of underlying semantic structure in text-based documents. Experimental results on various subsets of the 20-News groups text corpus and binary coded digits data are given by way of demonstration. Ata Kabán, Mark A. Girolami |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | The topographic organization and visualization of binary data using multivariate-Bernoulli latent variable modelsabstractA nonlinear latent variable model for the topographic organization and subsequent visualization of multivariate binary data is presented. The generative topographic mapping (GTM) is a nonlinear factor analysis model for continuous data which assumes an isotropic Gaussian noise model and performs uniform sampling from a two-dimensional (2-D) latent space. Despite the, success of the GTM when applied to continuous data the development of a similar model for discrete binary data has been hindered due, in part, to the nonlinear link function inherent in the binomial distribution which yields a log-likelihood that is nonlinear in the model parameters. The paper presents an effective method for the parameter estimation of a binary latent variable model-a binary version of the GTM-by adopting a variational approximation to the binomial likelihood. This approximation thus provides a log-likelihood which is quadratic in the model parameters and so obviates the necessity of an iterative M-step in the expectation maximization (EM) algorithm. The power of this method is demonstrated on two significant application domains, handwritten digit recognition and the topographic organization of semantically similar text-based documents. Mark A. Girolami |
IEEE Trans. Neural Networks | 1 |
| 2000 | A generative model for sparse discrete binary data with non-uniform categorical priors
Mark A. Girolami |
ESANN | 1 |
| 2000 | Initialized and Guided EM-Clustering of Sparse Binary Data with Application to Text Based DocumentsabstractWe investigate an alternative way of combining classification and clustering techniques for sparse binary data in order to reduce the amount of training samples required. Initializing EM from the available labels also reduces the algorithms' known dependency on the initialization, which is more evident in the case of sparse data. In addition, the two-valued Poisson class-model is proposed in this paper as a sparse variant of the usual binomial assumption. Our method can be seen as a fusion between generalized logistic regression and parametric mixture modeling. Comparative simulation results on subsets of the 20 Newsgroups' binary coded text corpora and binary handwritten digits data demonstrate the potential usefulness of the suggested method. Ata Kabán, Mark A. Girolami |
ICPR | 2 |
| 2000 | Probabilistic Hierarchical Clustering Method for Organizing Collections of Text DocumentsabstractA generic probabilistic framework for the unsupervised hierarchical clustering of large-scale sparse high-dimensional data collections is proposed. The framework is based on a hierarchical probabilistic mixture methodology. Two classes of models emerge from the analysis and these have been called symmetric and asymmetric models. For text data specifically both asymmetric and symmetric models based on multinomial and binomial distributions are most appropriate. An expectation maximisation parameter estimation method is provided for all of these models. An experimental comparison of the models is obtained for two extensive online document collections. Alexei Vinokourov, Mark A. Girolami |
ICPR | 2 |
| 1999 | Independent Component Analysis Using an Extended Infomax Algorithm for Mixed Sub-Gaussian and Super-Gaussian SourcesabstractAn extension of the infomax algorithm of Bell and Sejnowski (1995) is presented that is able blindly to separate mixed signals with sub- and supergaussian source distributions. This was achieved by using a simple type of learning rule first derived by Girolami (1997) by choosing negentropy as a projection pursuit index. Parameterized probability distributions that have sub- and supergaussian regimes were used to derive a general learning rule that preserves the simple architecture proposed by Bell and Sejnowski (1995), is optimized using the natural gradient by Amari (1998), and uses the stability analysis of Cardoso and Laheld (1996) to switch between sub- and supergaussian regimes. We demonstrate that the extended infomax algorithm is able to separate 20 sources with a variety of source distributions easily. Applied to high-dimensional data from electroencephalographic recordings, it is effective at separating artifacts such as eye blinks and line noise from weaker electrical signals that arise from sources in the brain. Te-Won Lee, Mark A. Girolami, Terrence J. Sejnowski |
Neural Comput. | 2 |
| 1999 | Blind source separation of more sources than mixtures using overcomplete representationsabstractEmpirical results were obtained for the blind source separation of more sources than mixtures using a previously proposed framework for learning overcomplete representations. This technique assumes a linear mixing model with additive noise and involves two steps: (1) learning an overcomplete representation for the observed data and (2) inferring sources given a sparse prior on the coefficients. We demonstrate that three speech signals can be separated with good fidelity given only two mixtures of the three signals. Similar results were obtained with mixtures of two speech signals and one music signal. Te-Won Lee, Michael S. Lewicki, Mark A. Girolami, Terrence J. Sejnowski |
IEEE Signal Process. Lett. | 3 |
| 1998 | Noise reduction and speech enhancement via temporal anti-Hebbian learningabstractTemporal extensions of both linear and nonlinear anti-Hebbian learning have been shown to be suited to the problem of blind separation of sources from their convolved mixtures. This paper presents a generalized form of anti-Hebbian learning for a partially connected recurrent network based on the maximum likelihood estimation principle. Inspired by features of the binaural unmasking effect the network and associated online adaptation are applied to the enhancement of speech, which is corrupted by interfering noise, competing speech and reverberation. Graded simulations based on speech corrupted with increasingly complex levels of reverberation are reported. It is shown that for high levels of reverberation the proposed method compares favorably with classical adaptive filter approaches to speech enhancement in real acoustic environments. Mark A. Girolami |
ICASSP | 1 |
| 1998 | A nonlinear model of the binaural cocktail party effect
Mark A. Girolami |
Neurocomputing | 1 |
| 1998 | An Alternative Perspective on Adaptive Independent Component Analysis AlgorithmsabstractThis article develops an extended independent component analysis algorithm for mixtures of arbitrary subgaussian and supergaussian sources. The gaussian mixture model of Pearson is employed in deriving a closed-form generic score function for strictly subgaussian sources. This is combined with the score function for a unimodal supergaussian density to provide a computationally simple yet powerful algorithm for performing independent component analysis on arbitrary mixtures of nongaussian sources. Mark A. Girolami |
Neural Comput. | 1 |
| 1998 | The Latent Variable Data Model for Exploratory Data Analysis and Visualisation: A Generalisation of the Nonlinear Infomax Algorithm
Mark A. Girolami |
Neural Process. Lett. | 1 |
| 1998 | A common neural-network model for unsupervised exploratory data analysis and independent component analysisabstractThis paper presents the derivation of an unsupervised learning algorithm, which enables the identification and visualization of latent structure within ensembles of high-dimensional data. This provides a linear projection of the data onto a lower dimensional subspace to identify the characteristic structure of the observations independent latent causes. The algorithm is shown to be a very promising tool for unsupervised exploratory data analysis and data visualization. Experimental results confirm the attractiveness of this technique for exploratory data analysis and an empirical comparison is made with the recently proposed generative topographic mapping (GTM) and standard principal component analysis (PCA). Based on standard probability density models a generic nonlinearity is developed which allows both 1) identification and visualization of dichotomised clusters inherent in the observed data and 2) separation of sources with arbitrary distributions from mixtures, whose dimensionality may be greater than that of number of sources. The resulting algorithm is therefore also a generalized neural approach to independent component analysis (ICA) and it is considered to be a promising method for analysis of real-world data that will consist of sub- and super-Gaussian components such as biomedical signals. Mark A. Girolami, Andrzej Cichocki, Shun-ichi Amari |
IEEE Trans. Neural Networks | 1 |
| 1997 | Independence is far from normal
Mark A. Girolami, Colin Fyfe |
ESANN | 1 |
| 1997 | Kurtosis extrema and identification of independent components: a neural network approachabstractWe propose a nonlinear self-organising network which solely employs computationally simple Hebbian and anti-Hebbian learning in approximating a linear independent component analysis (ICA). Current neural architectures and algorithms which perform parallel ICA are either restricted to positively kurtotic data distributions or data which exhibits one sign of kurtosis . We show that the proposed network is capable of separating mixtures of speech, noise and signals with both platykurtic (positive kurtosis) and leptokurtic (negative kurtosis) distributions in a blind manner. A simulation is reported which successfully separates a mixture of twenty sources of music, speech, noise and fundamental frequencies. Mark A. Girolami, Colin Fyfe |
ICASSP | 1 |
| 1997 | Stochastic ICA Contrast Maximisation Using Oja's Nonlinear PCA AlgorithmabstractIndependent Component Analysis (ICA) is an important extension of linear Principal Component Analysis (PCA). PCA performs a data transformation to provide independence to second order, that is, decorrelation. ICA transforms data to provide approximate independence up to and beyond second order yielding transformed data with fully factorable probability densities. The linear ICA transformation has been applied to the classical statistical signal-processing problem of Blind Separation of Sources (BSS), that is, separating unknown original source signals from a mixture whose mode of mixing is undetermined. In this paper it is shown that Oja's Nonlinear PCA algorithm performs a general stochastic online adaptive ICA. This analysis is corroborated with three simulations. The first separates unknown mixtures of original natural images, which have sub-Gaussian densities, the second separates linear mixtures of natural speech whose densities are super-Gaussian. Finally unknown mixtures of original images, which have both sub- and super-Gaussian densities are separated. Mark A. Girolami, Colin Fyfe |
Int. J. Neural Syst. | 1 |
| 1997 | An extended exploratory projection pursuit network with linear and nonlinear anti-hebbian lateral connections applied to the cocktail party problem
Mark A. Girolami, Colin Fyfe |
Neural Networks | 1 |
| 1996 | A Temporal Model of Linear Anti-Hebbian Learning
Mark A. Girolami, Colin Fyfe |
Neural Process. Lett. | 1 |