VLDB 2026 Research / reviewers in the wild / expert
Maurizio Filippone
dblp:35/5597
· DBLP profile ↗
63ranked-venue papers
10as first author
22since 2021 · last 2025
0000-0001-7294-472XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 10 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Classification by Coupling Data Mollification with Label SmoothingabstractIntroducing training-time augmentations is a key technique to enhance generalization and prepare deep neural networks against test-time corruptions. Inspired by the success of generative diffusion models, we propose a novel approach of coupling data mollification, in the form of image noising and blurring, with label smoothing to align predicted label confidences with image degradation. The method is simple to implement, introduces negligible overheads, and can be combined with existing augmentations. We demonstrate improved robustness and uncertainty quantification on the corrupted image benchmarks of CIFAR, TinyImageNet and ImageNet datasets. Markus Heinonen, Ba-Hien Tran, Michael Kampffmeyer, Maurizio Filippone |
AISTATS | 4 |
| 2025 | Unconditionally Calibrated Priors for Beta Mixture Density NetworksabstractMixture Density Networks (MDNs) allow to model arbitrarily complex mappings between inputs and mixture densities, enabling flexible conditional density estimation, at the risk of severe overfitting. A Bayesian approach can alleviate this problem by specifying a prior over the parameters of the neural network. However, these priors can be difficult to specify due to the lack of interpretability. We propose a novel neural network construction for conditional mixture densities that allows one to specify the prior in the predictive distribution domain. The construction is based on mapping the targets to the unit hypercube via a diffeomorphism, enabling the use of mixtures of Beta distributions. We prove that the prior predictive distributions are calibrated in the sense that they are equal to the unconditional density function defined by the diffeomorphism. Contrary to Bayesian Gaussian MDNs, which exhibit tied functional and distributional complexity, we show that our construction allows to decouple them. We propose an extension allowing to model correlations in the covariates via Gaussian copulas, potentially reducing the necessary number of mixture components. Our experiments show competitive performance on standard benchmarks with respect to the state of the art. Alix Lheritier, Maurizio Filippone |
AISTATS | 2 |
| 2025 | Zero-shot Model-based Reinforcement Learning using Large Language ModelsabstractThe emerging zero-shot capabilities of Large Language Models (LLMs) have led to their applications in areas extending well beyond natural language processing tasks.
In reinforcement learning, while LLMs have been extensively used in text-based environments, their integration with continuous state spaces remains understudied.
In this paper, we investigate how pre-trained LLMs can be leveraged to predict in context the dynamics of continuous Markov decision processes.
We identify handling multivariate data and incorporating the control signal as key challenges that limit the potential of LLMs' deployment in this setup and propose Disentangled In-Context Learning (DICL) to address them.
We present proof-of-concept applications in two reinforcement learning settings: model-based policy evaluation and data-augmented off-policy reinforcement learning, supported by theoretical analysis of the proposed methods.
Our experiments further demonstrate that our approach produces well-calibrated uncertainty estimates. We release the code at https://github.com/abenechehab/dicl. Abdelhakim Benechehab, Youssef Attia El Hili, Ambroise Odonnat, Oussama Zekri, Albert Thomas 0001, Giuseppe Paolo, Maurizio Filippone, Ievgen Redko, Balázs Kégl |
ICLR | 7 |
| 2025 | AdaPTS: Adapting Univariate Foundation Models to Probabilistic Multivariate Time Series ForecastingabstractPre-trained foundation models (FMs) have shown exceptional performance in univariate time series forecasting tasks. However, several practical challenges persist, including managing intricate dependencies among features and quantifying uncertainty in predictions. This study aims to tackle these critical limitations by introducing adapters—feature-space transformations that facilitate the effective use of pre-trained univariate time series FMs for multivariate tasks. Adapters operate by projecting multivariate inputs into a suitable latent space and applying the FM independently to each dimension. Inspired by the literature on representation learning and partially stochastic Bayesian neural networks, we present a range of adapters and optimization/inference strategies. Experiments conducted on both synthetic and real-world datasets confirm the efficacy of adapters, demonstrating substantial enhancements in forecasting accuracy and uncertainty quantification compared to baseline methods. Our framework, AdaPTS, positions adapters as a modular, scalable, and effective solution for leveraging time series FMs in multivariate contexts, thereby promoting their wider adoption in real-world applications. We release the code at https://github.com/abenechehab/AdaPTS. Abdelhakim Benechehab, Vasilii Feofanov, Giuseppe Paolo, Albert Thomas 0001, Maurizio Filippone, Balázs Kégl |
ICML | 5 |
| 2025 | Variational Inference for Quantum HyperNetworksabstractBinary Neural Networks (BiNNs), which employ single-bit precision weights, have emerged as a promising solution to reduce memory usage and power consumption while maintaining competitive performance in large-scale systems. However, training BiNNs remains a significant challenge due to the limitations of conventional training algorithms. Quantum HyperNetworks offer a novel paradigm for enhancing the optimization of BiNN by leveraging quantum computing. Specifically, a Variational Quantum Algorithm is employed to generate binary weights through quantum circuit measurements, while key quantum phenomena such as superposition and entanglement facilitate the exploration of a broader solution space. In this work, we establish a connection between this approach and Bayesian inference by deriving the Evidence Lower Bound (ELBO), when direct access to the output distribution is available (i.e., in simulations), and introducing a surrogate ELBO based on the Maximum Mean Discrepancy (MMD) metric for scenarios involving implicit distributions, as commonly encountered in practice. Our experimental results demonstrate that the proposed methods outperform standard Maximum Likelihood Estimation (MLE), improving trainability and generalization. Luca Nepote, Alix Lheritier, Nicolas Bondoux, Marios Kountouris, Maurizio Filippone |
IJCNN | 5 |
| 2024 | Position: Bayesian Deep Learning is Needed in the Age of Large-Scale AIabstractIn the current landscape of deep learning research, there is a predominant emphasis on achieving high predictive accuracy in supervised tasks involving large image and language datasets. However, a broader perspective reveals a multitude of overlooked metrics, tasks, and data types, such as uncertainty, active and continual learning, and scientific data, that demand attention. Bayesian deep learning (BDL) constitutes a promising avenue, offering advantages across these diverse settings. This paper posits that BDL can elevate the capabilities of deep learning. It revisits the strengths of BDL, acknowledges existing challenges, and highlights some exciting research avenues aimed at addressing these obstacles. Looking ahead, the discussion focuses on possible ways to combine large-scale foundation models with BDL to unlock their full potential. Theodore Papamarkou, Maria Skoularidou, Konstantina Palla, Laurence Aitchison, Julyan Arbel, David B. Dunson, Maurizio Filippone, Vincent Fortuin, Philipp Hennig, José Miguel Hernández-Lobato, Aliaksandr Hubin, Alexander Immer, Theofanis Karaletsos, Mohammad Emtiyaz Khan, Agustinus Kristiadi, Yingzhen Li, Stephan Mandt, Christopher Nemeth, Michael A. Osborne, Tim G. J. Rudner, David Rügamer, Yee Whye Teh, Max Welling, Andrew Gordon Wilson, Ruqi Zhang |
ICML | 7 |
| 2024 | Improved Random Features for Dot Product KernelsabstractDot product kernels, such as polynomial and exponential (softmax) kernels, are among the most widely used kernels in machine learning, as they enable modeling the interactions between input features, which is crucial in applications like computer vision, natural language processing, and recommender systems. We make several novel contributions for improving the efficiency of random feature approximations for dot product kernels, to make these kernels more useful in large scale learning. First, we present a generalization of existing random feature approximations for polynomial kernels, such as Rademacher and Gaussian sketches and TensorSRHT, using complex-valued random features. We show empirically that the use of complex features can significantly reduce the variances of these approximations. Second, we provide a theoretical analysis for understanding the factors affecting the efficiency of various random feature approximations, by deriving closed-form expressions for their variances. These variance formulas elucidate conditions under which certain approximations (e.g., TensorSRHT) achieve lower variances than others (e.g., Rademacher sketches), and conditions under which the use of complex features leads to lower variances than real features. Third, by using these variance formulas, which can be evaluated in practice, we develop a data-driven optimization approach to improve random feature approximations for general dot product kernels, which is also applicable to the Gaussian kernel. We describe the improvements brought by these contributions with extensive experiments on a variety of tasks and datasets. Jonas Wacker, Motonobu Kanagawa, Maurizio Filippone |
J. Mach. Learn. Res. | 3 |
| 2023 | Complex-to-Real Sketches for Tensor Products with Applications to the Polynomial KernelabstractRandomized sketches of a tensor product of $p$ vectors follow a tradeoff between statistical efficiency and computational acceleration. Commonly used approaches avoid computing the high-dimensional tensor product explicitly, resulting in a suboptimal dependence of $O(3^p)$ in the embedding dimension. We propose a simple Complex-to-Real (CtR) modification of well-known sketches that replaces real random projections by complex ones, incurring a lower $O(2^p)$ factor in the embedding dimension. The output of our sketches is real-valued, which renders their downstream use straightforward. In particular, we apply our sketches to $p$-fold self-tensored inputs corresponding to the feature maps of the polynomial kernel. We show that our method achieves state-of-the-art performance in terms of accuracy and speed compared to other randomized approximations from the literature. Jonas Wacker, Ruben Ohana, Maurizio Filippone |
AISTATS | 3 |
| 2023 | Fully Bayesian Autoencoders with Latent Sparse Gaussian ProcessesabstractWe present a fully Bayesian autoencoder model that treats both local latent variables and global decoder parameters in a Bayesian fashion. This approach allows for flexible priors and posterior approximations while keeping the inference costs low. To achieve this, we introduce an amortized MCMC approach by utilizing an implicit stochastic network to learn sampling from the posterior over local latent variables. Furthermore, we extend the model by incorporating a Sparse Gaussian Process prior over the latent space, allowing for a fully Bayesian treatment of inducing points and kernel hyperparameters and leading to improved scalability. Additionally, we enable Deep Gaussian Process priors on the latent space and the handling of missing data. We evaluate our model on a range of experiments focusing on dynamic representation learning and generative modeling, demonstrating the strong performance of our approach in comparison to existing methods that combine Gaussian Processes and autoencoders. Ba-Hien Tran, Babak Shahbaba, Stephan Mandt, Maurizio Filippone |
ICML | 4 |
| 2023 | Imposing Functional Priors on Bayesian Neural NetworksabstractInternational audience Bogdan L. Kozyrskiy, Dimitrios Milios, Maurizio Filippone |
ICPRAM | 3 |
| 2023 | Continuous-Time Functional Diffusion ProcessesabstractWe introduce Functional Diffusion Processes (FDPs), which generalize score-based diffusion models to infinite-dimensional function spaces. FDPs require a new mathematical framework to describe the forward and backward dynamics, and several extensions to derive practical training objectives. These include infinite-dimensional versions of Girsanov theorem, in order to be able to compute an ELBO, and of the sampling theorem, in order to guarantee that functional evaluations in a countable set of points are equivalent to infinite-dimensional functions. We use FDPs to build a new breed of generative models in function spaces, which do not require specialized network architectures, and that can work with any kind of continuous data.
Our results on real data show that FDPs achieve high-quality image generation, using a simple MLP architecture with orders of magnitude fewer parameters than existing diffusion models. Giulio Franzese, Giulio Corallo, Simone Rossi 0001, Markus Heinonen, Maurizio Filippone, Pietro Michiardi |
NeurIPS | 5 |
| 2023 | One-Line-of-Code Data Mollification Improves Optimization of Likelihood-based Generative ModelsabstractGenerative Models (GMs) have attracted considerable attention due to their tremendous success in various domains, such as computer vision where they are capable to generate impressive realistic-looking images. Likelihood-based GMs are attractive due to the possibility to generate new data by a single model evaluation. However, they typically achieve lower sample quality compared to state-of-the-art score-based Diffusion Models (DMs). This paper provides a significant step in the direction of addressing this limitation. The idea is to borrow one of the strengths of score-based DMs, which is the ability to perform accurate density estimation in low-density regions and to address manifold overfitting by means of data mollification. We propose a view of data mollification within likelihood-based GMs as a continuation method, whereby the optimization objective smoothly transitions from simple-to-optimize to the original target. Crucially, data mollification can be implemented by adding one line of code in the optimization loop, and we demonstrate that this provides a boost in generation quality of likelihood-based GMs, without computational overheads. We report results on real-world image data sets and UCI benchmarks with popular likelihood-based GMs, including variants of variational autoencoders and normalizing flows, showing large improvements in FID score and density estimation. Ba-Hien Tran, Giulio Franzese, Pietro Michiardi, Maurizio Filippone |
NeurIPS | 4 |
| 2022 | Revisiting the Effects of Stochasticity for Hamiltonian SamplersabstractWe revisit the theoretical properties of Hamiltonian stochastic differential equations (SDES) for Bayesian posterior sampling, and we study the two types of errors that arise from numerical SDE simulation: the discretization error and the error due to noisy gradient estimates in the context of data subsampling. Our main result is a novel analysis for the effect of mini-batches through the lens of differential operator splitting, revising previous literature results. The stochastic component of a Hamiltonian SDE is decoupled from the gradient noise, for which we make no normality assumptions. This leads to the identification of a convergence bottleneck: when considering mini-batches, the best achievable error rate is $\mathcal{O}(\eta^2)$, with $\eta$ being the integrator step size. Our theoretical results are supported by an empirical study on a variety of regression and classification tasks for Bayesian neural networks. Giulio Franzese, Dimitrios Milios, Maurizio Filippone, Pietro Michiardi |
ICML | 3 |
| 2022 | Locally Smoothed Gaussian Process RegressionabstractWe develop a novel framework to accelerate Gaussian process regression (GPR). In particular, we consider localization kernels at each data point to down-weigh the contributions from other data points that are far away, and we derive the GPR model stemming from the application of such localization operation. Through a set of experiments, we demonstrate the competitive performance of the proposed approach compared to full GPR, other localized models, and deep Gaussian processes. Crucially, these performances are obtained with considerable speedups compared to standard global GPR due to the sparsification effect of the Gram matrix induced by the localization operation. Davit Gogolashvili, Bogdan L. Kozyrskiy, Maurizio Filippone |
KES | 3 |
| 2022 | Variational Bootstrap for ClassificationabstractCarrying out Bayesian inference over parameters of statistical models is intractable when the likelihood and the prior are non-conjugate. Variational bootstrap provides a way to obtain samples from the posterior distribution over model parameters, where each sample is the solution of a task where the labels are perturbed. For Bayesian linear regression with a Gaussian likelihood, variational bootstrap yields samples from the exact posterior, whereas for nonlinear models with a Gaussian likelihood some guarantees of approaching the true posterior can be established. In this work, we extend variational bootstrap to the Bernoulli likelihood to tackle classification tasks. We use a transformation of the labels which allows us to turn the classification task into a regression one, and then we apply variational bootstrap to obtain samples from an approximate posterior distribution over the parameters of the model. Variational bootstrap allows us to employ advanced gradient optimization techniques which provide fast convergence. We provide experimental evidence that the proposed approach allows us to achieve classification accuracy and uncertainty estimation comparable with MCMC methods at a fraction of the cost. Bogdan L. Kozyrskiy, Dimitrios Milios, Maurizio Filippone |
KES | 3 |
| 2022 | Local Random Feature Approximations of the Gaussian KernelabstractA fundamental drawback of kernel-based statistical models is their limited scalability to large data sets, which requires resorting to approximations. In this work, we focus on the popular Gaussian kernel and on techniques to linearize kernel-based models by means of random feature approximations. In particular, we do so by studying a less explored random feature approximation based on Maclaurin expansions and polynomial sketches. We show that such approaches yield poor results when modelling high-frequency data, and we propose a novel localization scheme that improves kernel approximations and downstream performance significantly in this regime. We demonstrate these gains on a number of experiments involving the application of Gaussian process regression to synthetic and real-world data of different data sizes and dimensions. Jonas Wacker, Maurizio Filippone |
KES | 2 |
| 2022 | All You Need is a Good Functional Prior for Bayesian Deep LearningabstractThe Bayesian treatment of neural networks dictates that a prior distribution is specified over their weight and bias parameters. This poses a challenge because modern neural networks are characterized by a large number of parameters, and the choice of these priors has an uncontrolled effect on the induced functional prior, which is the distribution of the functions obtained by sampling the parameters from their prior distribution. We argue that this is a hugely limiting aspect of Bayesian deep learning, and this work tackles this limitation in a practical and effective way. Our proposal is to reason in terms of functional priors, which are easier to elicit, and to “tune” the priors of neural network parameters in a way that they reflect such functional priors. Gaussian processes offer a rigorous framework to define prior distributions over functions, and we propose a novel and robust framework to match their prior with the functional prior of neural networks based on the minimization of their Wasserstein distance. We provide vast experimental evidence that coupling these priors with scalable Markov chain Monte Carlo sampling offers systematically large performance improvements over alternative choices of priors and state-of-the-art approximate Bayesian deep learning approaches. We consider this work a considerable step in the direction of making the long-standing challenge of carrying out a fully Bayesian treatment of neural networks, including convolutional neural networks, a concrete possibility. Ba-Hien Tran, Simone Rossi 0001, Dimitrios Milios, Maurizio Filippone |
J. Mach. Learn. Res. | 4 |
| 2021 | Sparse Gaussian Processes Revisited: Bayesian Approaches to Inducing-Variable ApproximationsabstractVariational inference techniques based on inducing variables provide an elegant framework for scalable posterior estimation in Gaussian process (GP) models. Besides enabling scalability, one of their main advantages over sparse approximations using direct marginal likelihood maximization is that they provide a robust alternative for point estimation of the inducing inputs, i.e. the location of the inducing variables. In this work we challenge the common wisdom that optimizing the inducing inputs in the variational framework yields optimal performance. We show that, by revisiting old model approximations such as the fully-independent training conditionals endowed with powerful sampling-based inference methods, treating both inducing locations and GP hyper-parameters in a Bayesian way can improve performance significantly. Based on stochastic gradient Hamiltonian Monte Carlo, we develop a fully Bayesian approach to scalable GP and deep GP models, and demonstrate its state-of-the-art performance through an extensive experimental campaign across several regression and classification problems. Simone Rossi 0001, Markus Heinonen, Edwin V. Bonilla, Zheyang Shen, Maurizio Filippone |
AISTATS | 5 |
| 2021 | An Identifiable Double VAE For Disentangled RepresentationsabstractA large part of the literature on learning disentangled representations focuses on variational autoencoders (VAEs). Recent developments demonstrate that disentanglement cannot be obtained in a fully unsupervised setting without inductive biases on models and data. However, Khemakhem et al., AISTATS, 2020 suggest that employing a particular form of factorized prior, conditionally dependent on auxiliary variables complementing input observations, can be one such bias, resulting in an identifiable model with guarantees on disentanglement. Working along this line, we propose a novel VAE-based generative model with theoretical guarantees on identifiability. We obtain our conditional prior over the latents by learning an optimal representation, which imposes an additional strength on their regularization. We also extend our method to semi-supervised settings. Experimental results indicate superior performance with respect to state-of-the-art approaches, according to several established metrics proposed in the literature on disentanglement. Graziano Mita, Maurizio Filippone, Pietro Michiardi |
ICML | 2 |
| 2021 | Sparse within Sparse Gaussian Processes using Neighbor InformationabstractApproximations to Gaussian processes (GPs) based on inducing variables, combined with variational inference techniques, enable state-of-the-art sparse approaches to infer GPs at scale through mini-batch based learning. In this work, we further push the limits of scalability of sparse GPs by allowing large number of inducing variables without imposing a special structure on the inducing inputs. In particular, we introduce a novel hierarchical prior, which imposes sparsity on the set of inducing variables. We treat our model variationally, and we experimentally show considerable computational gains compared to standard sparse GPs when sparsity on the inducing variables is realized considering the nearest inducing inputs of a random mini-batch of the data. We perform an extensive experimental validation that demonstrates the effectiveness of our approach compared to the state-of-the-art. Our approach enables the possibility to use sparse GPs using a large number of inducing points without incurring a prohibitive computational cost. Gia-Lac Tran, Dimitrios Milios, Pietro Michiardi, Maurizio Filippone |
ICML | 4 |
| 2021 | Multimodal Variational Autoencoders for Sensor Fusion and Cross GenerationabstractThe cognitive system of humans, which allows them to create representations of their surroundings exploiting multiple senses, has inspired several applications to mimic this remarkable property. The key for learning rich representations of data collected by multiple, diverse sensors, is to design generative models that can ingest multimodal inputs, and merge them in a common space. This enables to: i) obtain a coherent generation of samples for all modalities, ii) enable cross-sensor generation, by using available modalities to generate missing ones and iii) exploit synergy across modalities, to increase reconstruction quality. In this work, we study multimodal variational autoencoders, and propose new methods for learning a joint representation that can both improve synergy and enable cross generation of missing sensor data. We evaluate these approaches on well-established datasets as well as on a new dataset that involves multimodal object detection with three modalities. Our results shed light on the role of joint posterior modeling and training objectives, indicating that even simple and efficient heuristics enable both synergy and cross generation properties to coexist. Matthieu Da Silva-Filarder, Andrea Ancora, Maurizio Filippone, Pietro Michiardi |
ICMLA | 3 |
| 2021 | Model Selection for Bayesian AutoencodersabstractWe develop a novel method for carrying out model selection for Bayesian autoencoders (BAEs) by means of prior hyper-parameter optimization. Inspired by the common practice of type-II maximum likelihood optimization and its equivalence to Kullback-Leibler divergence minimization, we propose to optimize the distributional sliced-Wasserstein distance (DSWD) between the output of the autoencoder and the empirical data distribution. The advantages of this formulation are that we can estimate the DSWD based on samples and handle high-dimensional problems. We carry out posterior estimation of the BAE parameters via stochastic gradient Hamiltonian Monte Carlo and turn our BAE into a generative model by fitting a flexible Dirichlet mixture model in the latent space. Thanks to this approach, we obtain a powerful alternative to variational autoencoders, which are the preferred choice in modern application of autoencoders for representation learning with uncertainty. We evaluate our approach qualitatively and quantitatively using a vast experimental campaign on a number of unsupervised learning tasks and show that, in small-data regimes where priors matter, our approach provides state-of-the-art results, outperforming multiple competitive baselines. Ba-Hien Tran, Simone Rossi 0001, Dimitrios Milios, Pietro Michiardi, Edwin V. Bonilla, Maurizio Filippone |
NeurIPS | 6 |
| 2020 | LIBRE: Learning Interpretable Boolean Rule EnsemblesabstractWe present a novel method—LIBRE—learn an interpretable classifier, which materializes as a set of Boolean rules. LIBRE uses an ensemble of bottom-up, weak learners operating on a random subset of features, which allows for the learning of rules that generalize well on unseen data even in imbalanced settings. Weak learners are combined with a simple union so that the final ensemble is also interpretable. Experimental results indicate that LIBRE efficiently strikes the right balance between prediction accuracy, which is competitive with black-box methods, and interpretability, which is often superior to alternative methods from the literature. Graziano Mita, Paolo Papotti, Maurizio Filippone, Pietro Michiardi |
AISTATS | 3 |
| 2020 | Kernel Computations from Large-Scale Random Features Obtained by Optical Processing UnitsabstractApproximating kernel functions with random features (RFs) has been a successful application of random projections for nonparametric estimation. However, performing random projections presents computational challenges for large-scale problems. Recently, a new optical hardware called Optical Processing Unit (OPU) has been developed for fast and energy-efficient computation of large-scale RFs in the analog domain. More specifically, the OPU performs the multiplication of input vectors by a large random matrix with complexvalued i.i.d. Gaussian entries, followed by the application of an element-wise squared absolute value operation - this last nonlinearity being intrinsic to the sensing process. In this paper, we show that this operation results in a dot-product kernel that has connections to the polynomial kernel, and we extend this computation to arbitrary powers of the feature map. Experiments demonstrate that the OPU kernel and its RF approximation achieve competitive performance in applications using kernel ridge regression and transfer learning for image classification. Crucially, thanks to the use of the OPU, these results are obtained with time and energy savings. Ruben Ohana, Jonas Wacker, Jonathan Dong, Sébastien Marmin, Florent Krzakala, Maurizio Filippone, Laurent Daudet |
ICASSP | 6 |
| 2020 | Walsh-Hadamard Variational Inference for Bayesian Deep LearningabstractOver-parameterized models, such as DeepNets and ConvNets, form a class of models that are routinely adopted in a wide variety of applications, and for which Bayesian inference is desirable but extremely challenging. Variational inference offers the tools to tackle this challenge in a scalable way and with some degree of flexibility on the approximation, but for overparameterized models this is challenging due to the over-regularization property of the variational objective. Inspired by the literature on kernel methods, and in particular on structured approximations of distributions of random matrices, this paper proposes Walsh-Hadamard Variational Inference (WHVI), which uses Walsh-Hadamardbased factorization strategies to reduce the parameterization and accelerate computations, thus avoiding over-regularization issues with the variational objective. Extensive theoretical and empirical analyses demonstrate that WHVI yields considerable speedups and model reductions compared to other techniques to carry out approximate inference for over-parameterized models, and ultimately show how advances in kernel methods can be translated into advances in approximate Bayesian inference for Deep Learning. Simone Rossi 0001, Sébastien Marmin, Maurizio Filippone |
NeurIPS | 3 |
| 2020 | Model Monitoring and Dynamic Model Selection in Travel Time-Series Forecasting
Rosa Candela, Pietro Michiardi, Maurizio Filippone, Maria A. Zuluaga |
ECML/PKDD (4) | 3 |
| 2019 | Calibrating Deep Convolutional Gaussian ProcessesabstractThe wide adoption of Convolutional Neural Networks CNNs in applications where decision-making under uncertainty is fundamental, has brought a great deal of attention to the ability of these models to accurately quantify the uncertainty in their predictions. Previous work on combining CNNs with Gaussian processes GPs has been developed under the assumption that the predictive probabilities of these models are well-calibrated. In this paper we show that, in fact, current combinations of CNNs and GPs are miscalibrated. We proposes a novel combination that considerably outperforms previous approaches on this aspect, while achieving state-of-the-art performance on image classification tasks. Gia-Lac Tran, Edwin V. Bonilla, John P. Cunningham, Pietro Michiardi, Maurizio Filippone |
AISTATS | 5 |
| 2019 | Good Initializations of Variational Bayes for Deep ModelsabstractStochastic variational inference is an established way to carry out approximate Bayesian inference for deep models flexibly and at scale. While there have been effective proposals for good initializations for loss minimization in deep learning, far less attention has been devoted to the issue of initialization of stochastic variational inference. We address this by proposing a novel layer-wise initialization strategy based on Bayesian linear models. The proposed method is extensively validated on regression and classification tasks, including Bayesian Deep Nets and Conv Nets, showing faster and better convergence compared to alternatives inspired by the literature on initializations for loss minimization. Simone Rossi 0001, Pietro Michiardi, Maurizio Filippone |
ICML | 3 |
| 2019 | Pseudo-Extended Markov chain Monte CarloabstractSampling from posterior distributions using Markov chain Monte Carlo (MCMC) methods can require an exhaustive number of iterations, particularly when the posterior is multi-modal as the MCMC sampler can become trapped in a local mode for a large number of iterations. In this paper, we introduce the pseudo-extended MCMC method as a simple approach for improving the mixing of the MCMC sampler for multi-modal posterior distributions. The pseudo-extended method augments the state-space of the posterior using pseudo-samples as auxiliary variables. On the extended space, the modes of the posterior are connected, which allows the MCMC sampler to easily move between well-separated posterior modes. We demonstrate that the pseudo-extended approach delivers improved MCMC sampling over the Hamiltonian Monte Carlo algorithm on multi-modal posteriors, including Boltzmann machines and models with sparsity-inducing priors. Christopher Nemeth, Fredrik Lindsten, Maurizio Filippone, James Hensman |
NeurIPS | 3 |
| 2018 | Constraining the Dynamics of Deep Probabilistic ModelsabstractWe introduce a novel generative formulation of deep probabilistic models implementing "soft" constraints on their function dynamics. In particular, we develop a flexible methodological framework where the modeled functions and derivatives of a given order are subject to inequality or equality constraints. We then characterize the posterior distribution over model and constraint parameters through stochastic variational inference. As a result, the proposed approach allows for accurate and scalable uncertainty quantification on the predictions and on all parameters. We demonstrate the application of equality constraints in the challenging problem of parameter inference in ordinary differential equation models, while we showcase the application of inequality constraints on the problem of monotonic regression of count data. The proposed approach is extensively tested in several experimental settings, leading to highly competitive results in challenging modeling applications, while offering high expressiveness, flexibility and scalability. Marco Lorenzi, Maurizio Filippone |
ICML | 2 |
| 2018 | Dirichlet-based Gaussian Processes for Large-scale Calibrated ClassificationabstractThis paper studies the problem of deriving fast and accurate classification algorithms with uncertainty quantification. Gaussian process classification provides a principled approach, but the corresponding computational burden is hardly sustainable in large-scale problems and devising efficient alternatives is a challenge. In this work, we investigate if and how Gaussian process regression directly applied to classification labels can be used to tackle this question. While in this case training is remarkably faster, predictions need to be calibrated for classification and uncertainty estimation. To this aim, we propose a novel regression approach where the labels are obtained through the interpretation of classification labels as the coefficients of a degenerate Dirichlet distribution. Extensive experimental results show that the proposed approach provides essentially the same accuracy and uncertainty quantification as Gaussian process classification while requiring only a fraction of computational resources. Dimitrios Milios, Raffaello Camoriano, Pietro Michiardi, Lorenzo Rosasco, Maurizio Filippone |
NeurIPS | 5 |
| 2018 | Deep Gaussian Process autoencoders for novelty detection
Remi Domingues, Pietro Michiardi, Jihane Zouaoui, Maurizio Filippone |
Mach. Learn. | 4 |
| 2018 | A comparative evaluation of outlier detection algorithms: Experiments and analyses
Remi Domingues, Maurizio Filippone, Pietro Michiardi, Jihane Zouaoui |
Pattern Recognit. | 2 |
| 2017 | Random Feature Expansions for Deep Gaussian ProcessesabstractThe composition of multiple Gaussian Processes as a Deep Gaussian Process DGP enables a deep probabilistic nonparametric approach to flexibly tackle complex machine learning problems with sound quantification of uncertainty. Existing inference approaches for DGP models have limited scalability and are notoriously cumbersome to construct. In this work we introduce a novel formulation of DGPs based on random feature expansions that we train using stochastic variational inference. This yields a practical learning framework which significantly advances the state-of-the-art in inference for DGPs, and enables accurate quantification of uncertainty. We extensively showcase the scalability and performance of our proposal on several datasets with up to 8 million observations, and various DGP architectures with up to 30 hidden layers. Kurt Cutajar, Edwin V. Bonilla, Pietro Michiardi, Maurizio Filippone |
ICML | 4 |
| 2017 | Mini-batch spectral clusteringabstractThe cost of computing the spectrum of Laplacian matrices hinders the application of spectral clustering to large data sets. While approximations recover computational tractability, they can potentially affect clustering performance. This paper proposes a practical approach to learn spectral clustering, where the spectrum of the Laplacian is recovered following a constrained optimization problem that we solve using adaptive mini-batch-based stochastic gradient optimization on Stiefel manifolds. Crucially, the proposed approach is formulated so that the memory footprint of the algorithm is low, the cost of each iteration is linear in the number of samples, and convergence to critical points of the objective function is guaranteed. Extensive experimental validation on data sets with up to half a million samples demonstrate its scalability and its ability to outperform state-of-the-art approximate methods to learn spectral clustering for a given computational budget. Maurizio Filippone |
IJCNN | 2 |
| 2017 | Entropic Trace Estimates for Log Determinants
Jack K. Fitzsimons, Diego Granziol, Kurt Cutajar, Michael A. Osborne, Maurizio Filippone, Stephen J. Roberts |
ECML/PKDD (1) | 5 |
| 2017 | Bayesian Inference of Log Determinants
Jack K. Fitzsimons, Kurt Cutajar, Maurizio Filippone, Michael A. Osborne, Stephen J. Roberts |
UAI | 3 |
| 2017 | AutoGP: Exploring the Capabilities and Limitations of Gaussian Process Models
Karl Krauth, Edwin V. Bonilla, Kurt Cutajar, Maurizio Filippone |
UAI | 4 |
| 2016 | Preconditioning Kernel MatricesabstractThe computational and storage complexity of kernel machines presents the primary barrier to their scaling to large, modern, datasets. A common way to tackle the scalability issue is to use the conjugate gradient algorithm, which relieves the constraints on both storage (the kernel matrix need not be stored) and computation (both stochastic gradients and parallelization can be used). Even so, conjugate gradient is not without its own issues: the conditioning of kernel matrices is often such that conjugate gradients will have poor convergence in practice. Preconditioning is a common approach to alleviating this issue. Here we propose preconditioned conjugate gradients for kernel machines, and develop a broad range of preconditioners particularly useful for kernel matrices. We describe a scalable approach to both solving kernel machines and learning their hyperparameters. We show this approach is exact in the limit of iterations and outperforms state-of-the-art approximations for a given computational budget. Kurt Cutajar, Michael A. Osborne, John P. Cunningham, Maurizio Filippone |
ICML | 4 |
| 2016 | Fast Parameter Inference in Nonlinear Dynamical Systems using Iterative Gradient MatchingabstractParameter inference in mechanistic models of coupled differential equations is a topical and challenging problem. We propose a new method based on kernel ridge regression and gradient matching, and an objective function that simultaneously encourages goodness of fit and penalises inconsistencies with the differential equations. Fast minimisation is achieved by exploiting partial convexity inherent in this function, and setting up an iterative algorithm in the vein of the EM algorithm. An evaluation of the proposed method on various benchmark data suggests that it compares favourably with state-of-the-art alternatives. Mu Niu, Simon Rogers, Maurizio Filippone, Dirk Husmeier |
ICML | 3 |
| 2016 | Looking Good With Flickr Faves: Gaussian Processes for Finding Difference Makers in Personality ImpressionsabstractFlickr allows its users to generate galleries of "faves", i.e., pictures that they have tagged as favourite. According to recent studies, the faves are predictive of the personality traits that people attribute to Flickr users. This article investigates the phenomenon and shows that faves allow one to predict whether a Flickr user is perceived to be above median or not with respect to each of the Big-Five Traits (accuracy up to 79\% depending on the trait). The classifier - based on Gaussian Processes with a new kernel designed for this work - allows one to identify the visual characteristics of faves that better account for the prediction outcome. Xiaoyu Xiong, Maurizio Filippone, Alessandro Vinciarelli |
ACM Multimedia | 2 |
| 2015 | Monte Carlo Strength Evaluation: Fast and Reliable Password CheckingabstractModern password guessing attacks adopt sophisticated probabilistic techniques that allow for orders of magnitude less guesses to succeed compared to brute force. Unfortunately, best practices and password strength evaluators failed to keep up: they are generally based on heuristic rules designed to defend against obsolete brute force attacks. Many passwords can only be guessed with significant effort, and motivated attackers may be willing to invest resources to obtain valuable passwords. However, it is eminently impractical for the defender to simulate expensive attacks against each user to accurately characterize their password strength. This paper proposes a novel method to estimate the number of guesses needed to find a password using modern attacks. The proposed method requires little resources, applies to a wide set of probabilistic models, and is characterised by highly desirable convergence properties. Matteo Dell'Amico, Maurizio Filippone |
CCS | 2 |
| 2015 | Enabling scalable stochastic gradient-based inference for Gaussian processes by employing the Unbiased LInear System SolvEr (ULISSE)abstractIn applications of Gaussian processes where quantification of uncertainty is of primary interest, it is necessary to accurately characterize the posterior distribution over covariance parameters. This paper proposes an adaptation of the Stochastic Gradient Langevin Dynamics algorithm to draw samples from the posterior distribution over covariance parameters with negligible bias and without the need to compute the marginal likelihood. In Gaussian process regression, this has the enormous advantage that stochastic gradients can be computed by solving linear systems only. A novel unbiased linear systems solver based on parallelizable covariance matrix-vector products is developed to accelerate the unbiased estimation of gradients. The results demonstrate the possibility to enable scalable and exact (in a Monte Carlo sense) quantification of uncertainty in Gaussian processes without imposing any special structure on the covariance or reducing the number of input vectors. Maurizio Filippone, Raphael Engler |
ICML | 1 |
| 2015 | MCMC for Variationally Sparse Gaussian ProcessesabstractGaussian process (GP) models form a core part of probabilistic machine learning. Considerable research effort has been made into attacking three issues with GP models: how to compute efficiently when the number of data is large; how to approximate the posterior when the likelihood is not Gaussian and how to estimate covariance function parameter posteriors. This paper simultaneously addresses these, using a variational approximation to the posterior which is sparse in sup- port of the function but otherwise free-form. The result is a Hybrid Monte-Carlo sampling scheme which allows for a non-Gaussian approximation over the function values and covariance parameters simultaneously, with efficient computations based on inducing-point sparse GPs. James Hensman, Alexander G. de G. Matthews, Maurizio Filippone, Zoubin Ghahramani |
NIPS | 3 |
| 2015 | On User Availability Prediction and Network ApplicationsabstractUser connectivity patterns in network applications are known to be heterogeneous and to follow periodic (daily and weekly) patterns. In many cases, the regularity and the correlation of those patterns is problematic: For network applications, many connected users create peaks of demand; in contrast, in peer-to-peer scenarios, having few users online results in a scarcity of available resources. On the other hand, since connectivity patterns exhibit a periodic behavior, they are to some extent predictable. This paper shows how this can be exploited to anticipate future user connectivity and to have applications proactively responding to it. We evaluate the probability that any given user will be online at any given time, and assess the prediction on 6-month availability traces from three different Internet applications. Building upon this, we show how our probabilistic approach makes it easy to evaluate and optimize the performance in a number of diverse network application models and to use them to optimize systems. In particular, we show how this approach can be used in distributed hash tables, friend-to-friend storage, and cache preloading for social networks, resulting in substantial gains in data availability and system efficiency at negligible costs. Matteo Dell'Amico, Maurizio Filippone, Pietro Michiardi, Yves Roudier |
IEEE/ACM Trans. Netw. | 2 |
| 2014 | Bayesian Inference for Gaussian Process Classifiers with Annealing and Pseudo-Marginal MCMCabstractKernel methods have revolutionized the fields of pattern recognition and machine learning. Their success, however, critically depends on the choice of kernel parameters. Using Gaussian process (GP) classification as a working example, this paper focuses on Bayesian inference of covariance (kernel) parameters using Markov chain Monte Carlo (MCMC) methods. The motivation is that, compared to standard optimization of kernel parameters, they have been systematically demonstrated to be superior in quantifying uncertainty in predictions. Recently, the Pseudo-Marginal MCMC approach has been proposed as a practical inference tool for GP models. In particular, it amounts in replacing the analytically intractable marginal likelihood by an unbiased estimate obtainable by approximate methods and importance sampling. After discussing the potential drawbacks in employing importance sampling, this paper proposes the application of annealed importance sampling. The results empirically demonstrate that compared to importance sampling, annealed importance sampling can reduce the variance of the estimate of the marginal likelihood exponentially in the number of data at a computational cost that scales only polynomially. The results on real data demonstrate that employing annealed importance sampling in the Pseudo-Marginal MCMC approach represents a step forward in the development of fully automated exact inference engines for GP models. Maurizio Filippone |
ICPR | 1 |
| 2014 | Pseudo-Marginal Bayesian Multiple-Class Multiple-Kernel Learning for Neuroimaging DataabstractIn clinical neuroimaging applications where subjects belong to one of multiple classes of disease states and multiple imaging sources are available, the aim is to achieve accurate classification while assessing the importance of the sources in the classification task. This work proposes the use of fully Bayesian multiple-class multiple-kernel learning based on Gaussian Processes, as it offers flexible classification capabilities and a sound quantification of uncertainty in parameter estimates and predictions. The exact inference of parameters and accurate quantification of uncertainty in Gaussian Process models, however, poses a computationally challenging problem. This paper proposes the application of advanced inference techniques based on Markov chain Monte Carlo and unbiased estimates of the marginal likelihood, and demonstrates their ability to accurately and efficiently carry out inference in their application on synthetic data and real clinical neuroimaging data. The results in this paper are important as they further work in the direction of achieving computationally feasible fully Bayesian models for a wide range of real world applications. Andrew D. O'Harney, Andre F. Marquand, Katya Rubia, Kaylita Chantiluke, Anna B. Smith, Ana Cubillo, Camilla Blain, Maurizio Filippone |
ICPR | 8 |
| 2014 | Pseudo-Marginal Bayesian Inference for Gaussian ProcessesabstractThe main challenges that arise when adopting Gaussian process priors in probabilistic modeling are how to carry out exact Bayesian inference and how to account for uncertainty on model parameters when making model-based predictions on out-of-sample data. Using probit regression as an illustrative working example, this paper presents a general and effective methodology based on the pseudo-marginal approach to Markov chain Monte Carlo that efficiently addresses both of these issues. The results presented in this paper show improvements over existing sampling methods to simulate from the posterior distribution over the parameters defining the covariance function of the Gaussian Process prior. This is particularly important as it offers a powerful tool to carry out full Bayesian inference of Gaussian Process based hierarchic statistical models in general. The results also demonstrate that Monte Carlo based integration of all model parameters is actually feasible in this class of models providing a superior quantification of uncertainty in predictions. Extensive comparisons with respect to state-of-the-art probabilistic classifiers confirm this assertion. Maurizio Filippone, Mark A. Girolami |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Predicting Continuous Conflict Perceptionwith Bayesian Gaussian ProcessesabstractConflict is one of the most important phenomena of social life, but it is still largely neglected by the computing community. This work proposes an approach that detects common conversational social signals (loudness, overlapping speech, etc.) and predicts the conflict level perceived by human observers in continuous, non-categorical terms. The proposed regression approach is fully Bayesian and it adopts automatic relevance determination to identify the social signals that influence most the outcome of the prediction. The experiments are performed over the SSPNet Conflict Corpus, a publicly available collection of 1,430 clips extracted from televised political debates (roughly 12 hours of material for 138 subjects in total). The results show that it is possible to achieve a correlation close to 0.8 between actual and predicted conflict perception. Samuel Kim, Fabio Valente, Maurizio Filippone, Alessandro Vinciarelli |
IEEE Trans. Affect. Comput. | 3 |
| 2013 | ODE parameter inference using adaptive gradient matching with Gaussian processesabstractParameter inference in mechanistic models based on systems of coupled differential equations is a topical yet computationally challenging problem, due to the need to follow each parameter adaptation with a numerical integration of the differential equations. Techniques based on gradient matching, which aim to minimize the discrepancy between the slope of a data interpolant and the derivatives predicted from the differential equations, offer a computationally appealing shortcut to the inference problem. The present paper discusses a method based on nonparametric Bayesian statistics with Gaussian processes due to Calderhead et al. (2008), and shows how inference in this model can be substantially improved by consistently sampling from the joint distribution of the ODE parameters and GP hyperparameters. We demonstrate the efficiency of our adaptive gradient matching technique on three benchmark systems, and perform a detailed comparison with the method in Calderhead et al. (2008) and the explicit ODE integration approach, both in terms of parameter inference accuracy and in terms of computational efficiency. Frank Dondelinger, Dirk Husmeier, Simon Rogers, Maurizio Filippone |
AISTATS | 4 |
| 2013 | A comparative evaluation of stochastic-based inference methods for Gaussian process models
Maurizio Filippone, Mingjun Zhong, Mark A. Girolami |
Mach. Learn. | 1 |
| 2013 | Editorial A Successful Change From TNN to TNNLS and a Very Successful YearabstractThis issue marks the first anniversary issue of IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS after it changed its name from IEEE TRANSACTIONS ON NEURAL NETWORKS. I am happy to report that we had a great year! The number of new submissions in a year exceeded 1,000 for the first time in the history of TNN/TNNLS. IEEE TNN had a very successful development for 22 years from 1990 to 2011, and we have good reasons to believe that IEEE TNNLS will have many more years of successful growth. Derong Liu 0001, Charles W. Anderson, Ahmad Taher Azar, Giorgio Battistelli, Eduardo Bayro-Corrochano, Cristiano Cervellera, David A. Elizondo, Maurizio Filippone, Giorgio Gnecco, Tingwen Huang, Weifeng Liu 0016, Wenlian Lu, Ana Madureira, Igor Skrjanc, Thomas Villmann, Q. M. Jonathan Wu, Shengli Xie 0001, Dong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2012 | Predicting the conflict level in television political debates: an approach based on crowdsourcing, nonverbal communication and gaussian processesabstractOne of the most recent trends in multimedia indexing is to represent data in terms of the social and psychological phenomena that users perceive. In such a perspective this article proposes an approach for the automatic detection of conflict level in television political debates. The proposed approach includes the use of crowdsourcing techniques for modeling the perception of data consumers, the extraction of (language independent) nonverbal behavioral cues and the application of regression techniques based on Gaussian Processes. The experiments have been performed over 1430 clips of 30 seconds extracted from 45 political debates (roughly 12 hours of material). The results show that a correlation up to 0.8 can be achieved between the actual and predicted conflict level. Samuel Kim, Maurizio Filippone, Fabio Valente, Alessandro Vinciarelli |
ACM Multimedia | 2 |
| 2012 | From speech to personality: mapping voice quality and intonation into personality differencesabstractFrom a cognitive point of view, personality perception corresponds to capturing individual differences and can be thought of as positioning the people around us in an ideal personality space. The more similar the personality of two individuals, the closer their position in the space. This work shows that the mutual position of two individuals in the personality space can be inferred from prosodic features. The experiments, based on ordinal regression techniques, have been performed over a corpus of 640 speech samples comprising 322 individuals assessed in terms of personality traits by 11 human judges, which is the largest database of this type in the literature. The results show that the mutual position of two individuals can be predicted with up to 80% accuracy. Gelareh Mohammadi, Antonio Origlia, Maurizio Filippone, Alessandro Vinciarelli |
ACM Multimedia | 3 |
| 2011 | Simulated annealing for supervised gene selection
Maurizio Filippone, Francesco Masulli, Stefano Rovetta |
Soft Comput. | 1 |
| 2010 | Information theoretic novelty detection
Maurizio Filippone, Guido Sanguinetti |
Pattern Recognit. | 1 |
| 2010 | Applying the Possibilistic c-Means Algorithm in Kernel-Induced SpacesabstractIn this paper, we study a kernel extension of the classic possibilisticc-means. In the proposed extension, we implicitly map input patterns into a possibly high-dimensional space by means of positive semidefinite kernels. In this new space, we model the mapped data by means of the possibilistic clustering algorithm. We study in more detail the special case where we model the mapped data using a single cluster only, since it turns out to have many interesting properties. The modeled memberships in kernel-induced spaces yield a modeling of generic shapes in the input space. We analyze in detail the connections to one-class support vector machines and kernel density estimation, thus, suggesting that the proposed algorithm can be used in many scenarios of unsupervised learning. In the experimental part, we analyze the stability and the accuracy of the proposed algorithm on some synthetic and real datasets. The results show high stability and good performances in terms of accuracy. Maurizio Filippone, Francesco Masulli, Stefano Rovetta |
IEEE Trans. Fuzzy Syst. | 1 |
| 2009 | Dealing with non-metric dissimilarities in fuzzy central clustering algorithms
Maurizio Filippone |
Int. J. Approx. Reason. | 1 |
| 2009 | Soft ranking in clustering
Stefano Rovetta, Francesco Masulli, Maurizio Filippone |
Neurocomputing | 3 |
| 2009 | A comparative evaluation of nonlinear dynamics methods for time series prediction
Francesco Camastra, Maurizio Filippone |
Neural Comput. Appl. | 2 |
| 2008 | A survey of kernel and spectral methods for clustering
Maurizio Filippone, Francesco Camastra, Francesco Masulli, Stefano Rovetta |
Pattern Recognit. | 1 |
| 2007 | Local Learning of Tide Level Time Series using a Fuzzy ApproachabstractForecasting the tide level in the Venezia lagoon is a very compelling task. In this work we propose a new approach to the learning of tide level time series based on the local learning procedure of Bottou and Vapnik, by considering the use of a fuzzy method for the selection of the closest patterns to the one to forecast. We made use also as learners of Support Vector Machines and of their ensembles based on Bagging and AdaBoost. The obtained forecasts of 500 randomly selected tide levels seem to be quite promising. Good performances are also noticed for forecasts of a set of 80 tide levels corresponding to exceptional periods with high tide and sea variabilities. The obtained forecasts of 80 selected tide levels compare very favorably with those of the baseline linear regressor model. Elio Canestrelli, P. Canestrelli, Marco Corazza, Maurizio Filippone, Silvio Giove, Francesco Masulli |
IJCNN | 4 |
| 2006 | Supervised Classification and Gene Selection Using Simulated AnnealingabstractGenomic data are often characterized by small cardinality and high dimensionality. For those data, a feature selection procedure could highlight the relevant genes and improve the classification results. In this paper we propose a wrapper approach to gene selection in classification of gene expression data using simulated annealing and SVM. The proposed approach can do global combinatorial searches through the space of possible input subsets, can handle cases with numerical, categorical or mixed inputs, and is able to find (sub-)optimal subsets of input variables giving very low classification errors. The method has been tested on the publicly available data sets Leukemia by Golub et al. and Colon by Alon at al. The experimental results highlight the capacity of the method to select minimal sets of relevant genes. Maurizio Filippone, Francesco Masulli, Stefano Rovetta |
IJCNN | 1 |