VLDB 2026 Research / reviewers in the wild / expert
Tom Heskes
dblp:03/1145 · also Tom M. Heskes
· DBLP profile ↗
113ranked-venue papers
24as first author
11since 2021 · last 2025
0000-0002-3398-5235ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 90 · 22 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 7Graphics, computer vision, multimedia, augmented reality and games · 7Theory of computation · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Offline Changepoint Detection With Gaussian ProcessesabstractThis work proposes Segmenting changepoint Gaussian process regression (SegCPGP), an offline changepoint detection method that integrates Gaussian process regression with the changepoint kernel, the likelihood ratio test and binary search. We use the spectral mixture kernel to detect various types of changes without prior knowledge of their type. SegCPGP outperforms state-of-the-art methods when detecting various change types in synthetic datasets; in real world changepoint detection datasets, it performs on par with its competitors. While its hypothesis test shows slight miscalibration, we find SegCPGP remains reasonably reliable. Janneke Verbeek, Tom Heskes, Yuliya Shapovalova |
UAI | 2 |
| 2025 | Confident neural network regression with Bootstrapped Deep Ensembles
Laurens Sluijterman, Eric Cator, Tom Heskes |
Neurocomputing | 3 |
| 2025 | Bias-variance decompositions: the exclusive privilege of Bregman divergencesabstractBias-variance decompositions are widely used to understand the generalization performance of machine learning models. While the squared error loss permits a straightforward decomposition, other loss functions - such as zero-one loss or $L_1$ loss - either fail to sum bias and variance to the expected loss or rely on definitions that lack the essential properties of meaningful bias and variance. Recent research has shown that clean decompositions can be achieved for the broader class of Bregman divergences, with the cross-entropy loss as a special case. However, the necessary and sufficient conditions for these decompositions remain an open question. In this paper, we address this question by studying continuous, nonnegative loss functions that satisfy the identity of indiscernibles (zero loss if and only if the two arguments are identical). We prove that so-called $g$-Bregman or rho-tau divergences are the only such loss functions that have a clean bias-variance decomposition. A $g$-Bregman divergence can be transformed into a standard Bregman divergence through an invertible change of variables. This makes the squared Euclidean distance, up to such a variable transformation, the only symmetric loss function with a clean bias-variance decomposition. Consequently, common metrics such as zero-one and $L_1$ losses cannot admit a clean bias-variance decomposition, explaining why previous attempts have failed. We also examine the impact of relaxing the restrictions on the loss functions and how this affects our results. Tom Heskes |
J. Mach. Learn. Res. | 1 |
| 2025 | Invariant neural architecture for learning term synthesis in instantiation provingabstractContains fulltext : 310648.pdf (Publisher’s version ) (Open Access) Jelle Piepenbrock, Josef Urban, Konstantin Korovin, Miroslav Olsák, Tom Heskes, Mikolás Janota |
J. Symb. Comput. | 5 |
| 2025 | Likelihood-ratio-based confidence intervals for neural networksabstractAbstract This paper introduces a first implementation of a novel likelihood-ratio-based approach for constructing confidence intervals for neural networks. Our method, called DeepLR, offers several qualitative advantages: most notably, the ability to construct asymmetric intervals that expand in regions with a limited amount of data, and the inherent incorporation of factors such as the amount of training time, network architecture, and regularization techniques. While acknowledging that the current implementation of the method is prohibitively expensive for many deep-learning applications, the high cost may already be justified in specific fields like medical predictions or astrophysics, where a reliable uncertainty estimate for a single prediction is essential. This work highlights the significant potential of a likelihood-ratio-based uncertainty estimate and establishes a promising avenue for future research. Laurens Sluijterman, Eric Cator, Tom Heskes |
Mach. Learn. | 3 |
| 2024 | Optimal training of Mean Variance Estimation neural networksabstractThis paper focusses on the optimal implementation of a Mean Variance Estimation network (MVE network) (Nix and Weigend, 1994). This type of network is often used as a building block for uncertainty estimation methods in a regression setting, for instance Concrete dropout (Gal et al., 2017) and Deep Ensembles (Lakshminarayanan et al., 2017). Specifically, an MVE network assumes that the data is produced from a normal distribution with a mean function and variance function. The MVE network outputs a mean and variance estimate and optimizes the network parameters by minimizing the negative loglikelihood. In our paper, we present two significant insights. Firstly, the convergence difficulties reported in recent work can be relatively easily prevented by following the simple yet often overlooked recommendation from the original authors that a warm-up period should be used. During this period, only the mean is optimized with a fixed variance. We demonstrate the effectiveness of this step through experimentation, highlighting that it should be standard practice. As a sidenote, we examine whether, after the warm-up, it is beneficial to fix the mean while optimizing the variance or to optimize both simultaneously. Here, we do not observe a substantial difference. Secondly, we introduce a novel improvement of the MVE network: separate regularization of the mean and the variance estimate. We demonstrate, both on toy examples, multiple benchmark UCI regression data sets, and on the UTKFace data set, that following the original recommendations and the novel separate regularization can lead to significant improvements. Laurens Sluijterman, Eric Cator, Tom Heskes |
Neurocomputing | 3 |
| 2024 | Unsupervised Anomaly Detection Algorithms on Real-world Data: How Many Do We Need?abstractIn this study we evaluate 33 unsupervised anomaly detection algorithms on 52 real-world multivariate tabular data sets, performing the largest comparison of unsupervised anomaly detection algorithms to date. On this collection of data sets, the EIF (Extended Isolation Forest) algorithm significantly outperforms the most other algorithms. Visualizing and then clustering the relative performance of the considered algorithms on all data sets, we identify two clear clusters: one with "local” data sets, and another with "global” data sets. "Local” anomalies occupy a region with low density when compared to nearby samples, while "global” occupy an overall low density region in the feature space. On the local data sets the $k$NN ($k$-nearest neighbor) algorithm comes out on top. On the global data sets, the EIF (extended isolation forest) algorithm performs the best. Also taking into consideration the algorithms' computational complexity, a toolbox with these two unsupervised anomaly detection algorithms suffices for finding anomalies in this representative collection of multivariate data sets. By providing access to code and data sets, our study can be easily reproduced and extended with more algorithms and/or data sets. Roel Bouman, Zaharah Bukhsh, Tom Heskes |
J. Mach. Learn. Res. | 3 |
| 2024 | How to evaluate uncertainty estimates in machine learning for regression?abstractAs neural networks become more popular, the need for accompanying uncertainty estimates increases. There are currently two main approaches to test the quality of these estimates. Most methods output a density. They can be compared by evaluating their loglikelihood on a test set. Other methods output a prediction interval directly. These methods are often tested by examining the fraction of test points that fall inside the corresponding prediction intervals. Intuitively, both approaches seem logical. However, we demonstrate through both theoretical arguments and simulations that both ways of evaluating the quality of uncertainty estimates have serious flaws. Firstly, both approaches cannot disentangle the separate components that jointly create the predictive uncertainty, making it difficult to evaluate the quality of the estimates of these components. Specifically, the quality of a confidence interval cannot reliably be tested by estimating the performance of a prediction interval. Secondly, the loglikelihood does not allow a comparison between methods that output a prediction interval directly and methods that output a density. A better loglikelihood also does not necessarily guarantee better prediction intervals, which is what the methods are often used for in practice. Moreover, the current approach to test prediction intervals directly has additional flaws. We show why testing a prediction or confidence interval on a single test set is fundamentally flawed. At best, marginal coverage is measured, implicitly averaging out overconfident and underconfident predictions. A much more desirable property is pointwise coverage, requiring the correct coverage for each prediction. We demonstrate through practical examples that these effects can result in favouring a method, based on the predictive uncertainty, that has undesirable behaviour of the confidence or prediction intervals. Finally, we propose a simulation-based testing approach that addresses these problems while still allowing easy comparison between different methods. This approach can be used for the development of new uncertainty quantification methods. Laurens Sluijterman, Eric Cator, Tom Heskes |
Neural Networks | 3 |
| 2023 | Automatic Inference of Fault Tree Models Via Multi-Objective Evolutionary AlgorithmsabstractFault tree analysis is a well-known technique in reliability engineering and risk assessment, which supports decision-making processes and the management of complex systems. Traditionally, fault tree (FT) models are built manually together with domain experts, considered a time-consuming process prone to human errors. With Industry 4.0, there is an increasing availability of inspection and monitoring data, making techniques that enable knowledge extraction from large data sets relevant. Thus, our goal with this work is to propose a data-driven approach to infer efficient FT structures that achieve a complete representation of the failure mechanisms contained in the failure data set without human intervention. Our algorithm, the FT-MOEA, based on multi-objective evolutionary algorithms, enables the simultaneous optimization of different relevant metrics such as the FT size, the error computed based on the failure data set and the Minimal Cut Sets. Our results show that, for six case studies from the literature, our approach successfully achieved automatic, efficient, and consistent inference of the associated FT models. We also present the results of a parametric analysis that tests our algorithm for different relevant conditions that influence its performance, as well as an overview of the data-driven methods used to automatically infer FT models. Lisandro Arturo Jimenez-Roa, Tom Heskes, Tiedo Tinga, Mariëlle Stoelinga |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2022 | Non-parametric synergy modeling of chemical compounds with Gaussian processesabstractBACKGROUND: Understanding the synergetic and antagonistic effects of combinations of drugs and toxins is vital for many applications, including treatment of multifactorial diseases and ecotoxicological monitoring. Synergy is usually assessed by comparing the response of drug combinations to a predicted non-interactive response from reference (null) models. Possible choices of null models are Loewe additivity, Bliss independence and the recently rediscovered Hand model. A different approach is taken by the MuSyC model, which directly fits a generalization of the Hill model to the data. All of these models, however, fit the dose-response relationship with a parametric model. RESULTS: We propose the Hand-GP model, a non-parametric model based on the combination of the Hand model with Gaussian processes. We introduce a new logarithmic squared exponential kernel for the Gaussian process which captures the logarithmic dependence of response on dose. From the monotherapeutic response and the Hand principle, we construct a null reference response and synergy is assessed from the difference between this null reference and the Gaussian process fitted response. Statistical significance of the difference is assessed from the confidence intervals of the Gaussian process fits. We evaluate performance of our model on a simulated data set from Greco, two simulated data sets of our own design and two benchmark data sets from Chou and Talalay. We compare the Hand-GP model to standard synergy models and show that our model performs better on these data sets. We also compare our model to the MuSyC model as an example of a recent method on these five data sets and on two-drug combination screens: Mott et al. anti-malarial screen and O'Neil et al. anti-cancer screen. We identify cases in which the HandGP model is preferred and cases in which the MuSyC model is preferred. CONCLUSION: The Hand-GP model is a flexible model to capture synergy. Its non-parametric and probabilistic nature allows it to model a wide variety of response patterns. Yuliya Shapovalova, Tom Heskes, Tjeerd Dijkstra |
BMC Bioinform. | 2 |
| 2021 | Probabilistic Modelling of Gait for Robust Passive Monitoring in Daily LifeabstractPassive monitoring in daily life may provide valuable insights into a person's health throughout the day. Wearable sensor devices play a key role in enabling such monitoring in a non-obtrusive fashion. However, sensor data collected in daily life reflect multiple health and behavior-related factors together. This creates the need for a structured principled analysis to produce reliable and interpretable predictions that can be used to support clinical diagnosis and treatment. In this work we develop a principled modelling approach for free-living gait (walking) analysis. Gait is a promising target for non-obtrusive monitoring because it is common and indicative of many different movement disorders such as Parkinson's disease (PD), yet its analysis has largely been limited to experimentally controlled lab settings. To locate and characterize stationary gait segments in free-living using accelerometers, we present an unsupervised probabilistic framework designed to segment signals into differing gait and non-gait patterns. We evaluate the approach using a new video-referenced dataset including 25 PD patients with motor fluctuations and 25 age-matched controls, performing unscripted daily living activities in and around their own houses. Using this dataset, we demonstrate the framework's ability to detect gait and predict medication induced fluctuations in PD patients based on free-living gait. We show that our approach is robust to varying sensor locations, including the wrist, ankle, trouser pocket and lower back. Yordan P. Raykov, Luc J. W. Evers, Reham Badawy, Bastiaan R. Bloem, Tom Heskes, Marjan J. Meinders, Kasper Claes, Max A. Little |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Causal Shapley Values: Exploiting Causal Knowledge to Explain Individual Predictions of Complex ModelsabstractShapley values underlie one of the most popular model-agnostic methods within explainable artificial intelligence. These values are designed to attribute the difference between a model's prediction and an average baseline to the different features used as input to the model. Being based on solid game-theoretic principles, Shapley values uniquely satisfy several desirable properties, which is why they are increasingly used to explain the predictions of possibly complex and highly non-linear machine learning models. Shapley values are well calibrated to a user’s intuition when features are independent, but may lead to undesirable, counterintuitive explanations when the independence assumption is violated. In this paper, we propose a novel framework for computing Shapley values that generalizes recent work that aims to circumvent the independence assumption. By employing Pearl's do-calculus, we show how these `causal' Shapley values can be derived for general causal graphs without sacrificing any of their desirable properties. Moreover, causal Shapley values enable us to separate the contribution of direct and indirect effects. We provide a practical implementation for computing causal Shapley values based on causal chain graphs when only partial information is available and illustrate their utility on a real-world example. Tom Heskes, E. M. C. Sijben, Ioan Gabriel Bucur, Tom Claassen |
NeurIPS | 1 |
| 2020 | MASSIVE: Tractable and Robust Bayesian Learning of Many-Dimensional Instrumental Variable ModelsabstractThe recent availability of huge, many-dimensional data sets, like those arising from genome-wide association studies (GWAS), provides many opportunities for strengthening causal inference. One popular approach is to utilize these many-dimensional measurements as instrumental variables (instruments) for improving the causal effect estimate between other pairs of variables. Unfortunately, searching for proper instruments in a many-dimensional set of candidates is a daunting task due to the intractable model space and the fact that we cannot directly test which of these candidates are valid, so most existing search methods either rely on overly stringent modeling assumptions or fail to capture the inherent model uncertainty in the selection process. We show that, as long as at least some of the candidates are (close to) valid, without knowing a priori which ones, they collectively still pose enough restrictions on the target interaction to obtain a reliable causal effect estimate. We propose a general and efficient causal inference algorithm that accounts for model uncertainty by performing Bayesian model averaging over the most promising many-dimensional instrumental variable models, while at the same time employing weaker assumptions regarding the data generating process. We showcase the efficiency, robustness and predictive performance of our algorithm through experimental results on both simulated and real-world data. Ioan Gabriel Bucur, Tom Claassen, Tom Heskes |
UAI | 3 |
| 2019 | Large-scale local causal inference of gene regulatory relationships
Ioan Gabriel Bucur, Tom Claassen, Tom Heskes |
Int. J. Approx. Reason. | 3 |
| 2019 | Hierarchical Bayesian inference for concurrent model fitting and comparison for group studiesabstractComputational modeling plays an important role in modern neuroscience research. Much previous research has relied on statistical methods, separately, to address two problems that are actually interdependent. First, given a particular computational model, Bayesian hierarchical techniques have been used to estimate individual variation in parameters over a population of subjects, leveraging their population-level distributions. Second, candidate models are themselves compared, and individual variation in the expressed model estimated, according to the fits of the models to each subject. The interdependence between these two problems arises because the relevant population for estimating parameters of a model depends on which other subjects express the model. Here, we propose a hierarchical Bayesian inference (HBI) framework for concurrent model comparison, parameter estimation and inference at the population level, combining previous approaches. We show that this framework has important advantages for both parameter estimation and model comparison theoretically and experimentally. The parameters estimated by the HBI show smaller errors compared to other methods. Model comparison by HBI is robust against outliers and is not biased towards overly simplistic models. Furthermore, the fully Bayesian approach of our theory enables researchers to make inference on group-level parameters by performing HBI t-test. Payam Piray, Amir Dezfouli, Tom Heskes, Michael J. Frank, Nathaniel D. Daw |
PLoS Comput. Biol. | 3 |
| 2019 | Stable Specification Search in Structural Equation Models with Latent VariablesabstractIn our previous study, we introduced stable specification search for cross-sectional data (S3C). It is an exploratory causal method that combines the concept of stability selection and multi-objective optimization to search for stable and parsimonious causal structures across the entire range of model complexities. S3C, however, is designed to model causal relations among observed variables. In this study, we extended S3C to S3C-Latent, to model linear causal relations between latent variables that are measured through observed proxies. We evaluated S3C-Latent on simulated data and compared the results to those of PC-MIMBuild, an extension of the PC algorithm, the state-of-the-art causal discovery method. The comparison shows that S3C-Latent achieved better performance. We also applied S3C-Latent to real-world data of children with attention deficit/hyperactivity disorder and data about measuring mental abilities among pupils. The results are consistent with those of previous studies. Ridho Rahmadi, Perry Groot, Tom Heskes |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2018 | Learning the Causal Structure of Copula Models with Latent Variables
Ruifei Cui, Perry Groot, Moritz Schauer, Tom Heskes |
UAI | 4 |
| 2018 | Bayesian data integration for quantifying the contribution of diverse measurements to parameter estimatesabstractMotivation: Computational models in biology are frequently underdetermined, due to limits in our capacity to measure biological systems. In particular, mechanistic models often contain parameters whose values are not constrained by a single type of measurement. It may be possible to achieve better model determination by combining the information contained in different types of measurements. Bayesian statistics provides a convenient framework for this, allowing a quantification of the reduction in uncertainty with each additional measurement type. We wished to explore whether such integration is feasible and whether it can allow computational models to be more accurately determined. Results: We created an ordinary differential equation model of cell cycle regulation in budding yeast and integrated data from 13 different studies covering different experimental techniques. We found that for some parameters, a single type of measurement, relative time course mRNA expression, is sufficient to constrain them. Other parameters, however, were only constrained when two types of measurements were combined, namely relative time course and absolute transcript concentration. Comparing the estimates to measurements from three additional, independent studies, we found that the degradation and transcription rates indeed matched the model predictions in order of magnitude. The predicted translation rate was incorrect however, thus revealing a deficiency in the model. Since this parameter was not constrained by any of the measurement types separately, it was only possible to falsify the model when integrating multiple types of measurements. In conclusion, this study shows that integrating multiple measurement types can allow models to be more accurately determined. Availability and implementation: The models and files required for running the inference are included in the Supplementary information. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Bram Thijssen, Tjeerd Dijkstra, Tom Heskes, Lodewyk F. A. Wessels |
Bioinform. | 3 |
| 2018 | The stablespec package for causal discovery on cross-sectional and longitudinal data in RabstractThe R package stablespec is an implementation of our method stable specification search . The method aims at causal discovery on both cross-sectional and longitudinal data through stable specification search in constrained structural equation models . Ridho Rahmadi, Perry Groot, Tom Heskes |
Neurocomputing | 3 |
| 2018 | A scalable preference model for autonomous decision-makingabstractEmerging domains such as smart electric grids require decisions to be made autonomously, based on the observed behaviors of large numbers of connected consumers. Existing approaches either lack the flexibility to capture nuanced, individualized preference profiles, or scale poorly with the size of the dataset. We propose a preference model that combines flexible Bayesian nonparametric priors—providing state-of-the-art predictive power—with well-justified structural assumptions that allow a scalable implementation. The Gaussian process scalable preference model via Kronecker factorization ( GaSPK ) model provides accurate choice predictions and principled uncertainty estimates as input to decision-making tasks. In consumer choice settings where alternatives are described by few key attributes, inference in our model is highly efficient and scalable to tens of thousands of choices. Markus Peters, Maytal Saar-Tsechansky, Wolfgang Ketter, Sinead Williamson, Perry Groot, Tom Heskes |
Mach. Learn. | 6 |
| 2017 | Robust Causal Estimation in the Large-Sample Limit without Strict FaithfulnessabstractCausal effect estimation from observational data is an important and much studied research topic. The instrumental variable (IV) and local causal discovery (LCD) patterns are canonical examples of settings where a closed-form expression exists for the causal effect of one variable on another, given the presence of a third variable. Both rely on faithfulness to infer that the latter only influences the target effect via the cause variable. In reality, it is likely that this assumption only holds approximately and that there will be at least some form of weak interaction. This brings about the paradoxical situation that, in the large-sample limit, no predictions are made, as detecting the weak edge invalidates the setting. We introduce an alternative approach by replacing strict faithfulness with a prior that reflects the existence of many ’weak’ (irrelevant) and ’strong’ interactions. We obtain a posterior distribution over the target causal effect estimator which shows that, in many cases, we can still make good estimates. We demonstrate the approach in an application on a simple linear-Gaussian setting, using the MultiNest sampling algorithm, and compare it with established techniques to show our method is robust even when strict faithfulness is violated. Ioan Gabriel Bucur, Tom Claassen, Tom Heskes |
AISTATS | 3 |
| 2017 | Robust Estimation of Gaussian Copula Causal Structure from Mixed Data with Missing ValuesabstractWe consider the problem of causal structure learning from data with missing values, assumed to be drawn from a Gaussian copula model. First, we extend the 'Rank PC' algorithm, designed for Gaussian copula models with purely continuous data (so-called nonparanormal models), to incomplete data by applying rank correlation to pairwise complete observations and replacing the sample size with an effective sample size in the conditional independence tests to account for the information loss from missing values. The resulting approach works when the data are missing completely at random (MCAR). Then, we propose a Gibbs sampling procedure to draw correlation matrix samples from mixed data under missingness at random (MAR). These samples are translated into an average correlation matrix, and an effective sample size, resulting in the 'Copula PC' algorithm for incomplete data. Simulation study shows that: 1) the usage of the effective sample size significantly improves the performance of 'Rank PC' and 'Copula PC'; 2) 'Copula PC' estimates a more accurate correlation matrix and causal structure than 'Rank PC' under MCAR and, even more so, under MAR. Also, we illustrate our methods on a real-world data set about gene expression. Ruifei Cui, Perry Groot, Tom Heskes |
ICDM | 3 |
| 2017 | RankProd 2.0: a refactored bioconductor package for detecting differentially expressed features in molecular profiling datasetsabstractMOTIVATION: The Rank Product (RP) is a statistical technique widely used to detect differentially expressed features in molecular profiling experiments such as transcriptomics, metabolomics and proteomics studies. An implementation of the RP and the closely related Rank Sum (RS) statistics has been available in the RankProd Bioconductor package for several years. However, several recent advances in the understanding of the statistical foundations of the method have made a complete refactoring of the existing package desirable. RESULTS: We implemented a completely refactored version of the RankProd package, which provides a more principled implementation of the statistics for unpaired datasets. Moreover, the permutation-based P -value estimation methods have been replaced by exact methods, providing faster and more accurate results. AVAILABILITY AND IMPLEMENTATION: RankProd 2.0 is available at Bioconductor ( https://www.bioconductor.org/packages/devel/bioc/html/RankProd.html ) and as part of the mzMatch pipeline ( http://www.mzmatch.sourceforge.net ). CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Francesco Del Carratore, Andris Jankevics, Rob Eisinga, Tom Heskes, Fangxin Hong, Rainer Breitling |
Bioinform. | 4 |
| 2017 | Exact p-values for pairwise comparison of Friedman rank sums, with application to comparing classifiersabstractBACKGROUND: The Friedman rank sum test is a widely-used nonparametric method in computational biology. In addition to examining the overall null hypothesis of no significant difference among any of the rank sums, it is typically of interest to conduct pairwise comparison tests. Current approaches to such tests rely on large-sample approximations, due to the numerical complexity of computing the exact distribution. These approximate methods lead to inaccurate estimates in the tail of the distribution, which is most relevant for p-value calculation. RESULTS: We propose an efficient, combinatorial exact approach for calculating the probability mass distribution of the rank sum difference statistic for pairwise comparison of Friedman rank sums, and compare exact results with recommended asymptotic approximations. Whereas the chi-squared approximation performs inferiorly to exact computation overall, others, particularly the normal, perform well, except for the extreme tail. Hence exact calculation offers an improvement when small p-values occur following multiple testing correction. Exact inference also enhances the identification of significant differences whenever the observed values are close to the approximate critical value. We illustrate the proposed method in the context of biological machine learning, were Friedman rank sum difference tests are commonly used for the comparison of classifiers over multiple datasets. CONCLUSIONS: We provide a computationally fast method to determine the exact p-value of the absolute rank sum difference of a pair of Friedman rank sums, making asymptotic tests obsolete. Calculation of exact p-values is easy to implement in statistical software and the implementation in R is provided in one of the Additional files and is also available at http://www.ru.nl/publish/pages/726696/friedmanrsd.zip . Rob Eisinga, Tom Heskes, Ben Pelzer, Manfred te Grotenhuis |
BMC Bioinform. | 2 |
| 2017 | Multi-Domain Transfer Component Analysis for Domain Generalization
Thomas Grubinger, Adriana Birlutiu, Holger Schöner, Thomas Natschläger, Tom Heskes |
Neural Process. Lett. | 5 |
| 2016 | Exploring Constraint: Simulating Self-Organization and Autogenesis in the Autogenic AutomatonabstractAbstract Origin of life theories often argue that molecular self-organization explains the spontaneous emergence of structural and dynamical constraints. How-ever, the preservation of these constraints over time is not well-explained because of the self-undermining and self-limiting nature of these same processes. A pro-cess called autogenesis has been proposed in which negative structural coupling between self-organized processes preserves the constraints thereby accumulated. This paper presents a computer simulation of this process (the Autogenic Automa-ton) and compares its behavior to the same self-organizing processes when uncou-pled. We demonstrate that this coupling produces a second-order constraint that can both resist dissipation and become replicated in new substrates over time. Terrence W. Deacon, Tom Heskes, Stefan Leijnen |
ALIFE | 2 |
| 2016 | Causal Discovery from Big Data - Mission (Im)possible?
Tom Heskes |
ICAART (1) | 1 |
| 2016 | Copula PC Algorithm for Causal Discovery from Mixed Data
Ruifei Cui, Perry Groot, Tom Heskes |
ECML/PKDD (2) | 3 |
| 2015 | Causal Discovery from Medical Data: Dealing with Missing Values and a Mixture of Discrete and Continuous Data
Elena Sokolova, Perry Groot, Tom Claassen, Daniel von Rhein, Jan K. Buitelaar, Tom Heskes |
AIME | 6 |
| 2015 | KeCo: Kernel-Based Online Co-agreement Algorithm
Laurens Wiel, Tom Heskes, Evgeni Levin |
Discovery Science | 2 |
| 2015 | Batch Steepest-Descent-Mildest-Ascent for Interactive Maximum Margin Clustering
Fabian Gieseke, Tapio Pahikkala, Tom Heskes |
IDA | 3 |
| 2015 | Bayesian Estimation of Conditional Independence Graphs Improves Functional Connectivity EstimatesabstractFunctional connectivity concerns the correlated activity between neuronal populations in spatially segregated regions of the brain, which may be studied using functional magnetic resonance imaging (fMRI). This coupled activity is conveniently expressed using covariance, but this measure fails to distinguish between direct and indirect effects. A popular alternative that addresses this issue is partial correlation, which regresses out the signal of potentially confounding variables, resulting in a measure that reveals only direct connections. Importantly, provided the data are normally distributed, if two variables are conditionally independent given all other variables, their respective partial correlation is zero. In this paper, we propose a probabilistic generative model that allows us to estimate functional connectivity in terms of both partial correlations and a graph representing conditional independencies. Simulation results show that this methodology is able to outperform the graphical LASSO, which is the de facto standard for estimating partial correlations. Furthermore, we apply the model to estimate functional connectivity for twenty subjects using resting-state fMRI data. Results show that our model provides a richer representation of functional connectivity as compared to considering partial correlations alone. Finally, we demonstrate how our approach can be extended in several ways, for instance to achieve data fusion by informing the conditional independence graph with data from probabilistic tractography. As our Bayesian formulation of functional connectivity provides access to the posterior distribution instead of only to point estimates, we are able to quantify the uncertainty associated with our results. This reveals that while we are able to infer a clear backbone of connectivity in our empirical results, the data are not accurately described by simply looking at the mode of the distribution over connectivity. The implication of this is that deterministic alternatives may misjudge connectivity results by drawing conclusions from noisy and limited data. Max Hinne, Ronald J. Janssen, Tom Heskes, Marcel van Gerven |
PLoS Comput. Biol. | 3 |
| 2015 | MAGMA: Generalized Gene-Set Analysis of GWAS DataabstractBy aggregating data for complex traits in a biologically meaningful way, gene and gene-set analysis constitute a valuable addition to single-marker analysis. However, although various methods for gene and gene-set analysis currently exist, they generally suffer from a number of issues. Statistical power for most methods is strongly affected by linkage disequilibrium between markers, multi-marker associations are often hard to detect, and the reliance on permutation to compute p-values tends to make the analysis computationally very expensive. To address these issues we have developed MAGMA, a novel tool for gene and gene-set analysis. The gene analysis is based on a multiple regression model, to provide better statistical performance. The gene-set analysis is built as a separate layer around the gene analysis for additional flexibility. This gene-set analysis also uses a regression structure to allow generalization to analysis of continuous properties of genes and simultaneous analysis of multiple gene sets and other gene properties. Simulations and an analysis of Crohn's Disease data are used to evaluate the performance of MAGMA and to compare it to a number of other gene and gene-set analysis tools. The results show that MAGMA has significantly more power than other tools for both the gene and the gene-set analysis, identifying more genes and gene sets associated with Crohn's Disease while maintaining a correct type 1 error rate. Moreover, the MAGMA analysis of the Crohn's Disease data was found to be considerably faster as well. Christiaan A. de Leeuw, Joris M. Mooij, Tom Heskes, Danielle Posthuma |
PLoS Comput. Biol. | 3 |
| 2015 | A Bayesian Framework for Combining Protein and Network Topology Information for Predicting Protein-Protein InteractionsabstractComputational methods for predicting protein-protein interactions are important tools that can complement high-throughput technologies and guide biologists in designing new laboratory experiments. The proteins and the interactions between them can be described by a network which is characterized by several topological properties. Information about proteins and interactions between them, in combination with knowledge about topological properties of the network, can be used for developing computational methods that can accurately predict unknown protein-protein interactions. This paper presents a supervised learning framework based on Bayesian inference for combining two types of information: i) network topology information, and ii) information related to proteins and the interactions between them. The motivation of our model is that by combining these two types of information one can achieve a better accuracy in predicting protein-protein interactions, than by using models constructed from these two types of information independently. Adriana Birlutiu, Florence d'Alché-Buc, Tom Heskes |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2014 | Motion history images for online speaker/signer diarizationabstractWe present a solution to the problem of online speaker/signer diarization - the task of determining who spoke/signed when?. Our solution is based on the idea that gestural activity (hands and body movement) is highly correlated with uttering activity. This correlation is necessarily true for sign languages and mostly true for spoken languages. The novel part of our solution is the use of motion history images (MHI) as a likelihood measure for probabilistically detecting uttering activities. MHI is an efficient representation of where and how motion occurred for a fixed period of time. We conducted experiments on 4.9 hours of the AMI meeting data and 1.4 hours of sign language dataset (Kata Kolok data). The best performance obtained is 15.70% for sign language and 31.90% for spoken language (measurements are in DER). These results show that our solution is applicable in real-world applications like video conferences and information retrieval. Binyam Gebrekidan Gebre, Peter Wittenburg, Tom Heskes, Sebastian Drude |
ICASSP | 3 |
| 2014 | Mutual Information Estimation with Random Forests
Mike Koeman, Tom Heskes |
ICONIP (2) | 2 |
| 2014 | Speaker diarization using gesture and speechabstractContains fulltext : 134621.pdf (Author’s version preprint ) (Open Access) Binyam Gebrekidan Gebre, Peter Wittenburg, Sebastian Drude, Marijn Huijbregts, Tom Heskes |
INTERSPEECH | 5 |
| 2014 | A comparative study of cell classifiers for image-based high-throughput screeningabstractBACKGROUND: Millions of cells are present in thousands of images created in high-throughput screening (HTS). Biologists could classify each of these cells into a phenotype by visual inspection. But in the presence of millions of cells this visual classification task becomes infeasible. Biologists train classification models on a few thousand visually classified example cells and iteratively improve the training data by visual inspection of the important misclassified phenotypes. Classification methods differ in performance and performance evaluation time. We present a comparative study of computational performance of gentle boosting, joint boosting CellProfiler Analyst (CPA), support vector machines (linear and radial basis function) and linear discriminant analysis (LDA) on two data sets of HT29 and HeLa cancer cells. RESULTS: For the HT29 data set we find that gentle boosting, SVM (linear) and SVM (RBF) are close in performance but SVM (linear) is faster than gentle boosting and SVM (RBF). For the HT29 data set the average performance difference between SVM (RBF) and SVM (linear) is 0.42 %. For the HeLa data set we find that SVM (RBF) outperforms other classification methods and is on average 1.41 % better in performance than SVM (linear). CONCLUSIONS: Our study proposes SVM (linear) for iterative improvement of the training data and SVM (RBF) for the final classifier to classify all unlabeled cells in the whole data set. Syed Abbas, Tjeerd Dijkstra, Tom Heskes |
BMC Bioinform. | 3 |
| 2014 | A fast algorithm for determining bounds and accurate approximate p-values of the rank product statistic for replicate experimentsabstractBACKGROUND: The rank product method is a powerful statistical technique for identifying differentially expressed molecules in replicated experiments. A critical issue in molecule selection is accurate calculation of the p-value of the rank product statistic to adequately address multiple testing. Both exact calculation and permutation and gamma approximations have been proposed to determine molecule-level significance. These current approaches have serious drawbacks as they are either computationally burdensome or provide inaccurate estimates in the tail of the p-value distribution. RESULTS: We derive strict lower and upper bounds to the exact p-value along with an accurate approximation that can be used to assess the significance of the rank product statistic in a computationally fast manner. The bounds and the proposed approximation are shown to provide far better accuracy over existing approximate methods in determining tail probabilities, with the slightly conservative upper bound protecting against false positives. We illustrate the proposed method in the context of a recently published analysis on transcriptomic profiling performed in blood. CONCLUSIONS: We provide a method to determine upper bounds and accurate approximate p-values of the rank product statistic. The proposed algorithm provides an order of magnitude increase in throughput as compared with current approaches and offers the opportunity to explore new application domains with even larger multiple testing issue. The R code is published in one of the Additional files and is available at http://www.ru.nl/publish/pages/726696/rankprodbounds.zip . Tom Heskes, Rob Eisinga, Rainer Breitling |
BMC Bioinform. | 1 |
| 2014 | Premise Selection for Mathematics by Corpus Analysis and Kernel Methods
Jesse Alama, Tom Heskes, Daniel Kühlwein, Evgeni Tsivtsivadze, Josef Urban |
J. Autom. Reason. | 2 |
| 2013 | The gesturer is the speakerabstractWe present and solve the speaker diarization problem in a novel way. We hypothesize that the gesturer is the speaker and that identifying the gesturer can be taken as identifying the active speaker. We provide evidence in support of the hypothesis from gesture literature and audio-visual synchrony studies. We also present a vision-only diarization algorithm that relies on gestures (i.e. upper body movements). Experiments carried out on 8.9 hours of a publicly available dataset (the AMI meeting data) show that diarization error rates as low as 15% can be achieved. Binyam Gebrekidan Gebre, Peter Wittenburg, Tom Heskes |
ICASSP | 3 |
| 2013 | Automatic sign language identificationabstractWe propose a Random-Forest based sign language identification system. The system uses low-level visual features and is based on the hypothesis that sign languages have varying distributions of phonemes (hand-shapes, locations and movements). We evaluated the system on two sign languages - British SL and Greek SL, both taken from a publicly available corpus, called Dicta Sign Corpus. Achieved average F1 scores are about 95% - indicating that sign languages can be identified with high accuracy using only low-level visual features. Binyam Gebrekidan Gebre, Peter Wittenburg, Tom Heskes |
ICIP | 3 |
| 2013 | Bayesian Probabilities for Constraint-Based Causal Discovery
Tom Claassen, Tom Heskes |
IJCAI | 2 |
| 2013 | Learning Sparse Causal Models is not NP-hard
Tom Claassen, Joris M. Mooij, Tom Heskes |
UAI | 3 |
| 2013 | Cyclic Causal Discovery from Continuous Equilibrium Data
Joris M. Mooij, Tom Heskes |
UAI | 2 |
| 2013 | Efficiently learning the preferences of peopleabstractThis paper presents a framework for optimizing the preference learning process. In many real-world applications in which preference learning is involved the available training data is scarce and obtaining labeled training data is expensive. Fortunately in many of the preference learning situations data is available from multiple subjects. We use the multi-task formalism to enhance the individual training data by making use of the preference information learned from other subjects. Furthermore, since obtaining labels is expensive, we optimally choose which data to ask a subject for labelling to obtain the most of information about her/his preferences. This paradigm—called active learning—has hardly been studied in a multi-task formalism. We propose an alternative for the standard criteria in active learning which actively chooses queries by making use of the available preference data from other subjects. The advantage of this alternative is the reduced computation costs and reduced time subjects are involved. We validate empirically our approach on three real-world data sets involving the preferences of people. Adriana Birlutiu, Perry Groot, Tom Heskes |
Mach. Learn. | 3 |
| 2013 | Bayesian Sparse Partial Least SquaresabstractPartial least squares (PLS) is a class of methods that makes use of a set of latent or unobserved variables to model the relation between (typically) two sets of input and output variables, respectively. Several flavors, depending on how the latent variables or components are computed, have been developed over the last years. In this letter, we propose a Bayesian formulation of PLS along with some extensions. In a nutshell, we provide sparsity at the input space level and an automatic estimation of the optimal number of latent components. We follow the variational approach to infer the parameter distributions. We have successfully tested the proposed methods on a synthetic data benchmark and on electrocorticogram data associated with several motor outputs in monkeys. Diego Vidaurre, Marcel van Gerven, Concha Bielza, Pedro Larrañaga, Tom Heskes |
Neural Comput. | 5 |
| 2012 | Online Co-regularized Algorithms
Tom de Ruijter, Evgeni Tsivtsivadze, Tom Heskes |
Discovery Science | 3 |
| 2012 | A Bayesian Approach to Constraint Based Causal Inference
Tom Claassen, Tom Heskes |
UAI | 2 |
| 2012 | Molecular Machines in the Synapse: Overlapping Protein Sets Control Distinct Steps in NeurosecretionabstractActivity regulated neurotransmission shapes the computational properties of a neuron and involves the concerted action of many proteins. Classical, intuitive working models often assign specific proteins to specific steps in such complex cellular processes, whereas modern systems theories emphasize more integrated functions of proteins. To test how often synaptic proteins participate in multiple steps in neurotransmission we present a novel probabilistic method to analyze complex functional data from genetic perturbation studies on neuronal secretion. Our method uses a mixture of probabilistic principal component analyzers to cluster genetic perturbations on two distinct steps in synaptic secretion, vesicle priming and fusion, and accounts for the poor standardization between different studies. Clustering data from 121 perturbations revealed that different perturbations of a given protein are often assigned to different steps in the release process. Furthermore, vesicle priming and fusion are inversely correlated for most of those perturbations where a specific protein domain was mutated to create a gain-of-function variant. Finally, two different modes of vesicle release, spontaneous and action potential evoked release, were affected similarly by most perturbations. This data suggests that the presynaptic protein network has evolved as a highly integrated supramolecular machine, which is responsible for both spontaneous and activity induced release, with a group of core proteins using different domains to act on multiple steps in the release process. L. Niels Cornelisse, Evgeni Tsivtsivadze, Marieke Meijer, Tjeerd Dijkstra, Tom Heskes, Matthijs Verhage |
PLoS Comput. Biol. | 5 |
| 2011 | A structure independent algorithm for causal discovery
Tom Claassen, Tom Heskes |
ESANN | 2 |
| 2011 | Learning of causal relations
John A. Quinn, Joris M. Mooij, Tom Heskes, Michael Biehl |
ESANN | 3 |
| 2011 | A Markov Random Field Approach to Neural Encoding and Decoding
Marcel van Gerven, Eric Maris, Tom Heskes |
ICANN (2) | 3 |
| 2011 | Learning from Multiple Annotators with Gaussian Processes
Perry Groot, Adriana Birlutiu, Tom Heskes |
ICANN (2) | 3 |
| 2011 | On Causal Discovery with Cyclic Additive Noise ModelsabstractWe study a particular class of cyclic causal models, where each variable is a (possibly nonlinear) function of its parents and additive noise. We prove that the causal graph of such models is generically identifiable in the bivariate, Gaussian-noise case. We also propose a method to learn such models from observational data. In the acyclic case, the method reduces to ordinary regression, but in the more challenging cyclic case, an additional term arises in the loss function, which makes it a special case of nonlinear independent component analysis. We illustrate the proposed method on synthetic data. Joris M. Mooij, Dominik Janzing, Tom Heskes, Bernhard Schölkopf |
NIPS | 3 |
| 2011 | Semantic Graph Kernels for Automated ReasoningabstractLearning reasoning techniques from previous knowledge is a largely underdeveloped area of automated reasoning. As large bodies of formal knowledge are becoming available to automated reasoners, state-of-the-art machine learning methods can provide powerful heuristics for problem-specific detection of relevant knowledge contained in the libraries. In this paper we develop a semantic graph kernel suitable for learning in structured mathematical domains. Our kernel incorporates contextual information about the features and unlike “random walk”-based graph kernels it is also applicable to sparse graphs. We evaluate the proposed semantic graph kernel on a subset of the large formal Mizar mathematical library. Our empirical evaluation demonstrates that graph kernels in general are particularly suitable for the automated reasoning domain and that in many cases our semantic graph kernel leads to improvement in performance compared to linear, Gaussian, latent semantic, and geometric graph kernels. Evgeni Tsivtsivadze, Josef Urban, Herman Geuvers, Tom Heskes |
SDM | 4 |
| 2011 | A Logical Characterization of Constraint-Based Causal Discovery
Tom Claassen, Tom Heskes |
UAI | 2 |
| 2011 | Properties of Bethe Free Energies and Message Passing in Gaussian ModelsabstractWe address the problem of computing approximate marginals in Gaussian probabilistic models by using mean field and fractional Bethe approximations. We define the Gaussian fractional Bethe free energy in terms of the moment parameters of the approximate marginals, derive a lower and an upper bound on the fractional Bethe free energy and establish a necessary condition for the lower bound to be bounded from below. It turns out that the condition is identical to the pairwise normalizability condition, which is known to be a sufficient condition for the convergence of the message passing algorithm. We show that stable fixed points of the Gaussian message passing algorithm are local minima of the Gaussian Bethe free energy. By a counterexample, we disprove the conjecture stating that the unboundedness of the free energy implies the divergence of the message passing algorithm. Botond Cseke, Tom Heskes |
J. Artif. Intell. Res. | 2 |
| 2011 | Approximate Marginals in Latent Gaussian Models
Botond Cseke, Tom Heskes |
J. Mach. Learn. Res. | 2 |
| 2011 | Predicting Preference Judgments of Individual Normal and Hearing-Impaired Listeners With Gaussian ProcessesabstractA probabilistic kernel approach to pairwise preference learning based on Gaussian processes is applied to predict preference judgments for sound quality degradation mechanisms that might be present in a hearing aid. Subjective sound quality comparisons for 14 normal-hearing and 18 hearing-impaired subjects were used for evaluating the predictive performance. Stimuli were sentences subjected to three kinds of distortion (additive noise, peak clipping, and center clipping) with eight levels of degradation for each distortion type. The kernel approach gives a significant improvement in preference predictions of hearing-impaired subjects by individualizing the learning process. A significant difference is shown between normal-hearing and hearing-impaired subjects, because of nonlinearities in the perception of hearing-impaired subjects. In particular, hearing-impaired subjects have significant nonlinear preference judgments when making pairwise comparisons between peak clipped sentences with different clipping thresholds. The probabilistic kernel approach is shown to be robust when generalizing over distortions and over subjects. Perry Groot, Tom Heskes, Tjeerd Dijkstra, James M. Kates |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Editorial: One Year as EiC, and Editorial-Board Changes at TNNabstractIAM ABOUT to start my second year of service as the Editor-in-Chief (EiC) of the IEEE TRANSACTIONS ON NEURAL NETWORKS (TNN). Needless to say, my first year as the EiC has been full of excitement and challenges. Transitioning this position from my predecessor to me went very smoothly during the months of September 2009 to January 2010. During the past year, we have accumulated 50+ Associate Editors (AEs) handling roughly 600 new submissions (not counting resubmissions and revised submissions). With the help of these AEs and my predecessor, I was quickly able to learn to do my job, and as such, the transition had very few glitches. The easy part of my job is checking whether a submission is in compliance with our guidelines and where it is within the scope of the TRANSACTIONS, before it is assigned to an AE for handling. The difficult part of my job has been dealing with some papers with three or more reviewers, all of whom agreed to review them but for some reason failed to respond to repeated automatic-review reminders. AEs handling these papers have to take several extra steps to remind reviewers through phone calls or e-mails, look for replacement reviewers, or review the papers themselves. Most authors have been appreciative of the work of the AEs and reviewers, and they accept our decisions without a problem. The backlog of papers has been kept short over the last year. We have maintained an organized printing and paperacceptance schedule, with papers typically printed in the journal within 2‐3 months of acceptance. Our page budget has been kept constant in the past few years (roughly 2060 pages per year), and we expect to hold the same page count for next year. Marco Baglietto, Lubica Benusková, Ivo Bukovsky, Tianping Chen, Tom Heskes, Kazushi Ikeda, Fakhri Karray, Rhee Man Kil, Robert Legenstein, Jinhu Lü 0001, Yunqian Ma, Malik Magdon-Ismail, Michael G. Paulin, Robi Polikar, Danil V. Prokhorov, Marco A. Wiering, Vicente Zarzoso |
IEEE Trans. Neural Networks | 5 |
| 2010 | Bayesian Monte Carlo for the Global Optimization of Expensive Functions
Perry Groot, Adriana Birlutiu, Tom Heskes |
ECAI | 3 |
| 2010 | Causal discovery in multiple models from different experimentsabstractA long-standing open research problem is how to use information from different experiments, including background knowledge, to infer causal relations. Recent developments have shown ways to use multiple data sets, provided they originate from identical experiments. We present the MCI-algorithm as the first method that can infer provably valid causal relations in the large sample limit from different experiments. It is fast, reliable and produces very clear and easily interpretable output. It is based on a result that shows that constraint-based causal discovery is decomposable into a candidate pair identification and subsequent elimination step that can be applied separately from different models. We test the algorithm on a variety of synthetic input model sets to assess its behavior and the quality of the output. The method shows promising signs that it can be adapted to suit causal discovery in real-world application areas as well, including large databases. Tom Claassen, Tom Heskes |
NIPS | 2 |
| 2010 | Multi-task preference learning with an application to hearing aid personalization
Adriana Birlutiu, Perry Groot, Tom Heskes |
Neurocomputing | 3 |
| 2010 | Neural Decoding with Hierarchical Generative ModelsabstractRecent research has shown that reconstruction of perceived images based on hemodynamic response as measured with functional magnetic resonance imaging (fMRI) is starting to become feasible. In this letter, we explore reconstruction based on a learned hierarchy of features by employing a hierarchical generative model that consists of conditional restricted Boltzmann machines. In an unsupervised phase, we learn a hierarchy of features from data, and in a supervised phase, we learn how brain activity predicts the states of those features. Reconstruction is achieved by sampling from the model, conditioned on brain activity. We show that by using the hierarchical generative model, we can obtain good-quality reconstructions of visual images of handwritten digits presented during an fMRI scanning session. Marcel van Gerven, Floris P. de Lange, Tom Heskes |
Neural Comput. | 3 |
| 2009 | Exploring the impact of alternative feature representations on BCI classification
Ali Bahramisharif, Marcel van Gerven, Tom Heskes |
ESANN | 3 |
| 2009 | Multi-task Preference learning with Gaussian Processes
Adriana Birlutiu, Perry Groot, Tom Heskes |
ESANN | 3 |
| 2009 | Bayesian Source Localization with the Multivariate Laplace PriorabstractWe introduce a novel multivariate Laplace (MVL) distribution as a sparsity promoting prior for Bayesian source localization that allows the specification of constraints between and within sources. We represent the MVL distribution as a scale mixture that induces a coupling between source variances instead of their means. Approximation of the posterior marginals using expectation propagation is shown to be very efficient due to properties of the scale mixture representation. The computational bottleneck amounts to computing the diagonal elements of a sparse matrix inverse. Our approach is illustrated using a mismatch negativity paradigm for which MEG data and a structural MRI have been acquired. We show that spatial coupling leads to sources which are active over larger cortical areas as compared with an uncoupled prior. Marcel van Gerven, Botond Cseke, Robert Oostenveld, Tom Heskes |
NIPS | 4 |
| 2009 | Gene regulation in the intraerythrocytic cycle of Plasmodium falciparumabstractMOTIVATION: To date, there is little knowledge about one of the processes fundamental to the biology of Plasmodium falciparum, gene regulation including transcriptional control. We use noisy threshold models to identify regulatory sequence elements explaining membership to a gene expression cluster where each cluster consists of genes active during the part of the developmental cycle inside a red blood cell. Our approach is both able to capture the combinatorial nature of gene regulation and to incorporate uncertainty about the functionality of putative regulatory sequence elements. RESULTS: We find a characteristic pattern where the most common motifs tend to be absent upstream of genes active in the first half of the cycle and present upstream of genes active in the second half. We find no evidence that motif's score, orientation, location and multiplicity improves prediction of gene expression. Through comparative genome analysis, we find a list of potential transcription factors and their associated motifs. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rasa Jurgelenaite, Tjeerd Dijkstra, Clemens H. M. Kocken, Tom Heskes |
Bioinform. | 4 |
| 2009 | Selecting features for BCI control based on a covert spatial attention paradigm
Marcel van Gerven, Ali Bahramisharif, Tom Heskes, Ole Jensen |
Neural Networks | 3 |
| 2008 | Bounds on the Bethe Free Energy for Gaussian Networks
Botond Cseke, Tom Heskes |
UAI | 2 |
| 2007 | Regulator Discovery from Gene Expression Time Series of Malaria Parasites: a Hierachical ApproachabstractWe introduce a hierarchical Bayesian model for the discovery of putative regulators from gene expression data only. The hierarchy incorporates the knowledge that there are just a few regulators that by themselves only regulate a handful of genes. This is implemented through a so-called spike-and-slab prior, a mixture of Gaussians with different widths, with mixing weights from a hierarchical Bernoulli model. For efficient inference we implemented expectation propagation. Running the model on a malaria parasite data set, we found four genes with significant homology to transcription factors in an amoebe, one RNA regulator and three genes of unknown function (out of the top ten genes considered). José Miguel Hernández-Lobato, Tjeerd Dijkstra, Tom Heskes |
NIPS | 3 |
| 2007 | Expectation Propagation for Rating Players in Sports Competitions
Adriana Birlutiu, Tom Heskes |
PKDD | 2 |
| 2007 | Predicting carcinoid heart disease with the noisy-threshold classifier
Marcel van Gerven, Rasa Jurgelenaite, Babs G. Taal, Tom Heskes, Peter J. F. Lucas |
Artif. Intell. Medicine | 4 |
| 2006 | EM Algorithm for Symmetric Causal Independence Models
Rasa Jurgelenaite, Tom Heskes |
ECML | 2 |
| 2006 | Convexity Arguments for Efficient Minimization of the Bethe and Kikuchi Free EnergiesabstractLoopy and generalized belief propagation are popular algorithms for approximate inference in Markov random fields and Bayesian networks. Fixed points of these algorithms have been shown to correspond to extrema of the Bethe and Kikuchi free energy, both of which are approximations of the exact Helmholtz free energy. However, belief propagation does not always converge, which motivates approaches that explicitly minimize the Kikuchi/Bethe free energy, such as CCCP and UPS. Here we describe a class of algorithms that solves this typically non-convex constrained minimization problem through a sequence of convex constrained minimizations of upper bounds on the Kikuchi free energy. Intuitively one would expect tighter bounds to lead to faster algorithms, which is indeed convincingly demonstrated in our simulations. Several ideas are applied to obtain tight convex bounds that yield dramatic speed-ups over CCCP. Tom Heskes |
J. Artif. Intell. Res. | 1 |
| 2005 | Novel approximations for inference in nonlinear dynamical systems using expectation propagation
Alexander Ypma, Tom Heskes |
Neurocomputing | 2 |
| 2005 | Change Point Problems in Linear Dynamical SystemsabstractWe study the problem of learning two regimes (we have a normal and a prefault regime in mind) based on a train set of non-Markovian observation sequences. Key to the model is that we assume that once the system switches from the normal to the prefault regime it cannot restore and will eventually result in a fault. We refer to the particular setting as semi-supervised since we assume the only information given to the learner is whether a particular sequence ended with a stop (implying that the sequence was generated by the normal regime) or with a fault (implying that there was a switch from the normal to the fault regime). In the latter case the particular time point at which a switch occurred is not known. The underlying model used is a switching linear dynamical system (SLDS). The constraints in the regime transition probabilities result in an exact inference procedure that scales quadratically with the length of a sequence. Maximum aposteriori (MAP) parameter estimates can be found using an expectation maximization (EM) algorithm with this inference algorithm in the E-step. For long sequences this will not be practically feasible and an approximate inference and an approximate EM procedure is called for. We describe a flexible class of approximations corresponding to different choices of clusters in a Kikuchi free energy with weak consistency constraints. Onno Zoeter, Tom Heskes |
J. Mach. Learn. Res. | 2 |
| 2004 | Novel approximations for inference and learning in nonlinear dynamical systems
Alexander Ypma, Tom Heskes |
ESANN | 2 |
| 2004 | On the Uniqueness of Loopy Belief Propagation Fixed PointsabstractWe derive sufficient conditions for the uniqueness of loopy belief propagation fixed points. These conditions depend on both the structure of the graph and the strength of the potentials and naturally extend those for convexity of the Bethe free energy. We compare them with (a strengthened version of) conditions derived elsewhere for pairwise potentials. We discuss possible implications for convergent algorithms, as well as for other approximate free energies. Tom Heskes |
Neural Comput. | 1 |
| 2003 | Multi-scale Switching Linear Dynamical Systems
Onno Zoeter, Tom Heskes |
ICANN | 2 |
| 2003 | Approximate Expectation MaximizationabstractWe discuss the integration of the expectation-maximization (EM) algorithm for maximum likelihood learning of Bayesian networks with belief propagation algorithms for approximate inference. Specifically we propose to combine the outer-loop step of convergent belief propagation algorithms with the M-step of the EM algorithm. This then yields an approximate EM algorithm that is essentially still double loop, with the important advantage of an inner loop that is guaranteed to converge. Simulations illustrate the merits of such an approach. Tom Heskes, Onno Zoeter, Wim Wiegerinck |
NIPS | 1 |
| 2003 | Approximate Inference and Constrained Optimization
Tom Heskes, Kees Albers, Hilbert J. Kappen |
UAI | 1 |
| 2003 | Task Clustering and Gating for Bayesian Multitask Learning
Bart Bakker, Tom Heskes |
J. Mach. Learn. Res. | 2 |
| 2003 | Optimising newspaper sales using neural-Bayesian technology
Tom Heskes, Jan-Joost Spanjers, Bart Bakker, Wim Wiegerinck |
Neural Comput. Appl. | 1 |
| 2003 | Clustering ensembles of neural network models
Bart Bakker, Tom Heskes |
Neural Networks | 2 |
| 2003 | Hierarchical Visualization of Time-Series Data Using Switching Linear Dynamical SystemsabstractWe propose a novel visualization algorithm for high-dimensional time-series data. In contrast to most visualization techniques, we do not assume consecutive data points to be independent. The basic model is a linear dynamical system which can be seen as a dynamic extension of a probabilistic principal component model. A further extension to a particular switching linear dynamical system allows a representation of complex data onto multiple and even a hierarchy of plots. Using sensible approximations based on expectation propagation, the projections can be performed in essentially the same order of complexity as their static counterpart. We apply our method on a real-world data set with sensor readings from a paper machine. Onno Zoeter, Tom Heskes |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Model Clustering for Neural Network Ensembles
Bart Bakker, Tom Heskes |
ICANN | 2 |
| 2002 | Stable Fixed Points of Loopy Belief Propagation Are Local Minima of the Bethe Free EnergyabstractWe extend recent work on the connection between loopy belief propagation and the Bethe free energy. Constrained minimization of the Bethe free energy can be turned into an unconstrained saddle-point problem. Both converging double-loop algorithms and standard loopy belief propagation can be inter- preted as attempts to solve this saddle-point problem. Stability analysis then leads us to conclude that stable (cid:12)xed points of loopy belief propagation must be (local) minima of the Bethe free energy. Perhaps surprisingly, the converse need not be the case: minima can be unstable (cid:12)xed points. We illustrate this with an example and discuss implications. Tom Heskes |
NIPS | 1 |
| 2002 | Fractional Belief PropagationabstractWe consider loopy belief propagation for approximate inference in prob- abilistic graphical models. A limitation of the standard algorithm is that clique marginals are computed as if there were no loops in the graph. To overcome this limitation, we introduce fractional belief propagation. Fractional belief propagation is formulated in terms of a family of ap- proximate free energies, which includes the Bethe free energy and the naive mean-field free as special cases. Using the linear response correc- tion of the clique marginals, the scale parameters can be tuned. Simula- tion results illustrate the potential merits of the approach. Wim Wiegerinck, Tom Heskes |
NIPS | 2 |
| 2002 | Expectation Propogation for Approximate Inference in Dynamic Bayesian Networks
Tom Heskes, Onno Zoeter |
UAI | 1 |
| 2002 | IPF for Discrete Chain Factor Graphs
Wim Wiegerinck, Tom Heskes |
UAI | 2 |
| 2002 | Approximate algorithms for neural-Bayesian approaches
Tom Heskes, Bart Bakker, Hilbert J. Kappen |
Theor. Comput. Sci. | 1 |
| 2001 | Self-organizing maps, vector quantization, and mixture modelingabstractSelf-organizing maps are popular algorithms for unsupervised learning and data visualization. Exploiting the link between vector quantization and mixture modeling, we derive expectation-maximization (EM) algorithms for self-organizing maps with and without missing values. We compare self-organizing maps with the elastic-net approach and explain why the former is better suited for the visualization of high-dimensional data. Several extensions and improvements are discussed. As an illustration we apply a self-organizing map based on a multinomial distribution to market basket analysis. Tom Heskes |
IEEE Trans. Neural Networks | 1 |
| 2000 | Empirical Bayes for Learning to Learn
Tom Heskes |
ICML | 1 |
| 2000 | General Bias/Variance Decomposition with Target Independent Variance of Error Functions Derived from the Exponential Family of DistributionsabstractAn important theoretical tool in machine learning is the bias/variance decomposition of the generalization error. It was introduced for the mean square error. The bias/variance decomposition includes the concept of the average predictor. The bias is the error of the average predictor, and the systematic part of the generalization error, while the variability around the average predictor is the variance. We present a large group of error functions with the same desirable properties as the bias/variance decomposition. The error functions are derived from the exponential family of distributions via the statistical deviance measure. We prove that this family of error functions contains all error functions decomposable in that manner. We state the connection between the bias/variance decomposition and the ambiguity decomposition and present a useful approximation of ambiguity that is quadratic in the ensemble coefficients. Jakob V. Hansen, Tom Heskes |
ICPR | 2 |
| 2000 | EM Algorithms for Self-Organizing MapsabstractSelf-organizing maps are popular algorithms for unsupervised learning and data visualization. Exploiting the link between vector quantization and mixture modeling, we derive EM algorithms for self-organizing maps with and without missing values. We compare self-organizing maps with the elastic-net approach and explain why the former is better suited for the visualization of high-dimensional data. Several extensions and improvements are discussed. Tom Heskes, Jan-Joost Spanjers, Wim Wiegerinck |
IJCNN (6) | 1 |
| 2000 | Input selection based on an ensemble
Piërre van de Laar, Tom Heskes |
Neurocomputing | 2 |
| 2000 | On "Natural" Learning and Pruning in Multilayered PerceptronsabstractSeveral studies have shown that natural gradient descent for on-line learning is much more efficient than standard gradient descent. In this article, we derive natural gradients in a slightly different manner and discuss implications for batch-mode learning and pruning, linking them to existing algorithms such as Levenberg-Marquardt optimization and optimal brain surgeon. The Fisher matrix plays an important role in all these algorithms. The second half of the article discusses a layered approximation of the Fisher matrix specific to multilayered perceptrons. Using this approximation rather than the exact Fisher matrix, we arrive at much faster "natural" learning algorithms and more robust pruning procedures. Tom Heskes |
Neural Comput. | 1 |
| 1999 | Model clustering by deterministic annealing
Bart Bakker, Tom Heskes |
ESANN | 2 |
| 1999 | Partial Retraining: A New Approach to Input Relevance DeterminationabstractIn this article we introduce partial retraining, an algorithm to determine the relevance of the input variables of a trained neural network. We place this algorithm in the context of other approaches to relevance determination. Numerical experiments on both artificial and real-world problems show that partial retraining outperforms its competitors, which include methods based on constant substitution, analysis of weight magnitudes, and "optimal brain surgeon". Piërre van de Laar, Tom Heskes, Stan C. A. M. Gielen |
Int. J. Neural Syst. | 2 |
| 1999 | Pruning Using Parameter and Neuronal MetricsabstractIn this article, we introduce a measure of optimality for architecture selection algorithms for neural networks: the distance from the original network to the new network in a metric defined by the probability distributions of all possible networks. We derive two pruning algorithms, one based on a metric in parameter space and the other based on a metric in neuron space, which are closely related to well-known architecture selection algorithms, such as GOBS. Our framework extends the theoretically range of validity of GOBS and therefore can explain results observed in previous experiments. In addition, we give some computational improvements for these algorithms. Piërre van de Laar, Tom Heskes |
Neural Comput. | 2 |
| 1998 | Solving a Huge Number of Similar Tasks: A Combination of Multi-Task Learning and a Hierarchical Bayesian Approach
Tom Heskes |
ICML | 1 |
| 1998 | Bias/Variance Decompositions for Likelihood-Based EstimatorsabstractThe bias/variance decomposition of mean-squared error is well understood and relatively straightforward. In this note, a similar simple decomposition is derived, valid for any kind of error measure that, when using the appropriate probability model, can be derived from a Kullback-Leibler divergence or log-likelihood. Tom Heskes |
Neural Comput. | 1 |
| 1997 | Input Selection with Partial Retraining
Piërre van de Laar, Stan C. A. M. Gielen, Tom Heskes |
ICANN | 3 |
| 1997 | Selecting Weighting Factors in Logarithmic Opinion Pools
Tom Heskes |
NIPS | 1 |
| 1997 | Task-Dependent Learning of Attention
Piërre van de Laar, Tom Heskes, Stan C. A. M. Gielen |
Neural Networks | 2 |
| 1996 | Practical Confidence and Prediction Intervals
Tom Heskes |
NIPS | 1 |
| 1996 | Balancing Between Bagging and Bumping
Tom Heskes |
NIPS | 1 |
| 1996 | How Dependencies between Successive Examples Affect On-Line LearningabstractWe study the dynamics of on-line learning for a large class of neural networks and learning rules, including backpropagation for multilayer perceptrons. In this paper, we focus on the case where successive examples are dependent, and we analyze how these dependencies affect the learning process. We define the representation error and the prediction error. The representation error measures how well the environment is represented by the network after learning. The prediction error is the average error that a continually learning network makes on the next example. In the neighborhood of a local minimum of the error surface, we calculate these errors. We find that the more predictable the example presentation, the higher the representation error, i.e., the less accurate the asymptotic representation of the whole environment. Furthermore we study the learning process in the presence of a plateau. Plateaus are flat spots on the error surface, which can severely slow down the learning process. In particular, they are notorious in applications with multilayer perceptrons. Our results, which are confirmed by simulations of a multilayer perceptron learning a chaotic time series using backpropagation, explain how dependencies between examples can help the learning process to escape from a plateau. Wim Wiegerinck, Tom Heskes |
Neural Comput. | 2 |
| 1996 | A theoretical comparison of batch-mode, on-line, cyclic, and almost-cyclic learningabstractWe study and compare different neural network learning strategies: batch-mode learning, online learning, cyclic learning, and almost-cyclic learning. Incremental learning strategies require less storage capacity than batch-mode learning. However, due to the arbitrariness in the presentation order of the training patterns, incremental learning is a stochastic process; whereas batch-mode learning is deterministic. In zeroth order, i.e., as the learning parameter eta tends to zero, all learning strategies approximate the same ordinary differential equation for convenience referred to as the "ideal behavior". Using stochastic methods valid for small learning parameters eta, we derive differential equations describing the evolution of the lowest-order deviations from this ideal behavior. We compute how the asymptotic misadjustment, measuring the average asymptotic distance from a stable fixed point of the ideal behavior, scales as a function of the learning parameter and the number of training patterns. Knowing the asymptotic misadjustment, we calculate the typical number of learning steps necessary to generate a weight within order epsilon of this fixed point, both with fixed and time-dependent learning parameters. We conclude that almost-cyclic learning (learning with random cycles) is a better alternative for batch-mode learning than cyclic learning (learning with a fixed cycle). Tom Heskes, Wim Wiegerinck |
IEEE Trans. Neural Networks | 1 |
| 1994 | Stochastics of on-line back-propagation
Tom Heskes |
ESANN | 1 |
| 1992 | Retrieval of pattern sequences at variable speeds in a neural network with delays
Tom Heskes, Stan C. A. M. Gielen |
Neural Networks | 1 |