EDBT 2026 Demo / reviewers in the wild / expert
Christopher C. Drovandi
dblp:29/9583 · also Chris Drovandi
· DBLP profile ↗
12ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0001-9222-8763ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Polynomial Stein Discrepancy for Assessing Moment ConvergenceabstractWe propose a novel method for measuring the discrepancy between a set of samples and a desired posterior distribution for Bayesian inference. Classical methods for assessing sample quality like the effective sample size are not appropriate for scalable Bayesian sampling algorithms, such as stochastic gradient Langevin dynamics, that are asymptotically biased. Instead, the gold standard is to use the kernel Stein Discrepancy (KSD), which is itself not scalable given its quadratic cost in the number of samples. The KSD and its faster extensions also typically suffer from the curse-of-dimensionality and can require extensive tuning. To address these limitations, we develop the polynomial Stein discrepancy (PSD) and an associated goodness-of-fit test. While the new test is not fully convergence-determining, we prove that it detects differences in the first $r$ moments for Gaussian targets. We empirically show that the test has higher power than its competitors in several examples, and at a lower computational cost. Finally, we demonstrate that the PSD can assist practitioners to select hyper-parameters of Bayesian sampling algorithms more efficiently than competitors. Narayan Srinivasan, Matthew Sutton, Christopher C. Drovandi, Leah F. South |
ICML | 3 |
| 2025 | A Unified Framework for Variable Selection in Model-Based Clustering with Missing Not at RandomabstractModel-based clustering integrated with variable selection is a powerful tool for uncovering latent structures within complex data. However, its effectiveness is often hindered by challenges such as identifying relevant variables that define heterogeneous subgroups and handling data that are missing not at random, a prevalent issue in fields like transcriptomics. While several notable methods have been proposed to address these problems, they typically tackle each issue in isolation, thereby limiting their flexibility and adaptability. This paper introduces a unified framework designed to address these challenges simultaneously. Our approach incorporates a data-driven penalty matrix into penalized clustering to enable more flexible variable selection, along with a mechanism that explicitly models the relationship between missingness and latent class membership. We demonstrate that, under certain regularity conditions, the proposed framework achieves both asymptotic consistency and selection consistency, even in the presence of missing data. This unified strategy significantly enhances the capability and efficiency of model-based clustering, advancing methodologies for identifying informative variables that define homogeneous subgroups in the presence of complex missing data patterns. The performance of the framework, including its computational efficiency, is evaluated through simulations and demonstrated using both synthetic and real-world transcriptomic datasets. Binh H. Ho, Long Nguyen Chi, TrungTin Nguyen, Van Ha Hoang, Christopher C. Drovandi |
NeurIPS | 6 |
| 2025 | Bayesian Score Calibration for Approximate ModelsabstractScientists continue to develop increasingly complex mechanistic models to reflect their knowledge more realistically. Statistical inference using these models can be challenging since the corresponding likelihood function is often intractable and model simulation may be computationally burdensome. Fortunately, in many of these situations it is possible to adopt a surrogate model or approximate likelihood function. It may be convenient to conduct Bayesian inference directly with a surrogate, but this can result in a posterior with poor uncertainty quantification. In this paper, we propose a new method for adjusting approximate posterior samples to reduce bias and improve posterior coverage properties. We do this by optimizing a transformation of the approximate posterior, the result of which maximizes a scoring rule. Our approach requires only a (fixed) small number of complex model simulations and is numerically stable. We develop supporting theory for our method and demonstrate beneficial corrections to approximate posteriors across several examples of increasing complexity. Joshua J. Bon, David J. Warne, David J. Nott, Christopher C. Drovandi |
J. Mach. Learn. Res. | 4 |
| 2024 | LooplessFluxSampler: an efficient toolbox for sampling the loopless flux solution space of metabolic modelsabstractBACKGROUND: Uniform random sampling of mass-balanced flux solutions offers an unbiased appraisal of the capabilities of metabolic networks. Unfortunately, it is impossible to avoid thermodynamically infeasible loops in flux samples when using convex samplers on large metabolic models. Current strategies for randomly sampling the non-convex loopless flux space display limited efficiency and lack theoretical guarantees. RESULTS: Here, we present LooplessFluxSampler, an efficient algorithm for exploring the loopless mass-balanced flux solution space of metabolic models, based on an Adaptive Directions Sampling on a Box (ADSB) algorithm. ADSB is rooted in the general Adaptive Direction Sampling (ADS) framework, specifically the Parallel ADS, for which theoretical convergence and irreducibility results are available for sampling from arbitrary distributions. By sampling directions that adapt to the target distribution, ADSB traverses more efficiently the sample space achieving faster mixing than other methods. Importantly, the presented algorithm is guaranteed to target the uniform distribution over convex regions, and it provably converges on the latter distribution over more general (non-convex) regions provided the sample can have full support. CONCLUSIONS: LooplessFluxSampler enables scalable statistical inference of the loopless mass-balanced solution space of large metabolic models. Grounded in a theoretically sound framework, this toolbox provides not only efficient but also reliable results for exploring the properties of the almost surely non-convex loopless flux space. Finally, LooplessFluxSampler includes a Markov Chain diagnostics suite for assessing the quality of the final sample and the performance of the algorithm. Pedro A. Saa, Sebastian Zapararte, Christopher C. Drovandi, Lars Keld Nielsen |
BMC Bioinform. | 3 |
| 2024 | Perlin noise generation of physiologically realistic cardiac fibrosisabstractFibrosis, a pathological increase in extracellular matrix proteins, is a significant health issue that hinders the function of many organs in the body, in some cases fatally. In the heart, fibrosis impacts on electrical propagation in a complex and poorly predictable fashion, potentially serving as a substrate for dangerous arrhythmias. Individual risk depends on the spatial manifestation of fibrotic tissue, and learning the spatial arrangement on the fine scale in order to predict these impacts still relies upon invasive ex vivo procedures. As a result, the effects of spatial variability on the symptomatic impact of cardiac fibrosis remain poorly understood. In this work, we address the issue of availability of such imaging data via a computational methodology for generating new realisations of cardiac fibrosis microstructure. Using the Perlin noise technique from computer graphics, together with an automated calibration process that requires only a single training image, we demonstrate successful capture of collagen texturing in four types of fibrosis microstructure observed in histological sections. We then use this generator to quantitatively analyse the conductive properties of these different types of cardiac fibrosis, as well as produce three-dimensional realisations of histologically-observed patterning. Owing to the generator's flexibility and automated calibration process, we also anticipate that it might be useful in producing additional realisations of other physiological structures. Brodie Lawson, Christopher C. Drovandi, Pamela M. Burrage, Alfonso Bueno-Orovio, Rodrigo Weber dos Santos, Blanca Rodríguez, Kerrie L. Mengersen, Kevin Burrage |
Medical Image Anal. | 2 |
| 2024 | Unlocking ensemble ecosystem modelling for large and complex networksabstractThe potential effects of conservation actions on threatened species can be predicted using ensemble ecosystem models by forecasting populations with and without intervention. These model ensembles commonly assume stable coexistence of species in the absence of available data. However, existing ensemble-generation methods become computationally inefficient as the size of the ecosystem network increases, preventing larger networks from being studied. We present a novel sequential Monte Carlo sampling approach for ensemble generation that is orders of magnitude faster than existing approaches. We demonstrate that the methods produce equivalent parameter inferences, model predictions, and tightly constrained parameter combinations using a novel sensitivity analysis method. For one case study, we demonstrate a speed-up from 108 days to 6 hours, while maintaining equivalent ensembles. Additionally, we demonstrate how to identify the parameter combinations that strongly drive feasibility and stability, drawing ecological insight from the ensembles. Now, for the first time, larger and more realistic networks can be practically simulated and analysed. Sarah A. Vollert, Christopher C. Drovandi, Matthew P. Adams |
PLoS Comput. Biol. | 2 |
| 2023 | Transport Reversible Jump ProposalsabstractReversible jump Markov chain Monte Carlo (RJMCMC) proposals that achieve reasonable acceptance rates and mixing are notoriously difficult to design in most applications. Inspired by recent advances in deep neural network-based normalizing flows and density estimation, we demonstrate an approach to enhance the efficiency of RJMCMC sampling by performing transdimensional jumps involving reference distributions. In contrast to other RJMCMC proposals, the proposed method is the first to apply a non-linear transport-based approach to construct efficient proposals between models with complicated dependency structures. It is shown that, in the setting where exact transports are used, our RJMCMC proposals have the desirable property that the acceptance probability depends only on the model probabilities. Numerical experiments demonstrate the efficacy of the approach. Laurence Davies, Robert Salomone, Matthew Sutton, Christopher C. Drovandi |
AISTATS | 4 |
| 2022 | Efficient inference and identifiability analysis for differential equation models with random parametersabstractHeterogeneity is a dominant factor in the behaviour of many biological processes. Despite this, it is common for mathematical and statistical analyses to ignore biological heterogeneity as a source of variability in experimental data. Therefore, methods for exploring the identifiability of models that explicitly incorporate heterogeneity through variability in model parameters are relatively underdeveloped. We develop a new likelihood-based framework, based on moment matching, for inference and identifiability analysis of differential equation models that capture biological heterogeneity through parameters that vary according to probability distributions. As our novel method is based on an approximate likelihood function, it is highly flexible; we demonstrate identifiability analysis using both a frequentist approach based on profile likelihood, and a Bayesian approach based on Markov-chain Monte Carlo. Through three case studies, we demonstrate our method by providing a didactic guide to inference and identifiability analysis of hyperparameters that relate to the statistical moments of model parameters from independent observed data. Our approach has a computational cost comparable to analysis of models that neglect heterogeneity, a significant improvement over many existing alternatives. We demonstrate how analysis of random parameter models can aid better understanding of the sources of heterogeneity from biological data. Alexander P. Browning, Christopher C. Drovandi, Ian W. Turner, Adrianne L. Jenner, Matthew J. Simpson |
PLoS Comput. Biol. | 2 |
| 2022 | Computationally efficient mechanism discovery for cell invasion with uncertainty quantificationabstractParameter estimation for mathematical models of biological processes is often difficult and depends significantly on the quality and quantity of available data. We introduce an efficient framework using Gaussian processes to discover mechanisms underlying delay, migration, and proliferation in a cell invasion experiment. Gaussian processes are leveraged with bootstrapping to provide uncertainty quantification for the mechanisms that drive the invasion process. Our framework is efficient, parallelisable, and can be applied to other biological problems. We illustrate our methods using a canonical scratch assay experiment, demonstrating how simply we can explore different functional forms and develop and test hypotheses about underlying mechanisms, such as whether delay is present. All code and data to reproduce this work are available at https://github.com/DanielVandH/EquationLearning.jl. Daniel J. Vandenheuvel, Christopher C. Drovandi, Matthew J. Simpson |
PLoS Comput. Biol. | 2 |
| 2021 | Inference of ventricular activation properties from non-invasive electrocardiographyabstractThe realisation of precision cardiology requires novel techniques for the non-invasive characterisation of individual patients’ cardiac function to inform therapeutic and diagnostic decision-making. Both electrocardiography and imaging are used for the clinical diagnosis of cardiac disease. The integration of multi-modal datasets through advanced computational methods could enable the development of the cardiac ‘digital twin’, a comprehensive virtual tool that mechanistically reveals a patient's heart condition from clinical data and simulates treatment outcomes. The adoption of cardiac digital twins requires the non-invasive efficient personalisation of the electrophysiological properties in cardiac models. This study develops new computational techniques to estimate key ventricular activation properties for individual subjects by exploiting the synergy between non-invasive electrocardiography, cardiac magnetic resonance (CMR) imaging and modelling and simulation. More precisely, we present an efficient sequential Monte Carlo approximate Bayesian computation-based inference method, integrated with Eikonal simulations and torso-biventricular models constructed based on clinical CMR imaging. The method also includes a novel strategy to treat combined continuous (conduction speeds) and discrete (earliest activation sites) parameter spaces and an efficient dynamic time warping-based ECG comparison algorithm. We demonstrate results from our inference method on a cohort of twenty virtual subjects with cardiac ventricular myocardial-mass volumes ranging from 74 cm3 to 171 cm3 and considering low versus high resolution for the endocardial discretisation (which determines possible locations of the earliest activation sites). Results show that our method can successfully infer the ventricular activation properties in sinus rhythm from non-invasive epicardial activation time maps and ECG recordings, achieving higher accuracy for the endocardial speed and sheet (transmural) speed than for the fibre or sheet-normal directed speeds. Julià Camps, Brodie Lawson, Christopher C. Drovandi, Ana Mincholé, Zhinuo J. Wang, Vicente Grau, Kevin Burrage, Blanca Rodríguez |
Medical Image Anal. | 3 |
| 2020 | Estimating a novel stochastic model for within-field disease dynamics of banana bunchy top virus via approximate Bayesian computationabstractThe Banana Bunchy Top Virus (BBTV) is one of the most economically important vector-borne banana diseases throughout the Asia-Pacific Basin and presents a significant challenge to the agricultural sector. Current models of BBTV are largely deterministic, limited by an incomplete understanding of interactions in complex natural systems, and the appropriate identification of parameters. A stochastic network-based Susceptible-Infected-Susceptible model has been created which simulates the spread of BBTV across the subsections of a banana plantation, parameterising nodal recovery, neighbouring and distant infectivity across summer and winter. Findings from posterior results achieved through Markov Chain Monte Carlo approach to approximate Bayesian computation suggest seasonality in all parameters, which are influenced by correlated changes in inspection accuracy, temperatures and aphid activity. This paper demonstrates how the model may be used for monitoring and forecasting of various disease management strategies to support policy-level decision making. Abhishek Varghese, Christopher C. Drovandi, Antonietta Mira, Kerrie L. Mengersen |
PLoS Comput. Biol. | 2 |
| 2015 | Melanoma Cell Colony Expansion Parameters Revealed by Approximate Bayesian ComputationabstractIn vitro studies and mathematical models are now being widely used to study the underlying mechanisms driving the expansion of cell colonies. This can improve our understanding of cancer formation and progression. Although much progress has been made in terms of developing and analysing mathematical models, far less progress has been made in terms of understanding how to estimate model parameters using experimental in vitro image-based data. To address this issue, a new approximate Bayesian computation (ABC) algorithm is proposed to estimate key parameters governing the expansion of melanoma cell (MM127) colonies, including cell diffusivity, D, cell proliferation rate, λ, and cell-to-cell adhesion, q, in two experimental scenarios, namely with and without a chemical treatment to suppress cell proliferation. Even when little prior biological knowledge about the parameters is assumed, all parameters are precisely inferred with a small posterior coefficient of variation, approximately 2-12%. The ABC analyses reveal that the posterior distributions of D and q depend on the experimental elapsed time, whereas the posterior distribution of λ does not. The posterior mean values of D and q are in the ranges 226-268 µm2h-1, 311-351 µm2h-1 and 0.23-0.39, 0.32-0.61 for the experimental periods of 0-24 h and 24-48 h, respectively. Furthermore, we found that the posterior distribution of q also depends on the initial cell density, whereas the posterior distributions of D and λ do not. The ABC approach also enables information from the two experiments to be combined, resulting in greater precision for all estimates of D and λ. Brenda N. Vo, Christopher C. Drovandi, Anthony N. Pettitt, Graeme J. Pettet |
PLoS Comput. Biol. | 2 |