Niel Hens

dblp:44/2253 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0003-1881-0637ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Algorithms for Markov Binomial Chains
abstract
We study algorithms to analyze a particular class of Markov population processes that is often used in epidemiology. More specifically, Markov binomial chains are the model that arises from stochastic time-discretizations of classical compartmental models. In this work we formalize this class of Markov population processes and focus on the problem of computing the expected time to termination in a given such model. Our theoretical contributions include proving that Markov binomial chains whose flow of individuals through compartments is acyclic almost surely terminate. We give a PSPACE algorithm for the problem of approximating the time to termination and a direct algorithm for the exact problem in the Blum-Shub-Smale model of computation. Finally, we provide a natural encoding of Markov binomial chains into a common input language for probabilistic model checkers. We implemented the latter encoding and present some initial empirical results showcasing what formal methods can do for practicing epidemiologists.
Alejandro Alarcón Gonzalez, Niel Hens, Tim Leys, Guillermo A. Pérez
Log. Methods Comput. Sci.2
2025 Nonparametric serial interval estimation with uniform mixtures
abstract
The serial interval of an infectious disease is a key instrument to understand transmission dynamics. Estimation of the serial interval distribution from illness onset data extracted from transmission pairs is challenging due to the presence of censoring and state-of-the-art methods mostly rely on parametric models. We present a fully data-driven methodology to estimate the serial interval distribution based on interval-censored serial interval data. The proposed nonparametric estimator of the cumulative distribution function of the serial interval is based on the class of uniform mixtures. Closed-form solutions are available for point estimates of different serial interval features and the bootstrap is used to construct confidence intervals. Algorithms underlying our approach are simple, stable, and computationally inexpensive, making them easily implementable in a programming language that is most familiar to a potential user. The nonparametric user-friendly routine is included in the EpiDelays package for ease of implementation. Our method complements existing parametric approaches for serial interval estimation and permits to analyze past, current, or future illness onset data streams following a set of best practices in epidemiological delay modeling.
Oswaldo Gressani, Niel Hens
PLoS Comput. Biol.2
2025 Measuring approximate functional dependencies: a comparative study
Marcel Parciak, Sebastiaan Weytjens, Niel Hens, Frank Neven, Liesbet M. Peeters, Stijn Vansummeren
VLDB J.3
2024 Measuring Approximate Functional Dependencies: A Comparative Study
abstract
Approximate functional dependencies (AFDs) are functional dependencies (FDs) that “almost” hold in a relation. While various measures have been proposed to quantify the level to which an FD holds approximately, they are difficult to compare and it is unclear which measure is preferable when one needs to discover FDs in real-world data, i.e., data that only approximately satisfies the FD. In response, this paper formally and qualitatively compares AFD measures. We obtain a formal comparison through a novel presentation of measures in terms of Shannon and logical entropy. Qualitatively, we perform a sensitivity analysis w.r.t. structural properties of input relations and quantitatively study the effectiveness of AFD measures for ranking AFDs on real world data. Based on this analysis, we give clear recommendations for the AFD measures to use in practice.
Marcel Parciak, Sebastiaan Weytjens, Niel Hens, Frank Neven, Liesbet M. Peeters, Stijn Vansummeren
ICDE3
2024 Exploring the Pareto front of multi-objective COVID-19 mitigation policies using reinforcement learning
abstract
Infectious disease outbreaks can have a disruptive impact on public health and societal processes. As decision-making in the context of epidemic mitigation is multi-dimensional hence complex, reinforcement learning in combination with complex epidemic models provides a methodology to design refined prevention strategies. Current research focuses on optimizing policies with respect to a single objective, such as the pathogen’s attack rate. However, as the mitigation of epidemics involves distinct, and possibly conflicting, criteria (i.a., mortality, morbidity, economic cost, well-being), a multi-objective decision approach is warranted to obtain balanced policies. To enhance future decision-making, we propose a deep multi-objective reinforcement learning approach by building upon a state-of-the-art algorithm called Pareto Conditioned Networks (PCN) to obtain a set of solutions for distinct outcomes of the decision problem. We consider different deconfinement strategies after the first Belgian lockdown within the COVID-19 pandemic and aim to minimize both COVID-19 cases (i.e., infections and hospitalizations) and the societal burden induced by the mitigation measures. As such, we connected a multi-objective Markov decision process with a stochastic compartment model designed to approximate the Belgian COVID-19 waves and explore reactive strategies. As these social mitigation measures are implemented in a continuous action space that modulates the contact matrix of the age-structured epidemic model, we extend PCN to this setting. We evaluate the solution set that PCN returns, and observe that it explored the whole range of possible social restrictions, leading to high-quality trade-offs, as it captured the problem dynamics. In this work, we demonstrate that multi-objective reinforcement learning adds value to epidemiological modeling and provides essential insights to balance mitigation policies.
Mathieu Reymond, Conor F. Hayes, Lander Willem, Roxana Radulescu, Steven Abrams, Diederik M. Roijers, Enda Howley, Patrick Mannion, Niel Hens, Ann Nowé, Pieter Libin
Expert Syst. Appl.9
2022 Inferring age-specific differences in susceptibility to and infectiousness upon SARS-CoV-2 infection based on Belgian social contact data
abstract
Several important aspects related to SARS-CoV-2 transmission are not well known due to a lack of appropriate data. However, mathematical and computational tools can be used to extract part of this information from the available data, like some hidden age-related characteristics. In this paper, we present a method to investigate age-specific differences in transmission parameters related to susceptibility to and infectiousness upon contracting SARS-CoV-2 infection. More specifically, we use panel-based social contact data from diary-based surveys conducted in Belgium combined with the next generation principle to infer the relative incidence and we compare this to real-life incidence data. Comparing these two allows for the estimation of age-specific transmission parameters. Our analysis implies the susceptibility in children to be around half of the susceptibility in adults, and even lower for very young children (preschooler). However, the probability of adults and the elderly to contract the infection is decreasing throughout the vaccination campaign, thereby modifying the picture over time.
Nicolas Franco, Pietro Coletti, Lander Willem, Leonardo Angeli, Adrien Lajot, Steven Abrams, Philippe Beutels, Christel Faes, Niel Hens
PLoS Comput. Biol.9
2022 EpiLPS: A fast and flexible Bayesian tool for estimation of the time-varying reproduction number
abstract
In infectious disease epidemiology, the instantaneous reproduction number [Formula: see text] is a time-varying parameter defined as the average number of secondary infections generated by an infected individual at time t. It is therefore a crucial epidemiological statistic that assists public health decision makers in the management of an epidemic. We present a new Bayesian tool (EpiLPS) for robust estimation of the time-varying reproduction number. The proposed methodology smooths the epidemic curve and allows to obtain (approximate) point estimates and credible intervals of [Formula: see text] by employing the renewal equation, using Bayesian P-splines coupled with Laplace approximations of the conditional posterior of the spline vector. Two alternative approaches for inference are presented: (1) an approach based on a maximum a posteriori argument for the model hyperparameters, delivering estimates of [Formula: see text] in only a few seconds; and (2) an approach based on a Markov chain Monte Carlo (MCMC) scheme with underlying Langevin dynamics for efficient sampling of the posterior target distribution. Case counts per unit of time are assumed to follow a negative binomial distribution to account for potential overdispersion in the data that would not be captured by a classic Poisson model. Furthermore, after smoothing the epidemic curve, a "plug-in'' estimate of the reproduction number can be obtained from the renewal equation yielding a closed form expression of [Formula: see text] as a function of the spline parameters. The approach is extremely fast and free of arbitrary smoothing assumptions. EpiLPS is applied on data of SARS-CoV-1 in Hong-Kong (2003), influenza A H1N1 (2009) in the USA and on the SARS-CoV-2 pandemic (2020-2021) for Belgium, Portugal, Denmark and France.
Oswaldo Gressani, Jacco Wallinga, Christian L. Althaus, Niel Hens, Christel Faes
PLoS Comput. Biol.4
2022 Different forms of superspreading lead to different outcomes: Heterogeneity in infectiousness and contact behavior relevant for the case of SARS-CoV-2
abstract
Superspreading events play an important role in the spread of several pathogens, such as SARS-CoV-2. While the basic reproduction number of the original Wuhan SARS-CoV-2 is estimated to be about 3 for Belgium, there is substantial inter-individual variation in the number of secondary cases each infected individual causes-with most infectious individuals generating no or only a few secondary cases, while about 20% of infectious individuals is responsible for 80% of new infections. Multiple factors contribute to the occurrence of superspreading events: heterogeneity in infectiousness, individual variations in susceptibility, differences in contact behavior, and the environment in which transmission takes place. While superspreading has been included in several infectious disease transmission models, research into the effects of different forms of superspreading on the spread of pathogens remains limited. To disentangle the effects of infectiousness-related heterogeneity on the one hand and contact-related heterogeneity on the other, we implemented both forms of superspreading in an individual-based model describing the transmission and spread of SARS-CoV-2 in a synthetic Belgian population. We considered its impact on viral spread as well as on epidemic resurgence after a period of social distancing. We found that the effects of superspreading driven by heterogeneity in infectiousness are different from the effects of superspreading driven by heterogeneity in contact behavior. On the one hand, a higher level of infectiousness-related heterogeneity results in a lower risk of an outbreak persisting following the introduction of one infected individual into the population. Outbreaks that did persist led to fewer total cases and were slower, with a lower peak which occurred at a later point in time, and a lower herd immunity threshold. Finally, the risk of resurgence of an outbreak following a period of lockdown decreased. On the other hand, when contact-related heterogeneity was high, this also led to fewer cases in total during persistent outbreaks, but caused outbreaks to be more explosive in regard to other aspects (such as higher peaks which occurred earlier, and a higher herd immunity threshold). Finally, the risk of resurgence of an outbreak following a period of lockdown increased. We found that these effects were conserved when testing combinations of infectiousness-related and contact-related heterogeneity.
Elise J. Kuylen, Andrea Torneri, Lander Willem, Pieter Libin, Steven Abrams, Pietro Coletti, Nicolas Franco, Frederik Verelst, Philippe Beutels, Jori Liesenborgs, Niel Hens
PLoS Comput. Biol.11
2021 Assessing the feasibility and effectiveness of household-pooled universal testing to control COVID-19 epidemics
abstract
Outbreaks of SARS-CoV-2 are threatening the health care systems of several countries around the world. The initial control of SARS-CoV-2 epidemics relied on non-pharmaceutical interventions, such as social distancing, teleworking, mouth masks and contact tracing. However, as pre-symptomatic transmission remains an important driver of the epidemic, contact tracing efforts struggle to fully control SARS-CoV-2 epidemics. Therefore, in this work, we investigate to what extent the use of universal testing, i.e., an approach in which we screen the entire population, can be utilized to mitigate this epidemic. To this end, we rely on PCR test pooling of individuals that belong to the same households, to allow for a universal testing procedure that is feasible with the limited testing capacity. We evaluate two isolation strategies: on the one hand pool isolation, where we isolate all individuals that belong to a positive PCR test pool, and on the other hand individual isolation, where we determine which of the individuals that belong to the positive PCR pool are positive, through an additional testing step. We evaluate this universal testing approach in the STRIDE individual-based epidemiological model in the context of the Belgian COVID-19 epidemic. As the organisation of universal testing will be challenging, we discuss the different aspects related to sample extraction and PCR testing, to demonstrate the feasibility of universal testing when a decentralized testing approach is used. We show through simulation, that weekly universal testing is able to control the epidemic, even when many of the contact reductions are relieved. Finally, our model shows that the use of universal testing in combination with stringent contact reductions could be considered as a strategy to eradicate the virus.
Pieter Libin, Lander Willem, Timothy Verstraeten, Andrea Torneri, Joris Vanderlocht, Niel Hens
PLoS Comput. Biol.6
2021 On realized serial and generation intervals given control measures: The COVID-19 pandemic case
abstract
The SARS-CoV-2 pathogen is currently spreading worldwide and its propensity for presymptomatic and asymptomatic transmission makes it difficult to control. The control measures adopted in several countries aim at isolating individuals once diagnosed, limiting their social interactions and consequently their transmission probability. These interventions, which have a strong impact on the disease dynamics, can affect the inference of the epidemiological quantities. We first present a theoretical explanation of the effect caused by non-pharmaceutical intervention measures on the mean serial and generation intervals. Then, in a simulation study, we vary the assumed efficacy of control measures and quantify the effect on the mean and variance of realized generation and serial intervals. The simulation results show that the realized serial and generation intervals both depend on control measures and their values contract according to the efficacy of the intervention strategies. Interestingly, the mean serial interval differs from the mean generation interval. The deviation between these two values depends on two factors. First, the number of undiagnosed infectious individuals. Second, the relationship between infectiousness, symptom onset and timing of isolation. Similarly, the standard deviations of realized serial and generation intervals do not coincide, with the former shorter than the latter on average. The findings of this study are directly relevant to estimates performed for the current COVID-19 pandemic. In particular, the effective reproduction number is often inferred using both daily incidence data and the generation interval. Failing to account for either contraction or mis-specification by using the serial interval could lead to biased estimates of the effective reproduction number. Consequently, this might affect the choices made by decision makers when deciding which control measures to apply based on the value of the quantity thereof.
Andrea Torneri, Pieter Libin, Gianpaolo Scalia Tomba, Christel Faes, James G. Wood, Niel Hens
PLoS Comput. Biol.6
2018 Heterogeneous computing for epidemiological model fitting and simulation
abstract
BACKGROUND: Over the last years, substantial effort has been put into enhancing our arsenal in fighting epidemics from both technological and theoretical perspectives with scientists from different fields teaming up for rapid assessment of potentially urgent situations. This paper focusses on the computational aspects of infectious disease models and applies commonly available graphics processing units (GPUs) for the simulation of these models. However, fully utilizing the resources of both CPUs and GPUs requires a carefully balanced heterogeneous approach. RESULTS: The contribution of this paper is twofold. First, an efficient GPU implementation for evaluating a small-scale ODE model; here, the basic S(usceptible)-I(nfected)-R(ecovered) model, is discussed. Second, an asynchronous particle swarm optimization (PSO) implementation is proposed where batches of particles are sent asynchronously from the host (CPU) to the GPU for evaluation. The ultimate goal is to infer model parameters that enable the model to correctly describe observed data. The particles of the PSO algorithm are candidate parameters of the model; finding the right one is a matter of optimizing the likelihood function which quantifies how well the model describes the observed data. By employing a heterogeneous approach, in which both CPU and GPU are kept busy with useful work, speedups of 10 to 12 times can be achieved on a moderate machine with a high-end consumer GPU as compared to a high-end system with 32 CPU cores. CONCLUSIONS: Utilizing GPUs for parameter inference can bring considerable increases in performance using average host systems with high-end consumer GPUs. Future studies should evaluate the benefit of using newer CPU and GPU architectures as well as applying this method to more complex epidemiological scenarios.
Thomas Kovac, Tom Haber, Frank Van Reeth, Niel Hens
BMC Bioinform.4
2016 Estimating Time of Infection Using Prior Serological and Individual Information Can Greatly Improve Incidence Estimation of Human and Wildlife Infections
abstract
Diseases of humans and wildlife are typically tracked and studied through incidence, the number of new infections per time unit. Estimating incidence is not without difficulties, as asymptomatic infections, low sampling intervals and low sample sizes can introduce large estimation errors. After infection, biomarkers such as antibodies or pathogens often change predictably over time, and this temporal pattern can contain information about the time since infection that could improve incidence estimation. Antibody level and avidity have been used to estimate time since infection and to recreate incidence, but the errors on these estimates using currently existing methods are generally large. Using a semi-parametric model in a Bayesian framework, we introduce a method that allows the use of multiple sources of information (such as antibody level, pathogen presence in different organs, individual age, season) for estimating individual time since infection. When sufficient background data are available, this method can greatly improve incidence estimation, which we show using arenavirus infection in multimammate mice as a test case. The method performs well, especially compared to the situation in which seroconversion events between sampling sessions are the main data source. The possibility to implement several sources of information allows the use of data that are in many cases already available, which means that existing incidence data can be improved without the need for additional sampling efforts or laboratory assays.
Benny Borremans, Niel Hens, Philippe Beutels, Herwig Leirs, Jonas Reijniers
PLoS Comput. Biol.2
2015 Optimizing agent-based transmission models for infectious diseases
abstract
BACKGROUND: Infectious disease modeling and computational power have evolved such that large-scale agent-based models (ABMs) have become feasible. However, the increasing hardware complexity requires adapted software designs to achieve the full potential of current high-performance workstations. RESULTS: We have found large performance differences with a discrete-time ABM for close-contact disease transmission due to data locality. Sorting the population according to the social contact clusters reduced simulation time by a factor of two. Data locality and model performance can also be improved by storing person attributes separately instead of using person objects. Next, decreasing the number of operations by sorting people by health status before processing disease transmission has also a large impact on model performance. Depending of the clinical attack rate, target population and computer hardware, the introduction of the sort phase decreased the run time from 26% up to more than 70%. We have investigated the application of parallel programming techniques and found that the speedup is significant but it drops quickly with the number of cores. We observed that the effect of scheduling and workload chunk size is model specific and can make a large difference. CONCLUSIONS: Investment in performance optimization of ABM simulator code can lead to significant run time reductions. The key steps are straightforward: the data structure for the population and sorting people on health status before effecting disease propagation. We believe these conclusions to be valid for a wide range of infectious disease ABMs. We recommend that future studies evaluate the impact of data management, algorithmic procedures and parallelization on model performance.
Lander Willem, Sean Stijven, Engelbert Tijskens, Philippe Beutels, Niel Hens, Jan Broeckhove
BMC Bioinform.5
2014 Active Learning to Understand Infectious Disease Models and Improve Policy Making
abstract
Modeling plays a major role in policy making, especially for infectious disease interventions but such models can be complex and computationally intensive. A more systematic exploration is needed to gain a thorough systems understanding. We present an active learning approach based on machine learning techniques as iterative surrogate modeling and model-guided experimentation to systematically analyze both common and edge manifestations of complex model runs. Symbolic regression is used for nonlinear response surface modeling with automatic feature selection. First, we illustrate our approach using an individual-based model for influenza vaccination. After optimizing the parameter space, we observe an inverse relationship between vaccination coverage and cumulative attack rate reinforced by herd immunity. Second, we demonstrate the use of surrogate modeling techniques on input-response data from a deterministic dynamic model, which was designed to explore the cost-effectiveness of varicella-zoster virus vaccination. We use symbolic regression to handle high dimensionality and correlated inputs and to identify the most influential variables. Provided insight is used to focus research, reduce dimensionality and decrease decision uncertainty. We conclude that active learning is needed to fully understand complex systems behavior. Surrogate models can be readily explored at no computational expense, and can also be used as emulator to improve rapid policy making in various settings.
Lander Willem, Sean Stijven, Ekaterina Vladislavleva, Jan Broeckhove, Philippe Beutels, Niel Hens
PLoS Comput. Biol.6
2012 Living on Three Time Scales: The Dynamics of Plasma Cell and Antibody Populations Illustrated for Hepatitis A Virus
abstract
Understanding the mechanisms involved in long-term persistence of humoral immunity after natural infection or vaccination is challenging and crucial for further research in immunology, vaccine development as well as health policy. Long-lived plasma cells, which have recently been shown to reside in survival niches in the bone marrow, are instrumental in the process of immunity induction and persistence. We developed a mathematical model, assuming two antibody-secreting cell subpopulations (short- and long-lived plasma cells), to analyze the antibody kinetics after HAV-vaccination using data from two long-term follow-up studies. Model parameters were estimated through a hierarchical nonlinear mixed-effects model analysis. Long-term individual predictions were derived from the individual empirical parameters and were used to estimate the mean time to immunity waning. We show that three life spans are essential to explain the observed antibody kinetics: that of the antibodies (around one month), the short-lived plasma cells (several months) and the long-lived plasma cells (decades). Although our model is a simplified representation of the actual mechanisms that govern individual immune responses, the level of agreement between long-term individual predictions and observed kinetics is reassuringly close. The quantitative assessment of the time scales over which plasma cells and antibodies live and interact provides a basis for further quantitative research on immunology, with direct consequences for understanding the epidemiology of infectious diseases, and for timing serum sampling in clinical trials of vaccines.
Mathieu Andraud, Olivier Lejeune, Jammbe Z. Musoro, Benson Ogunjimi, Philippe Beutels, Niel Hens
PLoS Comput. Biol.6