EDBT 2026 Demo / reviewers in the wild / expert
Philippe Lemey
dblp:43/3909
· DBLP profile ↗
24ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-2826-5353ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SERAPHIM 2.0: an extended toolbox for studying phylogenetically informed movementsabstractSUMMARY: We report the second version of the R package "seraphim", a toolbox developed to process and analyze the output of spatially explicit phylogeographic reconstructions. This approach - also known as continuous phylogeographic inference - is commonly used in molecular epidemiology to reconstruct the dispersal history and spatiotemporal dynamics of rapidly evolving pathogens. The "seraphim" package now implements a broad range of features including (i) visualization of phylogeographic inferences, (ii) estimation of lineage dispersal metrics, (iii) several phylogeographic simulators, and (iv) hypothesis testing procedures to investigate the impact of environmental factors on variables such as diffusion velocity, dispersal location, and dispersal frequency of phylogenetic lineages. AVAILABILITY AND IMPLEMENTATION: The package is openly available (https://github.com/sdellicour/seraphim) along with a series of tutorials describing the different analytical procedures it implements. Simon Dellicour, Nuno R. Faria, Rebecca Rose, Philippe Lemey, Oliver G. Pybus |
Bioinform. | 4 |
| 2026 | nf-core/viralmetagenome: A novel pipeline for untargeted viral genome reconstructionabstractMOTIVATION: Reconstructing eukaryotic viral genomes from metagenomic data is challenging due to their extensive diversity and potential genome segmentation. Current approaches often rely on labor-intensive manual curation for reference selection and scaffolding, limiting scalability for large studies or rapid outbreak response. We address the critical need for an automated, scalable pipeline for efficient viral metagenomic analysis without manual intervention. RESULTS: We present nf-core/viralmetagenome, a comprehensive Nextflow pipeline for the untargeted reconstruction and variant analysis of eukaryotic DNA and RNA viruses from short-read metagenomic or hybridisation capture enriched samples. The pipeline automates the entire process from read preprocessing to consensus generation, integrating multiple de novo assemblers, automated reference selection, and iterative consensus refinement. It features robust quality control, extensive documentation, and seamless portability via Docker and Singularity. We validated the pipeline on diverse simulated and real datasets, demonstrating its ability to recover high-quality genomes from complex metagenomic samples and resolve co-infections, making it a powerful tool for viral surveillance. AVAILABILITY: nf-core/viralmetagenome is freely available at https://github.com/nf-core/viralmetagenome with comprehensive documentation at https://nf-co.re/viralmetagenome. Archival code repository snapshots are published at zenodo with doi: https://doi.org/10.5281/zenodo.17524074. Joon Klaps, Philippe Lemey, Magda Bletsa, Liana Eleni Kafetzopoulou |
Bioinform. | 2 |
| 2026 | ViMOP: a user-friendly and field-applicable pipeline for untargeted viral genome nanopore sequencingabstractMOTIVATION: Untargeted, also known as metagenomic, nanopore sequencing is a powerful tool for virus genomic surveillance, particularly in resource-limited settings and when paired with the portability of the MinION device (Oxford Nanopore Technologies, ONT). However, a major bottleneck for global access is the absence of a user-friendly software capable of efficiently analyzing untargeted nanopore sequencing data to generate high-quality consensus genomes. RESULTS: We share ViMOP, a pipeline built on our long-term experience in nanopore field sequencing. The pipeline emphasizes field user-friendliness, flexibility and versatility to analyze reads generated directly from human clinical samples. The software assembles de novo contigs, matches contigs to known viral references and uses them to assemble consensus genomes. Executed with a single Nextflow command or via the EPI2ME Desktop interface (ONT), results are summarized in an HTML report. ViMOP, through its user-centered design, lowers the barrier to high-quality virus genome reconstruction and advances capacity for genomic surveillance. AVAILABILITY AND IMPLEMENTATION: ViMOP is freely available for non-commercial use (https://github.com/opr-group-bnitm/vimop and https://zenodo.org/records/17913089), along with the associated database (https://zenodo.org/records/17652512), the scripts used to generate it (https://zenodo.org/records/17632662) and benchmarking code (https://zenodo.org/records/17633185). Nils Peter Petersen, Mia Le, Annick Renevey, Ehizojie Emua, Sarah Ryter, Giuditta Annibaldis, Jacob Camara, Sanaba Boumbaly, Cyril Erameh, Tanja Laske, Jan Baumbach, Philippe Lemey, Stephan Günther 0003, Sophie Duraffour, Liana Eleni Kafetzopoulou |
Bioinform. | 12 |
| 2025 | HIPSTR: highest independent posterior subtree reconstruction in TreeAnnotator XabstractSUMMARY: In Bayesian phylogenetic and phylodynamic studies, it is common to summarize the posterior distribution of trees with a time-calibrated summary phylogeny. While the maximum clade credibility (MCC) tree is often used for this purpose, we here show that a novel summary tree method-the highest independent posterior subtree reconstruction, or (HIPSTR)-contains consistently higher supported clades over MCC. We also provide faster computational routines for estimating both summary trees in an updated version of TreeAnnotator X, an open-source software program that summarizes the information from a sample of trees and returns many helpful statistics such as individual clade credibilities contained in the summary tree. RESULTS: HIPSTR and MCC reconstructions on two Ebola virus and two SARS-CoV-2 datasets show that HIPSTR yields summary trees that consistently contain clades with higher support compared to MCC trees. The MCC trees regularly fail to include several clades with very high posterior probability (≥0.95) as well as a large number of clades with moderate to high posterior probability (≥50%), whereas HIPSTR-in particular its majority-rule extension MrHIPSTR-achieves near-perfect performance in this respect. HIPSTR and MrHIPSTR also exhibit favourable computational performance over MCC in TreeAnnotator X. Comparison to the recent CCD0-MAP algorithm yielded mixed results and requires a more in-depth investigation in follow-up studies. AVAILABILITY AND IMPLEMENTATION: TreeAnnotator X is available as part of the BEAST X (v10.5.0) software package, available at https://github.com/beast-dev/beast-mcmc/releases, and on Zenodo (DOI: https://doi.org/10.5281/zenodo.4895234). Guy Baele, Luiz Max Carvalho, Marius Brusselmans, Gytis Dudas, John T. McCrone, Philippe Lemey, Marc A. Suchard, Andrew Rambaut |
Bioinform. | 7 |
| 2024 | Many-core algorithms for high-dimensional gradients on phylogenetic treesabstractMOTIVATION: Advancements in high-throughput genomic sequencing are delivering genomic pathogen data at an unprecedented rate, positioning statistical phylogenetics as a critical tool to monitor infectious diseases globally. This rapid growth spurs the need for efficient inference techniques, such as Hamiltonian Monte Carlo (HMC) in a Bayesian framework, to estimate parameters of these phylogenetic models where the dimensions of the parameters increase with the number of sequences N. HMC requires repeated calculation of the gradient of the data log-likelihood with respect to (wrt) all branch-length-specific (BLS) parameters that traditionally takes O(N2) operations using the standard pruning algorithm. A recent study proposes an approach to calculate this gradient in O(N), enabling researchers to take advantage of gradient-based samplers such as HMC. The CPU implementation of this approach makes the calculation of the gradient computationally tractable for nucleotide-based models but falls short in performance for larger state-space size models, such as Markov-modulated and codon models. Here, we describe novel massively parallel algorithms to calculate the gradient of the log-likelihood wrt all BLS parameters that take advantage of graphics processing units (GPUs) and result in many fold higher speedups over previous CPU implementations. RESULTS: We benchmark these GPU algorithms on three computing systems using three evolutionary inference examples exploring complete genomes from 997 dengue viruses, 62 carnivore mitochondria and 49 yeasts, and observe a >128-fold speedup over the CPU implementation for codon-based models and >8-fold speedup for nucleotide-based models. As a practical demonstration, we also estimate the timing of the first introduction of West Nile virus into the continental Unites States under a codon model with a relaxed molecular clock from 104 full viral genomes, an inference task previously intractable. AVAILABILITY AND IMPLEMENTATION: We provide an implementation of our GPU algorithms in BEAGLE v4.0.0 (https://github.com/beagle-dev/beagle-lib), an open-source library for statistical phylogenetics that enables parallel calculations on multi-core CPUs and GPUs. We employ a BEAGLE-implementation using the Bayesian phylogenetics framework BEAST (https://github.com/beast-dev/beast-mcmc). Karthik Gangavarapu, Guy Baele, Mathieu Fourment, Philippe Lemey, Frederick A. Matsen IV, Marc A. Suchard |
Bioinform. | 5 |
| 2024 | spread.gl: visualizing pathogen dispersal in a high-performance browser applicationabstractMOTIVATION: Bayesian phylogeographic analyses are pivotal in reconstructing the spatio-temporal dispersal histories of pathogens. However, interpreting the complex outcomes of phylogeographic reconstructions requires sophisticated visualization tools. RESULTS: To meet this challenge, we developed spread.gl, an open-source, feature-rich browser application offering a smooth and intuitive visualization tool for both discrete and continuous phylogeographic inferences, including the animation of pathogen geographic dispersal through time. Spread.gl can render and combine the visualization of multiple layers that contain information extracted from the input phylogeny and diverse environmental data layers, enabling researchers to explore which environmental factors may have impacted pathogen dispersal patterns before conducting formal testing. We showcase the visualization features of spread.gl with representative examples, including the smooth animation of a phylogeographic reconstruction based on >17 000 SARS-CoV-2 genomic sequences. AVAILABILITY AND IMPLEMENTATION: Source code, installation instructions, example input data, and outputs of spread.gl are accessible at https://github.com/GuyBaele/SpreadGL. Nena Bollen, Samuel L. Hong, Marius Brusselmans, Fabiana Gambaro, Joon Klaps, Marc A. Suchard, Andrew Rambaut, Philippe Lemey, Simon Dellicour, Guy Baele |
Bioinform. | 9 |
| 2023 | Accelerating Bayesian inference of dependency between mixed-type biological traitsabstractInferring dependencies between mixed-type biological traits while accounting for evolutionary relationships between specimens is of great scientific interest yet remains infeasible when trait and specimen counts grow large. The state-of-the-art approach uses a phylogenetic multivariate probit model to accommodate binary and continuous traits via a latent variable framework, and utilizes an efficient bouncy particle sampler (BPS) to tackle the computational bottleneck-integrating many latent variables from a high-dimensional truncated normal distribution. This approach breaks down as the number of specimens grows and fails to reliably characterize conditional dependencies between traits. Here, we propose an inference pipeline for phylogenetic probit models that greatly outperforms BPS. The novelty lies in 1) a combination of the recent Zigzag Hamiltonian Monte Carlo (Zigzag-HMC) with linear-time gradient evaluations and 2) a joint sampling scheme for highly correlated latent variables and correlation matrix elements. In an application exploring HIV-1 evolution from 535 viruses, the inference requires joint sampling from an 11,235-dimensional truncated normal and a 24-dimensional covariance matrix. Our method yields a 5-fold speedup compared to BPS and makes it possible to learn partial correlations between candidate viral mutations and virulence. Computational speedup now enables us to tackle even larger problems: we study the evolution of influenza H1N1 glycosylations on around 900 viruses. For broader applicability, we extend the phylogenetic probit model to incorporate categorical traits, and demonstrate its use to study Aquilegia flower and pollinator co-evolution. Zhenyu Zhang 0019, Akihiko Nishimura, Nídia S. Trovão, Joshua L. Cherry, Andrew J. Holbrook, Philippe Lemey, Marc A. Suchard |
PLoS Comput. Biol. | 7 |
| 2020 | Incorporating heterogeneous sampling probabilities in continuous phylogeographic inference - Application to H5N1 spread in the Mekong regionabstractMOTIVATION: The potentially low precision associated with the geographic origin of sampled sequences represents an important limitation for spatially explicit (i.e. continuous) phylogeographic inference of fast-evolving pathogens such as RNA viruses. A substantial proportion of publicly available sequences is geo-referenced at broad spatial scale such as the administrative unit of origin, rather than more precise locations (e.g. geographic coordinates). Most frequently, such sequences are either discarded prior to continuous phylogeographic inference or arbitrarily assigned to the geographic coordinates of the centroid of their administrative area of origin for lack of a better alternative. RESULTS: We here implement and describe a new approach that allows to incorporate heterogeneous prior sampling probabilities over a geographic area. External data, such as outbreak locations, are used to specify these prior sampling probabilities over a collection of sub-polygons. We apply this new method to the analysis of highly pathogenic avian influenza H5N1 clade data in the Mekong region. Our method allows to properly include, in continuous phylogeographic analyses, H5N1 sequences that are only associated with large administrative areas of origin and assign them with more accurate locations. Finally, we use continuous phylogeographic reconstructions to analyse the dispersal dynamics of different H5N1 clades and investigate the impact of environmental factors on lineage dispersal velocities. AVAILABILITY AND IMPLEMENTATION: Our new method allowing heterogeneous sampling priors for continuous phylogeographic inference is implemented in the open-source multi-platform software package BEAST 1.10. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Simon Dellicour, Philippe Lemey, Jean Artois, Tommy T. Lam, Alice Fusaro, Isabella Monne, Giovanni Cattoli, Dmitry Kuznetsov, Ioannis Xenarios, Gwenaelle Dauphin, Wantanee Kalpravidh, Sophie von Dobschütz, Filip Claes, Scott H. Newman, Marc A. Suchard, Guy Baele, Marius Gilbert |
Bioinform. | 2 |
| 2018 | Bayesian Best-Arm Identification for Selecting Influenza Mitigation Strategies
Pieter Libin, Timothy Verstraeten, Diederik M. Roijers, Jelena Grujic, Kristof Theys, Philippe Lemey, Ann Nowé |
ECML/PKDD (3) | 6 |
| 2017 | Adaptive MCMC in Bayesian phylogenetics: an application to analyzing partitioned data in BEASTabstractMOTIVATION: Advances in sequencing technology continue to deliver increasingly large molecular sequence datasets that are often heavily partitioned in order to accurately model the underlying evolutionary processes. In phylogenetic analyses, partitioning strategies involve estimating conditionally independent models of molecular evolution for different genes and different positions within those genes, requiring a large number of evolutionary parameters that have to be estimated, leading to an increased computational burden for such analyses. The past two decades have also seen the rise of multi-core processors, both in the central processing unit (CPU) and Graphics processing unit processor markets, enabling massively parallel computations that are not yet fully exploited by many software packages for multipartite analyses. RESULTS: We here propose a Markov chain Monte Carlo (MCMC) approach using an adaptive multivariate transition kernel to estimate in parallel a large number of parameters, split across partitioned data, by exploiting multi-core processing. Across several real-world examples, we demonstrate that our approach enables the estimation of these multipartite parameters more efficiently than standard approaches that typically use a mixture of univariate transition kernels. In one case, when estimating the relative rate parameter of the non-coding partition in a heterochronous dataset, MCMC integration efficiency improves by > 14-fold. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of the BEAST code base, a widely used open source software package to perform Bayesian phylogenetic inference. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Guy Baele, Philippe Lemey, Andrew Rambaut, Marc A. Suchard |
Bioinform. | 2 |
| 2016 | SERAPHIM: studying environmental rasters and phylogenetically informed movementsabstractSERAPHIM ("Studying Environmental Rasters and PHylogenetically Informed Movements") is a suite of computational methods developed to study phylogenetic reconstructions of spatial movement in an environmental context. SERAPHIM extracts the spatio-temporal information contained in estimated phylogenetic trees and uses this information to calculate summary statistics of spatial spread and to visualize dispersal history. Most importantly, SERAPHIM enables users to study the impact of customized environmental variables on the spread of the study organism. Specifically, given an environmental raster, SERAPHIM computes environmental "weights" for each phylogeny branch, which represent the degree to which the environmental variable impedes (or facilitates) lineage movement. Correlations between movement duration and these environmental weights are then assessed, and the statistical significances of these correlations are evaluated using null distributions generated by a randomization procedure. SERAPHIM can be applied to any phylogeny whose nodes are annotated with spatial and temporal information. At present, such phylogenies are most often found in the field of emerging infectious diseases, but will become increasingly common in other biological disciplines as population genomic data grows. AVAILABILITY AND IMPLEMENTATION: SERAPHIM 1.0 is freely available from http://evolve.zoo.ox.ac.uk/ R package, source code, example files, tutorials and a manual are also available from this website. CONTACT: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Simon Dellicour, Rebecca Rose, Nuno R. Faria, Philippe Lemey, Oliver G. Pybus |
Bioinform. | 4 |
| 2014 | πBUSS: a parallel BEAST/BEAGLE utility for sequence simulation under complex evolutionary scenariosabstractBACKGROUND: Simulated nucleotide or amino acid sequences are frequently used to assess the performance of phylogenetic reconstruction methods. BEAST, a Bayesian statistical framework that focuses on reconstructing time-calibrated molecular evolutionary processes, supports a wide array of evolutionary models, but lacked matching machinery for simulation of character evolution along phylogenies. RESULTS: We present a flexible Monte Carlo simulation tool, called πBUSS, that employs the BEAGLE high performance library for phylogenetic computations to rapidly generate large sequence alignments under complex evolutionary models. πBUSS sports a user-friendly graphical user interface (GUI) that allows combining a rich array of models across an arbitrary number of partitions. A command-line interface mirrors the options available through the GUI and facilitates scripting in large-scale simulation studies. πBUSS may serve as an easy-to-use, standard sequence simulation tool, but the available models and data types are particularly useful to assess the performance of complex BEAST inferences. The connection with BEAST is further strengthened through the use of a common extensible markup language (XML), allowing to specify also more advanced evolutionary models. To support simulation under the latter, as well as to support simulation and analysis in a single run, we also add the πBUSS core simulation routine to the list of BEAST XML parsers. CONCLUSIONS: πBUSS offers a unique combination of flexibility and ease-of-use for sequence simulation under realistic evolutionary scenarios. Through different interfaces, πBUSS supports simulation studies ranging from modest endeavors for illustrative purposes to complex and large-scale assessments of evolutionary inference procedures. Applications are not restricted to the BEAST framework, or even time-measured evolutionary histories, and πBUSS can be connected to various other programs using standard input and output format. Filip Bielejec, Philippe Lemey, Guy Baele, Andrew Rambaut, Marc A. Suchard |
BMC Bioinform. | 2 |
| 2014 | The Genealogical Population Dynamics of HIV-1 in a Large Transmission Chain: Bridging within and among Host Evolutionary RatesabstractTransmission lies at the interface of human immunodeficiency virus type 1 (HIV-1) evolution within and among hosts and separates distinct selective pressures that impose differences in both the mode of diversification and the tempo of evolution. In the absence of comprehensive direct comparative analyses of the evolutionary processes at different biological scales, our understanding of how fast within-host HIV-1 evolutionary rates translate to lower rates at the between host level remains incomplete. Here, we address this by analyzing pol and env data from a large HIV-1 subtype C transmission chain for which both the timing and the direction is known for most transmission events. To this purpose, we develop a new transmission model in a Bayesian genealogical inference framework and demonstrate how to constrain the viral evolutionary history to be compatible with the transmission history while simultaneously inferring the within-host evolutionary and population dynamics. We show that accommodating a transmission bottleneck affords the best fit our data, but the sparse within-host HIV-1 sampling prevents accurate quantification of the concomitant loss in genetic diversity. We draw inference under the transmission model to estimate HIV-1 evolutionary rates among epidemiologically-related patients and demonstrate that they lie in between fast intra-host rates and lower rates among epidemiologically unrelated individuals infected with HIV subtype C. Using a new molecular clock approach, we quantify and find support for a lower evolutionary rate along branches that accommodate a transmission event or branches that represent the entire backbone of transmitted lineages in our transmission history. Finally, we recover the rate differences at the different biological scales for both synonymous and non-synonymous substitution rates, which is only compatible with the 'store and retrieve' hypothesis positing that viruses stored early in latently infected cells preferentially transmit or establish new infections upon reactivation. Bram Vrancken, Andrew Rambaut, Marc A. Suchard, Alexei J. Drummond, Guy Baele, Inge Derdelinckx, Eric Van Wijngaerden, Anne-Mieke Vandamme, Kristel Van Laethem, Philippe Lemey |
PLoS Comput. Biol. | 10 |
| 2013 | Bayesian evolutionary model testing in the phylogenomics era: matching model complexity with computational efficiencyabstractMOTIVATION: The advent of new sequencing technologies has led to increasing amounts of data being available to perform phylogenetic analyses, with genomic data giving rise to the field of phylogenomics. High-performance computing is becoming an indispensable research tool to fit complex evolutionary models, which take into account specific genomic properties, to large datasets. Here, we perform an extensive Bayesian phylogenetic model selection study, comparing codon and nucleotide substitution models, including codon position partitioning for nucleotide data as well gene-specific substitution models for both data types. For the best fitting partitioned models, we also compare independent partitioning with standard diffuse prior specification to conditional partitioning via hierarchical prior specification. To compare the different models, we use state-of-the-art marginal likelihood estimation techniques, including path sampling and stepping-stone sampling. RESULTS: We show that a full codon model best describes the features of a whole mitochondrial genome dataset, consisting of 12 protein-coding genes, but only when each gene is allowed to evolve under a separate codon model. However, when using hierarchical prior specification for the partition-specific parameters instead of independent diffuse priors, codon position partitioned nucleotide models can still outperform standard codon models. We demonstrate the feasibility of fitting such a combination of complex models using the BEAGLE library for BEAST in combination with recent graphics cards. We argue that development and use of such models needs to be accompanied by state-of-the-art marginal likelihood estimators because the more traditional and computationally less demanding estimators do not offer adequate accuracy. Guy Baele, Philippe Lemey |
Bioinform. | 2 |
| 2013 | Make the most of your samples: Bayes factor estimators for high-dimensional models of sequence evolutionabstractBACKGROUND: Accurate model comparison requires extensive computation times, especially for parameter-rich models of sequence evolution. In the Bayesian framework, model selection is typically performed through the evaluation of a Bayes factor, the ratio of two marginal likelihoods (one for each model). Recently introduced techniques to estimate (log) marginal likelihoods, such as path sampling and stepping-stone sampling, offer increased accuracy over the traditional harmonic mean estimator at an increased computational cost. Most often, each model's marginal likelihood will be estimated individually, which leads the resulting Bayes factor to suffer from errors associated with each of these independent estimation processes. RESULTS: We here assess the original 'model-switch' path sampling approach for direct Bayes factor estimation in phylogenetics, as well as an extension that uses more samples, to construct a direct path between two competing models, thereby eliminating the need to calculate each model's marginal likelihood independently. Further, we provide a competing Bayes factor estimator using an adaptation of the recently introduced stepping-stone sampling algorithm and set out to determine appropriate settings for accurately calculating such Bayes factors, with context-dependent evolutionary models as an example. While we show that modest efforts are required to roughly identify the increase in model fit, only drastically increased computation times ensure the accuracy needed to detect more subtle details of the evolutionary process. CONCLUSIONS: We show that our adaptation of stepping-stone sampling for direct Bayes factor calculation outperforms the original path sampling approach as well as an extension that exploits more samples. Our proposed approach for Bayes factor estimation also has preferable statistical properties over the use of individual marginal likelihood estimates for both models under comparison. Assuming a sigmoid function to determine the path between two competing models, we provide evidence that a single well-chosen sigmoid shape value requires less computational efforts in order to approximate the true value of the (log) Bayes factor compared to the original approach. We show that the (log) Bayes factors calculated using path sampling and stepping-stone sampling differ drastically from those estimated using either of the harmonic mean estimators, supporting earlier claims that the latter systematically overestimate the performance of high-dimensional models, which we show can lead to erroneous conclusions. Based on our results, we argue that highly accurate estimation of differences in model fit for high-dimensional models requires much more computational effort than suggested in recent studies on marginal likelihood estimation. Guy Baele, Philippe Lemey, Stijn Vansteelandt |
BMC Bioinform. | 2 |
| 2012 | A counting renaissance: combining stochastic mapping and empirical Bayes to quickly detect amino acid sites under positive selectionabstractAbstract Motivation: Statistical methods for comparing relative rates of synonymous and non-synonymous substitutions maintain a central role in detecting positive selection. To identify selection, researchers often estimate the ratio of these relative rates () at individual alignment sites. Fitting a codon substitution model that captures heterogeneity in across sites provides a reliable way to perform such estimation, but it remains computationally prohibitive for massive datasets. By using crude estimates of the numbers of synonymous and non-synonymous substitutions at each site, counting approaches scale well to large datasets, but they fail to account for ancestral state reconstruction uncertainty and to provide site-specific estimates. Results: We propose a hybrid solution that borrows the computational strength of counting methods, but augments these methods with empirical Bayes modeling to produce a relatively fast and reliable method capable of estimating site-specific values in large datasets. Importantly, our hybrid approach, set in a Bayesian framework, integrates over the posterior distribution of phylogenies and ancestral reconstructions to quantify uncertainty about site-specific estimates. Simulations demonstrate that this method competes well with more-principled statistical procedures and, in some cases, even outperforms them. We illustrate the utility of our method using human immunodeficiency virus, feline panleukopenia and canine parvovirus evolution examples. Availability: Renaissance counting is implemented in the development branch of BEAST, freely available at http://code.google.com/p/beast-mcmc/. The method will be made available in the next public release of the package, including support to set up analyses in BEAUti. Contact: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Philippe Lemey, Volodymyr M. Minin, Filip Bielejec, Sergei L. Kosakovsky Pond, Marc A. Suchard |
Bioinform. | 1 |
| 2011 | SPREAD: spatial phylogenetic reconstruction of evolutionary dynamicsabstractSUMMARY: SPREAD is a user-friendly, cross-platform application to analyze and visualize Bayesian phylogeographic reconstructions incorporating spatial-temporal diffusion. The software maps phylogenies annotated with both discrete and continuous spatial information and can export high-dimensional posterior summaries to keyhole markup language (KML) for animation of the spatial diffusion through time in virtual globe software. In addition, SPREAD implements Bayes factor calculation to evaluate the support for hypotheses of historical diffusion among pairs of discrete locations based on Bayesian stochastic search variable selection estimates. SPREAD takes advantage of multicore architectures to process large joint posterior distributions of phylogenies and their spatial diffusion and produces visualizations as compelling and interpretable statistical summaries for the different spatial projections. AVAILABILITY: SPREAD is licensed under the GNU Lesser GPL and its source code is freely available as a GitHub repository: https://github.com/phylogeography/SPREAD CONTACT: [email protected]. Filip Bielejec, Andrew Rambaut, Marc A. Suchard, Philippe Lemey |
Bioinform. | 4 |
| 2010 | RDP3: a flexible and fast computer program for analyzing recombinationabstractAbstract Summary: RDP3 is a new version of the RDP program for characterizing recombination events in DNA-sequence alignments. Among other novelties, this version includes four new recombination analysis methods (3SEQ, VISRD, PHYLRO and LDHAT), new tests for recombination hot-spots, a range of matrix methods for visualizing over-all patterns of recombination within datasets and recombination-aware ancestral sequence reconstruction. Complementary to a high degree of analysis flow automation, RDP3 also has a highly interactive and detailed graphical user interface that enables more focused hands-on cross-checking of results with a wide variety of newly implemented phylogenetic tree construction and matrix-based recombination signal visualization methods. The new RDP3 can accommodate large datasets and is capable of analyzing alignments ranging in size from 1000×10 kilobase sequences to 20×2 megabase sequences within 48 h on a desktop PC. Availability: RDP3 is available for free from its web site http://darwin.uvigo.es/rdp/rdp.html Contact: [email protected] Supplementary information: The RDP3 program manual contains detailed descriptions of the various methods it implements and a step-by-step guide describing how best to use these. Darren P. Martin, Philippe Lemey, Martin Lott, Vincent Moulton, David Posada, Pierre Lefeuvre |
Bioinform. | 2 |
| 2010 | Estimating the individualized HIV-1 genetic barrier to resistance using a nelfinavir fitness landscapeabstractBACKGROUND: Failure on Highly Active Anti-Retroviral Treatment is often accompanied with development of antiviral resistance to one or more drugs included in the treatment. In general, the virus is more likely to develop resistance to drugs with a lower genetic barrier. Previously, we developed a method to reverse engineer, from clinical sequence data, a fitness landscape experienced by HIV-1 under nelfinavir (NFV) treatment. By simulation of evolution over this landscape, the individualized genetic barrier to NFV resistance may be estimated for an isolate. RESULTS: We investigated the association of estimated genetic barrier with risk of development of NFV resistance at virological failure, in 201 patients that were predicted fully susceptible to NFV at baseline, and found that a higher estimated genetic barrier was indeed associated with lower odds for development of resistance at failure (OR 0.62 (0.45 - 0.94), per additional mutation needed, p = .02). CONCLUSIONS: Thus, variation in individualized genetic barrier to NFV resistance may impact effective treatment options available after treatment failure. If similar results apply for other drugs, then estimated genetic barrier may be a new clinical tool for choice of treatment regimen, which allows consideration of available treatment options after virological failure. Kristof Theys, Koen Deforche, Gertjan Beheydt, Yves Moreau, Kristel Van Laethem, Philippe Lemey, Ricardo Camacho, Soo-Yon Rhee, Robert W. Shafer, Eric Van Wijngaerden, Anne-Mieke Vandamme |
BMC Bioinform. | 6 |
| 2009 | Identifying recombinants in human and primate immunodeficiency virus sequence alignments using quartet scanningabstractBACKGROUND: Recombination has a profound impact on the evolution of viruses, but characterizing recombination patterns in molecular sequences remains a challenging endeavor. Despite its importance in molecular evolutionary studies, identifying the sequences that exhibit such patterns has received comparatively less attention in the recombination detection framework. Here, we extend a quartet-mapping based recombination detection method to enable identification of recombinant sequences without prior specifications of either query and reference sequences. Through simulations we evaluate different recombinant identification statistics and significance tests. We compare the quartet approach with triplet-based methods that employ additional heuristic tests to identify parental and recombinant sequences. RESULTS: Analysis of phylogenetic simulations reveal that identifying the descendents of relatively old recombination events is a challenging task for all methods available, and that quartet scanning performs relatively well compared to the triplet based methods. The use of quartet scanning is further demonstrated by analyzing both well-established and putative HIV-1 recombinant strains. In agreement with recent findings, we provide evidence that the presumed circulating recombinant CRF02_AG is a 'pure' lineage, whereas the presumed parental lineage subtype G has a recombinant origin. We also demonstrate HIV-1 intrasubtype recombination, confirm the hybrid origin of SIV in chimpanzees and further disentangle the recombinant history of SIV lineages in a primate immunodeficiency virus data set. CONCLUSION: Quartet scanning makes a valuable addition to triplet-based methods for identifying recombinant sequences without prior specifications of either query and reference sequences. The new method is available in the VisRD v.3.0 package http://www.cmp.uea.ac.uk/~vlm/visrd. Philippe Lemey, Martin Lott, Darren P. Martin, Vincent Moulton |
BMC Bioinform. | 1 |
| 2009 | Bayesian Phylogeography Finds Its RootsabstractAs a key factor in endemic and epidemic dynamics, the geographical distribution of viruses has been frequently interpreted in the light of their genetic histories. Unfortunately, inference of historical dispersal or migration patterns of viruses has mainly been restricted to model-free heuristic approaches that provide little insight into the temporal setting of the spatial dynamics. The introduction of probabilistic models of evolution, however, offers unique opportunities to engage in this statistical endeavor. Here we introduce a Bayesian framework for inference, visualization and hypothesis testing of phylogeographic history. By implementing character mapping in a Bayesian software that samples time-scaled phylogenies, we enable the reconstruction of timed viral dispersal patterns while accommodating phylogenetic uncertainty. Standard Markov model inference is extended with a stochastic search variable selection procedure that identifies the parsimonious descriptions of the diffusion process. In addition, we propose priors that can incorporate geographical sampling distributions or characterize alternative hypotheses about the spatial dynamics. To visualize the spatial and temporal information, we summarize inferences using virtual globe software. We describe how Bayesian phylogeography compares with previous parsimony analysis in the investigation of the influenza A H5N1 origin and H5N1 epidemiological linkage among sampling localities. Analysis of rabies in West African dog populations reveals how virus diffusion may enable endemic maintenance through continuous epidemic cycles. From these analyses, we conclude that our phylogeographic framework will make an important asset in molecular epidemiology that can be easily generalized to infer biogeogeography from genetic data for many organisms. Philippe Lemey, Andrew Rambaut, Alexei J. Drummond, Marc A. Suchard |
PLoS Comput. Biol. | 1 |
| 2008 | Estimation of an in vivo fitness landscape experienced by HIV-1 under drug selective pressure useful for prediction of drug resistance evolution during treatmentabstractMOTIVATION: HIV-1 antiviral resistance is a major cause of antiviral treatment failure. The in vivo fitness landscape experienced by the virus in presence of treatment could in principle be used to determine both the susceptibility of the virus to the treatment and the genetic barrier to resistance. We propose a method to estimate this fitness landscape from cross-sectional clinical genetic sequence data of different subtypes, by reverse engineering the required selective pressure for HIV-1 sequences obtained from treatment naive patients, to evolve towards sequences obtained from treated patients. The method was evaluated for recovering 10 random fictive selective pressures in simulation experiments, and for modeling the selective pressure under treatment with the protease inhibitor nelfinavir. RESULTS: The estimated fitness function under nelfinavir treatment considered fitness contributions of 114 mutations at 48 sites. Estimated fitness correlated significantly with the in vitro resistance phenotype in 519 matched genotype-phenotype pairs (R(2) = 0.47 (0.41 - 0.54)) and variation in predicted evolution under nelfinavir selective pressure correlated significantly with observed in vivo evolution during nelfinavir treatment for 39 mutations (with FDR = 0.05). AVAILABILITY: The software is available on request from the authors, and data sets are available from http://jose.med.kuleuven.be/~kdforc0/nfv-fitness-data/. Koen Deforche, Ricardo Camacho, Kristel Van Laethem, Philippe Lemey, Andrew Rambaut, Yves Moreau, Anne-Mieke Vandamme |
Bioinform. | 4 |
| 2007 | Synonymous Substitution Rates Predict HIV Disease Progression as a Result of Underlying Replication DynamicsabstractUpon HIV transmission, some patients develop AIDS in only a few months, while others remain disease free for 20 or more years. This variation in the rate of disease progression is poorly understood and has been attributed to host genetics, host immune responses, co-infection, viral genetics, and adaptation. Here, we develop a new "relaxed-clock" phylogenetic method to estimate absolute rates of synonymous and nonsynonymous substitution through time. We identify an unexpected association between the synonymous substitution rate of HIV and disease progression parameters. Since immune activation is the major determinant of HIV disease progression, we propose that this process can also determine viral generation times, by creating favourable conditions for HIV replication. These conclusions may apply more generally to HIV evolution, since we also observed an overall low synonymous substitution rate for HIV-2, which is known to be less pathogenic than HIV-1 and capable of tempering the detrimental effects of immune activation. Humoral immune responses, on the other hand, are the major determinant of nonsynonymous rate changes through time in the envelope gene, and our relaxed-clock estimates support a decrease in selective pressure as a consequence of immune system collapse. Philippe Lemey, Sergei L. Kosakovsky Pond, Alexei J. Drummond, Oliver G. Pybus, Beth Shapiro, Helena Barroso, Nuno Taveira, Andrew Rambaut |
PLoS Comput. Biol. | 1 |
| 2005 | SlidingBayes: exploring recombination using a sliding window approach based on Bayesian phylogenetic inferenceabstractAbstract Summary: We developed a software tool (SlidingBayes) for recombination analysis based on Bayesian phylogenetic inference. Sliding-Bayes provides a powerful approach for detecting potential recombination, especially between highly divergent sequences and complex HIV-1 recombinants for which simpler methods like neighbor joining (NJ) may be less powerful. SlidingBayes guides Markov Chain Monte Carlo (MCMC) sampling performed by MrBayes in a sliding window across the alignment (Bayesian scanning). The tool can be used for nucleotide and amino acid sequences and combines all the modeling possibilities of MrBayes with the ability to plot the posterior probability support for clustering of various combinations of taxa. Availability: SlidingBayes is available at http://www.kuleuven.ac.be/rega/cev/Software/ Contact: [email protected] Supplementary information: A quick guide and examples for SlidingBayes are available at http://www.kuleuven.ac.be/rega/cev/Software/ Dimitris Paraskevis, Koen Deforche, Philippe Lemey, Gkikas Magiorkinis, Angelos Hatzakis, Anne-Mieke Vandamme |
Bioinform. | 3 |