David Posada

dblp:94/359 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
1since 2021 · last 2022
0000-0003-1407-3406ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
13 papers
Bioinformatics and computational biology · 97% Computational science and engineering · 3%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 41% High-performance computing · 41% Cloud and datacenter computing · 18%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
phylogenetics
1.482022
Phylovar: toward scalable phylogeny-aware inference of single-nucleotide variations from single-cell DNA sequencing data · Bioinform. 2022
RecPhyloXML: a format for reconciled gene trees · Bioinform. 2018
jmodeltest.org: selection of nucleotide substitution models on the cloud · Bioinform. 2014
Bioinformatics and computational biology
genomics
0.612022
Phylovar: toward scalable phylogeny-aware inference of single-nucleotide variations from single-cell DNA sequencing data · Bioinform. 2022
Bioinformatics and computational biology › genomics › computational genomics
SNP detection
0.612022
Phylovar: toward scalable phylogeny-aware inference of single-nucleotide variations from single-cell DNA sequencing data · Bioinform. 2022
Bioinformatics and computational biology › phylogenetics
gene tree reconciliation
0.312018
RecPhyloXML: a format for reconciled gene trees · Bioinform. 2018
Bioinformatics and computational biology › next-generation sequencing
next-generation sequencing simulation
0.312018
NGSphy: phylogenomic simulation of next-generation sequencing data · Bioinform. 2018
Bioinformatics and computational biology › phylogenetics
phylogenomics
0.312018
NGSphy: phylogenomic simulation of next-generation sequencing data · Bioinform. 2018
Bioinformatics and computational biology
molecular evolution
0.232013
Protein evolution along phylogenetic histories under structurally constrained substitution models · Bioinform. 2013
GARD: a genetic algorithm for recombination detection · Bioinform. 2006
MODELTEST: testing the model of DNA substitution · Bioinform. 1998
Bioinformatics and computational biology › molecular evolution
recombination detection
0.232010
RDP3: a flexible and fast computer program for analyzing recombination · Bioinform. 2010
GARD: a genetic algorithm for recombination detection · Bioinform. 2006
RDP2: recombination detection and analysis from sequence alignments · Bioinform. 2005
Computational science and engineering
model selection
0.222011
ProtTest 3: fast selection of best-fit models of protein evolution · Bioinform. 2011
ProtTest: selection of best-fit models of protein evolution · Bioinform. 2005
Bioinformatics and computational biology
cancer genomics
0.212022
Phylovar: toward scalable phylogeny-aware inference of single-nucleotide variations from single-cell DNA sequencing data · Bioinform. 2022
Bioinformatics and computational biology › molecular evolution
recombination analysis
0.222010
RDP3: a flexible and fast computer program for analyzing recombination · Bioinform. 2010
RDP2: recombination detection and analysis from sequence alignments · Bioinform. 2005
High-performance computing › cluster computing
multi-core cluster computing
0.112011
ProtTest 3: fast selection of best-fit models of protein evolution · Bioinform. 2011
Parallel and multicore computing
parallel computing
0.112011
ProtTest 3: fast selection of best-fit models of protein evolution · Bioinform. 2011
Bioinformatics and computational biology › phylogenetics › evolutionary history reconstruction
ancestral sequence reconstruction
0.112010
RDP3: a flexible and fast computer program for analyzing recombination · Bioinform. 2010
Bioinformatics and computational biology › population genetics › recombination
recombination hotspot detection
0.112010
RDP3: a flexible and fast computer program for analyzing recombination · Bioinform. 2010
Bioinformatics and computational biology › molecular evolution
recombination breakpoint detection
0.122006
RDP2: recombination detection and analysis from sequence alignments · Bioinform. 2005
GARD: a genetic algorithm for recombination detection · Bioinform. 2006
Bioinformatics and computational biology › statistical genetics
genotype-phenotype association
0.122005
TreeScan: a bioinformatic application to search for genotype/phenotype associations using haplotype trees · Bioinform. 2005
Simulating haplotype blocks in the human genome · Bioinform. 2003
Bioinformatics and computational biology
population genetics
0.122005
Simulating haplotype blocks in the human genome · Bioinform. 2003
TreeScan: a bioinformatic application to search for genotype/phenotype associations using haplotype trees · Bioinform. 2005
Cloud and datacenter computing › datacenter services › online service systems › internet services
web server
0.112014
jmodeltest.org: selection of nucleotide substitution models on the cloud · Bioinform. 2014
Bioinformatics and computational biology › statistical genetics › haplotype analysis
haplotype-based association testing
0.112005
TreeScan: a bioinformatic application to search for genotype/phenotype associations using haplotype trees · Bioinform. 2005
Bioinformatics and computational biology › population genetics › population genetics simulation
haplotype simulation
0.012003
Simulating haplotype blocks in the human genome · Bioinform. 2003
Bioinformatics and computational biology › statistical genetics
association analysis
0.012003
Simulating haplotype blocks in the human genome · Bioinform. 2003

Methods — techniques the papers use, named apart from their topics

single-cell sequencing · 0.6phylogeny-guided variant calling · 0.6statistical distribution sampling · 0.3maximum likelihood · 0.2protein folding stability model · 0.2coalescent simulation · 0.2phylogenetic tree construction · 0.2matrix visualization · 0.1likelihood-based model selection · 0.1genetic algorithm · 0.1
YearPublicationVenuePosition
2022 Phylovar: toward scalable phylogeny-aware inference of single-nucleotide variations from single-cell DNA sequencing data
abstract
MOTIVATION: Single-nucleotide variants (SNVs) are the most common variations in the human genome. Recently developed methods for SNV detection from single-cell DNA sequencing data, such as SCIΦ and scVILP, leverage the evolutionary history of the cells to overcome the technical errors associated with single-cell sequencing protocols. Despite being accurate, these methods are not scalable to the extensive genomic breadth of single-cell whole-genome (scWGS) and whole-exome sequencing (scWES) data. RESULTS: Here, we report on a new scalable method, Phylovar, which extends the phylogeny-guided variant calling approach to sequencing datasets containing millions of loci. Through benchmarking on simulated datasets under different settings, we show that, Phylovar outperforms SCIΦ in terms of running time while being more accurate than Monovar (which is not phylogeny-aware) in terms of SNV detection. Furthermore, we applied Phylovar to two real biological datasets: an scWES triple-negative breast cancer data consisting of 32 cells and 3375 loci as well as an scWGS data of neuron cells from a normal human brain containing 16 cells and approximately 2.5 million loci. For the cancer data, Phylovar detected somatic SNVs with high or moderate functional impact that were also supported by bulk sequencing dataset and for the neuron dataset, Phylovar identified 5745 SNVs with non-synonymous effects some of which were associated with neurodegenerative diseases. AVAILABILITY AND IMPLEMENTATION: Phylovar is implemented in Python and is publicly available at https://github.com/NakhlehLab/Phylovar.
Mohammad Amin Edrisi, Monica V. Valecha, Sunkara B. V. Chowdary, Sergio Robledo, Huw A. Ogilvie, David Posada, Hamim Zafar, Luay Nakhleh
Bioinform.6
2018 RecPhyloXML: a format for reconciled gene trees
abstract
Motivation: A reconciliation is an annotation of the nodes of a gene tree with evolutionary events-for example, speciation, gene duplication, transfer, loss, etc.-along with a mapping onto a species tree. Many algorithms and software produce or use reconciliations but often using different reconciliation formats, regarding the type of events considered or whether the species tree is dated or not. This complicates the comparison and communication between different programs. Results: Here, we gather a consortium of software developers in gene tree species tree reconciliation to propose and endorse a format that aims to promote an integrative-albeit flexible-specification of phylogenetic reconciliations. This format, named recPhyloXML, is accompanied by several tools such as a reconciled tree visualizer and conversion utilities. Availability and implementation: http://phylariane.univ-lyon1.fr/recphyloxml/.
Wandrille Duchemin, Guillaume Gence, Anne-Muriel Arigon Chifolleau, Lars Arvestad, Mukul S. Bansal, Vincent Berry, Bastien Boussau, François Chevenet, Nicolas Comte, Adrián A. Davín, Christophe Dessimoz, David Dylus, Damir Hasic, Diego Mallo, Rémi Planel, David Posada, Céline Scornavacca, Gergely J. Szöllosi, Louxin Zhang, Eric Tannier, Vincent Daubin
Bioinform.16
2018 NGSphy: phylogenomic simulation of next-generation sequencing data
abstract
Motivation: Advances in sequencing technologies have made it feasible to obtain massive datasets for phylogenomic inference, often consisting of large numbers of loci from multiple species and individuals. The phylogenomic analysis of next-generation sequencing (NGS) data requires a complex computational pipeline where multiple technical and methodological decisions are necessary that can influence the final tree obtained, like those related to coverage, assembly, mapping, variant calling and/or phasing. Results: To assess the influence of these variables we introduce NGSphy, an open-source tool for the simulation of Illumina reads/read counts obtained from haploid/diploid individual genomes with thousands of independent gene families evolving under a common species tree. In order to resemble real NGS experiments, NGSphy includes multiple options to model sequencing coverage (depth) heterogeneity across species, individuals and loci, including off-target or uncaptured loci. For comprehensive simulations covering multiple evolutionary scenarios, parameter values for the different replicates can be sampled from user-defined statistical distributions. Availability and implementation: Source code, full documentation and tutorials including a 'Getting started' guide are available at http://github.com/merlyescalona/ngsphy. Supplementary information: Supplementary data are available at Bioinformatics online.
Merly Escalona, Sara Rocha, David Posada
Bioinform.3
2014 jmodeltest.org: selection of nucleotide substitution models on the cloud
abstract
Abstract Summary: The selection of models of nucleotide substitution is one of the major steps of modern phylogenetic analysis. Different tools exist to accomplish this task, among which jModelTest 2 (jMT2) is one of the most popular. Still, to deal with large DNA alignments with hundreds or thousands of loci, users of jMT2 need to have access to High Performance Computing clusters, including installation and configuration capabilities, conditions not always met. Here we present jmodeltest.org, a novel web server for the transparent execution of jMT2 across different platforms and for a wide range of users. Its main benefit is straightforward execution, avoiding any configuration/execution issues, and reducing significantly in most cases the time required to complete the analysis. Availability and implementation: jmodeltest.org is accessible using modern browsers, such as Firefox, Chrome, Opera, Safari and IE from http://jmodeltest.org. User registration is not mandatory, but users wanting to have additional functionalities, like access to previous analyses, have the possibility of opening a user account. Contact: [email protected]
Jose Manuel Santorum, Diego Darriba, Guillermo L. Taboada, David Posada
Bioinform.4
2013 Protein evolution along phylogenetic histories under structurally constrained substitution models
abstract
MOTIVATION: Models of molecular evolution aim at describing the evolutionary processes at the molecular level. However, current models rarely incorporate information from protein structure. Conversely, structure-based models of protein evolution have not been commonly applied to simulate sequence evolution in a phylogenetic framework, and they often ignore relevant evolutionary processes such as recombination. A simulation evolutionary framework that integrates substitution models that account for protein structure stability should be able to generate more realistic in silico evolved proteins for a variety of purposes. RESULTS: We developed a method to simulate protein evolution that combines models of protein folding stability, such that the fitness depends on the stability of the native state both with respect to unfolding and misfolding, with phylogenetic histories that can be either specified by the user or simulated with the coalescent under complex evolutionary scenarios, including recombination, demographics and migration. We have implemented this framework in a computer program called ProteinEvolver. Remarkably, comparing these models with empirical amino acid replacement models, we found that the former produce amino acid distributions closer to distributions observed in real protein families, and proteins that are predicted to be more stable. Therefore, we conclude that evolutionary models that consider protein stability and realistic evolutionary histories constitute a better approximation of the real evolutionary process.
Miguel Arenas, Helena G. Dos Santos, David Posada, Ugo Bastolla
Bioinform.3
2011 ProtTest 3: fast selection of best-fit models of protein evolution
abstract
UNLABELLED: We have implemented a high-performance computing (HPC) version of ProtTest that can be executed in parallel in multicore desktops and clusters. This version, called ProtTest 3, includes new features and extended capabilities. AVAILABILITY: ProtTest 3 source code and binaries are freely available under GNU license for download from http://darwin.uvigo.es/software/prottest3, linked to a Mercurial repository at Bitbucket (https://bitbucket.org/). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Diego Darriba, Guillermo L. Taboada, Ramón Doallo, David Posada
Bioinform.4
2010 RDP3: a flexible and fast computer program for analyzing recombination
abstract
Abstract Summary: RDP3 is a new version of the RDP program for characterizing recombination events in DNA-sequence alignments. Among other novelties, this version includes four new recombination analysis methods (3SEQ, VISRD, PHYLRO and LDHAT), new tests for recombination hot-spots, a range of matrix methods for visualizing over-all patterns of recombination within datasets and recombination-aware ancestral sequence reconstruction. Complementary to a high degree of analysis flow automation, RDP3 also has a highly interactive and detailed graphical user interface that enables more focused hands-on cross-checking of results with a wide variety of newly implemented phylogenetic tree construction and matrix-based recombination signal visualization methods. The new RDP3 can accommodate large datasets and is capable of analyzing alignments ranging in size from 1000×10 kilobase sequences to 20×2 megabase sequences within 48 h on a desktop PC. Availability: RDP3 is available for free from its web site http://darwin.uvigo.es/rdp/rdp.html Contact: [email protected] Supplementary information: The RDP3 program manual contains detailed descriptions of the various methods it implements and a step-by-step guide describing how best to use these.
Darren P. Martin, Philippe Lemey, Martin Lott, Vincent Moulton, David Posada, Pierre Lefeuvre
Bioinform.5
2010 Characterization of phylogenetic networks with NetTest
abstract
BACKGROUND: Typical evolutionary events like recombination, hybridization or gene transfer make necessary the use of phylogenetic networks to properly depict the evolution of DNA and protein sequences. Although several theoretical classes have been proposed to characterize these networks, they make stringent assumptions that will likely not be met by the evolutionary process. We have recently shown that the complexity of simulated networks is a function of the population recombination rate, and that at moderate and large recombination rates the resulting networks cannot be categorized. However, we do not know whether these results extend to networks estimated from real data. RESULTS: We introduce a web server for the categorization of explicit phylogenetic networks, including the most relevant theoretical classes developed so far. Using this tool, we analyzed statistical parsimony phylogenetic networks estimated from approximately 5,000 DNA alignments, obtained from the NCBI PopSet and Polymorphix databases. The level of characterization was correlated to nucleotide diversity, and a high proportion of the networks derived from these data sets could be formally characterized. CONCLUSIONS: We have developed a public web server, NetTest (freely available from the software section at http://darwin.uvigo.es), to formally characterize the complexity of phylogenetic networks. Using NetTest we found that most statistical parsimony networks estimated with the program TCS could be assigned to a known network class. The level of network characterization was correlated to nucleotide diversity and dependent upon the intra/interspecific levels, although no significant differences were detected among genes. More research on the properties of phylogenetic networks is clearly needed.
Miguel Arenas, Mateus Patricio, David Posada, Gabriel Valiente
BMC Bioinform.3
2009 An Evolutionary Model-Based Algorithm for Accurate Phylogenetic Breakpoint Mapping and Subtype Prediction in HIV-1
abstract
Genetically diverse pathogens (such as Human Immunodeficiency virus type 1, HIV-1) are frequently stratified into phylogenetically or immunologically defined subtypes for classification purposes. Computational identification of such subtypes is helpful in surveillance, epidemiological analysis and detection of novel variants, e.g., circulating recombinant forms in HIV-1. A number of conceptually and technically different techniques have been proposed for determining the subtype of a query sequence, but there is not a universally optimal approach. We present a model-based phylogenetic method for automatically subtyping an HIV-1 (or other viral or bacterial) sequence, mapping the location of breakpoints and assigning parental sequences in recombinant strains as well as computing confidence levels for the inferred quantities. Our Subtype Classification Using Evolutionary ALgorithms (SCUEAL) procedure is shown to perform very well in a variety of simulation scenarios, runs in parallel when multiple sequences are being screened, and matches or exceeds the performance of existing approaches on typical empirical cases. We applied SCUEAL to all available polymerase (pol) sequences from two large databases, the Stanford Drug Resistance database and the UK HIV Drug Resistance Database. Comparing with subtypes which had previously been assigned revealed that a minor but substantial (approximately 5%) fraction of pure subtype sequences may in fact be within- or inter-subtype recombinants. A free implementation of SCUEAL is provided as a module for the HyPhy package and the Datamonkey web server. Our method is especially useful when an accurate automatic classification of an unknown strain is desired, and is positioned to complement and extend faster but less accurate methods. Given the increasingly frequent use of HIV subtype information in studies focusing on the effect of subtype on treatment, clinical outcome, pathogenicity and vaccine design, the importance of accurate, robust and extensible subtyping procedures is clear.
Sergei L. Kosakovsky Pond, David Posada, Eric W. Stawiski, Colombe Chappey, Art F. Y. Poon, Gareth Hughes, Esther Fearnhill, Michael B. Gravenor, Andrew J. Leigh Brown, Simon D. W. Frost
PLoS Comput. Biol.2
2007 Recodon: Coalescent simulation of coding DNA sequences with recombination, migration and demography
abstract
BACKGROUND: Coalescent simulations have proven very useful in many population genetics studies. In order to arrive to meaningful conclusions, it is important that these simulations resemble the process of molecular evolution as much as possible. To date, no single coalescent program is able to simulate codon sequences sampled from populations with recombination, migration and growth. RESULTS: We introduce a new coalescent program, called Recodon, which is able to simulate samples of coding DNA sequences under complex scenarios in which several evolutionary forces can interact simultaneously (namely, recombination, migration and demography). The basic codon model implemented is an extension to the general time-reversible model of nucleotide substitution with a proportion of invariable sites and among-site rate variation. In addition, the program implements non-reversible processes and mixtures of different codon models. CONCLUSION: Recodon is a flexible tool for the simulation of coding DNA sequences under realistic evolutionary models. These simulations can be used to build parameter distributions for testing evolutionary hypotheses using experimental data. Recodon is written in C, can run in parallel, and is freely available from http://darwin.uvigo.es/.
Miguel Arenas, David Posada
BMC Bioinform.2
2006 GARD: a genetic algorithm for recombination detection
abstract
MOTIVATION: Phylogenetic and evolutionary inference can be severely misled if recombination is not accounted for, hence screening for it should be an essential component of nearly every comparative study. The evolution of recombinant sequences can not be properly explained by a single phylogenetic tree, but several phylogenies may be used to correctly model the evolution of non-recombinant fragments. RESULTS: We developed a likelihood-based model selection procedure that uses a genetic algorithm to search multiple sequence alignments for evidence of recombination breakpoints and identify putative recombinant sequences. GARD is an extensible and intuitive method that can be run efficiently in parallel. Extensive simulation studies show that the method nearly always outperforms other available tools, both in terms of power and accuracy and that the use of GARD to screen sequences for recombination ensures good statistical properties for methods aimed at detecting positive selection. AVAILABILITY: Freely available http://www.datamonkey.org/GARD/
Sergei L. Kosakovsky Pond, David Posada, Michael B. Gravenor, Christopher H. Woelk, Simon D. W. Frost
Bioinform.2
2005 ProtTest: selection of best-fit models of protein evolution
abstract
SUMMARY: Using an appropriate model of amino acid replacement is very important for the study of protein evolution and phylogenetic inference. We have built a tool for the selection of the best-fit model of evolution, among a set of candidate models, for a given protein sequence alignment. AVAILABILITY: ProtTest is available under the GNU license from http://darwin.uvigo.es
Federico Abascal, Rafael Zardoya, David Posada
Bioinform.3
2005 RDP2: recombination detection and analysis from sequence alignments
abstract
UNLABELLED: RDP2 is a Windows 95/XP program that examines nucleotide sequence alignments and attempts to identify recombinant sequences and recombination breakpoints using 10 published recombination detection methods, including GENECONV, BOOTSCAN, MAXIMUM chi(2), CHIMAERA and SISTER SCANNING. The program enables fast automated analysis of large alignments (up to 300 sequences containing 13 000 sites), and interactive exploration, management and verification of results with different recombination detection and tree drawing methods. AVAILABILITY: RDP2 is available free from the RDP2 website (http://darwin.uvigo.es/rdp/rdp.html) CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Detailed descriptions of RDP2 and the methods it implements are included in the program manual, which can be downloaded from the RDP2 website.
Darren P. Martin, Carolyn Williamson, David Posada
Bioinform.3
2005 TreeScan: a bioinformatic application to search for genotype/phenotype associations using haplotype trees
abstract
SUMMARY: We present the software implementation of the tree scanning method to detect associations between genetic haplotypes and quantitative traits, utilizing the evolutionary history of the haplotypes, in samples of unrelated individuals. AVAILABILITY: The program is available free of charge, under the GNU General Public License. A package including C source code, a Makefile, and Windows (DOS) and Macintosh binaries, can be downloaded from http://darwin.uvigo.es
David Posada, Taylor J. Maxwell, Alan R. Templeton
Bioinform.1
2003 Simulating haplotype blocks in the human genome
abstract
SUMMARY: A bioinformatic tool was written to simulate haplotypes and SNPs under a modified coalescent with recombination. The most important feature of this program is that it allows for the specification of non-homogeneous recombination rates, which results in the formation of the so-called 'haplotype blocks' of the human genome. The program also implements different mutation models and flexible demographic histories. The samples generated can be very useful to better understand the architecture of the human genome and to investigate its impact in association studies searching for disease genes. AVAILABILITY: The SNPsim package is available at http://www.evolgenics.com/software
David Posada, Carsten Wiuf
Bioinform.1
1998 MODELTEST: testing the model of DNA substitution
abstract
SUMMARY: The program MODELTEST uses log likelihood scores to establish the model of DNA evolution that best fits the data. AVAILABILITY: The MODELTEST package, including the source code and some documentation is available at http://bioag.byu. edu/zoology/crandall_lab/modeltest.html.
David Posada, Keith A. Crandall
Bioinform.1