VLDB 2026 Research / reviewers in the wild / expert
Kevin J. Liu
dblp:155/3741
· DBLP profile ↗
8ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Phylogenetic Placement of Aligned Genomes and Metagenomes With Non-Tree-Like Evolutionary HistoriesabstractPhylogenetic placement is the computational task that places a query taxon into a reference phylogeny using computational analysis of biomolecular sequence data or other evolutionary characters. A chief advantage of phylogenetic placement over de novo phylogenetic reconstruction is greatly reduced computational requirements, and the former has been applied in many different topics in phylogenetics. One of the more recent applications has been enabled by rapid advances in biomolecular sequencing technology: classification of genomes, metagenomes, and metagenome-assembled genomes (MAGs) in large-scale datasets produced by next-generation sequencing. A number of methods have been developed for this purpose, and all share the common simplifying assumption that a phylogenetic tree suffices for modeling the evolutionary history of all genomes and/or metagenomes under study. Another parallel development in today's post-genomic era is a greater understanding of the prevalence and importance of non-tree-like evolution in the Tree of Life – the evolutionary history of all life on Earth – which in fact may not be a tree at all. More general graph representations such as phylogenetic networks have therefore been proposed, and a new generation of phylogenetic network reconstruction methods are under active development. But the simplifying assumption made by phylogenetic tree placement methods is fundamentally at odds with the non-tree-like evolutionary histories of many microbes and other organisms. The consequences of this conflict are poorly understood. In this study, we conduct a comprehensive performance study to directly assess the impact of non-tree-like evolution on phylogenetic tree placement of genomes and metagenomes. Our study includes in silico simulation experiments as well as empirical data analyses. We find that the topological accuracy of phylogenetic tree placement degrades quickly as genomic sequence evolution becomes increasingly non-tree-like. We then introduce a new statistical method for phylogenetic network placement of genomes and metagenomes, which we refer to as NetPlacer version 0. Initial experiments with NetPlacer provide a proof-of-concept, but also point to the need for greater computational scalability. We conclude with thoughts on algorithmic techniques to enable fast and accurate phylogenetic network placement. Md Alamin, Kevin J. Liu |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | The Impact of Species Tree Estimation Error on Cophylogenetic ReconstructionabstractJust as a phylogeny encodes the evolutionary relationships among a group of organisms, a cophylogeny represents the coevolutionary relationships among symbiotic partners. Both are primarily reconstructed using computational analysis of biomolecular sequence data. The most widely used cophylogenetic reconstruction methods utilize an important simplifying assumption: species phylogenies for each set of coevolved taxa are required as input and assumed to be correct. Many studies have shown that this assumption is rarely - if ever - satisfied, and the consequences for cophylogenetic studies are poorly understood. To address this gap, we conduct a comprehensive performance study that quantifies the relationship between species tree estimation error and downstream cophylogenetic estimation accuracy. We study the performance of state-of-the-art methods for cophylogenetic reconstruction using in silico model-based simulations. Our investigation also assessed cophylogenetic reproducibility using genomic sequence data from two important models of symbiosis: soil-associated fungi and their endosymbiotic bacteria, and bobtail squid and their bioluminescent bacterial symbionts. Our findings conclusively demonstrate the major impact that upstream phylogenetic estimation error has on downstream cophylogenetic reconstruction. Relative to other experimental factors such as cophylogenetic estimation method choice and coevolutionary event costs, phylogenetic estimation error ranked highest in importance based on a random forest-based variable importance assessment. We conclude with practical guidance and future research directions. Among the many considerations needed for accurate cophylogenetic reconstruction - choice of computational method, method settings, sampling design, and others - just as much attention must be paid to careful species phylogeny estimation using modern best practices. Julia Zheng, Yuya Nishida, Alicja Okrasinska, Gregory M. Bonito, Elizabeth A. C. Heath-Heckman, Kevin J. Liu |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | A Statistical Optimization Technique to Inform Statistical Resampling Assessments of Phylogenetic Reconstruction Reliability
Byungho Lee, Kevin J. Liu |
BIBM | 2 |
| 2024 | Estimating confidence intervals on reconstructed cophylogenies using bootstrap resamplingabstractCophylogenies represent coevolutionary histories of two or more sets of coevolved taxa, and are used to study coevolution and other fundamental evolutionary processes. As with traditional phylogenies, cophylogenies are primarily reconstructed using computational analysis of DNA and other biomolecular sequence data. An essential question concerns the reliability of reconstructed phylogenies and cophylogenies.Statistical resampling offers a principled approach to evaluate statistical confidence for these tasks. We therefore apply bootstrap resampling – one of the most widely used non-parametric resampling techniques – to place confidence intervals on a reconstructed cophylogeny, which is the first such method to our knowledge. We validate the performance of the resulting reliability estimates in a simulation study as well as an empirical case study of bobtail squid and its bioluminescent endosymbionts. The utility of statistical resampling to assess cophylogenetic reconstruction reliability in an automated and data-driven manner points the way forward – both for future methods development and wider adoption in studies of symbioses and other forms of coevolution. Julia Zheng, Yuya Nishida, Elizabeth A. C. Heath-Heckman, Kevin J. Liu |
BIBM | 4 |
| 2021 | Build a better bootstrap and the RAWR shall beat a random path to your door: phylogenetic support estimation revisitedabstractMOTIVATION: The standard bootstrap method is used throughout science and engineering to perform general-purpose non-parametric resampling and re-estimation. Among the most widely cited and widely used such applications is the phylogenetic bootstrap method, which Felsenstein proposed in 1985 as a means to place statistical confidence intervals on an estimated phylogeny (or estimate 'phylogenetic support'). A key simplifying assumption of the bootstrap method is that input data are independent and identically distributed (i.i.d.). However, the i.i.d. assumption is an over-simplification for biomolecular sequence analysis, as Felsenstein noted. RESULTS: In this study, we introduce a new sequence-aware non-parametric resampling technique, which we refer to as RAWR ('RAndom Walk Resampling'). RAWR consists of random walks that synthesize and extend the standard bootstrap method and the 'mirrored inputs' idea of Landan and Graur. We apply RAWR to the task of phylogenetic support estimation. RAWR's performance is compared to the state-of-the-art using synthetic and empirical data that span a range of dataset sizes and evolutionary divergence. We show that RAWR support estimates offer comparable or typically superior type I and type II error compared to phylogenetic bootstrap support. We also conduct a re-analysis of large-scale genomic sequence data from a recent study of Darwin's finches. Our findings clarify phylogenetic uncertainty in a charismatic clade that serves as an important model for complex adaptive evolution. AVAILABILITY AND IMPLEMENTATION: Data and software are publicly available under open-source software and open data licenses at: https://gitlab.msu.edu/liulab/RAWR-study-datasets-and-scripts. Wei Wang 0337, Ahmad Hejasebazzi, Julia Zheng, Kevin J. Liu |
Bioinform. | 4 |
| 2019 | An Application of Random Walk Resampling to Phylogenetic HMM Inference and LearningabstractStatistical resampling methods are widely used for confidence interval placement and as a data perturbation technique for statistical inference and learning. An important assumption of popular resampling methods such as the standard bootstrap is that input observations are identically and independently distributed (i.i.d.). However, within the area of computational biology and bioinformatics, many different factors can contribute to intra-sequence dependence, such as recombination and other evolutionary processes governing sequence evolution. The SEquential RESampling (“SERES”) framework was previously proposed to relax the simplifying assumption of i.i.d. input observations. SERES resampling takes the form of random walks on an input of either aligned or unaligned biomolecular sequences. This study introduces the first application of SERES random walks on aligned sequence inputs and is also the first to demonstrate the utility of SERES as a data perturbation technique to yield improved statistical estimates. We focus on the classical problem of recombination-aware local genealogical inference. We show in a simulation study that coupling SERES resampling and re-estimation with recHMM, a hidden Markov model-based method, produces local genealogical inferences with consistent and often large improvements in terms of topological accuracy. We further evaluate method performance using an empirical HIV genome sequence dataset. Wei Wang 0337, Qiqige Wuyun, Kevin J. Liu |
BIBM | 3 |
| 2016 | A scalability study of phylogenetic network inference methods using empirical datasets and simulations involving a single reticulationabstractBACKGROUND: Branching events in phylogenetic trees reflect bifurcating and/or multifurcating speciation and splitting events. In the presence of gene flow, a phylogeny cannot be described by a tree but is instead a directed acyclic graph known as a phylogenetic network. Both phylogenetic trees and networks are typically reconstructed using computational analysis of multi-locus sequence data. The advent of high-throughput sequencing technologies has brought about two main scalability challenges: (1) dataset size in terms of the number of taxa and (2) the evolutionary divergence of the taxa in a study. The impact of both dimensions of scale on phylogenetic tree inference has been well characterized by recent studies; in contrast, the scalability limits of phylogenetic network inference methods are largely unknown. RESULTS: In this study, we quantify the performance of state-of-the-art phylogenetic network inference methods on large-scale datasets using empirical data sampled from natural mouse populations and a range of simulations using model phylogenies with a single reticulation. We find that, as in the case of phylogenetic tree inference, the performance of leading network inference methods is negatively impacted by both dimensions of dataset scale. In general, we found that topological accuracy degrades as the number of taxa increases; a similar effect was observed with increased sequence mutation rate. The most accurate methods were probabilistic inference methods which maximize either likelihood under coalescent-based models or pseudo-likelihood approximations to the model likelihood. The improved accuracy obtained with probabilistic inference methods comes at a computational cost in terms of runtime and main memory usage, which become prohibitive as dataset size grows past twenty-five taxa. None of the probabilistic methods completed analyses of datasets with 30 taxa or more after many weeks of CPU runtime. CONCLUSIONS: We conclude that the state of the art of phylogenetic network inference lags well behind the scope of current phylogenomic studies. New algorithmic development is critically needed to address this methodological gap. Hussein A. Hejase, Kevin J. Liu |
BMC Bioinform. | 2 |
| 2014 | An HMM-Based Comparative Genomic Framework for Detecting Introgression in EukaryotesabstractOne outcome of interspecific hybridization and subsequent effects of evolutionary forces is introgression, which is the integration of genetic material from one species into the genome of an individual in another species. The evolution of several groups of eukaryotic species has involved hybridization, and cases of adaptation through introgression have been already established. In this work, we report on PhyloNet-HMM-a new comparative genomic framework for detecting introgression in genomes. PhyloNet-HMM combines phylogenetic networks with hidden Markov models (HMMs) to simultaneously capture the (potentially reticulate) evolutionary history of the genomes and dependencies within genomes. A novel aspect of our work is that it also accounts for incomplete lineage sorting and dependence across loci. Application of our model to variation data from chromosome 7 in the mouse (Mus musculus domesticus) genome detected a recently reported adaptive introgression event involving the rodent poison resistance gene Vkorc1, in addition to other newly detected introgressed genomic regions. Based on our analysis, it is estimated that about 9% of all sites within chromosome 7 are of introgressive origin (these cover about 13 Mbp of chromosome 7, and over 300 genes). Further, our model detected no introgression in a negative control data set. We also found that our model accurately detected introgression and other evolutionary processes from synthetic data sets simulated under the coalescent model with recombination, isolation, and migration. Our work provides a powerful framework for systematic analysis of introgression while simultaneously accounting for dependence across sites, point mutations, recombination, and ancestral polymorphism. Kevin J. Liu, Jingxuan Dai, Kathy Truong, Michael H. Kohn, Luay Nakhleh |
PLoS Comput. Biol. | 1 |