EDBT 2026 Demo / reviewers in the wild / expert
Hongzhe Li
dblp:47/4887
· DBLP profile ↗
29ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | dGAMLSS: an exact, distributed algorithm to fit Generalized Additive Models for Location, Scale, and Shape for privacy-preserving population reference chartsabstractMOTIVATION: There is growing interest in estimating population reference ranges across age and sex to better identify atypical clinically-relevant measurements throughout the lifespan. For this task, the World Health Organization recommends using Generalized Additive Models for Location, Scale, and Shape (GAMLSS), which can model non-linear growth trajectories under complex distributions that address the heterogeneity in human populations.Fitting GAMLSS models requires large, generalizable sample sizes, especially for accurate estimation of extreme quantiles, but obtaining such multi-site data can be challenging due to privacy concerns and practical considerations. In settings where patient data cannot be shared, privacy-preserving distributed algorithms for federated learning can be used, but no such algorithm exists for GAMLSS. RESULTS: We propose distributed GAMLSS (dGAMLSS), a distributed algorithm that can fit GAMLSS models across multiple sites without sharing patient-level data. This includes specific considerations for the fitting of smooth functions at varying levels of communication efficiency. We demonstrate the effectiveness of dGAMLSS in constructing population reference charts across clinical, genomics, and neuroimaging settings and show that dGAMLSS is able to reproduce pooled reference charts and inference down to numerical differences. AVAILABILITY AND IMPLEMENTATION: An R package providing examples of the dGAMLSS algorithm, as well as functions for sharing and aggregating site-specific parameters, is available at https://github.com/hufengling/dGAMLSS. Fengling Hu, Jiayi Tong, Margaret Gardner, Lifespan Brain Chart Consortium, Andrew A. Chen, Richard A. I. Bethlehem, Jakob Seidlitz, Hongzhe Li, Aaron Alexander-Bloch, Yong Chen 0016, Russell T. Shinohara |
Bioinform. | 8 |
| 2025 | Wasserstein F-tests for Frechet regression on Bures-Wasserstein manifoldsabstractThis paper addresses regression analysis for covariance matrix-valued outcomes with Euclidean covariates, motivated by applications in single-cell genomics and neuroscience where covariance matrices are observed across many samples. Our analysis leverages Fréchet regression on the Bures-Wasserstein manifold to estimate the conditional Fréchet mean given covariates $x$. We establish a non-asymptotic uniform $\sqrt{n}$-rate of convergence (up to logarithmic factors) over covariates with $\|x\| \lesssim \sqrt{\log n}$ and derive a pointwise central limit theorem to enable statistical inference. For testing covariate effects, we devise a novel test whose null distribution converges to a weighted sum of independent chi-square distributions, with power guarantees against a sequence of contiguous alternatives. Simulations validate the accuracy of the asymptotic theory. Finally, we apply our methods to a single-cell gene expression dataset, revealing age-related changes in gene co-expression networks. Haoshu Xu, Hongzhe Li |
J. Mach. Learn. Res. | 2 |
| 2023 | Inference of microbial covariation networks using copula models with mixture marginsabstractMOTIVATION: Quantification of microbial covariations from 16S rRNA and metagenomic sequencing data is difficult due to their sparse nature. In this article, we propose using copula models with mixed zero-beta margins for the estimation of taxon-taxon covariations using data of normalized microbial relative abundances. Copulas allow for separate modeling of the dependence structure from the margins, marginal covariate adjustment, and uncertainty measurement. RESULTS: Our method shows that a two-stage maximum-likelihood approach provides accurate estimation of model parameters. A corresponding two-stage likelihood ratio test for the dependence parameter is derived and is used for constructing covariation networks. Simulation studies show that the test is valid, robust, and more powerful than tests based upon Pearson's and rank correlations. Furthermore, we demonstrate that our method can be used to build biologically meaningful microbial networks based on a dataset from the American Gut Project. AVAILABILITY AND IMPLEMENTATION: R package for implementation is available at https://github.com/rebeccadeek/CoMiCoN. Rebecca Deek, Hongzhe Li |
Bioinform. | 2 |
| 2023 | The Study of 2-D Magnetic Focusing Inversion Based on the Adjustable Exponential Minimum Support Stabilizing FunctionalabstractWhen the target magnetic structures have clear interfaces and significant petrophysical contrasts to the surrounding strata, the focusing (or sharp-boundary) inversion methods are preferred over the smooth inversion methods. The result of focusing inversions depends on the types of the focusing model constraint (stabilizer). In this study, we propose an adjustable exponential minimum support (AEMS) stabilizer, which is an extension of the previously introduced exponential minimum support (EMS) stabilizer. AEMS is similar to EMS, but it includes an adjustable focusing scheme to control the sharpness of the resulting structure. Using synthetic models, we compared the inversion results obtained with the AEMS method to those obtained by smoothing stabilizers and other focusing stabilizers. AEMS inversion results have a more accurate physical-property distribution and are closer to the true models, indicating an improved model recovery from the other inversion methods. The improvement is mainly attributed to the adjustable scheme of the focusing parameter in the AEMS method. Then, we inverted a field dataset using the AEMS inversion method to obtain a new reference model for copper-gold porphyry deposit detection and interpretation. Linjun Huang, Han Song, Zecheng Wang, Hongzhe Li, Delong Ma |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | A New High-Resolution Remote Sensing Monitoring Method for Nutrients in Coastal WatersabstractMariculture is an important offshore economic activity, and excessive farming can lead to the deterioration of sea ecology. The concentration of nutrients (mainly DIN (dissolved inorganic nitrogen) and PO4 (orthophosphate-phosphorous)) is the main factor characterizing the health condition of farmed seas. Conventional field monitoring methods are spatiotemporally limited, and remote sensing technology has the advantages of high spatial coverage and long time series monitoring. Thus, the Sentinel-3 reflectance data and the in situ measured data for the offshore waters of Wenzhou were matched simultaneously. Then, the matched dataset between the Sentinel-2 band and the in situ measured data were obtained through spectral correspondence conversion between Sentinel-2 and Sentinel-3, and a machine learning algorithm was used to build the inversion model with an independent validation process. The correlations between the concentration of nutrients, area of rafts and precipitation were assessed, and a strong positive correlation was found between the concentration of nutrients and the area of rafts, and a weak negative correlation was found between the former and precipitation. Difeng Wang, Shuping Pan, Hongzhe Li, Fang Gong, Haoyan Hu, Xianqiang He, Zhuoqi Zheng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | A compositional mediation model for a binary outcome: Application to microbiome studiesabstractMOTIVATION: The delicate balance of the microbiome is implicated in our health and is shaped by external factors, such as diet and xenobiotics. Therefore, understanding the role of the microbiome in linking external factors and our health conditions is crucial to translate microbiome research into therapeutic and preventative applications. RESULTS: We introduced a sparse compositional mediation model for binary outcomes to estimate and test the mediation effects of the microbiome utilizing the compositional algebra defined in the simplex space and a linear zero-sum constraint on probit regression coefficients. For this model with the standard causal assumptions, we showed that both the causal direct and indirect effects are identifiable. We further developed a method for sensitivity analysis for the assumption of the no unmeasured confounding effects between the mediator and the outcome. We conducted extensive simulation studies to assess the performance of the proposed method and applied it to real microbiome data to study mediation effects of the microbiome on linking fat intake to overweight/obesity. AVAILABILITY AND IMPLEMENTATION: An R package can be downloaded from https://github.com/mbsohn/cmmb. SUPPLEMENTARY INFORMATION: Supplementary files are available at Bioinformatics online. Michael B. Sohn, Jiarui Lu, Hongzhe Li |
Bioinform. | 3 |
| 2021 | MiRKAT: kernel machine regression-based global association tests for the microbiomeabstractSUMMARY: Distance-based tests of microbiome beta diversity are an integral part of many microbiome analyses. MiRKAT enables distance-based association testing with a wide variety of outcome types, including continuous, binary, censored time-to-event, multivariate, correlated and high-dimensional outcomes. Omnibus tests allow simultaneous consideration of multiple distance and dissimilarity measures, providing higher power across a range of simulation scenarios. Two measures of effect size, a modified R-squared coefficient and a kernel RV coefficient, are incorporated to allow comparison of effect sizes across multiple kernels. AVAILABILITY AND IMPLEMENTATION: MiRKAT is available on CRAN as an R package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nehemiah Wilson, Ni Zhao, Xiang Zhan, Hyunwook Koh, Weijia Fu, Hongzhe Li, Michael C. Wu, Anna M. Plantinga |
Bioinform. | 7 |
| 2021 | Optimal Structured Principal Subspace Estimation: Metric Entropy and Minimax RatesabstractDriven by a wide range of applications, several principal subspace estimation problems have been studied individually under different structural constraints. This paper presents a unified framework for the statistical analysis of a general structured principal subspace estimation problem which includes as special cases sparse PCA/SVD, non-negative PCA/SVD, subspace constrained PCA/SVD, and spectral clustering. General minimax lower and upper bounds are established to characterize the interplay between the information-geometric complexity of the constraint set for the principal subspaces, the signal-to-noise ratio (SNR), and the dimensionality. The results yield interesting phase transition phenomena concerning the rates of convergence as a function of the SNRs and the fundamental limit for consistent estimation. Applying the general results to the specific settings yields the minimax rates of convergence for those problems, including the previous unknown optimal rates for sparse SVD, non-negative PCA/SVD and subspace constrained PCA/SVD. T. Tony Cai, Hongzhe Li |
J. Mach. Learn. Res. | 2 |
| 2020 | Trust-driven Distributed Self-collaborative Security Architecture of IoT Based on Blockchain and Smart ContractsabstractAs Internet of Things (IoT) technology is growing fast, it is clearly forseen that the number of IoT devices and the scale of connections will be further expanded. IoT is capable of taking advantage of the existing network infrastructure effectively, so as to achieve data sharing between devices. However, the large scale and complexity of the network structure will bring potential security risks to IoT system. The traditional access control model is more complex and centralized. This paper aims to build a trust-driven distributed self-collaborative security architecture of IoT based on blockchain and smart contracts, which could overcome the single-point failure problem of centralized entities. Implementation of the security architecture shows that it still possesses characteristics including scalability, lightweight, and fine granularity, while the secure access control preserves the privacy of IoT devices. Hongzhe Li, Sirun Xu, Saifei Li, Guangcheng Sun, Lianshan Yan |
VTC Fall | 1 |
| 2018 | Medication class enrichment analysis: a novel algorithm to analyze multiple pharmacologic exposures simultaneously using electronic health record dataabstractObjective: Observational studies analyzing multiple exposures simultaneously have been limited by difficulty distinguishing relevant results from chance associations due to poor specificity. Set-based methods have been successfully used in genomics to improve signal-to-noise ratio. We present and demonstrate medication class enrichment analysis (MCEA), a signal-to-noise enhancement algorithm for observational data inspired by set-based methods. Materials and Methods: We used The Health Improvement Network database to study medications associated with Clostridium difficile infection (CDI). We performed case-control studies for each medication in The Health Improvement Network to obtain odds ratios (ORs) for association with CDI. We then calculated the association of each pharmacologic class with CDI using logistic regression and MCEA. We also performed simulation studies in which we assessed the sensitivity and specificity of logistic regression compared to MCEA for ORs 0.1-2.0. Results: When analyzing pharmacologic classes using logistic regression, 47 of 110 pharmacologic classes were identified as associated with CDI. When analyzing pharmacologic classes using MCEA, only fluoroquinolones, a class of antibiotics with biologically confirmed causation, and heparin products were associated with CDI. In simulation, MCEA had superior specificity compared to logistic regression across all tested effect sizes and equal or better sensitivity for all effect sizes besides those close to null. Discussion: Although these results demonstrate the promise of MCEA, additional studies that include inpatient administered medications are necessary for validation of the algorithm. Conclusions: In clinical and simulation studies, MCEA demonstrated superior sensitivity and specificity for identifying pharmacologic classes associated with CDI compared to logistic regression. Ravy K. Vajravelu, Frank I. Scott, Ronac Mamtani, Hongzhe Li, Jason H. Moore, James D. Lewis |
J. Am. Medical Informatics Assoc. | 4 |
| 2017 | A general framework for association analysis of microbial communities on a taxonomic treeabstractMotivation: : Association analysis of microbiome composition with disease-related outcomes provides invaluable knowledge towards understanding the roles of microbes in the underlying disease mechanisms. Proper analysis of sparse compositional microbiome data is challenging. Existing methods rely on strong assumptions on the data structure and fail to pinpoint the associated microbial communities. Results: : We develop a general framework to: (i) perform robust association tests for the microbial community that exhibits arbitrary inter-taxa dependencies; (ii) localize lineages on the taxonomic tree that are associated with covariates (e.g. disease status); and (iii) assess the overall association of the whole microbial community with the covariates. Unlike existing methods for microbiome association analysis, our framework does not make any distributional assumptions on the microbiome data; it allows for the adjustment of confounding variables and accommodates excessive zero observations; and it incorporates taxonomic information. We perform extensive simulation studies under a wide-range of scenarios to evaluate the new methods and demonstrate substantial power gain over existing methods. The advantages of the proposed framework are further demonstrated with real datasets from two microbiome studies. The relevant R package miLineage is publicly available. Availability and Implementation: : miLineage package, manual and tutorial are available at https://medschool.vanderbilt.edu/tang-lab/software/miLineage . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Zheng-Zheng Tang, Guanhua Chen 0002, Alexander V. Alekseyenko, Hongzhe Li |
Bioinform. | 4 |
| 2016 | Automated Source Code Instrumentation for Verifying Potential Vulnerabilities
Hongzhe Li, Jaesang Oh, Hakjoo Oh, Heejo Lee |
SEC | 1 |
| 2016 | A two-part mixed-effects model for analyzing longitudinal microbiome compositional dataabstractMOTIVATION: The human microbial communities are associated with many human diseases such as obesity, diabetes and inflammatory bowel disease. High-throughput sequencing technology has been widely used to quantify the microbial composition in order to understand its impacts on human health. Longitudinal measurements of microbial communities are commonly obtained in many microbiome studies. A key question in such microbiome studies is to identify the microbes that are associated with clinical outcomes or environmental factors. However, microbiome compositional data are highly skewed, bounded in [0,1), and often sparse with many zeros. In addition, the observations from repeated measures in longitudinal studies are correlated. A method that takes into account these features is needed for association analysis in longitudinal microbiome data. RESULTS: In this paper, we propose a two-part zero-inflated Beta regression model with random effects (ZIBR) for testing the association between microbial abundance and clinical covariates for longitudinal microbiome data. The model includes a logistic regression component to model presence/absence of a microbe in the samples and a Beta regression component to model non-zero microbial abundance, where each component includes a random effect to account for the correlations among the repeated measurements on the same subject. Both simulation studies and the application to real microbiome data have shown that ZIBR model outperformed the previously used methods. The method provides a useful tool for identifying the relevant taxa based on longitudinal or repeated measures in microbiome research. AVAILABILITY AND IMPLEMENTATION: https://github.com/chvlyl/ZIBR CONTACT: [email protected]. Eric Z. Chen, Hongzhe Li |
Bioinform. | 2 |
| 2016 | CLORIFI: software vulnerability discovery using code clone verificationabstractSummary Software vulnerability has long been considered an important threat to the system safety. A vulnerability is often reproduced because of the frequent code reuse by programmers. Security patches are usually not propagated to all code clones; however, they could be leveraged to discover unknown vulnerabilities. Static code auditing approaches are frequently proposed to scan source codes for security flaws; unfortunately, these approaches generate too many false positives. While dynamic execution analysis methods can precisely report vulnerabilities, they are ineffective in path exploration, which limits them to scale to large programs. With the purpose of detecting vulnerability in a scalable way with more preciseness, in this paper, we propose a novel mechanism, called software vulnerability discovery using Code Clone Verification (CLORIFI), that scalably discovers vulnerabilities in real world programs using code clone verification. In the beginning, we use a fast and scalable syntax‐based way to find code clones in program source codes based on released security patches. Subsequently, code clones are being verified using concolic testing to dramatically decrease the false positives. In addition, we mitigate the path explosion problem by backward sensitive data tracing in concolic execution. Experiments have been conducted with real‐world open‐source projects (recent Linux OS distributions and program packages). As a result, we found 7 real vulnerabilities out of 63 code clones from Ubuntu 14.04 LTS (Canonical, London, UK) and 10 vulnerabilities out of 40 code clones from CentOS 7.0 (The CentOS Project(community contributed)). Furthermore, we confirmed more code clone vulnerabilities in various versions of programs including Rsyslog (Open Source(Original author: Rainer Gerhards)), Apache (Apache Software Foundation, Forest Hill, Maryland, USA) and Firefox (Mozilla Corporation, Mountain View, California, USA). In order to evaluate the effectiveness of vulnerability verification in a systematic way, we also utilized Juliet Test Suite as measurement objects. The results show that CLORIFI achieves 98% accuracy with 0 false positives. Copyright © 2015 John Wiley & Sons, Ltd. Hongzhe Li, Hyuckmin Kwon, Jonghoon Kwon, Heejo Lee |
Concurr. Comput. Pract. Exp. | 1 |
| 2015 | glmgraph: an R package for variable selection and predictive modeling of structured genomic dataabstractUNLABELLED: One central theme of modern high-throughput genomic data analysis is to identify relevant genomic features as well as build up a predictive model based on selected features for various tasks such as personalized medicine. Correlating the large number of 'omics' features with a certain phenotype is particularly challenging due to small sample size (n) and high dimensionality (p). To address this small n, large p problem, various forms of sparse regression models have been proposed by exploiting the sparsity assumption. Among these, network-constrained sparse regression model is of particular interest due to its ability to utilize the prior graph/network structure in the omics data. Despite its potential usefulness for omics data analysis, no efficient R implementation is publicly available. Here we present an R software package 'glmgraph' that implements the graph-constrained regularization for both sparse linear regression and sparse logistic regression. We implement both the L1 penalty and minimax concave penalty for variable selection and Laplacian penalty for coefficient smoothing. Efficient coordinate descent algorithm is used to solve the optimization problem. We demonstrate the use of the package by applying it to a human microbiome dataset, where phylogeny structure among bacterial taxa is available. AVAILABILITY AND IMPLEMENTATION: 'glmgraph' is implemented in R and C++ Armadillo and publicly available under CRAN. Li Chen 0029, Han Liu 0001, Jean-Pierre A. Kocher, Hongzhe Li |
Bioinform. | 4 |
| 2015 | Power and sample-size estimation for microbiome studies using pairwise distances and PERMANOVAabstractMOTIVATION: The variation in community composition between microbiome samples, termed beta diversity, can be measured by pairwise distance based on either presence-absence or quantitative species abundance data. PERMANOVA, a permutation-based extension of multivariate analysis of variance to a matrix of pairwise distances, partitions within-group and between-group distances to permit assessment of the effect of an exposure or intervention (grouping factor) upon the sampled microbiome. Within-group distance and exposure/intervention effect size must be accurately modeled to estimate statistical power for a microbiome study that will be analyzed with pairwise distances and PERMANOVA. RESULTS: We present a framework for PERMANOVA power estimation tailored to marker-gene microbiome studies that will be analyzed by pairwise distances, which includes: (i) a novel method for distance matrix simulation that permits modeling of within-group pairwise distances according to pre-specified population parameters; (ii) a method to incorporate effects of different sizes within the simulated distance matrix; (iii) a simulation-based method for estimating PERMANOVA power from simulated distance matrices; and (iv) an R statistical software package that implements the above. Matrices of pairwise distances can be efficiently simulated to satisfy the triangle inequality and incorporate group-level effects, which are quantified by the adjusted coefficient of determination, omega-squared (ω2). From simulated distance matrices, available PERMANOVA power or necessary sample size can be estimated for a planned microbiome study. Brendan J. Kelly, Robert Gross, Kyle Bittinger, Scott A. Sherrill-Mix, James D. Lewis, Ronald G. Collman, Frederic D. Bushman, Hongzhe Li |
Bioinform. | 8 |
| 2014 | A change-point model for identifying 3′UTR switching by next-generation RNA sequencingabstractMOTIVATION: Next-generation RNA sequencing offers an opportunity to investigate transcriptome in an unprecedented scale. Recent studies have revealed widespread alternative polyadenylation (polyA) in eukaryotes, leading to various mRNA isoforms differing in their 3' untranslated regions (3'UTR), through which, the stability, localization and translation of mRNA can be regulated. However, very few, if any, methods and tools are available for directly analyzing this special alternative RNA processing event. Conventional methods rely on annotation of polyA sites; yet, such knowledge remains incomplete, and identification of polyA sites is still challenging. The goal of this article is to develop methods for detecting 3'UTR switching without any prior knowledge of polyA annotations. RESULTS: We propose a change-point model based on a likelihood ratio test for detecting 3'UTR switching. We develop a directional testing procedure for identifying dramatic shortening or lengthening events in 3'UTR, while controlling mixed directional false discovery rate at a nominal level. To our knowledge, this is the first approach to analyze 3'UTR switching directly without relying on any polyA annotations. Simulation studies and applications to two real datasets reveal that our proposed method is powerful, accurate and feasible for the analysis of next-generation RNA sequencing data. CONCLUSIONS: The proposed method will fill a void among alternative RNA processing analysis tools for transcriptome studies. It can help to obtain additional insights from RNA sequencing data by understanding gene regulation mechanisms through the analysis of 3'UTR switching. AVAILABILITY AND IMPLEMENTATION: The software is implemented in Java and can be freely downloaded from http://utr.sourceforge.net/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wei Wang 0080, Zhi Wei 0001, Hongzhe Li |
Bioinform. | 3 |
| 2013 | Software Vulnerability Detection Using Backward Trace Analysis and Symbolic ExecutionabstractSoftware vulnerability has long been considered an important threat to the safety of software systems. When source code is accessible, we can get much help from the information of source code to detect vulnerabilities. Static analysis has been used frequently to scan code for errors that cause security problems when source code is available. However, they often generate many false positives. Symbolic execution has also been proposed to detect vulnerabilities and has shown good performance in some researches. However, they are either ineffective in path exploration or could not scale well to large programs. During practical use, since most of paths are actually not related to security problems and software vulnerabilities are usually caused by the improper use of security-sensitive functions, the number of paths could be reduced by tracing sensitive data backwardly from security-sensitive functions so as to consider paths related to vulnerabilities only. What's more, in order to leave ourselves free from generating bug triggering test input, formal reasoning could be used by solving certain program conditions. In this research, we propose backward trace analysis and symbolic execution to detect vulnerabilities from source code. We first find out all the hot spot in source code file. Based on each hot spot, we construct a data flow tree so that we can get the possible execution traces. Afterwards, we do symbolic execution to generate program constraint(PC) and get security constraint(SC) from our predefined security requirements along each execution trace. A program constraint is a constraint imposed by program logic on program variables. A security constraint(SC) is a constraint on program variables that must be satisfied to ensure system security. Finally, this hot spot will be reported as a vulnerability if there is an assignment of values to program inputs which could satisfy PC but violates SC, in other words, satisfy PC Λ S̅C̅. We have implemented our approach and conducted experiments on test cases which we randomly choose from Juliet Test Suites provided by US National Security Agency(NSA). The results show that our approach achieves Precision value of 83.33%, Recall value of 90.90% and F1 Value of 86.95% which gains the best performance among competing tools. Moreover, our approach can efficiently mitigate path explosion problem in traditional symbolic execution. Hongzhe Li, Taebeom Kim, Munkhbayar Bat-Erdene, Heejo Lee |
ARES | 1 |
| 2012 | Associating microbiome composition with environmental covariates using generalized UniFrac distancesabstractMOTIVATION: The human microbiome plays an important role in human disease and health. Identification of factors that affect the microbiome composition can provide insights into disease mechanism as well as suggest ways to modulate the microbiome composition for therapeutical purposes. Distance-based statistical tests have been applied to test the association of microbiome composition with environmental or biological covariates. The unweighted and weighted UniFrac distances are the most widely used distance measures. However, these two measures assign too much weight either to rare lineages or to most abundant lineages, which can lead to loss of power when the important composition change occurs in moderately abundant lineages. RESULTS: We develop generalized UniFrac distances that extend the weighted and unweighted UniFrac distances for detecting a much wider range of biologically relevant changes. We evaluate the use of generalized UniFrac distances in associating microbiome composition with environmental covariates using extensive Monte Carlo simulations. Our results show that tests using the unweighted and weighted UniFrac distances are less powerful in detecting abundance change in moderately abundant lineages. In contrast, the generalized UniFrac distance is most powerful in detecting such changes, yet it retains nearly all its power for detecting rare and highly abundant lineages. The generalized UniFrac distance also has an overall better power than the joint use of unweighted/weighted UniFrac distances. Application to two real microbiome datasets has demonstrated gains in power in testing the associations between human microbiome and diet intakes and habitual smoking. AVAILABILITY: http://cran.r-project.org/web/packages/GUniFrac Kyle Bittinger, Emily S. Charlson, Christian Hoffmann 0005, James D. Lewis, Gary D. Wu, Ronald G. Collman, Frederic D. Bushman, Hongzhe Li |
Bioinform. | 9 |
| 2010 | Co-expression networks: graph properties and topological comparisonsabstractMOTIVATION: Microarray-based gene expression data have been generated widely to study different biological processes and systems. Gene co-expression networks are often used to extract information about groups of genes that are 'functionally' related or co-regulated. However, the structural properties of such co-expression networks have not been rigorously studied and fully compared with known biological networks. In this article, we aim at investigating the structural properties of co-expression networks inferred for the species Saccharomyces Cerevisiae and comparing them with the topological properties of the known, well-established transcriptional network, MIPS physical network and protein-protein interaction (PPI) network of yeast. RESULTS: These topological comparisons indicate that co-expression networks are not distinctly related with either the PPI or the MIPS physical interaction networks, showing important structural differences between them. When focusing on a more literal comparison, vertex by vertex and edge by edge, the conclusion is the same: the fact that two genes exhibit a high gene expression correlation degree does not seem to obviously correlate with the existence of a physical binding between the proteins produced by these genes or the existence of a MIPS physical interaction between the genes. The comparison of the yeast regulatory network with inferred yeast co-expression networks would suggest, however, that they could somehow be related. CONCLUSIONS: We conclude that the gene expression-based co-expression networks reflect more on the gene regulatory networks but less on the PPI or MIPS physical interaction networks. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ramon Xulvi-Brunet, Hongzhe Li |
Bioinform. | 2 |
| 2008 | Network-constrained regularization and variable selection for analysis of genomic dataabstractMOTIVATION: Graphs or networks are common ways of depicting information. In biology in particular, many different biological processes are represented by graphs, such as regulatory networks or metabolic pathways. This kind of a priori information gathered over many years of biomedical research is a useful supplement to the standard numerical genomic data such as microarray gene-expression data. How to incorporate information encoded by the known biological networks or graphs into analysis of numerical data raises interesting statistical challenges. In this article, we introduce a network-constrained regularization procedure for linear regression analysis in order to incorporate the information from these graphs into an analysis of the numerical data, where the network is represented as a graph and its corresponding Laplacian matrix. We define a network-constrained penalty function that penalizes the L(1)-norm of the coefficients but encourages smoothness of the coefficients on the network. RESULTS: Simulation studies indicated that the method is quite effective in identifying genes and subnetworks that are related to disease and has higher sensitivity than the commonly used procedures that do not use the pathway structure information. Application to one glioblastoma microarray gene-expression dataset identified several subnetworks on several of the Kyoto Encyclopedia of Genes and Genomes (KEGG) transcriptional pathways that are related to survival from glioblastoma, many of which were supported by published literatures. CONCLUSIONS: The proposed network-constrained regularization procedure efficiently utilizes the known pathway structures in identifying the relevant genes and the subnetworks that might be related to phenotype in a general regression framework. As more biological networks are identified and documented in databases, the proposed method should find more applications in identifying the subnetworks that are related to diseases and other biological processes. Caiyan Li, Hongzhe Li |
Bioinform. | 2 |
| 2008 | In Response to Comment on "Network-constrained regularization and variable selection for analysis of genomic data"abstractAbstract Contact: [email protected] Caiyan Li, Hongzhe Li |
Bioinform. | 2 |
| 2007 | Group SCAD regression analysis for microarray time course gene expression dataabstractMOTIVATION: Since many important biological systems or processes are dynamic systems, it is important to study the gene expression patterns over time in a genomic scale in order to capture the dynamic behavior of gene expression. Microarray technologies have made it possible to measure the gene expression levels of essentially all the genes during a given biological process. In order to determine the transcriptional factors (TFs) involved in gene regulation during a given biological process, we propose to develop a functional response model with varying coefficients in order to model the transcriptional effects on gene expression levels and to develop a group smoothly clipped absolute deviation (SCAD) regression procedure for selecting the TFs with varying coefficients that are involved in gene regulation during a biological process. RESULTS: Simulation studies indicated that such a procedure is quite effective in selecting the relevant variables with time-varying coefficients and in estimating the coefficients. Application to the yeast cell cycle microarray time course gene expression data set identified 19 of the 21 known TFs related to the cell cycle process. In addition, we have identified another 52 TFs that also have periodic transcriptional effects on gene expression during the cell cycle process. Compared to simple linear regression (SLR) analysis at each time point, our procedure identified more known cell cycle related TFs. CONCLUSIONS: The proposed group SCAD regression procedure is very effective for identifying variables with time-varying coefficients, in particular, for identifying the TFs that are related to gene expression over time. By identifying the TFs that are related to gene expression variations over time, the procedure can potentially provide more insight into the gene regulatory networks. Hongzhe Li |
Bioinform. | 3 |
| 2007 | A Markov random field model for network-based analysis of genomic dataabstractMOTIVATION: A central problem in genomic research is the identification of genes and pathways involved in diseases and other biological processes. The genes identified or the univariate test statistics are often linked to known biological pathways through gene set enrichment analysis in order to identify the pathways involved. However, most of the procedures for identifying differentially expressed (DE) genes do not utilize the known pathway information in the phase of identifying such genes. In this article, we develop a Markov random field (MRF)-based method for identifying genes and subnetworks that are related to diseases. Such a procedure models the dependency of the DE patterns of genes on the networks using a local discrete MRF model. RESULTS: Simulation studies indicated that the method is quite effective in identifying genes and subnetworks that are related to disease and has higher sensitivity and lower false discovery rates than the commonly used procedures that do not use the pathway structure information. Applications to two breast cancer microarray gene expression datasets identified several subnetworks on several of the KEGG transcriptional pathways that are related to breast cancer recurrence or survival due to breast cancer. CONCLUSIONS: The proposed MRF-based model efficiently utilizes the known pathway structures in identifying the DE genes and the subnetworks that might be related to phenotype. As more biological networks are identified and documented in databases, the proposed method should find more applications in identifying the subnetworks that are related to diseases and other biological processes. Zhi Wei 0001, Hongzhe Li |
Bioinform. | 2 |
| 2005 | Penalized Cox regression analysis in the high-dimensional and low-sample size settings, with applications to microarray gene expression dataabstractMOTIVATION: An important application of microarray technology is to relate gene expression profiles to various clinical phenotypes of patients. Success has been demonstrated in molecular classification of cancer in which the gene expression data serve as predictors and different types of cancer serve as a categorical outcome variable. However, there has been less research in linking gene expression profiles to the censored survival data such as patients' overall survival time or time to cancer relapse. It would be desirable to have models with good prediction accuracy and parsimony property. RESULTS: We propose to use the L(1) penalized estimation for the Cox model to select genes that are relevant to patients' survival and to build a predictive model for future prediction. The computational difficulty associated with the estimation in the high-dimensional and low-sample size settings can be efficiently solved by using the recently developed least-angle regression (LARS) method. Our simulation studies and application to real datasets on predicting survival after chemotherapy for patients with diffuse large B-cell lymphoma demonstrate that the proposed procedure, which we call the LARS-Cox procedure, can be used for identifying important genes that are related to time to death due to cancer and for building a parsimonious model for predicting the survival of future patients. The LARS-Cox regression gives better predictive performance than the L(2) penalized regression and a few other dimension-reduction based methods. CONCLUSIONS: We conclude that the proposed LARS-Cox procedure can be very useful in identifying genes relevant to survival phenotypes and in building a parsimonious predictive model that can be used for classifying future patients into clinically relevant high- and low-risk groups based on the gene expression profile and survival times of previous patients. Jiang Gui, Hongzhe Li |
Bioinform. | 2 |
| 2005 | Boosting proportional hazards models using smoothing splines, with applications to high-dimensional microarray dataabstractMOTIVATION: An important area of research in the postgenomics era is to relate high-dimensional genetic or genomic data to various clinical phenotypes of patients. Due to large variability in time to certain clinical events among patients, studying possibly censored survival phenotypes can be more informative than treating the phenotypes as categorical variables. Due to high dimensionality and censoring, building a predictive model for time to event is more difficult than the classification/linear regression problem. We propose to develop a boosting procedure using smoothing splines for estimating the general proportional hazards models. Such a procedure can potentially be used for identifying non-linear effects of genes on the risk of developing an event. RESULTS: Our empirical simulation studies showed that the procedure can indeed recover the true functional forms of the covariates and can identify important variables that are related to the risk of an event. Results from predicting survival after chemotherapy for patients with diffuse large B-cell lymphoma demonstrate that the proposed method can be used for identifying important genes that are related to time to death due to cancer and for building a parsimonious model for predicting the survival of future patients. In addition, there is clear evidence of non-linear effects of some genes on survival time. Hongzhe Li, Yihui Luan |
Bioinform. | 1 |
| 2004 | Dimension reduction methods for microarrays with application to censored survival dataabstractMOTIVATION: Recent research has shown that gene expression profiles can potentially be used for predicting various clinical phenotypes, such as tumor class, drug response and survival time. While there has been extensive studies on tumor classification, there has been less emphasis on other phenotypic features, in particular, patient survival time or time to cancer recurrence, which are subject to right censoring. We consider in this paper an analysis of censored survival time based on microarray gene expression profiles. RESULTS: We propose a dimension reduction strategy, which combines principal components analysis and sliced inverse regression, to identify linear combinations of genes, that both account for the variability in the gene expression levels and preserve the phenotypic information. The extracted gene combinations are then employed as covariates in a predictive survival model formulation. We apply the proposed method to a large diffuse large-B-cell lymphoma dataset, which consists of 240 patients and 7399 genes, and build a Cox proportional hazards model based on the derived gene expression components. The proposed method is shown to provide a good predictive performance for patient survival, as demonstrated by both the significant survival difference between the predicted risk groups and the receiver operator characteristics analysis. AVAILABILITY: R programs are available upon request from the authors. SUPPLEMENTARY INFORMATION: http://dna.ucdavis.edu/~hli/bioinfo-surv-supp.pdf. Lexin Li, Hongzhe Li |
Bioinform. | 2 |
| 2004 | Model-based methods for identifying periodically expressed genes based on time course microarray gene expression dataabstractMOTIVATION: The expressions of many genes associated with certain periodic biological and cell cycle processes such as circadian rhythm regulation are known to be rhythmic. Identification of the genes whose time course expressions are synchronized to certain periodic biological process may help to elucidate the molecular basis of many diseases, and these gene products may in turn represent drug targets relevant to those diseases. RESULTS: We propose in this paper a statistical framework based on a shape-invariant model together with a false discovery rate (FDR) procedure for identifying periodically expressed genes based on microarray time-course gene expression data and a set of known periodically expressed guide genes. We applied the proposed methods to the alpha-factor, cdc15 and cdc28 synchronized yeast cell cycle data sets and identified a total of 1010 cell-cycle-regulated genes at a FDR of 0.5% in at least one of the three data sets analyzed, including 89 (86%) of 104 known periodic transcripts. We also identified 344 and 201 circadian rhythmic genes in vivo in mouse heart and liver tissues with FDR of 10 and 2.5%, respectively. Our results also indicate that the shape-invariant model fits the data well and provides estimate of the common shape function and the relative phases for these periodically regulated genes. Yihui Luan, Hongzhe Li |
Bioinform. | 2 |
| 2003 | Clustering of time-course gene expression data using a mixed-effects model with B-splinesabstractMOTIVATION: Time-course gene expression data are often measured to study dynamic biological systems and gene regulatory networks. To account for time dependency of the gene expression measurements over time and the noisy nature of the microarray data, the mixed-effects model using B-splines was introduced. This paper further explores such mixed-effects model in analyzing the time-course gene expression data and in performing clustering of genes in a mixture model framework. RESULTS: After fitting the mixture model in the framework of the mixed-effects model using an EM algorithm, we obtained the smooth mean gene expression curve for each cluster. For each gene, we obtained the best linear unbiased smooth estimate of its gene expression trajectory over time, combining data from that gene and other genes in the same cluster. Simulated data indicate that the methods can effectively cluster noisy curves into clusters differing in either the shapes of the curves or the times to the peaks of the curves. We further demonstrate the proposed method by clustering the yeast genes based on their cell cycle gene expression data and the human genes based on the temporal transcriptional response of fibroblasts to serum. Clear periodic patterns and varying times to peaks are observed for different clusters of the cell-cycle regulated genes. Results of the analysis of the human fibroblasts data show seven distinct transcriptional response profiles with biological relevance. AVAILABILITY: Matlab programs are available on request from the authors. Yihui Luan, Hongzhe Li |
Bioinform. | 2 |