EDBT 2026 Demo / reviewers in the wild / expert
Ziyi Li 0001
dblp:143/8841-1
· DBLP profile ↗
18ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0001-8359-0533ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 5 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TUSCAN: Tumor segmentation and classification analysis in spatial transcriptomicsabstractThe identification of tumor cells is pivotal for understanding tumor heterogeneity and the tumor microenvironment. Recent advances in spatially resolved transcriptomics (SRT) have revolutionized the way that transcriptomic profiles are characterized and have enabled the simultaneous quantification of transcript locations in intact tissue samples. SRT is a promising alternative method to study gene expression patterns in spatial domains. Nevertheless, the precise detection of tumor regions within intact tissue remains a great challenge. A common strategy for identifying tumor cells is via tumor-specific marker gene expression signatures, which are highly dependent on marker accuracy. Another effective approach is through aneuploid copy number alterations, as most types of cancer exhibit copy number abnormalities. Here, we introduce a novel computational method, called TUSCAN (TUmor Segmentation and Classification ANalysis in spatial transcriptomics), which constructs a spatial copy number variation profile to improve the accuracy of tumor region identification. TUSCAN combines gene information from SRT data and hematoxylin-and-eosin-staining image to annotate tumor sections and other benign tissues. We benchmark the performance of TUSCAN and several existing methods through the application to multiple datasets from different SRT platforms. We demonstrate that TUSCAN can effectively delineate tumor regions, with improved accuracy compared to other approaches. Additionally, the output of TUSCAN provides interpretable clonal evolution inferences that may lead to novel insights into disease development and potential druggable targets. Chenxuan Zang, Charles C. Guo, Yaohong Wang, Ziyi Li 0001 |
PLoS Comput. Biol. | 5 |
| 2025 | PoweREST: Statistical power estimation for spatial transcriptomics experiments to detect differentially expressed genes between two conditionsabstractRecent advancements in spatial transcriptomics (ST) have significantly enhanced biological research in various domains. However, the high cost for current ST data generation techniques restricts the large-scale application of ST. Consequently, maximization of the use of available resources to achieve robust statistical power for ST data is a pressing need. One fundamental question in ST analysis is detection of differentially expressed genes (DEGs) under different conditions using ST data. Such DEG analyses are performed frequently, but their power calculations are rarely discussed in the literature. To address this gap, we developed PoweREST, a power estimation tool designed to support the power calculation for DEG detection with 10X Genomics Visium data. PoweREST enables power estimation both before any ST experiments and after preliminary data are collected, making it suitable for a wide variety of power analyses in ST studies. We also provide a user-friendly, program-free web application that allows users to interactively calculate and visualize study power along with relevant parameters. Lan Shui, Anirban Maitra, Ken Lau, Harsimran Kaur, Liang Li 0026, Ziyi Li 0001 |
PLoS Comput. Biol. | 7 |
| 2024 | SCIntRuler: guiding the integration of multiple single-cell RNA-seq datasets with a novel statistical metricabstractMOTIVATION: The growing number of single-cell RNA-seq (scRNA-seq) studies highlights the potential benefits of integrating multiple datasets, such as augmenting sample sizes and enhancing analytical robustness. Inherent diversity and batch discrepancies within samples or across studies continue to pose significant challenges for computational analyses. Questions persist in practice, lacking definitive answers: Should we use a specific integration method or opt for simply merging the datasets during joint analysis? Among all the existing data integration methods, which one is more suitable in specific scenarios? RESULT: To fill the gap, we introduce SCIntRuler, a novel statistical metric for guiding the integration of multiple scRNA-seq datasets. SCIntRuler helps researchers make informed decisions regarding the necessity of data integration and the selection of an appropriate integration method. Our simulations and real data applications demonstrate that SCIntRuler streamlines decision-making processes and facilitates the analysis of diverse scRNA-seq datasets under varying contexts, thereby alleviating the complexities associated with the integration of heterogeneous scRNA-seq datasets. AVAILABILITY AND IMPLEMENTATION: The implementation of our method is available on CRAN as an open-source R package with a user-friendly manual available: https://cloud.r-project.org/web/packages/SCIntRuler/index.html. Yue Lyu, Steven H. Lin, Hao Wu 0003, Ziyi Li 0001 |
Bioinform. | 4 |
| 2024 | Regional analysis to delineate intrasample heterogeneity with RegionalSTabstractMOTIVATION: Spatial transcriptomics has greatly contributed to our understanding of spatial and intra-sample heterogeneity, which could be crucial for deciphering the molecular basis of human diseases. Intra-tumor heterogeneity, e.g. may be associated with cancer treatment responses. However, the lack of computational tools for exploiting cross-regional information and the limited spatial resolution of current technologies present major obstacles to elucidating tissue heterogeneity. RESULTS: To address these challenges, we introduce RegionalST, an efficient computational method that enables users to quantify cell type mixture and interactions, identify sub-regions of interest, and perform cross-region cell type-specific differential analysis for the first time. Our simulations and real data applications demonstrate that RegionalST is an efficient tool for visualizing and analyzing diverse spatial transcriptomics data, thereby enabling accurate and flexible exploration of tissue heterogeneity. Overall, RegionalST provides a one-stop destination for researchers seeking to delve deeper into the intricacies of spatial transcriptomics data. AVAILABILITY AND IMPLEMENTATION: The implementation of our method is available as an open-source R/Bioconductor package with a user-friendly manual available at https://bioconductor.org/packages/release/bioc/html/RegionalST.html. Yue Lyu, Ziyi Li 0001 |
Bioinform. | 4 |
| 2023 | A novel statistical method for decontaminating T-cell receptor sequencing dataabstractThe T-cell receptor (TCR) repertoire is highly diverse among the population and plays an essential role in initiating multiple immune processes. TCR sequencing (TCR-seq) has been developed to profile the T cell repertoire. Similar to other high-throughput experiments, contamination can happen during several steps of TCR-seq, including sample collection, preparation and sequencing. Such contamination creates artifacts in the data, leading to inaccurate or even biased results. Most existing methods assume 'clean' TCR-seq data as the starting point with no ability to handle data contamination. Here, we develop a novel statistical model to systematically detect and remove contamination in TCR-seq data. We summarize the observed contamination into two sources, pairwise and cross-cohort. For both sources, we provide visualizations and summary statistics to help users assess the severity of the contamination. Incorporating prior information from 14 existing TCR-seq datasets with minimum contamination, we develop a straightforward Bayesian model to statistically identify contaminated samples. We further provide strategies for removing the impacted sequences to allow for downstream analysis, thus avoiding any need to repeat experiments. Our proposed model shows robustness in contamination detection compared with a few off-the-shelf detection methods in simulation studies. We illustrate the use of our proposed method on two TCR-seq datasets generated locally. Ruoxing Li, Mehmet Altan, Alexandre Reuben, Ruitao Lin, John V. Heymach, Runzhe Chen, Latasha Little, Shawna Hubert, Ziyi Li 0001 |
Briefings Bioinform. | 11 |
| 2023 | A comprehensive assessment of cell type-specific differential expression methods in bulk dataabstractAccounting for cell type compositions has been very successful at analyzing high-throughput data from heterogeneous tissues. Differential gene expression analysis at cell type level is becoming increasingly popular, yielding biomarker discovery in a finer granularity within a particular cell type. Although several computational methods have been developed to identify cell type-specific differentially expressed genes (csDEG) from RNA-seq data, a systematic evaluation is yet to be performed. Here, we thoroughly benchmark six recently published methods: CellDMC, CARseq, TOAST, LRCDE, CeDAR and TCA, together with two classical methods, csSAM and DESeq2, for a comprehensive comparison. We aim to systematically evaluate the performance of popular csDEG detection methods and provide guidance to researchers. In simulation studies, we benchmark available methods under various scenarios of baseline expression levels, sample sizes, cell type compositions, expression level alterations, technical noises and biological dispersions. Real data analyses of three large datasets on inflammatory bowel disease, lung cancer and autism provide evaluation in both the gene level and the pathway level. We find that csDEG calling is strongly affected by effect size, baseline expression level and cell type compositions. Results imply that csDEG discovery is a challenging task itself, with room to improvements on handling low signal-to-noise ratio and low expression genes. Guanqun Meng, Wen Tang 0003, Emina Huang, Ziyi Li 0001, Hao Feng 0005 |
Briefings Bioinform. | 4 |
| 2022 | A comprehensive comparison of supervised and unsupervised methods for cell type identification in single-cell RNA-seqabstractThe cell type identification is among the most important tasks in single-cell RNA-sequencing (scRNA-seq) analysis. Many in silico methods have been developed and can be roughly categorized as either supervised or unsupervised. In this study, we investigated the performances of 8 supervised and 10 unsupervised cell type identification methods using 14 public scRNA-seq datasets of different tissues, sequencing protocols and species. We investigated the impacts of a number of factors, including total amount of cells, number of cell types, sequencing depth, batch effects, reference bias, cell population imbalance, unknown/novel cell type, and computational efficiency and scalability. Instead of merely comparing individual methods, we focused on factors' impacts on the general category of supervised and unsupervised methods. We found that in most scenarios, the supervised methods outperformed the unsupervised methods, except for the identification of unknown cell types. This is particularly true when the supervised methods use a reference dataset with high informational sufficiency, low complexity and high similarity to the query dataset. However, such outperformance could be undermined by some undesired dataset properties investigated in this study, which lead to uninformative and biased reference datasets. In these scenarios, unsupervised methods could be comparable to supervised methods. Our study not only explained the cell typing methods' behaviors under different experimental settings but also provided a general guideline for the choice of method according to the scientific goal and dataset properties. Finally, our evaluation workflow is implemented as a modularized R pipeline that allows future evaluation of new methods. Availability: All the source codes are available at https://github.com/xsun28/scRNAIdent. Xiaochu Lin, Ziyi Li 0001, Hao Wu 0003 |
Briefings Bioinform. | 3 |
| 2022 | NeuCA web server: a neural network-based cell annotation tool with web-app and GUIabstractSUMMARY: Correctly annotating individual cell's type is an important initial step in single-cell RNA sequencing (scRNA-seq) data analysis. Here, we present NeuCA web server, a neural network-based scRNA-seq cell annotation tool with web-app portal and graphical user interface, for automatically assigning cell labels. NeuCA algorithm is accurate and exhaustive, maximizing the usage of measured cells for downstream analysis. NeuCA web server provides over 20 ready-to-use pre-trained classifiers for commonly used tissue types. As the first web-app tool with neural-network infrastructure implemented, NeuCA web will facilitate the research community in analyzing and annotating scRNA-seq data. AVAILABILITY AND IMPLEMENTATION: NeuCA web server is implemented with R Shiny application online at https://statbioinfo.shinyapps.io/NeuCA/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daoyu Duan, Sijia He, Emina Huang, Ziyi Li 0001, Hao Feng 0005 |
Bioinform. | 4 |
| 2022 | A machine learning-based method for automatically identifying novel cells in annotating single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) has been widely used to decompose complex tissues into functionally distinct cell types. The first and usually the most important step of scRNA-seq data analysis is to accurately annotate the cell labels. In recent years, many supervised annotation methods have been developed and shown to be more convenient and accurate than unsupervised cell clustering. One challenge faced by all the supervised annotation methods is the identification of the novel cell type, which is defined as the cell type that is not present in the training data, only exists in the testing data. Existing methods usually label the cells simply based on the correlation coefficients or confidence scores, which sometimes results in an excessive number of unlabeled cells. RESULTS: We developed a straightforward yet effective method combining autoencoder with iterative feature selection to automatically identify novel cells from scRNA-seq data. Our method trains an autoencoder with the labeled training data and applies the autoencoder to the testing data to obtain reconstruction errors. By iteratively selecting features that demonstrate a bi-modal pattern and reclustering the cells using the selected feature, our method can accurately identify novel cells that are not present in the training data. We further combined this approach with a support vector machine to provide a complete solution for annotating the full range of cell types. Extensive numerical experiments using five real scRNA-seq datasets demonstrated favorable performance of the proposed method over existing methods serving similar purposes. AVAILABILITY AND IMPLEMENTATION: Our R software package CAMLU is publicly available through the Zenodo repository (https://doi.org/10.5281/zenodo.7054422) or GitHub repository (https://github.com/ziyili20/CAMLU). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ziyi Li 0001, Yizhuo Wang 0002, Irene Ganan-Gomez, Simona Colla, Kim-Anh Do |
Bioinform. | 1 |
| 2022 | CondiS web app: imputation of censored lifetimes for machine learning-based survival analysisabstractSUMMARY: In the era of big data, machine learning techniques are widely applied to every area in biomedical research including survival analysis. It is well recognized that censoring, which is a common missing issue in survival time data, hampers the direct usage of these machine learning techniques. Here, we present CondiS, a web toolkit with graphical user interface to help impute the survival times for censored observations and predict the survival times for future enrolled patients. CondiS imputes a censored survival time based on its distribution conditional on its observed part. When covariates are available, CondiS-X incorporates this information to further increase the imputation accuracy. Users can also upload data of newly enrolled patients and predict their survival times. As the first web-app tool with an imputation function for censored lifetime data, CondiS web can facilitate conducting survival analysis with machine learning approaches. AVAILABILITY AND IMPLEMENTATION: CondiS is an open-source application implemented with Shiny in R, available free at: https://biostatistics.mdanderson.org/shinyapps/CondiS/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yizhuo Wang 0002, Christopher R. Flowers, Ziyi Li 0001, Xuelin Huang |
Bioinform. | 3 |
| 2022 | EDClust: an EM-MM hybrid method for cell clustering in multiple-subject single-cell RNA sequencingabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) has revolutionized biological research by enabling the measurement of transcriptomic profiles at the single-cell level. With the increasing application of scRNA-seq in larger-scale studies, the problem of appropriately clustering cells emerges when the scRNA-seq data are from multiple subjects. One challenge is the subject-specific variation; systematic heterogeneity from multiple subjects may have a significant impact on clustering accuracy. Existing methods seeking to address such effects suffer from several limitations. RESULTS: We develop a novel statistical method, EDClust, for multi-subject scRNA-seq cell clustering. EDClust models the sequence read counts by a mixture of Dirichlet-multinomial distributions and explicitly accounts for cell-type heterogeneity, subject heterogeneity and clustering uncertainty. An EM-MM hybrid algorithm is derived for maximizing the data likelihood and clustering the cells. We perform a series of simulation studies to evaluate the proposed method and demonstrate the outstanding performance of EDClust. Comprehensive benchmarking on four real scRNA-seq datasets with various tissue types and species demonstrates the substantial accuracy improvement of EDClust compared to existing methods. AVAILABILITY AND IMPLEMENTATION: The R package is freely available at https://github.com/weix21/EDClust. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ziyi Li 0001, Hongkai Ji, Hao Wu 0003 |
Bioinform. | 2 |
| 2022 | CondiS: A conditional survival distribution-based method for censored data imputation overcoming the hurdle in machine learning-based survival analysis
Yizhuo Wang 0002, Christopher R. Flowers, Ziyi Li 0001, Xuelin Huang |
J. Biomed. Informatics | 3 |
| 2021 | Complete deconvolution of DNA methylation signals from complex tissues: a geometric approachabstractMOTIVATION: It is a common practice in epigenetics research to profile DNA methylation on tissue samples, which is usually a mixture of different cell types. To properly account for the mixture, estimating cell compositions has been recognized as an important first step. Many methods were developed for quantifying cell compositions from DNA methylation data, but they mostly have limited applications due to lack of reference or prior information. RESULTS: We develop Tsisal, a novel complete deconvolution method which accurately estimate cell compositions from DNA methylation data without any prior knowledge of cell types or their proportions. Tsisal is a full pipeline to estimate number of cell types, cell compositions and identify cell-type-specific CpG sites. It can also assign cell type labels when (full or part of) reference panel is available. Extensive simulation studies and analyses of seven real datasets demonstrate the favorable performance of our proposed method compared with existing deconvolution methods serving similar purpose. AVAILABILITY AND IMPLEMENTATION: The proposed method Tsisal is implemented as part of the R/Bioconductor package TOAST at https://bioconductor.org/packages/TOAST. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hao Wu 0003, Ziyi Li 0001 |
Bioinform. | 3 |
| 2020 | Robust partial reference-free cell composition estimation from tissue expressionabstractMOTIVATION: In the analysis of high-throughput omics data from tissue samples, estimating and accounting for cell composition have been recognized as important steps. High cost, intensive labor requirements and technical limitations hinder the cell composition quantification using cell-sorting or single-cell technologies. Computational methods for cell composition estimation are available, but they are either limited by the availability of a reference panel or suffer from low accuracy. RESULTS: We introduce TOols for the Analysis of heterogeneouS Tissues TOAST/-P and TOAST/+P, two partial reference-free algorithms for estimating cell composition of heterogeneous tissues based on their gene expression profiles. TOAST/-P and TOAST/+P incorporate additional biological information, including cell-type-specific markers and prior knowledge of compositions, in the estimation procedure. Extensive simulation studies and real data analyses demonstrate that the proposed methods provide more accurate and robust cell composition estimation than existing methods. AVAILABILITY AND IMPLEMENTATION: The proposed methods TOAST/-P and TOAST/+P are implemented as part of the R/Bioconductor package TOAST at https://bioconductor.org/packages/TOAST. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ziyi Li 0001, Zhenxing Guo 0001, Hao Wu 0003 |
Bioinform. | 1 |
| 2020 | Detection of differentially methylated CpG sites between tumor samples with uneven tumor puritiesabstractMOTIVATION: Inference of differentially methylated (DM) CpG sites between two groups of tumor samples with different geno- or pheno-types is a critical step to uncover the epigenetic mechanism of tumorigenesis, and identify biomarkers for cancer subtyping. However, as a major source of confounding factor, uneven distributions of tumor purity between two groups of tumor samples will lead to biased discovery of DM sites if not properly accounted for. RESULTS: We here propose InfiniumDM, a generalized least square model to adjust tumor purity effect for differential methylation analysis. Our method is applicable to a variety of experimental designs including with or without normal controls, different sources of normal tissue contaminations. We compared our method with conventional methods including minfi, limma and limma corrected by tumor purity using simulated datasets. Our method shows significantly better performance at different levels of differential methylation thresholds, sample sizes, mean purity deviations and so on. We also applied the proposed method to breast cancer samples from TCGA database to further evaluate its performance. Overall, both simulation and real data analyses demonstrate favorable performance over existing methods serving similar purpose. AVAILABILITY AND IMPLEMENTATION: InfiniumDM is a part of R package InfiniumPurify, which is freely available from GitHub (https://github.com/Xiaoqizheng/InfiniumPurify). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ziyi Li 0001, Nana Wei, Hua-Jun Wu, Xiaoqi Zheng |
Bioinform. | 2 |
| 2019 | Dissecting differential signals in high-throughput data from complex tissuesabstractMOTIVATION: Samples from clinical practices are often mixtures of different cell types. The high-throughput data obtained from these samples are thus mixed signals. The cell mixture brings complications to data analysis, and will lead to biased results if not properly accounted for. RESULTS: We develop a method to model the high-throughput data from mixed, heterogeneous samples, and to detect differential signals. Our method allows flexible statistical inference for detecting a variety of cell-type specific changes. Extensive simulation studies and analyses of two real datasets demonstrate the favorable performance of our proposed method compared with existing ones serving similar purpose. AVAILABILITY AND IMPLEMENTATION: The proposed method is implemented as an R package and is freely available on GitHub (https://github.com/ziyili20/TOAST). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ziyi Li 0001, Zhijin Wu, Hao Wu 0003 |
Bioinform. | 1 |
| 2019 | Distributed learning from multiple EHR databases: Contextual embedding models for medical events
Ziyi Li 0001, Kirk Roberts, Xiaoqian Jiang, Qi Long |
J. Biomed. Informatics | 1 |
| 2017 | Incorporating biological information in sparse principal component analysis with application to genomic dataabstractBACKGROUND: Sparse principal component analysis (PCA) is a popular tool for dimensionality reduction, pattern recognition, and visualization of high dimensional data. It has been recognized that complex biological mechanisms occur through concerted relationships of multiple genes working in networks that are often represented by graphs. Recent work has shown that incorporating such biological information improves feature selection and prediction performance in regression analysis, but there has been limited work on extending this approach to PCA. In this article, we propose two new sparse PCA methods called Fused and Grouped sparse PCA that enable incorporation of prior biological information in variable selection. RESULTS: Our simulation studies suggest that, compared to existing sparse PCA methods, the proposed methods achieve higher sensitivity and specificity when the graph structure is correctly specified, and are fairly robust to misspecified graph structures. Application to a glioblastoma gene expression dataset identified pathways that are suggested in the literature to be related with glioblastoma. CONCLUSIONS: The proposed sparse PCA methods Fused and Grouped sparse PCA can effectively incorporate prior biological information in variable selection, leading to improved feature selection and more interpretable principal component loadings and potentially providing insights on molecular underpinnings of complex diseases. Ziyi Li 0001, Sandra Safo, Qi Long |
BMC Bioinform. | 1 |