Xiaoqi Zheng

dblp:54/7927 · DBLP profile ↗
← Back
27ranked-venue papers
3as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 15 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 StackAge: an ensemble-based clock for precise quantification of biological age using multi-omics data
abstract
Accurate quantification of biological age is essential for early risk stratification and intervention of chronic diseases. Here, we present StackAge, an ensemble-based biological aging clock that integrates large-scale plasma proteomic and metabolomic profiles from 30 376 participants in the UK Biobank. StackAge demonstrated high accuracy in age prediction (Pearson r ≈ 0.93 with chronological age) and substantially enhanced risk prediction for 12 chronic diseases, achieving AUCs exceeding 0.90 for type 2 diabetes, Alzheimer's disease, and chronic kidney disease. Notably, the incorporation of estimated aging rates consistently improved disease prediction beyond conventional omics and demographic features. Feature interpretation and pathway enrichment analyses revealed that aging-associated biomarkers were enriched in inflammation, metabolic stress, and extracellular matrix remodeling pathways. Mediation analysis further indicated that modifiable lifestyle factors may accelerate biological aging, thereby increasing susceptibility to cardiovascular, neurological, immune, and musculoskeletal disorders. Together, these findings establish a robust multi-omics framework for quantifying individual aging trajectories and highlight biological age as a clinically actionable indicator for precision prevention and health management of age-related diseases.
Yingyi Jiang, Yuan Fei, Xiaoqi Zheng, Yufang Qin
Briefings Bioinform.5
2026 Castl: robust identification of spatially variable genes in spatial transcriptomics via an ensemble-based framework
abstract
Spatially variable genes (SVGs) are essential for elucidating tissue organization within spatially resolved transcriptomics. While a number of computational methods have been developed for SVG identification, their reliance on algorithm-specific assumptions, such as predefined kernel functions or spatial neighborhood graphs, often results in substantial variability in sensitivity and inflated false discovery rates (FDRs) across heterogeneous datasets. To address this challenge, we here develop Castl, an ensemble-based framework for SVG identification that integrates multiple detection methods through statistically designed aggregation modules. Comprehensive evaluations on both simulated and real-world data demonstrate that Castl consistently identifies biologically meaningful spatial expression patterns, mitigates method-specific biases and effectively controls FDRs across various biological contexts, resolutions, and spatial technologies. This flexible, assumption-free framework offers a robust and standardized foundation for spatially informed feature discovery in complex biological systems.
Yiyi Yu, Ping-An He 0001, Xiaoqi Zheng
Briefings Bioinform.4
2026 Deep Prior Framework: integrating functional specificity with general plausibility for targeted protein evolution
abstract
The efficiency of directed protein evolution largely relies on computational methods to enrich mutants with high fitness. Traditional strategies, such as zero-shot approaches based on Protein Language Models (PLMs), primarily leverage general "plausibility" priors learned from natural sequences. However, in the absence of experimental feedback, their ability to guide evolution toward specific functions ("specificity") remains limited. Here, we introduce the Deep Prior Framework (DPF), a novel paradigm that integrates universal structural plausibility with task-oriented specificity priors. DPF incorporates an innovative Bernoulli-Attention (BATT) module within a Mixture of Experts architecture, enabling efficient screening of high-fitness mutants. Benchmarking on nine deep mutational scanning datasets, DPF outperforms existing methods in terms of PLM-based method. More importantly, our method also shows high performance on in silico directed evolution of Blastobotrys adeninivorans xanthine dehydrogenase (BaXD) without intermediate experimental feedback. Experimental validation of the top-ranked mutants showed an average activity enhancement of over four-fold compared with WT, with the best mutant achieving more than a nine-fold improvement. Furthermore, we applied DPF to a large-scale annotation of unreviewed sequences in UniProt Knowledgebase (UniProtKB). Of the 42 913 366 predicted samples (~21.56% of the total), 90.07% (38 934 581 proteins) were assigned high-confidence functional labels. In summary, this study demonstrates that DPF, by incorporating specificity-aware functional priors, can significantly advance efficient and targeted protein engineering.
Senxin Zhang, Yining Qin, Hanwen Zhu, Feilong Meng, Xiaoqi Zheng
Briefings Bioinform.6
2026 Spider: a flexible and unified framework for simulating spatial transcriptomics data
abstract
MOTIVATION: Spatial transcriptomics (ST) technologies provide valuable insights into cellular heterogeneity by simultaneously acquiring both gene expression profiles and cellular location information. However, the limited diversity and accuracy of "gold standard" datasets hindered the effectiveness and fairness of benchmarking rapidly growing ST analysis tools. RESULTS: To address this issue, we proposed Spider, a flexible and comprehensive framework for simulating ST data without requiring real ST data as a reference. By characterizing the spatial patterns using cell type proportions and transition matrix between adjacent cells, Spider can produce more realistic and diverse simulated data and offer enhanced modeling flexibility compared to existing simulation methods. Additionally, Spider provides interactive features for customizing the spatial domain, such as zone segmentation and integration of histology imaging data. Benchmark analyses demonstrate that Spider outperforms other simulation tools in preserving the spatial characteristics of real ST data and facilitating the evaluation of downstream analysis methods. Spider is implemented in Python and available at https://github.com/YANG-ERA/Spider. AVAILABILITY AND IMPLEMENTATION: All codes, simulated ST data in this paper are publicly available at https://github.com/YANG-ERA/Spider.
Nana Wei, Congcong Hu, Hua-Jun Wu, Xiaoqi Zheng
Bioinform.8
2026 CAM-Interacted Vision GNN for Multi-Label Medical Images
abstract
Vision Graph Neural Network (ViG) is designed to recognize different objects through graph-level processing. However, ViG constructs graphs with appearance-level neighbors and neglects the category semantic. The oversight results in the unintentional connection of patches that belong to different objects, thus affecting the distinctiveness of categories in multi-label medical image learning. Since the pixel-level annotations for images are not easily available, category-aware graphs can not be directly built. To solve this problem, we consider localizing category-specific regions using Class Activation Maps (CAMs), an effective way to highlight regions belonging to each category without requiring manual annotations. Specifically, we propose a CAM-interacted Vision GNN (CiV-GNN), in which category-aware graphs are formed to perform intra-category graph processing. CIV-GNN includes a Class-activated Patch Division (CAPD) module, which introduces CAMs as guidance for category-aware graph building. Furthermore, we develop a Multi-graph Interactive Processing (MIP) module to model the relations between category-aware graphs, promoting inter-category interaction learning. Experimental results show that CiV-GNN performs well in surgical tool localization and multi-label medical image classification. Specifically, for m2cai16-localization, CiV-GNN exhibits a 1.43% and 7.02% improvement in mAP50 and mAP50-95, respectively, compared to YOLOv8.
Jingchao Wang 0002, Baoyao Yang, Si-Qi Liu 0003, Xiaoqi Zheng, Wenbin Yao, Junxiang Chen
IEEE J. Biomed. Health Informatics4
2025 Image-assisted Label Connective Completion for Vessel Segmentation with Insufficient Annotations
abstract
Automatic and accurate vessel segmentation is crucial for disease diagnosis. Deep learning methods are widely used, but their promising results rely on accurately annotated data. Due to complex vessel morphology and low-contrast image, accurate vessel delineation poses a practical challenge, resulting in insufficient annotations, which is a prominent form of noisy labels. This paper proposes an Image-assisted Label Connective Completion method, which enhances label’s vessel information by images under the supervision of connectivity to address insufficient annotation issue. Specifically, we develop an Image-guided Vessel Enhancement module, which transmits structural information extracted from images based on label navigation to label space, promoting completion of missing annotated parts in original labels. In addition, a branch completion-connectivity loss is designed and introduced as an auxiliary supervision to prevent vessel branch disconnection during label completion. Experimental results on DRIVE, CHASE DB1 and DCA1 datasets demonstrate that our method outperforms existing noisy labels learning methods.
Xiaoqi Zheng, Baoyao Yang, Xiuwen Fang, Wenfang Yao, Mang Ye
ICASSP1
2024 BS-clock, advancing epigenetic age prediction with high-resolution DNA methylation bisulfite sequencing data
abstract
MOTIVATION: DNA methylation patterns provide precise and accurate estimates of biological age due to their robustness and predictable changes associated with aging processes. Although several methylation aging clocks have been developed in recent years, they are primarily designed for DNA methylation array data, which has limited CpG coverage and detection sensitivity compared to bisulfite sequencing data. RESULTS: Here, we present BS-clock, a novel DNA methylation clock for human aging based on bisulfite sequencing data. Using BS-seq data from 529 samples retrieved from four tissues, our BS-clock achieves higher correlations with chronological age in multiple tissue types compared to existing array-based clocks. Our study revealed age-dependent aging rates across different age stages and disease conditions, and overall low cross-tissue prediction capability by applying the model trained on one tissue type to others. In summary, BS-clock overcomes limitations of array-based techniques, offering genome-wide CpG site coverage and more robust and accurate aging quantification. This research paves the way for advanced epigenetic studies of aging and holds promise for developing targeted interventions to promote healthy aging. AVAILABILITY AND IMPLEMENTATION: All analysis codes for reproducing the results of the study are publicly available at https://github.com/hucongcong97/BS-clock.
Congcong Hu, Naiqian Zhang, Xiaoqi Zheng
Bioinform.5
2024 Embedding enhancement with foreground feature alignment and primitive knowledge for few-shot learning
Xiaoqi Zheng
Eng. Appl. Artif. Intell.1
2024 A novel hypergraph model for identifying and prioritizing personalized drivers in cancer
abstract
Cancer development is driven by an accumulation of a small number of driver genetic mutations that confer the selective growth advantage to the cell, while most passenger mutations do not contribute to tumor progression. The identification of these driver genes responsible for tumorigenesis is a crucial step in designing effective cancer treatments. Although many computational methods have been developed with this purpose, the majority of existing methods solely provided a single driver gene list for the entire cohort of patients, ignoring the high heterogeneity of driver events across patients. It remains challenging to identify the personalized driver genes. Here, we propose a novel method (PDRWH), which aims to prioritize the mutated genes of a single patient based on their impact on the abnormal expression of downstream genes across a group of patients who share the co-mutation genes and similar gene expression profiles. The wide experimental results on 16 cancer datasets from TCGA showed that PDRWH excels in identifying known general driver genes and tumor-specific drivers. In the comparative testing across five cancer types, PDRWH outperformed existing individual-level methods as well as cohort-level methods. Our results also demonstrated that PDRWH could identify both common and rare drivers. The personalized driver profiles could improve tumor stratification, providing new insights into understanding tumor heterogeneity and taking a further step toward personalized treatment. We also validated one of our predicted novel personalized driver genes on tumor cell proliferation by vitro cell-based assays, the promoting effect of the high expression of Low-density lipoprotein receptor-related protein 1 (LRP1) on tumor cell proliferation.
Naiqian Zhang, Fubin Ma, Yuxuan Pang, Chenye Wang, Yusen Zhang 0002, Xiaoqi Zheng
PLoS Comput. Biol.7
2023 TRAmHap: accurate prediction of transcriptional activity from DNA methylation haplotypes in bisulfite-sequencing data
abstract
Deoxyribonucleic acid (DNA) methylation (DNAm) is an important epigenetic mechanism that plays a role in chromatin structure and transcriptional regulation. Elucidating the relationship between DNAm and gene expression is of great importance for understanding its role in transcriptional regulation. The conventional approach is to construct machine-learning-based methods to predict gene expression based on mean methylation signals in promoter regions. However, this type of strategy only explains about 25% of gene expression variation, and hence is inadequate in elucidating the relationship between DNAm and transcriptional activity. In addition, using mean methylation as input features neglects the heterogeneity of cell populations that can be reflected by DNAm haplotypes. We here developed TRAmaHap, a novel deep-learning framework that predicts gene expression by utilizing the characteristics of DNAm haplotypes in proximal promoters and distal enhancers. Using benchmark data of human and mouse normal tissues, TRAmHap shows much higher accuracy than existing machine-learning based methods, by explaining 60~80% of gene expression variation across tissue types and disease conditions. Our model demonstrated that gene expression can be accurately predicted by DNAm patterns in promoters and long-range enhancers as far as 25 kb away from transcription start site, especially in the presence of intra-gene chromatin interactions.
Hanwen Zhu, Kangwen Cai, Leiqin Liu, Yaochen Xu, Xiaoqi Zheng
Briefings Bioinform.8
2023 ExosomePurity: tumour purity deconvolution in serum exosomes based on miRNA signatures
abstract
Exosomes cargo tumour-characterized biomolecules secreted from cancer cells and play a pivotal role in tumorigenesis and cancer progression, thus providing their potential for non-invasive cancer monitoring. Since cancer cell-derived exosomes are often mixed with those from healthy cells in liquid biopsy of tumour patients, accurately measuring the purity of tumour cell-derived exosomes is not only critical for the early detection but also essential for unbiased identification of diagnosis biomarkers. Here, we propose 'ExosomePurity', a tumour purity deconvolution model to estimate tumour purity in serum exosomes of cancer patients based on microribonucleic acid (miRNA)-Seq data. We first identify the differently expressed miRNAs as signature to distinguish cancer cell- from healthy cell-derived exosomes. Then, the deconvolution model was developed to estimate the proportions of cancer exosomes and normal exosomes in serum. The purity predicted by the model shows high correlation with actual purity in simulated data and actual data. Moreover, the model is robust under the different levels of noise background. The tumour purity was also used to correct differential expressed gene analysis. ExosomePurity empowers the research community to study non-invasive early diagnosis and to track cancer progression in cancers more efficiently. It is implemented in R and is freely available from GitHub (https://github.com/WangHYLab/ExosomePurity).
Yao Dai, Xiaoqi Zheng
Briefings Bioinform.8
2023 E-value: a superior alternative to P-value and its adjustments in DNA methylation studies
abstract
DNA methylation plays a crucial role in transcriptional regulation. Reduced representation bisulfite sequencing (RRBS) is a technique of increasing use for analyzing genome-wide methylation profiles. Many computational tools such as Metilene, MethylKit, BiSeq and DMRfinder have been developed to use RRBS data for the detection of the differentially methylated regions (DMRs) potentially involved in epigenetic regulations of gene expression. For DMR detection tools, as for countless other medical applications, P-values and their adjustments are among the most standard reporting statistics used to assess the statistical significance of biological findings. However, P-values are coming under increasing criticism relating to their questionable accuracy and relatively high levels of false positive or negative indications. Here, we propose a method to calculate E-values, as likelihood ratios falling into the null hypothesis over the entire parameter space, for DMR detection in RRBS data. We also provide the R package 'metevalue' as a user-friendly interface to implement E-value calculations into various DMR detection tools. To evaluate the performance of E-values, we generated various RRBS benchmarking datasets using our simulator 'RRBSsim' with eight samples in each experimental group. Our comprehensive benchmarking analyses showed that using E-values not only significantly improved accuracy, area under ROC curve and power, over that of P-values or adjusted P-values, but also reduced false discovery rates and type I errors. In applications using real RRBS data of CRL rats and a clinical trial on low-salt diet, the use of E-values detected biologically more relevant DMRs and also improved the negative association between DNA methylation and gene expression.
Yifan Yang 0001, Liyuan Zhou, Xiaoqi Zheng, Rongxian Yue, David L. Mattson, Srividya Kidambi, Mingyu Liang, Pengyuan Liu 0003, Xiaoqing Pan
Briefings Bioinform.5
2022 Purification of tumor methylomes through residual decomposition
abstract
Due to the high heterogeneity of tumor tissue, methylation profiles of tumor samples obtained in clinical experiments are always mixture signals from different cellular components, including cancer, normal and stromal cells, etc. Among them, the admixture of normal cells is deemed as a major confounding factor for many downstream analyses. Decomposing mixture signals into profiles of their primitive constituents is vital for accurate differential calling and patient grouping. However, methods for purification of tumor methylomes are still lacking, even given a reliable estimate of tumor purity. In this work, we present ResDec, a residual-decomposition linear regression model for tumor methylome purification. We systematically evaluated the performance of our method compared with existing methods on both simulation data and TCGA methylation samples. ResDec achieves consistently better performance under different scenarios, including different numbers of matched normal samples, perturbations of input tumor purities and matched normal methylomes.
Nana Wei, Yijing Zhu, Yating Nie, Shiyu Fan, Yuanchen Sun, Xiaoqi Zheng
BIBM6
2022 mHapTk: a comprehensive toolkit for the analysis of DNA methylation haplotypes
abstract
SUMMARY: Bisulfite sequencing remains the gold standard technique to detect DNA methylation profiles at single-nucleotide resolution. The DNA methylation status of CpG sites on the same fragment represents a discrete methylation haplotype (mHap). The mHap-level metrics were demonstrated to be promising cancer biomarkers and explain more gene expression variation than average methylation. However, most existing tools focus on average methylation and neglect mHap patterns. Here, we present mHapTk, a comprehensive python toolkit for the analysis of DNA mHap. It calculates eight mHap-level summary statistics in predefined regions or across individual CpG in a genome-wide manner. It identifies methylation haplotype blocks, in which methylations of pairwise CpGs are tightly correlated. Furthermore, mHap patterns can be visualized with the built-in functions in mHapTk or external tools such as IGV and deepTools. AVAILABILITY AND IMPLEMENTATION: https://jiantaoshi.github.io/mhaptk/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kangwen Cai, Leiqin Liu, Xiaoqi Zheng
Bioinform.5
2022 DriverRWH: discovering cancer driver genes by random walk on a gene mutation hypergraph
abstract
BACKGROUND: Recent advances in next-generation sequencing technologies have helped investigators generate massive amounts of cancer genomic data. A critical challenge in cancer genomics is identification of a few cancer driver genes whose mutations cause tumor growth. However, the majority of existing computational approaches underuse the co-occurrence mutation information of the individuals, which are deemed to be important in tumorigenesis and tumor progression, resulting in high rate of false positive. RESULTS: To make full use of co-mutation information, we present a random walk algorithm referred to as DriverRWH on a weighted gene mutation hypergraph model, using somatic mutation data and molecular interaction network data to prioritize candidate driver genes. Applied to tumor samples of different cancer types from The Cancer Genome Atlas, DriverRWH shows significantly better performance than state-of-art prioritization methods in terms of the area under the curve scores and the cumulative number of known driver genes recovered in top-ranked candidate genes. Besides, DriverRWH discovers several potential drivers, which are enriched in cancer-related pathways. DriverRWH recovers approximately 50% known driver genes in the top 30 ranked candidate genes for more than half of the cancer types. In addition, DriverRWH is also highly robust to perturbations in the mutation data and gene functional network data. CONCLUSION: DriverRWH is effective among various cancer types in prioritizes cancer driver genes and provides considerable improvement over other tools with a better balance of precision and sensitivity. It can be a useful tool for detecting potential driver genes and facilitate targeted cancer therapies.
Chenye Wang, Junhan Shi, Jiansheng Cai, Yusen Zhang 0002, Xiaoqi Zheng, Naiqian Zhang
BMC Bioinform.5
2022 Secuer: Ultrafast, scalable and accurate clustering of single-cell RNA-seq data
abstract
Identifying cell clusters is a critical step for single-cell transcriptomics study. Despite the numerous clustering tools developed recently, the rapid growth of scRNA-seq volumes prompts for a more (computationally) efficient clustering method. Here, we introduce Secuer, a Scalable and Efficient speCtral clUstERing algorithm for scRNA-seq data. By employing an anchor-based bipartite graph representation algorithm, Secuer enjoys reduced runtime and memory usage over one order of magnitude for datasets with more than 1 million cells. Meanwhile, Secuer also achieves better or comparable accuracy than competing methods in small and moderate benchmark datasets. Furthermore, we showcase that Secuer can also serve as a building block for a new consensus clustering method, Secuer-consensus, which again improves the runtime and scalability of state-of-the-art consensus clustering methods while also maintaining the accuracy. Overall, Secuer is a versatile, accurate, and scalable clustering framework suitable for small to ultra-large single-cell clustering tasks.
Nana Wei, Yating Nie, Xiaoqi Zheng, Hua-Jun Wu
PLoS Comput. Biol.4
2021 The DNA methylation haplotype (mHap) format and mHapTools
abstract
SUMMARY: Bisulfite sequencing (BS-seq) is currently the gold standard for measuring genome-wide DNA methylation profiles at single-nucleotide resolution. Most analyses focus on mean CpG methylation and ignore methylation states on the same DNA fragments [DNA methylation haplotypes (mHaps)]. Here, we propose mHap, a simple DNA mHap format for storing DNA BS-seq data. This format reduces the size of a BAM file by 40- to 140-fold while retaining complete read-level CpG methylation information. It is also compatible with the Tabix tool for fast and random access. We implemented a command-line tool, mHapTools, for converting BAM/SAM files from existing platforms to mHap files as well as post-processing DNA methylation data in mHap format. With this tool, we processed all publicly available human reduced representation bisulfite sequencing data and provided these data as a comprehensive mHap database. AVAILABILITY AND IMPLEMENTATION: https://jiantaoshi.github.io/mHap/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuhao Dan, Yaochen Xu, Xiaoqi Zheng
Bioinform.5
2021 TimNet: A text-image matching network integrating multi-stage feature extraction with multi-scale metrics
Xiaoqi Zheng, Yingfan Tao, Ruikai Zhang, Wenming Yang, Qingmin Liao
Neurocomputing1
2020 A comprehensive review of computational prediction of genome-wide features
abstract
There are significant correlations among different types of genetic, genomic and epigenomic features within the genome. These correlations make the in silico feature prediction possible through statistical or machine learning models. With the accumulation of a vast amount of high-throughput data, feature prediction has gained significant interest lately, and a plethora of papers have been published in the past few years. Here we provide a comprehensive review on these published works, categorized by the prediction targets, including protein binding site, enhancer, DNA methylation, chromatin structure and gene expression. We also provide discussions on some important points and possible future directions.
Tianlei Xu, Xiaoqi Zheng, Zhaohui S. Qin, Hao Wu 0003
Briefings Bioinform.2
2020 Detection of differentially methylated CpG sites between tumor samples with uneven tumor purities
abstract
MOTIVATION: Inference of differentially methylated (DM) CpG sites between two groups of tumor samples with different geno- or pheno-types is a critical step to uncover the epigenetic mechanism of tumorigenesis, and identify biomarkers for cancer subtyping. However, as a major source of confounding factor, uneven distributions of tumor purity between two groups of tumor samples will lead to biased discovery of DM sites if not properly accounted for. RESULTS: We here propose InfiniumDM, a generalized least square model to adjust tumor purity effect for differential methylation analysis. Our method is applicable to a variety of experimental designs including with or without normal controls, different sources of normal tissue contaminations. We compared our method with conventional methods including minfi, limma and limma corrected by tumor purity using simulated datasets. Our method shows significantly better performance at different levels of differential methylation thresholds, sample sizes, mean purity deviations and so on. We also applied the proposed method to breast cancer samples from TCGA database to further evaluate its performance. Overall, both simulation and real data analyses demonstrate favorable performance over existing methods serving similar purpose. AVAILABILITY AND IMPLEMENTATION: InfiniumDM is a part of R package InfiniumPurify, which is freely available from GitHub (https://github.com/Xiaoqizheng/InfiniumPurify). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ziyi Li 0001, Nana Wei, Hua-Jun Wu, Xiaoqi Zheng
Bioinform.5
2020 Deconvolution of heterogeneous tumor samples using partial reference signals
abstract
Deconvolution of heterogeneous bulk tumor samples into distinct cellular populations is an important yet challenging problem, particularly when only partial references are available. A common approach to dealing with this problem is to deconvolve the mixed signals using available references and leverage the remaining signal as a new cell component. However, as indicated in our simulation, such an approach tends to over-estimate the proportions of known cell types and fails to detect novel cell types. Here, we propose PREDE, a partial reference-based deconvolution method using an iterative non-negative matrix factorization algorithm. Our method is verified to be effective in estimating cell proportions and expression profiles of unknown cell types based on simulated datasets at a variety of parameter settings. Applying our method to TCGA tumor samples, we found that proportions of pure cancer cells better indicate different subtypes of tumor samples. We also detected several cell types for each cancer type whose proportions successfully predicted patient survival. Our method makes a significant contribution to deconvolution of heterogeneous tumor samples and could be widely applied to varieties of high throughput bulk data. PREDE is implemented in R and is freely available from GitHub (https://xiaoqizheng.github.io/PREDE).
Yufang Qin, Siwei Nan, Nana Wei, Hua-Jun Wu, Xiaoqi Zheng
PLoS Comput. Biol.7
2019 Comprehensive anticancer drug response prediction based on a simple cell line-drug complex network model
abstract
BACKGROUND: Accurate prediction of anticancer drug responses in cell lines is a crucial step to accomplish the precision medicine in oncology. Although many popular computational models have been proposed towards this non-trivial issue, there is still room for improving the prediction performance by combining multiple types of genome-wide molecular data. RESULTS: We first demonstrated an observation on the CCLE and GDSC datasets, i.e., genetically similar cell lines always exhibit higher response correlations to structurally related drugs. Based on this observation we built a cell line-drug complex network model, named CDCN model. It captures different contributions of all available cell line-drug responses through cell line similarities and drug similarities. We executed anticancer drug response prediction on CCLE and GDSC independently. The result is significantly superior to that of some existing studies. More importantly, our model could predict the response of new drug to new cell line with considerable performance. We also divided all possible cell lines into "sensitive" and "resistant" groups by their response values to a given drug, the prediction accuracy, sensitivity, specificity and goodness of fit are also very promising. CONCLUSION: CDCN model is a comprehensive tool to predict anticancer drug responses. Compared with existing methods, it is able to provide more satisfactory prediction results with less computational consumption.
Chuanying Liu, Xiaoqi Zheng, Yushuang Li
BMC Bioinform.3
2017 Accounting for tumor purity improves cancer subtype classification from DNA methylation data
abstract
MOTIVATION: Tumor sample classification has long been an important task in cancer research. Classifying tumors into different subtypes greatly benefits therapeutic development and facilitates application of precision medicine on patients. In practice, solid tumor tissue samples obtained from clinical settings are always mixtures of cancer and normal cells. Thus, the data obtained from these samples are mixed signals. The 'tumor purity', or the percentage of cancer cells in cancer tissue sample, will bias the clustering results if not properly accounted for. RESULTS: In this article, we developed a model-based clustering method and an R function which uses DNA methylation microarray data to infer tumor subtypes with the consideration of tumor purity. Simulation studies and the analyses of The Cancer Genome Atlas data demonstrate improved results compared with existing methods. AVAILABILITY AND IMPLEMENTATION: InfiniumClust is part of R package InfiniumPurify , which is freely available from CRAN ( https://cran.r-project.org/web/packages/InfiniumPurify/index.html ). CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hao Feng 0005, Hao Wu 0003, Xiaoqi Zheng
Bioinform.4
2015 Predicting tumor purity from methylation microarray data
abstract
MOTIVATION: In cancer genomics research, one important problem is that the solid tissue sample obtained from clinical settings is always a mixture of cancer and normal cells. The sample mixture brings complication in data analysis and results in biased findings if not correctly accounted for. Estimating tumor purity is of great interest, and a number of methods have been developed using gene expression, copy number variation or point mutation data. RESULTS: We discover that in cancer samples, the distributions of data from Illumina Infinium 450 k methylation microarray are highly correlated with tumor purities. We develop a simple but effective method to estimate purities from the microarray data. Analyses of the Cancer Genome Atlas lung cancer data demonstrate favorable performance of the proposed method. AVAILABILITY AND IMPLEMENTATION: The method is implemented in InfiniumPurify, which is freely available at https://bitbucket.org/zhengxiaoqi/infiniumpurify. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Naiqian Zhang, Hua-Jun Wu, Hao Wu 0003, Xiaoqi Zheng
Bioinform.6
2015 Predicting Anticancer Drug Responses Using a Dual-Layer Integrated Cell Line-Drug Network Model
abstract
The ability to predict the response of a cancer patient to a therapeutic agent is a major goal in modern oncology that should ultimately lead to personalized treatment. Existing approaches to predicting drug sensitivity rely primarily on profiling of cancer cell line panels that have been treated with different drugs and selecting genomic or functional genomic features to regress or classify the drug response. Here, we propose a dual-layer integrated cell line-drug network model, which uses both cell line similarity network (CSN) data and drug similarity network (DSN) data to predict the drug response of a given cell line using a weighted model. Using the Cancer Cell Line Encyclopedia (CCLE) and Cancer Genome Project (CGP) studies as benchmark datasets, our single-layer model with CSN or DSN and only a single parameter achieved a prediction performance comparable to the previously generated elastic net model. When using the dual-layer model integrating both CSN and DSN, our predicted response reached a 0.6 Pearson correlation coefficient with observed responses for most drugs, which is significantly better than the previous results using the elastic net model. We have also applied the dual-layer cell line-drug integrated network model to fill in the missing drug response values in the CGP dataset. Even though the dual-layer integrated cell line-drug network model does not specifically model mutation information, it correctly predicted that BRAF mutant cell lines would be more sensitive than BRAF wild-type cell lines to three MEK1/2 inhibitors tested.
Naiqian Zhang, Xiaoqi Zheng, Xiaole Shirley Liu
PLoS Comput. Biol.5
2014 Sequence-based identification of recombination spots using pseudo nucleic acid representation and recursive feature extraction by linear kernel SVM
abstract
BACKGROUND: Identification of the recombination hot/cold spots is critical for understanding the mechanism of recombination as well as the genome evolution process. However, experimental identification of recombination spots is both time-consuming and costly. Developing an accurate and automated method for reliably and quickly identifying recombination spots is thus urgently needed. RESULTS: Here we proposed a novel approach by fusing features from pseudo nucleic acid composition (PseNAC), including NAC, n-tier NAC and pseudo dinucleotide composition (PseDNC). A recursive feature extraction by linear kernel support vector machine (SVM) was then used to rank the integrated feature vectors and extract optimal features. SVM was adopted for identifying recombination spots based on these optimal features. To evaluate the performance of the proposed method, jackknife cross-validation test was employed on a benchmark dataset. The overall accuracy of this approach was 84.09%, which was higher (from 0.37% to 3.79%) than those of state-of-the-art tools. CONCLUSIONS: Comparison results suggested that linear kernel SVM is a useful vehicle for identifying recombination hot/cold spots.
Liqi Li, Sanjiu Yu, Yongsheng Li 0003, Xiaoqi Zheng, Shiwen Zhou 0002
BMC Bioinform.6
2013 Prioritization of candidate disease genes by topological similarity between disease and protein diffusion profiles
abstract
BACKGROUND: Identification of gene-phenotype relationships is a fundamental challenge in human health clinic. Based on the observation that genes causing the same or similar phenotypes tend to correlate with each other in the protein-protein interaction network, a lot of network-based approaches were proposed based on different underlying models. A recent comparative study showed that diffusion-based methods achieve the state-of-the-art predictive performance. RESULTS: In this paper, a new diffusion-based method was proposed to prioritize candidate disease genes. Diffusion profile of a disease was defined as the stationary distribution of candidate genes given a random walk with restart where similarities between phenotypes are incorporated. Then, candidate disease genes are prioritized by comparing their diffusion profiles with that of the disease. Finally, the effectiveness of our method was demonstrated through the leave-one-out cross-validation against control genes from artificial linkage intervals and randomly chosen genes. Comparative study showed that our method achieves improved performance compared to some classical diffusion-based methods. To further illustrate our method, we used our algorithm to predict new causing genes of 16 multifactorial diseases including Prostate cancer and Alzheimer's disease, and the top predictions were in good consistent with literature reports. CONCLUSIONS: Our study indicates that integration of multiple information sources, especially the phenotype similarity profile data, and introduction of global similarity measure between disease and gene diffusion profiles are helpful for prioritizing candidate disease genes. AVAILABILITY: Programs and data are available upon request.
Yufang Qin, Taigang Liu, Xiaoqi Zheng
BMC Bioinform.5