Qi Zhao 0009

dblp:05/490-9 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
5since 2021 · last 2023
0000-0002-8683-6145ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2023 dSCOPE: a software to detect sequences critical for liquid-liquid phase separation
abstract
Membrane-based cells are the fundamental structural and functional units of organisms, while evidences demonstrate that liquid-liquid phase separation (LLPS) is associated with the formation of membraneless organelles, such as P-bodies, nucleoli and stress granules. Many studies have been undertaken to explore the functions of protein phase separation (PS), but these studies lacked an effective tool to identify the sequence segments that critical for LLPS. In this study, we presented a novel software called dSCOPE (http://dscope.omicsbio.info) to predict the PS-driving regions. To develop the predictor, we curated experimentally identified sequence segments that can drive LLPS from published literature. Then sliding sequence window based physiological, biochemical, structural and coding features were integrated by random forest algorithm to perform prediction. Through rigorous evaluation, dSCOPE was demonstrated to achieve satisfactory performance. Furthermore, large-scale analysis of human proteome based on dSCOPE showed that the predicted PS-driving regions enriched various protein post-translational modifications and cancer mutations, and the proteins which contain predicted PS-driving regions enriched critical cellular signaling pathways. Taken together, dSCOPE precisely predicted the protein sequence segments critical for LLPS, with various helpful information visualized in the webserver to facilitate LLPS-related research.
Kai Yu 0011, Haoyang Cheng, Huai-Qiang Ju, Zhixiang Zuo, Qi Zhao 0009, Shiyang Kang, Zexian Liu
Briefings Bioinform.9
2022 Deciphering clonal dynamics and metastatic routines in a rare patient of synchronous triple-primary tumors and multiple metastases with MPTevol
abstract
Multiple primary tumor (MPT) is a special and rare cancer type, defined as more than two primary tumors presenting at the diagnosis in a single patient. The molecular characteristics and tumorigenesis of MPT remain unclear due to insufficient approaches. Here, we present MPTevol, a practical computational framework for comprehensively exploring the MPT from multiregion sequencing (MRS) experiments. To verify the utility of MPTevol, we performed whole-exome MRS for 33 samples of a rare patient with triple-primary tumors and three metastatic sites and systematically investigated clonal dynamics and metastatic routines. MPTevol assists in comparing genomic profiles across samples, detecting clonal evolutionary history and metastatic routines and quantifying the metastatic history. All triple-primary tumors were independent origins and their genomic characteristics were consistent with corresponding sporadic tumors, strongly supporting their independent tumorigenesis. We further showed two independent early monoclonal seeding events for the metastases in the ovary and uterus. We revealed that two ovarian metastases were disseminated from the same subclone of the primary tumor through undergoing whole-genome doubling processes, suggesting metastases-to-metastases seeding occurred when tumors had similar microenvironments. Surprisingly, according to the metastasis timing model of MPTevol, we found that primary tumors of about 0.058-0.124 cm diameter have been disseminating to distant organs, which is much earlier than conventional clinical views. We developed MPT-specialized analysis framework MPTevol and demonstrated its utility in explicitly resolving clonal evolutionary history and metastatic seeding routines with a rare MPT case. MPTevol is implemented in R and is available at https://github.com/qingjian1991/MPTevol under the GPL v3 license.
Qingjian Chen, Qi-Nian Wu, Yu-Ming Rong, Shixiang Wang, Zhixiang Zuo, Long Bai 0006, Shuqiang Yuan, Qi Zhao 0009
Briefings Bioinform.9
2022 MeRIPseqPipe: an integrated analysis pipeline for MeRIP-seq data based on Nextflow
abstract
SUMMARY: MeRIPseqPipe is an integrated and automatic pipeline that can provide users a friendly solution to perform in-depth mining of MeRIP-seq data. It integrates many functional analysis modules, range from basic processing to downstream analysis. All the processes are embedded in Nextflow with Docker support, which ensures high reproducibility and scalability of the analysis. MeRIPseqPipe is particularly suitable for analyzing a large number of samples at once with a simple command. The final output directory is structured based on each step and tool. And visualization reports containing various tables and plots are provided as HTML files. AVAILABILITY AND IMPLEMENTATION: MeRIPseqPipe is freely available at https://github.com/canceromics/MeRIPseqPipe. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaoqiong Bao, Kaiyu Zhu, Xuefei Liu, Zhihang Chen 0002, Ziwei Luo 0001, Qi Zhao 0009, Jian Ren 0002, Zhixiang Zuo
Bioinform.6
2022 DrugCVar: a platform for evidence-based drug annotation for genetic variants in cancer
abstract
MOTIVATION: Targeted therapy for cancer-related genetic variants is critical for precision medicine. Although several databases including The Clinical Interpretation of Variants in Cancer (CIViC), The Oncology Knowledge Base (OncoKB), The Cancer Genome Interpreter (CGI) and My Cancer Genome (MCG) provide clinical interpretations of variants in cancer, the clinical evidence was limited and miscellaneous. In this study, we developed the DrugCVar database, which integrated our manually curated cancer variant-drug targeting evidence from literature and the interpretations from the public resources. RESULTS: In total, 7830 clinical evidences for cancer variant-drug targeting were integrated and classified into 10 evidence tiers. Searching and browsing functions were provided for quick queries of cancer variant-drug targeting evidence. Also, batch annotation module was developed for user-provided massive genetic variants in various formats. Details, such as the mutation function, location of the variants in gene and protein structures and mutation statistics of queried genes in various tumor types, were also provided for further investigations. Thus, DrugCVar could serve as a comprehensive annotation tool to interpret potential drugs for cancer variants especially the massive ones from clinical cancer genomics studies. AVAILABILITY AND IMPLEMENTATION: The database is available at http://drugcvar.omicsbio.info. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhikai Qian, Kai Yu 0011, Yongqiang Zheng, Qi Zhao 0009, Zexian Liu
Bioinform.8
2021 gutMEGA: a database of the human gut MEtaGenome Atlas
abstract
The gut microbiota plays important roles in human health through regulating both physiological homeostasis and disease emergence. The accumulation of metagenomic sequencing studies enables us to better understand the temporal and spatial variations of the gut microbiota under different physiological and pathological conditions. However, it is inconvenient for scientists to query and retrieve published data; thus, a comprehensive resource for the quantitative gut metagenome is urgently needed. In this study, we developed gut MEtaGenome Atlas (gutMEGA), a well-annotated comprehensive database, to curate and host published quantitative gut microbiota datasets from Homo sapiens. By carefully curating the gut microbiota composition, phenotypes and experimental information, gutMEGA finally integrated 59 132 quantification events for 6457 taxa at seven different levels (kingdom, phylum, class, order, family, genus and species) under 776 conditions. Moreover, with various browsing and search functions, gutMEGA provides a fast and simple way for users to obtain the relative abundances of intestinal microbes among phenotypes. Overall, gutMEGA is a convenient and comprehensive resource for gut metagenome research, which can be freely accessed at http://gutmega.omicsbio.info.
Kai Yu 0011, Qi Zhao 0009, Zexian Liu, Xiaoxing Li
Briefings Bioinform.5
2020 Deep learning based prediction of reversible HAT/HDAC-specific lysine acetylation
abstract
Protein lysine acetylation regulation is an important molecular mechanism for regulating cellular processes and plays critical physiological and pathological roles in cancers and diseases. Although massive acetylation sites have been identified through experimental identification and high-throughput proteomics techniques, their enzyme-specific regulation remains largely unknown. Here, we developed the deep learning-based protein lysine acetylation modification prediction (Deep-PLA) software for histone acetyltransferase (HAT)/histone deacetylase (HDAC)-specific acetylation prediction based on deep learning. Experimentally identified substrates and sites of several HATs and HDACs were curated from the literature to generate enzyme-specific data sets. We integrated various protein sequence features with deep neural network and optimized the hyperparameters with particle swarm optimization, which achieved satisfactory performance. Through comparisons based on cross-validations and testing data sets, the model outperformed previous studies. Meanwhile, we found that protein-protein interactions could enrich enzyme-specific acetylation regulatory relations and visualized this information in the Deep-PLA web server. Furthermore, a cross-cancer analysis of acetylation-associated mutations revealed that acetylation regulation was intensively disrupted by mutations in cancers and heavily implicated in the regulation of cancer signaling. These prediction and analysis results might provide helpful information to reveal the regulatory mechanism of protein acetylation in various biological processes to promote the research on prognosis and treatment of cancers. Therefore, the Deep-PLA predictor and protein acetylation interaction networks could provide helpful information for studying the regulation of protein acetylation. The web server of Deep-PLA could be accessed at http://deeppla.cancerbio.info.
Kai Yu 0011, Yimeng Du, Xinjiao Gao, Qi Zhao 0009, Xiaoxing Li, Zexian Liu
Briefings Bioinform.6
2020 CrossICC: iterative consensus clustering of cross-platform gene expression data without adjusting batch effect
abstract
Unsupervised clustering of high-throughput gene expression data is widely adopted for cancer subtyping. However, cancer subtypes derived from a single dataset are usually not applicable across multiple datasets from different platforms. Merging different datasets is necessary to determine accurate and applicable cancer subtypes but is still embarrassing due to the batch effect. CrossICC is an R package designed for the unsupervised clustering of gene expression data from multiple datasets/platforms without the requirement of batch effect adjustment. CrossICC utilizes an iterative strategy to derive the optimal gene signature and cluster numbers from a consensus similarity matrix generated by consensus clustering. This package also provides abundant functions to visualize the identified subtypes and evaluate subtyping performance. We expected that CrossICC could be used to discover the robust cancer subtypes with significant translational implications in personalized care for cancer patients. AVAILABILITY AND IMPLEMENTATION: The package is implemented in R and available at GitHub (https://github.com/bioinformatist/CrossICC) and Bioconductor (http://bioconductor.org/packages/release/bioc/html/CrossICC.html) under the GPL v3 License.
Qi Zhao 0009, Yu Sun 0050, Hongwan Zhang, Xingyang Li, Kaiyu Zhu, Zexian Liu, Jian Ren 0002, Zhixiang Zuo
Briefings Bioinform.1
2015 IBS: an illustrator for the presentation and visualization of biological sequences
abstract
UNLABELLED: Biological sequence diagrams are fundamental for visualizing various functional elements in protein or nucleotide sequences that enable a summarization and presentation of existing information as well as means of intuitive new discoveries. Here, we present a software package called illustrator of biological sequences (IBS) that can be used for representing the organization of either protein or nucleotide sequences in a convenient, efficient and precise manner. Multiple options are provided in IBS, and biological sequences can be manipulated, recolored or rescaled in a user-defined mode. Also, the final representational artwork can be directly exported into a publication-quality figure. AVAILABILITY AND IMPLEMENTATION: The standalone package of IBS was implemented in JAVA, while the online service was implemented in HTML5 and JavaScript. Both the standalone package and online service are freely available at http://ibs.biocuckoo.org. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wenzhong Liu, Yubin Xie, Jiyong Ma, Xiaotong Luo, Zhixiang Zuo, Urs Lahrmann, Qi Zhao 0009, Yueyuan Zheng, Yong Zhao 0013, Yu Xue 0001, Jian Ren 0002
Bioinform.8