Chunquan Li 0002

dblp:75/127-2 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-4700-5496ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 8 since 2021
YearPublicationVenuePosition
2025 FunlncModel: integrating multi-omic features from upstream and downstream regulatory networks into a machine learning framework to identify functional lncRNAs
abstract
Accumulating evidence indicates that long noncoding RNAs (lncRNAs) play important roles in molecular and cellular biology. Although many algorithms have been developed to reveal their associations with complex diseases by using downstream targets, the upstream (epi)genetic regulatory information has not been sufficiently leveraged to predict the function of lncRNAs in various biological processes. Therefore, we present FunlncModel, a machine learning-based interpretable computational framework, which aims to screen out functional lncRNAs by integrating a large number of (epi)genetic features and functional genomic features from their upstream/downstream multi-omic regulatory networks. We adopted the random forest method to mine nearly 60 features in three categories from >2000 datasets across 11 data types, including transcription factors (TFs), histone modifications, typical enhancers, super-enhancers, methylation sites, and mRNAs. FunlncModel outperformed alternative methods for classification performance in human embryonic stem cell (hESC) (0.95 Area Under Curve (AUROC) and 0.97 Area Under the Precision-Recall Curve (AUPRC)). It could not only infer the most known lncRNAs that influence the states of stem cells, but also discover novel high-confidence functional lncRNAs. We extensively validated FunlncModel's efficacy by up to 27 cancer-related functional prediction tasks, which involved multiple cancer cell growth processes and cancer hallmarks. Meanwhile, we have also found that (epi)genetic regulatory features, such as TFs and histone modifications, serve as strong predictors for revealing the function of lncRNAs. Overall, FunlncModel is a strong and stable prediction model for identifying functional lncRNAs in specific cellular contexts. FunlncModel is available as a web server at https://bio.liclab.net/FunlncModel/.
Yan-Yu Li, Fengcui Qian, Guorui Zhang, Xue-Cang Li, Li-Wei Zhou, Zhengmin Yu, Wei Liu 0187, Qiu-Yu Wang, Chunquan Li 0002
Briefings Bioinform.9
2024 PEGCN: A Single-Cell Type Annotation model based on GCN with Pseudo Labels and Ensemble Learning
abstract
Single-cell type annotation helps to identify specific subpopulations of cells associated with a particular disease, thus providing therapeutic targets for precision medicine. Previous tasks on single-cell type annotation have been performed on the basis of RNA-seq sequencing data, and gradually some studies are now being conducted on scATAC-seq sequencing data as well. Although the researchs in this area has achieved considerable results, it is still constrained by the high dimensionality, high noise, and uneven sample distribution inherent in scATAC-seq datasets. In this paper, we propose a graph convolution model that incorporates ensemble learning. The utilization of Pseudo Labels for cell graph construction significantly mitigates the excessively high anisotropy index in the graph caused by noise, presenting a novel approach to enhancing graph quality. Furthermore, our method introduces a second innovation by effectively addressing the challenge of high-dimensional features in ATAC data through guided feature selection facilitated by ensemble learning, ultimately improving the overall performance and accuracy of the analysis. The model effectively reduces the impact of the previously mentioned problems on the scATAC-seq dataset, and improves the predictive performance of the model. The single-cell type annotation results exceed 90% of the ACC in different datasets and the Macro-average ACC exceeds that of other models.
Fengcui Qian, Chunquan Li 0002, Chunping Ouyang
BIBM5
2023 MGCT: Mutual-Guided Cross-Modality Transformer for Survival Outcome Prediction using Integrative Histopathology-Genomic Features
abstract
The rapidly emerging field of deep learning-based computational pathology has shown promising results in utilizing whole slide images (WSIs) to objectively prognosticate cancer patients. However, most prognostic methods are currently limited to either histopathology or genomics alone, which inevitably reduces their potential to accurately predict patient prognosis. Whereas integrating WSIs and genomic features presents three main challenges: (1) the enormous heterogeneity of gigapixel WSIs which can reach sizes as large as 150,000×150,000 pixels; (2) the absence of a spatially corresponding relationship between histopathology images and genomic molecular data; and (3) the existing early, late, and intermediate multimodal feature fusion strategies struggle to capture the explicit interactions between WSIs and genomics. To ameliorate these issues, we propose the Mutual-Guided Cross-Modality Transformer (MGCT), a weakly-supervised, attention-based multimodal learning framework that can combine histology features and genomic features to model the genotype-phenotype interactions within the tumor microenvironment. To validate the effectiveness of MGCT, we conduct experiments using nearly 3,600 gigapixel WSIs across five different cancer types sourced from The Cancer Genome Atlas (TCGA). Extensive experimental results consistently emphasize that MGCT outperforms the state-of-the-art (SOTA) methods.
Yunzan Liu, Hui Cui 0002, Chunquan Li 0002, Jiquan Ma
BIBM4
2023 CanMethdb: a database for genome-wide DNA methylation annotation in cancers
abstract
MOTIVATION: DNA methylation within gene body and promoters in cancer cells is well documented. An increasing number of studies showed that cytosine-phosphate-guanine (CpG) sites falling within other regulatory elements could also regulate target gene activation, mainly by affecting transcription factors (TFs) binding in human cancers. This led to the urgent need for comprehensively and effectively collecting distinct cis-regulatory elements and TF-binding sites (TFBS) to annotate DNA methylation regulation. RESULTS: We developed a database (CanMethdb, http://meth.liclab.net/CanMethdb/) that focused on the upstream and downstream annotations for CpG-genes in cancers. This included upstream cis-regulatory elements, especially those involving distal regions to genes, and TFBS annotations for the CpGs and downstream functional annotations for the target genes, computed through integrating abundant DNA methylation and gene expression profiles in diverse cancers. Users could inquire CpG-target gene pairs for a cancer type through inputting a genomic region, a CpG, a gene name, or select hypo/hypermethylated CpG sets. The current version of CanMethdb documented a total of 38 986 060 CpG-target gene pairs (with 6 769 130 unique pairs), involving 385 217 CpGs and 18 044 target genes, abundant cis-regulatory elements and TFs for 33 TCGA cancer types. CanMethdb might help biologists perform in-depth studies of target gene regulations based on DNA methylations in cancer. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://github.com/chunquanlipathway/CanMethdb. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jianmei Zhao, Fengcui Qian, Xuecang Li, Zhengmin Yu, Yanyu Li, Yongsan Yang, Qi Pan, Qiuyu Wang, Jian Zhang 0084, Guohua Wang 0001, Chunquan Li 0002
Bioinform.17
2023 SENIES: DNA Shape Enhanced Two-Layer Deep Learning Predictor for the Identification of Enhancers and Their Strength
abstract
Identifying enhancers is a critical task in bioinformatics due to their primary role in regulating gene expression. For this reason, various computational algorithms devoted to enhancer identification have been put forward over the years. More features are extracted from the single DNA sequences to boost the performance. Nevertheless, DNA structural information is neglected, which is an essential factor affecting the binding preferences of transcription factors to regulatory elements like enhancers. Here, we propose SENIES, a DNA shape enhanced deep learning predictor, to identify enhancers and their strength. The predictor consists of two layers where the first layer is for enhancer and non-enhancer identification, and the second layer is for predicting the strength of enhancers. Apart from two common sequence-derived features (i.e., one-hot and k-mer), DNA shape is introduced to describe the 3D structures of DNA sequences. Performance comparison with state-of-the-art methods conducted on public datasets demonstrates the effectiveness and robustness of our predictor. The code implementation of SENIES is publicly available at https://github.com/hlju-liye/SENIES.
Fanhui Kong, Hui Cui 0002, Fan Wang 0026, Chunquan Li 0002, Jiquan Ma
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 GREAP: a comprehensive enrichment analysis software for human genomic regions
abstract
The rapid development of genomic high-throughput sequencing has identified a large number of DNA regulatory elements with abundant epigenetics markers, which promotes the rapid accumulation of functional genomic region data. The comprehensively understanding and research of human functional genomic regions is still a relatively urgent work at present. However, the existing analysis tools lack extensive annotation and enrichment analytical abilities for these regions. Here, we designed a novel software, Genomic Region sets Enrichment Analysis Platform (GREAP), which provides comprehensive region annotation and enrichment analysis capabilities. Currently, GREAP supports 85 370 genomic region reference sets, which cover 634 681 107 regions across 11 different data types, including super enhancers, transcription factors, accessible chromatins, etc. GREAP provides widespread annotation and enrichment analysis of genomic regions. To reflect the significance of enrichment analysis, we used the hypergeometric test and also provided a Locus Overlap Analysis. In summary, GREAP is a powerful platform that provides many types of genomic region sets for users and supports genomic region annotations and enrichment analyses. In addition, we developed a customizable genome browser containing >400 000 000 customizable tracks for visualization. The platform is freely available at http://www.liclab.net/Greap/view/index.
Yongsan Yang, Fengcui Qian, Xuecang Li, Yanyu Li, Liwei Zhou, Qiuyu Wang, Xinyuan Zhou, Jian Zhang 0084, Zhengmin Yu, Ting Cui, Chenchen Feng, Desi Shang, Mengfei Sun, Yuexin Zhang, Huifang Tang, Chunquan Li 0002
Briefings Bioinform.19
2021 ComPAT: A Comprehensive Pathway Analysis Tools
Xiaojie Su, Chenchen Feng, Ziyu Ning, Qiuyu Wang, Yuexin Zhang, Ling Wei, Xinyuan Zhou, Chunquan Li 0002
ICIC (3)11
2021 TRlnc: a comprehensive database for human transcriptional regulatory information of lncRNAs
abstract
Long noncoding RNAs (lncRNAs) have been proven to play important roles in transcriptional processes and biological functions. With the increasing study of human diseases and biological processes, information in human H3K27ac ChIP-seq, ATAC-seq and DNase-seq datasets is accumulating rapidly, resulting in an urgent need to collect and process data to identify transcriptional regulatory regions of lncRNAs. We therefore developed a comprehensive database for human regulatory information of lncRNAs (TRlnc, http://bio.licpathway.net/TRlnc), which aimed to collect available resources of transcriptional regulatory regions of lncRNAs and to annotate and illustrate their potential roles in the regulation of lncRNAs in a cell type-specific manner. The current version of TRlnc contains 8 683 028 typical enhancers/super-enhancers and 32 348 244 chromatin accessibility regions associated with 91 906 human lncRNAs. These regions are identified from over 900 human H3K27ac ChIP-seq, ATAC-seq and DNase-seq samples. Furthermore, TRlnc provides the detailed genetic and epigenetic annotation information within transcriptional regulatory regions (promoter, enhancer/super-enhancer and chromatin accessibility regions) of lncRNAs, including common SNPs, risk SNPs, eQTLs, linkage disequilibrium SNPs, transcription factors, methylation sites, histone modifications and 3D chromatin interactions. It is anticipated that the use of TRlnc will help users to gain in-depth and useful insights into the transcriptional regulatory mechanisms of lncRNAs.
Yanyu Li, Xuecang Li, Yongsan Yang, Fengcui Qian, Zhidong Tang, Jianmei Zhao, Jian Zhang 0084, Xuefeng Bai 0002, Yong Jiang 0004, Jianyuan Zhou, Yuexin Zhang, Liwei Zhou, Jianjun Xie, Enmin Li, Qiuyu Wang, Chunquan Li 0002
Briefings Bioinform.17
2020 HiFreSP: A novel high-frequency sub-pathway mining approach to identify robust prognostic gene signatures
abstract
With the increasing awareness of heterogeneity in cancers, better prediction of cancer prognosis is much needed for more personalized treatment. Recently, extensive efforts have been made to explore the variations in gene expression for better prognosis. However, the prognostic gene signatures predicted by most existing methods have little robustness among different datasets of the same cancer. To improve the robustness of the gene signatures, we propose a novel high-frequency sub-pathways mining approach (HiFreSP), integrating a randomization strategy with gene interaction pathways. We identified a six-gene signature (CCND1, CSF3R, E2F2, JUP, RARA and TCF7) in esophageal squamous cell carcinoma (ESCC) by HiFreSP. This signature displayed a strong ability to predict the clinical outcome of ESCC patients in two independent datasets (log-rank test, P = 0.0045 and 0.0087). To further show the predictive performance of HiFreSP, we applied it to two other cancers: pancreatic adenocarcinoma and breast cancer. The identified signatures show high predictive power in all testing datasets of the two cancers. Furthermore, compared with the two popular prognosis signature predicting methods, the least absolute shrinkage and selection operator penalized Cox proportional hazards model and the random survival forest, HiFreSP showed better predictive accuracy and generalization across all testing datasets of the above three cancers. Lastly, we applied HiFreSP to 8137 patients involving 20 cancer types in the TCGA database and found high-frequency prognosis-associated pathways in many cancers. Taken together, HiFreSP shows higher prognostic capability and greater robustness, and the identified signatures provide clinical guidance for cancer prognosis. HiFreSP is freely available via GitHub: https://github.com/chunquanlipathway/HiFreSP.
Jianmei Zhao, Xuecang Li, Chenchen Feng, Fengcui Qian, Yuejuan Liu, Jian Zhang 0084, Bo Ai 0005, Ziyu Ning, Wei Liu 0187, Xuefeng Bai 0002, Zhiyong Wu 0009, Xiue Xu, Zhidong Tang, Qi Pan, Liyan Xu, Chunquan Li 0002, Qiuyu Wang, Enmin Li
Briefings Bioinform.20
2019 TRCirc: a resource for transcriptional regulation information of circRNAs
abstract
In recent years, high-throughput genomic technologies like chromatin immunoprecipitation sequencing (ChIp-seq) and transcriptome sequencing (RNA-seq) have been becoming both more refined and less expensive, making them more accessible. Many circular RNAs (circRNAs) that originate from back-spliced exons have been identified in various cell lines across different species. However, the regulatory mechanism for transcription of circRNAs remains unclear. Therefore, there is an urgent need to construct a database detailing the transcriptional regulation of circRNAs. TRCirc (http://www.licpathway.net/TRCirc) provides a resource for efficient retrieval, browsing and visualization of transcriptional regulation information of circRNAs. The current version of TRCirc documents 92 375 circRNAs and 161 transcription factors (TFs) from more than 100 cell types and together represent more than 765 000 TF-circRNA regulatory relationships. Furthermore, TRCirc provides other regulatory information about transcription of circRNAs, including their expression, methylation levels, H3K27ac signals in regulation regions and super-enhancers associated with circRNAs. TRCirc provides a convenient, user-friendly interface to search, browse and visualize detailed information about these circRNAs.
Zhidong Tang, Xuecang Li, Jianmei Zhao, Fengcui Qian, Chenchen Feng, Yanyu Li, Jian Zhang 0084, Yong Jiang 0004, Yongsan Yang, Qiuyu Wang, Chunquan Li 0002
Briefings Bioinform.11
2015 Characterizing and optimizing human anticancer drug targets based on topological properties in the context of biological pathways
Jian Zhang 0084, Yan Wang 0023, Desi Shang, Fulong Yu, Wei Liu 0187, Chenchen Feng, Qiuyu Wang, Yanjun Xu, Yuejuan Liu, Xuefeng Bai 0002, Xuecang Li, Chunquan Li 0002
J. Biomed. Informatics13
2014 The detection of risk pathways, regulated by miRNAs, via the integration of sample-matched miRNA-mRNA profiles and pathway structure
Jing Li 0115, Chunquan Li 0002, Junwei Han 0003, Chunlong Zhang, Desi Shang, Qianlan Yao, Yanjun Xu, Wei Liu 0187, Meng Zhou 0003, Haixiu Yang, Xia Li 0004
J. Biomed. Informatics2
2013 Topologically inferring risk-active pathways toward precise cancer classification by directed random walk
abstract
MOTIVATION: The accurate prediction of disease status is a central challenge in clinical cancer research. Microarray-based gene biomarkers have been identified to predict outcome and outperform traditional clinical parameters. However, the robustness of the individual gene biomarkers is questioned because of their little reproducibility between different cohorts of patients. Substantial progress in treatment requires advances in methods to identify robust biomarkers. Several methods incorporating pathway information have been proposed to identify robust pathway markers and build classifiers at the level of functional categories rather than of individual genes. However, current methods consider the pathways as simple gene sets but ignore the pathway topological information, which is essential to infer a more robust pathway activity. RESULTS: Here, we propose a directed random walk (DRW)-based method to infer the pathway activity. DRW evaluates the topological importance of each gene by capturing the structure information embedded in the directed pathway network. The strategy of weighting genes by their topological importance greatly improved the reproducibility of pathway activities. Experiments on 18 cancer datasets showed that the proposed method yielded a more accurate and robust overall performance compared with several existing gene-based and pathway-based classification methods. The resulting risk-active pathways are more reliable in guiding therapeutic selection and the development of pathway-specific therapeutic strategies. AVAILABILITY: DRW is freely available at http://210.46.85.180:8080/DRWPClass/
Wei Liu 0187, Chunquan Li 0002, Yanjun Xu, Haixiu Yang, Qianlan Yao, Junwei Han 0003, Desi Shang, Chunlong Zhang, Yun Xiao 0001, Xia Li 0004
Bioinform.2
2011 DOSim: An R package for similarity between diseases based on Disease Ontology
abstract
BACKGROUND: The construction of the Disease Ontology (DO) has helped promote the investigation of diseases and disease risk factors. DO enables researchers to analyse disease similarity by adopting semantic similarity measures, and has expanded our understanding of the relationships between different diseases and to classify them. Simultaneously, similarities between genes can also be analysed by their associations with similar diseases. As a result, disease heterogeneity is better understood and insights into the molecular pathogenesis of similar diseases have been gained. However, bioinformatics tools that provide easy and straight forward ways to use DO to study disease and gene similarity simultaneously are required. RESULTS: We have developed an R-based software package (DOSim) to compute the similarity between diseases and to measure the similarity between human genes in terms of diseases. DOSim incorporates a DO-based enrichment analysis function that can be used to explore the disease feature of an independent gene set. A multilayered enrichment analysis (GO and KEGG annotation) annotation function that helps users explore the biological meaning implied in a newly detected gene module is also part of the DOSim package. We used the disease similarity application to demonstrate the relationship between 128 different DO cancer terms. The hierarchical clustering of these 128 different cancers showed modular characteristics. In another case study, we used the gene similarity application on 361 obesity-related genes. The results revealed the complex pathogenesis of obesity. In addition, the gene module detection and gene module multilayered annotation functions in DOSim when applied on these 361 obesity-related genes helped extend our understanding of the complex pathogenesis of obesity risk phenotypes and the heterogeneity of obesity-related diseases. CONCLUSIONS: DOSim can be used to detect disease-driven gene modules, and to annotate the modules for functions and pathways. The DOSim package can also be used to visualise DO structure. DOSim can reflect the modular characteristic of disease related genes and promote our understanding of the complex pathogenesis of diseases. DOSim is available on the Comprehensive R Archive Network (CRAN) or http://bioinfo.hrbmu.edu.cn/dosim.
Binsheng Gong, Tao Liu 0031, Chunquan Li 0002, Shaoqi Rao, Xia Li 0004
BMC Bioinform.7