Fengcui Qian

dblp:236/0448 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0003-3817-4690ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021
YearPublicationVenuePosition
2025 FunlncModel: integrating multi-omic features from upstream and downstream regulatory networks into a machine learning framework to identify functional lncRNAs
abstract
Accumulating evidence indicates that long noncoding RNAs (lncRNAs) play important roles in molecular and cellular biology. Although many algorithms have been developed to reveal their associations with complex diseases by using downstream targets, the upstream (epi)genetic regulatory information has not been sufficiently leveraged to predict the function of lncRNAs in various biological processes. Therefore, we present FunlncModel, a machine learning-based interpretable computational framework, which aims to screen out functional lncRNAs by integrating a large number of (epi)genetic features and functional genomic features from their upstream/downstream multi-omic regulatory networks. We adopted the random forest method to mine nearly 60 features in three categories from >2000 datasets across 11 data types, including transcription factors (TFs), histone modifications, typical enhancers, super-enhancers, methylation sites, and mRNAs. FunlncModel outperformed alternative methods for classification performance in human embryonic stem cell (hESC) (0.95 Area Under Curve (AUROC) and 0.97 Area Under the Precision-Recall Curve (AUPRC)). It could not only infer the most known lncRNAs that influence the states of stem cells, but also discover novel high-confidence functional lncRNAs. We extensively validated FunlncModel's efficacy by up to 27 cancer-related functional prediction tasks, which involved multiple cancer cell growth processes and cancer hallmarks. Meanwhile, we have also found that (epi)genetic regulatory features, such as TFs and histone modifications, serve as strong predictors for revealing the function of lncRNAs. Overall, FunlncModel is a strong and stable prediction model for identifying functional lncRNAs in specific cellular contexts. FunlncModel is available as a web server at https://bio.liclab.net/FunlncModel/.
Yan-Yu Li, Fengcui Qian, Guorui Zhang, Xue-Cang Li, Li-Wei Zhou, Zhengmin Yu, Wei Liu 0187, Qiu-Yu Wang, Chunquan Li 0002
Briefings Bioinform.2
2024 PEGCN: A Single-Cell Type Annotation model based on GCN with Pseudo Labels and Ensemble Learning
abstract
Single-cell type annotation helps to identify specific subpopulations of cells associated with a particular disease, thus providing therapeutic targets for precision medicine. Previous tasks on single-cell type annotation have been performed on the basis of RNA-seq sequencing data, and gradually some studies are now being conducted on scATAC-seq sequencing data as well. Although the researchs in this area has achieved considerable results, it is still constrained by the high dimensionality, high noise, and uneven sample distribution inherent in scATAC-seq datasets. In this paper, we propose a graph convolution model that incorporates ensemble learning. The utilization of Pseudo Labels for cell graph construction significantly mitigates the excessively high anisotropy index in the graph caused by noise, presenting a novel approach to enhancing graph quality. Furthermore, our method introduces a second innovation by effectively addressing the challenge of high-dimensional features in ATAC data through guided feature selection facilitated by ensemble learning, ultimately improving the overall performance and accuracy of the analysis. The model effectively reduces the impact of the previously mentioned problems on the scATAC-seq dataset, and improves the predictive performance of the model. The single-cell type annotation results exceed 90% of the ACC in different datasets and the Macro-average ACC exceeds that of other models.
Fengcui Qian, Chunquan Li 0002, Chunping Ouyang
BIBM2
2023 CanMethdb: a database for genome-wide DNA methylation annotation in cancers
abstract
MOTIVATION: DNA methylation within gene body and promoters in cancer cells is well documented. An increasing number of studies showed that cytosine-phosphate-guanine (CpG) sites falling within other regulatory elements could also regulate target gene activation, mainly by affecting transcription factors (TFs) binding in human cancers. This led to the urgent need for comprehensively and effectively collecting distinct cis-regulatory elements and TF-binding sites (TFBS) to annotate DNA methylation regulation. RESULTS: We developed a database (CanMethdb, http://meth.liclab.net/CanMethdb/) that focused on the upstream and downstream annotations for CpG-genes in cancers. This included upstream cis-regulatory elements, especially those involving distal regions to genes, and TFBS annotations for the CpGs and downstream functional annotations for the target genes, computed through integrating abundant DNA methylation and gene expression profiles in diverse cancers. Users could inquire CpG-target gene pairs for a cancer type through inputting a genomic region, a CpG, a gene name, or select hypo/hypermethylated CpG sets. The current version of CanMethdb documented a total of 38 986 060 CpG-target gene pairs (with 6 769 130 unique pairs), involving 385 217 CpGs and 18 044 target genes, abundant cis-regulatory elements and TFs for 33 TCGA cancer types. CanMethdb might help biologists perform in-depth studies of target gene regulations based on DNA methylations in cancer. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://github.com/chunquanlipathway/CanMethdb. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jianmei Zhao, Fengcui Qian, Xuecang Li, Zhengmin Yu, Yanyu Li, Yongsan Yang, Qi Pan, Qiuyu Wang, Jian Zhang 0084, Guohua Wang 0001, Chunquan Li 0002
Bioinform.2
2022 GREAP: a comprehensive enrichment analysis software for human genomic regions
abstract
The rapid development of genomic high-throughput sequencing has identified a large number of DNA regulatory elements with abundant epigenetics markers, which promotes the rapid accumulation of functional genomic region data. The comprehensively understanding and research of human functional genomic regions is still a relatively urgent work at present. However, the existing analysis tools lack extensive annotation and enrichment analytical abilities for these regions. Here, we designed a novel software, Genomic Region sets Enrichment Analysis Platform (GREAP), which provides comprehensive region annotation and enrichment analysis capabilities. Currently, GREAP supports 85 370 genomic region reference sets, which cover 634 681 107 regions across 11 different data types, including super enhancers, transcription factors, accessible chromatins, etc. GREAP provides widespread annotation and enrichment analysis of genomic regions. To reflect the significance of enrichment analysis, we used the hypergeometric test and also provided a Locus Overlap Analysis. In summary, GREAP is a powerful platform that provides many types of genomic region sets for users and supports genomic region annotations and enrichment analyses. In addition, we developed a customizable genome browser containing >400 000 000 customizable tracks for visualization. The platform is freely available at http://www.liclab.net/Greap/view/index.
Yongsan Yang, Fengcui Qian, Xuecang Li, Yanyu Li, Liwei Zhou, Qiuyu Wang, Xinyuan Zhou, Jian Zhang 0084, Zhengmin Yu, Ting Cui, Chenchen Feng, Desi Shang, Mengfei Sun, Yuexin Zhang, Huifang Tang, Chunquan Li 0002
Briefings Bioinform.2
2021 TRlnc: a comprehensive database for human transcriptional regulatory information of lncRNAs
abstract
Long noncoding RNAs (lncRNAs) have been proven to play important roles in transcriptional processes and biological functions. With the increasing study of human diseases and biological processes, information in human H3K27ac ChIP-seq, ATAC-seq and DNase-seq datasets is accumulating rapidly, resulting in an urgent need to collect and process data to identify transcriptional regulatory regions of lncRNAs. We therefore developed a comprehensive database for human regulatory information of lncRNAs (TRlnc, http://bio.licpathway.net/TRlnc), which aimed to collect available resources of transcriptional regulatory regions of lncRNAs and to annotate and illustrate their potential roles in the regulation of lncRNAs in a cell type-specific manner. The current version of TRlnc contains 8 683 028 typical enhancers/super-enhancers and 32 348 244 chromatin accessibility regions associated with 91 906 human lncRNAs. These regions are identified from over 900 human H3K27ac ChIP-seq, ATAC-seq and DNase-seq samples. Furthermore, TRlnc provides the detailed genetic and epigenetic annotation information within transcriptional regulatory regions (promoter, enhancer/super-enhancer and chromatin accessibility regions) of lncRNAs, including common SNPs, risk SNPs, eQTLs, linkage disequilibrium SNPs, transcription factors, methylation sites, histone modifications and 3D chromatin interactions. It is anticipated that the use of TRlnc will help users to gain in-depth and useful insights into the transcriptional regulatory mechanisms of lncRNAs.
Yanyu Li, Xuecang Li, Yongsan Yang, Fengcui Qian, Zhidong Tang, Jianmei Zhao, Jian Zhang 0084, Xuefeng Bai 0002, Yong Jiang 0004, Jianyuan Zhou, Yuexin Zhang, Liwei Zhou, Jianjun Xie, Enmin Li, Qiuyu Wang, Chunquan Li 0002
Briefings Bioinform.5
2020 HiFreSP: A novel high-frequency sub-pathway mining approach to identify robust prognostic gene signatures
abstract
With the increasing awareness of heterogeneity in cancers, better prediction of cancer prognosis is much needed for more personalized treatment. Recently, extensive efforts have been made to explore the variations in gene expression for better prognosis. However, the prognostic gene signatures predicted by most existing methods have little robustness among different datasets of the same cancer. To improve the robustness of the gene signatures, we propose a novel high-frequency sub-pathways mining approach (HiFreSP), integrating a randomization strategy with gene interaction pathways. We identified a six-gene signature (CCND1, CSF3R, E2F2, JUP, RARA and TCF7) in esophageal squamous cell carcinoma (ESCC) by HiFreSP. This signature displayed a strong ability to predict the clinical outcome of ESCC patients in two independent datasets (log-rank test, P = 0.0045 and 0.0087). To further show the predictive performance of HiFreSP, we applied it to two other cancers: pancreatic adenocarcinoma and breast cancer. The identified signatures show high predictive power in all testing datasets of the two cancers. Furthermore, compared with the two popular prognosis signature predicting methods, the least absolute shrinkage and selection operator penalized Cox proportional hazards model and the random survival forest, HiFreSP showed better predictive accuracy and generalization across all testing datasets of the above three cancers. Lastly, we applied HiFreSP to 8137 patients involving 20 cancer types in the TCGA database and found high-frequency prognosis-associated pathways in many cancers. Taken together, HiFreSP shows higher prognostic capability and greater robustness, and the identified signatures provide clinical guidance for cancer prognosis. HiFreSP is freely available via GitHub: https://github.com/chunquanlipathway/HiFreSP.
Jianmei Zhao, Xuecang Li, Chenchen Feng, Fengcui Qian, Yuejuan Liu, Jian Zhang 0084, Bo Ai 0005, Ziyu Ning, Wei Liu 0187, Xuefeng Bai 0002, Zhiyong Wu 0009, Xiue Xu, Zhidong Tang, Qi Pan, Liyan Xu, Chunquan Li 0002, Qiuyu Wang, Enmin Li
Briefings Bioinform.6
2019 TRCirc: a resource for transcriptional regulation information of circRNAs
abstract
In recent years, high-throughput genomic technologies like chromatin immunoprecipitation sequencing (ChIp-seq) and transcriptome sequencing (RNA-seq) have been becoming both more refined and less expensive, making them more accessible. Many circular RNAs (circRNAs) that originate from back-spliced exons have been identified in various cell lines across different species. However, the regulatory mechanism for transcription of circRNAs remains unclear. Therefore, there is an urgent need to construct a database detailing the transcriptional regulation of circRNAs. TRCirc (http://www.licpathway.net/TRCirc) provides a resource for efficient retrieval, browsing and visualization of transcriptional regulation information of circRNAs. The current version of TRCirc documents 92 375 circRNAs and 161 transcription factors (TFs) from more than 100 cell types and together represent more than 765 000 TF-circRNA regulatory relationships. Furthermore, TRCirc provides other regulatory information about transcription of circRNAs, including their expression, methylation levels, H3K27ac signals in regulation regions and super-enhancers associated with circRNAs. TRCirc provides a convenient, user-friendly interface to search, browse and visualize detailed information about these circRNAs.
Zhidong Tang, Xuecang Li, Jianmei Zhao, Fengcui Qian, Chenchen Feng, Yanyu Li, Jian Zhang 0084, Yong Jiang 0004, Yongsan Yang, Qiuyu Wang, Chunquan Li 0002
Briefings Bioinform.4