VLDB 2026 Research / reviewers in the wild / expert
Pengfei Liang 0002
dblp:217/0325-2
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2023
0000-0001-9541-9485ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A computational framework of routine test data for the cost-effective chronic disease predictionabstractChronic diseases, because of insidious onset and long latent period, have become the major global disease burden. However, the current chronic disease diagnosis methods based on genetic markers or imaging analysis are challenging to promote completely due to high costs and cannot reach universality and popularization. This study analyzed massive data from routine blood and biochemical test of 32 448 patients and developed a novel framework for cost-effective chronic disease prediction with high accuracy (AUC 87.32%). Based on the best-performing XGBoost algorithm, 20 classification models were further constructed for 17 types of chronic diseases, including 9 types of cancers, 5 types of cardiovascular diseases and 3 types of mental illness. The highest accuracy of the model was 90.13% for cardia cancer, and the lowest was 76.38% for rectal cancer. The model interpretation with the SHAP algorithm showed that CREA, R-CV, GLU and NEUT% might be important indices to identify the most chronic diseases. PDW and R-CV are also discovered to be crucial indices in classifying the three types of chronic diseases (cardiovascular disease, cancer and mental illness). In addition, R-CV has a higher specificity for cancer, ALP for cardiovascular disease and GLU for mental illness. The association between chronic diseases was further revealed. At last, we build a user-friendly explainable machine-learning-based clinical decision support system (DisPioneer: http://bioinfor.imu.edu.cn/dispioneer) to assist in predicting, classifying and treating chronic diseases. This cost-effective work with simple blood tests will benefit more people and motivate clinical implementation and further investigation of chronic diseases prevention and surveillance program. Mingzhu Liu, Qilemuge Xi, Yuchao Liang, Haicheng Li, Pengfei Liang 0002, Temuqile Temuqile, Yongchun Zuo |
Briefings Bioinform. | 6 |
| 2022 | iProbiotics: a machine learning platform for rapid identification of probiotic properties from whole-genome primary sequencesabstractLactic acid bacteria consortia are commonly present in food, and some of these bacteria possess probiotic properties. However, discovery and experimental validation of probiotics require extensive time and effort. Therefore, it is of great interest to develop effective screening methods for identifying probiotics. Advances in sequencing technology have generated massive genomic data, enabling us to create a machine learning-based platform for such purpose in this work. This study first selected a comprehensive probiotics genome dataset from the probiotic database (PROBIO) and literature surveys. Then, k-mer (from 2 to 8) compositional analysis was performed, revealing diverse oligonucleotide composition in strain genomes and apparently more probiotic (P-) features in probiotic genomes than non-probiotic genomes. To reduce noise and improve computational efficiency, 87 376 k-mers were refined by an incremental feature selection (IFS) method, and the model achieved the maximum accuracy level at 184 core features, with a high prediction accuracy (97.77%) and area under the curve (98.00%). Functional genomic analysis using annotations from gene ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG) and Rapid Annotation using Subsystem Technology (RAST) databases, as well as analysis of genes associated with host gastrointestinal survival/settlement, carbohydrate utilization, drug resistance and virulence factors, revealed that the distribution of P-features was biased toward genes/pathways related to probiotic function. Our results suggest that the role of probiotics is not determined by a single gene, but by a combination of k-mer genomic components, providing new insights into the identification and underlying mechanisms of probiotics. This work created a novel and free online bioinformatic tool, iProbiotics, which would facilitate rapid screening for probiotics. Haicheng Li, Lei Zheng 0009, Jinzhao Li, Pengfei Liang 0002, Lai-Yu Kwok, Yongchun Zuo, Wenyi Zhang 0004, Heping Zhang |
Briefings Bioinform. | 6 |
| 2021 | Dppa2/4 as a trigger of signaling pathways to promote zygote genome activation by binding to CG-rich regionabstractDevelopmental pluripotency-associated 2 (Dppa2) and developmental pluripotency-associated 4 (Dppa4) as positive drivers were helpful for transcriptional regulation of zygotic genome activation (ZGA). Here, we systematically assessed the cooperative interplay of Dppa2 and Dppa4 in regulating cell pluripotency and found that simultaneous overexpression of Dppa2/4 can make induced pluripotent stem cells closer to embryonic stem cells (ESCs). Compared with other pluripotency transcription factors, Dppa2/4 can regulate majorities of signaling pathways by binding on CG-rich region of proximal promoter (0-500 bp), of which 85% and 77% signaling pathways were significantly activated by Dppa2 and Dppa4, respectively. Notably, Dppa2/4 also can dramatically trigger the decisive signaling pathways for facilitating ZGA, including Hippo, MAPK and TGF-beta signaling pathways and so on. At last, we found alkaline phosphatase, placental-like 2 (Alppl2) was completely silenced when Dppa2 and 4 single- or double-knockout in ESC, which is consistent with Dux. Moreover, Alppl2 was significantly activated in mouse 2-cell embryos and 4-8 cells stage of human embryos, further predicted that Alppl2 was directly regulated by Dppa2/4 as a ZGA candidate driver to facilitate pre-embryonic development. Hanshuang Li, Chunshen Long, Jinzhu Xiang, Pengfei Liang 0002, Xueling Li, Yongchun Zuo |
Briefings Bioinform. | 4 |
| 2021 | HelPredictor models single-cell transcriptome to predict human embryo lineage allocationabstractThe in-depth understanding of cellular fate decision of human preimplantation embryos has prompted investigations on how changes in lineage allocation, which is far from trivial and remains a time-consuming task by experimental methods. It is desirable to develop a novel effective bioinformatics strategy to consider transitions of coordinated embryo lineage allocation and stage-specific patterns. There are rapidly growing applications of machine learning models to interpret complex datasets for identifying candidate development-related factors and lineage-determining molecular events. Here we developed the first machine learning platform, HelPredictor, that integrates three feature selection methods, namely, principal components analysis, F-score algorithm and squared coefficient of variation, and four classical machine learning classifiers that different combinations of methods and classifiers have independent outputs by increment feature selection method. With application to single-cell sequencing data of human embryo, HelPredictor not only achieved 94.9% and 90.9% respectively with cross-validation and independent test, but also fast classified different embryonic lineages and their development trajectories using less HelPredictor-predicted factors. The above-mentioned candidate lineage-specific genes were discussed in detail and were clustered for exploring transitions of embryonic heterogeneity. Our tool can fast and efficiently reveal potential lineage-specific and stage-specific biomarkers and provide insights into how advanced computational tools contribute to development research. The source code is available at https://github.com/liameihao/HelPredictor. Pengfei Liang 0002, Lei Zheng 0009, Chunshen Long, Wuritu Yang, Yongchun Zuo |
Briefings Bioinform. | 1 |
| 2021 | eHSCPr discriminating the cell identity involved in endothelial to hematopoietic transitionabstractMOTIVATION: Hematopoietic stem cells (HSCs) give rise to all blood cells and play a vital role throughout the whole lifespan through their pluripotency and self-renewal properties. Accurately identifying the stages of early HSCs is extremely important, as it may open up new prospects for extracorporeal blood research. Existing experimental techniques for identifying the early stages of HSCs development are time-consuming and expensive. Machine learning has shown its excellence in massive single-cell data processing and it is desirable to develop related computational models as good complements to experimental techniques. RESULTS: In this study, we presented a novel predictor called eHSCPr specifically for predicting the early stages of HSCs development. To reveal the distinct genes at each developmental stage of HSCs, we compared F-score with three state-of-art differential gene selection methods (limma, DESeq2, edgeR) and evaluated their performance. F-score captured the more critical surface markers of endothelial cells and hematopoietic cells, and the area under receiver operating characteristic curve (ROC) value was 0.987. Based on SVM, the 10-fold cross-validation accuracy of eHSCpr in the independent dataset and the training dataset reached 94.84% and 94.19%, respectively. Importantly, we performed transcription analysis on the F-score gene set, which indeed further enriched the signal markers of HSCs development stages. eHSCPr can be a powerful tool for predicting early stages of HSCs development, facilitating hypothesis-driven experimental design and providing crucial clues for the in vitro blood regeneration studies. AVAILABILITY AND IMPLEMENTATION: http://bioinfor.imu.edu.cn/ehscpr. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hao Wang 0109, Pengfei Liang 0002, Lei Zheng 0009, Chunshen Long, Hanshuang Li, Yongchun Zuo |
Bioinform. | 2 |