EDBT 2026 Demo / reviewers in the wild / expert
Yongchun Zuo
dblp:43/8348
· DBLP profile ↗
16ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-6065-7835ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Refining cell subpopulation identification and cell fate mapping using cell energy
Chunshen Long, Hanshuang Li, Qilemuge Xi, Yongchun Zuo |
Expert Syst. Appl. | 6 |
| 2024 | SpaNCMG: improving spatial domains identification of spatial transcriptomics using neighborhood-complementary mixed-view graph convolutional networkabstractThe advancement of spatial transcriptomics (ST) technology contributes to a more profound comprehension of the spatial properties of gene expression within tissues. However, due to challenges of high dimensionality, pronounced noise and dynamic limitations in ST data, the integration of gene expression and spatial information to accurately identify spatial domains remains challenging. This paper proposes a SpaNCMG algorithm for the purpose of achieving precise spatial domain description and localization based on a neighborhood-complementary mixed-view graph convolutional network. The algorithm enables better adaptation to ST data at different resolutions by integrating the local information from KNN and the global structure from r-radius into a complementary neighborhood graph. It also introduces an attention mechanism to achieve adaptive fusion of different reconstructed expressions, and utilizes KPCA method for dimensionality reduction. The application of SpaNCMG on five datasets from four sequencing platforms demonstrates superior performance to eight existing advanced methods. Specifically, the algorithm achieved highest ARI accuracies of 0.63 and 0.52 on the datasets of the human dorsolateral prefrontal cortex and mouse somatosensory cortex, respectively. It accurately identified the spatial locations of marker genes in the mouse olfactory bulb tissue and inferred the biological functions of different regions. When handling larger datasets such as mouse embryos, the SpaNCMG not only identified the main tissue structures but also explored unlabeled domains. Overall, the good generalization ability and scalability of SpaNCMG make it an outstanding tool for understanding tissue structure and disease mechanisms. Our codes are available at https://github.com/ZhihaoSi/SpaNCMG. Zhihao Si, Hanshuang Li, Wenjing Shang, Lingjiao Kong, Chunshen Long, Yongchun Zuo, Zhenxing Feng |
Briefings Bioinform. | 7 |
| 2024 | Integrating somatic mutation profiles with structural deep clustering network for metabolic stratification in pancreatic cancer: a comprehensive analysis of prognostic and genomic landscapesabstractPancreatic cancer is a globally recognized highly aggressive malignancy, posing a significant threat to human health and characterized by pronounced heterogeneity. In recent years, researchers have uncovered that the development and progression of cancer are often attributed to the accumulation of somatic mutations within cells. However, cancer somatic mutation data exhibit characteristics such as high dimensionality and sparsity, which pose new challenges in utilizing these data effectively. In this study, we propagated the discrete somatic mutation data of pancreatic cancer through a network propagation model based on protein-protein interaction networks. This resulted in smoothed somatic mutation profile data that incorporate protein network information. Based on this smoothed mutation profile data, we obtained the activity levels of different metabolic pathways in pancreatic cancer patients. Subsequently, using the activity levels of various metabolic pathways in cancer patients, we employed a deep clustering algorithm to establish biologically and clinically relevant metabolic subtypes of pancreatic cancer. Our study holds scientific significance in classifying pancreatic cancer based on somatic mutation data and may provide a crucial theoretical basis for the diagnosis and immunotherapy of pancreatic cancer patients. Honghao Li, Dongqing Su, Yuqiang Xiong, Haodong Wei, Hongmei Sun, Qilemuge Xi, Yongchun Zuo |
Briefings Bioinform. | 10 |
| 2023 | A computational framework of routine test data for the cost-effective chronic disease predictionabstractChronic diseases, because of insidious onset and long latent period, have become the major global disease burden. However, the current chronic disease diagnosis methods based on genetic markers or imaging analysis are challenging to promote completely due to high costs and cannot reach universality and popularization. This study analyzed massive data from routine blood and biochemical test of 32 448 patients and developed a novel framework for cost-effective chronic disease prediction with high accuracy (AUC 87.32%). Based on the best-performing XGBoost algorithm, 20 classification models were further constructed for 17 types of chronic diseases, including 9 types of cancers, 5 types of cardiovascular diseases and 3 types of mental illness. The highest accuracy of the model was 90.13% for cardia cancer, and the lowest was 76.38% for rectal cancer. The model interpretation with the SHAP algorithm showed that CREA, R-CV, GLU and NEUT% might be important indices to identify the most chronic diseases. PDW and R-CV are also discovered to be crucial indices in classifying the three types of chronic diseases (cardiovascular disease, cancer and mental illness). In addition, R-CV has a higher specificity for cancer, ALP for cardiovascular disease and GLU for mental illness. The association between chronic diseases was further revealed. At last, we build a user-friendly explainable machine-learning-based clinical decision support system (DisPioneer: http://bioinfor.imu.edu.cn/dispioneer) to assist in predicting, classifying and treating chronic diseases. This cost-effective work with simple blood tests will benefit more people and motivate clinical implementation and further investigation of chronic diseases prevention and surveillance program. Mingzhu Liu, Qilemuge Xi, Yuchao Liang, Haicheng Li, Pengfei Liang 0002, Temuqile Temuqile, Yongchun Zuo |
Briefings Bioinform. | 11 |
| 2022 | iProbiotics: a machine learning platform for rapid identification of probiotic properties from whole-genome primary sequencesabstractLactic acid bacteria consortia are commonly present in food, and some of these bacteria possess probiotic properties. However, discovery and experimental validation of probiotics require extensive time and effort. Therefore, it is of great interest to develop effective screening methods for identifying probiotics. Advances in sequencing technology have generated massive genomic data, enabling us to create a machine learning-based platform for such purpose in this work. This study first selected a comprehensive probiotics genome dataset from the probiotic database (PROBIO) and literature surveys. Then, k-mer (from 2 to 8) compositional analysis was performed, revealing diverse oligonucleotide composition in strain genomes and apparently more probiotic (P-) features in probiotic genomes than non-probiotic genomes. To reduce noise and improve computational efficiency, 87 376 k-mers were refined by an incremental feature selection (IFS) method, and the model achieved the maximum accuracy level at 184 core features, with a high prediction accuracy (97.77%) and area under the curve (98.00%). Functional genomic analysis using annotations from gene ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG) and Rapid Annotation using Subsystem Technology (RAST) databases, as well as analysis of genes associated with host gastrointestinal survival/settlement, carbohydrate utilization, drug resistance and virulence factors, revealed that the distribution of P-features was biased toward genes/pathways related to probiotic function. Our results suggest that the role of probiotics is not determined by a single gene, but by a combination of k-mer genomic components, providing new insights into the identification and underlying mechanisms of probiotics. This work created a novel and free online bioinformatic tool, iProbiotics, which would facilitate rapid screening for probiotics. Haicheng Li, Lei Zheng 0009, Jinzhao Li, Pengfei Liang 0002, Lai-Yu Kwok, Yongchun Zuo, Wenyi Zhang 0004, Heping Zhang |
Briefings Bioinform. | 8 |
| 2021 | Dppa2/4 as a trigger of signaling pathways to promote zygote genome activation by binding to CG-rich regionabstractDevelopmental pluripotency-associated 2 (Dppa2) and developmental pluripotency-associated 4 (Dppa4) as positive drivers were helpful for transcriptional regulation of zygotic genome activation (ZGA). Here, we systematically assessed the cooperative interplay of Dppa2 and Dppa4 in regulating cell pluripotency and found that simultaneous overexpression of Dppa2/4 can make induced pluripotent stem cells closer to embryonic stem cells (ESCs). Compared with other pluripotency transcription factors, Dppa2/4 can regulate majorities of signaling pathways by binding on CG-rich region of proximal promoter (0-500 bp), of which 85% and 77% signaling pathways were significantly activated by Dppa2 and Dppa4, respectively. Notably, Dppa2/4 also can dramatically trigger the decisive signaling pathways for facilitating ZGA, including Hippo, MAPK and TGF-beta signaling pathways and so on. At last, we found alkaline phosphatase, placental-like 2 (Alppl2) was completely silenced when Dppa2 and 4 single- or double-knockout in ESC, which is consistent with Dux. Moreover, Alppl2 was significantly activated in mouse 2-cell embryos and 4-8 cells stage of human embryos, further predicted that Alppl2 was directly regulated by Dppa2/4 as a ZGA candidate driver to facilitate pre-embryonic development. Hanshuang Li, Chunshen Long, Jinzhu Xiang, Pengfei Liang 0002, Xueling Li, Yongchun Zuo |
Briefings Bioinform. | 6 |
| 2021 | HelPredictor models single-cell transcriptome to predict human embryo lineage allocationabstractThe in-depth understanding of cellular fate decision of human preimplantation embryos has prompted investigations on how changes in lineage allocation, which is far from trivial and remains a time-consuming task by experimental methods. It is desirable to develop a novel effective bioinformatics strategy to consider transitions of coordinated embryo lineage allocation and stage-specific patterns. There are rapidly growing applications of machine learning models to interpret complex datasets for identifying candidate development-related factors and lineage-determining molecular events. Here we developed the first machine learning platform, HelPredictor, that integrates three feature selection methods, namely, principal components analysis, F-score algorithm and squared coefficient of variation, and four classical machine learning classifiers that different combinations of methods and classifiers have independent outputs by increment feature selection method. With application to single-cell sequencing data of human embryo, HelPredictor not only achieved 94.9% and 90.9% respectively with cross-validation and independent test, but also fast classified different embryonic lineages and their development trajectories using less HelPredictor-predicted factors. The above-mentioned candidate lineage-specific genes were discussed in detail and were clustered for exploring transitions of embryonic heterogeneity. Our tool can fast and efficiently reveal potential lineage-specific and stage-specific biomarkers and provide insights into how advanced computational tools contribute to development research. The source code is available at https://github.com/liameihao/HelPredictor. Pengfei Liang 0002, Lei Zheng 0009, Chunshen Long, Wuritu Yang, Yongchun Zuo |
Briefings Bioinform. | 6 |
| 2021 | Modular arrangements of sequence motifs determine the functional diversity of KDM proteinsabstractHistone lysine demethylases (KDMs) play a vital role in regulating chromatin dynamics and transcription. KDM proteins are given modular activities by its sequence motifs with obvious roles division, which endow the complex and diverse functions. In our review, according to functional features, we classify sequence motifs into four classes: catalytic motifs, targeting motifs, regulatory motifs and potential motifs. JmjC, as the main catalytic motif, combines to Fe2+ and α-ketoglutarate by residues H-D/E-H and S-N-N/Y-K-N/Y-T/S. Targeting motifs make catalytic motifs recognize specific methylated lysines, such as PHD that helps KDM5 to demethylate H3K4me3. Regulatory motifs consist of a functional network. For example, NLS, Ser-rich, TPR and JmjN motifs regulate the nuclear localization. And interactions through the CW-type-C4H2C2-SWIRM are necessary to the demethylase activity of KDM1B. Additionally, many conservative domains that have potential functions but no deep exploration are reviewed for the first time. These conservative domains are usually amino acid-rich regions, which have great research value. The arrangements of four types of sequence motifs generate that KDM proteins diversify toward modular activities and biological functions. Finally, we draw a blueprint of functional mechanisms to discuss the modular activity of KDMs. Zerong Wang, Baofang Xu, Ruixia Tian, Yongchun Zuo |
Briefings Bioinform. | 5 |
| 2021 | Clinical significance and immunogenomic landscape analyses of the immune cell signature based prognostic model for patients with breast cancerabstractBreast cancer is one of the most common types of cancers and the leading cause of death from malignancy among women worldwide. Tumor-infiltrating lymphocytes are a source of important prognostic biomarkers for breast cancer patients. In this study, based on the tumor-infiltrating lymphocytes in the tumor immune microenvironment, a risk score prognostic model was developed in the training cohort for risk stratification and prognosis prediction in breast cancer patients. The prognostic value of this risk score prognostic model was also verified in the two testing cohorts and the TCGA pan cancer cohort. Nomograms were also established in the training and testing cohorts to validate the clinical use of this model. Relationships between the risk score, intrinsic molecular subtypes, immune checkpoints, tumor-infiltrating immune cell abundances and the response to chemotherapy and immunotherapy were also evaluated. Based on these results, we can conclude that this risk score model could serve as a robust prognostic biomarker, provide therapeutic benefits for the development of novel chemotherapy and immunotherapy, and may be helpful for clinical decision making in breast cancer patients. Yuqiang Xiong, Dongqing Su, Chunlu Yu, Yiyin Cao, Yi Pan 0007, Qianzi Lu, Yongchun Zuo |
Briefings Bioinform. | 9 |
| 2021 | Immune cell infiltration-based signature for prognosis and immunogenomic analysis in breast cancerabstractBreast cancer is one of the most human malignant diseases and the leading cause of cancer-related death in the world. However, the prognostic and therapeutic benefits of breast cancer patients cannot be predicted accurately by the current stratifying system. In this study, an immune-related prognostic score was established in 22 breast cancer cohorts with a total of 6415 samples. An extensive immunogenomic analysis was conducted to explore the relationships between immune score, prognostic significance, infiltrating immune cells, cancer genotypes and potential immune escape mechanisms. Our analysis revealed that this immune score was a promising biomarker for estimating overall survival in breast cancer. This immune score was associated with important immunophenotypic factors, such as immune escape and mutation load. Further analysis revealed that patients with high immune scores exhibited therapeutic benefits from chemotherapy and immunotherapy. Based on these results, we can conclude that this immune score may be a useful tool for overall survival prediction and treatment guidance for patients with breast cancer. Chunlu Yu, Yiyin Cao, Yongchun Zuo |
Briefings Bioinform. | 5 |
| 2021 | RaacLogo: a new sequence logo generator by using reduced amino acid clustersabstractSequence logos give a fast and concise display in visualizing consensus sequence. Protein exhibits greater complexity and diversity than DNA, which usually affects the graphical representation of the logo. Reduced amino acids perform powerful ability for simplifying complexity of sequence alignment, which motivated us to establish RaacLogo. As a new sequence logo generator by using reduced amino acid alphabets, RaacLogo can easily generate many different simplified logos tailored to users by selecting various reduced amino acid alphabets that consisted of more than 40 clustering algorithms. This current web server provides 74 types of reduced amino acid alphabet, which were manually extracted to generate 673 reduced amino acid clusters (RAACs) for dealing with protein alignment. A two-dimensional selector was proposed for easily selecting desired RAACs with underlying biology knowledge. It is anticipated that the RaacLogo web server will play more high-potential roles for protein sequence alignment, topological estimation and protein design experiments. RaacLogo is freely available at http://bioinfor.imu.edu.cn/raaclogo. Lei Zheng 0009, Wuritu Yang, Yongchun Zuo |
Briefings Bioinform. | 5 |
| 2021 | eHSCPr discriminating the cell identity involved in endothelial to hematopoietic transitionabstractMOTIVATION: Hematopoietic stem cells (HSCs) give rise to all blood cells and play a vital role throughout the whole lifespan through their pluripotency and self-renewal properties. Accurately identifying the stages of early HSCs is extremely important, as it may open up new prospects for extracorporeal blood research. Existing experimental techniques for identifying the early stages of HSCs development are time-consuming and expensive. Machine learning has shown its excellence in massive single-cell data processing and it is desirable to develop related computational models as good complements to experimental techniques. RESULTS: In this study, we presented a novel predictor called eHSCPr specifically for predicting the early stages of HSCs development. To reveal the distinct genes at each developmental stage of HSCs, we compared F-score with three state-of-art differential gene selection methods (limma, DESeq2, edgeR) and evaluated their performance. F-score captured the more critical surface markers of endothelial cells and hematopoietic cells, and the area under receiver operating characteristic curve (ROC) value was 0.987. Based on SVM, the 10-fold cross-validation accuracy of eHSCpr in the independent dataset and the training dataset reached 94.84% and 94.19%, respectively. Importantly, we performed transcription analysis on the F-score gene set, which indeed further enriched the signal markers of HSCs development stages. eHSCPr can be a powerful tool for predicting early stages of HSCs development, facilitating hypothesis-driven experimental design and providing crucial clues for the in vitro blood regeneration studies. AVAILABILITY AND IMPLEMENTATION: http://bioinfor.imu.edu.cn/ehscpr. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hao Wang 0109, Pengfei Liang 0002, Lei Zheng 0009, Chunshen Long, Hanshuang Li, Yongchun Zuo |
Bioinform. | 6 |
| 2019 | Function determinants of TET proteins: the arrangements of sequence motifs with specific codesabstractThe ten-eleven translocation (TET) proteins play a crucial role in promoting locus-specific reversal of DNA methylation, a type of chromatin modification. Considerable evidences have demonstrated that the sequence motifs with specific codes are important to determine the functions of domains and active sites. Here, we surveyed major studies and reviews regarding the multiple functions of the TET proteins and established the patterns of the motif arrangements that determine the functions of TET proteins. First, we summarized the functional sequence basis of TET proteins and identified the new functional motifs based on the phylogenetic relationship. Next, we described the sequence characteristics of the functional motifs in detail and provided an overview of the relationship between the sequence motifs and the functions of TET proteins, including known functions and potential functions. Finally, we highlighted that sequence motifs with diverse post-translational modifications perform unique functions, and different selection pressures lead to different arrangements of sequence motifs, resulting in different paralogs and isoforms. Guangpeng Li, Yongchun Zuo |
Briefings Bioinform. | 3 |
| 2017 | Dynamic and modular gene regulatory networks drive the development of gametogenesisabstractGametogenesis is a complex process, which includes mitosis and meiosis and results in the production of ovum and sperm. The development of gametogenesis is dynamic and needs many different genes to work synergistically, but it is lack of global perspective research about this process. In this study, we detected the dynamic process of gametogenesis from the perspective of systems biology based on protein-protein interaction networks (PPINs) and functional analysis. Results showed that gametogenesis genes have strong synergistic effects in PPINs within and between different phases during the development. Addition to the synergistic effects on molecular networks, gametogenesis genes showed functional consistency within and between different phases, which provides the further evidence about the dynamic process during the development of gametogenesis. At last, we detected and provided the core molecular modules of different phases about gametogenesis. The gametogenesis genes and related modules can be obtained from our Web site Gametogenesis Molecule Online (GMO, http://gametsonline.nwsuaflmz.com/index.php), which is freely accessible. GMO may be helpful for the reference and application of these genes and modules in the future identification of key genes about gametogenesis. Summary, this work provided a computational perspective and frame to the analysis of the gametogenesis dynamics and modularity in both human and mouse. Dongxue Che, Weiyang Bai, Leijie Li, Guiyou Liu, Liangcai Zhang, Yongchun Zuo, Shiheng Tao, Jinlian Hua, Mingzhi Liao |
Briefings Bioinform. | 7 |
| 2017 | PseKRAAC: a flexible web server for generating pseudo K-tuple reduced amino acids compositionabstractThe reduced amino acids perform powerful ability for both simplifying protein complexity and identifying functional conserved regions. However, dealing with different protein problems may need different kinds of cluster methods. Encouraged by the success of pseudo-amino acid composition algorithm, we developed a freely available web server, called PseKRAAC (the pseudo K-tuple reduced amino acids composition). By implementing reduced amino acid alphabets, the protein complexity can be significantly simplified, which leads to decrease chance of overfitting, lower computational handicap and reduce information redundancy. PseKRAAC delivers more capability for protein research by incorporating three crucial parameters that describes protein composition. Users can easily generate many different modes of PseKRAAC tailored to their needs by selecting various reduced amino acids alphabets and other characteristic parameters. It is anticipated that the PseKRAAC web server will become a very useful tool in computational proteomics and protein sequence analysis. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at http://bigdata.imu.edu.cn/psekraac CONTACTS: [email protected] or [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Yongchun Zuo, Yingli Chen, Guangpeng Li, Zhenhe Yan |
Bioinform. | 1 |
| 2013 | PreDNA: accurate prediction of DNA-binding sites in proteins by integrating sequence and geometric structure informationabstractMOTIVATION: Protein-DNA interactions often take part in various crucial processes, which are essential for cellular function. The identification of DNA-binding sites in proteins is important for understanding the molecular mechanisms of protein-DNA interaction. Thus, we have developed an improved method to predict DNA-binding sites by integrating structural alignment algorithm and support vector machine-based methods. RESULTS: Evaluated on a new non-redundant protein set with 224 chains, the method has 80.7% sensitivity and 82.9% specificity in the 5-fold cross-validation test. In addition, it predicts DNA-binding sites with 85.1% sensitivity and 85.3% specificity when tested on a dataset with 62 protein-DNA complexes. Compared with a recently published method, BindN+, our method predicts DNA-binding sites with a 7% better area under the receiver operating characteristic curve value when tested on the same dataset. Many important problems in cell biology require the dense non-linear interactions between functional modules be considered. Thus, our prediction method will be useful in detecting such complex interactions. Qianzhong Li, Shuai Liu 0002, Yongchun Zuo |
Bioinform. | 5 |