VLDB 2026 Research / reviewers in the wild / expert
Yuan Zhu 0005
dblp:35/5929-5
· DBLP profile ↗
14ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-1372-9135ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TPLN: Cell-Type-Specific Gene Regulatory Network Inference Based on Poisson Log-Normal ModelabstractThe rapid advancement of single-cell RNA sequencing (scRNA-seq) technology has enabled the inference of gene regulatory networks (GRNs) at cell-type resolution, providing unprecedented insights into cellular heterogeneity. However, most existing methods focus on inferring GRNs at the singlecell level and lack a systematic framework for modeling cell-type-specific regulatory architectures. Furthermore, the inherent count-based nature and pronounced sparsity of scRNA-seq data pose significant challenges for graph-based models in accurately reconstructing GRNs for individual cell types. To overcome these limitations, we introduce TPLN, a hierarchical modeling framework based on the Poisson Log-Normal (PLN) model, which jointly addresses the count distribution and sparsity characteristics of scRNA-seq data. TPLN unifies cell-type-specific GRN inference with topology-aware gene module identification within a single algorithmic framework, facilitating the discovery of key regulatory interactions and functionally coherent modules across diverse cell types. Extensive evaluations on both simulated and real-world datasets show that TPLN consistently outperforms state-of-the-art methods-such as Glasso, SparLLM, PLNet, and GENIE3-in terms of GRN inference accuracy. Additionally, TPLN demonstrates superior performance in module identification and downstream functional enrichment analysis. Wenfei Fu, Yuan Zhu 0005 |
BIBM | 4 |
| 2025 | D3Impute: Dropout-aware discrimination, distribution-aware modeling, and density-guide imputation for scRNA-seq dataabstractSingle-cell RNA sequencing (scRNA-seq) has revolutionized the study of cellular heterogeneity. A major challenge, however, lies in the prevalence of non-biological zeros-false measurements caused by technical limitations that mask a cell's true transcriptome. This fundamental issue of distinguishing these artifacts from true biological zeros, where a gene is genuinely absent, remains a key hurdle for computational methods, as misclassification can distort biological signals during data recovery. To overcome this, we introduce D3Impute, a discriminative imputation framework built on three key innovations: (1) a distribution-aware normalization step that adapts to dataset-specific characteristics while preserving meaningful biological variation; (2) a dual-network discriminator that uses bulk RNA-seq data as a biological reference to accurately identify non-biological zeros while retaining the true biological zeros; and (3) a density-guided imputation engine that recovers expression values while maintaining local cellular neighborhood structures. Through comprehensive benchmarking against 12 state-of-the-art methods across six diverse datasets, D3Impute demonstrates consistent and significant improvements in essential downstream analyses, including cell clustering, trajectory inference, and differential expression detection. Furthermore, we provide an extensive practical evaluation of D3Impute, demonstrating its robustness across varying data qualities and providing clear guidelines for optimal application. By offering a robust, biologically informed, and user-oriented solution, D3Impute not only enhances scRNA-seq data analysis but also offers a generalizable framework for handling zero-inflated data in computational biology. Linfeng Jiang, Yuan Zhu 0005 |
PLoS Comput. Biol. | 4 |
| 2025 | Preoperative Prediction of Microvascular Invasion in Hepatocellular Carcinoma From Multi-Sequence Magnetic Resonance Imaging Based on Deep Fusion Representation LearningabstractRecent studies have identified microvascular invasion (MVI) as the most vital independent biomarker associated with early tumor recurrence. With advancements in medical technology, several computational methods have been developed to predict preoperative MVI using diverse medical images. These existing methods rely on human experience, attribute selection or clinical trial testing, which is often time-consuming and labor-intensive. Leveraging the advantages of deep learning, this study presents a novel end-to-end algorithm for predicting MVI prior to surgery. We devised a series of data preprocessing strategies to fully extract multi-view features from the data while preserving peritumoral information. Notably, a new multi-branch deep fused feature algorithm based on ResNet (DFFResNet) is introduced, which combines Magnetic Resonance Images (MRI) from different sequences to enhance information complementarity and integration. We conducted prediction experiments on a dataset from the Radiology Department of the First Hospital of Lanzhou University, comprising 117 individuals and seven MRI sequences. The model was trained on 80% of the data using 10-fold cross-validation, and the remaining 20% were used for testing. This evaluation was processed in two cases: CROI, containing samples with a complete region of interest (ROI), and PROI, containing samples with a partial ROI region. The robustness results from repeated experiments at both image and patient levels demonstrate the superior performance and improved generalization of the proposed method compared to alternative models. Our approach yields highly competitive prediction results even when the ROI region outline is incomplete, offering a novel and effective multi-sequence fused strategy for predicting preoperative MVI. Haishu Ma, Lingzhi Sun, Shinan Wang, Lulu Lu, Yong He 0003, Yuan Zhu 0005 |
IEEE J. Biomed. Health Informatics | 8 |
| 2022 | scDEA: differential expression analysis in single-cell RNA-sequencing data via ensemble learningabstractThe identification of differentially expressed genes between different cell groups is a crucial step in analyzing single-cell RNA-sequencing (scRNA-seq) data. Even though various differential expression analysis methods for scRNA-seq data have been proposed based on different model assumptions and strategies recently, the differentially expressed genes identified by them are quite different from each other, and the performances of them depend on the underlying data structures. In this paper, we propose a new ensemble learning-based differential expression analysis method, scDEA, to produce a more stable and accurate result. scDEA integrates the P-values obtained from 12 individual differential expression analysis methods for each gene using a P-value combination method. Comprehensive experiments show that scDEA outperforms the state-of-the-art individual methods with different experimental settings and evaluation metrics. We expect that scDEA will serve a wide range of users, including biologists, bioinformaticians and data scientists, who need to detect differentially expressed genes in scRNA-seq data. Hui-Sheng Li, Le Ou-Yang, Yuan Zhu 0005, Hong Yan 0001, Xiao-Fei Zhang |
Briefings Bioinform. | 3 |
| 2021 | Differential network analysis by simultaneously considering changes in gene interactions and gene expressionabstractMOTIVATION: Differential network analysis is an important tool to investigate the rewiring of gene interactions under different conditions. Several computational methods have been developed to estimate differential networks from gene expression data, but most of them do not consider that gene network rewiring may be driven by the differential expression of individual genes. New differential network analysis methods that simultaneously take account of the changes in gene interactions and changes in expression levels are needed. RESULTS: : In this article, we propose a differential network analysis method that considers the differential expression of individual genes when identifying differential edges. First, two hypothesis test statistics are used to quantify changes in partial correlations between gene pairs and changes in expression levels for individual genes. Then, an optimization framework is proposed to combine the two test statistics so that the resulting differential network has a hierarchical property, where a differential edge can be considered only if at least one of the two involved genes is differentially expressed. Simulation results indicate that our method outperforms current state-of-the-art methods. We apply our method to identify the differential networks between the luminal A and basal-like subtypes of breast cancer and those between acute myeloid leukemia and normal samples. Hub nodes in the differential networks estimated by our method, including both differentially and nondifferentially expressed genes, have important biological functions. AVAILABILITY AND IMPLEMENTATION: All the datasets underlying this article are publicly available. Processed data and source code can be accessed through the Github repository at https://github.com/Zhangxf-ccnu/chNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jia-Juan Tu, Le Ou-Yang, Yuan Zhu 0005, Hong Yan 0001, Hong Qin 0008, Xiao-Fei Zhang |
Bioinform. | 3 |
| 2019 | KSLIC: K-mediods Clustering Based Simple Linear Iterative Clustering
Houwang Zhang, Yuan Zhu 0005 |
PRCV (2) | 2 |
| 2019 | A unified formulation of a class of graph matching techniques
Yuan Zhu 0005, Jiufeng Zhou, Hong Yan 0001 |
Pattern Recognit. | 1 |
| 2017 | Inference of cellular level signaling networks using single-cell gene expression data in Caenorhabditis elegans reveals mechanisms of cell fate specificationabstractMOTIVATION: Cell fate specification plays a key role to generate distinct cell types during metazoan development. However, most of the underlying signaling networks at cellular level are not well understood. Availability of time lapse single-cell gene expression data collected throughout Caenorhabditis elegans embryogenesis provides an excellent opportunity for investigating signaling networks underlying cell fate specification at systems, cellular and molecular levels. RESULTS: We propose a framework to infer signaling networks at cellular level by exploring the single-cell gene expression data. Through analyzing the expression data of nhr-25 , a hypodermis-specific transcription factor, in every cells of both wild-type and mutant C.elegans embryos through RNAi against 55 genes, we have inferred a total of 23 genes that regulate (activate or inhibit) nhr-25 expression in cell-specific fashion. We also infer the signaling pathways consisting of each of these genes and nhr-25 based on a probabilistic graphical model for the selected five founder cells, 'ABarp', 'ABpla', 'ABpra', 'Caa' and 'Cpa', which express nhr-25 and mostly develop into hypodermis. By integrating the inferred pathways, we reconstruct five signaling networks with one each for the five founder cells. Using RNAi gene knockdown as a validation method, the inferred networks are able to predict the effects of the knockdown genes. These signaling networks in the five founder cells are likely to ensure faithful hypodermis cell fate specification in C.elegans at cellular level. AVAILABILITY AND IMPLEMENTATION: All source codes and data are available at the github repository https://github.com/xthuang226/Worm_Single_Cell_Data_and_Codes.git . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuan Zhu 0005, Leanne Lai Chan, Zhongying Zhao 0002, Hong Yan 0001 |
Bioinform. | 2 |
| 2016 | Protein complex detection based on partially shared multi-view clusteringabstractBACKGROUND: Protein complexes are the key molecular entities to perform many essential biological functions. In recent years, high-throughput experimental techniques have generated a large amount of protein interaction data. As a consequence, computational analysis of such data for protein complex detection has received increased attention in the literature. However, most existing works focus on predicting protein complexes from a single type of data, either physical interaction data or co-complex interaction data. These two types of data provide compatible and complementary information, so it is necessary to integrate them to discover the underlying structures and obtain better performance in complex detection. RESULTS: In this study, we propose a novel multi-view clustering algorithm, called the Partially Shared Multi-View Clustering model (PSMVC), to carry out such an integrated analysis. Unlike traditional multi-view learning algorithms that focus on mining either consistent or complementary information embedded in the multi-view data, PSMVC can jointly explore the shared and specific information inherent in different views. In our experiments, we compare the complexes detected by PSMVC from single data source with those detected from multiple data sources. We observe that jointly analyzing multi-view data benefits the detection of protein complexes. Furthermore, extensive experiment results demonstrate that PSMVC performs much better than 16 state-of-the-art complex detection techniques, including ensemble clustering and data integration techniques. CONCLUSIONS: In this work, we demonstrate that when integrating multiple data sources, using partially shared multi-view clustering model can help to identify protein complexes which are not readily identifiable by conventional single-view-based methods and other integrative analysis methods. All the results and source codes are available on https://github.com/Oyl-CityU/PSMVC . Le Ou-Yang, Xiao-Fei Zhang, Dao-Qing Dai, Meng-Yun Wu, Yuan Zhu 0005, Hong Yan 0001 |
BMC Bioinform. | 5 |
| 2016 | Regularized logistic regression with network-based pairwise interaction for biomarker identification in breast cancerabstractBACKGROUND: To facilitate advances in personalized medicine, it is important to detect predictive, stable and interpretable biomarkers related with different clinical characteristics. These clinical characteristics may be heterogeneous with respect to underlying interactions between genes. Usually, traditional methods just focus on detection of differentially expressed genes without taking the interactions between genes into account. Moreover, due to the typical low reproducibility of the selected biomarkers, it is difficult to give a clear biological interpretation for a specific disease. Therefore, it is necessary to design a robust biomarker identification method that can predict disease-associated interactions with high reproducibility. RESULTS: In this article, we propose a regularized logistic regression model. Different from previous methods which focus on individual genes or modules, our model takes gene pairs, which are connected in a protein-protein interaction network, into account. A line graph is constructed to represent the adjacencies between pairwise interactions. Based on this line graph, we incorporate the degree information in the model via an adaptive elastic net, which makes our model less dependent on the expression data. Experimental results on six publicly available breast cancer datasets show that our method can not only achieve competitive performance in classification, but also retain great stability in variable selection. Therefore, our model is able to identify the diagnostic and prognostic biomarkers in a more robust way. Moreover, most of the biomarkers discovered by our model have been verified in biochemical or biomedical researches. CONCLUSIONS: The proposed method shows promise in the diagnosis of disease pathogenesis with different clinical characteristics. These advances lead to more accurate and stable biomarker discovery, which can monitor the functional changes that are perturbed by diseases. Based on these predictions, researchers may be able to provide suggestions for new therapeutic approaches. Meng-Yun Wu, Xiao-Fei Zhang, Dao-Qing Dai, Le Ou-Yang, Yuan Zhu 0005, Hong Yan 0001 |
BMC Bioinform. | 5 |
| 2016 | Comparative analysis of housekeeping and tissue-specific driver nodes in human protein interaction networksabstractBACKGROUND: Several recent studies have used the Minimum Dominating Set (MDS) model to identify driver nodes, which provide the control of the underlying networks, in protein interaction networks. There may exist multiple MDS configurations in a given network, thus it is difficult to determine which one represents the real set of driver nodes. Because these previous studies only focus on static networks and ignore the contextual information on particular tissues, their findings could be insufficient or even be misleading. RESULTS: In this study, we develop a Collective-Influence-corrected Minimum Dominating Set (CI-MDS) model which takes into account the collective influence of proteins. By integrating molecular expression profiles and static protein interactions, 16 tissue-specific networks are established as well. We then apply the CI-MDS model to each tissue-specific network to detect MDS proteins. It generates almost the same MDSs when it is solved using different optimization algorithms. In addition, we classify MDS proteins into Tissue-Specific MDS (TS-MDS) proteins and HouseKeeping MDS (HK-MDS) proteins based on the number of tissues in which they are expressed and identified as MDS proteins. Notably, we find that TS-MDS proteins and HK-MDS proteins have significantly different topological and functional properties. HK-MDS proteins are more central in protein interaction networks, associated with more functions, evolving more slowly and subjected to a greater number of post-translational modifications than TS-MDS proteins. Unlike TS-MDS proteins, HK-MDS proteins significantly correspond to essential genes, ageing genes, virus-targeted proteins, transcription factors and protein kinases. Moreover, we find that besides HK-MDS proteins, many TS-MDS proteins are also linked to disease related genes, suggesting the tissue specificity of human diseases. Furthermore, functional enrichment analysis reveals that HK-MDS proteins carry out universally necessary biological processes and TS-MDS proteins usually involve in tissue-dependent functions. CONCLUSIONS: Our study uncovers key features of TS-MDS proteins and HK-MDS proteins, and is a step forward towards a better understanding of the controllability of human interactomes. Xiao-Fei Zhang, Le Ou-Yang, Dao-Qing Dai, Meng-Yun Wu, Yuan Zhu 0005, Hong Yan 0001 |
BMC Bioinform. | 5 |
| 2015 | Local Topology Preserved Tensor Models for Graph MatchingabstractThis paper proposes local topology preserved features in graph matching problem based on tensor technique. Many tensor based works paid much attention on catching many invariant feature tuples while local information for every single point to improve matching performance is also important. Here our proposed Local Topology Preserved Tensor (LTPT) models not only take into account of the neighbor structure but also employ the three-order tensor technique to keep the geometric consistency. Extensive experiments on the synthetic and real datasets show that LTPT performs better than the state-of-the-art graph matching methods. Jiufeng Zhou, Hong Yan 0001, Yuan Zhu 0005 |
SMC | 3 |
| 2015 | Determining minimum set of driver nodes in protein-protein interaction networksabstractBACKGROUND: Recently, several studies have drawn attention to the determination of a minimum set of driver proteins that are important for the control of the underlying protein-protein interaction (PPI) networks. In general, the minimum dominating set (MDS) model is widely adopted. However, because the MDS model does not generate a unique MDS configuration, multiple different MDSs would be generated when using different optimization algorithms. Therefore, among these MDSs, it is difficult to find out the one that represents the true driver set of proteins. RESULTS: To address this problem, we develop a centrality-corrected minimum dominating set (CC-MDS) model which includes heterogeneity in degree and betweenness centralities of proteins. Both the MDS model and the CC-MDS model are applied on three human PPI networks. Unlike the MDS model, the CC-MDS model generates almost the same sets of driver proteins when we implement it using different optimization algorithms. The CC-MDS model targets more high-degree and high-betweenness proteins than the uncorrected counterpart. The more central position allows CC-MDS proteins to be more important in maintaining the overall network connectivity than MDS proteins. To indicate the functional significance, we find that CC-MDS proteins are involved in, on average, more protein complexes and GO annotations than MDS proteins. We also find that more essential genes, aging genes, disease-associated genes and virus-targeted genes appear in CC-MDS proteins than in MDS proteins. As for the involvement in regulatory functions, the sets of CC-MDS proteins show much stronger enrichment of transcription factors and protein kinases. The results about topological and functional significance demonstrate that the CC-MDS model can capture more driver proteins than the MDS model. CONCLUSIONS: Based on the results obtained, the CC-MDS model presents to be a powerful tool for the determination of driver proteins that can control the underlying PPI networks. The software described in this paper and the datasets used are available at https://github.com/Zhangxf-ccnu/CC-MDS . Xiao-Fei Zhang, Le Ou-Yang, Yuan Zhu 0005, Meng-Yun Wu, Dao-Qing Dai |
BMC Bioinform. | 3 |
| 2013 | Identification of DNA-Binding and Protein-Binding Proteins Using Enhanced Graph Wavelet FeaturesabstractInteractions between biomolecules play an essential role in various biological processes. For predicting DNA-binding or protein-binding proteins, many machine-learning-based techniques have used various types of features to represent the interface of the complexes, but they only deal with the properties of a single atom in the interface and do not take into account the information of neighborhood atoms directly. This paper proposes a new feature representation method for biomolecular interfaces based on the theory of graph wavelet. The enhanced graph wavelet features (EGWF) provides an effective way to characterize interface feature through adding physicochemical features and exploiting a graph wavelet formulation. Particularly, graph wavelet condenses the information around the center atom, and thus enhances the discrimination of features of biomolecule binding proteins in the feature space. Experiment results show that EGWF performs effectively for predicting DNA-binding and protein-binding proteins in terms of Matthew's correlation coefficient (MCC) score and the area value under the receiver operating characteristic curve (AUC). Yuan Zhu 0005, Weiqiang Zhou, Dao-Qing Dai, Hong Yan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |