VLDB 2026 Research / reviewers in the wild / expert
Le Ou-Yang
dblp:146/9908
· DBLP profile ↗
64ranked-venue papers
11as first author
33since 2021 · last 2026
0000-0003-4007-4568ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 50 · 9 first-author · 25 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 5 since 2021Computer networks · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SeOMLR: one-step multi-view latent representation with self-weighted ensemble learning for multi-omics cancer subtypingabstractMOTIVATION: Accurate cancer subtyping is critically important for cancer treatment due to significant molecular heterogeneity. While existing methods with multi-omics integration have achieved some success in cancer subtype identification by leveraging the rich information provided by multi-omics data, most approaches remain limited by an overemphasis on cross-omics consistency at the expense of intra-omics specificity. Furthermore, a two-step scheme is often adopted to extract cluster structure from a consistency matrix or a continuous indicator matrix by k-means, which inevitably leads to information loss and unstable clusters. RESULTS: To overcome these issues, we propose seOMLR, a one-step multi-view latent representation method with self-weighted ensemble learning for cancer subtyping. Using relaxed exclusivity constraints and consistency regularization terms, seOMLR exploits the specificity and consistency of multi-omics data by building a sparse low-rank self-representation framework. Simultaneously, a self-weighted ensemble strategy is introduced to adaptively incorporate prior subtyping information from other methods, indirectly promoting specificity and consistency learning. Moreover, the discrete clustering structure is subsequently extracted via spectral rotation to avoid information loss and cluster instability. Through joint iterative optimization of fusion and clustering, seOMLR enhances subtyping accuracy. Experiments on both simulated datasets and eight real multi-omics cancer datasets from TCGA demonstrate that seOMLR outperforms competing methods, achieving efficient multi-omics data fusion and providing computational framework support for cancer subtyping research. AVAILABILITY AND IMPLEMENTATION: Supplementary data are available at Bioinformatics online. Wenjing Song, Yesen Sun, Le Ou-Yang |
Bioinform. | 3 |
| 2026 | FEQIN: A feature-enhanced query interaction network for breast ultrasound segmentation
Xiling Luo, Le Ou-Yang |
Pattern Recognit. | 2 |
| 2026 | GFLearn: Generalized Feature Learning for Drug-Target Binding Affinity PredictionabstractPredicting drug-target binding affinity is critical for drug discovery, as it helps identify promising drug candidates and predict their effectiveness. Recent advancements in deep learning have made significant progress in tackling this task. However, existing methods heavily rely on training data, and their performance is often limited when predicting binding affinities for new drugs and targets. To address this challenge, we propose a novel Generalized Feature Learning (GFLearn) model for drug-target binding affinity prediction. By integrating Graph Neural Networks (GNNs) with a self-supervised invariant feature learning module, our GFLearn model can extract robust and highly generalizable features from both drugs and targets, significantly enhancing prediction performance. This innovation enables the model to effectively predict binding affinities for previously unseen drugs or targets, while also mitigating the common issue of prediction performance degradation caused by shifts in data distribution. Extensive experiments were conducted on two diverse datasets across three challenging scenarios: new drugs, new targets, and combinations of both. Comparisons with state-of-the-art methods demonstrated that our GFLearn model consistently outperformed others, showcasing its robustness across various prediction tasks. Additionally, cross-dataset evaluations and noise perturbation experiments further validated the model's generalizability across different data distributions. Case studies on two drug-target pairs, Canertinib-PIK3C2G and MLN8054-FLT1, provided further evidence of GFLearn's ability to make accurate binding affinity predictions, offering valuable insights for drug screening and repurposing efforts. Zibo Huang, Xinrui Weng, Le Ou-Yang |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | LineGRN: A Line Graph Neural Network for Gene Regulatory Network InferenceabstractGene regulatory networks (GRNs) depicts the complex interactions between transcription factors and target genes, offering profound insight into deciphering the mechanisms of cellular processes. The advancement of single-cell RNA sequencing (scRNA-seq) technologies has provided a crucial perspective for inferring GRNs at single-cell resolution, leading to the development of numerous computational methods for GRN inference. However, most existing methods fail to adequately capture the association patterns between gene pairs, and the low-degree-node-dominated topology of prior GRNs imposes fundamental limitations on information propagation. In this study, we propose LineGRN, a novel line graph neural network framework for inferring GRNs from scRNA-seq data. By modeling the neighborhood relationships between gene pairs, LineGRN effectively preserves interaction signals within the topological structure. Moreover, the line graph transformation produces a high-degree-node-dominated local network topology, which enables more efficient information propagation. Comprehensive experiments on real datasets demonstrate that LineGRN significantly outperforms seven state-of-the-art methods. Furthermore, LineGRN exhibits low sensitivity to parameter variations and noise interference. Notably, case studies provide empirical evidence of the model's ability to uncover potential TF-target regulatory associations. Weiming Yu, Le Ou-Yang |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | TransGRN: A Transfer Learning-Based Framework for Inferring Gene Regulatory Networks Across Cell LinesabstractInferring gene regulatory networks (GRNs) is critical for understanding the mechanisms that govern cellular behavior. Advances in single-cell RNA sequencing (scRNA-seq) have enabled GRN analysis at single-cell resolution and stimulated the development of many computational methods. However, most existing approaches depend heavily on extensive prior regulatory information, which limits their effectiveness in few-shot settings where such data for the target cell line are scarce or unavailable. To address this challenge, we propose TransGRN, a transfer learning-based method for inferring gene regulatory networks (GRNs) across cell lines. TransGRN adopts a cross-cell-line pre-training strategy that combines scRNA-seq data from multiple source cell lines with biological knowledge obtained from large language models. In addition, it includes a regulatory interaction extraction module that integrates gene expression profiles with semantic information. By transferring generalizable gene-gene regulatory patterns from source to target cell lines, TransGRN achieves state-of-the-art performance in both benchmark tests and few-shot GRN inference tasks. Weiming Yu, Le Ou-Yang |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | LGFFM: A Localized and Globalized Frequency Fusion Model for Ultrasound Image SegmentationabstractAccurate segmentation of ultrasound images plays a critical role in disease screening and diagnosis. Recently, neural network-based methods have garnered significant attention for their potential in improving ultrasound image segmentation. However, these methods still face significant challenges, primarily due to inherent issues in ultrasound images, such as low resolution, speckle noise, and artifacts. Additionally, ultrasound image segmentation encompasses a wide range of scenarios, including organ segmentation (e.g., cardiac and fetal head) and lesion segmentation (e.g., breast cancer and thyroid nodules), making the task highly diverse and complex. Existing methods are often designed for specific segmentation scenarios, which limits their flexibility and ability to meet the diverse needs across various scenarios. To address these challenges, we propose a novel Localized and Globalized Frequency Fusion Model (LGFFM) for ultrasound image segmentation. Specifically, we first design a Parallel Bi-Encoder (PBE) architecture that integrates Local Feature Blocks (LFB) and Global Feature Blocks (GLB) to enhance feature extraction. Additionally, we introduce a Frequency Domain Mapping Module (FDMM) to capture texture information, particularly high-frequency details such as edges. Finally, a Multi-Domain Fusion (MDF) method is developed to effectively integrate features across different domains. We conduct extensive experiments on eight representative public ultrasound datasets across four different types. The results demonstrate that LGFFM outperforms current state-of-the-art methods in both segmentation accuracy and generalization performance. Xiling Luo, Yi Wang 0031, Le Ou-Yang |
IEEE Trans. Medical Imaging | 3 |
| 2025 | GeCC: Generalized Contrastive Clustering with Domain Shifts ModelingabstractContrastive clustering performs clustering and data representation in a unified model, where instance- and cluster-level constrastive learning are conducted simultaneously. However, commonly-used data augmentation methods make contrastive mechanism effect but may cause representation learning getting stuck in domain-specific information, which further deteriorates clustering performance and limits generalization ability. To this end, we propose a new framework, named Generalized Contrastive Clustering with domain shifts modeling (GeCC), which can integrate diverse domain knowledge to improve the clustering performance. Specifically, we first design a cluster-guided domain shifts modeling module to synthesize a reference view with diverse domain information. Then, we introduce instance representation and cluster assignment contrastive modules with well-designed attention weights to guide the representation learning and clustering. In this way, our method can maximize the extraction of cluster-related information and avoid over-fitting domain-specific features. Experimental results on four benchmark datasets demonstrate that our proposed method consistently outperforms other state-of-the-art methods. Wenhui Wu 0001, Le Ou-Yang, Ran Wang 0001, Debby Dan Wang |
AAAI | 3 |
| 2025 | DUIMC: Deep Unbalanced Incomplete Multi-View Clustering via Graph Constrained Imputation and Contrastive LearningabstractDue to the frequent occurrence of missing views in real-world multi-view data, incomplete multi-view clustering (IMVC) has attracted significant attention. However, most existing IMVC methods overlook the fact that incomplete data in practical applications often exhibits varying missing rates across different views, rendering their mechanisms ineffective under such conditions. Although several works based on conventional learning methods have been proposed to solve unbalanced incomplete multi-view clustering (UIMVC), their performance is limited by their shallow feature representation and over-sophisticated optimization procedure. In this paper, we propose Deep Unbalanced Incomplete Multi-view Clustering via Graph Constrained Imputation and Contrastive Learning (DUIMC) to address UIMVC with deep learning paradigm. Specifically, DUIMC introduces a novel differentiable imputation layer for dynamically handling unbalanced incompleteness and integrates it with multi-view contrastive clustering into a unified deep representation learning framework. Furthermore, bi-level graph constraints are imposed on imputation and representation learning to preserve local consistency at both the feature and instance levels. In addition, we develop adaptive fusion mechanisms to adaptively restrain the impact aroused by information unbalance among views. Extensive experimental results on five benchmark datasets demonstrate DUIMC's superior clustering performance over several traditional state-of-the-art approaches. Wenhui Wu 0001, Guanqi Wen, Le Ou-Yang, Ran Wang 0001, Sam Kwong |
ACM Multimedia | 3 |
| 2025 | Reinforcement learning-guided sample selection for improved identification of phage receptor binding proteinsabstractAbstract Receptor binding proteins (RBPs) are critical for bacteriophage infection, mediating the initial recognition and attachment to specific bacterial surface receptors, thereby determining host range and infection specificity. Despite their functional importance, RBPs exhibit high sequence diversity, vary significantly in length (200–2000+ amino acids), and remain poorly annotated in current databases. This low conservation limits the effectiveness of homology-based search methods, which often fail to detect distantly related RBPs. Deep learning models offer a more flexible alternative by learning complex sequence patterns, but face challenges in training due to extreme class imbalance—RBPs are vastly outnumbered by non-RBP phage proteins, leading to inefficient learning and biased predictions. To address this, we propose a deep learning framework enhanced with a reinforcement learning-based sample selection strategy that dynamically prioritizes the most informative negative samples during training. Evaluated on a curated phage protein dataset, our method achieves a precision of 0.9455 and an F1-score of 0.7309, outperforming existing homology-based and deep learning approaches. Our results demonstrate that integrating strategic sample selection significantly improves RBP detection in diverse phage genomes. Xiling Luo, Le Ou-Yang, Yanni Sun, Jiayu Shang 0001 |
Briefings Bioinform. | 2 |
| 2025 | GCLink: a graph contrastive link prediction framework for gene regulatory network inferenceabstractMOTIVATION: Gene regulatory networks (GRNs) unveil the intricate interactions among genes, pivotal in elucidating the complex biological processes within cells. The advent of single-cell RNA-sequencing (scRNA-seq) enables the inference of GRNs at single-cell resolution. However, the majority of current supervised network inference methods typically concentrate on predicting pairwise gene regulatory interaction, thus failing to fully exploit correlations among all genes and exhibiting limited generalization performance. RESULTS: To address these issues, we propose a graph contrastive link prediction (GCLink) model to infer potential gene regulatory interactions from scRNA-seq data. Based on known gene regulatory interactions and scRNA-seq data, GCLink introduces a graph contrastive learning strategy to aggregate the feature and neighborhood information of genes to learn their representations. This approach reduces the dependence of our model on sample size and enhance its ability in predicting potential gene regulatory interactions. Extensive experiments on real scRNA-seq datasets demonstrate that GCLink outperforms other state-of-the-art methods in most cases. Furthermore, by pretraining GCLink on a source cell line with abundant known regulatory interactions and fine-tuning it on a target cell line with limited amount of known interactions, our GCLink model exhibits good performance in GRN inference, demonstrating its effectiveness in inferring GRNs from datasets with limited known interactions. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/Yoyiming/GCLink. Weiming Yu, Zerun Lin, Miaofang Lan, Le Ou-Yang |
Bioinform. | 4 |
| 2025 | GRESS: Grouping Belief-Based Deep Contrastive Subspace ClusteringabstractThe self-expressive coefficient plays a crucial role in the self-expressiveness-based subspace clustering method. To enhance the precision of the self-expressive coefficient, we propose a novel deep subspace clustering method, named grouping belief-based deep contrastive subspace clustering (GRESS), which integrates the clustering information and higher-order relationship into the coefficient matrix. Specifically, we develop a deep contrastive subspace clustering module to enhance the learning of both self-expressive coefficients and cluster representations simultaneously. This approach enables the derivation of relatively noiseless self-expressive similarities and cluster-based similarities. To enable interaction between these two types of similarities, we propose a unique grouping belief-based affinity refinement module. This module leverages grouping belief to uncover the higher-order relationships within the similarity matrix, and integrates the well-designed noisy similarity suppression and similarity increment regularization to eliminate redundant connections while complete absent information. Extensive experimental results on four benchmark datasets validate the superiority of our proposed method GRESS over several state-of-the-art methods. Wenhui Wu 0001, Le Ou-Yang, Ran Wang 0001, Sam Kwong |
IEEE Trans. Cybern. | 3 |
| 2024 | Pathway Activity Autoencoders for Enhanced Omics Analysis and Clinical InterpretabilityabstractLarge-scale analyses of omics data are crucial for advancing precision medicine and personalised treatments. Current methods to unravel cancer development and prognosis exhibit a dichotomy between interpretability and representational power. Most either rely on post-hoc interpretability techniques with black-box models or use linear models that might miss complex biological interactions. We propose a novel configurable prior-knowledge-based deep auto-encoding framework called PAAE and its generative variant PAVAE, for analyzing cancer RNA-seq data. Our method constrains its learned internal representation with biological pathways, providing interpretability without sacrificing predictive power. Our model is tested on 3 different downstream tasks: cancer subtype classification, survival analysis and unsupervised clustering of the learned features. Our models outperform baselines while having orders of magnitude less parameters than naive models. Extensive interpretability analyses, including task-relevant feature identification demonstrate our model’s effectiveness at identifying underlying biological signals in an unsupervised fashion. Visualisations are used to highlight the intrinsic interpretability of our models. The source code of this study is available at github.com/phcavelar/pathwayae Pedro H. C. Avelar, Le Ou-Yang, Min Wu 0008, Sophia Tsoka |
BIBM | 2 |
| 2024 | Clustering single-cell multi-omics data via graph regularized multi-view ensemble learningabstractMOTIVATION: Single-cell clustering plays a crucial role in distinguishing between cell types, facilitating the analysis of cell heterogeneity mechanisms. While many existing clustering methods rely solely on gene expression data obtained from single-cell RNA sequencing techniques to identify cell clusters, the information contained in mono-omic data is often limited, leading to suboptimal clustering performance. The emergence of single-cell multi-omics sequencing technologies enables the integration of multiple omics data for identifying cell clusters, but how to integrate different omics data effectively remains challenging. In addition, designing a clustering method that performs well across various types of multi-omics data poses a persistent challenge due to the data's inherent characteristics. RESULTS: In this paper, we propose a graph-regularized multi-view ensemble clustering (GRMEC-SC) model for single-cell clustering. Our proposed approach can adaptively integrate multiple omics data and leverage insights from multiple base clustering results. We extensively evaluate our method on five multi-omics datasets through a series of rigorous experiments. The results of these experiments demonstrate that our GRMEC-SC model achieves competitive performance across diverse multi-omics datasets with varying characteristics. AVAILABILITY AND IMPLEMENTATION: Implementation of GRMEC-SC, along with examples, can be found on the GitHub repository: https://github.com/polarisChen/GRMEC-SC. Fuqun Chen, Guanhua Zou, Yongxian Wu, Le Ou-Yang |
Bioinform. | 4 |
| 2024 | MARS: a motif-based autoregressive model for retrosynthesis predictionabstractMOTIVATION: Retrosynthesis is a critical task in drug discovery, aimed at finding a viable pathway for synthesizing a given target molecule. Many existing approaches frame this task as a graph-generating problem. Specifically, these methods first identify the reaction center, and break a targeted molecule accordingly to generate the synthons. Reactants are generated by either adding atoms sequentially to synthon graphs or by directly adding appropriate leaving groups. However, both of these strategies have limitations. Adding atoms results in a long prediction sequence that increases the complexity of generation, while adding leaving groups only considers those in the training set, which leads to poor generalization. RESULTS: In this paper, we propose a novel end-to-end graph generation model for retrosynthesis prediction, which sequentially identifies the reaction center, generates the synthons, and adds motifs to the synthons to generate reactants. Given that chemically meaningful motifs fall between the size of atoms and leaving groups, our model achieves lower prediction complexity than adding atoms and demonstrates superior performance than adding leaving groups. We evaluate our proposed model on a benchmark dataset and show that it significantly outperforms previous state-of-the-art models. Furthermore, we conduct ablation studies to investigate the contribution of each component of our proposed model to the overall performance on benchmark datasets. Experiment results demonstrate the effectiveness of our model in predicting retrosynthesis pathways and suggest its potential as a valuable tool in drug discovery. AVAILABILITY AND IMPLEMENTATION: All code and data are available at https://github.com/szu-ljh2020/MARS. Jiahan Liu, Chaochao Yan, Yang Yu 0010, Chan Lu, Junzhou Huang, Le Ou-Yang, Peilin Zhao |
Bioinform. | 6 |
| 2024 | Equivariant Line Graph Neural Network for Protein-Ligand Binding Affinity PredictionabstractBinding affinity prediction of three-dimensional (3D) protein-ligand complexes is critical for drug repositioning and virtual drug screening. Existing approaches usually transform a 3D protein-ligand complex to a two-dimensional (2D) graph, and then use graph neural networks (GNNs) to predict its binding affinity. However, the node and edge features of the 2D graph are extracted based on invariant local coordinate systems of the 3D complex. As a result, these approaches can not fully learn the global information of the complex, such as the physical symmetry and the topological information of bonds. To address these issues, we propose a novel Equivariant Line Graph Network (ELGN) for binding affinity prediction of 3D protein-ligand complexes. The proposed ELGN firstly adds a super node to the 3D complex, and then builds a line graph based on the 3D complex. After that, ELGN uses a new E(3)-equivariant network layer to pass the messages between nodes and edges based on the global coordinate system of the 3D complex. Experimental results on two real datasets demonstrate the effectiveness of ELGN over several state-of-the-art baselines. Yiqiang Yi, Kangfei Zhao, Le Ou-Yang, Peilin Zhao |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | MoVAE: A Variational AutoEncoder for Molecular Graph GenerationabstractMolecule generation plays an important role in accelerating drug discovery. In recent years, many molecule generation methods have been proposed based on variational autoencoders (VAEs), due to its advantages in latent manifold representation learning and training stability. However, most of the existing VAE-based models require tedious graph matching operations during training, and tend to generate invalid molecules. To overcome these limitations, in this paper, we propose a novel molecular graph variational autoencoder (MoVAE). Firstly, to avoid complicated graph matching, the proposed MoVAE only encodes and decodes all the nodes and edges individually. Secondly, to improve the generation validity, it adversarially trains the model by treating the encoder and decoder as the discriminator and generator. In addition, to generate molecules with various target conditions, the MoVAE also introduces drug property constraints and valence histogram constraints. Experiment results on two real datasets show that our model outperforms almost all the state-of-the-art algorithms. Zerun Lin, Lixin Duan, Le Ou-Yang, Peilin Zhao |
SDM | 4 |
| 2023 | Inferring gene regulatory networks from single-cell gene expression data via deep multi-view contrastive learningabstractThe inference of gene regulatory networks (GRNs) is of great importance for understanding the complex regulatory mechanisms within cells. The emergence of single-cell RNA-sequencing (scRNA-seq) technologies enables the measure of gene expression levels for individual cells, which promotes the reconstruction of GRNs at single-cell resolution. However, existing network inference methods are mainly designed for data collected from a single data source, which ignores the information provided by multiple related data sources. In this paper, we propose a multi-view contrastive learning (DeepMCL) model to infer GRNs from scRNA-seq data collected from multiple data sources or time points. We first represent each gene pair as a set of histogram images, and then introduce a deep Siamese convolutional neural network with contrastive loss to learn the low-dimensional embedding for each gene pair. Moreover, an attention mechanism is introduced to integrate the embeddings extracted from different data sources and different neighbor gene pairs. Experimental results on synthetic and real-world datasets validate the effectiveness of our contrastive learning and attention mechanisms, demonstrating the effectiveness of our model in integrating multiple data sources for GRN inference. Zerun Lin, Le Ou-Yang |
Briefings Bioinform. | 2 |
| 2023 | Reading Multilevel 2-D Barcodes Using a Machine Learning ApproachabstractThis article addresses the reading problem of multilevel 2-D barcodes over a print-and-capture (PC) channel. The prior reading schemes have different limitations to hinder their applications, e.g., suffering from quantization error, being sensitive to the predetermined decision boundaries, and being sensitive to the selection of initial parameters. In this article, we introduce a machine learning approach to address the above limitations using a new ensemble clustering (EC) algorithm. Based on the new EC algorithm, we propose two reading schemes of a multilevel 2-D barcode. Specifically, the first proposed scheme is named the EC reading scheme. In the EC reading scheme, we introduce a weighted ensemble mechanism to assign different weights to different base clustering results. Then, we propose the second scheme, named the enhanced EC (EEC) reading scheme, to further improve the reading performance with the help of the reference symbols. We implement our approach and conduct extensive performance comparisons through an actual excremental platform under various multilevel 2-D barcodes and various capturing devices. From experimental results, we observe that both proposed reading schemes have better performance than the prior reading schemes. Moreover, the EEC reading scheme has better performance than the EC reading scheme, and their performance gap becomes more apparent as the distortion of a PC channel increases. Jiaheng Zhang, Le Ou-Yang, Changsheng Chen 0001, Ning Xie 0007 |
IEEE Internet Things J. | 3 |
| 2023 | Self-representative kernel concept factorization
Wenhui Wu 0001, Ran Wang 0001, Le Ou-Yang |
Knowl. Based Syst. | 4 |
| 2023 | scTSSR2: Imputing Dropout Events for Single-Cell RNA Sequencing Using Fast Two-Side Self-RepresentationabstractThe single-cell RNA sequencing (scRNA-seq) technique begins a new era by revealing gene expression patterns at single-cell resolution, enabling studies of heterogeneity and transcriptome dynamics of complex tissues at single-cell resolution. However, existing large proportion of dropout events may hinder downstream analyses. Thus imputation of dropout events is an important step in analyzing scRNA-seq data. We develop scTSSR2, a new imputation method that combines matrix decomposition with the previously developed two-side sparse self-representation, leading to fast two-side sparse self-representation to impute dropout events in scRNA-seq data. The comparisons of computational speed and memory usage among different imputation methods show that scTSSR2 has distinct advantages in terms of computational speed and memory usage. Comprehensive downstream experiments show that scTSSR2 outperforms the state-of-the-art imputation methods. A user-friendly R package scTSSR2 is developed to denoise the scRNA-seq data to improve the data quality. Le Ou-Yang, Hong Yan 0001, Xiao-Fei Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | scMIC: A Deep Multi-Level Information Fusion Framework for Clustering Single-Cell Multi-Omics DataabstractCell type identification is a crucial step towards the study of cellular heterogeneity and biological processes. Advances in single-cell sequencing technology have enabled the development of a variety of clustering methods for cell type identification. However, most of existing methods are designed for clustering single omic data such as single-cell RNA-sequencing (scRNA-seq) data. The accumulation of single-cell multi-omics data provides a great opportunity to integrate different omics data for cell clustering, but also raise new computational challenges for existing methods. How to integrate multi-omics data and leverage their consensus and complementary information to improve the accuracy of cell clustering still remains a challenge. In this study, we propose a new deep multi-level information fusion framework, named scMIC, for clustering single-cell multi-omics data. Our model can integrate the attribute information of cells and the potential structural relationship among cells from local and global levels, and reduce redundant information between different omics from cell and feature levels, leading to more discriminative representations. Moreover, the proposed multiple collaborative supervised clustering strategy is able to guide the learning process of the core encoding part by learning the high-confidence target distribution, which facilitates the interaction between the clustering part and the representation learning part, as well as the information exchange between omics, and finally obtain more robust clustering results. Experiments on seven single-cell multi-omics datasets show the superiority of scMIC over existing state-of-the-art methods. Youlin Zhan, Jiahan Liu, Le Ou-Yang |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | scDEA: differential expression analysis in single-cell RNA-sequencing data via ensemble learningabstractThe identification of differentially expressed genes between different cell groups is a crucial step in analyzing single-cell RNA-sequencing (scRNA-seq) data. Even though various differential expression analysis methods for scRNA-seq data have been proposed based on different model assumptions and strategies recently, the differentially expressed genes identified by them are quite different from each other, and the performances of them depend on the underlying data structures. In this paper, we propose a new ensemble learning-based differential expression analysis method, scDEA, to produce a more stable and accurate result. scDEA integrates the P-values obtained from 12 individual differential expression analysis methods for each gene using a P-value combination method. Comprehensive experiments show that scDEA outperforms the state-of-the-art individual methods with different experimental settings and evaluation metrics. We expect that scDEA will serve a wide range of users, including biologists, bioinformaticians and data scientists, who need to detect differentially expressed genes in scRNA-seq data. Hui-Sheng Li, Le Ou-Yang, Yuan Zhu 0005, Hong Yan 0001, Xiao-Fei Zhang |
Briefings Bioinform. | 2 |
| 2022 | Matrix factorization for biomedical link prediction and scRNA-seq data imputation: an empirical surveyabstractAdvances in high-throughput experimental technologies promote the accumulation of vast number of biomedical data. Biomedical link prediction and single-cell RNA-sequencing (scRNA-seq) data imputation are two essential tasks in biomedical data analyses, which can facilitate various downstream studies and gain insights into the mechanisms of complex diseases. Both tasks can be transformed into matrix completion problems. For a variety of matrix completion tasks, matrix factorization has shown promising performance. However, the sparseness and high dimensionality of biomedical networks and scRNA-seq data have raised new challenges. To resolve these issues, various matrix factorization methods have emerged recently. In this paper, we present a comprehensive review on such matrix factorization methods and their usage in biomedical link prediction and scRNA-seq data imputation. Moreover, we select representative matrix factorization methods and conduct a systematic empirical comparison on 15 real data sets to evaluate their performance under different scenarios. By summarizing the experimental results, we provide general guidelines for selecting matrix factorization methods for different biomedical matrix completion tasks and point out some future directions to further improve the performance for biomedical link prediction and scRNA-seq data imputation. Le Ou-Yang, Zi-Chao Zhang 0001, Min Wu 0008 |
Briefings Bioinform. | 1 |
| 2022 | DEMOC: a deep embedded multi-omics learning approach for clustering single-cell CITE-seq dataabstractAdvances in single-cell RNA sequencing (scRNA-seq) technologies has provided an unprecedent opportunity for cell-type identification. As clustering is an effective strategy towards cell-type identification, various computational approaches have been proposed for clustering scRNA-seq data. Recently, with the emergence of cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq), the cell surface expression of specific proteins and the RNA expression on the same cell can be captured, which provides more comprehensive information for cell analysis. However, existing single cell clustering algorithms are mainly designed for single-omic data, and have difficulties in handling multi-omics data with diverse characteristics efficiently. In this study, we propose a novel deep embedded multi-omics clustering with collaborative training (DEMOC) model to perform joint clustering on CITE-seq data. Our model can take into account the characteristics of transcriptomic and proteomic data, and make use of the consistent and complementary information provided by different data sources effectively. Experiment results on two real CITE-seq datasets demonstrate that our DEMOC model not only outperforms state-of-the-art single-omic clustering methods, but also achieves better and more stable performance than existing multi-omics clustering methods. We also apply our model on three scRNA-seq datasets to assess the performance of our model in rare cell-type identification, novel cell-subtype detection and cellular heterogeneity analysis. Experiment results illustrate the effectiveness of our model in discovering the underlying patterns of data. Guanhua Zou, Yilong Lin, Tianyang Han, Le Ou-Yang |
Briefings Bioinform. | 4 |
| 2022 | Identifying Gene Network Rewiring Based on Partial CorrelationabstractIt is an important task to learn how gene regulatory networks change under different conditions. Several Gaussian graphical model-based methods have been proposed to deal with this task by inferring differential networks from gene expression data. However, most existing methods define the differential networks as the difference of precision matrices, which may include false differential edges caused by the change of conditional variances. In addition, prior information about the condition-specific networks and the differential networks can be obtained from other domains. It is useful to incorporate prior information into differential network analysis. In this study, we propose a new differential network analysis method to address the above challenges. Instead of using the precision matrices, we define the differential networks as the difference of partial correlations, which can exclude the spurious differential edges due to the variants of conditional variances. Furthermore, prior information from multiple hypothesis testing is incorporated using a weighted fused penalty. Simulation studies show that our method outperforms the competing methods. We also apply our method to identify the differential network between luminal A and basal-like subtypes of breast cancers and the differential network between acute myeloid leukemia tumors and normal samples. The hub genes in the differential networks identified by our method carry out important biological functions. Yuting Tan 0001, Le Ou-Yang, Xingpeng Jiang, Hong Yan 0001, Xiao-Fei Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Inferring Gene Co-Expression Networks by Incorporating Prior Protein-Protein Interaction NetworksabstractInferring gene co-expression networks from high-throughput gene expression data is an important task in bioinformatics. Many gene networks often exhibit modular structures. Although several Gaussian graphical model-based methods have been developed to estimate gene co-expression networks by incorporating the modular structural prior, none of them takes into account the modular structures captured by the prior networks (e.g., protein interaction networks). In this study, we propose a novel prior network-dependent gene network inference (pGNI) method to estimate gene co-expression networks by integrating gene expression data and prior protein interaction network data. The underlying modular structure is learned from both sets of data. Through simulation studies, we demonstrate the feasibility and effectiveness of our method. We also apply our method to two real datasets. The modular structures in the networks estimated by our method are biological significant. Meng-Guo Wang, Le Ou-Yang, Hong Yan 0001, Xiao-Fei Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | Recent advances in network-based methods for disease gene predictionabstractDisease-gene association through genome-wide association study (GWAS) is an arduous task for researchers. Investigating single nucleotide polymorphisms that correlate with specific diseases needs statistical analysis of associations. Considering the huge number of possible mutations, in addition to its high cost, another important drawback of GWAS analysis is the large number of false positives. Thus, researchers search for more evidence to cross-check their results through different sources. To provide the researchers with alternative and complementary low-cost disease-gene association evidence, computational approaches come into play. Since molecular networks are able to capture complex interplay among molecules in diseases, they become one of the most extensively used data for disease-gene association prediction. In this survey, we aim to provide a comprehensive and up-to-date review of network-based methods for disease gene prediction. We also conduct an empirical analysis on 14 state-of-the-art methods. To summarize, we first elucidate the task definition for disease gene prediction. Secondly, we categorize existing network-based efforts into network diffusion methods, traditional machine learning methods with handcrafted graph features and graph representation learning methods. Thirdly, an empirical analysis is conducted to evaluate the performance of the selected methods across seven diseases. We also provide distinguishing findings about the discussed methods based on our empirical analysis. Finally, we highlight potential research directions for future studies on disease gene prediction. Sezin Kircali Ata, Min Wu 0008, Yuan Fang 0001, Le Ou-Yang, Chee Keong Kwoh 0001, Xiaoli Li 0001 |
Briefings Bioinform. | 4 |
| 2021 | WDNE: an integrative graphical model for inferring differential networks from multi-platform gene expression data with missing valuesabstractThe mechanisms controlling biological process, such as the development of disease or cell differentiation, can be investigated by examining changes in the networks of gene dependencies between states in the process. High-throughput experimental methods, like microarray and RNA sequencing, have been widely used to gather gene expression data, which paves the way to infer gene dependencies based on computational methods. However, most differential network analysis methods are designed to deal with fully observed data, but missing values, such as the dropout events in single-cell RNA-sequencing data, are frequent. New methods are needed to take account of these missing values. Moreover, since the changes of gene dependencies may be driven by certain perturbed genes, considering the changes in gene expression levels may promote the identification of gene network rewiring. In this study, a novel weighted differential network estimation (WDNE) model is proposed to handle multi-platform gene expression data with missing values and take account of changes in gene expression levels. Simulation studies demonstrate that WDNE outperforms state-of-the-art differential network estimation methods. When applied WDNE to infer differential gene networks associated with drug resistance in ovarian tumors, cell differentiation and breast tumor heterogeneity, the hub genes in the estimated differential gene networks can provide important insights into the underlying mechanisms. Furthermore, a Matlab toolbox, differential network analysis toolbox, was developed to implement the WDNE model and visualize the estimated differential networks. Le Ou-Yang, Dehan Cai, Xiao-Fei Zhang, Hong Yan 0001 |
Briefings Bioinform. | 1 |
| 2021 | Differential network analysis by simultaneously considering changes in gene interactions and gene expressionabstractMOTIVATION: Differential network analysis is an important tool to investigate the rewiring of gene interactions under different conditions. Several computational methods have been developed to estimate differential networks from gene expression data, but most of them do not consider that gene network rewiring may be driven by the differential expression of individual genes. New differential network analysis methods that simultaneously take account of the changes in gene interactions and changes in expression levels are needed. RESULTS: : In this article, we propose a differential network analysis method that considers the differential expression of individual genes when identifying differential edges. First, two hypothesis test statistics are used to quantify changes in partial correlations between gene pairs and changes in expression levels for individual genes. Then, an optimization framework is proposed to combine the two test statistics so that the resulting differential network has a hierarchical property, where a differential edge can be considered only if at least one of the two involved genes is differentially expressed. Simulation results indicate that our method outperforms current state-of-the-art methods. We apply our method to identify the differential networks between the luminal A and basal-like subtypes of breast cancer and those between acute myeloid leukemia and normal samples. Hub nodes in the differential networks estimated by our method, including both differentially and nondifferentially expressed genes, have important biological functions. AVAILABILITY AND IMPLEMENTATION: All the datasets underlying this article are publicly available. Processed data and source code can be accessed through the Github repository at https://github.com/Zhangxf-ccnu/chNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jia-Juan Tu, Le Ou-Yang, Yuan Zhu 0005, Hong Yan 0001, Hong Qin 0008, Xiao-Fei Zhang |
Bioinform. | 2 |
| 2021 | EnTSSR: A Weighted Ensemble Learning Method to Impute Single-Cell RNA Sequencing DataabstractThe advancements of single-cell RNA sequencing (scRNA-seq) technologies have provided us unprecedented opportunities to characterize cellular states and investigate the mechanisms of complex diseases. Due to technical issues such as dropout events, scRNA-seq data contains excess of false zero counts, which has a substantial impact on the downstream analyses. Although several computational approaches have been proposed to impute dropout events in scRNA-seq data, there is no strong consensus on which is the best approach. In this study, we propose a novel weighted ensemble learning method, named EnTSSR, to impute dropout events in scRNA-seq data. By using a multi-view two-side sparse self-representation framework, our model can exploit the consensus similarities between genes and between cells based on the imputed results of various imputation methods. Moreover, we introduce a weighted ensemble strategy to leverage the information captured by various imputation methods effectively. Down-sampling experiments, clustering analysis, differential expression analysis and cell trajectory inference are carried out to evaluate the performance of our proposed model. Experiment results demonstrate that our EnTSSR can effectively recover the true expression pattern of scRNA-seq data. Yilong Lin, Chongbin Yuan, Xiao-Fei Zhang, Le Ou-Yang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | WMLRR: A Weighted Multi-View Low Rank Representation to Identify Cancer Subtypes From Multiple Types of Omics DataabstractThe identification of cancer subtypes is of great importance for understanding the heterogeneity of tumors and providing patients with more accurate diagnoses and treatments. However, it is still a challenge to effectively integrate multiple omics data to establish cancer subtypes. In this paper, we propose an unsupervised integration method, named weighted multi-view low rank representation (WMLRR), to identify cancer subtypes from multiple types of omics data. Given a group of patients described by multiple omics data matrices, we first learn a unified affinity matrix which encodes the similarities among patients by exploring the sparsity-consistent low-rank representations from the joint decompositions of multiple omics data matrices. Unlike existing subtype identification methods that treat each omics data matrix equally, we assign a weight to each omics data matrix and learn these weights automatically through the optimization process. Finally, we apply spectral clustering on the learned affinity matrix to identify cancer subtypes. Experiment results show that the survival times between our identified cancer subtypes are significantly different, and our predicted survivals are more accurate than other state-of-the-art methods. In addition, some clinical analyses of the diseases also demonstrate the effectiveness of our method in identifying molecular subtypes with biological significance and clinical relevance. Yesen Sun, Le Ou-Yang, Dao-Qing Dai |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | Time-Varying Differential Network Analysis for Revealing Network Rewiring over Cancer ProgressionabstractTo reveal how gene regulatory networks change over cancer development, multiple time-varying differential networks between adjacent cancer stages should be estimated simultaneously. Since the network rewiring may be driven by the perturbation of certain individual genes, there may be some hub nodes shared by these differential networks. Although several methods have been developed to estimate differential networks from gene expression data, most of them are designed for estimating a single differential network, which neglect the similarities between different differential networks. In this article, we propose a new Gaussian graphical model-based method to jointly estimate multiple time-varying differential networks for identifying network rewiring over cancer development. A D-trace loss is used to determine the differential networks. A tree-structured group Lasso penalty is designed to identify the common hub nodes shared by different differential networks and the specific hub nodes unique to individual differential networks. Simulation experiment results demonstrate that our method outperforms other state-of-the-art techniques in most cases. We also apply our method to The Cancer Genome Atlas data to explore gene network rewiring over different breast cancer stages. Hub nodes in the estimated differential networks rediscover well known genes associated with the development and progression of breast cancer. Le Ou-Yang, Hong Yan 0001, Xiao-Fei Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | A Joint Graphical Model for Inferring Gene Networks Across Multiple Subpopulations and Data TypesabstractReconstructing gene networks from gene expression data is a long-standing challenge. In most applications, the observations can be divided into several distinct but related subpopulations and the gene expression measurements can be collected from multiple data types. Most existing methods are designed to estimate a single gene network from a single dataset. These methods may be suboptimal since they do not exploit the similarities and differences among different subpopulations and data types. In this article, we propose a joint graphical model to estimate the multiple gene networks simultaneously. Our model decomposes each subpopulation-specific gene network as a sum of common and unique components and imposes a group lasso penalty on gene networks corresponding to different data types. The gene network variations across subpopulations can be learned automatically by the decompositions of networks, and the similarities and differences among data types can be captured by the group lasso penalty. The simulation studies demonstrate that our method outperforms the state-of-the-art methods. We also apply our method to the cancer genome atlas breast cancer datasets to reconstruct subtype-specific gene networks. Hub nodes in the estimated subnetworks unique to individual cancer subtypes rediscover well-known genes associated with breast cancer subtypes and provide interesting predictions. Xiao-Fei Zhang, Le Ou-Yang, Xiaohua Hu 0001, Hong Yan 0001 |
IEEE Trans. Cybern. | 2 |
| 2020 | scTSSR: gene expression recovery for single-cell RNA sequencing using two-side sparse self-representationabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) methods make it possible to reveal gene expression patterns at single-cell resolution. Due to technical defects, dropout events in scRNA-seq will add noise to the gene-cell expression matrix and hinder downstream analysis. Therefore, it is important for recovering the true gene expression levels before carrying out downstream analysis. RESULTS: In this article, we develop an imputation method, called scTSSR, to recover gene expression for scRNA-seq. Unlike most existing methods that impute dropout events by borrowing information across only genes or cells, scTSSR simultaneously leverages information from both similar genes and similar cells using a two-side sparse self-representation model. We demonstrate that scTSSR can effectively capture the Gini coefficients of genes and gene-to-gene correlations observed in single-molecule RNA fluorescence in situ hybridization (smRNA FISH). Down-sampling experiments indicate that scTSSR performs better than existing methods in recovering the true gene expression levels. We also show that scTSSR has a competitive performance in differential expression analysis, cell clustering and cell trajectory inference. AVAILABILITY AND IMPLEMENTATION: The R package is available at https://github.com/Zhangxf-ccnu/scTSSR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Le Ou-Yang, Xing-Ming Zhao, Hong Yan 0001, Xiao-Fei Zhang |
Bioinform. | 2 |
| 2020 | Joint reconstruction of multiple gene networks by simultaneously capturing inter-tumor and intra-tumor heterogeneityabstractMOTIVATION: Reconstruction of cancer gene networks from gene expression data is important for understanding the mechanisms underlying human cancer. Due to heterogeneity, the tumor tissue samples for a single cancer type can be divided into multiple distinct subtypes (inter-tumor heterogeneity) and are composed of non-cancerous and cancerous cells (intra-tumor heterogeneity). If tumor heterogeneity is ignored when inferring gene networks, the edges specific to individual cancer subtypes and cell types cannot be characterized. However, most existing network reconstruction methods do not simultaneously take inter-tumor and intra-tumor heterogeneity into account. RESULTS: In this article, we propose a new Gaussian graphical model-based method for jointly estimating multiple cancer gene networks by simultaneously capturing inter-tumor and intra-tumor heterogeneity. Given gene expression data of heterogeneous samples for different cancer subtypes, a non-cancerous network shared across different cancer subtypes and multiple subtype-specific cancerous networks are estimated jointly. Tumor heterogeneity can be revealed by the difference in the estimated networks. The performance of our method is first evaluated using simulated data, and the results indicate that our method outperforms other state-of-the-art methods. We also apply our method to The Cancer Genome Atlas breast cancer data to reconstruct non-cancerous and subtype-specific cancerous gene networks. Hub nodes in the networks estimated by our method perform important biological functions associated with breast cancer development and subtype classification. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/Zhangxf-ccnu/NETI2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jia-Juan Tu, Le Ou-Yang, Hong Yan 0001, Xiao-Fei Zhang, Hong Qin 0008 |
Bioinform. | 2 |
| 2020 | A graph regularized generalized matrix factorization model for predicting links in biomedical bipartite networksabstractMOTIVATION: Predicting potential links in biomedical bipartite networks can provide useful insights into the diagnosis and treatment of complex diseases and the discovery of novel drug targets. Computational methods have been proposed recently to predict potential links for various biomedical bipartite networks. However, existing methods are usually rely on the coverage of known links, which may encounter difficulties when dealing with new nodes without any known link information. RESULTS: In this study, we propose a new link prediction method, named graph regularized generalized matrix factorization (GRGMF), to identify potential links in biomedical bipartite networks. First, we formulate a generalized matrix factorization model to exploit the latent patterns behind observed links. In particular, it can take into account the neighborhood information of each node when learning the latent representation for each node, and the neighborhood information of each node can be learned adaptively. Second, we introduce two graph regularization terms to draw support from affinity information of each node derived from external databases to enhance the learning of latent representations. We conduct extensive experiments on six real datasets. Experiment results show that GRGMF can achieve competitive performance on all these datasets, which demonstrate the effectiveness of GRGMF in prediction potential links in biomedical bipartite networks. AVAILABILITY AND IMPLEMENTATION: The package is available at https://github.com/happyalfred2016/GRGMF. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zi-Chao Zhang 0001, Xiao-Fei Zhang, Min Wu 0008, Le Ou-Yang, Xing-Ming Zhao, Xiaoli Li 0001 |
Bioinform. | 4 |
| 2020 | Self-weighted adaptive structure learning for ASD diagnosis via multi-template multi-center representation
Fanglin Huang, Ee-Leng Tan, Peng Yang 0011, Le Ou-Yang, Jiuwen Cao, Tianfu Wang 0001, Bai Ying Lei |
Medical Image Anal. | 5 |
| 2020 | Sparse regularized low-rank tensor regression with applications in genomic data analysis
Le Ou-Yang, Xiao-Fei Zhang, Hong Yan 0001 |
Pattern Recognit. | 1 |
| 2020 | Differential Network Analysis via Weighted Fused Conditional Gaussian Graphical ModelabstractThe development and prognosis of complex diseases usually involves changes in regulatory relationships among biomolecules. Understanding how the regulatory relationships change with genetic alterations can help to reveal the underlying biological mechanisms for complex diseases. Although several models have been proposed to estimate the differential network between two different states, they are not suitable to deal with situations where the molecules of interest are affected by other covariates. Nor can they make use of prior information that provides insights about the structures of biomolecular networks. In this study, we introduce a novel weighted fused conditional Gaussian graphical model to jointly estimate two state-specific biomolecular regulatory networks and their difference between two different states. Unlike previous differential network estimation methods, our model can take into account the related covariates and the prior network information when inferring differential networks. The effectiveness of our proposed model is first evaluated based on simulation studies. Experiment results demonstrate that our model outperforms other state-of-the-art differential networks estimation models in all cases. We then apply our model to identify the differential gene network between two subtypes of glioblastoma based on gene expression and miRNA expression data. Our model is able to discover known mechanisms of glioblastoma and provide interesting predictions. Le Ou-Yang, Xiao-Fei Zhang, Xiaohua Hu 0001, Hong Yan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Identifying Gene Network Rewiring Using Robust Differential Graphical Model with Multivariate $t$t-DistributionabstractIdentifying gene network rewiring under different biological conditions is important for understanding the mechanisms underlying complex diseases. Gaussian graphical models, which assume the data follow the multivariate normal distribution, are widely used to identify gene network rewiring. However, the normality assume often fails in reality since the data are contaminated by extreme outliers in general. In this study, we propose a new robust differential graphical model to identify gene network rewiring between two conditions based on the multivariate t-distribution. The multivariate t-distribution is more robust to outliers than the normal distribution since it has heavy tails and allows values far from the mean. A fused lasso penalty is used to borrow information across conditions to improve the results. We develop an expectation maximization algorithm to solve the optimization model. Experiment results on simulated data show that our method outperforms the state-of-the-art methods. Our method is also applied to identify gene network rewiring between luminal A and basal-like subtypes of breast cancer, and gene network rewiring between the proneural and mesenchymal subtypes of glioblastoma. Several key genes which drive gene network rewiring are discovered. Le Ou-Yang, Xiaohua Hu 0001, Xiao-Fei Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | A Machine Learning Approach to Phase Reference Estimation With NoiseabstractThis paper concerns the problem of phase reference estimation with noise, introduced by the imperfect phase-locked loop (PLL) circuit, or the imperfect channel estimation, or both. Prior solutions for suppressing phase noise focus on improving the accuracy of phase reference estimation. The accuracy of phase reference estimation is not high enough due to the following two limitations. First, since the PLL circuit works in radio-frequency (RF), a PLL circuit with high accuracy leads to high cost and high complexity, which makes the deployment difficult. Second, as data rates increase and wireless channels become more complex, the receiver is more difficult to obtain an ideal channel estimation and the negative effect of phase noise becomes more apparent. In this paper, we propose a machine learning approach to mitigate the negative effect of phase noise by using clustering algorithms. The key intuition of our approach is that the clustering algorithm can adaptively trace the shifted constellation point due to the phase noise. Our approach is adaptive because it can adaptively find each received symbol belongs to its original constellation point if the phase noise is not too large, e.g., no larger than $0.25 \pi $ . While the shifted distance is not too large, we can map the received symbols into the correct constellation point to mitigate the negative effect of phase noise. Instead of directly using conventional clustering algorithms into the proposed machine learning approach, we propose a new weighted ensemble clustering algorithm to further improve the performance of our approach. In comparison with prior approaches based on RF circuits, our approach has comparable reception performance but with low complexity and low cost. Our experimental results show that, for a QPSK system, our approach improves the demodulation performance and the decoding performance about 10 dB, 8 dB under BCH codes, and 3 dB under Turbo codes, respectively. Even the demodulation performance of our approach without channel coding is better than the decoding performance of the system with channel coding about 5 dB under BCH codes. Ning Xie 0007, Le Ou-Yang, Alex X. Liu |
IEEE Trans. Commun. | 2 |
| 2020 | Spectrum Sharing in mmWave Cellular Networks Using Clustering AlgorithmsabstractThis paper concerns the problem of spectrum sharing in mmWave cellular networks, where multiple network operators are granted access to the same spectrum resources. Prior spectrum sharing schemes for mmWave cellular networks have two limitations: high coordination overhead and high computational complexity. In this paper, we propose two spectrum sharing schemes for the mobile scenario and the stationary scenario, respectively. In the mobile scenario, we propose a Spectrum sharing scheme using Clustering Algorithms to find the Optimal Positions of Mobile BSes (SCA-OPM). In the stationary scenario, we propose a Spectrum sharing scheme using Clustering Algorithms to Select the most appropriate Subset from all BSes (SCA-SSB). Moreover, for each newly-arriving mobile terminal (MT), we propose two online spectrum sharing schemes based on our SCA-OPM scheme and SCA-SSB scheme, which effectively saves the overall overhead and complexity. Our experimental results show that the SCA-OPM scheme has the best performance, while the SCA-SSB scheme has the same performance as that of the prior scheme but with lower overhead and lower complexity. When the MT density is 200 MTs/km2, for the sum-rate with 10-2bits/s/Hz, the performance gap between the SCA-OPM scheme and the remaining schemes achieves about 19 dB. Moreover, our online schemes have the same performance as those of their counterpart schemes but with lower overhead and lower complexity for the scenario of a few newly-arriving MTs. Ning Xie 0007, Le Ou-Yang, Alex X. Liu |
IEEE/ACM Trans. Netw. | 2 |
| 2020 | A Machine Learning Approach to Blind Multi-Path Classification for Massive MIMO SystemsabstractThis paper concerns the problem of the multi-path classification in a multi-user multi-input multi-output (MIMO) system. We propose a machine learning approach to achieve a blind multi-path classification in the uplink (UL) scheme of a multi-user massive MIMO system. Note that the “blind” term in our approach represents the achievement of the multi-path classification without different pilot sequences on different users, without prior channel state information (CSI) at each user, and without any exploiting the special properties of the received signal. Specifically, our approach consists of two phases. In the first phase, multiple users transmit communication-requests to the base station (BS) for message transmission. The BS only estimates the scaled large-scale path loss of each user, which is determined by the distance between the transmitter and the receiver and is independent of the number of multi-path. Then, the BS compares the difference of the scaled large-scale path loss between any two users. If the difference is sufficiently large, the BS notifies all users to permit their simultaneous message transmissions. However, if the difference is small, the BS notifies each user to slightly adjust their transmission power and then permits their simultaneous message transmissions as well. In the second phase, all users simultaneously transmit their messages using the same radio resource. Then the BS selects one predetermined constellation point from the received pilot symbols as the input of clustering algorithms. According to the clustering results, the BS classifies each multi-path into a specific user. The key intuition of our approach is that the clustering algorithms can generate multiple cluster centroids and each cluster centroid represents the average reception power of each user. Moreover, we use a weighted ensemble clustering algorithm to further improve the performance of our approach. We implemented our approach and conducted extensive performance comparison. Our experimental results show that, when the received signal-to-noise ratio (SNR) is more than 13 dB, our approach with the weighted ensemble clustering algorithm can correctly classify all multi-path to the corresponding user and the output SNR can be improved by 3.2 dB, where we consider three users in an 8PSK system and each user possess 50 multi-path. Ning Xie 0007, Le Ou-Yang, Alex X. Liu |
IEEE/ACM Trans. Netw. | 2 |
| 2019 | Improve L2-normalized Softmax with Exponential Moving AverageabstractIn this paper, we propose an effective training method to improve the performance of L2-normalized softmax for convolutional neural networks. Recent studies of deep learning show that by L2-normalizing the input features of softmax, the accuracy of CNN can be increased. Several works proposed novel loss functions based on the L2-normalized softmax. A common property shared by these modified normalized softmax models is that an extra set of parameters is introduced as the class centers. Although the physical meaning of this parameter is clear, few attentions have been paid to how to learn these class centers, which limits further improvement. In this paper, we address the problem of learning the class centers in the L2-normalized softmax. By treating the CNN training process as a time series, we propose a novel learning algorithm that combines the generally used gradient descent with the exponential moving average. Extensive experiments show that our model not only achieves better performance but also has a higher tolerance to the imbalance data. Xuefei Zhe, Le Ou-Yang, Hong Yan 0001 |
IJCNN | 2 |
| 2019 | DiffNetFDR: differential network analysis with false discovery rate controlabstractSUMMARY: To identify biological network rewiring under different conditions, we develop a user-friendly R package, named DiffNetFDR, to implement two methods developed for testing the difference in different Gaussian graphical models. Compared to existing tools, our methods have the following features: (i) they are based on Gaussian graphical models which can capture the changes of conditional dependencies; (ii) they determine the tuning parameters in a data-driven manner; (iii) they take a multiple testing procedure to control the overall false discovery rate; and (iv) our approach defines the differential network based on partial correlation coefficients so that the spurious differential edges caused by the variants of conditional variances can be excluded. We also develop a Shiny application to provide easier analysis and visualization. Simulation studies are conducted to evaluate the performance of our methods. We also apply our methods to two real gene expression datasets. The effectiveness of our methods is validated by the biological significance of the identified differential networks. AVAILABILITY AND IMPLEMENTATION: R package and Shiny app are available at https://github.com/Zhangxf-ccnu/DiffNetFDR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiao-Fei Zhang, Le Ou-Yang, Xiaohua Hu 0001, Hong Yan 0001 |
Bioinform. | 2 |
| 2019 | EnImpute: imputing dropout events in single-cell RNA-sequencing data via ensemble learningabstractSUMMARY: Imputation of dropout events that may mislead downstream analyses is a key step in analyzing single-cell RNA-sequencing (scRNA-seq) data. We develop EnImpute, an R package that introduces an ensemble learning method for imputing dropout events in scRNA-seq data. EnImpute combines the results obtained from multiple imputation methods to generate a more accurate result. A Shiny application is developed to provide easier implementation and visualization. Experiment results show that EnImpute outperforms the individual state-of-the-art methods in almost all situations. EnImpute is useful for correcting the noisy scRNA-seq data before performing downstream analysis. AVAILABILITY AND IMPLEMENTATION: The R package and Shiny application are available through Github at https://github.com/Zhangxf-ccnu/EnImpute. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiao-Fei Zhang, Le Ou-Yang, Xing-Ming Zhao, Xiaohua Hu 0001, Hong Yan 0001 |
Bioinform. | 2 |
| 2019 | Predicting synthetic lethal interactions in human cancers using graph regularized self-representative matrix factorizationabstractBACKGROUND: Synthetic lethality has attracted a lot of attentions in cancer therapeutics due to its utility in identifying new anticancer drug targets. Identifying synthetic lethal (SL) interactions is the key step towards the exploration of synthetic lethality in cancer treatment. However, biological experiments are faced with many challenges when identifying synthetic lethal interactions. Thus, it is necessary to develop computational methods which could serve as useful complements to biological experiments. RESULTS: In this paper, we propose a novel graph regularized self-representative matrix factorization (GRSMF) algorithm for synthetic lethal interaction prediction. GRSMF first learns the self-representations from the known SL interactions and further integrates the functional similarities among genes derived from Gene Ontology (GO). It can then effectively predict potential SL interactions by leveraging the information provided by known SL interactions and functional annotations of genes. Extensive experiments on the synthetic lethal interaction data downloaded from SynLethDB database demonstrate the superiority of our GRSMF in predicting potential synthetic lethal interactions, compared with other competing methods. Moreover, case studies of novel interactions are conducted in this paper for further evaluating the effectiveness of GRSMF in synthetic lethal interaction prediction. CONCLUSIONS: In this paper, we demonstrate that by adaptively exploiting the self-representation of original SL interaction data, and utilizing functional similarities among genes to enhance the learning of self-representation matrix, our GRSMF could predict potential SL interactions more accurately than other state-of-the-art SL interaction prediction methods. Jiang Huang, Min Wu 0008, Le Ou-Yang, Zexuan Zhu 0001 |
BMC Bioinform. | 4 |
| 2019 | Inferring Gene Network Rewiring by Combining Gene Expression and Gene Mutation DataabstractGene dependency networks often undergo changes with respect to different disease states. Understanding how these networks rewire between two different disease states is an important task in genomic research. Although many computational methods have been proposed to undertake this task via differential network analysis, most of them are designed for a predefined data type. With the development of the high throughput technologies, gene activity measurements can be collected from different aspects (e.g., mRNA expression and DNA mutation). These different data types might share some common characteristics and include certain unique properties of data type. New methods are needed to explore the similarity and difference between differential networks estimated from different data types. In this study, we develop a new differential network inference model which identifies gene network rewiring by combining gene expression and gene mutation data. Similarities and differences between different data types are learned via a group bridge penalty function. Simulation studies have demonstrated that our method consistently outperforms the competing methods. We also apply our method to identify gene network rewiring associated with ovarian cancer platinum resistance from The Cancer Genome Atlas data. There are certain differential edges common to both data types and some differential edges unique to individual data types. Hub genes in the differential networks inferred by our method play important roles in ovarian cancer drug resistance. Jia-Juan Tu, Le Ou-Yang, Xiaohua Hu 0001, Xiao-Fei Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | Joint Learning of Multiple Differential Networks With Latent VariablesabstractGraphical models have been widely used to learn the conditional dependence structures among random variables. In many controlled experiments, such as the studies of disease or drug effectiveness, learning the structural changes of graphical models under two different conditions is of great importance. However, most existing graphical models are developed for estimating a single graph and based on a tacit assumption that there is no missing relevant variables, which wastes the common information provided by multiple heterogeneous data sets and underestimates the influence of latent/unobserved relevant variables. In this paper, we propose a joint differential network analysis (JDNA) model to jointly estimate multiple differential networks with latent variables from multiple data sets. The JDNA model is built on a penalized D-trace loss function, with group lasso or generalized fused lasso penalties. We implement a proximal gradient-based alternating direction method of multipliers to tackle the corresponding convex optimization problems. Extensive simulation experiments demonstrate that JDNA model outperforms state-of-the-art methods in estimating the structural changes of graphical models. Moreover, a series of experiments on several real-world data sets have been performed and experiment results consistently show that our proposed JDNA model is effective in identifying differential networks under different conditions. Le Ou-Yang, Xiao-Fei Zhang, Xing-Ming Zhao, Debby Dan Wang, Fu Lee Wang, Bai Ying Lei, Hong Yan 0001 |
IEEE Trans. Cybern. | 1 |
| 2018 | RepLong: de novo repeat identification using long read sequencing dataabstractMotivation: The identification of repetitive elements is important in genome assembly and phylogenetic analyses. The existing de novo repeat identification methods exploiting the use of short reads are impotent in identifying long repeats. Since long reads are more likely to cover repeat regions completely, using long reads is more favorable for recognizing long repeats. Results: In this study, we propose a novel de novo repeat elements identification method namely RepLong based on PacBio long reads. Given that the reads mapped to the repeat regions are highly overlapped with each other, the identification of repeat elements is equivalent to the discovery of consensus overlaps between reads, which can be further cast into a community detection problem in the network of read overlaps. In RepLong, we first construct a network of read overlaps based on pair-wise alignment of the reads, where each vertex indicates a read and an edge indicates a substantial overlap between the corresponding two reads. Secondly, the communities whose intra connectivity is greater than the inter connectivity are extracted based on network modularity optimization. Finally, representative reads in each community are extracted to form the repeat library. Comparison studies on Drosophila melanogaster and human long read sequencing data with genome-based and short-read-based methods demonstrate the efficiency of RepLong in identifying long repeats. RepLong can handle lower coverage data and serve as a complementary solution to the existing methods to promote the repeat identification performance on long-read sequencing data. Availability and implementation: The software of RepLong is freely available at https://github.com/ruiguo-bio/replong. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Rui Guo 0012, Yan-Ran Li 0001, Shan He 0001, Le Ou-Yang, Zexuan Zhu 0001 |
Bioinform. | 4 |
| 2018 | DiffGraph: an R package for identifying gene network rewiring using differential graphical modelsabstractSummary: We develop DiffGraph, an R package that integrates four influential differential graphical models for identifying gene network rewiring under two different conditions from gene expression data. The input and output of different models are packaged in the same format, making it convenient for users to compare different models using a wide range of datasets and carry out follow-up analysis. Furthermore, the inferred differential networks can be visualized both non-interactively and interactively. The package is useful for identifying gene network rewiring from input datasets, comparing the predictions of different methods and visualizing the results. Availability and implementation: The package is available at https://github.com/Zhangxf-ccnu/DiffGraph. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Xiao-Fei Zhang, Le Ou-Yang, Xiaohua Hu 0001, Hong Yan 0001 |
Bioinform. | 2 |
| 2018 | Identifying Gene Network Rewiring by Integrating Gene Expression and Gene Network DataabstractExploring the rewiring pattern of gene regulatory networks between different pathological states is an important task in bioinformatics. Although a number of computational approaches have been developed to infer differential networks from high-throughput data, most of them only focus on gene expression data. The valuable static gene regulatory network data accumulated in recent biomedical researches are neglected. In this study, we propose a new Gaussian graphical model-based method to infer differential networks by integrating gene expression and static gene regulatory network data. We first evaluate the empirical performance of our method by comparing with the state-of-the-art methods using simulation data. We also apply our method to The Cancer Genome Atlas data to identify gene network rewiring between ovarian cancers with different platinum responses, and rewiring between breast cancers of luminal A subtype and basal-like subtype. Hub genes in the estimated differential networks rediscover known genes associated with platinum resistance in ovarian cancer and signatures of the breast cancer intrinsic subtypes. Le Ou-Yang, Xiaohua Hu 0001, Xiao-Fei Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | Incorporating prior information into differential network analysis using non-paranormal graphical modelsabstractMOTIVATION: Understanding how gene regulatory networks change under different cellular states is important for revealing insights into network dynamics. Gaussian graphical models, which assume that the data follow a joint normal distribution, have been used recently to infer differential networks. However, the distributions of the omics data are non-normal in general. Furthermore, although much biological knowledge (or prior information) has been accumulated, most existing methods ignore the valuable prior information. Therefore, new statistical methods are needed to relax the normality assumption and make full use of prior information. RESULTS: We propose a new differential network analysis method to address the above challenges. Instead of using Gaussian graphical models, we employ a non-paranormal graphical model that can relax the normality assumption. We develop a principled model to take into account the following prior information: (i) a differential edge less likely exists between two genes that do not participate together in the same pathway; (ii) changes in the networks are driven by certain regulator genes that are perturbed across different cellular states and (iii) the differential networks estimated from multi-view gene expression data likely share common structures. Simulation studies demonstrate that our method outperforms other graphical model-based algorithms. We apply our method to identify the differential networks between platinum-sensitive and platinum-resistant ovarian tumors, and the differential networks between the proneural and mesenchymal subtypes of glioblastoma. Hub nodes in the estimated differential networks rediscover known cancer-related regulator genes and contain interesting predictions. AVAILABILITY AND IMPLEMENTATION: The source code is at https://github.com/Zhangxf-ccnu/pDNA. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiao-Fei Zhang, Le Ou-Yang, Hong Yan 0001 |
Bioinform. | 2 |
| 2017 | A multi-network clustering method for detecting protein complexes from multiple heterogeneous networksabstractBACKGROUND: The accurate identification of protein complexes is important for the understanding of cellular organization. Up to now, computational methods for protein complex detection are mostly focus on mining clusters from protein-protein interaction (PPI) networks. However, PPI data collected by high-throughput experimental techniques are known to be quite noisy. It is hard to achieve reliable prediction results by simply applying computational methods on PPI data. Behind protein interactions, there are protein domains that interact with each other. Therefore, based on domain-protein associations, the joint analysis of PPIs and domain-domain interactions (DDI) has the potential to obtain better performance in protein complex detection. As traditional computational methods are designed to detect protein complexes from a single PPI network, it is necessary to design a new algorithm that could effectively utilize the information inherent in multiple heterogeneous networks. RESULTS: In this paper, we introduce a novel multi-network clustering algorithm to detect protein complexes from multiple heterogeneous networks. Unlike existing protein complex identification algorithms that focus on the analysis of a single PPI network, our model can jointly exploit the information inherent in PPI and DDI data to achieve more reliable prediction results. Extensive experiment results on real-world data sets demonstrate that our method can predict protein complexes more accurately than other state-of-the-art protein complex identification algorithms. CONCLUSIONS: In this work, we demonstrate that the joint analysis of PPI network and DDI network can help to improve the accuracy of protein complex detection. Le Ou-Yang, Hong Yan 0001, Xiao-Fei Zhang |
BMC Bioinform. | 1 |
| 2017 | Protein Complex Detection via Effective Integration of Base Clustering Solutions and Co-Complex Affinity ScoresabstractWith the increasing availability of protein interaction data, various computational methods have been developed to predict protein complexes. However, different computational methods may have their own advantages and limitations. Ensemble clustering has thus been studied to minimize the potential bias and risk of individual methods and generate prediction results with better coverage and accuracy. In this paper, we extend the traditional ensemble clustering by taking into account the co-complex affinity scores and present an Ensemble H ierarchical Clustering framework (EnsemHC) to detect protein complexes. First, we construct co-cluster matrices by integrating the clustering results with the co-complex evidences. Second, we sum up the constructed co-cluster matrices to derive a final ensemble matrix via a novel iterative weighting scheme. Finally, we apply the hierarchical clustering to generate protein complexes from the final ensemble matrix. Experimental results demonstrate that our EnsemHC performs better than its base clustering methods and various existing integrative methods. In addition, we also observed that integrating the clusters and co-complex affinity scores from different data sources will improve the prediction performance, e.g., integrating the clusters from TAP data and co-complex affinities from binary PPI data achieved the best performance in our experiments. Min Wu 0008, Le Ou-Yang, Xiaoli Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2016 | Identifying protein complexes via multi-network clusteringabstractThe detection of protein complexes from protein-protein interaction (PPI) networks is an important step toward understanding the functional organization within cells. A great number of graph clustering algorithms have been proposed to undertake this task. Since PPI data collected by high-throughput technologies is quite noisy, simply applying graph clustering algorithms on PPI data is generally not adequate to achieve reliable prediction results. Behind protein interactions, there are protein domains that interact with each other. Jointly exploiting protein-protein interactions and domain-domain interactions (DDI) have the potential to increase the accuracy of protein complex detection. However, traditional graph clustering algorithms focus on clustering proteins within a single PPI network, and cannot make use of information inherent in other heterogeneous networks. In this paper, we proposed a novel generative model to perform multi-network clustering. Unlike previous protein complex detection algorithms that can only utilize the information within a single PPI network, our model is a flexible framework that can take into account PPIs, DDIs and domain-protein associations to achieve more consistent and reliable clustering results. Experiment results on real data demonstrate that our method performs much better than state-of-the-art protein complex detection techniques. Le Ou-Yang, Hong Yan 0001, Xiao-Fei Zhang |
BIBM | 1 |
| 2016 | A two-layer integration framework for protein complex detectionabstractBACKGROUND: Protein complexes carry out nearly all signaling and functional processes within cells. The study of protein complexes is an effective strategy to analyze cellular functions and biological processes. With the increasing availability of proteomics data, various computational methods have recently been developed to predict protein complexes. However, different computational methods are based on their own assumptions and designed to work on different data sources, and various biological screening methods have their unique experiment conditions, and are often different in scale and noise level. Therefore, a single computational method on a specific data source is generally not able to generate comprehensive and reliable prediction results. RESULTS: In this paper, we develop a novel Two-layer INtegrative Complex Detection (TINCD) model to detect protein complexes, leveraging the information from both clustering results and raw data sources. In particular, we first integrate various clustering results to construct consensus matrices for proteins to measure their overall co-complex propensity. Second, we combine these consensus matrices with the co-complex score matrix derived from Tandem Affinity Purification/Mass Spectrometry (TAP) data and obtain an integrated co-complex similarity network via an unsupervised metric fusion method. Finally, a novel graph regularized doubly stochastic matrix decomposition model is proposed to detect overlapping protein complexes from the integrated similarity network. CONCLUSIONS: Extensive experimental results demonstrate that TINCD performs much better than 21 state-of-the-art complex detection techniques, including ensemble clustering and data integration techniques. Le Ou-Yang, Min Wu 0008, Xiao-Fei Zhang, Dao-Qing Dai, Xiaoli Li 0001, Hong Yan 0001 |
BMC Bioinform. | 1 |
| 2016 | Protein complex detection based on partially shared multi-view clusteringabstractBACKGROUND: Protein complexes are the key molecular entities to perform many essential biological functions. In recent years, high-throughput experimental techniques have generated a large amount of protein interaction data. As a consequence, computational analysis of such data for protein complex detection has received increased attention in the literature. However, most existing works focus on predicting protein complexes from a single type of data, either physical interaction data or co-complex interaction data. These two types of data provide compatible and complementary information, so it is necessary to integrate them to discover the underlying structures and obtain better performance in complex detection. RESULTS: In this study, we propose a novel multi-view clustering algorithm, called the Partially Shared Multi-View Clustering model (PSMVC), to carry out such an integrated analysis. Unlike traditional multi-view learning algorithms that focus on mining either consistent or complementary information embedded in the multi-view data, PSMVC can jointly explore the shared and specific information inherent in different views. In our experiments, we compare the complexes detected by PSMVC from single data source with those detected from multiple data sources. We observe that jointly analyzing multi-view data benefits the detection of protein complexes. Furthermore, extensive experiment results demonstrate that PSMVC performs much better than 16 state-of-the-art complex detection techniques, including ensemble clustering and data integration techniques. CONCLUSIONS: In this work, we demonstrate that when integrating multiple data sources, using partially shared multi-view clustering model can help to identify protein complexes which are not readily identifiable by conventional single-view-based methods and other integrative analysis methods. All the results and source codes are available on https://github.com/Oyl-CityU/PSMVC . Le Ou-Yang, Xiao-Fei Zhang, Dao-Qing Dai, Meng-Yun Wu, Yuan Zhu 0005, Hong Yan 0001 |
BMC Bioinform. | 1 |
| 2016 | Regularized logistic regression with network-based pairwise interaction for biomarker identification in breast cancerabstractBACKGROUND: To facilitate advances in personalized medicine, it is important to detect predictive, stable and interpretable biomarkers related with different clinical characteristics. These clinical characteristics may be heterogeneous with respect to underlying interactions between genes. Usually, traditional methods just focus on detection of differentially expressed genes without taking the interactions between genes into account. Moreover, due to the typical low reproducibility of the selected biomarkers, it is difficult to give a clear biological interpretation for a specific disease. Therefore, it is necessary to design a robust biomarker identification method that can predict disease-associated interactions with high reproducibility. RESULTS: In this article, we propose a regularized logistic regression model. Different from previous methods which focus on individual genes or modules, our model takes gene pairs, which are connected in a protein-protein interaction network, into account. A line graph is constructed to represent the adjacencies between pairwise interactions. Based on this line graph, we incorporate the degree information in the model via an adaptive elastic net, which makes our model less dependent on the expression data. Experimental results on six publicly available breast cancer datasets show that our method can not only achieve competitive performance in classification, but also retain great stability in variable selection. Therefore, our model is able to identify the diagnostic and prognostic biomarkers in a more robust way. Moreover, most of the biomarkers discovered by our model have been verified in biochemical or biomedical researches. CONCLUSIONS: The proposed method shows promise in the diagnosis of disease pathogenesis with different clinical characteristics. These advances lead to more accurate and stable biomarker discovery, which can monitor the functional changes that are perturbed by diseases. Based on these predictions, researchers may be able to provide suggestions for new therapeutic approaches. Meng-Yun Wu, Xiao-Fei Zhang, Dao-Qing Dai, Le Ou-Yang, Yuan Zhu 0005, Hong Yan 0001 |
BMC Bioinform. | 4 |
| 2016 | Comparative analysis of housekeeping and tissue-specific driver nodes in human protein interaction networksabstractBACKGROUND: Several recent studies have used the Minimum Dominating Set (MDS) model to identify driver nodes, which provide the control of the underlying networks, in protein interaction networks. There may exist multiple MDS configurations in a given network, thus it is difficult to determine which one represents the real set of driver nodes. Because these previous studies only focus on static networks and ignore the contextual information on particular tissues, their findings could be insufficient or even be misleading. RESULTS: In this study, we develop a Collective-Influence-corrected Minimum Dominating Set (CI-MDS) model which takes into account the collective influence of proteins. By integrating molecular expression profiles and static protein interactions, 16 tissue-specific networks are established as well. We then apply the CI-MDS model to each tissue-specific network to detect MDS proteins. It generates almost the same MDSs when it is solved using different optimization algorithms. In addition, we classify MDS proteins into Tissue-Specific MDS (TS-MDS) proteins and HouseKeeping MDS (HK-MDS) proteins based on the number of tissues in which they are expressed and identified as MDS proteins. Notably, we find that TS-MDS proteins and HK-MDS proteins have significantly different topological and functional properties. HK-MDS proteins are more central in protein interaction networks, associated with more functions, evolving more slowly and subjected to a greater number of post-translational modifications than TS-MDS proteins. Unlike TS-MDS proteins, HK-MDS proteins significantly correspond to essential genes, ageing genes, virus-targeted proteins, transcription factors and protein kinases. Moreover, we find that besides HK-MDS proteins, many TS-MDS proteins are also linked to disease related genes, suggesting the tissue specificity of human diseases. Furthermore, functional enrichment analysis reveals that HK-MDS proteins carry out universally necessary biological processes and TS-MDS proteins usually involve in tissue-dependent functions. CONCLUSIONS: Our study uncovers key features of TS-MDS proteins and HK-MDS proteins, and is a step forward towards a better understanding of the controllability of human interactomes. Xiao-Fei Zhang, Le Ou-Yang, Dao-Qing Dai, Meng-Yun Wu, Yuan Zhu 0005, Hong Yan 0001 |
BMC Bioinform. | 2 |
| 2015 | Determining minimum set of driver nodes in protein-protein interaction networksabstractBACKGROUND: Recently, several studies have drawn attention to the determination of a minimum set of driver proteins that are important for the control of the underlying protein-protein interaction (PPI) networks. In general, the minimum dominating set (MDS) model is widely adopted. However, because the MDS model does not generate a unique MDS configuration, multiple different MDSs would be generated when using different optimization algorithms. Therefore, among these MDSs, it is difficult to find out the one that represents the true driver set of proteins. RESULTS: To address this problem, we develop a centrality-corrected minimum dominating set (CC-MDS) model which includes heterogeneity in degree and betweenness centralities of proteins. Both the MDS model and the CC-MDS model are applied on three human PPI networks. Unlike the MDS model, the CC-MDS model generates almost the same sets of driver proteins when we implement it using different optimization algorithms. The CC-MDS model targets more high-degree and high-betweenness proteins than the uncorrected counterpart. The more central position allows CC-MDS proteins to be more important in maintaining the overall network connectivity than MDS proteins. To indicate the functional significance, we find that CC-MDS proteins are involved in, on average, more protein complexes and GO annotations than MDS proteins. We also find that more essential genes, aging genes, disease-associated genes and virus-targeted genes appear in CC-MDS proteins than in MDS proteins. As for the involvement in regulatory functions, the sets of CC-MDS proteins show much stronger enrichment of transcription factors and protein kinases. The results about topological and functional significance demonstrate that the CC-MDS model can capture more driver proteins than the MDS model. CONCLUSIONS: Based on the results obtained, the CC-MDS model presents to be a powerful tool for the determination of driver proteins that can control the underlying PPI networks. The software described in this paper and the datasets used are available at https://github.com/Zhangxf-ccnu/CC-MDS . Xiao-Fei Zhang, Le Ou-Yang, Yuan Zhu 0005, Meng-Yun Wu, Dao-Qing Dai |
BMC Bioinform. | 2 |
| 2015 | Detecting Protein Complexes from Signed Protein-Protein Interaction NetworksabstractIdentification of protein complexes is fundamental for understanding the cellular functional organization. With the accumulation of physical protein-protein interaction (PPI) data, computational detection of protein complexes from available PPI networks has drawn a lot of attentions. While most of the existing protein complex detection algorithms focus on analyzing the physical protein-protein interaction network, none of them take into account the "signs" (i.e., activation-inhibition relationships) of physical interactions. As the "signs" of interactions reflect the way proteins communicate, considering the "signs" of interactions can not only increase the accuracy of protein complex identification, but also deepen our understanding of the mechanisms of cell functions. In this study, we proposed a novel Signed Graph regularized Nonnegative Matrix Factorization (SGNMF) model to identify protein complexes from signed PPI networks. In our experiments, we compared the results collected by our model on signed PPI networks with those predicted by the state-of-the-art complex detection techniques on the original unsigned PPI networks. We observed that considering the "signs" of interactions significantly benefits the detection of protein complexes. Furthermore, based on the predicted complexes, we predicted a set of signed complex-complex interactions for each dataset, which provides a novel insight of the higher level organization of the cell. All the experimental results and codes can be downloaded from http://mail.sysu.edu.cn/home/[email protected]/dai/others/SGNMF.zip. Le Ou-Yang, Dao-Qing Dai, Xiao-Fei Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2014 | Detecting temporal protein complexes from dynamic protein-protein interaction networksabstractBACKGROUND: Proteins dynamically interact with each other to perform their biological functions. The dynamic operations of protein interaction networks (PPI) are also reflected in the dynamic formations of protein complexes. Existing protein complex detection algorithms usually overlook the inherent temporal nature of protein interactions within PPI networks. Systematically analyzing the temporal protein complexes can not only improve the accuracy of protein complex detection, but also strengthen our biological knowledge on the dynamic protein assembly processes for cellular organization. RESULTS: In this study, we propose a novel computational method to predict temporal protein complexes. Particularly, we first construct a series of dynamic PPI networks by joint analysis of time-course gene expression data and protein interaction data. Then a Time Smooth Overlapping Complex Detection model (TS-OCD) has been proposed to detect temporal protein complexes from these dynamic PPI networks. TS-OCD can naturally capture the smoothness of networks between consecutive time points and detect overlapping protein complexes at each time point. Finally, a nonnegative matrix factorization based algorithm is introduced to merge those very similar temporal complexes across different time points. CONCLUSIONS: Extensive experimental results demonstrate the proposed method is very effective in detecting temporal protein complexes than the state-of-the-art complex detection techniques. Le Ou-Yang, Dao-Qing Dai, Xiaoli Li 0001, Min Wu 0008, Xiao-Fei Zhang |
BMC Bioinform. | 1 |
| 2014 | Detecting overlapping protein complexes based on a generative model with functional and topological propertiesabstractBACKGROUND: Identification of protein complexes can help us get a better understanding of cellular mechanism. With the increasing availability of large-scale protein-protein interaction (PPI) data, numerous computational approaches have been proposed to detect complexes from the PPI networks. However, most of the current approaches do not consider overlaps among complexes or functional annotation information of individual proteins. Therefore, they might not be able to reflect the biological reality faithfully or make full use of the available domain-specific knowledge. RESULTS: In this paper, we develop a Generative Model with Functional and Topological Properties (GMFTP) to describe the generative processes of the PPI network and the functional profile. The model provides a working mechanism for capturing the interaction structures and the functional patterns of proteins. By combining the functional and topological properties, we formulate the problem of identifying protein complexes as that of detecting a group of proteins which frequently interact with each other in the PPI network and have similar annotation patterns in the functional profile. Using the idea of link communities, our method naturally deals with overlaps among complexes. The benefits brought by the functional properties are demonstrated by real data analysis. The results evaluated using four criteria with respect to two gold standards show that GMFTP has a competitive performance over the state-of-the-art approaches. The effectiveness of detecting overlapping complexes is also demonstrated by analyzing the topological and functional features of multi- and mono-group proteins. CONCLUSIONS: Based on the results obtained in this study, GMFTP presents to be a powerful approach for the identification of overlapping protein complexes using both the PPI network and the functional profile. The software can be downloaded from http://mail.sysu.edu.cn/home/[email protected]/dai/others/GMFTP.zip. Xiao-Fei Zhang, Dao-Qing Dai, Le Ou-Yang, Hong Yan 0001 |
BMC Bioinform. | 3 |