Xiaoke Ma 0001

dblp:81/9706-1 · DBLP profile ↗
← Back
76ranked-venue papers
13as first author
61since 2021 · last 2026
0000-0002-5604-7137ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 38 · 8 first-author · 29 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 20 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SA2E: spatial-aware auto-encoder for cell type deconvolution of spatial transcriptomics data
abstract
MOTIVATION: Spatial transcriptomics (ST) technologies measure gene expression together with spatial locations, but each spot typically contains a mixture of cell types, posing a challenge for downstream analysis. Cell-type deconvolution aims to infer spot-wise cell-type proportions by integrating single-cell RNA-seq (scRNA-seq) and ST data. Many existing methods construct cell-type signatures from predefined marker genes, which can limit performance when marker information is incomplete or unavailable. RESULTS: To address this limitation, we propose a spatial-aware auto-encoder framework (SA2E) for cell-type deconvolution without requiring predefined cell-type biomarkers. SA2E learns latent spot representations using a spatially regularized auto-encoder that preserves the local topology of the spot spatial graph. Based on these representations, SA2E learns cell-type signatures by enforcing them to reconstruct ST expression. In our framework, simulated ST data with known proportions are used for supervised pretraining, while real ST data are optimized using the reconstruction objective. Extensive experiments on simulated and real ST datasets demonstrate that SA2E outperforms state-of-the-art deconvolution baselines. AVAILABILITY AND IMPLEMENTATION: The code of SA2E is available at Github (https://github.com/xkmaxidian/SA2E) and Zenodo (DOI: 10.5281/zenodo.18765467).
Yaxiong Ma, Zengfa Dou, Yuhong Zha, Xiaoke Ma 0001
Bioinform.4
2026 Multi-view spectral clustering via high-frequency truncation and graph wavelet modulation
Shouye Ke, Guohua Huang, Xiaoke Ma 0001
Knowl. Based Syst.4
2026 Dual-branch Bayesian adversarial learning for multi-view clustering
Guohua Huang, Xiaoke Ma 0001
Knowl. Based Syst.4
2026 Network models for bridging denoising and identifying spatial domains of spatially resolved transcriptomics
abstract
Spatially resolved transcriptomics (SRT) enables the simultaneous capture of gene expression profiles and spatial localization, providing valuable insights into tissue architecture. However, the preservation of spatial information requires additional experimental procedures, which often introduce substantial technical noise. Existing methods typically perform denoising and spatial domain identification in separate steps, leading to suboptimal performance and limiting their applicability. To address this limitation, we propose an integrative network model, stACN ( spatial transcriptomics Attribute Cell Network), that jointly denoises gene expression data and identifies spatial domains in SRT. Specifically, stACN first learns clean dual cell networks using a graph noise model, and then derives compatible cell features through joint tensor decomposition of the denoised networks. Experimental results demonstrate that stACN effectively enhances data quality, as measured by clustering agreement with reference annotations (Adjusted Rand Index, ARI), and facilitates spatial domain analysis in SRT datasets.
Haiyue Wang, Wensheng Zhang 0002, Zaiyi Liu, Xiaoke Ma 0001
PLoS Comput. Biol.4
2026 Robust Algorithm With Contrastive Learning for Identifying Spatial Domains From Noised Spatial Transcriptomics Data
abstract
Spatial transcriptomics (ST) technologies capture transcriptomics of genens with spatial context, enabling systematic exploration of micro-environment of tissues that is highly associated with spatial domains. And, noise of ST data poses a great challenge on designing algorithms for identifying spatial domains, whereas available methods remove noise by employing pre-processing procedure, resulting in the undesirable performance. To overcome this limitation, we propose a robust and joint framework, called jNFACL (joint Network-based Feature-Affinity Contrastive Learning), for identifying spatial domains of noised ST data, where deniosing of ST data and identifying spatial domains are simultaneously integrated. Specifically, jNFACL first constructs expression and spatial graphs with transcriptomics and spatial coordinates of spots, which removes heterogeneity of ST data. And, jNFACL separates noise of ST data by jointly projecting these constructed graphs into the shared subspace, where noise of ST data is separated from feature level with nonnegative matrix factorization. To further enhance quality of features of spots, contrastive learning is adopted to leverages spatial neighborhoods by pulling similar spots together and pushing dissimilar spots apart, where self-supervision information is incorporated, thereby improving the characterization and identification of spatial domains. Experiments on various datasets with different noise levels from multiple platforms and species demonstrate that jNFACL is much more accurate and robust than state-of-the-art methods, providing alternatives for analyzing noised ST data.
Yaxiong Ma, Peifeng Liang, Xiaoke Ma 0001
IEEE Trans. Comput. Biol. Bioinform.5
2026 MPT-MIL: Multimodal Aware Prompt Tuning for Prediction of Cancer Survival
abstract
As a critical statistical technique in oncology, survival prediction is used to estimate the probability of survival or time-to-event outcomes. Identifying survival-related factors from pathology and genomic data is a key approach for analyzing survival outcomes. However, current methods face several challenges, such as the suboptimal adaptation of pre-trained vision foundation models to specific tasks during feature extraction from whole slide images (WSIs), and the fact that many pathology-based models fail to integrate repetitive gene expression information during pre-training. In this study, we propose a plug-and-play multiple instance learning (MIL)-based foundation model tuning strategy to adapt vision foundation models for downstream tasks and incorporate knowledge from genomic data. Specifically, we introduce Task-specific Instance Selection, which utilizes zero-shot learning to efficiently select task-relevant WSI regions, improving tuning efficiency and reducing interference from irrelevant tissue areas. Additionally, we develop a multi-model prompt token for model fine-tuning, which integrates genetic information into the prompt-tuning process and transfers new modality information to pre-trained vision foundation models. To further enhance the model's ability to learn genetic information during fine-tuning, we introduce a Gene Distribution Aware Task as an auxiliary task to the traditional survival task. This auxiliary task helps the model better perceive multimodal information. Extensive experimental results on three public TCGA datasets demonstrate that our model outperforms all previous MIL-based methodologies and fine-tuning approaches in terms of performance.
Ruofan Zhang, Mengjie Fang, Zipei Wang, Xuebin Xie, Xiaoke Ma 0001, Jie Tian 0001, Di Dong
IEEE J. Biomed. Health Informatics6
2025 FissionVAE: Federated Non-IID Image Generation with Latent Space and Decoder Decomposition
abstract
Federated learning is a machine learning paradigm that enables decentralized clients to collaboratively learn a shared model while keeping all the training data local. While considerable research has focused on federated image generation, particularly Generative Adversarial Networks, Variational Autoencoders have received less attention. In this paper, we address the challenges of non-IID (independently and identically distributed) data environments featuring multiple groups of images of different types. Non-IID data distributions can lead to difficulties in maintaining a consistent latent space and can also result in local generators with disparate texture features being blended during aggregation. We thereby introduce FissionVAE that decouples the latent space and constructs decoder branches tailored to individual client groups. This method allows for customized learning that aligns with the unique data distributions of each group. Additionally, we incorporate hierarchical VAEs and demonstrate the use of heterogeneous decoder architectures within FissionVAE. We also explore strategies for setting the latent prior distributions to enhance the decoupling process. To evaluate our approach, we assemble two composite datasets: the first combines MNIST and FashionMNIST; the second comprises RGB datasets of cartoon and human faces, wild animals, marine vessels, and remote sensing images. Our experiments demonstrate that FissionVAE greatly improves generation quality on these datasets compared to baseline federated VAE models.
Hanchi Ren, Jingjing Deng 0001, Xianghua Xie, Xiaoke Ma 0001
IJCAI5
2025 Denoising spatially resolved transcriptomics with consistency of heterogeneous spatial coordinates, transcription, and morphology
abstract
Spatially resolved transcriptomics (SRT) simultaneously captures spatial coordinates, pathological features, and transcriptional profiles of cells within intact tissues, offering unprecedented opportunities to explore tissue architecture. However, SRT data often suffer from substantial technical noise introduced by experimental procedures, posing challenges for downstream analyses. To overcome these challenges, we introduce a Multiview Denoising framework for Spatial Transcriptomics (MvDST), which integrates a deep autoencoder and self-supervised learning to jointly reconstruct expression profiles, denoise features, and enforce cross-view consistency, effectively reducing technical noise, and heterogeneity. As a result, MvDST reliably and accurately delineates tissue subgroups across simulated datasets under various perturbations. In real cancer datasets, it distinguishes tumor-associated domains, identifies region-specific marker genes, and reveals intra-tumoral heterogeneity. Furthermore, we validate the robustness of MvDST across multiple spatial transcriptomics platforms, including 10 $\times $ Visium, STARmap, and osmFISH. Overall, these results demonstrate that MvDST can serve as a crucial initial step for the analysis of spatially resolved transcriptomics data.
Haiyue Wang, Shaoqing Feng, Xiaoke Ma 0001
Briefings Bioinform.4
2025 Enhancing and accelerating cell type deconvolution of large-scale spatial transcriptomics slices with dual network model
abstract
MOTIVATION: Cell type deconvolution deciphers spatial distribution of mRNA transcripts at single cell level by integrating single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics data to infer mixture of cell types of spots in slices. Current algorithms are criticized for neglecting connection between scRNA-seq and spatial transcriptomics data, as well as time-consuming, hampering their application to large-scale datasets. RESULTS: In this study, we propose a joint learning nonnegative matrix factorization algorithm for fast cell type deconvolution (aka jMF2D), which integrates scRNA-seq and spatial transcriptomics data with network models. To bridge scRNA-seq and spatial transcriptomics data, jMF2D jointly learns cell type similarity network to enhance quality of signatures of cell types, thereby promoting accuracy and efficiency of deconvolution. Experiments demonstrate that jMF2D outperforms state-of-the-art baselines in terms of accuracy by saving about 90% running time on various datasets generated by different platforms. Furthermore, it can also facilitates the identification of spatial domains and bio-marker genes, providing an efficient and effective model for analyzing spatial transcriptomics data. AVAILABILITY AND IMPLEMENTATION: The software is coded using python, and is free available for academic https://github.com/xkmaxidian/jMF2D.
Yuhong Zha, Shaoqing Feng, Quan Zou 0001, Xiaoke Ma 0001
Bioinform.5
2025 Sparse representation for restoring images by exploiting topological structure of graph of patches
abstract
Abstract Image restoration poses a significant challenge, aiming to accurately recover damaged images by delving into their inherent characteristics. Various models and algorithms have been explored by researchers to address different types of image distortions, including sparse representation, grouped sparse representation, and low‐rank self‐representation. The grouped sparse representation algorithm leverages the prior knowledge of non‐local self‐similarity and imposes sparsity constraints to maintain texture information within images. To further exploit the intrinsic properties of images, this study proposes a novel low‐rank representation‐guided grouped sparse representation image restoration algorithm. This algorithm integrates self‐representation models and trace optimization techniques to effectively preserve the original image structure, thereby enhancing image restoration performance while retaining the original texture and structural information. The proposed method was evaluated on image denoising and deblocking tasks across several datasets, demonstrating promising results.
Yaxian Gao, Zhaoyuan Cai, Xianghua Xie, Jingjing Deng 0001, Zengfa Dou, Xiaoke Ma 0001
IET Image Process.6
2025 One-step Multi-view Spectral Clustering with Subspaces Fusion on Grassmann manifold
Zengfa Dou, Haodong Ren, Yaxiong Ma, Guohua Huang, Xiaoke Ma 0001
Neurocomputing6
2025 Learning multi-level topology representation for multi-view clustering with deep non-negative matrix factorization
Zengfa Dou, Weiming Hou, Xianghua Xie, Xiaoke Ma 0001
Neural Networks5
2025 Self-Supervised Graph Embedding Clustering
abstract
Manifold learning and $K$K-means are two powerful techniques for data analysis in the field of artificial intelligence. When used for label learning, a promising strategy is to combine them directly and optimize both models simultaneously. However, a significant drawback of this approach is that it represents a naive and crude integration, requiring the optimization of all variables in both models without achieving a truly essential combination. Additionally, it introduces an extra hyperparameter and cannot ensure cluster balance. These challenges motivate us to explore whether a meaningful integration can be developed for dimensionality reduction clustering. In this paper, we propose a novel self-supervised manifold clustering framework that reformulates the two models into a unified framework, eliminating the need for additional hyperparameters while achieving dimensionality reduction clustering. Specifically, by analyzing the relationship between $K$K-means and manifold learning, we construct a meaningful low-dimensional manifold clustering model that directly produces the label matrix of the data. The label information is then used to guide the learning of the manifold structure, ensuring consistency between the manifold structure and the labels. Notably, we identify a valuable role of ${\ell _{2,p}}$ℓ2,p-norm regularization in clustering: maximizing the ${\ell _{2,p}}$ℓ2,p-norm naturally maintains class balance during clustering, and we provide a theoretical proof of this property. Extensive experimental results demonstrate the efficiency of our proposed model.
Fangfang Li 0005, Quanxue Gao, Xiaoke Ma 0001, Ming Yang 0024, Cheng Deng 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Group link prediction in bipartite graphs with graph neural networks
Shijie Luo 0001, He Li 0006, Xiaoke Ma 0001, Jiangtao Cui, Shaojie Qiao, Jae Soo Yoo
Pattern Recognit.4
2025 Local High-Order Graph Learning for Multi-View Clustering
abstract
As the accumulation of multi-view data continues to grow, multi-view clustering has become increasingly important in research fields like data mining. However, current methods have been criticized for their unsatisfactory performance, such as insufficient exploration of intra-view high-order relationships and poor characterization of inter-view diverse features. To overcome these challenges, we propose a novel approach called Local High-order Graph Learning for Multi-View Clustering (LHGL_MVC). Our method aims to explore high-order relationships within a view while also considering diverse information between views. In LHGL_MVC, we learn the initial graphs of each view through self-representation, which are decomposed into consistent and diverse parts to better capture the diversity of different views. Based on consistent parts, we propose a novel local high-order graph learning approach to more effectively explore high-order relationships between samples within each view. At the same time, we leverage high-order relationships between views using the rotated tensor nuclear norm. Finally, we obtain a unified graph for clustering by fusing all consistent affinity graphs and their high-order graphs with adaptive weights. All procedures are integrated into an overall objective function, which mutually promotes during the optimization process. The comprehensive experiments conducted on eleven real-world datasets demonstrate that LHGL_MVC significantly outperforms existing algorithms in various measurements, highlighting the superiority of the proposed method.
Qiang Lin 0001, Yaxiong Ma, Xiaoke Ma 0001
IEEE Trans. Big Data4
2025 Identifying Spatial Domains From Spatial Multi-Omics Data With Graph Mutual Information and Deep Subspace Learning
abstract
Spatial omics technologies enable the measurement of multiple molecular characterizations from the same tissue section while preserving spatial information, providing unprecedented opportunities to elucidate the relationship between cellular localization and tissue function. Spatial domain identification, which segments intact tissues into functionally distinct regions, is a fundamental task in spatial omics analysis. However, existing approaches are often limited to single-omics data or neglect spatial context, facing substantial limitations when extended to spatial multi-omics data. In this paper, we propose SIMID (Spatial domain Identification via graph Mutual Information and Deep subspace learning), a framework that integrates heterogeneous molecular profiles with spatial information to identify spatial domains. Specifically, a graph mutual information encoder is employed to capture cellular spatial proximity and molecular profile similarity, generating omics-specific cell embeddings for each omics layer. The deep subspace learning is then employed to construct cell network for each omics layer, converting heterogeneous multi-omics data into a homogeneous cell multi-layer network. SIMID further employs the low-rank and discriminative constraints to decompose the cell multi-layer network into consistent and complementary structures, providing an effective strategy for domain identification from spatial multi-omics data. Experimental results on both simulated and real-world spatial multi-omics datasets demonstrate that SIMID consistently outperforms existing methods and precisely reveals spatial domains from spatial multi-omics data.
Yaxiong Ma, Xiaoke Ma 0001
IEEE Trans. Comput. Biol. Bioinform.5
2024 TUT4CRS: Time-aware User-preference Tracking for Conversational Recommendation System
abstract
The Conversational Recommendation System (CRS) aims to capture user dynamic preferences and provide item recommendations based on multi-turn conversations. However, effectively modeling these dynamic preferences faces challenges due to conversational limitations, which mainly manifests as limited turns in a conversation (quantity aspect) and low compliance with queries (quality aspect). Previous studies often address these challenges in isolation, overlooking their interconnected nature. The fundamental issue underlying both problems lies in the potential abrupt changes in user preferences, to which CRS may not respond promptly. We acknowledge that user preferences are influenced by temporal factors, serving as a bridge between conversation quantity and quality. Therefore, we propose a more comprehensive CRS framework called Time-aware User-preference Tracking for Conversational Recommendation System (TUT4CRS), leveraging time dynamics to tackle both issues simultaneously. Specifically, we construct a global time interaction graph to incorporate rich external information and establish a local time-aware weight graph based on this information to adeptly select queries and effectively model user dynamic preferences. Extensive experiments on two real-world datasets validate that TUT4CRS can significantly improve recommendation performance while reducing the number of conversation turns.
Dongxiao He, Jinghan Zhang 0010, Xiaobao Wang, Meng Ge, Zhiyong Feng 0002, Longbiao Wang, Xiaoke Ma 0001
ACM Multimedia7
2024 Image restoration with group sparse representation and low-rank group residual learning
abstract
Abstract Image restoration, as a fundamental research topic of image processing, is to reconstruct the original image from degraded signal using the prior knowledge of image. Group sparse representation (GSR) is powerful for image restoration; it however often leads to undesirable sparse solutions in practice. In order to improve the quality of image restoration based on GSR, the sparsity residual model expects the representation learned from degraded images to be as close as possible to the true representation. In this article, a group residual learning based on low‐rank self‐representation is proposed to automatically estimate the true group sparse representation. It makes full use of the relation among patches and explores the subgroup structures within the same group, which makes the sparse residual model have better interpretation furthermore, results in high‐quality restored images. Extensive experimental results on two typical image restoration tasks (image denoising and deblocking) demonstrate that the proposed algorithm outperforms many other popular or state‐of‐the‐art image restoration methods.
Zhaoyuan Cai, Xianghua Xie, Jingjing Deng 0001, Zengfa Dou, Bo Tong, Xiaoke Ma 0001
IET Image Process.6
2024 FedBoosting: Federated learning with gradient protected boosting for text recognition
abstract
Conventional machine learning methodologies require the centralization of data for model training, which may be infeasible in situations where data sharing limitations are imposed due to concerns such as privacy and gradient protection. The Federated Learning (FL) framework enables the collaborative learning of a shared model without necessitating the centralization or sharing of data among the data proprietors. Nonetheless, in this paper, we demonstrate that the generalization capability of the joint model is suboptimal for Non-Independent and Non-Identically Distributed (Non-IID) data, particularly when employing the Federated Averaging (FedAvg) strategy as a result of the weight divergence phenomenon. Consequently, we present a novel boosting algorithm for FL to address both the generalization and gradient leakage challenges, as well as to facilitate accelerated convergence in gradient-based optimization. Furthermore, we introduce a secure gradient sharing protocol that incorporates Homomorphic Encryption (HE) and Differential Privacy (DP) to safeguard against gradient leakage attacks. Our empirical evaluation demonstrates that the proposed Federated Boosting (FedBoosting) technique yields significant enhancements in both prediction accuracy and computational efficiency in the visual text recognition task on publicly available benchmarks.
Hanchi Ren, Jingjing Deng 0001, Xianghua Xie, Xiaoke Ma 0001
Neurocomputing4
2024 Clustering dynamic networks by discriminating roles of vertices and capturing temporality with subsequent feature projection
Yaxiong Ma, Zengfa Dou, Guohua Huang, Xiaoke Ma 0001
Knowl. Based Syst.5
2024 Contrastive and adversarial regularized multi-level representation learning for incomplete multi-view clustering
Haiyue Wang, Wensheng Zhang 0002, Xiaoke Ma 0001
Neural Networks3
2024 Learning deep representation and discriminative features for clustering of multi-layer networks
Xiaoke Ma 0001, Quan Wang 0006, Maoguo Gong, Quanxue Gao
Neural Networks2
2024 Graph Contrastive Learning for Clustering of Multi-Layer Networks
abstract
Multi-layer networks precisely model complex systems in society and nature with various types of interactions, and identifying conserved modules that are well-connected in all layers is of great significance for revealing their structure-function relationships. Current algorithms are criticized for either ignoring the intrinsic relations among various layers, or failing to learn discriminative features. To attack these limitations, a novel graph contrastive learning framework for clustering of multi-layer networks is proposed by joining nonnegative matrix factorization and graph contrastive learning (called jNMF-GCL), where the intrinsic structure and discriminative of features are simultaneously addressed. Specifically, features of vertices are firstly learned by preserving the conserved structure in multi-layer networks with matrix factorization, and then jNMF-GCL learns an affinity structure of vertices by manipulating features of various layers. To enhance quality of features, contrastive learning is executed by selecting the positive and negative samples from the constructed affinity graph, which significantly improves discriminative of features. Finally, jNMF-GCL incorporates feature learning, construction of affinity graph, contrastive learning and clustering into an overall objective, where global and local structural information are seamlessly fused, providing a more effective way to describe structure of multi-layer networks. Extensive experiments conducted on both artificial and real-world networks have shown the superior performance of jNMF-GCL over state-of-the-art models across various metrics.
Xiaoke Ma 0001
IEEE Trans. Big Data2
2024 Learning Consistency and Specificity of Cells From Single-Cell Multi-Omic Data
abstract
Advancements in single-cell technologies concomitantly develop the epigenomic and transcriptomic profiles at the cell levels, providing opportunities to explore the potential biological mechanisms. Even though significant efforts have been dedicated to them, it remains challenging for the integration analysis of multi-omic data of single-cell because of the heterogeneity, complicated coupling and interpretability of data. To handle these issues, we propose a novel self-representation Learning-based Multi-omics data Integrative Clustering algorithm (sLMIC) for the integration of single-cell epigenomic profiles (DNA methylation or scATAC-seq) and transcriptomic (scRNA-seq), which the consistent and specific features of cells are explicitly extracted facilitating the cell clustering. Specifically, sLMIC constructs a graph for each type of single-cell data, thereby transforming omics data into multi-layer networks, which effectively removes heterogeneity of omic data. Then, sLMIC employs the low-rank and exclusivity constraints to separate the self-representation of cells into two parts, i.e., the shared and specific features, which explicitly characterize the consistency and diversity of omic data, providing an effective strategy to model the structure of cell types. Feature extraction and cell clustering are jointly formulated as an overall objective function, where latent features of data are obtained under the guidance of cell clustering. The extensive experimental results on 13 multi-omics datasets of single-cell from diverse organisms and tissues indicate that sLMIC observably exceeds the advanced algorithms regarding various measurements.
Haiyue Wang, Zaiyi Liu, Xiaoke Ma 0001
IEEE J. Biomed. Health Informatics3
2024 Noised Multi-Layer Networks Clustering With Graph Denoising and Structure Learning
abstract
Multi-layer networks treat various types of interactions at each level to model complex systems in nature and society, and clustering of them is of great significance for revealing mechanisms of systems. Vast majority of current algorithms focus on identifying the common communities in clear multi-layer networks, and few attempt has been devoted to the detection of layer-specific communities in noised ones. To address these issues, a joint learning algorithm withGraphDenoising andStructureLearning (calledGDSL) for the detection of layer-specific communities in noised multi-layer networks is proposed, which simultaneously integrates graph denoising, structure learning, and module detection. To remove noise of networks, GDSL re-constructs affinity graphs for the original ones by preserving community structure. To enhance robustness and discriminative of features, GDSL explores the relations of features among various layers with the Hilbert-Schmidt Independence Criterion and structure learning. Finally, GDSL joins all these procedures with an objective function, and deduces optimization rules. The results show that GDSL not only significantly outperforms baselines but also enhances the robustness of the algorithm, providing an effective model for community detection in noised multi-layer networks.
Wensheng Zhang 0002, Maoguo Gong, Xiaoke Ma 0001
IEEE Trans. Knowl. Data Eng.4
2023 AnomMAN: Detect anomalies on multi-view attributed networks
He Li 0006, Wanyuan Zhang, Xiaoke Ma 0001, Jiangtao Cui, Jae Soo Yoo
Inf. Sci.5
2023 Learning specific and conserved features of multi-layer networks
Xiaoke Ma 0001, Wensheng Zhang 0002, He Li 0006, Yanni Li, Jiangtao Cui
Inf. Sci.3
2023 Multi-View Clustering for Integration of Gene Expression and Methylation Data With Tensor Decomposition and Self-Representation Learning
abstract
The accumulated DNA methylation and gene expression provide a great opportunity to exploit the epigenetic patterns of genes, which is the foundation for revealing the underlying mechanisms of biological systems. Current integrative algorithms are criticized for undesirable performance because they fail to address the heterogeneity of expression and methylation data, and the intrinsic relations among them. To solve this issue, a novel multi-view clustering with self-representation learning and low-rank tensor constraint (MCSL-LTC) is proposed for the integration of gene expression and DNA methylation data, which are treated as complementary views. Specifically, MCSL-LTC first learns the low-dimensional features for each view with the linear projection, and then these features are fused in a unified tensor space with low-rank constraints. In this case, the complementary information of various views is precisely captured, where the heterogeneity of omic data is avoided, thereby enhancing the consistency of different views. Finally, MCSL-LTC obtains a consensus cluster of genes reflecting the structure and features of various views. Experimental results demonstrate that the proposed approach outperforms state-of-the-art baselines in terms of accuracy on both the social and cancer data, which provides an effective and efficient method for the integration of heterogeneous genomic data.
Weimin Hou, Zaiyi Liu, Xiaoke Ma 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 Layer-Specific Modules Detection in Cancer Multi-Layer Networks
abstract
Multi-layer networks provide an effective and efficient tool to model and characterize complex systems with multiple types of interactions, which differ greatly from the traditional single-layer networks. Graph clustering in multi-layer networks is highly non-trivial since it is difficult to balance the connectivity of clusters and the connection of various layers. The current algorithms for the layer-specific clusters are criticized for the low accuracy and sensitivity to the perturbation of networks. To overcome these issues, a novel algorithm for the layer-specific module in multi-layer networks based on nonnegative matrix factorization (LSNMF) is proposed by explicitly exploring the specific features of vertices. LSNMF first extract features of vertices in multi-layer networks by using nonnegative matrix factorization (NMF) and then decompose features of vertices into the common and specific components. The orthogonality constraint is imposed on the specific components to ensure the specificity of features of vertices, which provides a better strategy to characterize and model the structure of layer-specific modules. The extensive experiments demonstrate that the proposed algorithm dramatically outperforms state-of-the-art baselines in terms of various measurements. Furthermore, LSNMF efficiently extracts stage-specific modules, which are more likely to enrich the known functions, and also associate with the survival time of patients.
Xiaoke Ma 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2023 Network-Based Structural Learning Nonnegative Matrix Factorization Algorithm for Clustering of scRNA-Seq Data
abstract
Single-cell RNA sequencing (scRNA-seq) measures expression profiles at the single-cell level, which sheds light on revealing the heterogeneity and functional diversity among cell populations. The vast majority of current algorithms identify cell types by directly clustering transcriptional profiles, which ignore indirect relations among cells, resulting in an undesirable performance on cell type discovery and trajectory inference. Therefore, there is a critical need for inferring cell types and trajectories by exploiting the interactions among cells. In this study, we propose a network-based structural learning nonnegative matrix factorization algorithm (aka SLNMF) for the identification of cell types in scRNA-seq, which is transformed into a constrained optimization problem. SLNMF first constructs the similarity network for cells and then extracts latent features of the cells by exploiting the topological structure of the cell-cell network. To improve the clustering performance, the structural constraint is imposed on the model to learn the latent features of cells by preserving the structural information of the networks, thereby significantly improving the performance of algorithms. Finally, we track the trajectory of cells by exploring the relationships among cell types. Fourteen scRNA-seq datasets are adopted to validate the performance of algorithms with the number of single cells varying from 49 to 26,484. The experimental results demonstrate that SLNMF significantly outperforms fifteen state-of-the-art methods with 15.32% improvement in terms of accuracy, and it accurately identifies the trajectories of cells. The proposed model and methods provide an effective strategy to analyze scRNA-seq data. (The software is coded using matlab, and is freely available for academic https://github.com/xkmaxidian/SLNMF).
Xiaoke Ma 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2023 Multi-View Clustering With Graph Learning for scRNA-Seq Data
abstract
Advances in single-cell biotechnologies have generated the single-cell RNA sequencing (scRNA-seq) of gene expression profiles at cell levels, providing an opportunity to study cellular distribution. Although significant efforts developed in their analysis, many problems remain in studying cell types distribution because of the heterogeneity, high dimensionality, and noise of scRNA-seq. In this study, a multi-view clustering with graph learning algorithm (MCGL) for scRNA-seq data is proposed, which consists of multi-view learning, graph learning, and cell type clustering. In order to avoid a single feature space of scRNA-seq being inadequate to comprehensively characterize the functions of cells, MCGL constructs the multiple feature spaces and utilizes multi-view learning to comprehensively characterize scRNA-seq data from different perspectives. MCGL adaptively learns the similarity graphs of cells that overcome the dependence on fixed similarity, transforming scRNA-seq analysis into the analysis of multi-view clustering. MCGL decomposes the networks of cells into view-specific and common networks in multi-view learning, which better characterizes the topological relationship of cells. MCGL simultaneously utilizes multiple types of cell-cell networks and fully exploits the connection relationship between cells through the complementarity between networks to improve clustering performance. The graph learning, graph factorization, and cell-type clustering processes are accomplished simultaneously under one optimization framework. The performance of the MCGL algorithm is validated with ten scRNA-seq datasets from different scales, and experimental results imply that the proposed algorithm significantly outperforms fourteen state-of-the-art scRNA-seq algorithms.
Wensheng Zhang 0002, Weimin Hou, Xiaoke Ma 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 Joint Learning of Feature Extraction and Clustering for Large-Scale Temporal Networks
abstract
Temporal networks are ubiquitous in nature and society, and tracking the dynamics of networks is fundamental for investigating the mechanisms of systems. Dynamic communities in temporal networks simultaneously reflect the topology of the current snapshot (clustering accuracy) and historical ones (clustering drift). Current algorithms are criticized for their inability to characterize the dynamics of networks at the vertex level, independence of feature extraction and clustering, and high time complexity. In this study, we solve these problems by proposing a novel joint learning model for dynamic community detection in temporal networks (also known as jLMDC) via joining feature extraction and clustering. This model is formulated as a constrained optimization problem. Vertices are classified into dynamic and static groups by exploring the topological structure of temporal networks to fully exploit their dynamics at each time step. Then, jLMDC updates the features of dynamic vertices by preserving features of static ones during optimization. The advantage of jLMDC is that features are extracted under the guidance of clustering, promoting performance, and saving the running time of the algorithm. Finally, we extend jLMDC to detect the overlapping dynamic community in temporal networks. The experimental results on 11 temporal networks demonstrate that jLMDC improves accuracy up to 8.23% and saves 24.89% of running time on average compared to state-of-the-art methods.
Dongyuan Li, Xiaoke Ma 0001, Maoguo Gong
IEEE Trans. Cybern.2
2023 Clustering of Multilayer Networks Using Joint Learning Algorithm With Orthogonality and Specificity of Features
abstract
Complex systems in nature and society consist of various types of interactions, where each type of interaction belongs to a layer, resulting in the so-called multilayer networks. Identifying specific modules for each layer is of great significance for revealing the structure-function relations in multilayer networks. However, the available approaches are criticized undesirable because they fail to explicitly the specificity of modules, and balance the specificity and connectivity of modules. To overcome these drawbacks, we propose an accurate and flexible algorithm by joint learning matrix factorization and sparse representation (jMFSR) for specific modules in multilayer networks, where matrix factorization extracts features of vertices and sparse representation discovers specific modules. To exploit the discriminative latent features of vertices in multilayer networks, jMFSR incorporates linear discriminant analysis (LDA) into non-negative matrix factorization (NMF) to learn features of vertices that distinguish the categories. To explicitly measure the specificity of features, jMFSR decomposes features of vertices into common and specific parts, thereby enhancing the quality of features. Then, jMFSR jointly learns feature extraction, common-specific feature factorization, and clustering of multilayer networks. The experiments on 11 datasets indicate that jMFSR significantly outperforms state-of-the-art baselines in terms of various measurements.
Maoguo Gong, Xiaoke Ma 0001
IEEE Trans. Cybern.3
2023 DMGF-Net: An Efficient Dynamic Multi-Graph Fusion Network for Traffic Prediction
abstract
Traffic prediction is the core task of intelligent transportation system (ITS) and accurate traffic prediction can greatly improve the utilization of public resources. Dynamic interaction of multiple spatial relationships will influence the accuracy of traffic prediction. However, many existing methods only consider static spatial relationships, which restricts the accuracy of the prediction. To address the above problem, in this article, we propose the Dynamic Multi-Graph Fusion Network (DMGF-Net) to model the spatial-temporal correlations in traffic network. In the DMGF-Net, the fusion graph is designed to leverage and extract the various spatial correlations between different regions by fusing spatial graph, semantic graph, and spatial-semantic graph. Further, to dynamically learn the importance of different neighbors, we design the Dynamic Spatial-Temporal Unit (DSTU), which can adjust the aggregation weights of different neighbors by combining the convolution operation and the attention mechanism. It can selectively aggregate spatial-temporal features from different neighbors. Extensive experiments on three datasets demonstrate that effectiveness of our model, especially on PEMS08, our model achieves an increase of about 8.55% and 7.55% in terms of MAE and RMSE than the static model STGCN.
He Li 0006, Duo Jin, Xiaoke Ma 0001, Jiangtao Cui, De-Shuang Huang, Shaojie Qiao, Jae Soo Yoo
ACM Trans. Knowl. Discov. Data5
2022 Learning deep features and topological structure of cells for clustering of scRNA-sequencing data
abstract
Single-cell RNA sequencing (scRNA-seq) measures gene transcriptome at the cell level, paving the way for the identification of cell subpopulations. Although deep learning has been successfully applied to scRNA-seq data, these algorithms are criticized for the undesirable performance and interpretability of patterns because of the noises, high-dimensionality and extraordinary sparsity of scRNA-seq data. To address these issues, a novel deep learning subspace clustering algorithm (aka scGDC) for cell types in scRNA-seq data is proposed, which simultaneously learns the deep features and topological structure of cells. Specifically, scGDC extends auto-encoder by introducing a self-representation layer to extract deep features of cells, and learns affinity graph of cells, which provide a better and more comprehensive strategy to characterize structure of cell types. To address heterogeneity of scRNA-seq data, scGDC projects cells of various types onto different subspaces, where types, particularly rare cell types, are well discriminated by utilizing generative adversarial learning. Furthermore, scGDC joins deep feature extraction, structural learning and cell type discovery, where features of cells are extracted under the guidance of cell types, thereby improving performance of algorithms. A total of 15 scRNA-seq datasets from various tissues and organisms with the number of cells ranging from 56 to 63 103 are adopted to validate performance of algorithms, and experimental results demonstrate that scGDC significantly outperforms 14 state-of-the-art methods in terms of various measurements (on average 25.51% by improvement), where (rare) cell types are significantly associated with topology of affinity graph of cells. The proposed model and algorithm provide an effective strategy for the analysis of scRNA-seq data (The software is coded using python, and is freely available for academic https://github.com/xkmaxidian/scGDC).
Haiyue Wang, Xiaoke Ma 0001
Briefings Bioinform.2
2022 Learning discriminative and structural samples for rare cell types with deep generative model
abstract
Cell types (subpopulations) serve as bio-markers for the diagnosis and therapy of complex diseases, and single-cell RNA-sequencing (scRNA-seq) measures expression of genes at cell level, paving the way for the identification of cell types. Although great efforts have been devoted to this issue, it remains challenging to identify rare cell types in scRNA-seq data because of the few-shot problem, lack of interpretability and separation of generating samples and clustering of cells. To attack these issues, a novel deep generative model for leveraging the small samples of cells (aka scLDS2) is proposed by precisely estimating the distribution of different cells, which discriminate the rare and non-rare cell types with adversarial learning. Specifically, to enhance interpretability of samples, scLDS2 generates the sparse faked samples of cells with $\ell _1$-norm, where the relations among cells are learned, facilitating the identification of cell types. Furthermore, scLDS2 directly obtains cell types from the generated samples by learning the block structure such that cells belonging to the same types are similar to each other with the nuclear-norm. scLDS2 joins the generation of samples, classification of the generated and truth samples for cells and feature extraction into a unified generative framework, which transforms the rare cell types detection problem into a classification problem, paving the way for the identification of cell types with joint learning. The experimental results on 20 datasets demonstrate that scLDS2 significantly outperforms 17 state-of-the-art methods in terms of various measurements with 25.12% improvement in adjusted rand index on average, providing an effective strategy for scRNA-seq data with rare cell types. (The software is coded using python, and is freely available for academic https://github.com/xkmaxidian/scLDS2).
Haiyue Wang, Xiaoke Ma 0001
Briefings Bioinform.2
2022 Network-based integrative analysis of single-cell transcriptomic and epigenomic data for cell types
abstract
Advances in single-cell biotechnologies simultaneously generate the transcriptomic and epigenomic profiles at cell levels, providing an opportunity for investigating cell fates. Although great efforts have been devoted to either of them, the integrative analysis of single-cell multi-omics data is really limited because of the heterogeneity, noises and sparsity of single-cell profiles. In this study, a network-based integrative clustering algorithm (aka NIC) is present for the identification of cell types by fusing the parallel single-cell transcriptomic (scRNA-seq) and epigenomic profiles (scATAC-seq or DNA methylation). To avoid heterogeneity of multi-omics data, NIC automatically learns the cell-cell similarity graphs, which transforms the fusion of multi-omics data into the analysis of multiple networks. Then, NIC employs joint non-negative matrix factorization to learn the shared features of cells by exploiting the structure of learned cell-cell similarity networks, providing a better way to characterize the features of cells. The graph learning and integrative analysis procedures are jointly formulated as an optimization problem, and then the update rules are derived. Thirteen single-cell multi-omics datasets from various tissues and organisms are adopted to validate the performance of NIC, and the experimental results demonstrate that the proposed algorithm significantly outperforms the state-of-the-art methods in terms of various measurements. The proposed algorithm provides an effective strategy for the integrative analysis of single-cell multi-omics data (The software is coded using Matlab, and is freely available for academic https://github.com/xkmaxidian/NIC ).
Wensheng Zhang 0002, Xiaoke Ma 0001
Briefings Bioinform.3
2022 Predicting combinations of drugs by exploiting graph embedding of heterogeneous networks
abstract
BACKGROUND: Drug combination, offering an insight into the increased therapeutic efficacy and reduced toxicity, plays an essential role in the therapy of many complex diseases. Although significant efforts have been devoted to the identification of drugs, the identification of drug combination is still a challenge. The current algorithms assume that the independence of feature selection and drug prediction procedures, which may result in an undesirable performance. RESULTS: To address this issue, we develop a novel Semi-supervised Heterogeneous Network Embedding algorithm (called SeHNE) to predict the combination patterns of drugs by exploiting the graph embedding. Specifically, the ATC similarity of drugs, drug-target, and protein-protein interaction networks are integrated to construct the heterogeneous networks. Then, SeHNE jointly learns drug features by exploiting the topological structure of heterogeneous networks and predicting drug combination. One distinct advantage of SeHNE is that features of drugs are extracted under the guidance of classification, which improves the quality of features, thereby enhancing the performance of prediction of drugs. Experimental results demonstrate that the proposed algorithm is more accurate than state-of-the-art methods on various data, implying that the joint learning is promising for the identification of drug combination. CONCLUSIONS: The proposed model and algorithm provide an effective strategy for the prediction of combinatorial patterns of drugs, implying that the graph-based drug prediction is promising for the discovery of drugs.
Shiyin Tan, Zengfa Dou, Xiaoke Ma 0001
BMC Bioinform.5
2022 A hybrid method of detecting flame from video stream
abstract
Abstract In this paper, a method of detecting flame from video stream is proposed exploiting the characteristics of the disordered movement, rapid deformation and intense colour of the flame. Firstly, the frame difference between video frame and background frame is calculated to obtain the main part of the moving object, and the difference between frames is calculated frame by frame in time series to obtain the deformation part of the moving object, and then the sum of cumulative difference between frames and the background difference between frames are added to generate a binary image containing the moving object and the deformed part. Secondly, the binary image is morphologically opened, and rectangular segmentation is carried out to obtain multiple suspicious flame regions. Finally, in the light of the intense colour of the flame, the corresponding area is extracted from the original picture by using the segmentation rectangle, and the colour statistics of the area are carried out to further judge whether there is a burning flame in the area. The experimental results show that the algorithm can accurately detect the burning area of flame in the real scene and eliminate the light interference and the movement interference.
Zengfa Dou, Xiaoke Ma 0001, Xianghua Xie, Chubing Guo
IET Image Process.2
2022 Clustering of noised and heterogeneous multi-view data with graph learning and projection decomposition
Haiyue Wang, Wensheng Zhang 0002, Xiaoke Ma 0001
Knowl. Based Syst.3
2022 Multi-view clustering with constructed bipartite graph in embedding space
Benhui Zhang 0001, Xiaoke Ma 0001
Knowl. Based Syst.2
2022 Joint multi-label learning and feature extraction for temporal link prediction
Xiaoke Ma 0001, Shiyin Tan, Xianghua Xie, Xiaoxiong Zhong, Jingjing Deng 0001
Pattern Recognit.1
2022 Multi-View Clustering With Self-Representation and Structural Constraint
abstract
Multi-view data effectively model and characterize the underlying complex systems, and multi-view clustering is of great significance for revealing the mechanisms of systems, which groups objects into different clusters with high intra-cluster and low inter-cluster similarity for all views. Current algorithms are criticized for undesirable performance because they solely focus on either the shared features or correlation of objects, failing to address the heterogeneity and structural constraint of various views. To overcome these problems, a novelMulti-viewClustering withSelf-representation andStructuralConstraint (MCSSC) is proposed, which is a network-based method by fusing matrix factorization and low-rank representation of various views. Specifically, to remove heterogeneity of multi-view data, a network is constructed for each view, which casts the multi-view clustering into the multi-layer networks clustering problem. To extract the shared features of multiple views, MCSSC factorizes matrices associated with networks by projecting them into a common space and jointly learns an affinity graph for objects in multiple views with self-representation. To facilitate the clustering, the structural constraint is imposed on the affinity graph, where the clusters are identified. Extensive experiments demonstrate that MCSSC significantly outperforms the state-of-the-art in terms of accuracy, implying that the superiority of the proposed method.
Xiaoke Ma 0001, Wensheng Zhang 0002, He Li 0006, Yanni Li, Jiangtao Cui
IEEE Trans. Big Data2
2022 A Two-Phase Method to Balance the Result of Distributed Graph Repartitioning
abstract
With the increase in popularity of graph structured data arising in different areas such as Web, social network, communication network, knowledge graph, etc., there is a growing need for partitioning and repartitioning large graph data in a distributed system. However, the existing graph repartitioning methods are known for poor efficiency in the distributed environment and most of them lack a balance mechanism between edge cut and load balance. In this article, we introduce a new two-phase method to improve the result of distributed graph repartitioning. We first design a local method to identify all the potential candidate vertices that could improve the graph repartitioning result in load balance and edge cut at once in each partition locally. After that, we propose to migrate the selected vertices among the given initial partitions to improve the result of graph repartitioning. During this procedure, we propose to adopt a synchronous vertex migration method to balance both the edge cuts and load balance problems. Extensive experimental results demonstrate that the proposed method is more efficient than the existing methods in several aspects such as communication cost, running time, edge cut, and load balance. We also run SSSP and PageRank applications based on the graph repartitioning result on Giraph to indicate the efficiency of the proposed method.
He Li 0006, Jiangtao Cui, Xiaoke Ma 0001, Shaojie Qiao, Xindong Wu 0001
IEEE Trans. Big Data5
2022 Clustering of Cancer Attributed Networks by Dynamically and Jointly Factorizing Multi-Layer Graphs
abstract
The accumulated omic data provides an opportunity to exploit the mechanisms of cancers and poses a challenge for their integrative analysis. Although extensive efforts have been devoted to address this issue, the current algorithms result in undesirable performance because of the complexity of patterns and heterogeneity of data. In this study, the ultimate goal is to propose an effective and efficient algorithm (called NMF-DEC) to identify clusters by integrating the interactome and transcriptome data. By treating the expression profiles of genes as attributes of vertices in the gene interaction networks, we transform the integrative analysis of omic data into clustering of attributed networks. To circumvent the heterogeneity, we construct a similarity network for the attributes of genes and cast it into the common module detection problem in multi-layer networks. The NMF-DEC explores the relation between attributes and topological structure of networks by jointly factorizing the similarity and interaction networks with the same basis. In this optimization, the interaction network is dynamically updated and the information of attributes is dynamically incorporated, providing a better strategy to characterize the structure of modules in attributed networks. Extensive experiments indicate that compared with state-of-the-art baselines, NMF-DEC is more accurate on social network, and show better performance on cancer attributed networks, implying the superiority of the proposed methods for the integrative analysis of omic data.
Xiaoke Ma 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 Dynamic Module Detection in Temporal Attributed Networks of Cancers
abstract
Tracking the dynamic modules (modules change over time) during cancer progression is essential for studying cancer pathogenesis, diagnosis, and therapy. However, current algorithms only focus on detecting dynamic modules from temporal cancer networks without integrating the heterogeneous genomic data, thereby resulting in undesirable performance. To attack this issue, we propose a novel algorithm (akaTANMF) to detect dynamic modules in cancer temporal attributed networks, which integrates the temporal networks and gene attributes. To obtain the dynamic modules, the temporality and gene attributed are incorporated into an overall objective function, which transforms the dynamic module detection into an optimization problem. TANMF jointly decomposes the snapshots at two subsequent time steps to obtain the latent features of dynamic modules, where the attributes are fused via regulations. Furthermore, the$L_{1}$constraint is imposed to improve the robustness. Experimental results demonstrate that TANMF is more accurate than state-of-the-art methods in terms of accuracy. By applying TANMF to breast cancer data, the obtained dynamic modules are more enriched by the known pathways and associated with patients’ survival time. The proposed model and algorithm provide an effective way for the integrative analysis of heterogeneous omics.
Dongyuan Li, Shuyao Zhang, Xiaoke Ma 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 An Integrative Framework of Heterogeneous Genomic Data for Cancer Dynamic Modules Based on Matrix Decomposition
abstract
Cancer progression is dynamic, and tracking dynamic modules is promising for cancer diagnosis and therapy. Accumulated genomic data provide us an opportunity to investigate the underlying mechanisms of cancers. However, as far as we know, no algorithm has been designed for dynamic modules by integrating heterogeneous omics data. To address this issue, we propose an integrative framework for dynamic module detection based on regularized nonnegative matrix factorization method (DrNMF) by integrating the gene expression and protein interaction network. To remove the heterogeneity of genomic data, we divide the samples of expression profiles into groups to construct gene co-expression networks. To characterize the dynamics of modules, the temporal smoothness framework is adopted, in which the gene co-expression network at the previous stage and protein interaction network are incorporated into the objective function of DrNMF via regularization. The experimental results demonstrate that DrNMF is superior to state-of-the-art methods in terms of accuracy. For breast cancer data, the obtained dynamic modules are more enriched by the known pathways, and can be used to predict the stages of cancers and survival time of patients. The proposed model and algorithm provide an effective integrative analysis of heterogeneous genomic data for cancer progression.
Xiaoke Ma 0001, Peng Gang Sun, Maoguo Gong
IEEE ACM Trans. Comput. Biol. Bioinform.1
2022 POTAM: A Parallel Optimal Task Allocation Mechanism for Large-Scale Delay Sensitive Mobile Edge Computing
abstract
Design an optimization model for task management among Mobile Terminal (MT), Macro cell Base Station (MBS), and multiple Small cell Base Stations (SBS) for the large-scale Mobile Edge Computing (MEC) system, is a challenging issue due to the large number of tasks and SBSs. Inspired by this, we propose a Parallel Optimal Task Allocation Mechanism (POTAM) framework for MEC, which includes Device to Device (D2D)-enabled computing, MBS computing and Edge Computation Resource Distribution (ECRD) computing. In POTAM, we exploit a parallel multi-block Alternating Direction Method of Multipliers (ADMM) based method to model both requirements of delay and energy consumptions, which formulates the task allocation under these requirements as a nonlinear 0–1 integer programming problem. To solve this problem, we develop an efficient combination of conjugate gradient, Newton and linear search techniques based algorithm with Logarithmic Smoothing and Cyclic Block coordinate Gradient Projection (CBGP) methods, which can guarantee convergence and reduce computational complexity with a good scalability. In order to allocate task cooperatively, an optimal approach is proposed, ECRD-A, which is used to find the shortest path among each node. Numerical results demonstrate the effectiveness of the POTAM and it can effectively reduce delay and energy consumption for a large-scale MEC system.
Xiaoxiong Zhong, Xinghan Wang 0001, Tingting Yang 0001, Yuanyuan Yang 0001, Yang Qin 0001, Xiaoke Ma 0001
IEEE Trans. Commun.6
2022 Edge Repartitioning via Structure-Aware Group Migration
abstract
Graph partitioning is a mandatory step in distributed graph computing systems. Some existing systems use edge partitioning methods to partition static graphs. However, the structure of the real-world graphs changes dynamically, which leads to unnecessary vertex replicas and load imbalance, reducing the performance of graph computation. In this article, we focus on improving the lower partitioning quality caused by the dynamics of the graph structure. We propose an edge repartitioning algorithm via structure-aware group migration (SAGM-ER). We define a special structure edge group (EG) consisting of multiple edges, which can reduce vertex replicas by migrating to other partitions. In repartitioning, we search for EGs in parallel by a method based on a structure-aware priority and then migrate EGs to reduce vertex replicas. Compared to the state of the art, SAGM-ER can reduce more vertex replicas. We implement SAGM-ER on Powergraph, which reduces the redundant replicas by 63.33%, thus reducing executing time and communication costs in graph computation by 33.72% and 37.51%, respectively.
He Li 0006, Xiaoke Ma 0001, Jiangtao Cui, Jae Soo Yoo
IEEE Trans. Comput. Soc. Syst.4
2022 Deep Reinforcement Learning-based Trajectory Pricing on Ride-hailing Platforms
abstract
Dynamic pricing plays an important role in solving the problems such as traffic load reduction, congestion control, and revenue improvement. Efficient dynamic pricing strategies can increase capacity utilization, total revenue of service providers, and the satisfaction of both passengers and drivers. Many proposed dynamic pricing technologies focus on short-term optimization and face poor scalability in modeling long-term goals for the limitations of solution optimality and prohibitive computation. In this article, a deep reinforcement learning framework is proposed to tackle the dynamic pricing problem for ride-hailing platforms. A soft actor-critic (SAC) algorithm is adopted in the reinforcement learning framework. First, the dynamic pricing problem is translated into a Markov Decision Process (MDP) and is set up in continuous action spaces, which is no need for the discretization of action space. Then, a new reward function is obtained by the order response rate and the KL-divergence between supply distribution and demand distribution. Experiments and case studies demonstrate that the proposed method outperforms the baselines in terms of order response rate and total revenue.
Longji Huang, Meijuan Liu, He Li 0006, Qinglin Tan, Xiaoke Ma 0001, Jiangtao Cui, De-Shuang Huang
ACM Trans. Intell. Syst. Technol.6
2021 Multi-view Clustering for the Integration Analysis of Gene Expression and Methylation Data
abstract
The accumulated gene expression and DNA methylation data provide a great opportunity to exploit the mechanisms of biological systems. Current algorithms for the integration of gene expression and methylation data are characterized for undesirable performance because they fail to address the latent relations in the heterogeneous data. To solve this problem, we propose a novel multi-view clustering with self-representation learning and low-rank tensor constraint (MCSL-LTC), where the gene expression and DNA methylation data are treated as complementary views, and MCSL-LTC obtains a consensus partitioning reflecting the structure and features of various views. Specifically, self-representation learning is employed to explore the low-dimensional subspace structures embedded in different views, where the tensor norm is adopted to smooth different views, therefore improving the quality of features. Experimental results demonstrate that the proposed approach outperforms state-of-the-art baselines in terms of accuracy on both the social and cancer data, provides an effective and efficient method for the integration of heterogeneous genomic data.
Xiaoke Ma 0001
BIBM3
2021 Transfer Learning for Gene Ranking across Cancers
abstract
Understanding the mechanisms of the disease and discovering the underlying pathogenic genes is very important for human health and applications. Therefore, it is vital to rank potential genes and find out the interrelationships between diseases and genes. Current algorithms rely heavily on different types of biological networks, which come from biological experiments and related literature. As a result, they will get unsatisfactory results when similar diseases’ information is insufficient. To overcome this problem, we propose the Transfer Learning for Gene Ranking across Cancers algorithm (called GRTL). Specifically, GRTL decomposes the source and target gene expression data jointly. The feature matrices are also decomposed into the shared part and the special part. We also constrain the special part with label information and let them orthogonal to ensure diversity. Then, GRTL automatically learns an affinity graph via common knowledge. Finally, we prioritize genes with gene ranking algorithm. Experimental results demonstrate that the proposed algorithm is more accurate than state-of-the-art methods on the mutation and cancer dataset. The proposed model provides a novel way to find new interrelationships between genes and similar diseases.
Xiaoke Ma 0001
BIBM3
2021 Temporal Link Prediction for Cancer Networks using Structural Consistency Regularized Non-negative Matrix Factorization
abstract
Genes are basic biology system units that execute critical biological processes by interacting with others. However, interactions among genes associated with cancer progression are incomplete, thereby hampering the downstream network analysis. Although many algorithms have been devoted to link prediction on static networks, efforts for temporal link prediction on dynamic networks associated with cancer progression are limited. To address this problem, a novel structural consistency nonnegative matrix factorization(aka SCNMF) for the temporal link prediction on cancer networks is proposed, where features of vertices at the current time are learned under the temporal smoothness framework. Specifically, SCNMF jointly factorizes three successive snapshots using nonnegative matrix factorization (NMF) and learns the shared features of vertices as global consistency, where the original geometrical structure is preserved in the latent space. In this case, SCNMF simultaneously fuses the topological structure of the current snapshot and the temporality of snapshots at the previous and subsequent time. Experimental results demonstrate that the proposed algorithm outperforms state-of-the-art baselines in terms of accuracy, providing an effective alternative for temporal link prediction on cancer networks.
Xiaoke Ma 0001
BIBM3
2021 jSRC: a flexible and accurate joint learning algorithm for clustering of single-cell RNA-sequencing data
abstract
Single-cell RNA-sequencing (scRNA-seq) explores the transcriptome of genes at cell level, which sheds light on revealing the heterogeneity and dynamics of cell populations. Advances in biotechnologies make it possible to generate scRNA-seq profiles for large-scale cells, requiring effective and efficient clustering algorithms to identify cell types and informative genes. Although great efforts have been devoted to clustering of scRNA-seq, the accuracy, scalability and interpretability of available algorithms are not desirable. In this study, we solve these problems by developing a joint learning algorithm [a.k.a. joints sparse representation and clustering (jSRC)], where the dimension reduction (DR) and clustering are integrated. Specifically, DR is employed for the scalability and joint learning improves accuracy. To increase the interpretability of patterns, we assume that cells within the same type have similar expression patterns, where the sparse representation is imposed on features. We transform clustering of scRNA-seq into an optimization problem and then derive the update rules to optimize the objective of jSRC. Fifteen scRNA-seq datasets from various tissues and organisms are adopted to validate the performance of jSRC, where the number of single cells varies from 49 to 110 824. The experimental results demonstrate that jSRC significantly outperforms 12 state-of-the-art methods in terms of various measurements (on average 20.29% by improvement) with fewer running time. Furthermore, jSRC is efficient and robust across different scRNA-seq datasets from various tissues. Finally, jSRC also accurately identifies dynamic cell types associated with progression of COVID-19. The proposed model and methods provide an effective strategy to analyze scRNA-seq data (the software is coded using MATLAB and is free for academic purposes; https://github.com/xkmaxidian/jSRC).
Zaiyi Liu, Xiaoke Ma 0001
Briefings Bioinform.3
2021 TLGP: a flexible transfer learning algorithm for gene prioritization based on heterogeneous source domain
abstract
BACKGROUND: Gene prioritization (gene ranking) aims to obtain the centrality of genes, which is critical for cancer diagnosis and therapy since keys genes correspond to the biomarkers or targets of drugs. Great efforts have been devoted to the gene ranking problem by exploring the similarity between candidate and known disease-causing genes. However, when the number of disease-causing genes is limited, they are not applicable largely due to the low accuracy. Actually, the number of disease-causing genes for cancers, particularly for these rare cancers, are really limited. Therefore, there is a critical needed to design effective and efficient algorithms for gene ranking with limited prior disease-causing genes. RESULTS: In this study, we propose a transfer learning based algorithm for gene prioritization (called TLGP) in the cancer (target domain) without disease-causing genes by transferring knowledge from other cancers (source domain). The underlying assumption is that knowledge shared by similar cancers improves the accuracy of gene prioritization. Specifically, TLGP first quantifies the similarity between the target and source domain by calculating the affinity matrix for genes. Then, TLGP automatically learns a fusion network for the target cancer by fusing affinity matrix, pathogenic genes and genomic data of source cancers. Finally, genes in the target cancer are prioritized. The experimental results indicate that the learnt fusion network is more reliable than gene co-expression network, implying that transferring knowledge from other cancers improves the accuracy of network construction. Moreover, TLGP outperforms state-of-the-art approaches in terms of accuracy, improving at least 5%. CONCLUSION: The proposed model and method provide an effective and efficient strategy for gene ranking by integrating genomic data from various cancers.
Zuheng Xia, Jingjing Deng 0001, Xianghua Xie, Maoguo Gong, Xiaoke Ma 0001
BMC Bioinform.6
2021 Identification of dynamic community in temporal network via joint learning graph representation and nonnegative matrix factorization
Dongyuan Li, Qiang Lin 0001, Xiaoke Ma 0001
Neurocomputing3
2021 Joint nonnegative matrix factorization and network embedding for graph co-clustering
Xiaoke Ma 0001
Neurocomputing2
2021 Detecting dynamic community by fusing network embedding and nonnegative matrix factorization
Dongyuan Li, Xiaoxiong Zhong, Zengfa Dou, Maoguo Gong, Xiaoke Ma 0001
Knowl. Based Syst.5
2021 Identification of multi-layer networks community by fusing nonnegative matrix factorization and topological structural information
Changzhou Ma, Qiang Lin 0001, Xiaoke Ma 0001
Knowl. Based Syst.4
2021 Clustering Heterogeneous Information Network by Joint Graph Embedding and Nonnegative Matrix Factorization
abstract
Many complex systems derived from nature and society consist of multiple types of entities and heterogeneous interactions, which can be effectively modeled as heterogeneous information network (HIN). Structural analysis of heterogeneous networks is of great significance by leveraging the rich semantic information of objects and links in the heterogeneous networks. And, clustering heterogeneous networks aims to group vertices into classes, which sheds light on revealing the structure–function relations of the underlying systems. The current algorithms independently perform the feature extraction and clustering, which are criticized for not fully characterizing the structure of clusters. In this study, we propose a learning model by joint Graph Embedding and Nonnegative Matrix Factorization (aka GEjNMF ), where feature extraction and clustering are simultaneously learned by exploiting the graph embedding and latent structure of networks. We formulate the objective function of GEjNMF and transform the heterogeneous network clustering problem into a constrained optimization problem, which is effectively solved by l 0 -norm optimization. The advantage of GEjNMF is that features are selected under the guidance of clustering, which improves the performance and saves the running time of algorithms at the same time. The experimental results on three benchmark heterogeneous networks demonstrate that GEjNMF achieves the best performance with the least running time compared with the best state-of-the-art methods. Furthermore, the proposed algorithm is robust across heterogeneous networks from various fields. The proposed model and method provide an effective alternative for heterogeneous network clustering.
Benhui Zhang 0001, Maoguo Gong, Xiaoke Ma 0001
ACM Trans. Knowl. Discov. Data4
2021 Group Reassignment for Dynamic Edge Partitioning
abstract
Graph partitioning is a mandatory step in large-scale distributed graph processing. When partitioning real-world power-law graphs, the edge partitioning algorithm performs better than the traditional vertex partitioning algorithm, because it can cut a single vertex into multiple replicas to apportion the computation. Many advanced edge partitioning methods are designed for partitioning a static graph from scratch. However, the real-world graph structure changes continuously, which leads to a decrease in partition quality and affects the performance of the graph applications. Some studies are devoted to offline repartitioning or batch incremental partitioning, but how to deal with dynamics in real-time is still worthy of in-depth study. In this article, we discuss the impact of dynamic change on partition and discover that both insertion and deletion will lead to local suboptimal partitioning, which is the reason for the degradation of partition quality. As a solution, a dynamic edge partitioning algorithm is proposed to partition dynamics in real-time. Specifically, we deal with dynamics by a distributed stream and improve partition quality by reassigning some closely connected edges. Experiments show that it is robust to initial partition quality, dynamic scale and type, and distributed scale. Compared with the state-of-the-art dynamic partitioner, it can reduce vertex-cuts by 29.5 percent. Compared with the repartitioning algorithms, it can save the partitioning time by 91.0 percent. Applied on the graph task, it can reduce the increase of communication cost and the increase of the total time of task by 41.5 and 71.4 percent.
He Li 0006, Jiangtao Cui, Xiaoke Ma 0001, Senzhang Wang, Jae Soo Yoo, Philip S. Yu
IEEE Trans. Parallel Distributed Syst.5
2020 SeHNE: Semi-supervised Heterogeneous Network Embedding for Drug Combination
abstract
Drug combinations, offering increased therapeutic efficacy and reduced toxicity, play an important role in therapy of many complex diseases. Although great efforts have been devoted to the prediction of single drugs, the identification of drug combination is really limited. The current algorithms assume the independence of features and prediction, resulting in an undesirable performance. To address this issue, we develop a novel semisupervised heterogeneous network embedding algorithm (called SeHNE) to predict drug combinations, where ATC similarity of drugs, drug-target and protein-protein interaction (PPI) networks are integrated to construct heterogeneous network. SeHNE jointly learns features of drugs by exploiting the topological structure of heterogeneous networks, and prediction of drug combination. One typical advantage of SeHNE is that features are extracted under the guidance of classification, thereby improving the accuracy of algorithms. Experimental results demonstrate that proposed algorithm is more accurate than state-of-the-art methods on the dataset we collected, and the re-training process could improve the accuracy of classifier.
Shiyin Tan, Xiaoke Ma 0001
BIBM2
2020 A Task Allocation Framework for Large-Scale Mobile Edge Computing
abstract
We consider the problem of intelligent and efficient task allocation mechanism in large-scale mobile edge computing (MEC), which can reduce delay and energy consumption in a parallel and distributed optimization. In this paper, we study the joint optimization model to consider cooperative task management mechanism among mobile terminals (MT), macro cell base station (MBS), and multiple small cell base station (SBS) for large-scale MEC applications. We propose a parallel multi-block Alternating Direction Method of Multipliers (ADMM) based method to model both requirements of low delay and low energy consumption in the MEC system which formulates the task allocation under those requirements as a nonlinear 0-1 integer programming problem. To solve the optimization problem, we develop an efficient combination of conjugate gradient, Newton and linear search techniques based algorithm with Logarithmic Smoothing (for global variables updating) and the Cyclic Block coordinate Gradient Projection (CBGP, for local variables updating) methods, which can guarantee convergence and reduce computational complexity with a good scalability. Numerical results demonstrate the effectiveness of the proposed mechanism and it can effectively reduce delay and energy consumption for a large-scale MEC system.
Xinghan Wang 0001, Xiaoxiong Zhong, Yanbin Zheng, Xiaoke Ma 0001, Tingting Yang 0001, Genglin Zhang
GLOBECOM4
2020 Joint learning dimension reduction and clustering of single-cell RNA-sequencing data
abstract
MOTIVATION: Single-cell RNA-sequencing (scRNA-seq) profiles transcriptome of individual cells, which enables the discovery of cell types or subtypes by using unsupervised clustering. Current algorithms perform dimension reduction before cell clustering because of noises, high-dimensionality and linear inseparability of scRNA-seq data. However, independence of dimension reduction and clustering fails to fully characterize patterns in data, resulting in an undesirable performance. RESULTS: In this study, we propose a flexible and accurate algorithm for scRNA-seq data by jointly learning dimension reduction and cell clustering (aka DRjCC), where dimension reduction is performed by projected matrix decomposition and cell type clustering by non-negative matrix factorization. We first formulate joint learning of dimension reduction and cell clustering into a constrained optimization problem and then derive the optimization rules. The advantage of DRjCC is that feature selection in dimension reduction is guided by cell clustering, significantly improving the performance of cell type discovery. Eleven scRNA-seq datasets are adopted to validate the performance of algorithms, where the number of single cells varies from 49 to 68 579 with the number of cell types ranging from 3 to 14. The experimental results demonstrate that DRjCC significantly outperforms 13 state-of-the-art methods in terms of various measurements on cell type clustering (on average 17.44% by improvement). Furthermore, DRjCC is efficient and robust across different scRNA-seq datasets from various tissues. The proposed model and methods provide an effective strategy to analyze scRNA-seq data. AVAILABILITY AND IMPLEMENTATION: The software is coded using matlab, and is free available for academic https://github.com/xkmaxidian/DRjCC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaoke Ma 0001
Bioinform.2
2020 Co-regularized nonnegative matrix factorization for evolving community detection in dynamic networks
Xiaoke Ma 0001, Benhui Zhang 0001, Changzhou Ma, Zhiyu Ma
Inf. Sci.1
2020 Detecting community in attributed networks by dynamically exploring node attributes and topological structure
Xiaoxiong Zhong, Maoguo Gong, Xiaoke Ma 0001
Knowl. Based Syst.5
2019 An Integrative Framework for Protein Interaction Network and Methylation Data to Discover Epigenetic Modules
abstract
DNA methylation is a critical epigenetic modification that plays an important role in cancers. The available algorithms fail to fully characterize epigenetic modules. To address this issue, we first characterize the epigenetic module as a group of well-connected genes in the protein interaction network and are also co-methylated based on gene methylation profiles. Then, the epigenetic module discovery problem is transformed into an optimization problem. Then, a regularized nonnegative matrix factorization algorithm for methylation modules (RNMF-MM) is presented, where the co-methylation constraint is treated as a regularizer. Using the artificial networks with known module structure, we demonstrate that the proposed algorithm outperforms state-of-the-art approaches in terms of accuracy. On the basis of breast cancer methylation data and protein interaction network, the RNMF-MM algorithm discovers methylation modules that are significantly more enriched by the known pathways than those obtained by other algorithms. These modules serve as biomarkers for predicting cancer stages and estimating survival time of patients. The proposed model and algorithm provide an effective way for the integrative analysis of protein interaction network and methylation data.
Xiaoke Ma 0001, Penggang Sun, Zhong-Yuan Zhang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 Community Detection in Multi-Layer Networks Using Joint Nonnegative Matrix Factorization
abstract
Many complex systems are composed of coupled networks through different layers, where each layer represents one of many possible types of interactions. A fundamental question is how to extract communities in multi-layer networks. The current algorithms either collapses multi-layer networks into a single-layer network or extends the algorithms for single-layer networks by using consensus clustering. However, these approaches have been criticized for ignoring the connection among various layers, thereby resulting in low accuracy. To attack this problem, a quantitative function (multi-layer modularity density) is proposed for community detection in multi-layer networks. Afterward, we prove that the trace optimization of multi-layer modularity density is equivalent to the objective functions of algorithms, such as kernel$K$-means, nonnegative matrix factorization (NMF), spectral clustering and multi-view clustering, for multi-layer networks, which serves as the theoretical foundation for designing algorithms for community detection. Furthermore, aSemi-SupervisedjointNonnegativeMatrixFactorization algorithm (S2-jNMF) is developed by simultaneously factorizing matrices that are associated with multi-layer networks. Unlike the traditional semi-supervised algorithms, the partial supervision is integrated into the objective of the S2-jNMF algorithm. Finally, through extensive experiments on both artificial and real world networks, we demonstrate that the proposed method outperforms the state-of-the-art approaches for community detection in multi-layer networks.
Xiaoke Ma 0001, Di Dong, Quan Wang 0006
IEEE Trans. Knowl. Data Eng.1
2018 Identifying Condition-Specific Modules by Clustering Multiple Networks
abstract
Condition-specific modules in multiple networks must be determined to reveal the underlying molecular mechanisms of diseases. Current algorithms exhibit limitations such as low accuracy and high sensitivity to the number of networks because these algorithms discover condition-specific modules in multiple networks by separating specificity and modularity of modules. To overcome these limitations, we characterize condition-specific module as a group of genes whose connectivity is strong in the corresponding network and weak in other networks; this strategy can accurately depict the topological structure of condition-specific modules. We then transform the condition-specific module discovery problem into a clustering problem in multiple networks. We develop an efficient heuristic algorithm for the Specific Modules in Multiple Networks (SMMN), which discovers the condition-specific modules by considering multiple networks. By using the artificial networks, we demonstrate that SMMN outperforms state-of-the-art methods. In breast cancer networks, stage-specific modules discovered by SMMN are more discriminative in predicting cancer stages than those obtained by other techniques. In pan-cancer networks, cancer-specific modules are more likely to associate with survival time of patients, which is critical for cancer therapy.
Xiaoke Ma 0001, Penggang Sun, Guimin Qin
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 Extracting Stage-Specific and Dynamic Modules Through Analyzing Multiple Networks Associated with Cancer Progression
abstract
Determining the dynamics of pathways associated with cancer progression is critical for understanding the etiology of diseases. Advances in biological technology have facilitated the simultaneous genomic profiling of multiple patients at different clinical stages, thus generating the dynamic genomic data for cancers. Such data provide enable investigation of the dynamics of related pathways. However, methods for integrative analysis of dynamic genomic data are inadequate. In this study, we develop a novel nonnegative matrix factorization algorithm for dynamic modules ( NMF-DM), which simultaneously analyzes multiple networks for the identification of stage-specific and dynamic modules. NMF-DM applies the temporal smoothness framework by balancing the networks at the current stage and the previous stage. Experimental results indicate that the NMF-DM algorithm is more accurate than the state-of-the-art methods in artificial dynamic networks. In breast cancer networks, NMF-DM reveals the dynamic modules that are important for cancer stage transitions. Furthermore, the stage-specific and dynamic modules have distinct topological and biochemical properties. Finally, we demonstrate that the stage-specific modules significantly improve the accuracy of cancer stage prediction. The proposed algorithm provides an effective way to explore the time-dependent cancer genomic data.
Xiaoke Ma 0001, Wanxin Tang, Peizhuo Wang, Xingli Guo, Lin Gao 0006
IEEE ACM Trans. Comput. Biol. Bioinform.1
2017 Multiple network algorithm for epigenetic modules via the integration of genome-wide DNA methylation and gene expression data
abstract
BACKGROUND: With the increase in the amount of DNA methylation and gene expression data, the epigenetic mechanisms of cancers can be extensively investigate. Available methods integrate the DNA methylation and gene expression data into a network by specifying the anti-correlation between them. However, the correlation between methylation and expression is usually unknown and difficult to determine. RESULTS: To address this issue, we present a novel multiple network framework for epigenetic modules, namely, Epigenetic Module based on Differential Networks (EMDN) algorithm, by simultaneously analyzing DNA methylation and gene expression data. The EMDN algorithm prevents the specification of the correlation between methylation and expression. The accuracy of EMDN algorithm is more efficient than that of modern approaches. On the basis of The Cancer Genome Atlas (TCGA) breast cancer data, we observe that the EMDN algorithm can recognize positively and negatively correlated modules and these modules are significantly more enriched in the known pathways than those obtained by other algorithms. These modules can serve as bio-markers to predict breast cancer subtypes by using methylation profiles, where positively and negatively correlated modules are of equal importance in the classification of cancer subtypes. Epigenetic modules also estimate the survival time of patients, and this factor is critical for cancer therapy. CONCLUSIONS: The proposed model and algorithm provide an effective method for the integrative analysis of DNA methylation and gene expression. The algorithm is freely available as an R-package at https://github.com/william0701/EMDN .
Xiaoke Ma 0001, Zaiyi Liu, Wanxin Tang
BMC Bioinform.1
2017 Evolutionary Nonnegative Matrix Factorization Algorithms for Community Detection in Dynamic Networks
abstract
Discovering evolving communities in dynamic networks is essential to important applications such as analysis for dynamic web content and disease progression. Evolutionary clustering uses the temporal smoothness framework that simultaneously maximizes the clustering accuracy at the current time step and minimizes the clustering drift between two successive time steps. In this paper, we propose two evolutionary nonnegative matrix factorization (ENMF) frameworks for detecting dynamic communities. To address the theoretical relationship among evolutionary clustering algorithms, we first prove the equivalence relationship between ENMF and optimization of evolutionary modularity density. Then, we extend the theory by proving the equivalence between evolutionary spectral clustering and ENMF, which serves as the theoretical foundation for hybrid algorithms. Based on the equivalence, we propose a semi-supervised ENMF (sE-NMF) by incorporating a priori information into ENMF. Unlike the traditional semi-supervised algorithms, a priori information is integrated into the objective function of the algorithm. The main advantage of the proposed algorithm is to escape the local optimal solution without increasing time complexity. The experimental results over a number of artificial and real world dynamic networks illustrate that the proposed method is not only more accurate but also more robust than the state-of-the-art approaches.
Xiaoke Ma 0001, Di Dong
IEEE Trans. Knowl. Data Eng.1
2015 Revealing Pathway Dynamics in Heart Diseases by Analyzing Multiple Differential Networks
abstract
Development of heart diseases is driven by dynamic changes in both the activity and connectivity of gene pathways. Understanding these dynamic events is critical for understanding pathogenic mechanisms and development of effective treatment. Currently, there is a lack of computational methods that enable analysis of multiple gene networks, each of which exhibits differential activity compared to the network of the baseline/healthy condition. We describe the iMDM algorithm to identify both unique and shared gene modules across multiple differential co-expression networks, termed M-DMs (multiple differential modules). We applied iMDM to a time-course RNA-Seq dataset generated using a murine heart failure model generated on two genotypes. We showed that iMDM achieves higher accuracy in inferring gene modules compared to using single or multiple co-expression networks. We found that condition-specific M-DMs exhibit differential activities, mediate different biological processes, and are enriched for genes with known cardiovascular phenotypes. By analyzing M-DMs that are present in multiple conditions, we revealed dynamic changes in pathway activity and connectivity across heart failure conditions. We further showed that module dynamics were correlated with the dynamics of disease phenotypes during the development of heart failure. Thus, pathway dynamics is a powerful measure for understanding pathogenesis. iMDM provides a principled way to dissect the dynamics of gene pathways and its relationship to the dynamics of disease phenotype. With the exponential growth of omics data, our method can aid in generating systems-level insights into disease progression.
Xiaoke Ma 0001, Georgios Karamanlidis, Chi Fung Lee, Lorena Garcia-Menendez, Rong Tian, Kai Tan 0001
PLoS Comput. Biol.1
2014 Modeling disease progression using dynamics of pathway connectivity
abstract
MOTIVATION: Disease progression is driven by dynamic changes in both the activity and connectivity of molecular pathways. Understanding these dynamic events is critical for disease prognosis and effective treatment. Compared with activity dynamics, connectivity dynamics is poorly explored. RESULTS: We describe the M-module algorithm to identify gene modules with common members but varied connectivity across multiple gene co-expression networks (aka M-modules). We introduce a novel metric to capture the connectivity dynamics of an entire M-module. We find that M-modules with dynamic connectivity have distinct topological and biochemical properties compared with static M-modules and hub genes. We demonstrate that incorporation of module connectivity dynamics significantly improves disease stage prediction. We identify different sets of M-modules that are important for specific disease stage transitions and offer new insights into the molecular events underlying disease progression. Besides modeling disease progression, the algorithm and metric introduced here are broadly applicable to modeling dynamics of molecular pathways. AVAILABILITY AND IMPLEMENTATION: M-module is implemented in R. The source code is freely available at http://www.healthcare.uiowa.edu/labs/tan/M-module.zip.
Xiaoke Ma 0001, Kai Tan 0001
Bioinform.1
2012 Predicting protein complexes in protein interaction networks using a core-attachment algorithm based on graph communicability
Xiaoke Ma 0001, Lin Gao 0006
Inf. Sci.1
2010 A centrality measure based on spectral optimization of modularity density
Lidong Fu, Lin Gao 0006, Xiaoke Ma 0001
Sci. China Inf. Sci.3