Maozu Guo 0001

dblp:38/336-1 · also Mao-Zu Guo 0001 · DBLP profile ↗
← Back
128ranked-venue papers
5as first author
58since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 75 · 2 first-author · 38 since 2021Artificial intelligence and machine learning · 33 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 18 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 GL 2 T -Diff: Medical image translation via spatial-frequency fusion diffusion models
Dong Sui, Nanting Song, Yacong Li, Maozu Guo 0001, Kuanquan Wang, Gongning Luo
Comput. Vis. Image Underst.6
2026 AM-A2C: Cloud computing task scheduling algorithm based on attention mechanism and action mutation
Yuzhang Zhang, Le Tian 0001, Maozu Guo 0001
Future Gener. Comput. Syst.3
2026 Weakly supervised single-stage crack detection based on multi-scale feature fusion
abstract
Cracks pose a significant threat to road and building safety, making effective detection of cracks on road surfaces a focus of research both domestically and internationally. Deep learning-based methods often require extensive pixel-level annotations, posing significant labor costs. We propose a single-stage weakly supervised crack segmentation model based on multi-scale feature fusion. The model is built on a single-stage weakly supervised segmentation framework, which reduces model complexity. It utilizes a multi-scale feature fusion module (PPM) to integrate features at different scales, enhancing the ability to extract features from cracks of various sizes. The combination of the Domain Restriction Suppression (DRS) module and pixel affinity convolution is employed to optimize pseudo-pixel annotations. In addition, we propose a joint loss function to mitigate sample imbalance between crack and non-crack pixels. Compared to other two-stage weakly supervised segmentation models, our model is simpler and more effective, achieving excellent results on the Deep Crack and Crack500 datasets, surpassing most weakly supervised crack segmentation models in terms of Recall (Re), F-score (F1), and mean Intersection-over-Union (mIoU), achieving similar effects to fully supervised crack segmentation models. This demonstrates the effectiveness and robustness of our model.
Yacong Li, Maozu Guo 0001, Dong Sui
Intell. Data Anal.4
2026 CRTSE: Clustering and Reinforcement Learning Based Task Scheduling Algorithm for Edge Computing
abstract
Edge computing can overcome many shortcomings of traditional cloud computing and provide high quality computing services. However, it need to face the challenges of node heterogeneity and task diversity, which leads to the performance of edge computing systems relying on proper task scheduling. Existing task scheduling algorithms are usually designed based on mathematical models that are closely related to the internal details of edge computing systems, resulting in poor universality and limited quality of scheduling decisions. In this paper, we address these issues by proposing a Clustering and Reinforcement Learning Based Task Scheduling Algorithm for Edge Computing (CRTSE), which aims to shorten the task processing time and improve the energy efficiency of the system. The algorithm uses Adaptive Graph Auto-Encoder (AdaGAE) based clustering algorithm to cluster computing tasks and classify the computing tasks submitted to the edge computing system based on the clustering results. When making scheduling decisions, an independent deep reinforcement learning model is used for each class of computing tasks to obtain targeted scheduling preference information, and the scheduling decision is made based on these preference information. CRTSE possesses good universality due to the feature of not relying on mathematical models that are closely related to the internal details of edge computing systems, and has the ability to adapt well to diverse computing tasks and to learn continuously during the interaction with the system. Simulation experiments based on real data show that CRTSE can shorten the average task processing time of the edge computing system by up to 56.00% and reduce the average energy consumption of the edge computing system by up to 16.10% compared to existing excellent scheduling algorithms.
Le Tian 0001, Maozu Guo 0001
IEEE Trans. Cloud Comput.3
2026 Ctfnet: toward high generalization medical image segmentation via coarse-to-fine structures for multi-center datasets
Dong Sui, Sitong Bao, Donghui Lei, Yacong Li, Maozu Guo 0001, Xiangyu Li 0004, Kuanquan Wang, Gongning Luo
Vis. Comput.5
2025 FMPNet: A Multi-Features Fusion Framework for Predicting Neoadjuvant Chemoradiotherapy Efficacy in Locally Advanced Rectal Cancer
abstract
Neoadjuvant chemoradiotherapy (nCRT) is the standard treatment for locally advanced rectal cancer (LARC). However, substantial inter-patient variability in response to nCRT poses a significant challenge for accurately predicting treatment outcomes based on preoperative data, thereby complicating clinical decision-making. With the rapid advancement of artificial intelligence technologies, there has been growing interest in leveraging AI for predictive modeling in cancer therapy. In this study, we propose a novel multi-modal prediction framework, FMPNet, which integrates preoperative magnetic resonance imaging (MRI) and whole-slide image (WSI) features to predict nCRT efficacy. Specifically, for WSI processing, we develop an efficient tumor cell segmentation strategy and incorporate a deep subspace clustering mechanism into the feature extraction pipeline to enhance the model's representational capacity. Comprehensive experiments demonstrate that FMPNet consistently outperforms ten other feature fusion-based prediction models on both internal test sets and external validation cohorts across multiple metrics, including accuracy, precision, recall, F1-score and ROC-AUC curves. These results not only confirm the superior performance of our model but also underscore its potential to support more accurate and personalized clinical decision-making for patients with LARC.
Dong Sui, Nanting Song, Zhehao Xu, Yacong Li, Maozu Guo 0001, Gongning Luo, Kuanquan Wang, Henggui Zhang
BIBM5
2025 A Dual-Domain Framework with Wavelet Attention for Cardiac Ultrasound Image Quality Assessment
abstract
Automated quality assessment of cardiac ultrasound images is crucial for ensuring diagnostic accuracy and clinical decision-making reliability. Hospitals face significant challenges in efficiently screening ultrasound image quality, where manual expert review is time-consuming and subjective. However, dedicated methods for ultrasound image quality assessment remain scarce, with most adapted from natural image quality metrics that fail to capture clinically relevant diagnostic factors. In this paper, we propose a novel dual-domain framework that models both spatial anatomical features and frequency-domain spectral characteristics using specialized neural modules. Our approach incorporates cardiac-specific attention mechanisms and wavelet-based artifact detection to enable comprehensive, clinically aligned evaluation. Extensive experiments on a largescale clinical dataset demonstrate the superiority of our method, achieving 88.1 % overall accuracy, a macro-averaged F1-score of 88.1 %, and 100 % precision and recall in detecting diagnostically unacceptable images. The proposed framework is designed to meet clinical reliability standards, paving the way for safe integration into diagnostic workflows.
Dong Sui, Zhehao Xu, Nanting Song, Yacong Li, Maozu Guo 0001, Gongning Luo, Kuanquan Wang, Henggui Zhang
BIBM5
2025 Multi-Scale Feature Fusion Network for the Prediction of Protein-Protein Binding Affinity Changes upon Mutations
abstract
Accurate prediction of changes in protein-protein binding affinity influenced by mutations (i.e.,$\Delta\Delta G)$is essential for understanding the structure and function of proteins and elucidating the underlying mechanisms of complex diseases. We introduce a novel multi-scale feature fusion network for protein-protein$\Delta\Delta G$prediction, aiming to reduce the dependency of previous methods on intricate biological features and expert-driven knowledge, and demonstrate its effectiveness in the challenging and important domain of protein data. Specifically, we employ multi-scale modeling of protein complexes, and introduce self-attention mechanisms and various extraction modules to comprehensively capture the features. The proposed method effectively learns the complete biological regulation of complexes and incorporates interactions between amino acids. Extensive experiments on three benchmark datasets demonstrate that our proposed framework significantly outperforms the state-of-the-art methods for$\Delta\Delta G$prediction while also providing excellent performance in the domain of membrane protein design.
Hao Zhang 0128, Yang Liu 0006, Limin Yu, Zejie Wang, Maozu Guo 0001
BIBM6
2025 DM3diff: A novel multi-center, multi-modality and multi-source medical image segmentation framework based on DWT embeded diffusion model
Dong Sui, Yacong Li, Maozu Guo 0001, Xiangyu Li 0004, Kuanquan Wang, Gongning Luo
Knowl. Based Syst.4
2025 DHAG-DTA: Dynamic Hierarchical Affinity Graph Model for Drug-Target Binding Affinity Prediction
abstract
Computational methods for predicting drug-target binding affinity (DTA) are critical for large-scale screening of prospective therapeutic compounds during drug discovery. Deep neural networks (DNNs) have recently shown significant promise for DTA prediction. By leveraging available data for training, DNNs can expand the use of DTA prediction to situations where only sequence information is available for potential drug molecules and their targets, and there is no prior knowledge regarding the molecular geometric conformations. We propose DHAG-DTA, a general dynamic hierarchical affinity graph DNN approach, for DTA prediction using molecular sequence information and already known drug-target interactions. DHAG-DTA introduces a two-level hierarchical graph structure: at the upper level, interactions between drug and target molecules are represented via an affinity graph and at the lower level, embedded molecular graphs represent interactions within the individual molecules. This allows for integration of information from both inter and intra molecular interactions for DTA prediction, which has also been addressed in other recent independent work. The fundamental innovations introduced by DHAG-DTA include: (a) a single overall hierarchical graph that allows better assimilation of information during the learning process compared with loosely-coupled individual graphs, (b) dynamic determination of the affinity graph structure via the introduction of unlabeled edges and a maximum entropy criterion for active edge selection, (c) skip connections in the DNN for fusing intra and inter molecular information, and (d) fusion of both model-based and similarity-based feature embeddings to get robust embeddings of unseen molecules. Experimental results on two common benchmark datasets demonstrate that DHAG-DTA outperforms other existing models on multiple evaluation metrics, achieving state-of-the-art performance.
Cheng Wang 0049, Yang Liu 0006, Shitao Song, Gaurav Sharma 0001, Maozu Guo 0001
IEEE Trans. Comput. Biol. Bioinform.7
2025 Syn-Net: A Synchronous Frequency-Perception Fusion Network for Breast Tumor Segmentation in Ultrasound Images
abstract
Accurate breast tumor segmentation in ultrasound images is a crucial step in medical diagnosis and locating the tumor region. However, segmentation faces numerous challenges due to the complexity of ultrasound images, similar intensity distributions, variable tumor morphology, and speckle noise. To address these challenges and achieve precise segmentation of breast tumors in complex ultrasound images, we propose a Synchronous Frequency-perception Fusion Network (Syn-Net). Initially, we design a synchronous dual-branch encoder to extract local and global feature information simultaneously from complex ultrasound images. Secondly, we introduce a novel Frequency- perception Cross-Feature Fusion (FrCFusion) Block, which utilizes Discrete Cosine Transform (DCT) to learn all-frequency features and effectively fuse local and global features while mitigating issues arising from similar intensity distributions. In addition, we develop a Full-Scale Deep Supervision method that not only corrects the influence of speckle noise on segmentation but also effectively guides decoder features towards the ground truth. We conduct extensive experiments on three publicly available ultrasound breast tumor datasets. Comparison with 14 state-of-the-art deep learning segmentation methods demonstrates that our approach exhibits superior sensitivity to different ultrasound images, variations in tumor size and shape, speckle noise, and similarity in intensity distribution between surrounding tissues and tumors. On the BUSI and Dataset B datasets, our method achieves better Dice scores compared to state-of-the-art methods, indicating superior performance in ultrasound breast tumor segmentation.
Guangzhe Zhao, Xingguo Zhu, Feihu Yan, Maozu Guo 0001
IEEE J. Biomed. Health Informatics5
2025 Select Your Own Counterparts: Self-Supervised Graph Contrastive Learning With Positive Sampling
abstract
Contrastive learning (CL) has emerged as a powerful approach for self-supervised learning. However, it suffers from sampling bias, which hinders its performance. While the mainstream solutions, hard negative mining (HNM) and supervised CL (SCL), have been proposed to mitigate this critical issue, they do not effectively address graph CL (GCL). To address it, we propose graph positive sampling (GPS) and three contrastive objectives. The former is a novel learning paradigm designed to leverage the inherent properties of graphs for improved GCL models, which utilizes four complementary similarity measurements, including node centrality, topological distance, neighborhood overlapping, and semantic distance, to select positive counterparts for each node. Notably, GPS operates without relying on true labels and enables preprocessing applications. The latter aims to fuse positive samples and enhance representative selection in the semantic space. We release three node-level models with GPS and conduct extensive experiments on public datasets. The results demonstrate the superiority of GPS over state-of-the-art (SOTA) baselines and debiasing methods. In addition, the GPS has also been proven to be versatile, adaptive, and flexible.
Zehong Wang, Donghua Yu, Shigen Shen, Shichao Zhang 0001, Huawen Liu, Shuang Yao, Maozu Guo 0001
IEEE Trans. Neural Networks Learn. Syst.7
2024 MHANDTI: Drug-Target Interaction Prediction Model Based on Heterogeneous Graph Multi-Hop Attention Networks
abstract
Predicting drug-target interactions (DTI) has become an important step in the drug discovery and drug repositioning process. The biological identification of DTI incurs significant financial and temporal costs, and the deployment of computational methodologies for DTI prediction can substantially curtail both the duration and economic expenditure of drug discovery or repositioning. Drawing from disparate data sources can offer a comprehensive overview and diverse perspectives for predicting drug-target interactions. However, existing methods for processing drug and target information are unable to capture long-range dependency across heterogeneous graphs. To capture more comprehensive drug and target features for DTI prediction, this article proposes a drug-target interaction prediction model based on heterogeneous graph multi-hop attention networks, called MHANDTI. Concretely, the word-level deep convolutional neural network is used to obtain the structural features of drugs and targets as attributes of drug and target entity nodes in heterogeneous graphs. For the processing of heterogeneous graphs, MHANDTI designs a multi-hop attention diffusion layer (MHADL) to establish a connection between the non-direct connection nodes, and summarize the characteristics of the adjacent nodes of the center nodes in several hops. A series of results indicate that the performance of this method is superior to the latest existing frameworks.
Linlin Xing, Longbo Zhang, Hongzhen Cai, Maozu Guo 0001
BIBM6
2024 FTMSNet: Towards boundary-aware polyp segmentation framework based on hybrid Fourier Transform and Multi-scale Subtraction
abstract
Colorectal cancer is one of the diseases with the highest incidence and mortality rates worldwide, posing a severe threat to human life and health. Polyps are the primary cause of colorectal cancer, Colonoscopy is the gold standard for diagnosing colorectal polyps. Accurate polyp segmentation is crucial for patient diagnosis, treatment, and prognosis. However, the irregular shapes and low contrast between colorectal polyps and normal tissue make polyp segmentation a challenging task. Although deep learning-based methods have achieved promising results in this task, few approaches focus on the boundary information of colorectal polyps. In this study, to address the issue of edge blurring in polyp segmentation, we propose a boundary-aware polyp segmentation method based on a hybrid Fourier Transform and Multi-scale Convolutional neural network, referred to as FTMSNet. We introduce the Fourier Transform Module (FTM), which utilizes the Fourier Transform to retain only high-frequency information (such as boundary) in the frequency domain. By leveraging boundary information, our method achieves more precise and clearer boundary delineation for colorectal polyp segmentation. Simultaneously, the Multi-scale Feature Denoising Decoder (MFDD) we introduced is devised to mitigate noise interference during multi-scale information fusion. We validated the performance of FTMSNet on five publicly available datasets. Extensive experimental results demonstrate that our approach surpasses the state-of-the-art polyp segmentation methods.
Guangzhe Zhao, Feihu Yan, Maozu Guo 0001
BIBM5
2024 YMamba: A Dual-Branch Network Fusing State Space Model and CNN for Medical Image Segmentation
abstract
Intelligent technologies like deep learning have significantly improved medical image segmentation, enhancing clinical decision-making and reducing healthcare costs. However, CNN-based methods face limitations due to constrained receptive fields and semantic information loss in deeper layers. Consequently, to address these issues, we innovatively propose the YMamba model which employs a parallel hybrid architecture of CNN and VMamba, enhancing the modeling capability of distant features while maintaining local feature detail textures in medical images without introducing additional parameters. Additionally, our proposed DBFM module employs an enhanced attention strategy to integrate the strengths of both methods more effectively, reinforcing image feature representation and mitigating background noise. Finally, the MCFFD module receives shallow and deep features from the fusion module, addressing the challenge of size variation in target segmentation regions within medical images. Extensive experiments demonstrate that YMamba achieves state-of-the-art results on four medical image datasets BUSI, DDTI, TN3K, and ISIC2016.
Guangzhe Zhao, Feihu Yan, Maozu Guo 0001
BIBM5
2024 V-GMR: a variational autoencoder-based heterogeneous graph multi-behavior recommendation model
Haoqin Yang, Ran Rang, Linlin Xing, Longbo Zhang, Hongzhen Cai, Maozu Guo 0001
Appl. Intell.6
2024 Equivariant score-based generative diffusion framework for 3D molecules
abstract
BACKGROUND: Molecular biology is crucial for drug discovery, protein design, and human health. Due to the vastness of the drug-like chemical space, depending on biomedical experts to manually design molecules is exceedingly expensive. Utilizing generative methods with deep learning technology offers an effective approach to streamline the search space for molecular design and save costs. This paper introduces a novel E(3)-equivariant score-based diffusion framework for 3D molecular generation via SDEs, aiming to address the constraints of unified Gaussian diffusion methods. Within the proposed framework EMDS, the complete diffusion is decomposed into separate diffusion processes for distinct components of the molecular feature space, while the modeling processes also capture the complex dependency among these components. Moreover, angle and torsion angle information is integrated into the networks to enhance the modeling of atom coordinates and utilize spatial information more effectively. RESULTS: Experiments on the widely utilized QM9 dataset demonstrate that our proposed framework significantly outperforms the state-of-the-art methods in all evaluation metrics for 3D molecular generation. Additionally, ablation experiments are conducted to highlight the contribution of key components in our framework, demonstrating the effectiveness of the proposed framework and the performance improvements of incorporating angle and torsion angle information for molecular generation. Finally, the comparative results of distribution show that our method is highly effective in generating molecules that closely resemble the actual scenario. CONCLUSION: Through the experiments and comparative results, our framework clearly outperforms previous 3D molecular generation methods, exhibiting significantly better capacity for modeling chemically realistic molecules. The excellent performance of EMDS in 3D molecular generation brings novel and encouraging opportunities for tackling challenging biomedical molecule and protein scenarios.
Hao Zhang 0128, Yang Liu 0006, Cheng Wang 0049, Maozu Guo 0001
BMC Bioinform.5
2024 A MLP-Mixer and mixture of expert model for remaining useful life prediction of lithium-ion batteries
abstract
Abstract Accurately predicting the Remaining Useful Life (RUL) of lithium-ion batteries is crucial for battery management systems. Deep learning-based methods have been shown to be effective in predicting RUL by leveraging battery capacity time series data. However, the representation learning of features such as long-distance sequence dependencies and mutations in capacity time series still needs to be improved. To address this challenge, this paper proposes a novel deep learning model, the MLP-Mixer and Mixture of Expert (MMMe) model, for RUL prediction. The MMMe model leverages the Gated Recurrent Unit and Multi-Head Attention mechanism to encode the sequential data of battery capacity to capture the temporal features and a re-zero MLP-Mixer model to capture the high-level features. Additionally, we devise an ensemble predictor based on a Mixture-of-Experts (MoE) architecture to generate reliable RUL predictions. The experimental results on public datasets demonstrate that our proposed model significantly outperforms other existing methods, providing more reliable and precise RUL predictions while also accurately tracking the capacity degradation process. Our code and dataset are available at the website of github.
Lingling Zhao, Shitao Song, Pengyan Wang, Chunyu Wang 0002, Junjie Wang 0005, Maozu Guo 0001
Frontiers Comput. Sci.6
2024 A Recommendation Approach Based on Heterogeneous Network and Dynamic Knowledge Graph
abstract
Besides data sparsity and cold start, recommender systems often face the problems of selection bias and exposure bias. These problems influence the accuracy of recommendations and easily lead to overrecommendations. This paper proposes a recommendation approach based on heterogeneous network and dynamic knowledge graph (HN-DKG). The main steps include (1) determining the implicit preferences of users according to user’s cross-domain and cross-platform behaviors to form multimodal nodes and then building a heterogeneous knowledge graph; (2) Applying an improved multihead attention mechanism of the graph attention network (GAT) to realize the relationship enhancement of multimodal nodes and constructing a dynamic knowledge graph; and (3) Leveraging RippleNet to discover user’s layered potential interests and rating candidate items. In which, some mechanisms, such as user seed clusters, propagation blocking, and random seed mechanisms, are designed to obtain more accurate and diverse recommendations. In this paper, the public datasets are used to evaluate the performance of algorithms, and the experimental results show that the proposed method has good performance in the effectiveness and diversity of recommendations. On the MovieLens-1M dataset, the proposed model is 18%, 9%, and 2% higher than KGAT on F1, NDCG@10, and AUC and 20%, 2%, and 0.9% higher than RippleNet, respectively. On the Amazon Book dataset, the proposed model is 12%, 3%, and 2.5% higher than NFM on F1, NDCG@10, and AUC and 0.8%, 2.3%, and 0.35% higher than RippleNet, respectively.
Shanshan Wan, Yuquan Wu, Linhu Xiao, Maozu Guo 0001
Int. J. Intell. Syst.5
2024 BAMRE: Joint extraction model of Chinese medical entities and relations based on Biaffine transformation with relation attention
Linlin Xing, Longbo Zhang, Hongzhen Cai, Maozu Guo 0001
J. Biomed. Informatics6
2024 MIMR: Modality-Invariance Modeling and Refinement for unsupervised visible-infrared person re-identification
Zhiqi Pang, Chunyu Wang 0002, Honghu Pan, Lingling Zhao, Junjie Wang 0005, Maozu Guo 0001
Knowl. Based Syst.6
2024 Gradient-Based Local Causal Structure Learning
abstract
Finding the causal structure from a set of variables given observational data is a crucial task in many scientific areas. Most algorithms focus on discovering the global causal graph but few efforts have been made toward the local causal structure (LCS), which is of wide practical significance and easier to obtain. LCS learning faces the challenges of neighborhood determination and edge orientation. Available LCS algorithms build on conditional independence (CI) tests, they suffer the poor accuracy due to noises, various data generation mechanisms, and small-size samples of real-world applications, where CI tests do not work. In addition, they can only find the Markov equivalence class, leaving some edges undirected. In this article, we propose a GradieNt-based LCS learning approach (GraN-LCS) to determine neighbors and orient edges simultaneously in a gradient-descent way, and, thus, to explore LCS more accurately. GraN-LCS formulates the causal graph search as minimizing an acyclicity regularized score function, which can be optimized by efficient gradient-based solvers. GraN-LCS constructs a multilayer perceptron (MLP) to simultaneously fit all other variables with respect to a target variable and defines an acyclicity-constrained local recovery loss to promote the exploration of local graphs and to find out direct causes and effects of the target variable. To improve the efficacy, it applies preliminary neighborhood selection (PNS) to sketch the raw causal structure and further incorporates an$l_{1}$-norm-based feature selection on the first layer of MLP to reduce the scale of candidate variables and to pursue sparse weight matrix. GraN-LCS finally outputs LCS based on the sparse weighted adjacency matrix learned from MLPs. We conduct experiments on both synthetic and real-world datasets and verify its efficacy by comparing against state-of-the-art baselines. A detailed ablation study investigates the impact of key components of GraN-LCS and the results prove their contribution.
Jiaxuan Liang, Jun Wang 0035, Guoxian Yu, Carlotta Domeniconi, Xiangliang Zhang 0001, Maozu Guo 0001
IEEE Trans. Cybern.6
2024 Sentence Bag Graph Formulation for Biomedical Distant Supervision Relation Extraction
abstract
We introduce a novel graph-based framework for alleviating key challenges in distantly-supervised relation extraction and demonstrate its effectiveness in the challenging and important domain of biomedical data. Specifically, we propose a graph view of sentence bags referring to an entity pair, which enables message-passing based aggregation of information related to the entity pair over the sentence bag. The proposed framework alleviates the common problem of noisy labeling in distantly supervised relation extraction and also effectively incorporates inter-dependencies between sentences within a bag. Extensive experiments on two large-scale biomedical relation datasets and the widely utilized NYT dataset demonstrate that our proposed framework significantly outperforms the state-of-the-art methods for biomedical distant supervision relation extraction while also providing excellent performance for relation extraction in the general text mining domain.
Hao Zhang 0128, Yang Liu 0006, Tianming Liang, Gaurav Sharma 0001, Maozu Guo 0001
IEEE Trans. Knowl. Data Eng.7
2024 Personalized Federated Few-Shot Learning
abstract
Personalized federated learning (PFL) learns a personalized model for each client in a decentralized manner, where each client owns private data that are not shared and data among clients are non-independent and identically distributed (i.i.d.) However, existing PFL solutions assume that clients have sufficient training samples to jointly induce personalized models. Thus, existing PFL solutions cannot perform well in a few-shot scenario, where most or all clients only have a handful of samples for training. Furthermore, existing few-shot learning (FSL) approaches typically need centralized training data; as such, these FSL methods are not applicable in decentralized scenarios. How to enable PFL with limited training samples per client is a practical but understudied problem. In this article, we propose a solution called personalized federated few-shot learning (pFedFSL) to tackle this problem. Specifically, pFedFSL learns a personalized and discriminative feature space for each client by identifying which models perform well on which clients, without exposing local data of clients to the server and other clients, and which clients should be selected for collaboration with the target client. In the learned feature spaces, each sample is made closer to samples of the same category and farther away from samples of different categories. Experimental results on four benchmark datasets demonstrate that pFedFSL outperforms competitive baselines across different settings.
Guoxian Yu, Jun Wang 0035, Carlotta Domeniconi, Maozu Guo 0001, Xiangliang Zhang 0001, Li-Zhen Cui 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Reinforcement Causal Structure Learning on Order Graph
abstract
Learning directed acyclic graph (DAG) that describes the causality of observed data is a very challenging but important task. Due to the limited quantity and quality of observed data, and non-identifiability of causal graph, it is almost impossible to infer a single precise DAG. Some methods approximate the posterior distribution of DAGs to explore the DAG space via Markov chain Monte Carlo (MCMC), but the DAG space is over the nature of super-exponential growth, accurately characterizing the whole distribution over DAGs is very intractable. In this paper, we propose Reinforcement Causal Structure Learning on Order Graph (RCL-OG) that uses order graph instead of MCMC to model different DAG topological orderings and to reduce the problem size. RCL-OG first defines reinforcement learning with a new reward mechanism to approximate the posterior distribution of orderings in an efficacy way, and uses deep Q-learning to update and transfer rewards between nodes. Next, it obtains the probability transition model of nodes on order graph, and computes the posterior probability of different orderings. In this way, we can sample on this model to obtain the ordering with high probability. Experiments on synthetic and benchmark datasets show that RCL-OG provides accurate posterior probability approximation and achieves better results than competitive causal discovery algorithms.
Dezhi Yang, Guoxian Yu, Jun Wang 0035, Zhengtian Wu, Maozu Guo 0001
AAAI5
2023 A Novel Effectiveness Assessment Framework for Neoadjuvant Chemoradiotherapy of Locally Advanced Rectal Cancer Based on Multi-modal Intelligence
abstract
Neoadjuvant chemoradiotherapy (nCRT) is the stan-dard treatment for locally advanced rectal cancer (LARC). With the development of artificial intelligence, an increasing number of studies have begun to explore its application in cancer treatment prediction. However, the prior methods exhibit considerable variability even with slight modifications to the input data, which could potentially undermine the reliability of the results. In this paper, we proposed RP-Net, a novel multi-modal fusion-based framework that combines feature information from magnetic resonance imaging (MRI) and whole slide images (WSI), establishing a relationship to map the therapeutic effectiveness of nCRT for LARC. We investigated the relationship of the tumour region and its periphery tissues, and demonstrated the validity of the proposed framework that involving 11 different combinations of modalities. The experimental results revealed that it has achieved higher prediction accuracy compared to the four intra-categories single-modal combinations and outperformed the two intra-categories multi-modal combinations. When compared to the other four inter-categories multi-modal combinations, the fusion features get accuracy of 2 % ~ 6% improvement respectively.
Dong Sui, Weifeng Liu 0010, Maozu Guo 0001, Gongning Luo, Kuanquan Wang
BIBM4
2023 Causal Discovery by Graph Attention Reinforcement Learning
abstract
Discovery the causal structure graph among a set of variables is a fundamental but difficult task in many empirical sciences. Reinforcement learning based causal discovery from observed data achieves prominent results. However, previous algorithms lack interpretability and efficiency, and ignore the prior knowledge of causal structure. To solve these problems, we propose GARL that leverages graph attention network to embed the structure information and the prior knowledge, and reinforcement learning to search the variable ordering with the best score. GARL takes the structure information and prior knowledge as the computational skeleton of attention to obtain the embedded representation of variables, and then generates variable orderings through the designed ordering model. In addition, the structure information is used to form the DAG corresponding to the variable ordering, which reduces the computational difficulty and improves the efficiency. GARL generates DAGs in the reinforcement learning framework, and uses the score of DAG as the reward to optimize the network structure to search the DAG with the best score. Experimental results on synthetic and real datasets show that our GARL has obvious advantages in multi-node operation efficiency, and competitive results with competitive baselines.
Dezhi Yang, Guoxian Yu, Jun Wang 0035, Zhongmin Yan, Maozu Guo 0001
SDM5
2023 Predicting microbe-drug associations with structure-enhanced contrastive learning and self-paced negative sampling strategy
abstract
MOTIVATION: Predicting the associations between human microbes and drugs (MDAs) is one critical step in drug development and precision medicine areas. Since discovering these associations through wet experiments is time-consuming and labor-intensive, computational methods have already been an effective way to tackle this problem. Recently, graph contrastive learning (GCL) approaches have shown great advantages in learning the embeddings of nodes from heterogeneous biological graphs (HBGs). However, most GCL-based approaches don't fully capture the rich structure information in HBGs. Besides, fewer MDA prediction methods could screen out the most informative negative samples for effectively training the classifier. Therefore, it still needs to improve the accuracy of MDA predictions. RESULTS: In this study, we propose a novel approach that employs the Structure-enhanced Contrastive learning and Self-paced negative sampling strategy for Microbe-Drug Association predictions (SCSMDA). Firstly, SCSMDA constructs the similarity networks of microbes and drugs, as well as their different meta-path-induced networks. Then SCSMDA employs the representations of microbes and drugs learned from meta-path-induced networks to enhance their embeddings learned from the similarity networks by the contrastive learning strategy. After that, we adopt the self-paced negative sampling strategy to select the most informative negative samples to train the MLP classifier. Lastly, SCSMDA predicts the potential microbe-drug associations with the trained MLP classifier. The embeddings of microbes and drugs learning from the similarity networks are enhanced with the contrastive learning strategy, which could obtain their discriminative representations. Extensive results on three public datasets indicate that SCSMDA significantly outperforms other baseline methods on the MDA prediction task. Case studies for two common drugs could further demonstrate the effectiveness of SCSMDA in finding novel MDA associations. AVAILABILITY: The source code is publicly available on GitHub https://github.com/Yue-Yuu/SCSMDA-master.
Zhen Tian 0004, Haichuan Fang, Weixin Xie, Maozu Guo 0001
Briefings Bioinform.5
2023 Cooperative driver pathways discovery by multiplex network embedding
abstract
Cooperative driver pathways discovery helps researchers to study the pathogenesis of cancer. However, most discovery methods mainly focus on genomics data, and neglect the known pathway information and other related multi-omics data; thus they cannot faithfully decipher the carcinogenic process. We propose CDPMiner (Cooperative Driver Pathways Miner) to discover cooperative driver pathways by multiplex network embedding, which can jointly model relational and attribute information of multi-type molecules. CDPMiner first uses the pathway topology to quantify the weight of genes in different pathways, and optimizes the relations between genes and pathways. Then it constructs an attributed multiplex network consisting of micro RNAs, long noncoding RNAs, genes and pathways, embeds the network through deep joint matrix factorization to mine more essential information for pathway-level analysis and reconstructs the pathway interaction network. Finally, CDPMiner leverages the reconstructed network and mutation data to define the driver weight between pathways to discover cooperative driver pathways. Experimental results on Breast invasive carcinoma and Stomach adenocarcinoma datasets show that CDPMiner can effectively fuse multi-omics data to discover more driver pathways, which indeed cooperatively trigger cancers and are valuable for carcinogenesis analysis. Ablation study justifies CDPMiner for a more comprehensive analysis of cancer by fusing multi-omics data.
Jun Wang 0035, Zhengtian Wu, Maozu Guo 0001, Guoxian Yu
Briefings Bioinform.4
2023 A gene regulatory network inference model based on pseudo-siamese network
abstract
MOTIVATION: Gene regulatory networks (GRNs) arise from the intricate interactions between transcription factors (TFs) and their target genes during the growth and development of organisms. The inference of GRNs can unveil the underlying gene interactions in living systems and facilitate the investigation of the relationship between gene expression patterns and phenotypic traits. Although several machine-learning models have been proposed for inferring GRNs from single-cell RNA sequencing (scRNA-seq) data, some of these models, such as Boolean and tree-based networks, suffer from sensitivity to noise and may encounter difficulties in handling the high noise and dimensionality of actual scRNA-seq data, as well as the sparse nature of gene regulation relationships. Thus, inferring large-scale information from GRNs remains a formidable challenge. RESULTS: This study proposes a multilevel, multi-structure framework called a pseudo-Siamese GRN (PSGRN) for inferring large-scale GRNs from time-series expression datasets. Based on the pseudo-Siamese network, we applied a gated recurrent unit to capture the time features of each TF and target matrix and learn the spatial features of the matrices after merging by applying the DenseNet framework. Finally, we applied a sigmoid function to evaluate interactions. We constructed two maize sub-datasets, including gene expression levels and GRNs, using existing open-source maize multi-omics data and compared them to other GRN inference methods, including GENIE3, GRNBoost2, nonlinear ordinary differential equations, CNNC, and DGRNS. Our results show that PSGRN outperforms state-of-the-art methods. This study proposed a new framework: a PSGRN that allows GRNs to be inferred from scRNA-seq data, elucidating the temporal and spatial features of TFs and their target genes. The results show the model's robustness and generalization, laying a theoretical foundation for maize genotype-phenotype associations with implications for breeding work.
Maozu Guo 0001
BMC Bioinform.2
2023 Differential Gene Expression Prediction by Ensemble Deep Networks on Histone Modification Data
abstract
Predicting differential gene expression (DGE) from Histone modifications (HM) signal is crucial to understand how HM controls cell functional heterogeneity through influencing differential gene regulation. Most existing prediction methods use fixed-length bins to represent HM signals and transmit these bins into a single machine learning model to predict differential expression genes of single cell type or cell type pair. However, the inappropriate bin length may cause the splitting of the important HM segment and lead to information loss. Furthermore, the bias of single learning model may limit the prediction accuracy. Considering these problems, in this paper, we proposes an Ensemble deep neural networks framework for predicting Differential Gene Expression (EnDGE). EnDGE employs different feature extractors on input HM signal data with different bin lengths and fuses the feature vectors for DGE prediction. Ensemble multiple learning models with different HM signal cutting strategies helps to keep the integrity and consistency of genetic information in each signal segment, and offset the bias of individual models. Besides the popular feature extractors, we also propose a new Residual Network based model with higher prediction accuracy to increase the diversity of feature extractors. Experiments on the real datasets from the Roadmap Epigenome Project (REMC) show that for all cell type pairs, EnDGE significantly outperforms the state-of-the-art baselines for differential gene expression prediction.
Zimo Huang, Jun Wang 0035, Zhongmin Yan, Maozu Guo 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 Distantly-Supervised Long-Tailed Relation Extraction Using Constraint Graphs
abstract
Label noise and long-tailed distributions are two major challenges in distantly supervised relation extraction. Recent studies have shown great progress on denoising, but paid little attention to the problem of long-tailed relations. In this paper, we introduce a constraint graph to model the dependencies between relation labels. On top of that, we further propose a novel constraint graph-based relation extraction framework(CGRE) to handle the two challenges simultaneously. CGRE employs graph convolution networks to propagate information from data-rich relation nodes to data-poor relation nodes, and thus boosts the representation learning of long-tailed relations. To further improve the noise immunity, a constraint-aware attention module is designed in CGRE to integrate the constraint information. Extensive experimental results indicate that CGRE achieves significant improvements over the previous methods for both denoising and long-tailed relation extraction.
Tianming Liang, Yang Liu 0006, Hao Zhang 0128, Gaurav Sharma 0001, Maozu Guo 0001
IEEE Trans. Knowl. Data Eng.6
2023 Directed Acyclic Graph Learning on Attributed Heterogeneous Network
abstract
Learning the directed acyclic graph (DAG) among causal variables is a fundamental pre-task in causal discovery. Available DAG learning solutions canonically focus on homogeneous nodes with multiple variables and assume i.i.d. samples, how to learn DAG on typical attributed heterogeneous network (AHN) composed with different types of inter-dependent nodes and diverse attributes is a practical but more difficult task. In this paper, we propose HetDAG to identify DAG among nodes from heterogeneous network. HetDAG first embeds different types of node attributes and aggregates these embeddings as the node's raw representation. Then it uses contrastive learning with prior network structure to explore latent relationships between nodes and update the representation. Next, HetDAG introduces an attention-based DAG learning module that takes node representations as input to search DAG and orient edges between nodes. To the best of our knowledge, HetDAG is the first study to learn DAG on heterogeneous networks. Extensive experiments on both semi-synthetic and real data show that HetDAG can learn DAG in an efficacy way and outperforms the state-of-the-art approaches. The results on real biological networks confirm that HetDAG can find out the causal relations between lncRNAs and miRNAs.
Jiaxuan Liang, Jun Wang 0035, Guoxian Yu, Wei Guo 0017, Carlotta Domeniconi, Maozu Guo 0001
IEEE Trans. Knowl. Data Eng.6
2022 A new method for predicting plant proteins function based on multi label classification algorithm
abstract
Protein function annotation is an important content of bioinformatics. It is unrealistic to experimentally annotate a large number of protein functions, so automated prediction of the functions of proteins is required. Studying plant proteins can help us cultivate new plant varieties. However, the existing methods mainly aim at predicting the single function of plant proteins. Most proteins only have sequence information. Therefore, the multi-function prediction method of plant protein based on sequence has become the focus of research. In this study, we propose a method based on sequence to predict the function of plant proteins, called PlantGO. We describe the functional annotation problem as a multi-label classification problem using Gene Ontology (GO) terminology, and predict multiple functions for plant proteins. PlantGO extracts features from three aspects and performs feature fusion respectively. To avoid redundant information in features, PlantGO uses the random forest to select features and obtains three groups of optimal features. PlantGO uses a multi-label learning algorithm MLKNN to predict protein function, which proves the effectiveness and accuracy of the algorithm. Finally, PlantGO adopts the ensemble learning method to integrate all models into a unified model, further improving the prediction performance. Therefore, PlantGO has achieved excellent performance. The performance metrics are Accuracy (0.936), F1 Score (0.944), Ranking loss (0.011), which is better than the previous method in the independent test. We have developed an online server to predict the functions of plant proteins.
Juan Wang 0011, Maozu Guo 0001
BIBM3
2022 Flexible ConvNext Block Based Multi-task Learning Framework for Liver MRI Images Analysis
abstract
Liver cancer is the second leading cause of death all over the world in the 2020s’, and the incidence rate has been growing on a global scale and become a serious threat to human life. Early diagnosis of liver cancer from medical images can allow the patients to receive better treatment and achieve good outcomes. Although medical imaging approaches have made significant progress over the past decades, there are still great demands to reconstruct the network structure for the adaptation to downstream tasks. It remains a great challenge for liver tumor identification from MRI images. Recently, self-attention mechanism based transformer models can capture long-range dependencies, which make them perform well on many medical image analysis tasks. Such as Segformer and TransUNet, since lacking the translation in-variance and inductive bias of CNNs, they are still needed large-scale training to fill the gap, especially in the field of medical image analysis domain. In this study, we incorporate a novel flexible Condeathblock as a feature extractor to perform liver images analysis and propose a new analytical framework CXNet to extract discriminative multi-scale visual representations. The experiments results demonstrated our framework outperforms the state-of the-art models on three datasets, including a publicly available liver segmentation dataset as well as two in-house liver tumor classification/segmentation datasets. On the 3DIRCADb dataset, our CXNet outperforms the UNet model by 2.82%, 2.73%, and 4.46% in terms of the Jaccard similarity coefficient, Dice coefficient, and accuracy, and outperforms the TransUNet model by 8.49%, 5.35%,which 3.62%, respectively. Code is available at https://github.com/SPECTRELWF/CXNet
Dong Sui, Weifeng Liu 0010, Maozu Guo 0001, Gongning Luo, Kuanquan Wang
BIBM3
2022 Phenotype Prediction by Heterogeneous Molecular Network Embedding
abstract
Phenotype prediction aims to infer the traits of living organisms based on genomics data, which has important applications in biology such as cancer subtype diagnosis and crop breedings. Traditional phenotype prediction approaches only learn the low-dimensional representation of samples, but ignore the interaction between biomolecules and cannot use the topology structure of heterogeneous molecular networks. Furthermore, most of them lack interpretability and do not effectively identify key biomolecules associated with phenotypes. In this paper, we propose a heterogeneous network embedding based solution (PhenoHNE) to predict phenotype by fusing topology information of molecules. PhenoHNE firstly utilizes variational graph autoencode (VGAE) to obtain the embedding representation of molecules in the heterogeneous network. Secondly, PhenoHNE adopts multilayer perceptron (MLP) with attention mechanism to learn sample representation. Finally, it fuses molecular embedding representation and sample representation to predict the phenotype of samples. In this way, the genetics information of heterogeneous molecular network can further guide the feature learning of samples and improve the prediction performance. Experimental results on human and maize datasets confirm that PhenoHNE outperforms competitive methods by a large margin under different evaluation protocols, and it also can effectively identify the key molecules associated with phenotypes of interests.
Haojiang Tan, Jun Wang 0035, Guoxian Yu, Wei Guo 0017, Maozu Guo 0001
BIBM5
2022 Sub-Loc: predicting protein sub-mitochondrial localization based on sequence embedding
abstract
Mitochondria are subcellular organelles existing in most eukaryotic organisms. They have a pivotal role in lots of bio-chemical processes for cells. Proteins in different compartments of mitochondria have their transport routes. Locating proteins in mitochondria can provide a solid foundation for mitochondrial pathologies. So far, there have been several computational methods for solving the issue. However, their accuracy is so low that they do not identify the localization precisely. We develop an unsupervised learning model to represent mitochondrial proteins as n-dimensional vectors, called sequence embedding, which can learn the global and context information from mitochondrial proteins. We design a new model, called Sub-Loc, to predict the sub-mitochondrial localization of proteins using the SVM classifier and the sequence embedding method. The sequence embedding method and the Sub-Loc are tested by experiments. Experimental results show the sequence embedding method remarkably enhances the performance of prediction for sub-mitochondrial localization of proteins compared with other feature representations. The Sub-Loc outperforms other approaches for predicting sub-mitochondrial localization of proteins.
Juan Wang 0011, Haodong Bian, Maozu Guo 0001
BIBM4
2022 GraphCDA: a hybrid graph representation learning framework based on GCN and GAT for predicting disease-associated circRNAs
abstract
MOTIVATION: CircularRNA (circRNA) is a class of noncoding RNA with high conservation and stability, which is considered as an important disease biomarker and drug target. Accumulating pieces of evidence have indicated that circRNA plays a crucial role in the pathogenesis and progression of many complex diseases. As the biological experiments are time-consuming and labor-intensive, developing an accurate computational prediction method has become indispensable to identify disease-related circRNAs. RESULTS: We presented a hybrid graph representation learning framework, named GraphCDA, for predicting the potential circRNA-disease associations. Firstly, the circRNA-circRNA similarity network and disease-disease similarity network were constructed to characterize the relationships of circRNAs and diseases, respectively. Secondly, a hybrid graph embedding model combining Graph Convolutional Networks and Graph Attention Networks was introduced to learn the feature representations of circRNAs and diseases simultaneously. Finally, the learned representations were concatenated and employed to build the prediction model for identifying the circRNA-disease associations. A series of experimental results demonstrated that GraphCDA outperformed other state-of-the-art methods on several public databases. Moreover, GraphCDA could achieve good performance when only using a small number of known circRNA-disease associations as the training set. Besides, case studies conducted on several human diseases further confirmed the prediction capability of GraphCDA for predicting potential disease-related circRNAs. In conclusion, extensive experimental results indicated that GraphCDA could serve as a reliable tool for exploring the regulatory role of circRNAs in complex diseases.
Qiguo Dai, Zhaowei Wang 0005, Xiaodong Duan, Maozu Guo 0001
Briefings Bioinform.5
2022 Predicting miRNA-disease associations using an ensemble learning framework with resampling method
abstract
MOTIVATION: Accumulating evidences have indicated that microRNA (miRNA) plays a crucial role in the pathogenesis and progression of various complex diseases. Inferring disease-associated miRNAs is significant to explore the etiology, diagnosis and treatment of human diseases. As the biological experiments are time-consuming and labor-intensive, developing effective computational methods has become indispensable to identify associations between miRNAs and diseases. RESULTS: We present an Ensemble learning framework with Resampling method for MiRNA-Disease Association (ERMDA) prediction to discover potential disease-related miRNAs. Firstly, the resampling strategy is proposed for building multiple different balanced training subsets to address the challenge of sample imbalance within the database. Then, ERMDA extracts miRNA and disease feature representations by integrating miRNA-miRNA similarities, disease-disease similarities and experimentally verified miRNA-disease association information. Next, the feature selection approach is applied to reduce the redundant information and increase the diversity among these subsets. Lastly, ERMDA constructs an individual learner on each subset to yield primitive outcomes, and the soft voting method is introduced for making the final decision based on the prediction results of individual learners. A series of experimental results demonstrates that ERMDA outperforms other state-of-the-art methods on both balanced and unbalanced testing sets. Besides, case studies conducted on the three human diseases further confirm the ERMDA's prediction capability for identifying potential disease-related miRNAs. In conclusion, these experimental results demonstrate that our method can serve as an effective and reliable tool for researchers to explore the regulatory role of miRNAs in complex diseases.
Qiguo Dai, Zhaowei Wang 0005, Xiaodong Duan, Jinmiao Song, Maozu Guo 0001
Briefings Bioinform.6
2022 Differentially expressed genes prediction by multiple self-attention on epigenetics data
abstract
Predicting differentially expressed genes (DEGs) from epigenetics signal data is the key to understand how epigenetics controls cell functional heterogeneity by gene regulation. This knowledge can help developing 'epigenetics drugs' for complex diseases like cancers. Most of existing machine learning-based methods suffer defects in prediction accuracy, interpretability or training speed. To address these problems, in this paper, we propose a Multiple Self-Attention model for predicting DEGs on Epigenetic data (Epi-MSA). Epi-MSA first uses convolutional neural networks for neighborhood bins information embedding, and then employs multiple self-attention encoders on different input epigenetics factors data to learn which locations of genes are important for predicting DEGs. Next it trains a soft attention module to pick out which epigenetics factors are significant. The attention mechanism makes the model interpretable, and the pure matrix operation of self-attention enables the model to be parallel calculated and speeds up the training. Experiments on datasets from the Roadmap Epigenome Project and BluePrint Data Analysis Portal (BDAP) show that the performance of Epi-MSA is better than existing competitive methods, and Epi-MSA also has a smaller standard deviation, which shows that Epi-MSA is effective and stable. In addition, Epi-MSA has a good interpretability, this is confirmed by referring its attention weight matrix with existing biological knowledge.
Zimo Huang, Jun Wang 0035, Zhongmin Yan, Maozu Guo 0001
Briefings Bioinform.4
2022 ELSSI: parallel SNP-SNP interactions detection by ensemble multi-type detectors
abstract
With the development of high-throughput genotyping technology, single nucleotide polymorphism (SNP)-SNP interactions (SSIs) detection has become an essential way for understanding disease susceptibility. Various methods have been proposed to detect SSIs. However, given the disease complexity and bias of individual SSI detectors, these single-detector-based methods are generally unscalable for real genome-wide data and with unfavorable results. We propose a novel ensemble learning-based approach (ELSSI) that can significantly reduce the bias of individual detectors and their computational load. ELSSI randomly divides SNPs into different subsets and evaluates them by multi-type detectors in parallel. Particularly, ELSSI introduces a four-stage pipeline (generate, score, switch and filter) to iteratively generate new SNP combination subsets from SNP subsets, score the combination subset by individual detectors, switch high-score combinations to other detectors for re-scoring, then filter out combinations with low scores. This pipeline makes ELSSI able to detect high-order SSIs from large genome-wide datasets. Experimental results on various simulated and real genome-wide datasets show the superior efficacy of ELSSI to state-of-the-art methods in detecting SSIs, especially for high-order ones. ELSSI is applicable with moderate PCs on the Internet and flexible to assemble new detectors. The code of ELSSI is available at https://www.sdu-idea.cn/codes.php?name=ELSSI.
Xia Cao, Yuantao Feng, Maozu Guo 0001, Guoxian Yu, Jun Wang 0035
Briefings Bioinform.4
2022 Isoform function prediction by Gene Ontology embedding
abstract
MOTIVATION: High-resolution annotation of gene functions is a central task in functional genomics. Multiple proteoforms translated from alternatively spliced isoforms from a single gene are actual function performers and greatly increase the functional diversity. The specific functions of different isoforms can decipher the molecular basis of various complex diseases at a finer granularity. Multi-instance learning (MIL)-based solutions have been developed to distribute gene(bag)-level Gene Ontology (GO) annotations to isoforms(instances), but they simply presume that a particular annotation of the gene is responsible by only one isoform, neglect the hierarchical structures and semantics of massive GO terms (labels), or can only handle dozens of terms. RESULTS: We propose an efficacy approach IsofunGO to differentiate massive functions of isoforms by GO embedding. Particularly, IsofunGO first introduces an attributed hierarchical network to model massive GO terms, and a GO network embedding strategy to learn compact representations of GO terms and project GO annotations of genes into compressed ones, this strategy not only explores and preserves hierarchy between GO terms but also greatly reduces the prediction load. Next, it develops an attention-based MIL network to fuse genomics and transcriptomics data of isoforms and predict isoform functions by referring to compressed annotations. Extensive experiments on benchmark datasets demonstrate the efficacy of IsofunGO. Both the GO embedding and attention mechanism can boost the performance and interpretability. AVAILABILITYAND IMPLEMENTATION: The code of IsofunGO is available at http://www.sdu-idea.cn/codes.php?name=IsofunGO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sichao Qiu, Guoxian Yu, Xudong Lu 0001, Carlotta Domeniconi, Maozu Guo 0001
Bioinform.5
2022 Weakly Supervised Cross-Modal Hashing
abstract
Cross-modal hashing can efficiently retrieve data across different modalities and has been successfully applied in various domains. Although many supervised cross-modal hashing methods have been proposed, they generally focus on two modals only and assume that the labels of training data are sufficient and complete. This assumption is not practical in real scenarios. In this article, we propose the weakly supervised cross-modal hashing (WCHash), which takes into account the widely witnessed weakly supervised information (incompleteandinsufficient labels) of training data. Specifically, WCHash first optimizes a latent central modality with respect to other modalities. Next, it uses an efficient multi-label weak-label method to enrich the labels of training data and measures the semantic similarity between data points based on the enriched labels. After that, it uses this similarity to guide the correlation maximization between the respective data modals and the central modal and thus achieves the hash functions for cross-modal retrieval. Experimental results on real-world datasets demonstrate that WCHash is more efficient and effective than related state-of-the-art cross-modal hashing methods. WCHash can significantly reduce the complexity of cross-modal hashing on three or more modalities.
Xuanwu Liu, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Guoqiang Xiao 0001, Maozu Guo 0001
IEEE Trans. Big Data6
2022 EpiMC: Detecting Epistatic Interactions Using Multiple Clusterings
abstract
Detecting single nucleotide polymorphisms (SNPs) interactions is crucial to identify susceptibility genes associated with complex human diseases in genome-wide association studies. Clustering-based approaches are widely used in reducing search space and exploring potential relationships between SNPs in epistasis analysis. However, these approaches all only use a single measure to filter out nonsignificant SNP combinations, which may be significant ones from another perspective. In this paper, we propose a two-stage approach named EpiMC (Epistatic Interactions detection based on Multiple Clusterings) that employs multiple clusterings to obtain more precise candidate sets and more comprehensively detect high-order interactions based on these sets. In the first stage, EpiMC proposes a matrix factorization based multiple clusterings algorithm to generate multiple diverse clusterings, each of which divide all SNPs into different clusters. This stage aims to reduce the chance of filtering out potential candidates overlooked by a single clustering and groups associated SNPs together from different clustering perspectives. In the next stage, EpiMC considers both the single-locus effects and interaction effects to select high-quality disease associated SNPs, and then uses Jaccard similarity to get candidate sets. Finally, EpiMC uses exhaustive search on the obtained small candidate sets to precisely detect epsitatic interactions. Extensive simulation experiments show that EpiMC has a better performance in detecting high-order interactions than state-of-the-art solutions. On the Wellcome Trust Case Control Consortium (WTCCC) dataset, EpiMC detects several significant epistatic interactions associated with breast cancer (BC) and age-related macular degeneration (AMD), which again corroborate the effectiveness of EpiMC.
Jun Wang 0035, Maozu Guo 0001, Guoxian Yu
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 Tissue Specificity Based Isoform Function Prediction
abstract
Alternative splicing enables a gene spliced into different isoforms and hence protein variants. Identifying individual functions of these isoforms help deciphering the functional diversity of proteins. Although much efforts have been made for automatic gene function prediction, few efforts have been moved toward computational isoform function prediction, mainly due to the unavailable (or scanty) functional annotations of isoforms. Existing efforts directly combine multiple RNA-seq datasets without account of the important tissue specificity of alternative splicing. To bridge this gap, we introduce a novel approach called TS-Isofun to predict the functions of isoforms by integrating multiple functional association networks with respect to tissue specificity. TS-Isofun first constructs tissue-specific isoform functional association networks using multiple RNA-seq datasets from tissue-wise. Next, TS-Isofun assigns weights to these networks and models the tissue specificity by selectively integrating them with adaptive weights. It then introduces a joint matrix factorization-based data fusion model to leverage the integrated network, gene-level data and functional annotations of genes to infer the functions of isoforms. To achieve coherent weight assignment and isoform function prediction, TS-Isofun jointly optimizes the weights of individual networks and the isoform function prediction in a unified objective function. Experimental results show that TS-Isofun significantly outperforms state-of-the-art methods and the account of tissue specificity contributes to more accurate isoform function prediction.
Guoxian Yu, Qiuyue Huang, Xiangliang Zhang 0001, Maozu Guo 0001, Jun Wang 0035
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 DeepIDA: Predicting Isoform-Disease Associations by Data Fusion and Deep Neural Networks
abstract
Alternative splicing produces different isoforms from the same gene locus, it is an important mechanism for regulating gene expression and proteome diversity. Although the prediction of gene(ncRNA)-disease associations has been extensively studied, few (or no) computational solutions have been proposed for the prediction of isoform-disease association (IDA) at a large scale, mainly due to the lack of disease annotations of isoforms. However, increasing evidences confirm the associations between diseases and isoforms, which can more precisely uncover the pathology of complex diseases. Therefore, it is highly desirable to predict IDAs. To bridge this gap, we propose a deep neural network based solution (DeepIDA) to fuse multi-type genomics and transcriptomics data to predict IDAs. Particularly, DeepIDA uses gene-isoform relations to dispatch gene-disease associations to isoforms. In addition, it utilizes two DNN sub-networks with different structures to capture nucleotide and expression features of isoforms, Gene Ontology data and miRNA target data, respectively. After that, these two sub-networks are merged in a dense layer to predict IDAs. The experimental results on public datasets show that DeepIDA can effectively predict IDAs with AUPRC (area under the precision-recall curve) of 0.9141, macro F-measure of 0.9155, G-mean of 0.9278 and balanced accuracy of 0.9303 across 732 diseases, which are much higher than those of competitive methods. Further study on sixteen isoform-disease association cases again corroborates the superiority of DeepIDA. The code of DeepIDA is available at http://mlda.swu.edu.cn/codes.php?name=DeepIDA.
Guoxian Yu, Yeqian Yang, Yangyang Yan, Maozu Guo 0001, Xiangliang Zhang 0001, Jun Wang 0035
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 DeepRCI: Predicting ATP-Binding Proteins Using the Residue-Residue Contact Information
abstract
Adenine-5'-triphosphate (ATP) is a direct energy source for various activities of tissues and cells in the body. The release of ATP energies requires the assistance of ATP-binding proteins. Therefore, the identification of ATP-binding proteins is of great significance for the research on organisms. So far, there are several methods for predicting ATP-binding proteins. However, the accuracies of these methods are so low that the predicted proteins are inaccurate. Here, we designed a novel method, called as DeepRCI (based on Deep convolutional neural network and Residue-residue Contact Information), for predicting ATP-binding proteins. In order to maximize the performance of our method, we experimented with different hyperparameters and finally chose a 12-depth-512-filters deep convolutional neural network with an input size of 448*448. By using this model, DeepRCI achieved an accuracy of 93.61% on the test set which means a significant improvement of 11.78% over the state-of-the-art methods. We also compared the performance of residue-residue contact information datasets with different noise levels which are mainly due to gaps in the multiple sequence alignment. Compared with the low-noise dataset, the prediction accuracy on the high-noise dataset is reduced by 6.78%, which affects the performance of DeepRCI to a certain extent. We believe that with the increase of sequence data, this problem will eventually be solved. Finally, we provide a web service of DeepRCI which link can be obtained in Data Availability.
Yulan Zhao, Juan Wang 0011, Maozu Guo 0001
IEEE J. Biomed. Health Informatics4
2021 Genome-Phenome Association Prediction by Deep Factorizing Heterogeneous Molecular Network
abstract
Genome-phenome association (GPA) play a crucial part in deciphering the complex pathology of phenotypes (i.e., traits and diseases). Heterogeneous network-based GPA solutions can model the complex connections between multi-types of molecules by viewing them as nodes, and give better results than using single/two-omics data alone, but they overwhelmingly need to project other molecules toward homogeneous gene/phenotype nodes for data fusion and prediction, such projections result in information loss. Matrix factorization based data fusion can avoid such projection by integrating multi-type data in a coherent way, but they typically perform linear factorization and cannot mine the nonlinear relationships between molecules, which compromise the GPA analysis. Furthermore, most of them can not synergy network topology and node attribution information in a principle way. In this paper, we propose a deep matrix factorization based solution (DeepGPA) to predict GPAs by fusing heterogeneous molecular network and diverse attributes of nodes. DeepGPA performs deep matrix factorization on the block adjacency matrices of heterogeneous network in a cooperative manner to obtain the nonlinear representations of different moleclues. In addition, it performs low-rank representation learning on the attribute data with the shared nonlinear representations. In this way, both the network topology and node attributes are jointly mined to explore the representations of molecules and complex interplays between molecules and phenotypes. DeepGPA then uses the representational vectors of gene and phenotype nodes to predict GPAs. Experimental results on Maize datasets confirm t hat Deep DPA out performs competitive methods by a large margin under different evaluation protocols.
Haojiang Tan, Sichao Qiu, Jun Wang 0035, Guoxian Yu, Wei Guo 0017, Maozu Guo 0001
BIBM6
2021 Maize Epistasis Detection by Multi-class Quantitative Multifactor Dimensionality Reduction
abstract
Maize (Zea mays ssp. mays) is one of the most important food crops in the world, it is critical to explore the genetic architecture for improving yield in Maize. Existing methods focusing on the associations of single loci and phenotype may cause the missing heritablity. Furthermore, it is common that we only have partial phenotype information. In this paper, we transform maize epistasis detection into a correlation feature selection problem on multi-class data, and propose a Multi-class Quantitative Multifactor Dimensionality Reduction method called Epi-MQMDR to detect high order SNP interactions related to maize phenotype. First, we use semi-supervised learning to predict the trait values of samples with unknown phenotypes to enrich the genetic information. Then, we introduce an MDR-based algorithm for multi-class samples to detect SNP interactions associated with quantitative trait. We further discretize continuous phenotypic values of samples as labels and construct contingency tables from the classification results of SNP combination genotypes and sample labels to detect epistasis. Experiments on simulated models and real Zea mays datasets prove the efficacy of Epi-MQMDR on epistasis detection.
Jun Wang 0035, Guoxian Yu, Beibei Xin, Maozu Guo 0001
BIBM5
2021 Drug-Target Interaction Prediction Based on Gaussian Interaction Profile and Information Entropy
Lina Liu 0011, Shuang Yao, Zhaoyun Ding, Maozu Guo 0001, Donghua Yu, Keli Hu
ISBRA4
2021 Pathogenic gene prediction based on network embedding
abstract
In disease research, the study of gene-disease correlation has always been an important topic. With the emergence of large-scale connected data sets in biology, we use known correlations between the entities, which may be from different sets, to build a biological heterogeneous network and propose a new network embedded representation algorithm to calculate the correlation between disease and genes, using the correlation score to predict pathogenic genes. Then, we conduct several experiments to compare our method to other state-of-the-art methods. The results reveal that our method achieves better performance than the traditional methods.
Yang Liu 0006, Chunyu Wang 0002, Maozu Guo 0001
Briefings Bioinform.5
2021 Integrating multi-scale neighbouring topologies and cross-modal similarities for drug-protein interaction prediction
abstract
MOTIVATION: Identifying the proteins that interact with drugs can reduce the cost and time of drug development. Existing computerized methods focus on integrating drug-related and protein-related data from multiple sources to predict candidate drug-target interactions (DTIs). However, multi-scale neighboring node sequences and various kinds of drug and protein similarities are neither fully explored nor considered in decision making. RESULTS: We propose a drug-target interaction prediction method, DTIP, to encode and integrate multi-scale neighbouring topologies, multiple kinds of similarities, associations, interactions related to drugs and proteins. We firstly construct a three-layer heterogeneous network to represent interactions and associations across drug, protein, and disease nodes. Then a learning framework based on fully-connected autoencoder is proposed to learn the nodes' low-dimensional feature representations within the heterogeneous network. Secondly, multi-scale neighbouring sequences of drug and protein nodes are formulated by random walks. A module based on bidirectional gated recurrent unit is designed to learn the neighbouring sequential information and integrate the low-dimensional features of nodes. Finally, we propose attention mechanisms at feature level, neighbouring topological level and similarity level to learn more informative features, topologies and similarities. The prediction results are obtained by integrating neighbouring topologies, similarities and feature attributes using a multiple layer CNN. Comprehensive experimental results over public dataset demonstrated the effectiveness of our innovative features and modules. Comparison with other state-of-the-art methods and case studies of five drugs further validated DTIP's ability in discovering the potential candidate drug-related proteins.
Ping Xuan, Hui Cui 0002, Tiangang Zhang, Maozu Guo 0001, Toshiya Nakaguchi
Briefings Bioinform.5
2021 DMIL-IsoFun: predicting isoform function using deep multi-instance learning
abstract
MOTIVATION: Alternative splicing creates the considerable proteomic diversity and complexity on relatively limited genome. Proteoforms translated from alternatively spliced isoforms of a gene actually execute the biological functions of this gene, which reflect the functional knowledge of genes at a finer granular level. Recently, some computational approaches have been proposed to differentiate isoform functions using sequence and expression data. However, their performance is far from being desirable, mainly due to the imbalance and lack of annotations at isoform-level, and the difficulty of modeling gene-isoform relations. RESULT: We propose a deep multi-instance learning-based framework (DMIL-IsoFun) to differentiate the functions of isoforms. DMIL-IsoFun firstly introduces a multi-instance learning convolution neural network trained with isoform sequences and gene-level annotations to extract the feature vectors and initialize the annotations of isoforms, and then uses a class-imbalance Graph Convolution Network to refine the annotations of individual isoforms based on the isoform co-expression network and extracted features. Extensive experimental results show that DMIL-IsoFun improves the Smin and Fmax of state-of-the-art solutions by at least 29.6% and 40.8%. The effectiveness of DMIL-IsoFun is further confirmed on a testbed of human multiple-isoform genes, and maize isoforms related with photosynthesis. AVAILABILITY AND IMPLEMENTATION: The code and data are available at http://www.sdu-idea.cn/codes.php?name=DMIL-Isofun. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Guoxian Yu, Guangjie Zhou, Xiangliang Zhang 0001, Carlotta Domeniconi, Maozu Guo 0001
Bioinform.5
2021 Imbalance deep multi-instance learning for predicting isoform-isoform interactions
abstract
Multi-instance learning (MIL) can model complex bags (samples) that are further made of diverse instances (subsamples). In typical MIL, the labels of bags are known while those of individual instances are unknown and to be specified. In this paper we propose an imbalanced deep multi-instance learning approach (IDMIL-III) and apply it to predict genome-wide isoform–isoform interactions (IIIs). This prediction task is crucial for precisely understanding the interactome between proteoforms and to reveal their functional diversity. The current solutions typically formulate the prediction of IIIs as a MIL problem by pairing two genes as a “bag” and any two isoforms spliced from these two genes as “instances.” The key instances (interacting isoform pairs) trigger the label of the positive (interacting) gene bags, which is important for identifying the IIIs. Furthermore, the prediction task was simplified as a balanced classification problem, which in practice is a rather imbalanced one. To address these issues, IDMIL-III fuses RNA-seq, nucleotide sequence, amino acid sequence and exon array data, and further introduces a novel loss function to separately model the loss of positive pairs and of negative pairs, and thus to avoid the expected loss dominated by majority negative pairs. In addition, it includes an attention strategy to identify positive isoform pairs from a positive gene bag. Extensive experimental results prove the effectiveness of IDMIL-III on predicting IIIs. Particularly, IDMIL-III achieves an F1 value as 95.4%, at least 3.8% higher than those of competitive methods at the gene-level; and obtains an F1 as 29.8%, at least 2.4% higher than the state-of-the-art methods at the isoform-level. The code of IDMIL-III is available at http://mlda.swu.edu.cn/codes.php?name=IDMIL-III.
Guoxian Yu, Jun Wang 0035, Hong Zhang 0030, Xiangliang Zhang 0001, Maozu Guo 0001
Int. J. Intell. Syst.6
2021 Noise-robust Deep Cross-Modal Hashing
Guoxian Yu, Hong Zhang 0030, Maozu Guo 0001, Li-Zhen Cui 0001, Xiangliang Zhang 0001
Inf. Sci.4
2021 CDPath: Cooperative Driver Pathways Discovery Using Integer Linear Programming and Markov Clustering
abstract
Discovering driver pathways is an essential task to understand the pathogenesis of cancer and to design precise treatments for cancer patients. Increasing evidences have been indicating that multiple pathways often function cooperatively in carcinogenesis. In this study, we propose an approach called CDPath to discover cooperative driver pathways. CDPath first uses Integer Linear Programming to explore driver core modules from mutation profiles by enforcing co-occurrence and functional interaction relations between modules, and by maximizing the mutual exclusivity and coverage within modules. Next, to enforce cooperation of pathways and help the follow-up exact cooperative driver pathways discovery, it performs Markov clustering on pathway-pathway interaction network to cluster pathways. After that, it identifies pathways in different modules but in the same clusters as cooperative driver pathways. We apply CDPath on two TCGA datasets: breast cancer (BRCA) and endometrial cancer (UCEC). The results show that CDPath can identify known (i.e., TP53) and potential driver genes (i.e., SPTBN2). In addition, the identified cooperative driver pathways are related with the target cancer, and they are involved with carcinogenesis and several key biological processes. CDPath can uncover more potential biological associations between pathways (over 100 percent) and more cooperative driver pathways (over 200 percent) than competitive approaches. The demo codes of CDPath are available at http://mlda.swu.edu.cn/codes.php?name=CDPath.
Ziying Yang, Guoxian Yu, Maozu Guo 0001, Jiantao Yu, Xiangliang Zhang 0001, Jun Wang 0035
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 Cross-Species Protein Function Prediction with Asynchronous-Random Walk
abstract
Protein function prediction is a fundamental task in the post-genomic era. Available functional annotations of proteins are incomplete and the annotations of two homologous species are complementary to each other. However, how to effectively leverage mutually complementary annotations of different species to further boost the prediction performance is still not well studied. In this paper, we propose a cross-species protein function prediction approach by performing Asynchronous Random Walk on a heterogeneous network (AsyRW). AsyRW first constructs a heterogeneous network to integrate multiple functional association networks derived from different biological data, established homology-relationships between proteins from different species, known annotations of proteins and Gene Ontology (GO). To account for the intrinsic structures of intra- and inter-species of proteins and that of GO, AsyRW quantifies the individual walk lengths of each network node using the gravity-like theory, and then performs asynchronous-random walk with the individual length to predict associations between proteins and GO terms. Experiments on annotations archived in different years show that individual walk length and asynchronous-random walk can effectively leverage the complementary annotations of different species, AsyRW has a significantly improved performance to other related and competitive methods. The codes of AsyRW are available at: http://mlda.swu.edu.cn/codes.php?name=AsyRW.
Yingwen Zhao, Jun Wang 0035, Maozu Guo 0001, Xiangliang Zhang 0001, Guoxian Yu
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 CrowdWT: Crowdsourcing via Joint Modeling of Workers and Tasks
abstract
Crowdsourcing is a relatively inexpensive and efficient mechanism to collect annotations of data from the open Internet. Crowdsourcing workers are paid for the provided annotations, but the task requester usually has a limited budget. It is desirable to wisely assign the appropriate task to the right workers, so the overall annotation quality is maximized while the cost is reduced. In this article, we propose a novel task assignment strategy (CrowdWT) to capture the complex interactions between tasks and workers, and properly assign tasks to workers. CrowdWT first develops a Worker Bias Model (WBM) to jointly model the worker’s bias, the ground truths of tasks, and the task features. WBM constructs a mapping between task features and worker annotations to dynamically assign the task to a group of workers, who are more likely to give correct annotations for the task. CrowdWT further introduces a Task Difficulty Model (TDM), which builds a Kernel ridge regressor based on task features to quantify the intrinsic difficulty of tasks and thus to assign the difficult tasks to more reliable workers. Finally, CrowdWT combines WBM and TDM into a unified model to dynamically assign tasks to a group of workers and recall more reliable and even expert workers to annotate the difficult tasks. Our experimental results on two real-world datasets and two semi-synthetic datasets show that CrowdWT achieves high-quality answers within a limited budget, and has the best performance against competitive methods.<?vsp -1.5pt?>
Jinzheng Tu 0002, Guoxian Yu, Jun Wang 0035, Carlotta Domeniconi, Maozu Guo 0001, Xiangliang Zhang 0001
ACM Trans. Knowl. Discov. Data5
2020 A Stacked Ensemble Learning Framework with Heterogeneous Feature Combinations for Predicting ncRNA-Protein Interaction
abstract
The interaction between ncRNA and protein is a kind of crucial molecular activities in a cell. Developing computational methods to predict ncRNA-protein interactions has attracted increasing attentions in recent years. In this work, a novel stacked ensemble learning framework is presented for predicting ncRNA-protein interaction based on heterogeneous feature combinations, named HFC-RPI. Firstly, the compositional features of k-mer with different orders were extracted from the primary sequence and secondary structure of RNA and protein respectively. Secondly, we trained a set of base learners using a variety of heterogeneous combinations of the extracted features respectively. Thirdly, the prediction results of these base learners were employed to train the stacked learner, which output the final prediction result at the higher layer in HFC-RPI. Moreover, in order to improve the generalization of HFC-RPI, when training the base learners, a cross-validation based method was applied. Extensive experimental results showed that the proposed learning framework HFC-RPI was effective and feasible for predicting the interaction of ncRNA and protein. By comparing with state-of-the-art methods, HFC-RPI was superior to them on most performance evaluation metrics.
Qiguo Dai, Zhaowei Wang 0005, Jinmiao Song, Xiaodong Duan, Maozu Guo 0001, Zhen Tian 0004
BIBM5
2020 Cooperative Driver Pathway Discovery by Hierarchical Clustering and Link Prediction
abstract
Identifying driver pathway is a critical step to uncover the natural laws of the occurrence and progression of disease. Many studies show that multiple pathways often function cooperatively in carcinogenesis. However, how to computationally identify cooperative driver pathways of cancers is not well studied yet. Existing cooperative driver pathway identification methods either suffer from single type of genetic information source or computation difficulty. In this paper, we proposed a method (CDPLP) based on hierarchical clustering and link prediction. CDPLP firstly devises a new similarity metric to quantity the exclusivity and co-expression of two gene modules, and thus to obtain gene sets with exclusivity by hierarchical clustering. Next, it uses link prediction on the pathway-pathway interaction network to replenish the interactions between pathways. After that, CDPLP combines the gene sets and updated pathway network to discover the pathway pairs with high functional interaction and occurrence as cooperative pathways. CDPLP can make full use of multiple genetic information sources such as the mutation data, gene-gene interaction data and pathway-pathway network, and facilitate the optimization solution. We evaluated the performance of CDPLP on TCGA breast cancer (BRCA) dataset and compared it with other popular methods. The results show that cooperative driver pathways identified by CDPLP are highly associated with the target cancer, and are involved with carcinogenesis and several key biological processes.
Sufang Li, Jun Wang 0035, Maozu Guo 0001, Xiangliang Zhang 0001
BIBM3
2020 Epistasis Detection using Heterogeneous Bio-molecular Network
abstract
Detecting epistasis between single nucleotide poly-morphisms (SNPs) is crucial to explain the missing heritability of complex diseases in genome-wide association studies (GWAS). Many methods have been proposed for detecting SNP interactions, most of them only focus on reducing search space but ignore the relations of SNP with other bio-molecules. In this paper, we proposed a heterogeneous molecular network based method called EpiNet to detect high-order SNP interactions. EpiNet firstly uses samples (case/control) data to construct an SNP statistical network to capture the SNP distribution information in samples. In addition, EpiNet applies meta-path based similarity search in a heterogeneous molecular network composed with SNPs, genes, lncRNAs, miRNAs and diseases to construct an SNP relational network, which mines diverse associations between molecules and diseases to supplement the SNP statistical network and search the most associated SNPs. Next, EpiNet combines the two networks to get a composite SNP network, utilizes modularity based clustering to group all SNPs into different clusters and finally detects SNP interactions within each cluster. Simulation experiments on two-locus and three-locus disease models show that EpiNet has a better performance than competitive methods, even without the heterogeneous network. It also shows expressive power to identify high-order SNP interactions from real WTCCC breast cancer data.
Jun Wang 0035, Guoxian Yu, Li-Zhen Cui 0001, Maozu Guo 0001
BIBM5
2020 Partial Multi-label Learning with Label and Feature Collaboration
Guoxian Yu, Jun Wang 0035, Maozu Guo 0001
DASFAA (1)4
2020 EpIntMC: Detecting Epistatic Interactions Using Multiple Clusterings
Guoxian Yu, Maozu Guo 0001, Jun Wang 0035
ISBRA4
2020 Isoform function prediction based on bi-random walks on a heterogeneous network
abstract
MOTIVATION: Alternative splicing contributes to the functional diversity of protein species and the proteoforms translated from alternatively spliced isoforms of a gene actually execute the biological functions. Computationally predicting the functions of genes has been studied for decades. However, how to distinguish the functional annotations of isoforms, whose annotations are essential for understanding developmental abnormalities and cancers, is rarely explored. The main bottleneck is that functional annotations of isoforms are generally unavailable and functional genomic databases universally store the functional annotations at the gene level. RESULTS: We propose IsoFun to accomplish Isoform Function prediction based on bi-random walks on a heterogeneous network. IsoFun firstly constructs an isoform functional association network based on the expression profiles of isoforms derived from multiple RNA-seq datasets. Next, IsoFun uses the available Gene Ontology annotations of genes, gene-gene interactions and the relations between genes and isoforms to construct a heterogeneous network. After this, IsoFun performs a tailored bi-random walk on the heterogeneous network to predict the association between GO terms and isoforms, thus accomplishing the prediction of GO annotations of isoforms. Experimental results show that IsoFun significantly outperforms the state-of-the-art algorithms and improves the area under the receiver-operating curve (AUROC) and the area under the precision-recall curve (AUPRC) by 17% and 44% at the gene-level, respectively. We further validated the performance of IsoFun on the genes ADAM15 and BCL2L1. IsoFun accurately differentiates the functions of respective isoforms of these two genes. AVAILABILITY AND IMPLEMENTATION: The code of IsoFun is available at http://mlda.swu.edu.cn/codes.php? name=IsoFun. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Guoxian Yu, Keyao Wang, Carlotta Domeniconi, Maozu Guo 0001, Jun Wang 0035
Bioinform.4
2020 Predicting functions of maize proteins using graph convolutional network
abstract
BACKGROUND: Maize (Zea mays ssp. mays L.) is the most widely grown and yield crop in the world, as well as an important model organism for fundamental research of the function of genes. The functions of Maize proteins are annotated using the Gene Ontology (GO), which has more than 40000 terms and organizes GO terms in a direct acyclic graph (DAG). It is a huge challenge to accurately annotate relevant GO terms to a Maize protein from such a large number of candidate GO terms. Some deep learning models have been proposed to predict the protein function, but the effectiveness of these approaches is unsatisfactory. One major reason is that they inadequately utilize the GO hierarchy. RESULTS: To use the knowledge encoded in the GO hierarchy, we propose a deep Graph Convolutional Network (GCN) based model (DeepGOA) to predict GO annotations of proteins. DeepGOA firstly quantifies the correlations (or edges) between GO terms and updates the edge weights of the DAG by leveraging GO annotations and hierarchy, then learns the semantic representation and latent inter-relations of GO terms in the way by applying GCN on the updated DAG. Meanwhile, Convolutional Neural Network (CNN) is used to learn the feature representation of amino acid sequences with respect to the semantic representations. After that, DeepGOA computes the dot product of the two representations, which enable to train the whole network end-to-end coherently. Extensive experiments show that DeepGOA can effectively integrate GO structural information and amino acid information, and then annotates proteins accurately. CONCLUSIONS: Experiments on Maize PH207 inbred line and Human protein sequence dataset show that DeepGOA outperforms the state-of-the-art deep learning based methods. The ablation study proves that GCN can employ the knowledge of GO and boost the performance. Codes and datasets are available at http://mlda.swu.edu.cn/codes.php?name=DeepGOA .
Guangjie Zhou, Jun Wang 0035, Xiangliang Zhang 0001, Maozu Guo 0001, Guoxian Yu
BMC Bioinform.4
2020 Feature selection with missing labels based on label compression and local feature correlation
Guoxian Yu, Maozu Guo 0001, Jun Wang 0035
Neurocomputing3
2020 Weakly supervised semantic segmentation by iterative superpixel-CRF refinement with initial clues guiding
Yang Liu 0006, GuoJun Liu, Maozu Guo 0001
Neurocomputing4
2020 Multi-label crowd consensus via joint matrix factorization
Jinzheng Tu 0002, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Guoqiang Xiao 0001, Maozu Guo 0001
Knowl. Inf. Syst.6
2020 NMFGO: Gene Function Prediction via Nonnegative Matrix Factorization with Gene Ontology
abstract
Gene Ontology (GO) is a controlled vocabulary of terms that describe molecule function, biological roles, and cellular locations of gene products (i.e., proteins and RNAs), it hierarchically organizes more than 43,000 GO terms via the direct acyclic graph. A gene is generally annotated with several of these GO terms. Therefore, accurately predicting the association between genes and massive terms is a difficult challenge. To combat with this challenge, we propose an matrix factorization based approach called NMFGO. NMFGO stores the available GO annotations of genes in a gene-term association matrix and adopts an ontological structure based taxonomic similarity measure to capture the GO hierarchy. Next, it factorizes the association matrix into two low-rank matrices via nonnegative matrix factorization regularized with the GO hierarchy. After that, it employs a semantic similarity based k nearest neighbor classifier in the low-rank matrices approximated subspace to predict gene functions. Empirical study on three model species (S. cerevisiae, H. sapiens, and A. thaliana) shows that NMFGO is robust to the input parameters and achieves significantly better prediction performance than GIC, TO, dRW- kNN, and NtN, which were re-implemented based on the instructions of the original papers. The supplementary file and demo codes of NMFGO are available at http://mlda.swu.edu.cn/codes.php?name=NMFGO.
Guoxian Yu, Keyao Wang, Guangyuan Fu, Maozu Guo 0001, Jun Wang 0035
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 Accurate and Efficient Indoor Pathfinding Based on Building Information Modeling Data
abstract
Pathfinding is a fundamental problem for many areas, e.g., robotics, automation, computer-aided design, and computer graphics. Although outdoor pathfinding is fledged, indoor pathfinding remains a challenge due to the lack of indoor maps. Currently, some efforts have utilized building information modeling (BIM) to generate either the grid-based map or the topological map. However, either the grid-based map or the topological map is not sufficient to provide accurate and efficient pathfinding service. This article proposes a novel grid-topological map and develops an accurate and efficient indoor pathfinding scheme based on BIM. The grid-topological map is modeled jointly adopting the advantages of both the grid-based map and the topological map. First, the grid-based map is generated using the BIM data by extracting and mapping geometric and semantic data into planar grids. Second, a grid thinning algorithm is proposed to produce the topological map directly using the grid-based map. Third, a grid-topological map is presented by combining both the grid-based map and topological map. On top of the grid-topological map, an accurate and efficient pathfinding algorithm is developed. Empirical studies proved the effectiveness of the grid-topological map, as well as the accuracy and efficiency of the proposed indoor pathfinding algorithm.
Qingsheng Xie, Maozu Guo 0001, Jichao Zhao, Jia Wang 0015
IEEE Trans. Ind. Informatics3
2019 Ranking-Based Deep Cross-Modal Hashing
abstract
Cross-modal hashing has been receiving increasing interests for its low storage cost and fast query speed in multi-modal data retrievals. However, most existing hashing methods are based on hand-crafted or raw level features of objects, which may not be optimally compatible with the coding process. Besides, these hashing methods are mainly designed to handle simple pairwise similarity. The complex multilevel ranking semantic structure of instances associated with multiple labels has not been well explored yet. In this paper, we propose a ranking-based deep cross-modal hashing approach (RDCMH). RDCMH firstly uses the feature and label information of data to derive a semi-supervised semantic ranking list. Next, to expand the semantic representation power of hand-crafted features, RDCMH integrates the semantic ranking information into deep cross-modal hashing and jointly optimizes the compatible parameters of deep feature representations and of hashing functions. Experiments on real multi-modal datasets show that RDCMH outperforms other competitive baselines and achieves the state-of-the-art performance in cross-modal retrieval applications.
Xuanwu Liu, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Yazhou Ren 0001, Maozu Guo 0001
AAAI6
2019 Multiple Independent Subspace Clusterings
abstract
Multiple clustering aims at discovering diverse ways of organizing data into clusters. Despite the progress made, it’s still a challenge for users to analyze and understand the distinctive structure of each output clustering. To ease this process, we consider diverse clusterings embedded in different subspaces, and analyze the embedding subspaces to shed light into the structure of each clustering. To this end, we provide a two-stage approach called MISC (Multiple Independent Subspace Clusterings). In the first stage, MISC uses independent subspace analysis to seek multiple and statistical independent (i.e. non-redundant) subspaces, and determines the number of subspaces via the minimum description length principle. In the second stage, to account for the intrinsic geometric structure of samples embedded in each subspace, MISC performs graph regularized semi-nonnegative matrix factorization to explore clusters. It additionally integrates the kernel trick into matrix factorization to handle non-linearly separable clusters. Experimental results on synthetic datasets show that MISC can find different interesting clusterings from the sought independent subspaces, and it also outperforms other related and competitive approaches on real-world datasets.
Jun Wang 0035, Carlotta Domeniconi, Guoxian Yu, Guoqiang Xiao 0001, Maozu Guo 0001
AAAI6
2019 Multi-View Multi-Instance Multi-Label Learning Based on Collaborative Matrix Factorization
abstract
Multi-view Multi-instance Multi-label Learning (M3L) deals with complex objects encompassing diverse instances, represented with different feature views, and annotated with multiple labels. Existing M3L solutions only partially explore the inter or intra relations between objects (or bags), instances, and labels, which can convey important contextual information for M3L. As such, they may have a compromised performance.\ In this paper, we propose a collaborative matrix factorization based solution called M3Lcmf. M3Lcmf first uses a heterogeneous network composed of nodes of bags, instances, and labels, to encode different types of relations via multiple relational data matrices. To preserve the intrinsic structure of the data matrices, M3Lcmf collaboratively factorizes them into low-rank matrices, explores the latent relationships between bags, instances, and labels, and selectively merges the data matrices. An aggregation scheme is further introduced to aggregate the instance-level labels into bag-level and to guide the factorization. An empirical study on benchmark datasets show that M3Lcmf outperforms other related competitive solutions both in the instance-level and bag-level prediction.
Yuying Xing, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Zili Zhang 0001, Maozu Guo 0001
AAAI6
2019 CoPath: discovering cooperative driver pathways using greedy mutual exclusivity and bi-clustering
abstract
The carcinogenesis is typically involved with the cooperations of multiple pathways. Therefore, discovering cooperative driver pathways can provide more precise therapy to patients. Existing cooperative driver pathway identification methods can only identify few (or previously well-known) driver pathways, because of noisy or single type knowledge, or of the insufficient attention to pathway cooperations. To address these problems, we develop a novel approach called CoPath that leverages genomic alteration profiles and bi-clustering to discover cooperative driver pathways. Based on the mutation profiles reconstructed from somatic mutations and copy number variations, CoPath firstly uses a greedy search on the gene signaling network to identify mutually exclusive modules of genes, which have common downstream events in the network. Next, it incorporates the identified mutually exclusive modules and gene interaction network as prior knowledge, and introduces a dual regularized bi-clustering method on gene expression data to cluster cooperative mutually exclusive modules. Finally, it identifies the modules in the same cluster as cooperative driver pathways. Extensive experiments on real cancer genomics dataset (Breast from TCGA) shows that CoPath can not only detect individual driver pathways but also uncover more relations between potential driver genes and pathways than related competitive approaches.
Ziying Yang, Guoxian Yu, Jiantao Yu, Maozu Guo 0001, Jun Wang 0035
BIBM4
2019 DMIL-III: Isoform-isoform interaction prediction using deep multi-instance learning method
abstract
Alternative splicing modulates protein-protein and other ligand interactions, it results in proteoforms, translated from isoforms that are alternatively spliced from the same gene, to interact with different partners and have distinct or even opposing functions. Therefore, systematically identifying protein-protein interaction at the isoform-level is crucial to explore the function of proteoforms. Constructing the isoform-level interaction network currently is prohibited by the lack of a large golden set of experimentally validated interacting isoforms, which enable computationally predicting isoform-isoform interactions. In this paper, a deep convolution neural network based multi-instance learning approach called DMIL-III is proposed to predict isoform interactions. DMIL-III takes a gene pair as `bag' and two isoforms of the pairwise genes as the `instance' of the bag. DMIL-III follows the principle of multi-instance learning that at least one isoform-isoform interaction exists for a positive gene pair and none interacting isoforms occurs for a negative gene pair. DMIL-III integrates RNA-seq, nucleotide sequence, domain-domain interaction and exon array data. Experimental results indicate that DMIL-III achieves a superior performance with Accuracy of 93% on single-instance gene bags and of 94% on multi-instance gene bags, which are at least 14% and 29% higher than those of state-of-the-art methods. In addition, we further test DMIL-III on a set of experimentally confirmed isoform-isoform interactions and obtain an Accuracy of 65%, which is at least 10% higher than those of comparing methods at the isoform-level. All these results show the effectiveness of DMIL-III for predicting isoform-isoform interactions.
Guoxian Yu, Jun Wang 0035, Maozu Guo 0001, Xiangliang Zhang 0001
BIBM4
2019 Selective Matrix Factorization for Multi-relational Data Fusion
Yuehui Wang, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Xiangliang Zhang 0001, Maozu Guo 0001
DASFAA (1)6
2019 Discovering Multiple Co-Clusterings in Subspaces
abstract
Multiple clustering approaches aim at exploring alternative ways of organizing a given collection of data into various clusters from different perspectives. Although multiple one-way clusterings have been studied for more than a decade, how to explore alternative two-way clusterings (or co-clusterings) still remains an untouched topic, and an important one from an application standpoint. To solve this interesting but yet unexplored topic, we assume the existence of alternative co-clusterings embedded in different subspaces and simultaneously pursue multiple co-clusterings therein. We initially specify a subspace indicator matrix for each feature subspace, and employ matrix tri-factorization to seek row-wise and column-wise cluster indicator matrices in each subspace. To ensure diversity, we quantify the redundancy between pairwise co-clusterings using the cluster indicator and the subspace indicator matrices. We further introduce a unified objective function to simultaneously account for the two pursues, and an alternating optimization solution to iteratively optimize cluster indicator and feature indicator matrices. Our empirical study shows that the proposed solution can explore multiple meaningful co-clusterings and generally achieves better results than state-of-the-art methods.
Shixin Yao, Guoxian Yu, Jun Wang 0035, Carlotta Domeniconi, Maozu Guo 0001
SDM6
2019 A review of metrics measuring dissimilarity for rooted phylogenetic networks
abstract
A rooted phylogenetic network is an important structure in the description of evolutionary relationships. Computing the distance (topological dissimilarity) between two rooted phylogenetic networks is a fundamental in phylogenic analysis. During the past few decades, several polynomial-time computable metrics have been described. Here, we give a comprehensive review and analysis on those metrics, including the correlation among metrics and the distribution of distance values computed by each metric. Moreover, we describe the software and website, CDRPN (Computing Distance for Rooted Phylogenetic Networks), for measuring the topological dissimilarity between rooted phylogenetic networks. AVAILABILITY: http://bioinformatics.imu.edu.cn/distance/. CONTACT: [email protected].
Juan Wang 0011, Maozu Guo 0001
Briefings Bioinform.2
2019 Drug repositioning based on individual bi-random walks on a heterogeneous network
abstract
BACKGROUND: Traditional drug research and development is high cost, time-consuming and risky. Computationally identifying new indications for existing drugs, referred as drug repositioning, greatly reduces the cost and attracts ever-increasing research interests. Many network-based methods have been proposed for drug repositioning and most of them apply random walk on a heterogeneous network consisted with disease and drug nodes. However, these methods generally adopt the same walk-length for all nodes, and ignore the different contributions of different nodes. RESULTS: In this study, we propose a drug repositioning approach based on individual bi-random walks (DR-IBRW) on the heterogeneous network. DR-IBRW firstly quantifies the individual work-length of random walks for each node based on the network topology and knowledge that similar drugs tend to be associated with similar diseases. To account for the inner structural difference of the heterogeneous network, it performs bi-random walks with the quantified walk-lengths, and thus to identify new indications for approved drugs. Empirical study on public datasets shows that DR-IBRW achieves a much better drug repositioning performance than other related competitive methods. CONCLUSIONS: Using individual random walk-lengths for different nodes of heterogeneous network indeed boosts the repositioning performance. DR-IBRW can be easily generalized to prioritize links between nodes of a network.
Yuehui Wang, Maozu Guo 0001, Yazhou Ren 0001, Lianyin Jia, Guoxian Yu
BMC Bioinform.2
2019 OutDet: an algorithm for extracting the outer surfaces of building information models for integration with geographic information systems
abstract
The integration of Building Information Modelling (BIM) and geographic information systems (GIS) is a promising but challenging topic to solve problems in construction industry. However, loading and rendering rich BIM geometric data and large-scale GIS spatial information in a unified system is still technologically challenging. Current efforts mainly simplify the geometry in BIM models, or convert BIM geometric data to a lower level of detail (LOD). By noticing that only exterior features of BIM models are visible from outdoor observation points, culling BIM interior facilities can dramatically reduce the computational burden when visualizing BIM models in GIS. This study explores the outline detection problem and presents the OutDet algorithm, which selects representative observation points, transforms & projects the BIM geometric data into the same coordinate system, and detects the visible facilities. Empirical study results show that OutDet can cull a large portion of unnecessary features when rendering BIM models in GIS. The use of outlines of BIM models is not an alternative but rather a supplementary approach for current solutions. Jointly using LOD and outer surface can help improve the efficiency of integrated BIM-GIS visualization.Because OutDet retains BIM geometry and semantics, it can be applied to more BIM-GIS applications..
Jichao Zhao, Jia Wang 0015, Dingding Su, Maozu Guo 0001
Int. J. Geogr. Inf. Sci.7
2019 Drug-target interaction data cluster analysis based on improving the density peaks clustering algorithm
abstract
Since drug-target data have neither class labels nor the cluster number information, they are not suitable for clustering algorithms that require predefined parameters determined by comparing clustering results with real class labels. Density peaks clustering (DPC) is a density-based clustering alg orithm that can determine the number of clusters without requiring class labels. However, the predefined cutoff distance of local density limits its wide application. Therefore, this paper proposes an improved local density method based on a cutoff distance sequence that overcomes the limitations of DPC and can be successful applied to drug-target data. We also introduce multiple-dimensional scaling based on drug and target similarity and perform intuitive graph analysis of the two most significant differentiation features. Drugs of the Enzyme, GPCR, Ion Channel, and Nuclear Receptor 4 standard datasets are identified as 6, 6, 3, and 5 clusters by an improved algorithm, respectively, and similarly, their targets are identified be 5, 5, 8, and 4 clusters. Drug-target data clustering results of the improved algorithm are more reasonable than the results of the fast K-medoids and hierarchical clustering algorithms.
Maozu Guo 0001, Donghua Yu, GuoJun Liu, Shuang Cheng
Intell. Data Anal.1
2019 Variational inference with Gaussian mixture model and householder flow
GuoJun Liu, Yang Liu 0006, Maozu Guo 0001
Neural Networks3
2019 Details in the evaluation of circular RNA detection tools: Reply to Chen and Chuang
abstract
Chia-Ying Chen and Trees-Juen Chuang (referred as CYC & TJC below) recently submitted their comment [1] on our previous paper [2].In their paper, they scrutinized the CircBase [3] candidates that we used and pointed out several weak points of our paper.In summary, they suggested that the positive dataset we derived from CircBase required further evaluation.They also indicated that using all of these candidates as our dataset was not appropriate.They further suggested that three main confounding factors may affect our assessment of circRNA detection tools and that their performances should be re-evaluated.Before we begin to discuss their comment, we will briefly introduce the positive dataset we used.First, as stated in our previous paper, the 14,689 candidates detected in HeLa cells were downloaded from CircBase and reported by the study of Salzman et al. [4].These candidates were not identified with the use of find_circ [5] tool.As described in the study of Salzman et al. [4], all UCSC annotated exons in scrambled order were used to construct a custom database and identify circRNA candidates.Second, in our positive dataset, constant coverage of 10× for the intervening sequence and a minimum of two read pairs (paired-end simulated reads) to cross the back-spliced junction sites were generated for each candidate.Now, we will discuss the three confounding factors they listed in their paper.First, they suggested to remove 1046 candidates with unannotated exon boundaries from the positive dataset, especially candidates without canonical splice signals, such as GT-AG, GC-AG, or AT-AC, for the junctions.As mentioned above, CircBase-deposited circRNA candidates that we used were identified by Salzman et al. [4]; the candidates identified by their method should all match the exon boundaries.The discrepancies may be caused by inconsistent gene annotation files used.Salzman et al. [4] used UCSC known genes [6], whereas CYC & TJC used NCBI RefSeq-identified mRNA annotation files.We manually checked several candidates marked with "junctions with unannotated exon boundaries" in CYC & TJC's Supplemental Dataset S1.The junction sites of these candidates were annotated as exon boundaries in UCSC known genes annotation file (http://hgdownload.soe.ucsc.edu/goldenPath/hg19/database/knownGene.txt.gz).Thus, detection of circRNAs with annotated exon boundaries relies on the gene annotation files used, and novel candidates may be missed because of the incompleteness of the current database [7].For example, Szabo et al. [7] reinforced an annotation-based algorithm with a de novo module and discovered a validated circRNA from the not-fully-annotated RMST gene and several U12 cir-cRNAs produced from unannotated boundaries.Such case was also demonstrated by Xiao-Ou Zhang et al. [8].They detected thousands of novel exons (non-RefSeq, non-Ensembl, or non-UCSC known genes) in circRNAs by using an updated CIRCexplorere2 tool, and several of them were confirmed by Northern blot analysis and Sanger sequencing after RT-PCR [8].Other examples were shown by Salzman et al. [4], they found several noncoding RNA genes expressed
Xiangxiang Zeng, Maozu Guo 0001, Quan Zou 0001
PLoS Comput. Biol.3
2018 Weighted matrix factorization based data fusion for predicting lncRNA-disease associations
Guoxian Yu, Yuehui Wang, Jun Wang 0035, Guangyuan Fu, Maozu Guo 0001, Carlotta Domeniconi
BIBM5
2018 Multi-label Answer Aggregation Based on Joint Matrix Factorization
abstract
Crowdsourcing is a useful and economic approach to data annotation. To obtain annotation of high quality, various aggregation approaches have been developed, which take into account different factors that impact the quality of aggregated answers. However, existing methods generally focus on single-label (multi-class and binary) tasks, and they ignore the inter-correlation between labels, and thus may have compromised quality. In this paper, we introduce a Multi-Label answer aggregation approach based on Joint Matrix Factorization (ML-JMF). ML-JMF selectively and jointly factorizes the sample-label association matrices collected from different annotators into products of individual and shared low-rank matrices. As such, it takes advantage of the robustness of low-rank matrix approximation to noise, and reduces the impact of unreliable annotators by assigning small (zero) weights to their annotation matrices. In addition, it takes advantage of the correlation among labels by leveraging the shared low-rank matrix, and of the similarity between annotators using the individual low-rank matrices to guide the factorization. ML-JMF pursues the low-rank matrices via a unified objective function, and introduces an iterative technique to optimize it. ML-JMF finally uses the optimized low-rank matrices and weights to infer the ground-truth labels. Our experimental results on multi-label datasets show that ML-JMF outperforms competitive methods in inferring ground truth labels. Our approach can identify unreliable annotators, and is robust against their misleading answers through the assignment of small (zero) weights to their annotation.
Jinzheng Tu 0002, Guoxian Yu, Carlotta Domeniconi, Jun Wang 0035, Guoqiang Xiao 0001, Maozu Guo 0001
ICDM6
2018 Active Framework by Sparsity Exploitation for Constructing a Training Set
Maozu Guo 0001, Weining Wu, Yang Liu 0006
ICIC (1)1
2018 Predicting protein-protein interactions using high-quality non-interacting pairs
abstract
BACKGROUND: Identifying protein-protein interactions (PPIs) is of paramount importance for understanding cellular processes. Machine learning-based approaches have been developed to predict PPIs, but the effectiveness of these approaches is unsatisfactory. One major reason is that they randomly choose non-interacting protein pairs (negative samples) or heuristically select non-interacting pairs with low quality. RESULTS: To boost the effectiveness of predicting PPIs, we propose two novel approaches (NIP-SS and NIP-RW) to generate high quality non-interacting pairs based on sequence similarity and random walk, respectively. Specifically, the known PPIs collected from public databases are used to generate the positive samples. NIP-SS then selects the top-m dissimilar protein pairs as negative examples and controls the degree distribution of selected proteins to construct the negative dataset. NIP-RW performs random walk on the PPI network to update the adjacency matrix of the network, and then selects protein pairs not connected in the updated network as negative samples. Next, we use auto covariance (AC) descriptor to encode the feature information of amino acid sequences. After that, we employ deep neural networks (DNNs) to predict PPIs based on extracted features, positive and negative examples. Extensive experiments show that NIP-SS and NIP-RW can generate negative samples with higher quality than existing strategies and thus enable more accurate prediction. CONCLUSIONS: The experimental results prove that negative datasets constructed by NIP-SS and NIP-RW can reduce the bias and have good generalization ability. NIP-SS and NIP-RW can be used as a plugin to boost the effectiveness of PPIs prediction. Codes and datasets are available at http://mlda.swu.edu.cn/codes.php?name=NIP .
Guoxian Yu, Maozu Guo 0001, Jun Wang 0035
BMC Bioinform.3
2018 An improved K-medoids algorithm based on step increasing and optimizing medoids
Donghua Yu, GuoJun Liu, Maozu Guo 0001
Expert Syst. Appl.3
2018 Identification and prioritization of differentially expressed genes for time-series gene expression data
Linlin Xing, Maozu Guo 0001, Chunyu Wang 0002
Frontiers Comput. Sci.2
2018 Weakly supervised semantic segmentation based on EM algorithm with localization clues
Yang Liu 0006, GuoJun Liu, Deming Zhai, Maozu Guo 0001
Neurocomputing5
2018 Parametric local multiview hamming distance metric learning
Deming Zhai, Xianming Liu 0005, Hong Chang 0001, Yi Zhen, Xilin Chen 0001, Maozu Guo 0001, Wen Gao 0001
Pattern Recognit.6
2017 Refine gene functional similarity network based on interaction networks
abstract
BACKGROUND: In recent years, biological interaction networks have become the basis of some essential study and achieved success in many applications. Some typical networks such as protein-protein interaction networks have already been investigated systematically. However, little work has been available for the construction of gene functional similarity networks so far. In this research, we will try to build a high reliable gene functional similarity network to promote its further application. RESULTS: Here, we propose a novel method to construct and refine the gene functional similarity network. It mainly contains three steps. First, we establish an integrated gene functional similarity networks based on different functional similarity calculation methods. Then, we construct a referenced gene-gene association network based on the protein-protein interaction networks. At last, we refine the spurious edges in the integrated gene functional similarity network with the help of the referenced gene-gene association network. Experiment results indicate that the refined gene functional similarity network (RGFSN) exhibits a scale-free, small world and modular architecture, with its degrees fit best to power law distribution. In addition, we conduct protein complex prediction experiment for human based on RGFSN and achieve an outstanding result, which implies it has high reliability and wide application significance. CONCLUSIONS: Our efforts are insightful for constructing and refining gene functional similarity networks, which can be applied to build other high quality biological networks.
Zhen Tian 0004, Maozu Guo 0001, Chunyu Wang 0002
BMC Bioinform.2
2017 A comprehensive overview and evaluation of circular RNA detection tools
abstract
Circular RNA (circRNA) is mainly generated by the splice donor of a downstream exon joining to an upstream splice acceptor, a phenomenon known as backsplicing. It has been reported that circRNA can function as microRNA (miRNA) sponges, transcriptional regulators, or potential biomarkers. The availability of massive non-polyadenylated transcriptomes data has facilitated the genome-wide identification of thousands of circRNAs. Several circRNA detection tools or pipelines have recently been developed, and it is essential to provide useful guidelines on these pipelines for users, including a comprehensive and unbiased comparison. Here, we provide an improved and easy-to-use circRNA read simulator that can produce mimicking backsplicing reads supporting circRNAs deposited in CircBase. Moreover, we compared the performance of 11 circRNA detection tools on both simulated and real datasets. We assessed their performance regarding metrics such as precision, sensitivity, F1 score, and Area under Curve. It is concluded that no single method dominated on all of these metrics. Among all of the state-of-the-art tools, CIRI, CIRCexplorer, and KNIFE, which achieved better balanced performance between their precision and sensitivity, compared favorably to the other methods.
Xiangxiang Zeng, Maozu Guo 0001, Quan Zou 0001
PLoS Comput. Biol.3
2016 Epistasis detection using a permutation-based Gradient Boosting Machine
abstract
Detecting single nucleotide polymorphism (SNP) epistasis contributes to understand disease susceptibility and discover disease pathogenesis underlying complex disease. In this paper, we propose an approach called permutation-based Gradient Boosting Machine (pGBM) to detect pure epistasis by estimating the power of a GBM classifier which is influenced by permuting SNP pairs. pGBM is based on two permutation strategies and gradient boosting machine model. To extend pGBM to detect pure epistasis well on unbalanced dataset, average AUC difference value is chosen as the metric that quantifies the SNP interactions intensity. The experiment results demonstrate that our method has a high success rate with both balanced/unbalanced simulation and real dataset. In addition, pGBM shows great potential to detect pure SNP epistasis to uncover more complex disease pathogenesis.
Kai Che, Maozu Guo 0001, Lei Wang 0085, Yin Zhang 0009
BIBM3
2016 Revealing protein functions based on relationships of interacting proteins and GO terms
abstract
numerous computational methods predicted protein function based on the protein-protein interaction (PPI) network. These methods supposed that two proteins share the same function if they interact with each other. However, it is reported by recent studies that the functions of two interacting proteins may be related but different. In this paper, the functional relationship between interacting proteins is studied and a novel method, called as GoDIN, is advanced to annotate functions of the interacting protein in Gene Ontology (GO) context. It is assumed that the functional difference between interacting proteins can be expressed by semantic difference between GO term and its relatives. Thus, the method uses GO term and its relatives to annotate the interacting proteins separately according to their functional roles in the PPI network. The method is validated by a series of experiments and compared with the related method. The experimental results confirm the assumption and suggest that GoDIN is effective on predicting functions of protein.
Zhixia Teng, Maozu Guo 0001, Zhen Tian 0004, Kai Che
BIBM2
2016 Constructing an integrated gene similarity network for the identification of disease genes
abstract
Discovering novel genes that are involved in human diseases is a challenging task. In recent years, several computational approaches have been proposed to prioritize candidate disease genes. Most of these methods are mainly based on protein-protein interaction (PPI) networks. However, since these PPI networks contain false positives and only cover less half of known human genes, their reliability and coverage are both very low. Therefore, it is highly necessary to fuse multiple genomic data to construct a reliable gene similarity network and then infer disease genes on the whole genomic scale. Here, we proposed a novel method, named RWRB, to infer causal genes of interested disease. First, we construct five individual gene (protein) similarity networks based on multiple genomic data of human genes. Then, an integrated gene similarity network (IGSN) is reconstructed based on similarity network fusion (SNF) method. Finally, we employ the random walk with restart algorithm on the phenotype-gene bilayer network, which combines phenotype similarity network, IGSN as well as the phenotype-gene association network, to prioritize candidate disease genes. We investigate the effectiveness of RWRB through leave-one-out cross-validation methods in inferring phenotype-gene relationships. Results show that RWRB is more accurate than state-of-the-art methods on most evaluation metrics. Further analysis shows that the success of RWRB is benefited from IGSN which has a wider coverage and higher reliability comparing with current PPI networks.
Zhen Tian 0004, Maozu Guo 0001, Chunyu Wang 0002, Linlin Xing, Lei Wang 0085, Yin Zhang 0009
BIBM2
2016 Reconstructing gene regulatory network based on candidate auto selection method
abstract
The reconstruction of gene regulatory network (GRN) is a great challenge in systems biology and bioinformatics, and methods based on Bayesian network (BN) draw most of attention because of its inherent probability characteristics. As NP-hard problems, most of the BN methods often adopt the heuristic search, but they are time-consuming for biological networks with a large number of nodes. To solve this problem, this paper presents a Candidate Auto Selection algorithm (CAS) based on mutual information and breakpoint detection to limit the search space in order to accelerate the learning process. The proposed algorithm automatically restricts the neighbors of each node to a small set of candidates before structure learning. Then based on CAS algorithm, we propose a globally optimal greedy search method (CAS+G), which focuses on finding the high-scoring network structure, and a local learning method (CAS+L), which focuses on faster learning the structure with small loss of quality. Results show that the proposed CAS algorithm can effectively identify the neighbor nodes of each node. In the experiments, the CAS+G method outperforms the state-of-the-art method on simulation data for inferring GRNs, and the CAS+L method is significantly faster than the state-of-the-art method with little loss of accuracy. Hence, the CAS based algorithms are more suitable for GRN inference.
Linlin Xing, Maozu Guo 0001, Chunyu Wang 0002, Lei Wang 0085, Yin Zhang 0009
BIBM2
2016 SGFSC: speeding the gene functional similarity calculation based on hash tables
abstract
BACKGROUND: In recent years, many measures of gene functional similarity have been proposed and widely used in all kinds of essential research. These methods are mainly divided into two categories: pairwise approaches and group-wise approaches. However, a common problem with these methods is their time consumption, especially when measuring the gene functional similarities of a large number of gene pairs. The problem of computational efficiency for pairwise approaches is even more prominent because they are dependent on the combination of semantic similarity. Therefore, the efficient measurement of gene functional similarity remains a challenging problem. RESULTS: To speed current gene functional similarity calculation methods, a novel two-step computing strategy is proposed: (1) establish a hash table for each method to store essential information obtained from the Gene Ontology (GO) graph and (2) measure gene functional similarity based on the corresponding hash table. There is no need to traverse the GO graph repeatedly for each method with the help of the hash table. The analysis of time complexity shows that the computational efficiency of these methods is significantly improved. We also implement a novel Speeding Gene Functional Similarity Calculation tool, namely SGFSC, which is bundled with seven typical measures using our proposed strategy. Further experiments show the great advantage of SGFSC in measuring gene functional similarity on the whole genomic scale. CONCLUSIONS: The proposed strategy is successful in speeding current gene functional similarity calculation methods. SGFSC is an efficient tool that is freely available at http://nclab.hit.edu.cn/SGFSC . The source code of SGFSC can be downloaded from http://pan.baidu.com/s/1dFFmvpZ .
Zhen Tian 0004, Chunyu Wang 0002, Maozu Guo 0001, Zhixia Teng
BMC Bioinform.3
2016 A robust local sparse coding method for image classification with Histogram Intersection Kernel
Yang Liu 0006, GuoJun Liu, Maozu Guo 0001, Zhiyong Pan
Neurocomputing4
2016 MiRTDL: A Deep Learning Approach for miRNA Target Prediction
abstract
MicroRNAs (miRNAs) regulate genes that are associated with various diseases. To better understand miRNAs, the miRNA regulatory mechanism needs to be investigated and the real targets identified. Here, we present miRTDL, a new miRNA target prediction algorithm based on convolutional neural network (CNN). The CNN automatically extracts essential information from the input data rather than completely relying on the input dataset generated artificially when the precise miRNA target mechanisms are poorly known. In this work, the constraint relaxing method is first used to construct a balanced training dataset to avoid inaccurate predictions caused by the existing unbalanced dataset. The miRTDL is then applied to 1,606 experimentally validated miRNA target pairs. Finally, the results show that our miRTDL outperforms the existing target prediction algorithms and achieves significantly higher sensitivity, specificity and accuracy of 88.43, 96.44, and 89.98 percent, respectively. We also investigate the miRNA target mechanism, and the results show that the complementation features are more important than the others.
Shuang Cheng, Maozu Guo 0001, Chunyu Wang 0002, Yang Liu 0006, Xuejian Wu
IEEE ACM Trans. Comput. Biol. Bioinform.2
2015 Topic Network: Topic Model with Deep Learning for Image Classification
abstract
As a representative deep learning model, Convolutional Neural Networks (CNNs) can provide good features to represent the objects in image, and has made a great achievement in image classification and object detection. However, CNNs requires resizing the input images to a fixed size, which may affect the performance of the model due to information loss and distortion. To overcome the limitation, we replace the last pooling layer with topic model-LDA (Latent Dirichlet Allocation) to get a fixed-size output without resizing the input images, and we call it Topic Network. With Topic Network, the input images can be images of an arbitrary size and ratio without resizing, but the output is a k-dimension vector which represents the distribution of topics in image (k is the number of topics). Topic Network performs well in image classification task on Caltech101 and VOC2007 datasets.
Zhiyong Pan, Yang Liu 0006, GuoJun Liu, Maozu Guo 0001
KSEM4
2015 HAlign: Fast multiple similar DNA/RNA sequence alignment based on the centre star strategy
abstract
Abstract Motivation: Multiple sequence alignment (MSA) is important work, but bottlenecks arise in the massive MSA of homologous DNA or genome sequences. Most of the available state-of-the-art software tools cannot address large-scale datasets, or they run rather slowly. The similarity of homologous DNA sequences is often ignored. Lack of parallelization is still a challenge for MSA research. Results: We developed two software tools to address the DNA MSA problem. The first employed trie trees to accelerate the centre star MSA strategy. The expected time complexity was decreased to linear time from square time. To address large-scale data, parallelism was applied using the hadoop platform. Experiments demonstrated the performance of our proposed methods, including their running time, sum-of-pairs scores and scalability. Moreover, we supplied two massive DNA/RNA MSA datasets for further testing and research. Availability and implementation: The codes, tools and data are accessible free of charge at http://datamining.xmu.edu.cn/software/halign/. Contact: [email protected] or [email protected]
Quan Zou 0001, Qinghua Hu, Maozu Guo 0001, Guohua Wang 0001
Bioinform.3
2015 Harmonious competition learning for Gaussian mixtures
GuoJun Liu, Xianglong Tang, Maozu Guo 0001, Yang Liu 0006
Neurocomputing3
2015 Actively constructing an effective training set by expected gain maximization criterion
Weining Wu, Shaobin Huang, Maozu Guo 0001
Neurocomputing3
2014 Identification of functional miRNA regulatory modules and their associations via dynamic miRNA regulatory function
abstract
MicroRNAs (miRNAs) are small non-coding RNAs which cause target genes degradation or translational inhibition. Constructing functional miRNAs regulatory module can be a significant step towards the discovery of their regulatory roles in various development programs. In this paper, we present a Correlated Correspondence Regulatory Module model which builds on modified Correlated Topic Model (CTM). We apply the proposed method to the expression profiles of miRNAs and genes on 89 human cancer samples. The approach computationally predicts miRNA-gene interactions according to the negative or positive correlation relationship between miRNA and gene expression data and identifies functional miRNA regulatory modules from which we can infer multiple and dynamic miRNA function according to the known elements, the result shows consistency with published literature and database. Furthermore, a miRNA regulatory network is constructed in order to study the associations among various regulatory modules, we can detect evolution of miRNA function in biological process according these associations, which solve restriction of traditional methods that only focus on static miRNA function in single regulatory module. Online services can be accessed at the website (http://nclab.hit.edu.cn/CCRM).
Shuang Cheng, Maozu Guo 0001, Chunyu Wang 0002, Yang Liu 0006
BIBM2
2014 Hidden conditional random field for lung nodule detection
abstract
Lung nodule detection in thin section computerized tomography (CT) images is a useful but challenging task in the development of computer aided diagnosis (CAD) system for lung cancer. In order to improve sensitivity and reduce false positive, we consider a 3D nodule as a 2D region of interest (ROI) sequence and utilize a discriminative sequence model called hidden conditional random field to capture the correlations and transitions of a nodule's ROIs on several consecutive slices. First, we use region growing and thresholding to segment lung parenchyma. Second, selective enhancement filter is employed on 2D images to get 2D ROIs and after that, we match these ROIs on consecutive images based on a simple but effective criteria to get 2D ROI sequence(3D candidate) of a nodule. Third, given these ROI sequences, hidden conditional random field is devised to classify whether some 3D candidates are nodules or not based on these sequences. The proposed system is validated on 24 patients' scans which contain 59 nodules in total from Lung Image Database Consortium (LIDC) dataset. Experimental results demonstrate that our approach achieves high sensitivity and reduces false positive significantly.
Yang Liu 0006, Maozu Guo 0001, Ping Li 0013
ICIP3
2014 Inferring the soybean (Glycine max) microRNA functional network based on target gene network
abstract
MOTIVATION: The rapid accumulation of microRNAs (miRNAs) and experimental evidence for miRNA interactions has ushered in a new area of miRNA research that focuses on network more than individual miRNA interaction, which provides a systematic view of the whole microRNome. So it is a challenge to infer miRNA functional interactions on a system-wide level and further draw a miRNA functional network (miRFN). A few studies have focused on the well-studied human species; however, these methods can neither be extended to other non-model organisms nor take fully into account the information embedded in miRNA-target and target-target interactions. Thus, it is important to develop appropriate methods for inferring the miRNA network of non-model species, such as soybean (Glycine max), without such extensive miRNA-phenotype associated data as miRNA-disease associations in human. RESULTS: Here we propose a new method to measure the functional similarity of miRNAs considering both the site accessibility and the interactive context of target genes in functional gene networks. We further construct the miRFNs of soybean, which is the first study on soybean miRNAs on the network level and the core methods can be easily extended to other species. We found that miRFNs of soybean exhibit a scale-free, small world and modular architecture, with their degrees fit best to power-law and exponential distribution. We also showed that miRNA with high degree tends to interact with those of low degree, which reveals the disassortativity and modularity of miRFNs. Our efforts in this study will be useful to further reveal the soybean miRNA-miRNA and miRNA-gene interactive mechanism on a systematic level. AVAILABILITY AND IMPLEMENTATION: A web tool for information retrieval and analysis of soybean miRFNs and the relevant target functional gene networks can be accessed at SoymiRNet: http://nclab.hit.edu.cn/SoymiRNet.
Yungang Xu, Maozu Guo 0001, Chunyu Wang 0002, Yang Liu 0006
Bioinform.2
2014 CPL: Detecting Protein Complexes by Propagating Labels on Protein-Protein Interaction Network
Qiguo Dai, Maozu Guo 0001, Zhixia Teng, Chunyu Wang 0002
J. Comput. Sci. Technol.2
2013 MLPA: Detecting overlapping communities by multi-label propagation approach
abstract
The identification of communities is an important step in understanding of the complex network. Comparative studies suggest that the development of accurate and efficient methods to infer the communities is still in its early stages. Label propagation algorithm (LPA) that detects communities by propagating labels among vertices, attracts a great deal of attention recently. However, the communities detected by most LPAs are disjointed. Due to communities are often overlapping in real world networks, we show a multi-label propagation algorithm (MLPA) to detect overlapping communities. The inspiration is that the more people are familiar, the more they trust each other. To simulate the confidence of human communication, propagating intensity (PI) is defined to describe the confidence extent of the label propagated by neighboring vertices. The PI is then used to guide the propagation, with the purpose to make the detection more accurate. The results of extensive experiments both on synthetic and real networks show that the proposed MLPA outperforms many other methods. The effectiveness of MLPA can be attributed to its multi-label propagating strategy.
Qiguo Dai, Maozu Guo 0001, Yang Liu 0006
IEEE Congress on Evolutionary Computation2
2013 Effective constructing training sets for object detection
abstract
This paper addresses the problem of building up effective training sets at minimal labeling cost for object detection. This problem occurs in the situation that the part-based detector is trained on a group of positive examples with bounding box labels, but the images selected by uniform sampling do not reflect the desired training distribution and need additional labeling cost in order to obtain enough positive examples. We study the active training process in which some object windows are sampled from a pool of unlabeled candidate windows, and then their corresponding bounding annotations are queried. We derive an effective training set by selecting a group of most uncertain object windows according to the current detector. Our approach has been empirically demonstrated on the object detection task of PASCAL VOC dataset. The experiment results show that our proposed algorithm outperforms common uniform sampling within the same labeling cost.
Weining Wu, Yang Liu 0006, Wei Zeng 0006, Maozu Guo 0001, Chunyu Wang 0002
ICIP4
2013 Measuring gene functional similarity based on group-wise comparison of GO terms
abstract
MOTIVATION: Compared with sequence and structure similarity, functional similarity is more informative for understanding the biological roles and functions of genes. Many important applications in computational molecular biology require functional similarity, such as gene clustering, protein function prediction, protein interaction evaluation and disease gene prioritization. Gene Ontology (GO) is now widely used as the basis for measuring gene functional similarity. Some existing methods combined semantic similarity scores of single term pairs to estimate gene functional similarity, whereas others compared terms in groups to measure it. However, these methods may make error-prone judgments about gene functional similarity. It remains a challenge that measuring gene functional similarity reliably. RESULT: We propose a novel method called SORA to measure gene functional similarity in GO context. First of all, SORA computes the information content (IC) of a term making use of semantic specificity and coverage. Second, SORA measures the IC of a term set by means of combining inherited and extended IC of the terms based on the structure of GO. Finally, SORA estimates gene functional similarity using the IC overlap ratio of term sets. SORA is evaluated against five state-of-the-art methods in the file on the public platform for collaborative evaluation of GO-based semantic similarity measure. The carefully comparisons show SORA is superior to other methods in general. Further analysis suggests that it primarily benefits from the structure of GO, which implies expressive information about gene function. SORA offers an effective and reliable way to compare gene function. AVAILABILITY: The web service of SORA is freely available at http://nclab.hit.edu.cn/SORA/
Zhixia Teng, Maozu Guo 0001, Qiguo Dai, Chunyu Wang 0002, Ping Xuan
Bioinform.2
2013 Lnetwork: an efficient and effective method for constructing phylogenetic networks
abstract
MOTIVATION: The evolutionary history of species is traditionally represented with a rooted phylogenetic tree. Each tree comprises a set of clusters, i.e. subsets of the species that are descended from a common ancestor. When rooted phylogenetic trees are built from several different datasets (e.g. from different genes), the clusters are often conflicting. These conflicting clusters cannot be expressed as a simple phylogenetic tree; however, they can be expressed in a phylogenetic network. Phylogenetic networks are a generalization of phylogenetic trees that can account for processes such as hybridization, horizontal gene transfer and recombination, which are difficult to represent in standard tree-like models of evolutionary histories. There is currently a large body of research aimed at developing appropriate methods for constructing phylogenetic networks from cluster sets. The Cass algorithm can construct a much simpler network than other available methods, but is extremely slow for large datasets or for datasets that need lots of reticulate nodes. The networks constructed by Cass are also greatly dependent on the order of input data, i.e. it generally derives different phylogenetic networks for the same dataset when different input orders are used. RESULTS: In this study, we introduce an improved Cass algorithm, Lnetwork, which can construct a phylogenetic network for a given set of clusters. We show that Lnetwork is significantly faster than Cass and effectively weakens the influence of input data order. Moreover, we show that Lnetwork can construct a much simpler network than most of the other available methods. AVAILABILITY: Lnetwork has been built as a Java software package and is freely available at http://nclab.hit.edu.cn/∼wangjuan/Lnetwork/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Juan Wang 0011, Maozu Guo 0001, Yang Liu 0006, Chunyu Wang 0002, Linlin Xing, Kai Che
Bioinform.2
2013 A probabilistic model of active learning with multiple noisy oracles
Weining Wu, Yang Liu 0006, Maozu Guo 0001, Chunyu Wang 0002
Neurocomputing3
2012 Phase transition and New Fitness Function based Genetic Inductive Logic Programming algorithm
abstract
A new genetic inductive logic programming (GILP for short) algorithm named PT-NFF-GILP (Phase Transition and New Fitness Function based Genetic Inductive Logic Programming) is proposed in this paper. Based on phase transition of the covering test, PT-NFF-GILP randomly generates initial population in phase transition region instead of the whole space of candidate clauses. Moreover, a new fitness function, which not only considers the number of examples covered by rules, but also considers the ratio of the examples covered by rules to the training examples, is defined in PT-NFF-GILP. The new fitness function measures the quality of firstorder rules more precisely, and enhances the search performance of algorithm. Experiments on ten learning problems show that: 1) the new method of generating initial population can effectively reduce iteration number and enhance predictive accuracy of GILP algorithm; 2) the new fitness function measures the quality of first-order rules more precisely and avoids generating over-specific hypothesis; 3) The performance of PT-NFF-GILP is better than other algorithms compared with it, such as G-NET, KFOIL and NFOIL.
Yanjuan Li, Maozu Guo 0001
IEEE Congress on Evolutionary Computation2
2012 A new relational Tri-training system with adaptive data editing for inductive logic programming
Yanjuan Li, Maozu Guo 0001
Knowl. Based Syst.2
2012 Feature Selection for Monotonic Classification
abstract
Monotonic classification is a kind of special task in machine learning and pattern recognition. Monotonicity constraints between features and decision should be taken into account in these tasks. However, most existing techniques are not able to discover and represent the ordinal structures in monotonic datasets. Thus, they are inapplicable to monotonic classification. Feature selection has been proven effective in improving classification performance and avoiding overfitting. To the best of our knowledge, no technique has been specially designed to select features in monotonic classification until now. In this paper, we introduce a function, which is called rank mutual information, to evaluate monotonic consistency between features and decision in monotonic tasks. This function combines the advantages of dominance rough sets in reflecting ordinal structures and mutual information in terms of robustness. Then, rank mutual information is integrated with the search strategy of min-redundancy and max-relevance to compute optimal subsets of features. A collection of numerical experiments are given to show the effectiveness of the proposed technique.
Qinghua Hu, Lei Zhang 0006, David Zhang 0001, Yanping Song, Maozu Guo 0001, Daren Yu
IEEE Trans. Fuzzy Syst.6
2012 Rank Entropy-Based Decision Trees for Monotonic Classification
abstract
In many decision making tasks, values of features and decision are ordinal. Moreover, there is a monotonic constraint that the objects with better feature values should not be assigned to a worse decision class. Such problems are called ordinal classification with monotonicity constraint. Some learning algorithms have been developed to handle this kind of tasks in recent years. However, experiments show that these algorithms are sensitive to noisy samples and do not work well in real-world applications. In this work, we introduce a new measure of feature quality, called rank mutual information (RMI), which combines the advantage of robustness of Shannon's entropy with the ability of dominance rough sets in extracting ordinal structures from monotonic data sets. Then, we design a decision tree algorithm (REMT) based on rank mutual information. The theoretic and experimental analysis shows that the proposed algorithm can get monotonically consistent decision trees, if training samples are monotonically consistent. Its performance is still good when data are contaminated with noise.
Qinghua Hu, Xunjian Che, Lei Zhang 0006, David Zhang 0001, Maozu Guo 0001, Daren Yu
IEEE Trans. Knowl. Data Eng.5
2011 PlantMiRNAPred: efficient classification of real and pseudo plant pre-miRNAs
abstract
MOTIVATION: MicroRNAs (miRNAs) are a set of short (21-24 nt) non-coding RNAs that play significant roles as post-transcriptional regulators in animals and plants. While some existing methods use comparative genomic approaches to identify plant precursor miRNAs (pre-miRNAs), others are based on the complementarity characteristics between miRNAs and their target mRNAs sequences. However, they can only identify the homologous miRNAs or the limited complementary miRNAs. Furthermore, since the plant pre-miRNAs are quite different from the animal pre-miRNAs, all the ab initio methods for animals cannot be applied to plants. Therefore, it is essential to develop a method based on machine learning to classify real plant pre-miRNAs and pseudo genome hairpins. RESULTS: A novel classification method based on support vector machine (SVM) is proposed specifically for predicting plant pre-miRNAs. To make efficient prediction, we extract the pseudo hairpin sequences from the protein coding sequences of Arabidopsis thaliana and Glycine max, respectively. These pseudo pre-miRNAs are extracted in this study for the first time. A set of informative features are selected to improve the classification accuracy. The training samples are selected according to their distributions in the high-dimensional sample space. Our classifier PlantMiRNAPred achieves >90% accuracy on the plant datasets from eight plant species, including A.thaliana, Oryza sativa, Populus trichocarpa, Physcomitrella patens, Medicago truncatula, Sorghum bicolor, Zea mays and G.max. The superior performance of the proposed classifier can be attributed to the extracted plant pseudo pre-miRNAs, the selected training dataset and the carefully selected features. The ability of PlantMiRNAPred to discern real and pseudo pre-miRNAs provides a viable method for discovering new non-homologous plant pre-miRNAs.
Ping Xuan, Maozu Guo 0001, Yangchao Huang, Yufei Huang 0001
Bioinform.2
2011 A new co-training-style random forest for computer aided diagnosis
Maozu Guo 0001
J. Intell. Inf. Syst.2
2010 Two-stage clustering based effective sample selection for classification of pre-miRNAs
abstract
To solve the class imbalance problem in classification of pre-miRNAs with ab initio method, a novel sample selection method is proposed according to the characteristics of pre-miRNAs. Real/pseudo pre-miRNAs are clustered based on their stem similarity and their distribution in high dimensional sample space respectively. The training samples are selected according to the sample density of each cluster. Experimental results are validated by the cross validation and other testing datasets composed of human real/pseudo pre-miRNAs. When compared with the previous study, microPred, our classifier miRNAPred is nearly 12% greater in total accuracy. Our sample selection algorithm is useful to construct more efficient classifier for classification of real pre-miRNAs and pseudo hairpin sequences.
Ping Xuan, Maozu Guo 0001, Jun Wang 0035, Yingpeng Han
BIBM2
2010 Simple sequence-based kernels do not predict protein-protein interactions
abstract
MOTIVATION: A number of methods have been reported that predict protein-protein interactions (PPIs) with high accuracy using only simple sequence-based features such as amino acid 3mer content. This is surprising, given that many protein interactions have high specificity that depends on detailed atomic recognition between physiochemically complementary surfaces. Are the reported high accuracies realistic? RESULTS: We find that the reported accuracies of the predictions are significantly over-estimated, and strongly dependent on the structure of the training and testing datasets used. The choice of which protein pairs are deemed as non-interactions in the training data has a variable impact on the accuracy estimates, and the accuracies can be artificially inflated by a bias towards dominant samples in the positive data which result from the presence of hub proteins in the protein interaction network. To address this bias, we propose a positive set-specific method to create a 'balanced' negative set maintaining the degree distribution for each protein, leading to the conclusion that simple sequence-based features contain insufficient information to be useful for predicting PPIs, but that protein domain-based features have some predictive value. AVAILABILITY: Our method, named 'BRS-nonint', is available at http://www.bioinformatics.leeds.ac.uk/BRS-nonint/. All the datasets used in this study are derived from publicly available data, and are available at http://www.bioinformatics.leeds.ac.uk/BRS-nonint/PPI_RandomBalance.html CONTACT: [email protected]; [email protected].
Jiantao Yu, Maozu Guo 0001, Chris J. Needham, Yangchao Huang, Lu Cai, David R. Westhead
Bioinform.2
2010 Information entropy for ordinal classification
Qinghua Hu, Maozu Guo 0001, Daren Yu
Sci. China Inf. Sci.2
2010 Fuzzy preference based rough sets
Qinghua Hu, Daren Yu, Maozu Guo 0001
Inf. Sci.3
2009 CGTS: a site-clustering graph based tagSNP selection algorithm in genotype data
abstract
BACKGROUND: Recent studies have shown genetic variation is the basis of the genome-wide disease association research. However, due to the high cost on genotyping large number of single nucleotide polymorphisms (SNPs), it is essential to choose a small subset of informative SNPs (tagSNPs), which are able to capture most variation in a population, to represent the rest SNPs. Several methods have been proposed to find the minimum set of tagSNPs, but most of them still have some disadvantages such as information loss and block-partition limit. RESULTS: This paper proposes a new hybrid method named CGTS which combines the ideas of the clustering and the graph algorithms to select tagSNPs on genotype data. This method aims to maximize the number of the discarding nontagSNPs in the given set. CGTS integrates the information of the LD association and the genotype diversity using the site graphs, discards redundant SNPs using the algorithm based on these graph structures. The clustering algorithm is used to reduce the running time of CGTS. The efficiency of the algorithm and quality of solutions are evaluated on biological data and the comparisons with three popular selecting methods are shown in the paper. CONCLUSION: Our theoretical analysis and experimental results show that our algorithm CGTS is not only more efficient than other methods but also can be get higher accuracy in tagSNP selection.
Jun Wang 0035, Maozu Guo 0001, Chunyu Wang 0002
BMC Bioinform.2
2009 A hybrid clustering and graph based algorithm for tagSNP selection
Maozu Guo 0001, Jun Wang 0035, Chunyu Wang 0002, Yang Liu 0006
Soft Comput.1
2009 Editorial
Liang Zhao 0001, Maozu Guo 0001, Lipo Wang 0001
Soft Comput.2
2008 A topological transformation in evolutionary tree search methods based on maximum likelihood combining p-ECR and neighbor joining
abstract
BACKGROUND: Inference of evolutionary trees using the maximum likelihood principle is NP-hard. Therefore, all practical methods rely on heuristics. The topological transformations often used in heuristics are Nearest Neighbor Interchange (NNI), Subtree Prune and Regraft (SPR) and Tree Bisection and Reconnection (TBR). However, these topological transformations often fall easily into local optima, since there are not many trees accessible in one step from any given tree. Another more exhaustive topological transformation is p-Edge Contraction and Refinement (p-ECR). However, due to its high computation complexity, p-ECR has rarely been used in practice. RESULTS: To make the p-ECR move more efficient, this paper proposes a new method named p-ECRNJ. The main idea of p-ECRNJ is to use neighbor joining (NJ) to refine the unresolved nodes produced in p-ECR. CONCLUSION: Experiments with real datasets show that p-ECRNJ can find better trees than the best known maximum likelihood methods so far and can efficiently improve local topological transforms in reasonable time.
Maozu Guo 0001, Jian-Fu Li, Yang Liu 0006
BMC Bioinform.1
2004 A new Q-learning algorithm based on the metropolis criterion
abstract
The balance between exploration and exploitation is one of the key problems of action selection in Q-learning. Pure exploitation causes the agent to reach the locally optimal policies quickly, whereas excessive exploration degrades the performance of the Q-learning algorithm even if it may accelerate the learning process and allow avoiding the locally optimal policies. In this paper, finding the optimum policy in Q-learning is described as search for the optimum solution in combinatorial optimization. The Metropolis criterion of simulated annealing algorithm is introduced in order to balance exploration and exploitation of Q-learning, and the modified Q-learning algorithm based on this criterion, SA-Q-learning, is presented. Experiments show that SA-Q-learning converges more quickly than Q-learning or Boltzmann exploration, and that the search does not suffer of performance degradation due to excessive exploration.
Maozu Guo 0001, Yang Liu 0006, Jacek Malec
IEEE Trans. Syst. Man Cybern. Part B1