EDBT 2026 Demo / reviewers in the wild / expert
Min Li 0007
dblp:82/0-7
· DBLP profile ↗
247ranked-venue papers
28as first author
141since 2021 · last 2026
0000-0002-0188-1394ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 228 · 28 first-author · 125 since 2021Artificial intelligence and machine learning · 17 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-IdentificationabstractMultimodal pretraining has revolutionized visual understanding, but its impact on video-based person re-identification (ReID) remains underexplored. Existing approaches often rely on video-text pairs, yet suffer from two fundamental limitations: (1) lack of genuine multimodal pretraining, and (2) text poorly captures fine-grained temporal motion—an essential cue for distinguishing identities in video. In this work, we take a bold departure from text-based paradigms by introducing the first skeleton-driven pretraining framework for ReID. To achieve this, we propose Contrastive Skeleton-Image Pretraining for ReID (CSIP-ReID), a novel two-stage method that leverages skeleton sequences as a spatiotemporally informative modality aligned with video frames. In the first stage, we employ contrastive learning to align skeleton and visual features at sequence level. In the second stage, we introduce a dynamic Prototype Fusion Updater (PFU) to refine multimodal identity prototypes, fusing motion and appearance cues. Moreover, we propose a Skeleton Guided Temporal Modeling (SGTM) module that distills temporal cues from skeleton data and integrates them into visual features. Extensive experiments demonstrate that CSIP-ReID achieves new state-of-the-art results on standard video ReID benchmarks (MARS, LS-VID, iLIDS-VID). Moreover, it exhibits strong generalization to skeleton-only ReID tasks (BIWI, IAS), significantly outperforming previous methods. CSIP-ReID pioneers an annotation-free and motion-aware pretraining paradigm for ReID, opening a new frontier in multimodal representation learning. Rifen Lin, Alex Jinpeng Wang, Jiawei Mo, Min Li 0007 |
AAAI | 4 |
| 2026 | From Charts to Code: A Hierarchical Benchmark for Multimodal ModelsabstractJiahao Tang, Henry Hengyuan Zhao, Lijian Wu, Zijian Zhang, Yifei Tao, Dongxing Mao, Yang Wan, Jingru Tan, Min Zeng, Min Li, Alex Jinpeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiahao Tang, Hengyuan Zhao, Lijian Wu, Yifei Tao, Dongxing Mao, Yang Wan, Jingru Tan, Min Zeng 0004, Min Li 0007, Alex Jinpeng Wang |
ACL (1) | 10 |
| 2026 | DIMIR: Deep Incomplete Multi-view Information Recovery for Breast Cancer Subtype Classification
Wei Lan 0001, Yinghao Liu, Xuhua Yan, Qingfeng Chen, Liangliang Liu 0001, Min Li 0007, Yi Pan 0001 |
ISBRA (1) | 7 |
| 2026 | A Single-Cell Perturbation Analysis Framework Integrating Metabolic Constraints and Chain-Based Interpretability
Ruiqing Zheng, Ju Xiang, Min Li 0007 |
ISBRA (2) | 5 |
| 2026 | SRLST: a unified multimodal representation learning framework for spatial transcriptomics analysisabstractMOTIVATION: Spatial transcriptomics (ST) enables molecular profiling within native tissue architecture, yet accurate delineation of spatial domains in ST data is challenging, as it demands the coordinated integration of transcriptomic, spatial, and tissue histological information. RESULTS: We present SRLST, an unsupervised representation learning framework that holistically harmonize these three complementary data modalities to precisely uncover tissue organization. SRLST employs a dual-graph variational autoencoding strategy to jointly model spatial proximity and morphological relations, fusing these with gene-expression embeddings into a unified latent space. Across distinct experimental datasets, SRLST consistently outperforms existing methods in delineating cortical organization, identifying small discontinuous tissue compartments, and capturing complex intratumor heterogeneity. AVAILABILITY AND IMPLEMENTATION: The code implementation of the SRLST algorithm is available at https://github.com/lanbiolab/SRLST. Wei Lan 0001, Tongsheng Ling, Guohang He, Xuhua Yan, Ruiqing Zheng, Min Li 0007, Shirui Pan, Yi Pan 0001 |
Bioinform. | 7 |
| 2026 | NumMolFormer: an explicit functional group number-guided framework for structure-based drug designabstractMOTIVATION: Rational molecule generation that balances binding affinity with favorable physicochemical properties remains a formidable challenge in structure-based drug design. The number of functional groups is a key determinant, as over-functionalization compromises physicochemical properties, whereas under-functionalization reduces binding affinity. However, current methods are limited in their capacity to incorporate the constraint. RESULTS: To address this, we present NumMolFormer, a Transformer-based framework designed to explicitly model functional group numbers. NumMolFormer adopts a dual-sequence input strategy, integrated with a numerical embedding module and a dual-stream differential attention mechanism, allowing molecular structures and functional group numbers to be encoded separately. This formulation alleviates the inherent limitations of standard Transformer in handling numerical information. In addition, we construct a large-scale dataset of 18 million molecules with functional group annotations for molecular pre-training, and further fine-tune the model using a combination of self-supervised learning and reinforcement learning under protein pocket constraints. The results demonstrate that NumMolFormer effectively leverages functional group information to generate molecules with improved binding affinity, synthetic accessibility, and drug-likeness compared to baseline methods. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at http://www.github.com/zengzhicun/nummolformer. Zhicun Zeng, Yifan Wu 0008, Zhangli Lu, Min Li 0007 |
Bioinform. | 4 |
| 2026 | A causal inference framework for identifying essential genes to enhance drug synergy predictionabstractMOTIVATION: Identifying synergistic drug combinations holds promise for more effective treatment strategies. Recent deep learning methods such as Transformers and Graph Neural Networks have shown improved predictive performance, but most of them integrate drug and cell line representations without explicitly modelling the causal effects of genes in mediating drug responses. RESULTS: We introduce CADS (Causal Adjustment for Drug Synergy), a deep learning framework that explicitly models the gene-drug causal relationships to improve both prediction accuracy and biological interpretability. CADS integrates multi-omics data with a learnable gene-selection mechanism that performs causal backdoor adjustment, enabling both drug synergy prediction and causal gene discovery. Across multiple benchmark datasets, CADS consistently achieves superior performance compared with state-of-the-art drug synergy prediction models. In addition, downstream analyses on case studies demonstrate that the inferred gene causal scores can recover clinically validated cancer-related genes involved in drug combinations. These results demonstrate that explicitly modelling causal genetic effects can enhance the reliability and interpretability of drug synergy prediction. AVAILABILITY AND IMPLEMENTATION: The source code of CADS can be found at https://github.com/HuaiwuZhang/causalDC. Huaiwu Zhang, Xinliang Sun, Jianxin Wang 0001, Min Li 0007, Jing Tang 0002 |
Bioinform. | 4 |
| 2026 | Drug target prediction from perturbation transcriptomics via a biological function-guided hypergraph siamese networkabstractMOTIVATION: Understanding how small molecules modulate cellular states remains a critical challenge in drug discovery. The advent of perturbation transcriptomics offers new avenues for elucidating drug-target interactions by capturing cellular transcriptional responses to perturbations. RESULTS: In this study, we propose BioHSNet, a biological function-guided hypergraph siamese network for inferring drug-target interactions from perturbation transcriptomics. BioHSNet utilizes hyperedge representations of functionally grouped gene expression to capture higher-order functional relationships, and integrates compound structural information into the model to bridge chemical structure and functional response. Experimental results demonstrate that BioHSNet outperforms other transcriptome-based methods on the Broad Institute's L1000 datasets, particularly in cold start scenarios. The case study further demonstrates its practical utility for target prediction and drug screening. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/Zxinyizhang/BioHSNet. Xinliang Sun, Jiuxu Yang, Min Li 0007 |
Bioinform. | 4 |
| 2026 | Large language models meet NLP: a surveyabstractAbstract While large language models (LLMs) like ChatGPT have shown impressive capabilities in Natural Language Processing (NLP) tasks, a systematic investigation of their potential in this field remains largely unexplored. This study aims to address this gap by exploring the following questions. (1) How are LLMs currently applied to NLP tasks in the literature ? (2) Have traditional NLP tasks already been solved with LLMs ? (3) What is the future of the LLMs for NLP ? To answer these questions, we take the first step to provide a comprehensive overview of LLMs in NLP. Specifically, we first introduce a unified taxonomy including (1) parameter-frozen paradigm and (2) parameter-tuning paradigm to offer a unified perspective for understanding the current progress of LLMs in NLP. Furthermore, we summarize the new frontiers and the corresponding challenges, aiming to inspire further groundbreaking advancements. We hope this work offers valuable insights into {the potential and limitations} of LLMs, while also serving as a practical guide for building effective LLMs in NLP. Libo Qin 0001, Qiguang Chen, Xiachong Feng, Yang Wu 0010, Yongheng Zhang 0001, Min Li 0007, Wanxiang Che, Philip S. Yu |
Frontiers Comput. Sci. | 7 |
| 2026 | EssLM-MoE: A mixture-of-experts-enhanced framework for protein essentiality prediction using fused protein language models
Min Zeng 0004, Qianpei Liu, Wenkang Wang, Fuhao Zhang, Fei Guo 0001, Min Li 0007 |
Neurocomputing | 8 |
| 2026 | Guest Editorial for the 21th Asia Pacific Bioinformatics Conference
Min Li 0007, Feng Luo 0001, Yi-Ping Phoebe Chen |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2026 | WEmarker: Breast Cancer-Specific Prognostic Analysis With Weighted Multiplex Network EmbeddingabstractIn clinical trials, prognostic biomarkers have become essential for guiding treatment decisions after breast cancer surgery. Network-based methods have gained notable attention to reveal marker genes, but many existing methods only focus on a single network, which inevitably neglects the incompleteness of interaction relationships within the network. Even when based upon the multiplex network, most of methods directly integrate the multiplex network into an aggregated network and do not take into account the inherent noise in the biological networks, which can not preserve the topological structure of each original network very well. In this study, we propose a novel method, WEmarker, for breast cancer-specific prognostic analysis. WEmarker reduces the noise level of biological networks and quantifies the probability of interactions between genes, and represents the nodes in the weighted multiplex network as vectors while efficiently retaining the structure information of these networks for identifying prognostic biomarkers. The results show that WEmarker outperforms comparative methods and the case study also demonstrates that biomarkers identified by WEmarker have reliable biological interpretability for breast cancer prognosis. Xingyi Li 0003, Zhelin Zhao, Huihui Kong, Min Li 0007, Xuequn Shang 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2026 | MoChat: Joints-Grouped Spatio-Temporal Grounding Multimodal Large Language Model for Multi-Turn Motion Comprehension and DescriptionabstractDespite continuous advancements in deep learning for understanding human motion, existing models often struggle to accurately identify action timing and specific body parts, typically supporting only single-round interaction. This limitation is particularly pronounced in home exercise monitoring, neurological disorder assessment, and rehabilitation, where precise motion analysis is crucial for ensuring exercise efficacy, detecting early signs of neurological conditions, and guiding personalized recovery programs. In this paper, we propose MoChat, a multimodal large language model capable of spatio-temporal grounding of human motion and multi-turn dialogue understanding. To achieve this, we first group spatial features in skeleton frames according to human anatomical structures and process them through a Joints-Grouped Skeleton Encoder. The encoder's outputs are fused with large language model embeddings to generate spatio-aware representations. A cross-attention-based Regression Head module is then designed to align hidden-layer embeddings and skeletal sequence embeddings, enabling precise temporal grounding. Furthermore, we develop a pipeline for temporal grounding task to extract timestamps from skeleton-text pairs and construct a multi-turn instruction dialogues for spatial grounding task. Finally, various task instructions are generated for jointly training. Experimental results demonstrate that MoChat achieves state-of-the-art performance across multiple metrics in motion understanding tasks, making it as the first model capable of fine-grained spatio-temporal grounding of human motion. Jiawei Mo, Yixuan Chen 0019, Rifen Lin, Yongkang Ni, Feng Liang 0004, Min Zeng 0004, Xiping Hu, Min Li 0007 |
IEEE J. Biomed. Health Informatics | 8 |
| 2026 | DPGOK: A Deep Learning-Based Method for Protein Function Prediction by Fusing GO Knowledge With Protein FeaturesabstractAccurately predicting protein functions is critical for understanding disease mechanisms and discovering potential drug targets. Gene Ontology (GO), with its hierarchical and semantic information, provides valuable context that can be integrated to improve prediction accuracy. Recently, several existing methods have attempted to integrate GO knowledge with protein sequence features for function prediction. However, these methods ignore the fact that GO embeddings should be tailored to proteins to reflect protein-specific functional relevance. To address this limitation, we proposed DPGOK, a deep learning-based method that fused protein-aware GO representations with protein features for function prediction. DPGOK first learns GO semantic representations with a knowledge graph loss and further generates protein-aware GO embeddings under the guidance of protein features. Results show that DPGOK outperforms state-of-the-art methods across all GO domains. Additional experiments demonstrated that DPGOK is capable of discovering hierarchically deeper and more informative functions for target proteins. Ablation studies revealed that the knowledge graph loss we introduced contributes to more stable and semantically coherent GO representations across different domains. Finally, we find that the predictive performance can be further improved when DPGOK is combined with homology-based approaches. Qiurong Yang, Wenkang Wang, Wei Fan 0010, Ruiqing Zheng, Min Li 0007 |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Highly Undersampled MRI Reconstruction via a Single Posterior Sampling of Diffusion ModelsabstractIncoherent k-space undersampling and deep learning-based reconstruction methods have shown great success in accelerating MRI. However, the performance of most previous methods will degrade dramatically under high acceleration factors, e.g., $8\times $ or higher. Recently, denoising diffusion models (DM) have demonstrated promising results in solving this issue; however, one major drawback of the DM methods is the long inference time due to a dramatic number of iterative reverse posterior sampling steps. In this work, a Single Step Diffusion Model-based reconstruction framework, namely SSDM-MRI, is proposed for restoring MRI images from highly undersampled k-space. The proposed method achieves one-step reconstruction by first training a conditional DM and then iteratively distilling this model four times using an iterative selective distillation algorithm, which works synergistically with a shortcut reverse sampling strategy for model inference. Comprehensive experiments were carried out on both publicly available fastMRI brain and knee images, as well as an in-house multi-echo GRE (QSM) subject. Overall, the results showed that SSDM-MRI outperformed other methods in terms of numerical metrics (e.g., PSNR and SSIM), error maps, image fine details, and latent susceptibility information hidden in MRI phase images. In addition, the reconstruction time for a ${320}\times {320}$ brain slice of SSDM-MRI is only 0.45 second, which is only comparable to that of a simple U-net, making it a highly effective solution for MRI reconstruction tasks. Jin Liu 0012, Shanshan Shan, Chunyi Liu, Min Li 0007, Feng Liu 0005, G. Bruce Pike, Hongfu Sun, Yang Gao 0030 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Divide-Solve-Combine: An Interpretable and Accurate Prompting Framework for Zero-shot Multi-Intent DetectionabstractZero-shot multi-intent detection is capable of capturing multiple intents within a single utterance without any training data, which gains increasing attention. Building on the success of large language models (LLM), dominant approaches in the literature explore prompting techniques to enable zero-shot multi-intent detection. While significant advancements have been witnessed, the existing prompting approaches still face two major issues: lacking explicit reasoning and lacking interpretability. Therefore, in this paper, we introduce a Divide-Solve-Combine Prompting (DSCP) to address the above issues. Specifically, DSCP explicitly decomposes multi-intent detection into three components including (1) single-intent division prompting is utilized to decompose an input query into distinct sub-sentences, each containing a single intent; (2) intent-by-intent solution prompting is applied to solve each sub-sentence recurrently; and (3) multi-intent combination prompting is employed for combining each sub-sentence result to obtain the final multi-intent result. By decomposition, DSCP allows the model to track the explicit reasoning process and improve the interpretability. In addition, we propose an interactive divide-solve-combine prompting (Inter-DSCP) to naturally capture the interaction capabilities of large language models. Experimental results on two standard multi-intent benchmarks (i.e., MixATIS and MixSNIPS) reveal that both DSCP and Inter-DSCP obtain substantial improvements over baselines, achieving superior performance and higher interpretability. Libo Qin 0001, Qiguang Chen, Jingxuan Zhou, Hao Fei 0003, Wanxiang Che, Min Li 0007 |
AAAI | 7 |
| 2025 | CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have recently demonstrated amazing success in multi-modal tasks, including advancements in Multi-modal Chain-of-Thought (MCoT) reasoning. Despite these successes, current benchmarks still follow a traditional paradigm with multi-modal input and text-modal output, which leads to significant drawbacks such as missing visual operations and vague expressions. Motivated by this, we introduce a novel Chain of Multi-modal Thought (CoMT) benchmark to address these limitations. Different from the traditional MCoT benchmark, CoMT requires both multi-modal input and multi-modal reasoning output, aiming to mimic human-like reasoning that inherently integrates visual operation. Specifically, CoMT consists of four categories: (1) Visual Creation, (2) Visual Deletion, (3) Visual Update, and (4) Visual Selection to comprehensively explore complex visual operations and concise expression in real scenarios. We evaluate various LVLMs and strategies on CoMT, revealing some key insights into the capabilities and limitations of the current approaches. We hope that CoMT can inspire more research on introducing multi-modal generation into the reasoning process. Zihui Cheng, Qiguang Chen, Hao Fei 0003, Wanxiang Che, Min Li 0007, Libo Qin 0001 |
AAAI | 7 |
| 2025 | GateFuseNet: An Adaptive 3D Multimodal Neuroimaging Fusion Network for Parkinson's Disease DiagnosisabstractAccurate diagnosis of Parkinson's disease (PD) from MRI remains challenging due to symptom variability and pathological heterogeneity. Most existing methods rely on conventional magnitude-based MRI modalities, such as T1weighted images (T1w), which are less sensitive to PD pathology than Quantitative Susceptibility Mapping (QSM), a phasebased MRI technique that quantifies iron deposition in deep gray matter nuclei. In this study, we propose GateFuseNet, an adaptive 3D multimodal fusion network that integrates QSM and T1w images for PD diagnosis. The core innovation lies in a gated fusion module that learns modality-specific attention weights and channel-wise gating vectors for selective feature modulation. This hierarchical gating mechanism enhances ROIaware features while suppressing irrelevant signals. Experimental results show that our method outperforms three existing state-of-the-art approaches, achieving 85.00 % accuracy and 92.06% AUC. Ablation studies further validate the contributions of ROI guidance, multimodal integration, and fusion positioning. Grad-CAM visualizations confirm the model's focus on clinically relevant pathological regions. The source codes and pretrained models can be found at https://github.com/YangGaoUQ/GateFuseNet Hongfu Sun, Ruiqing Zheng, Min Zeng 0004, Min Li 0007, Yang Gao 0030 |
BIBM | 8 |
| 2025 | DeepDICI: Accurately Predicting Drug-Ion Channel Interactions via Deep Learning with Dynamic Structural FeaturesabstractIon channels are critical targets in drug development and play an important role in the treatment of various diseases. However, existing prediction methods either generalize across all drug targets or focus narrowly on specific ion channel proteins, leading to limitations in accuracy and versatility. In this study, we propose DeepDICI, a novel deep learning framework that integrates topology-aware protein embeddings with spatial-channel interaction modeling to overcome these challenges. Our approach leverages a LLaMA3-style language model to capture longrange structural dependencies between transmembrane segments and cytoplasmic domains, effectively simulating the grammar of channel folding. In addition, a Cross-Attention network dynamically aligns drug substructures with channel functional domains, enabling the precise identification of state-dependent binding interfaces. Comprehensive evaluations in multiple scenarios demonstrate that DeepDICI achieves the highest precision of prediction$(P R C>0.98)$and maintains exceptional robustness under cold start conditions, showing a 28.9% improvement in terms of recall for novel chemical scaffolds and a 14.3% increase in terms of accuracy for unseen channel targets. All the results indicate that DeepDICI represents a unified and interpretable paradigm for ion channel-targeted drug discovery, bridging computational predictions with biophysical insights into structuredynamic protein systems. The code and datasets for DeepDICI are freely available at https://github.com/CSUBioGroup/DeepDICI. Zhangli Lu, Zhicun Zeng, Ruiqing Zheng, Min Zeng 0004, Min Li 0007 |
BIBM | 6 |
| 2025 | ADESys: A Modular System for ADE Identification Research with LLM-RAG IntegrationabstractIdentifying Adverse Drug Events (ADE) from clinical course records presents a long-standing challenge in the field of natural language processing, primarily due to the semantic complexity and domain-specificity of clinical texts. The development of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) offers new possibilities for addressing this challenge. However, when applying these technologies to ADE identification research, key constraints exist, including the construction of domain knowledge bases, efficient retrieval, precise prompting, and high clinician usage threshold. To address these constraints, this paper presents ADESys, a system specifically designed to support research on ADE identification tasks through the integration of LLM and RAG. ADESys adopts a modular architecture covering dataset management and annotation, knowledge organization, LLM model integration, prompt engineering, and experimental validation. By offering an open and extensible platform, ADESys enables in-depth exploration of LLM-RAG components, serving as a comprehensive tool for advancing ADE identification research. Jun-Long Ma, Min Li 0007 |
BIBM | 4 |
| 2025 | DP-GPT: GPT-Driven Gene Text Feature Embedding Fused with Gene Expression Data for Depression PredictionabstractIn recent years, with the improvement of living standards, the prevalence of depression has been steadily increasing, making it a growing public health concern. Gene expression data can reveal links between genes and diseases. Studies have shown that gene expression in depression patients differs significantly from healthy individuals, offering potential for early detection. However, existing methods often depend on selecting differentially expressed genes, which may overlook important signals from other genes and are vulnerable to batch effects, limiting model generalization. To address these limitations, we propose DP-GPT, a GPT-driven gene text feature embedding framework fused with gene expression data for depression prediction. DP-GPT integrates gene expression data with features extracted by GPT. Specifically, gene names and summaries are obtained from the NCBI Gene database, then embedded using GPT to generate feature vectors. These are fused with sample gene expression data and fed into a classifier for prediction. Extensive experiments show that DP-GPT achieves superior performance in depression prediction. The source code can be obtained from https://github.com/CSUBioGroup/DP-GPT. Min Zeng 0004, Junyu Gao 0004, Qianpei Liu, Fuhao Zhang, Ruiqing Zheng, Min Li 0007 |
BIBM | 7 |
| 2025 | Accurate Residue-Level Prediction of Linear Interacting Peptides with Auxiliary Multi-Scale Anchor DetectionabstractIntrinsically disordered regions (IDRs) play essential roles in cellular signaling and regulation, primarily mediating interactions through short peptide motifs. Linear Interacting Peptides (LIPs) are a recently defined class of bindingassociated IDRs that undergo disorder-to-order transitions upon binding. LIPs encompass several well-characterized subclasses, including molecular recognition features (MoRFs) and short linear motifs (SLiMs). However, many current predictors focus exclusively on MoRFs prediction and exhibit limited effectiveness in identifying the broader class of LIPs. To address this limitation, we propose LipPredictor, the first deep learning framework specifically designed for LIP prediction. The model incorporates protein language model embeddings as input to a shared feature extraction module, which is composed of convolutional neural networks and a multi-head attention layer. It further adopts a dual-branch architecture that couples residue-level classification with auxiliary multi-scale anchor detection, enabling accurate identification of LIPs across diverse segment lengths through joint training. These results demonstrate its robustness and effectiveness in predicting LIPs. The source code can be obtained from https://github.com/Chenxi-Xia/LipPredictor. Fuhao Zhang, Chenxi Xia, Mingxin Dong, Min Zeng 0004, Jian Zhang 0020, Min Li 0007 |
BIBM | 6 |
| 2025 | Contrastive Learning-Based Method for Single-Cell Multi-omics Data Clustering
Zhenlan Liang, Ruiqing Zheng, Huayu Tao, Min Li 0007 |
ISBRA (1) | 5 |
| 2025 | DDLB: Using the Protein Language Model and Hierarchical Architecture to Improve Disordered Lipid-Binding Residues Prediction
Chaojin Wu, Fuhao Zhang, Pengzhen Jia, Min Zeng 0004, Min Li 0007 |
ISBRA (1) | 5 |
| 2025 | Towards Multi-resolution Spatiotemporal Graph Learning for Medical Time Series ClassificationabstractMedical time series has been playing a vital role in real-world healthcare systems as valuable information in monitoring health conditions of patients. Traditional methods towards medical time series classification rely on handcrafted feature extraction and statistical methods; with the recent advancement of artificial intelligence, the machine learning and deep learning methods have become more popular. However, existing methods often fail to fully model the complex spatial dynamics under different scales, which ignore the dynamic multi-resolution spatial and temporal joint inter-dependencies. Moreover, they are less likely to consider the special baseline wander problem as well as the multi-view characteristics of medical time series, which largely hinders their prediction performance. To address these limitations, we propose a Multi-resolution Spatiotemporal Graph Learning framework, MedGNN, for medical time series classification. Specifically, we first propose to construct multi-resolution adaptive graph structures to learn dynamic multi-scale embeddings. Then, to address the baseline wander problem, we propose Difference Attention Networks to operate self-attention mechanisms on the finite difference for temporal modeling. Moreover, to learn the multi-view characteristics, we utilize the Frequency Convolution Networks to capture complementary information of medical time series from the frequency domain. In addition, we introduce the Multi-resolution Graph Transformer architecture to model the dynamic dependencies and fuse the information from different resolutions. Finally, we have conducted extensive experiments on multiple medical real-world datasets that demonstrate the superior performance of our method. Our Code is available at this repository: https://github.com/aikunyi/MedGNN. Wei Fan 0010, Jingru Fei, Dingyu Guo, Kun Yi 0001, Xiaozhuang Song, Haolong Xiang, Hangting Ye, Min Li 0007 |
WWW | 8 |
| 2025 | scHLens: a web server for hierarchically and interactively exploring single cell RNA-seq dataabstractWith the great advancement of single-cell transcriptome technologies, the identification of cellular heterogeneity from scRNA-seq data has become an important task in biomedical research. There are several challenges associated with the existing analysis methods: (i) The reliance on command-line interfaces creates a substantial technical barrier for researchers lacking computational expertise; (ii) existing methods or platforms usually lack flexibility in workflow customization, forcing users into rigid analytical pipelines; (iii) hierarchical cellular subtypes challenge conventional clustering, as fixed-resolution analyses prevent the detection of biologically subtype cells. Here, we develop a hierarchical and interactive web server named scHLens. scHLens supports a user-defined analysis pipeline and hierarchical exploration mode, providing various visualization views and interaction operations. The three case studies demonstrate scHLens's ability to identify cellular heterogeneity. The online web server version is freely available at http://schlens.csuligroup.com, while the Docker version is available at https://hub.docker.com/r/zhiweideng975/schlens, and the source code can be obtained at https://github.com/ZhiweiDeng459/scHLens. Jiazhi Xia, Zhiwei Deng, Min Li 0007, Ruiqing Zheng |
Briefings Bioinform. | 4 |
| 2025 | RNALoc-LM: RNA subcellular localization prediction using pre-trained RNA language modelabstractMOTIVATION: Accurately predicting RNA subcellular localization is crucial for understanding the cellular functions and regulatory mechanisms of RNAs. Although many computational methods have been developed to predict the subcellular localization of lncRNAs, miRNAs, and circRNAs, very few of them are designed to simultaneously predict the subcellular localization of multiple types of RNAs. In addition, the emergence of pre-trained RNA language model has shown remarkable performance in various bioinformatics tasks, such as structure prediction and functional annotation. Despite these advancements, there remains a significant gap in applying pre-trained RNA language models specifically for predicting RNA subcellular localization. RESULTS: In this study, we proposed RNALoc-LM, the first interpretable deep-learning framework that leverages a pre-trained RNA language model for predicting RNA subcellular localization. RNALoc-LM uses a pre-trained RNA language model to encode RNA sequences, then captures local patterns and long-range dependencies through TextCNN and BiLSTM modules. A multi-head attention mechanism is used to focus on important regions within the RNA sequences. The results demonstrate that RNALoc-LM significantly outperforms both deep-learning baselines and existing state-of-the-art predictors. Additionally, motif analysis highlights RNALoc-LM's potential for discovering important motifs, while an ablation study confirms the effectiveness of the RNA sequence embeddings generated by the pre-trained RNA language model. AVAILABILITY AND IMPLEMENTATION: The RNALoc-LM web server is available at http://csuligroup.com:8000/RNALoc-LM. The source code can be obtained from https://github.com/CSUBioGroup/RNALoc-LM. Min Zeng 0004, Chengqian Lu, Rui Yin 0002, Fei Guo 0001, Min Li 0007 |
Bioinform. | 7 |
| 2025 | Leveraging protein language models for cross-variant CRISPR/Cas9 sgRNA activity predictionabstractMOTIVATION: Accurate prediction of single-guide RNA (sgRNA) activity is crucial for optimizing the CRISPR/Cas9 gene-editing system, as it directly influences the efficiency and accuracy of genome modifications. However, existing prediction methods mainly rely on large-scale experimental data of a single Cas9 variant to construct Cas9 protein (variants)-specific sgRNA activity prediction models, which limits their generalization ability and prediction performance across different Cas9 protein (variants), as well as their scalability to the continuously discovered new variants. RESULTS: In this study, we proposed PLM-CRISPR, a novel deep learning-based model that leverages protein language models to capture Cas9 protein (variants) representations for cross-variant sgRNA activity prediction. PLM-CRISPR uses tailored feature extraction modules for both sgRNA and protein sequences, incorporating a cross-variant training strategy and a dynamic feature fusion mechanism to effectively model their interactions. Extensive experiments demonstrate that PLM-CRISPR outperforms existing methods across datasets spanning seven Cas9 protein (variants) in three real-world scenarios, demonstrating its superior performance in handling data-scarce situations, including cases with few or no samples for novel variants. Comparative analyses with traditional machine learning and deep learning models further confirm the effectiveness of PLM-CRISPR. Additionally, motif analysis reveals that PLM-CRISPR accurately identifies high-activity sgRNA sequence patterns across diverse Cas9 protein (variants). Overall, PLM-CRISPR provides a robust, scalable, and generalizable solution for sgRNA activity prediction across diverse Cas9 protein (variants). AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/CSUBioGroup/PLM-CRISPR. Yalin Hou, Ruiqing Zheng, Fuhao Zhang, Fei Guo 0001, Min Li 0007, Min Zeng 0004 |
Bioinform. | 6 |
| 2025 | 2OMe-LM: predicting 2′-O-methylation sites in human RNA using a pre-trained RNA language modelabstractMOTIVATION: 2'-O-methylation (2OMe) is a common post-transcriptional modification in RNA that plays a crucial role in regulating gene expression and is implicated in various biological processes and diseases. Computational methods offer an efficient alternative to the time-consuming and costly experimental identification of 2OMe sites. Recent advancements in RNA pre-trained language models have revolutionized RNA bioinformatics. However, there remains a gap in their application specifically for predicting 2OMe sites. RESULTS: In the study, we propose a novel deep learning framework, 2OMe-LM, for predicting 2OMe sites in RNA. 2OMe-LM integrates RNA sequence features derived from RNA pre-trained language models with those obtained from the word2vec technique. Then, 2OMe-LM employs fully connected layers and a bidirectional long short-term memory network to process the two types of features separately, followed by a feature fusion module for the final prediction. Additionally, an attention block is incorporated to provide the interpretability of the prediction results. The results demonstrate that 2OMe-LM significantly outperforms existing state-of-the-art predictors, with features from RNA pre-trained language models proving to be critical. Motif analysis further demonstrates 2OMe-LM's potential for discovering 2OMe-related motifs. AVAILABILITY AND IMPLEMENTATION: The 2OMe-LM web server is available at https://csuligroup.com:9200/2OMe-LM. The source code can be obtained from https://github.com/CSUBioGroup/2OMe-LM. Qianpei Liu, Min Zeng 0004, Chengqian Lu, Shichao Kan, Fei Guo 0001, Min Li 0007 |
Bioinform. | 7 |
| 2025 | TPepRet: a deep learning model for characterizing T-cell receptors-antigen binding patternsabstractMOTIVATION: T-cell receptors (TCRs) elicit and mediate the adaptive immune response by recognizing antigenic peptides, a process pivotal for cancer immunotherapy, vaccine design, and autoimmune disease management. Understanding the intricate binding patterns between TCRs and peptides is critical for advancing these clinical applications. While several computational tools have been developed, they neglect the directional semantics inherent in sequence data, which are essential for accurately characterizing TCR-peptide interactions. RESULTS: To address this gap, we develop TPepRet, an innovative model that integrates subsequence mining with semantic integration capabilities. TPepRet combines the strengths of the Bidirectional Gated Recurrent Unit (BiGRU) network for capturing bidirectional sequence dependencies with the Large Language Model framework to analyze subsequences and global sequences comprehensively, which enables TPepRet to accurately decipher the semantic binding relationship between TCRs and peptides. We have evaluated TPepRet to a range of challenging scenarios, including performance benchmarking against other tools using diverse datasets, analysis of peptide binding preferences, characterization of T cells clonal expansion, identification of true binder in complex environments, assessment of key binding sites through alanine scanning, validation against expression rates from large-scale datasets, and ability to screen SARS-CoV-2 TCRs. The comprehensive results suggest that TPepRet outperforms existing tools. We believe TPepRet will become an effective tool for understanding TCR-peptide binding in clinical treatment. AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/CSUBioGroup/TPepRet.git. Meng Wang 0067, Wei Fan 0010, Tianrui Wu, Min Li 0007 |
Bioinform. | 4 |
| 2025 | PathActMarker: an R package for inferring pathway activity of complex diseases
Xingyi Li 0003, Zhelin Zhao, Xingyu Liao, Min Li 0007, Xuequn Shang 0001 |
Frontiers Comput. Sci. | 6 |
| 2025 | Enhancing ICD classification with semantic embedding rectification and long-tail refinement
Yuhao Wu 0002, Yifan Wu 0008, Wei Fan 0010, Min Li 0007 |
Knowl. Based Syst. | 4 |
| 2025 | SpaNN: Spatial Transcriptomic Data Enhancement Using Deep Neural NetworkabstractSpatial transcriptomic sequencing technology is a powerful tool that combines gene expression data with their physical locations in tissues or organs, providing researchers with unprecedented spatial resolution of cellular molecular functions. Currently, spatial transcriptomic sequencing based on in situ hybridization and imaging can obtain cell location information and transcriptome profiles at single-cell resolution, but it only detects a limited number of genes, which restricts its application in exploring whole-genome expression patterns. Therefore, it is essential to predict the spatial distribution of undetected genes in their spatial transcriptomic data. Here, we introduce a novel data enhancement technique, named SpaNN, which predicts transcriptome expression levels in spatial context. SpaNN employs a custom-designed similarity loss that leverages location information from spatial transcriptomic data to train a deep neural network. This network captures joint embeddings and uses a weighted k-nearest-neighbor approach to predict the unmeasured genes spatial expression levels. Our experiments show that SpaNN not only recovers the expression levels of unmeasured genes but also enhances cell clustering and visualization. Additionally, sensitivity and scalability analyses confirm that SpaNN is robust to parameter variations and can handle large-scale datasets effectively. Wenkang Wang, Ruiqing Zheng, Min Li 0007 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | A Hypergraph Convolutional Network With Explicit High-Order Interaction Information Extraction for Drug RepositioningabstractDrug repositioning, a promising strategy in drug development, aims to identify new indications for existing drugs while reducing costs and safety risks. Leveraging their unique advantages in modeling higher-order relations among nodes, hypergraphs and hypergraph neural networks (HGNN) have become increasingly popular in drug repositioning. However, most HGNN-based methods overlook the diverse relations generated during the convolution and do not explicitly model high-order interactions, limiting their ability to capture high-order interaction information adequately. To address these limitations, we propose HGCNDR, a hypergraph convolutional network with explicit high-order interaction extraction for drug repositioning. HGCNDR introduces a relation-aware hypergraph convolution operation to handle distinct relation types and a Hadamard product-based strategy to effectively model high-order interactions among drugs and diseases, efficiently extracting the resulting high-order interaction information. Specifically, HGCNDR constructs two feature graphs and a hypergraph based on drug similarity features, disease similarity features, and drug-disease association networks. HGCNDR then employs graph convolutional networks to extract embeddings from the feature graphs, while using the relation-aware hypergraph convolution operation and the strategy to extract structural and high-order interaction information embeddings from the hypergraph. Additionally, to preserve the common semantics between the embeddings extracted from the feature graphs and the hypergraph, HGCNDR introduces a consistency constraint. The experimental results demonstrate that HGCNDR has competitive performance compared to several baseline methods. Moreover, case studies on Alzheimer's disease and Breast carcinoma confirm that HGCNDR can retrieve more actual drug-disease associations in the top prediction results. Xiang Du, Xinliang Sun, Min Zeng 0004, Min Li 0007 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | Enhancing Protein Function Prediction Through the Fusion of Multi-Type Biological Knowledge With Protein Language Model and Graph Neural NetworkabstractProteins play crucial roles in diverse biological functions. Accurately annotating their functions is essential for understanding cellular mechanisms and developing therapies for complex diseases. Computational methods have been proposed as alternatives to labor-intensive and expensive experimental approaches. Existing computational methods have demonstrated that protein evolution information and Protein-Protein Interactions (PPIs) are essential for protein function prediction. However, traditional computational approaches for generating evolution information are time-consuming. On the other hand, proteins lacking interactions are ignored in previous studies. To address these limitations, we propose a novel deep learning framework, named DeepFMB, which incorporates multi-type biological knowledge. DeepFMB leverages a pre-trained protein language model to extract evolution information. Moreover, DeepFMB generates PPI-related features and orthology-related features using graph neural networks on the constructed PPI and orthology networks. Then, these multi-type features are fused adaptively for protein function prediction. Compared to eight state-of-the-art methods, DeepFMB outperforms all of them in terms of F-max and AUPR. Additionally, with the combination of sequence similarity-based inference, our predicted model predicts protein functions more accurately. Experimental results also validate the superior performance of our methods in predicting low-frequency GO terms. Ablation studies demonstrate that the multi-type biological knowledge we use is highly relevant to protein functions. Wenkang Wang, Yunyan Shuai, Min Zeng 0004, Min Li 0007 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | DISHIC: An Effective Method for Identifying Differential Interaction in Single-Cell Hi-CabstractSingle-cell three-dimensional (3D) genomics plays a vital role in understanding the spatial organization of chromatin within the cell nucleus. It provides insights into how chromatin structures influence regulatory mechanisms and cellular functions that are specific to different cell types and states. Analyzing differential interactions-the fundamental units of the 3D genome-at the single-cell level is crucial for extracting this valuable information. However, single-cell 3D genome data such as single-cell High-throughput Chromatin Conformation Capture (scHi-C) data are inherently sparse, noisy and heterogeneous, posing significant challenges that limit the downstream analyses. Although existing computational methods offer helpful solutions for identifying differential interactions at the single-cell level, most of them do not explicitly account for intrinsic statistical properties of the data, nor do they comprehensively incorporate both within-sample and between-sample covariates. In this study, we developed a statistical method named DISHIC (Differential Interaction analysis in single-cell Hi-C) to perform differential interaction analysis in scHi-C data. DISHIC leverages the Zero-Inflated Negative Binomial-based Wavelet (ZINB-WaVE) model, making it well-suited for high-dimensional zero-inflated count data with high dispersion. It models each bin pair in each sample independently while incorporating both bin-pair-level and cell-level covariates, allowing it to effectively capture noise and heterogeneity in sparse scHi-C data. This approach can detect differential interactions with greater accuracy and reliability while also allowing users to flexibly define covariates as input based on their specific needs. We evaluated the performance of DISHIC using both real and simulated datasets, demonstrating its enhanced effectiveness compared to existing state-of-the-art methods under various conditions. Furthermore, we conducted a comprehensive case study analyzing multi-omics datasets from two types of glial cells. This study explored the intricate relationships among chromatin interactions, gene expressions, and epigenetic modifications, providing new insights into cell-type-specific regulatory mechanisms. Yi-chao Zhao, Ruiqing Zheng, Pengzhen Jia, Min Li 0007 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | DRGCL: Drug Repositioning via Semantic-Enriched Graph Contrastive LearningabstractDrug repositioning greatly reduces drug development costs and time by discovering new indications for existing drugs. With the development of technology and large-scale biological databases, computational drug repositioning has increasingly attracted remarkable attention, which can narrow down repositioning candidates. Recently, graph neural networks (GNNs) have been widely used and achieved promising results in drug repositioning. However, the existing GNNs based methods usually focus on modeling the complex drug-disease association graph, but ignore the semantic information on the graph, which may lead to a lack of consistency of global topology information and local semantic information for the learned features. To alleviate the above challenge, we propose a novel drug repositioning model based on graph contrastive learning, termed DRGCL. First, we treat the known drug-disease associations as the topology graph. Second, we select the top- similar neighbor from drug/disease similarity information to construct the semantic graph rather than use the traditional data augmentation strategy, thereby maximally retaining rich semantic information. Finally, we pull closer to embedding consistency of the different embedding spaces by graph contrastive learning to enhance the topology and semantic feature on the graph. We have evaluated DRGCL on four benchmark datasets and the experiment results show that the proposed DRGCL is superior to the state-of-the-art methods. Especially, the average result of DRGCL is 11.92% higher than that of the second-best method in terms of AUPRC. The case studies further demonstrate the reliability of DRGCL. Xiao Jia 0020, Xinliang Sun, Min Li 0007 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | TransScore: A Graph Model for Pose Scoring and Affinity Prediction Based on Transformer Convolution NetworkabstractPredicting the interaction of protein and compound is an important task in drug discovery. Molecular docking has been a fundamental and vital computer-aid tool for digging potential interaction of the protein-compound pair. With the recent great success of artificial intelligence (AI), the scoring function, as a fundamental part of molecular docking, has been achieving much better performance by incorporating AI-based models. However, the AI-based models usually focus on a single prediction task (e.g., affinity prediction), which is limited by their lack of extensibility. Moreover, the performance of AI-based models usually declines in cold start scenarios, thus compromising the robustness. To this end, we propose a novel deep learning-based graph model based on the transformer convolution network for pose scoring and affinity prediction. TransScore captures the intrinsic characteristics of protein-compound poses by employing the self-attention mechanism, which achieves superior performances in both cold and warm scenarios for the pose-scoring task. The outstanding performance is also shown in imbalanced datasets, which demonstrates the robustness of TransScore. In addition, the gated residual algorithm in TransScore enhances the model to adapt to diverse related tasks. In particular, in the affinity prediction task, we have observed consistent improvements in warm/cold start scenarios. Moreover, it is noticeable that TransScore excels in both accuracy and precision, accurately predicting affinities and their relative ordering. We also conducted an analysis on carbonic anhydrase II, which bears out that TransScore can elaborate the interaction mechanism of the protein-ligand pair, suggesting the potential application of TransScore in drug discovery. Chuqi Lei, Wenkang Wang, Wei Fan 0010, Zhangli Lu, Jing Tang 0002, Min Li 0007 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | CellCircLoc: Deep Neural Network for Predicting and Explaining Cell Line-Specific CircRNA Subcellular LocalizationabstractThe subcellular localization of circular RNAs (circRNAs) is crucial for understanding their functional relevance and regulatory mechanisms. CircRNA subcellular localization exhibits variations across different cell lines, demonstrating the diversity and complexity of circRNA regulation within distinct cellular contexts. However, existing computational methods for predicting circRNA subcellular localization often ignore the importance of cell line specificity and instead train a general model on aggregated data from all cell lines. Considering the diversity and context-dependent behavior of circRNAs across different cell lines, it is imperative to develop cell line-specific models to accurately predict circRNA subcellular localization. In the study, we proposed CellCircLoc, a sequence-based deep learning model for circRNA subcellular localization prediction, which is trained for different cell lines. CellCircLoc utilizes a combination of convolutional neural networks, Transformer blocks, and bidirectional long short-term memory to capture both sequence local features and long-range dependencies within the sequences. In the Transformer blocks, CellCircLoc uses an attentive convolution mechanism to capture the importance of individual nucleotides. Extensive experiments demonstrate the effectiveness of CellCircLoc in accurately predicting circRNA subcellular localization across different cell lines, outperforming other computational models that do not consider cell line specificity. Moreover, the interpretability of CellCircLoc facilitates the discovery of important motifs associated with circRNA subcellular localization. Min Zeng 0004, Jingwei Lu, Chengqian Lu, Shichao Kan, Fei Guo 0001, Min Li 0007 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | S3 Agent: Unlocking the Power of VLLM for Zero-Shot Multi-Modal Sarcasm DetectionabstractMulti-modal sarcasm detection involves determining whether a given multi-modal input conveys sarcastic intent by analyzing the underlying sentiment. Recently, vision large language models have shown remarkable success on various of multi-modal tasks. Inspired by this, we systematically investigate the impact of vision large language models in zero-shot multi-modal sarcasm detection task. Furthermore, to capture different perspectives of sarcastic expressions, we propose a multi-view agent framework, S 3 Agent, designed to enhance zero-shot multi-modal sarcasm detection by leveraging three critical perspectives: superficial expression , semantic information , and sentiment expression . Our experiments on the MMSD2.0 dataset, which involves six models and four prompting strategies, demonstrate that our approach achieves state-of-the-art performance. Our method achieves an average improvement of 13.2% in accuracy. Moreover, we evaluate our method on the text-only sarcasm detection task, where it also surpasses baseline approaches. Peng Wang 0168, Yongheng Zhang 0001, Hao Fei 0001, Qiguang Chen, Jiasheng Si, Wenpeng Lu, Min Li 0007, Libo Qin 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2024 | spGCLF: a versatile deep graph contrastive learning framework for spatial transcriptomics analysisabstractThe rapid development of spatial transcriptomics (ST) has revolutionized the study of spatial heterogeneity and increased the demand for comprehensive methods to effectively characterize spatial domains. As a prerequisite for ST data analysis, spatial domain characterization is a critical step for downstream analysis and biological interpretation. Here, we propose a deep graph contrastive learning framework (spG-CLF) for spatial transcriptomics, which leverages contrastive learning through graph convolutional networks to effectively learn features among spots, addressing the issue of poor cross-platform generalizability. By alternately performing attribute and topological denoising, our method specifically reduces noise in ST data. We demonstrate that s p GCLF is effective for spatial domain identification, trajectory inference, multi-slice integration, and type annotation. Xiang Chen 0029, Junnan Yu, Min Li 0007 |
BIBM | 3 |
| 2024 | DP-BERT: a pre-trained deep language model for depression prediction using microarray dataabstractIn recent years, the increasing number of individuals diagnosed with depression and the growing awareness of its impact on modern society have highlighted the significance of accurate depression diagnosis. Microarray data has played a crucial role in uncovering the genetic mechanisms underlying depression. However, existing methods for depression prediction using microarray data often rely on the selection of differentially expressed genes. This approach disregards important information from other genes and is susceptible to batch effects, thereby limiting generalizability and model stability. To address these limitations, we propose DP-BERT, a depression prediction model based on Bidirectional Encoder Representations from Transformers (BERT). DP-BERT follows a pre-training and fine-tuning paradigm, leveraging a large amount of unlabeled microarray data from diverse sequencing platforms for pretraining to extract comprehensive genetic-level representations of psychiatric disorders. Subsequently, supervised fine-tuning is performed for depression prediction. Experimental results demonstrate that the pre-trained model achieves superior performance in depression prediction. The source code can be obtained from https://github.com/CSUBioGroup/DP-BERT. Junyu Gao 0004, Min Zeng 0004, Fang Wang 0028, Ruiqing Zheng, Jin Liu 0012, Fei Guo 0001, Min Li 0007 |
BIBM | 8 |
| 2024 | UniSleepPos: Sleep Posture Identification System Utilizing Millimeter-wave RadarabstractSleep posture identification is crucial for accurately assessing sleep quality and diagnosing related diseases. In the realm of non-intrusive sleep monitoring, non-contact technologies are becoming increasingly mainstream. Millimeter-wave radar is frequently utilized in sleep posture identification due to its high resolution, strong penetration, and excellent sensitivity. However, traditional radar-based methods for sleep posture identification often struggle with reliability when dealing with diverse individuals and complex sleep environments. To address these challenges, we propose UniSleepPos, which designs a novel dual-view fusion mechanism to integrate depression and elevation angle signals obtained from radar, thus accurately capturing the posture information of the monitored subject in three-dimensional space. Furthermore, we combine sleep posture identification with individual characteristics, utilizing existing individual labels as prior knowledge to assist in sleep posture identification. The integration of prior knowledge provides a valuable information source for the model, helping to enhance its understanding of the data and improve its performance. We collected sleep posture data from eight volunteers using millimeter-wave radar devices under various environmental conditions. Leave-one-subject-out experiments were conducted to validate the effectiveness of UniSleepPos. The results indicated that UniSleepPos significantly outperforms existing methods, demonstrating its potential for practical applications. Min Li 0007, Chu He, Junbin Mao, Min Zeng 0004, Jin Liu 0012 |
BIBM | 2 |
| 2024 | ComLMEss: Combining multiple protein language models enables accurate essential protein predictionabstractAccurately predicting essential proteins is vital for comprehending organism survival, aiding in drug discovery, and informing strategies for treating diseases. While previous computational methods for essential protein prediction have predominantly focused on network-based approaches, recent advancements have seen rapid development in sequence-based prediction methods. However, existing sequence-based prediction methods tend to focus only on sequence-level features, ignoring other biological information at diverse levels. To make use of the diverse information across various biological levels, in this study, we introduce ComLMEss, a novel deep learning framework that combines three protein language models. ComLMEss integrates ProtTrans, ESMFold and OntoProtein, which contain different levels of biological information include protein sequence, conservation, structural, and functional information. ComLMEss employs convolutional neural networks and transformer structure to refine and contextualize the representations from three language models, enabling accurate and robust predictions. Experimental results demonstrate that ComLMEss consistently outperforms existing methods. Ablation studies confirm that the effectiveness of combining different language models focus on different biological information. All results underscore the potential of ComLMEss in essential protein prediction. The source code can be obtained at https://github.com/CSUBioGroup/ComLMEss. Fuhao Zhang, Ruiqing Zheng, Fei Guo 0001, Min Li 0007, Min Zeng 0004 |
BIBM | 6 |
| 2024 | scGDCC: Graph-based Dual Contrastive Calibration for Single Cell MultiOmics ClusteringabstractCell clustering is vital for studying cellular heterogeneity and understanding biological mechanisms. With the advancement of sequencing technologies, it is now possible to obtain multiomics data from single cells, such as ATAC-seq and RNA-seq. Compared to single-omics data, multiomics data offer a more comprehensive view of the cellular landscape. Although several single-cell multiomics clustering methods have been developed, the sparsity and complexity of multiomics data make clustering a challenging computational task. This paper proposes a single cell multiomics clustering method called scGDCC, which is based on graph neural networks and dual contrastive calibration. scGDCC utilizes graph neural networks to capture the neighborhood information of cells and employs dual contrastive calibration to achieve more consistent joint representations of cells. Experiments on five dual-omics datasets (ATAC-seq and RNA-seq) and two triple-omics datasets (ATAC-seq, RNA-seq, and protein) demonstrate the superiority of this clustering method. Additionally, visualization experiments further validate the effectiveness of the joint cellular representations. Huayu Tao, Xinliang Sun, Min Li 0007, Ruiqing Zheng |
BIBM | 4 |
| 2024 | multiTAD: an Attention-Based Deep Learning Model for Identifying TAD Boundaries through Multi-Size Feature IntegrationabstractTopologically associating domains (TADs) are fundamental 3D genome structures that facilitate key gene regulatory interactions. The boundaries of TADs are rich in functional elements critical for maintaining structural integrity, making their identification essential for understanding the relationship between genome organization and gene expression. However, existing algorithms for identifying TAD boundaries often rely on fixed boundary sizes and neglect the varying predictive power of different features. To address these limitations, we introduce multiTAD, an advanced attention-based deep learning model that leverages 12 epigenetic signals to accurately detect TAD boundaries of diverse sizes. multiTAD significantly outperforms mainstream approaches, revealing distinct boundary size preferences across different cell lines. Additionally, multiTAD demonstrates strong cross-cell line predictive capabilities, further highlighting its broad applicability in genomic research. Hanyu Luo, Yajing Deng, Min Li 0007 |
BIBM | 5 |
| 2024 | Heterogeneous network impulsive dynamics for identifying disease-associated genesabstractIdentifying disease-associated genes (DAGs) is important for the research of complex diseases, and network-based methods have been a powerful and elegant strategy for this topic. Genes and their products perform biological functions through synergy in biological networks, but mining useful information from the networks remains an open issue. Therefore, we propose a novel heterogeneous network impulsive dynamics model to identify DAGs more effectively. It inspires a heterogeneous network impulsive dynamical process by imposing impulsive signals at specific nodes, in an enhanced dual-layer heterogeneous network. Then, it extracts the impulsive dynamical signatures of responses of nodes to the impulsive signals so as to infer DAGs. A series of experiments confirm that this model has good performance of inferring DAGs under different conditions, and case studies further demonstrate its effectiveness. Furthermore, a user-friendly web platform is provided to facilitate prioritization and analysis of DAGs. It may become a useful tool for studying complex diseases and relevant genes. Ju Xiang, Shengkai Chen, Lin-Cong-Hua Wang, Xiangmao Meng, Min Li 0007 |
BIBM | 5 |
| 2024 | Aligning Multimodal Biomedical Images and Language via One Large Vision-Language ModelabstractLarge Vision-Language Models (LVLMs) have garnered substantial attention in the biomedical image analysis domain due to their robust vision understanding capabilities. However, current methods rely heavily on dataset- and modality-specific fine-tuning. This involves tuning separate models for each dataset and biomedical modality. In this paper, we introduce a method for aligning multimodal biomedical images and language using a single LVLM, dubbed UniMed-LVLM. Specifically, we devise a General Projection Module (GPM) by integrating multiple image projection branches and implementing dynamic routing between the vision encoder and language decoder within the LLaVA-Med framework. Subsequently, we progressively align multiple biomedical modalities using a Parameter-Efficient Fine-Tuning (PEFT) technique known as Low-Rank Adaptation (LoRA). The model is initially trained on the LLaVA-Med dataset and then fine-tuned on four biomedical image analysis datasets: PathVqa, Slake, VqaRad, and Fitzpatrick17k, enabling the simultaneous analysis of radiology, pathology, and dermatology images. A single model is fine-tuned on three modalities across these datasets and evaluated on all test sets. Experimental results show that UniMed-LVLM improves the average evaluation score by 1.88% across the four datasets, validating its effectiveness in handling multimodal biomedical images. Min Zeng 0004, Jinfeng Ding, Yixiong Liang, Ruiqing Zheng, Min Li 0007, Shichao Kan |
BIBM | 7 |
| 2024 | An Effective Tool for Differential Interaction Analysis in Single-Cell Hi-C DataabstractThe three-dimensional (3D) genome refers to exploring the spatial arrangement of chromatins within the cell nucleus. Analyzing differential interactions, the basic units of the 3D genome, particularly at the single-cell level, is crucial for revealing cell-type-specific functions and states. However, the sparsity and heterogeneity of single-cell 3D genome data pose significant challenges. Although existing methods offer helpful solutions, most of them lack the design for data distribution and cell-specific characteristics. Here, we developed a new method, DISHIC, for identifying differential interactions in single-cell high-throughput Chromatin Conformation Capture (scHi-C) data. Based on the ZINB-WaVE model, DISHIC independently models each bin pair, accounting for both bin-pair-level and cell-level covariates. We validated DISHIC's effectiveness and demonstrated its enhanced effectiveness compared to other methods. The code is available at https://github.com/zhaoyichao777/DISHIC. Yi-chao Zhao, Ruiqing Zheng, Pengzhen Jia, Min Li 0007 |
BIBM | 5 |
| 2024 | LabCLIP: Label-Enhanced Clip for Improving Zero-Shot Text ClassificationabstractZero-shot text classification aims to handle the text classification task without any annotated training data, which can greatly alleviate the data scarcity problem. Current dominant approaches follow a novel text-image matching paradigm, reformulating zero-shot text classification into a text-image matching problem, which can capture the visual image information and show promising performance. Nevertheless, existing text-image matching approaches solely focus on the visual image information, ignoring the semantic knowledge embedded in the text labels. To address the challenge, in the work, we present a label-enhanced CLIP framework (Lab-CLIP) for zero-shot text classification to consider both the visual image and text label semantic information simultaneously. Specifically, LabCLIP first converts the label into the corresponding image, and then injects the text label into the corresponding label image to explicitly capture the label semantic knowledge. We conduct experiments on 8 publicly available zero-shot text classification datasets and experimental results indicate that LabCLIP outperforms previous approaches on all datasets (with 4.3% improvement on average). In addition, we provide extensive analysis on exploring how to effectively incorporate the text label information. Yongheng Zhang 0001, Peng Wang 0168, Qiguang Chen, Jingxuan Zhou, Yongmei Michelle Wang, Min Li 0007, Libo Qin 0001 |
ICASSP | 6 |
| 2024 | scCoRR: A Data-Driven Self-correction Framework for Labeled scRNA-Seq Data
Yongxin He, Jin Liu 0012, Min Li 0007, Ruiqing Zheng |
ISBRA (2) | 3 |
| 2024 | LoopNetica: Predicting Chromatin Loops Using Convolutional Neural Networks and Attention Mechanisms
Hanyu Luo, Min Li 0007 |
ISBRA (3) | 5 |
| 2024 | MSMK: Multiscale Module Kernel for Identifying Disease-Related Genes
Ju Xiang, Shengkai Chen, Xiangmao Meng, Ruiqing Zheng, Min Li 0007 |
ISBRA (1) | 6 |
| 2024 | What Factors Affect Multi-Modal In-Context Learning? An In-Depth ExplorationabstractRecently, rapid advancements in Multi-Modal In-Context Learning (MM-ICL) have achieved notable success, which is capable of achieving superior performance across various tasks without requiring additional parameter tuning. However, the underlying rules for the effectiveness of MM-ICL remain under-explored. To fill this gap, this work aims to investigate the research question: "_What factors affect the performance of MM-ICL?_" To this end, we investigate extensive experiments on the three core steps of MM-ICL including demonstration retrieval, demonstration ordering, and prompt construction using 6 vision large language models and 20 strategies. Our findings highlight (1) the necessity of a multi-modal retriever for demonstration retrieval, (2) the importance of intra-demonstration ordering over inter-demonstration ordering, and (3) the enhancement of task comprehension through introductory instructions in prompts. We hope this study can serve as a foundational guide for optimizing MM-ICL strategies in future research. Libo Qin 0001, Qiguang Chen, Hao Fei 0003, Zhi Chen 0006, Min Li 0007, Wanxiang Che |
NeurIPS | 5 |
| 2024 | Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal LearningabstractTraining models with longer in-context lengths is a significant challenge for multimodal machine learning due to substantial GPU memory and computational costs. This exploratory study does not present state-of-the-art models; rather, it introduces an innovative method designed to increase in-context text length in multi-modality large language models (MLLMs) efficiently. We present \ModelFullName (\ModelName), which processes long in-context text using visual tokens. This technique significantly reduces GPU memory usage and floating point operations (FLOPs). For instance, our method expands the pre-training in-context length from 256 to 2048 tokens with fewer FLOPs for a 56 billion parameter MOE model. Experimental results demonstrate that \ModelName enhances OCR capabilities and delivers superior performance on common downstream benchmarks for in-context few-shot evaluation. Additionally, \ModelName proves effective for long context inference, achieving results comparable to full text input while maintaining computational efficiency. Alex Jinpeng Wang, Min Li 0007, Zheng Shou 0001 |
NeurIPS | 4 |
| 2024 | scMLC: an accurate and robust multiplex community detection method for single-cell multi-omics dataabstractClustering cells based on single-cell multi-modal sequencing technologies provides an unprecedented opportunity to create high-resolution cell atlas, reveal cellular critical states and study health and diseases. However, effectively integrating different sequencing data for cell clustering remains a challenging task. Motivated by the successful application of Louvain in scRNA-seq data, we propose a single-cell multi-modal Louvain clustering framework, called scMLC, to tackle this problem. scMLC builds multiplex single- and cross-modal cell-to-cell networks to capture modal-specific and consistent information between modalities and then adopts a robust multiplex community detection method to obtain the reliable cell clusters. In comparison with 15 state-of-the-art clustering methods on seven real datasets simultaneously measuring gene expression and chromatin accessibility, scMLC achieves better accuracy and stability in most datasets. Synthetic results also indicate that the cell-network-based integration strategy of multi-omics data is superior to other strategies in terms of generalization. Moreover, scMLC is flexible and can be extended to single-cell sequencing data with more than two modalities. Ruiqing Zheng, Jin Liu 0012, Min Li 0007 |
Briefings Bioinform. | 4 |
| 2024 | stAA: adversarial graph autoencoder for spatial clustering task of spatially resolved transcriptomicsabstractWith the development of spatially resolved transcriptomics technologies, it is now possible to explore the gene expression profiles of single cells while preserving their spatial context. Spatial clustering plays a key role in spatial transcriptome data analysis. In the past 2 years, several graph neural network-based methods have emerged, which significantly improved the accuracy of spatial clustering. However, accurately identifying the boundaries of spatial domains remains a challenging task. In this article, we propose stAA, an adversarial variational graph autoencoder, to identify spatial domain. stAA generates cell embedding by leveraging gene expression and spatial information using graph neural networks and enforces the distribution of cell embeddings to a prior distribution through Wasserstein distance. The adversarial training process can make cell embeddings better capture spatial domain information and more robust. Moreover, stAA incorporates global graph information into cell embeddings using labels generated by pre-clustering. Our experimental results show that stAA outperforms the state-of-the-art methods and achieves better clustering results across different profiling platforms and various resolutions. We also conducted numerous biological analyses and found that stAA can identify fine-grained structures in tissues, recognize different functional subtypes within tumors and accurately identify developmental trajectories. Zhaoyu Fang, Ruiqing Zheng, Jin A, Mingzhu Yin, Min Li 0007 |
Briefings Bioinform. | 6 |
| 2024 | A comprehensive review of protein-centric predictors for biomolecular interactions: from proteins to nucleic acids and beyondabstractProteins interact with diverse ligands to perform a large number of biological functions, such as gene expression and signal transduction. Accurate identification of these protein-ligand interactions is crucial to the understanding of molecular mechanisms and the development of new drugs. However, traditional biological experiments are time-consuming and expensive. With the development of high-throughput technologies, an increasing amount of protein data is available. In the past decades, many computational methods have been developed to predict protein-ligand interactions. Here, we review a comprehensive set of over 160 protein-ligand interaction predictors, which cover protein-protein, protein-nucleic acid, protein-peptide and protein-other ligands (nucleotide, heme, ion) interactions. We have carried out a comprehensive analysis of the above four types of predictors from several significant perspectives, including their inputs, feature profiles, models, availability, etc. The current methods primarily rely on protein sequences, especially utilizing evolutionary information. The significant improvement in predictions is attributed to deep learning methods. Additionally, sequence-based pretrained models and structure-based approaches are emerging as new trends. Pengzhen Jia, Fuhao Zhang, Chaojin Wu, Min Li 0007 |
Briefings Bioinform. | 4 |
| 2024 | CAKE: a flexible self-supervised framework for enhancing cell visualization, clustering and rare cell identificationabstractSingle cell sequencing technology has provided unprecedented opportunities for comprehensively deciphering cell heterogeneity. Nevertheless, the high dimensionality and intricate nature of cell heterogeneity have presented substantial challenges to computational methods. Numerous novel clustering methods have been proposed to address this issue. However, none of these methods achieve the consistently better performance under different biological scenarios. In this study, we developed CAKE, a novel and scalable self-supervised clustering method, which consists of a contrastive learning model with a mixture neighborhood augmentation for cell representation learning, and a self-Knowledge Distiller model for the refinement of clustering results. These designs provide more condensed and cluster-friendly cell representations and improve the clustering performance in term of accuracy and robustness. Furthermore, in addition to accurately identifying the major type cells, CAKE could also find more biologically meaningful cell subgroups and rare cell types. The comprehensive experiments on real single-cell RNA sequencing datasets demonstrated the superiority of CAKE in visualization and clustering over other comparison methods, and indicated its extensive application in the field of cell heterogeneity analysis. Contact: Ruiqing Zheng. ([email protected]). Jin Liu 0012, Weixing Zeng, Shichao Kan, Min Li 0007, Ruiqing Zheng |
Briefings Bioinform. | 4 |
| 2024 | TripHLApan: predicting HLA molecules binding peptides based on triple coding matrix and transfer learningabstractHuman leukocyte antigen (HLA) recognizes foreign threats and triggers immune responses by presenting peptides to T cells. Computationally modeling the binding patterns between peptide and HLA is very important for the development of tumor vaccines. However, it is still a big challenge to accurately predict HLA molecules binding peptides. In this paper, we develop a new model TripHLApan for predicting HLA molecules binding peptides by integrating triple coding matrix, BiGRU + Attention models, and transfer learning strategy. We have found the main interaction site regions between HLA molecules and peptides, as well as the correlation between HLA encoding and binding motifs. Based on the discovery, we make the preprocessing and coding closer to the natural biological process. Besides, due to the input being based on multiple types of features and the attention module focused on the BiGRU hidden layer, TripHLApan has learned more sequence level binding information. The application of transfer learning strategies ensures the accuracy of prediction results under special lengths (peptides in length 8) and model scalability with the data explosion. Compared with the current optimal models, TripHLApan exhibits strong predictive performance in various prediction environments with different positive and negative sample ratios. In addition, we validate the superiority and scalability of TripHLApan's predictive performance using additional latest data sets, ablation experiments and binding reconstitution ability in the samples of a melanoma patient. The results show that TripHLApan is a powerful tool for predicting the binding of HLA-I and HLA-II molecular peptides for the synthesis of tumor vaccines. TripHLApan is publicly available at https://github.com/CSUBioGroup/TripHLApan.git. Meng Wang 0067, Chuqi Lei, Jianxin Wang 0001, Yaohang Li, Min Li 0007 |
Briefings Bioinform. | 5 |
| 2024 | A comprehensive computational benchmark for evaluating deep learning-based protein function prediction approachesabstractProteins play an important role in life activities and are the basic units for performing functions. Accurately annotating functions to proteins is crucial for understanding the intricate mechanisms of life and developing effective treatments for complex diseases. Traditional biological experiments struggle to keep pace with the growing number of known proteins. With the development of high-throughput sequencing technology, a wide variety of biological data provides the possibility to accurately predict protein functions by computational methods. Consequently, many computational methods have been proposed. Due to the diversity of application scenarios, it is necessary to conduct a comprehensive evaluation of these computational methods to determine the suitability of each algorithm for specific cases. In this study, we present a comprehensive benchmark, BeProf, to process data and evaluate representative computational methods. We first collect the latest datasets and analyze the data characteristics. Then, we investigate and summarize 17 state-of-the-art computational methods. Finally, we propose a novel comprehensive evaluation metric, design eight application scenarios and evaluate the performance of existing methods on these scenarios. Based on the evaluation, we provide practical recommendations for different scenarios, enabling users to select the most suitable method for their specific needs. All of these servers can be obtained from https://csuligroup.com/BEPROF and https://github.com/CSUBioGroup/BEPROF. Wenkang Wang, Yunyan Shuai, Qiurong Yang, Fuhao Zhang, Min Zeng 0004, Min Li 0007 |
Briefings Bioinform. | 6 |
| 2024 | scMAE: a masked autoencoder for single-cell RNA-seq clusteringabstractMOTIVATION: Single-cell RNA sequencing has emerged as a powerful technology for studying gene expression at the individual cell level. Clustering individual cells into distinct subpopulations is fundamental in scRNA-seq data analysis, facilitating the identification of cell types and exploration of cellular heterogeneity. Despite the recent development of many deep learning-based single-cell clustering methods, few have effectively exploited the correlations among genes, resulting in suboptimal clustering outcomes. RESULTS: Here, we propose a novel masked autoencoder-based method, scMAE, for cell clustering. scMAE perturbs gene expression and employs a masked autoencoder to reconstruct the original data, learning robust and informative cell representations. The masked autoencoder introduces a masking predictor, which captures relationships among genes by predicting whether gene expression values are masked. By integrating this masking mechanism, scMAE effectively captures latent structures and dependencies in the data, enhancing clustering performance. We conducted extensive comparative experiments using various clustering evaluation metrics on 15 scRNA-seq datasets from different sequencing platforms. Experimental results indicate that scMAE outperforms other state-of-the-art methods on these datasets. In addition, scMAE accurately identifies rare cell types, which are challenging to detect due to their low abundance. Furthermore, biological analyses confirm the biological significance of the identified cell subpopulations. AVAILABILITY AND IMPLEMENTATION: The source code of scMAE is available at: https://zenodo.org/records/10465991. Zhaoyu Fang, Ruiqing Zheng, Min Li 0007 |
Bioinform. | 3 |
| 2024 | Assembling spatial clustering framework for heterogeneous spatial transcriptomics data with GRAPHDeepabstractMOTIVATION: Spatial clustering is essential and challenging for spatial transcriptomics' data analysis to unravel tissue microenvironment and biological function. Graph neural networks are promising to address gene expression profiles and spatial location information in spatial transcriptomics to generate latent representations. However, choosing an appropriate graph deep learning module and graph neural network necessitates further exploration and investigation. RESULTS: In this article, we present GRAPHDeep to assemble a spatial clustering framework for heterogeneous spatial transcriptomics data. Through integrating 2 graph deep learning modules and 20 graph neural networks, the most appropriate combination is decided for each dataset. The constructed spatial clustering method is compared with state-of-the-art algorithms to demonstrate its effectiveness and superiority. The significant new findings include: (i) the number of genes or proteins of spatial omics data is quite crucial in spatial clustering algorithms; (ii) the variational graph autoencoder is more suitable for spatial clustering tasks than deep graph infomax module; (iii) UniMP, SAGE, SuperGAT, GATv2, GCN, and TAG are the recommended graph neural networks for spatial clustering tasks; and (iv) the used graph neural network in the existent spatial clustering frameworks is not the best candidate. This study could be regarded as desirable guidance for choosing an appropriate graph neural network for spatial clustering. AVAILABILITY AND IMPLEMENTATION: The source code of GRAPHDeep is available at https://github.com/narutoten520/GRAPHDeep. The studied spatial omics data are available at https://zenodo.org/record/8141084. Zhaoyu Fang, Lining Zhang, Dong-Sheng Cao 0001, Min Li 0007, Mingzhu Yin |
Bioinform. | 6 |
| 2024 | BertSNR: an interpretable deep learning framework for single-nucleotide resolution identification of transcription factor binding sites based on DNA language modelabstractMOTIVATION: Transcription factors are pivotal in the regulation of gene expression, and accurate identification of transcription factor binding sites (TFBSs) at high resolution is crucial for understanding the mechanisms underlying gene regulation. The task of identifying TFBSs from DNA sequences is a significant challenge in the field of computational biology today. To address this challenge, a variety of computational approaches have been developed. However, these methods face limitations in their ability to achieve high-resolution identification and often lack interpretability. RESULTS: We propose BertSNR, an interpretable deep learning framework for identifying TFBSs at single-nucleotide resolution. BertSNR integrates sequence-level and token-level information by multi-task learning based on pre-trained DNA language models. Benchmarking comparisons show that our BertSNR outperforms the existing state-of-the-art methods in TFBS predictions. Importantly, we enhanced the interpretability of the model through attentional weight visualization and motif analysis, and discovered the subtle relationship between attention weight and motif. Moreover, BertSNR effectively identifies TFBSs in promoter regions, facilitating the study of intricate gene regulation. AVAILABILITY AND IMPLEMENTATION: The BertSNR source code can be found at https://github.com/lhy0322/BertSNR. Hanyu Luo, Min Zeng 0004, Rui Yin 0002, Pingjian Ding, Lingyun Luo, Min Li 0007 |
Bioinform. | 7 |
| 2024 | Drug repositioning with adaptive graph convolutional networksabstractMOTIVATION: Drug repositioning is an effective strategy to identify new indications for existing drugs, providing the quickest possible transition from bench to bedside. With the rapid development of deep learning, graph convolutional networks (GCNs) have been widely adopted for drug repositioning tasks. However, prior GCNs based methods exist limitations in deeply integrating node features and topological structures, which may hinder the capability of GCNs. RESULTS: In this study, we propose an adaptive GCNs approach, termed AdaDR, for drug repositioning by deeply integrating node features and topological structures. Distinct from conventional graph convolution networks, AdaDR models interactive information between them with adaptive graph convolution operation, which enhances the expression of model. Concretely, AdaDR simultaneously extracts embeddings from node features and topological structures and then uses the attention mechanism to learn adaptive importance weights of the embeddings. Experimental results show that AdaDR achieves better performance than multiple baselines for drug repositioning. Moreover, in the case study, exploratory analyses are offered for finding novel drug-disease associations. AVAILABILITY AND IMPLEMENTATION: The soure code of AdaDR is available at: https://github.com/xinliangSun/AdaDR. Xinliang Sun, Xiao Jia 0020, Zhangli Lu, Jing Tang 0002, Min Li 0007 |
Bioinform. | 5 |
| 2024 | Plug-and-Play latent feature editing for orientation-adaptive quantitative susceptibility mapping neural networksabstractQuantitative susceptibility mapping (QSM) is a post-processing technique for deriving tissue magnetic susceptibility distribution from MRI phase measurements. Deep learning (DL) algorithms hold great potential for solving the ill-posed QSM reconstruction problem. However, a significant challenge facing current DL-QSM approaches is their limited adaptability to magnetic dipole field orientation variations during training and testing. In this work, we propose a novel Orientation-Adaptive Latent Feature Editing (OA-LFE) module to learn the encoding of acquisition orientation vectors and seamlessly integrate them into the latent features of deep networks. Importantly, it can be directly Plug-and-Play (PnP) into various existing DL-QSM architectures, enabling reconstructions of QSM from arbitrary magnetic dipole orientations. Its effectiveness is demonstrated by combining the OA-LFE module into our previously proposed phase-to-susceptibility single-step instant QSM (iQSM) network, which was initially tailored for pure-axial acquisitions. The proposed OA-LFE-empowered iQSM, which we refer to as iQSM+, is trained in a simulated-supervised manner on a specially-designed simulation brain dataset. Comprehensive experiments are conducted on simulated and in vivo human brain datasets, encompassing subjects ranging from healthy individuals to those with pathological conditions. These experiments involve various MRI platforms (3T and 7T) and aim to compare our proposed iQSM+ against several established QSM reconstruction frameworks, including the original iQSM. The iQSM+ yields QSM images with significantly improved accuracies and mitigates artifacts, surpassing other state-of-the-art DL-QSM algorithms. The PnP OA-LFE module’s versatility was further demonstrated by its successful application to xQSM, a distinct DL-QSM network for dipole inversion. In conclusion, this work introduces a new DL paradigm, allowing researchers to develop innovative QSM methods without requiring a complete overhaul of their existing architectures. Yang Gao 0030, Shanshan Shan, Pengfei Rong, Min Li 0007, Alan H. Wilman, G. Bruce Pike, Feng Liu 0005, Hongfu Sun |
Medical Image Anal. | 6 |
| 2024 | scMoMtF: An interpretable multitask learning framework for single-cell multi-omics data analysisabstractWith the rapidly development of biotechnology, it is now possible to obtain single-cell multi-omics data in the same cell. However, how to integrate and analyze these single-cell multi-omics data remains a great challenge. Herein, we introduce an interpretable multitask framework (scMoMtF) for comprehensively analyzing single-cell multi-omics data. The scMoMtF can simultaneously solve multiple key tasks of single-cell multi-omics data including dimension reduction, cell classification and data simulation. The experimental results shows that scMoMtF outperforms current state-of-the-art algorithms on these tasks. In addition, scMoMtF has interpretability which allowing researchers to gain a reliable understanding of potential biological features and mechanisms in single-cell multi-omics data. Wei Lan 0001, Tongsheng Ling, Qingfeng Chen, Ruiqing Zheng, Min Li 0007, Yi Pan 0001 |
PLoS Comput. Biol. | 5 |
| 2024 | Dopcc: Detecting Overlapping Protein Complexes via Multi-Metrics and Co-Core Attachment MethodabstractIdentification of protein complex is an important issue in the field of system biology, which is crucial to understanding the cellular organization and inferring protein functions. Recently, many computational methods have been proposed to detect protein complexes from protein-protein interaction (PPI) networks. However, most of these methods only focus on local information of proteins in the PPI network, which are easily affected by the noise in the PPI network. Meanwhile, it's still challenging to detect protein complexes, especially for overlapping cases. To address these issues, we propose a new method, named Dopcc, to detect overlapping protein complexes by constructing a multi-metrics network according to different strategies. First, we adopt the Jaccard coefficient to measure the neighbor similarity between proteins and denoise the PPI network. Then, we propose a new strategy, integrating hierarchical compressing with network embedding, to capture the high-order structural similarity between proteins. Further, a new co-core attachment strategy is proposed to detect overlapping protein complexes from multi-metrics. The experimental results show that our proposed method, Dopcc, outperforms the other eight state-of-the-art methods in terms of F-measure, MMR, and Composite Score on two yeast datasets. Wenkang Wang, Xiangmao Meng, Ju Xiang, Hayat Dino Bedru, Min Li 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | DyNRW: Time-Series Dynamical Networks for Identifying HCC-Related GenesabstractHepatocellular carcinoma (HCC) is a multifactorial and highly complex disease. Gaining a comprehensive understanding of the genetic factors associated with HCC is crucial for unraveling its intricate pathogenesis and identifying potential biomarkers. Recent advances suggest that biological network-based strategies are useful in prioritizing genes associated with diseases. However, existing network models predominantly rely on static biological networks, which to some extent hinder the modeling of dynamic biological processes and restrict the predictive capacity of disease genes. Therefore, we proposed the Dynamic Networks Random Walk (DyNRW) method to identify HCC-related genes. DyNRW overcomes the limitations of previous static network-based methods and models the dynamic regulatory relationships between genes in the biological system by constructing a time-series dynamic network based on the progression of HCC. To construct a more comprehensive biological network model, DyNRW presents an effective background-temporal multi-layer network framework to combine both static and dynamic network information. DyNRW extends the random-walk process to the multi-layer network, enabling the extraction of gene scores associated with HCC. According to the experimental results, DyNRW demonstrates better performance and stability compared to other state-of-the-art algorithms, and yields a set of promising candidate genes, many of which are confirmed by further biological validation. Jin A, Ju Xiang, Min Li 0007 |
BIBM | 3 |
| 2023 | A deep graph convolution network with attention for clustering scRNA-seq dataabstractIdentifying cell types is a primary objective in single-cell RNA sequencing (scRNA-seq) analysis, and clustering is a widely used method for achieving this goal. However, the abundance of data and the presence of noise pose significant challenges for single-cell clustering. We propose a novel method called scGCNClustering, which leverages a deep graph convolution network (GCN) with attention. scGCNClustering comprises two core models. The first model is a deep GCN encoder that effectively preserves both the topological structure and feature information in scRNA-seq data, enabling accurate cell segregation. Additionally, we incorporate an attention module into the GCN encoder to capture precise global similarities between cells. The second model is a deep zero-inflated negative binomial (ZINB) decoder, which approximates the true distribution of scRNA-seq data. Extensive analysis conducted on six real scRNA-seq datasets demonstrates that scGCNClustering achieves promising performance in scRNA-seq clustering. Xiang Chen 0029, Junnan Yu, Min Li 0007 |
BIBM | 4 |
| 2023 | DILM-ICD: A Deep Iterative Learning Model for Automatic ICD CodingabstractAutomatic International Classification of Disease (ICD) coding plays a crucial role in assigning ICD codes to electronic medical records. This task presents a challenging multi-label text classification problem due to the vast number of ICD codes and the imbalanced label distribution. However, accurately predicting all labels simultaneously is extremely difficult for such the large label space. In this paper, we propose a novel model called Deep Iterative Learning Model (DILM-ICD), which uses an iterative learning framework to perform automatic ICD coding task. The iterative learning framework can refine the prediction results by repeating the iteration modules, which simulates the human-like coding process. In addition, we propose a multi-head text-label matching mechanism, which combines the embedding ICD description information to better match the relationship between text and label. The combination of the iterative learning framework with the multi-head text-label matching mechanism enables the model pay attention to lowfrequency ICD codes. DILM-ICD is evaluated on the MIMIC-III-full dataset and MIMIC-III-50 dataset. The experimental results show that DILM-ICD achieves state-of-the-art results across multiple evaluation metrics, which demonstrates the effectiveness of our proposed model. Weiyan Qiu, Yifan Wu 0008, Kunying Niu, Min Zeng 0004, Min Li 0007 |
BIBM | 6 |
| 2023 | Protein function prediction using graph neural network with multi-type biological knowledgeabstractProteins play crucial roles in diverse biological functions, and accurately annotating their functions is essential for understanding cellular mechanisms and developing therapies for complex diseases. Computational methods have been proposed as alternatives to laborious experimental approaches. However, existing network-based methods focus on the protein-protein interaction (PPI) networks, while the proteins without interactions are ignored. To address this limitation, we propose a novel deep learning framework for protein function prediction, named PFP-GMB, which incorporates multi-type biological knowledge to consider the proteins not present in the PPI networks. PFP-GMB leverages a pre-trained protein language model to extract sequence representations. Moreover, PPIs and orthology relationships are used to generate functional related features via graph neural networks and attention mechanisms. Finally, these multi-type features are fused for protein function prediction. Compared to eight state-of-the-art methods, PFP-GMB outperforms all of them in terms of F-max and AUPR. The ablation studies further confirm the relevance and significance of the multi-type biological knowledge incorporated into PFP-GMB for protein function prediction. Yuyan Shuai, Wenkang Wang, Min Zeng 0004, Min Li 0007 |
BIBM | 5 |
| 2023 | A flexible gene regulatory network reconstruction method based on autoencoder and graph attention networkabstractReconstruction of gene regulatory networks from gene expression profile have been an important challenge task in system biology for decades. Recently, with the advancement of single cell RNA-seq technology, the studies in this field turn from bulk gene expression data to scRNA-seq data. However, the complexity of regulatory relationships and high noise in scRNA-seq introduce further challenges in addressing this issue. In this study, we proposed a flexible gene regulatory network reconstruction method based on autoencoder and graph attention network, called scGiant. scGiant incorporates autoencoder to capture the non-linear representation of genes with graph attention network to learn regulatory relationships among genes. To evaluate the performance of scGiant, we compared it with seven state-of-the-art GRN inference algorithms on four real single-cell RNA sequencing datasets, and the results demonstrate that scGiant is superior in accuracy and scalability. The inferred core GRN of CD8+ naïve T cells also demonstrates its potential in practical biological applications. Ruiqing Zheng, Yanping Zeng, Weixing Zeng, Min Li 0007 |
BIBM | 4 |
| 2023 | End-to-end Task-oriented Dialogue: A Survey of Tasks, Methods, and Future DirectionsabstractEnd-to-end task-oriented dialogue (EToD) can directly generate responses in an end-to-end fashion without modular training, which attracts escalating popularity.The advancement of deep neural networks, especially the successful use of large pre-trained models, has further led to significant progress in EToD research in recent years.In this paper, we present a thorough review and provide a unified perspective to summarize existing approaches as well as recent trends to advance the development of EToD research.The contributions of this paper can be summarized: (1) First survey: to our knowledge, we take the first step to present a thorough survey of this research field; (2) New taxonomy: we first introduce a unified perspective for EToD, including (i) Modularly EToD and (ii) Fully EToD; (3) New Frontiers: we discuss some potential frontier areas as well as the corresponding challenges, hoping to spur breakthrough research in EToD field; (4) Abundant resources: we build a public website 1 , where EToD researchers could directly access the recent progress.We hope this work can serve as a thorough reference for the EToD research community.EToD Modularly EToD ( §3.1) Libo Qin 0001, Wenbo Pan 0001, Qiguang Chen, Lizi Liao, Zhou Yu 0005, Yue Zhang 0004, Wanxiang Che, Min Li 0007 |
EMNLP | 8 |
| 2023 | Singularformer: Learning to Decompose Self-Attention to Linearize the Complexity of TransformerabstractTransformers achieve excellent performance in a variety of domains since they can capture long-distance dependencies through the self-attention mechanism. However, self-attention is computationally costly due to its quadratic complexity and high memory consumption. In this paper, we propose a novel Transformer variant (Singularformer) that uses neural networks to learn the singular value decomposition process of the attention matrix to design a linear-complexity and memory-efficient global self-attention mechanism. Specifically, we decompose the attention matrix into the product of three matrix factors based on singular value decomposition and design neural networks to learn these matrix factors, then the associative law of matrix multiplication is used to linearize the calculation of self-attention. The above procedure allows us to compute self-attention as two-dimensional reduction processes in the first and second token dimensional spaces, followed by a multi-head self-attention computational process on the first dimensional reduced token features. Experimental results on 8 real-world datasets demonstrate that Singularformer performs favorably against the other Transformer variants with lower time and space complexity. Our source code is publicly available at https://github.com/CSUBioGroup/Singularformer. Yifan Wu 0008, Shichao Kan, Min Zeng 0004, Min Li 0007 |
IJCAI | 4 |
| 2023 | Bubble: a fast single-cell RNA-seq imputation using an autoencoder constrained by bulk RNA-seq dataabstractSingle-cell RNA-sequencing technology (scRNA-seq) brings research to single-cell resolution. However, a major drawback of scRNA-seq is large sparsity, i.e. expressed genes with no reads due to technical noise or limited sequence depth during the scRNA-seq protocol. This phenomenon is also called 'dropout' events, which likely affect downstream analyses such as differential expression analysis, the clustering and visualization of cell subpopulations, cellular trajectory inference, etc. Therefore, there is a need to develop a method to identify and impute these dropout events. We propose Bubble, which first identifies dropout events from all zeros based on expression rate and coefficient of variation of genes within cell subpopulation, and then leverages an autoencoder constrained by bulk RNA-seq data to only impute those values. Unlike other deep learning-based imputation methods, Bubble fuses the matched bulk RNA-seq data as a constraint to reduce the introduction of false positive signals. Using simulated and several real scRNA-seq datasets, we demonstrate that Bubble enhances the recovery of missing values, gene-to-gene and cell-to-cell correlations, and reduces the introduction of false positive signals. Regarding some crucial downstream analyses of scRNA-seq data, Bubble facilitates the identification of differentially expressed genes, improves the performance of clustering and visualization, and aids the construction of cellular trajectory. More importantly, Bubble provides fast and scalable imputation with minimal memory usage. Xuhua Yan, Ruiqing Zheng, Min Li 0007 |
Briefings Bioinform. | 4 |
| 2023 | DKADE: a novel framework based on deep learning and knowledge graph for identifying adverse drug events and related medicationsabstractAdverse drug events (ADEs) are common in clinical practice and can cause significant harm to patients and increase resource use. Natural language processing (NLP) has been applied to automate ADE detection, but NLP systems become less adaptable when drug entities are missing or multiple medications are specified in clinical narratives. Additionally, no Chinese-language NLP system has been developed for ADE detection due to the complexity of Chinese semantics, despite ˃10 million cases of drug-related adverse events occurring annually in China. To address these challenges, we propose DKADE, a deep learning and knowledge graph-based framework for identifying ADEs. DKADE infers missing drug entities and evaluates their correlations with ADEs by combining medication orders and existing drug knowledge. Moreover, DKADE can automatically screen for new adverse drug reactions. Experimental results show that DKADE achieves an overall F1-score value of 91.13%. Furthermore, the adaptability of DKADE is validated using real-world external clinical data. In summary, DKADE is a powerful tool for studying drug safety and automating adverse event monitoring. Ze-Ying Feng, Jun-Long Ma, Min Li 0007, Ge-Fei He, Dong-Sheng Cao 0001 |
Briefings Bioinform. | 4 |
| 2023 | GraphLncLoc: long non-coding RNA subcellular localization prediction using graph convolutional networks based on sequence to graph transformationabstractThe subcellular localization of long non-coding RNAs (lncRNAs) is crucial for understanding lncRNA functions. Most of existing lncRNA subcellular localization prediction methods use k-mer frequency features to encode lncRNA sequences. However, k-mer frequency features lose sequence order information and fail to capture sequence patterns and motifs of different lengths. In this paper, we proposed GraphLncLoc, a graph convolutional network-based deep learning model, for predicting lncRNA subcellular localization. Unlike previous studies encoding lncRNA sequences by using k-mer frequency features, GraphLncLoc transforms lncRNA sequences into de Bruijn graphs, which transforms the sequence classification problem into a graph classification problem. To extract the high-level features from the de Bruijn graph, GraphLncLoc employs graph convolutional networks to learn latent representations. Then, the high-level feature vectors derived from de Bruijn graph are fed into a fully connected layer to perform the prediction task. Extensive experiments show that GraphLncLoc achieves better performance than traditional machine learning models and existing predictors. In addition, our analyses show that transforming sequences into graphs has more distinguishable features and is more robust than k-mer frequency features. The case study shows that GraphLncLoc can uncover important motifs for nucleus subcellular localization. GraphLncLoc web server is available at http://csuligroup.com:8000/GraphLncLoc/. Min Li 0007, Baoying Zhao, Rui Yin 0002, Chengqian Lu, Fei Guo 0001, Min Zeng 0004 |
Briefings Bioinform. | 1 |
| 2023 | A review of enzyme design in catalytic stability by artificial intelligenceabstractThe design of enzyme catalytic stability is of great significance in medicine and industry. However, traditional methods are time-consuming and costly. Hence, a growing number of complementary computational tools have been developed, e.g. ESMFold, AlphaFold2, Rosetta, RosettaFold, FireProt, ProteinMPNN. They are proposed for algorithm-driven and data-driven enzyme design through artificial intelligence (AI) algorithms including natural language processing, machine learning, deep learning, variational autoencoder/generative adversarial network, message passing neural network (MPNN). In addition, the challenges of design of enzyme catalytic stability include insufficient structured data, large sequence search space, inaccurate quantitative prediction, low efficiency in experimental validation and a cumbersome design process. The first principle of the enzyme catalytic stability design is to treat amino acids as the basic element. By designing the sequence of an enzyme, the flexibility and stability of the structure are adjusted, thus controlling the catalytic stability of the enzyme in a specific industrial environment or in an organism. Common indicators of design goals include the change in denaturation energy (ΔΔG), melting temperature (ΔTm), optimal temperature (Topt), optimal pH (pHopt), etc. In this review, we summarized and evaluated the enzyme design in catalytic stability by AI in terms of mechanism, strategy, data, labeling, coding, prediction, testing, unit, integration and prospect. Yongfan Ming, Wenkang Wang, Rui Yin 0002, Min Zeng 0004, Shizhe Tang, Min Li 0007 |
Briefings Bioinform. | 7 |
| 2023 | A comprehensive assessment and comparison of tools for HLA class I peptide-binding predictionabstractHuman leukocyte antigen class I (HLA-I) molecules bind intracellular peptides produced by protein hydrolysis and present them to the T cells for immune recognition and response. Prediction of peptides that bind HLA-I molecules is very important in immunotherapy. A growing number of computational predictors have been developed in recent years. We survey a comprehensive collection of 27 tools focusing on their input and output data characteristics, key aspects of the underlying predictive models and their availability. Moreover, we evaluate predictive performance for eight representative predictors. We consider a wide spectrum of relevant aspects including allele-specific analysis, influence of negative to positive data ratios and runtime. We also curate high-quality benchmark datasets based on analysis of the consistency of the data labels. Results reveal that each considered method provides accurate results, which can be explained by our analysis that finds that their predictive models capture meaningful binding motifs. Although some methods are overall more accurate than others, we find that none of them is universally superior. We provide a comprehensive comparison of the convenience as well as the accuracy of the methods under specific prediction scenarios, such as for specific alleles, metrics of predictive performance and constraints on runtime. Our systematic and broad analysis provides informative clues to the users to identify the most suitable tools for a given prediction scenario and for the developers to design future methods. Meng Wang 0067, Lukasz A. Kurgan, Min Li 0007 |
Briefings Bioinform. | 3 |
| 2023 | RLBind: a deep learning method to predict RNA-ligand binding sitesabstractIdentification of RNA-small molecule binding sites plays an essential role in RNA-targeted drug discovery and development. These small molecules are expected to be leading compounds to guide the development of new types of RNA-targeted therapeutics compared with regular therapeutics targeting proteins. RNAs can provide many potential drug targets with diverse structures and functions. However, up to now, only a few methods have been proposed. Predicting RNA-small molecule binding sites still remains a big challenge. New computational model is required to better extract the features and predict RNA-small molecule binding sites more accurately. In this paper, a deep learning model, RLBind, was proposed to predict RNA-small molecule binding sites from sequence-dependent and structure-dependent properties by combining global RNA sequence channel and local neighbor nucleotides channel. To our best knowledge, this research was the first to develop a convolutional neural network for RNA-small molecule binding sites prediction. Furthermore, RLBind also can be used as a potential tool when the RNA experimental tertiary structure is not available. The experimental results show that RLBind outperforms other state-of-the-art methods in predicting binding sites. Therefore, our study demonstrates that the combination of global information for full-length sequences and local information for limited local neighbor nucleotides in RNAs can improve the model's predictive performance for binding sites prediction. All datasets and resource codes are available at https://github.com/KailiWang1/RLBind. Renyi Zhou, Yifan Wu 0008, Min Li 0007 |
Briefings Bioinform. | 4 |
| 2023 | Inferring single-cell gene regulatory network by non-redundant mutual informationabstractGene regulatory network plays a crucial role in controlling the biological processes of living creatures. Deciphering the complex gene regulatory networks from experimental data remains a major challenge in system biology. Recent advances in single-cell RNA sequencing technology bring massive high-resolution data, enabling computational inference of cell-specific gene regulatory networks (GRNs). Many relevant algorithms have been developed to achieve this goal in the past years. However, GRN inference is still less ideal due to the extra noises involved in pseudo-time information and large amounts of dropouts in datasets. Here, we present a novel GRN inference method named Normi, which is based on non-redundant mutual information. Normi manipulates these problems by employing a sliding size-fixed window approach on the entire trajectory and conducts average smoothing strategy on the gene expression of the cells in each window to obtain representative cells. To further alleviate the impact of dropouts, we utilize the mixed KSG estimator to quantify the high-order time-delayed mutual information among genes, then filter out the redundant edges by adopting Max-Relevance and Min Redundancy algorithm. Moreover, we determined the optimal time delay for each gene pair by distance correlation. Normi outperforms other state-of-the-art GRN inference methods on both simulated data and single-cell RNA sequencing (scRNA-seq) datasets, demonstrating its superiority in robustness. The performance of Normi in real scRNA-seq data further reveals its ability to identify the key regulators and crucial biological processes. Yanping Zeng, Yongxin He, Ruiqing Zheng, Min Li 0007 |
Briefings Bioinform. | 4 |
| 2023 | DeepCellEss: cell line-specific essential protein prediction with attention-based interpretable deep learningabstractMOTIVATION: Protein essentiality is usually accepted to be a conditional trait and strongly affected by cellular environments. However, existing computational methods often do not take such characteristics into account, preferring to incorporate all available data and train a general model for all cell lines. In addition, the lack of model interpretability limits further exploration and analysis of essential protein predictions. RESULTS: In this study, we proposed DeepCellEss, a sequence-based interpretable deep learning framework for cell line-specific essential protein predictions. DeepCellEss utilizes a convolutional neural network and bidirectional long short-term memory to learn short- and long-range latent information from protein sequences. Further, a multi-head self-attention mechanism is used to provide residue-level model interpretability. For model construction, we collected extremely large-scale benchmark datasets across 323 cell lines. Extensive computational experiments demonstrate that DeepCellEss yields effective prediction performance for different cell lines and outperforms existing sequence-based methods as well as network-based centrality measures. Finally, we conducted some case studies to illustrate the necessity of considering specific cell lines and the superiority of DeepCellEss. We believe that DeepCellEss can serve as a useful tool for predicting essential proteins across different cell lines. AVAILABILITY AND IMPLEMENTATION: The DeepCellEss web server is available at http://csuligroup.com:8000/DeepCellEss. The source code and data underlying this study can be obtained from https://github.com/CSUBioGroup/DeepCellEss. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Min Zeng 0004, Fuhao Zhang, Fang-Xiang Wu, Min Li 0007 |
Bioinform. | 5 |
| 2023 | GraphscoreDTA: optimized graph neural network for protein-ligand binding affinity predictionabstractMOTIVATION: Computational approaches for identifying the protein-ligand binding affinity can greatly facilitate drug discovery and development. At present, many deep learning-based models are proposed to predict the protein-ligand binding affinity and achieve significant performance improvement. However, protein-ligand binding affinity prediction still has fundamental challenges. One challenge is that the mutual information between proteins and ligands is hard to capture. Another challenge is how to find and highlight the important atoms of the ligands and residues of the proteins. RESULTS: To solve these limitations, we develop a novel graph neural network strategy with the Vina distance optimization terms (GraphscoreDTA) for predicting protein-ligand binding affinity, which takes the combination of graph neural network, bitransport information mechanism and physics-based distance terms into account for the first time. Unlike other methods, GraphscoreDTA can not only effectively capture the protein-ligand pairs' mutual information but also highlight the important atoms of the ligands and residues of the proteins. The results show that GraphscoreDTA significantly outperforms existing methods on multiple test sets. Furthermore, the tests of drug-target selectivity on the cyclin-dependent kinase and the homologous protein families demonstrate that GraphscoreDTA is a reliable tool for protein-ligand binding affinity prediction. AVAILABILITY AND IMPLEMENTATION: The resource codes are available at https://github.com/CSUBioGroup/GraphscoreDTA. Renyi Zhou, Jing Tang 0002, Min Li 0007 |
Bioinform. | 4 |
| 2023 | scNCL: transferring labels from scRNA-seq to scATAC-seq data with neighborhood contrastive regularizationabstractMOTIVATION: scATAC-seq has enabled chromatin accessibility landscape profiling at the single-cell level, providing opportunities for determining cell-type-specific regulation codes. However, high dimension, extreme sparsity, and large scale of scATAC-seq data have posed great challenges to cell-type identification. Thus, there has been a growing interest in leveraging the well-annotated scRNA-seq data to help annotate scATAC-seq data. However, substantial computational obstacles remain to transfer information from scRNA-seq to scATAC-seq, especially for their heterogeneous features. RESULTS: We propose a new transfer learning method, scNCL, which utilizes prior knowledge and contrastive learning to tackle the problem of heterogeneous features. Briefly, scNCL transforms scATAC-seq features into gene activity matrix based on prior knowledge. Since feature transformation can cause information loss, scNCL introduces neighborhood contrastive learning to preserve the neighborhood structure of scATAC-seq cells in raw feature space. To learn transferable latent features, scNCL uses a feature projection loss and an alignment loss to harmonize embeddings between scRNA-seq and scATAC-seq. Experiments on various datasets demonstrated that scNCL not only realizes accurate and robust label transfer for common types, but also achieves reliable detection of novel types. scNCL is also computationally efficient and scalable to million-scale datasets. Moreover, we prove scNCL can help refine cell-type annotations in existing scATAC-seq atlases. AVAILABILITY AND IMPLEMENTATION: The source code and data used in this paper can be found in https://github.com/CSUBioGroup/scNCL-release. Xuhua Yan, Ruiqing Zheng, Jinmiao Chen, Min Li 0007 |
Bioinform. | 4 |
| 2023 | CLAIRE: contrastive learning-based batch correction framework for better balance between batch mixing and preservation of cellular heterogeneityabstractMOTIVATION: Integration of growing single-cell RNA sequencing datasets helps better understand cellular identity and function. The major challenge for integration is removing batch effects while preserving biological heterogeneities. Advances in contrastive learning have inspired several contrastive learning-based batch correction methods. However, existing contrastive-learning-based methods exhibit noticeable ad hoc trade-off between batch mixing and preservation of cellular heterogeneities (mix-heterogeneity trade-off). Therefore, a deliberate mix-heterogeneity trade-off is expected to yield considerable improvements in scRNA-seq dataset integration. RESULTS: We develop a novel contrastive learning-based batch correction framework, CIAIRE, which achieves superior mix-heterogeneity trade-off. The key contributions of CLAIRE are proposal of two complementary strategies: construction strategy and refinement strategy, to improve the appropriateness of positive pairs. Construction strategy dynamically generates positive pairs by augmenting inter-batch mutual nearest neighbors (MNN) with intra-batch k-nearest neighbors (KNN), which improves the coverage of positive pairs for the whole distribution of shared cell types between batches. Refinement strategy aims to automatically reduce the potential false positive pairs from the construction strategy, which resorts to the memory effect of deep neural networks. We demonstrate that CLAIRE possesses superior mix-heterogeneity trade-off over existing contrastive learning-based methods. Benchmark results on six real datasets also show that CLAIRE achieves the best integration performance against eight state-of-the-art methods. Finally, comprehensive experiments are conducted to validate the effectiveness of CLAIRE. AVAILABILITY AND IMPLEMENTATION: The source code and data used in this study can be found in https://github.com/CSUBioGroup/CLAIRE-release. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xuhua Yan, Ruiqing Zheng, Fang-Xiang Wu, Min Li 0007 |
Bioinform. | 4 |
| 2023 | LncLocFormer: a Transformer-based deep learning model for multi-label lncRNA subcellular localization prediction by using localization-specific attention mechanismabstractMOTIVATION: There is mounting evidence that the subcellular localization of lncRNAs can provide valuable insights into their biological functions. In the real world of transcriptomes, lncRNAs are usually localized in multiple subcellular localizations. Furthermore, lncRNAs have specific localization patterns for different subcellular localizations. Although several computational methods have been developed to predict the subcellular localization of lncRNAs, few of them are designed for lncRNAs that have multiple subcellular localizations, and none of them take motif specificity into consideration. RESULTS: In this study, we proposed a novel deep learning model, called LncLocFormer, which uses only lncRNA sequences to predict multi-label lncRNA subcellular localization. LncLocFormer utilizes eight Transformer blocks to model long-range dependencies within the lncRNA sequence and shares information across the lncRNA sequence. To exploit the relationship between different subcellular localizations and find distinct localization patterns for different subcellular localizations, LncLocFormer employs a localization-specific attention mechanism. The results demonstrate that LncLocFormer outperforms existing state-of-the-art predictors on the hold-out test set. Furthermore, we conducted a motif analysis and found LncLocFormer can capture known motifs. Ablation studies confirmed the contribution of the localization-specific attention mechanism in improving the prediction performance. AVAILABILITY AND IMPLEMENTATION: The LncLocFormer web server is available at http://csuligroup.com:9000/LncLocFormer. The source code can be obtained from https://github.com/CSUBioGroup/LncLocFormer. Min Zeng 0004, Yifan Wu 0008, Rui Yin 0002, Chengqian Lu, Junwen Duan, Min Li 0007 |
Bioinform. | 7 |
| 2023 | Retrieve and rerank for automated ICD coding via Contrastive Learning
Kunying Niu, Yifan Wu 0008, Yaohang Li, Min Li 0007 |
J. Biomed. Informatics | 4 |
| 2023 | ViPal: A framework for virulence prediction of influenza viruses with prior viral knowledge using genomic sequences
Rui Yin 0002, Zihan Luo 0001, Pei Zhuang, Min Zeng 0004, Min Li 0007, Zhuoyi Lin, Chee Keong Kwoh 0001 |
J. Biomed. Informatics | 5 |
| 2023 | Guest Editors' Introduction to the Special Section on Bioinformatics Research and Applications
Zhipeng Cai 0001, Min Li 0007, Pavel Skums, Yanjie Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | A Deep Learning Framework for Predicting Protein Functions With Co-Occurrence of GO TermsabstractThe understanding of protein functions is critical to many biological problems such as the development of new drugs and new crops. To reduce the huge gap between the increase of protein sequences and annotations of protein functions, many methods have been proposed to deal with this problem. These methods use Gene Ontology (GO) to classify the functions of proteins and consider one GO term as a class label. However, they ignore the co-occurrence of GO terms that is helpful for protein function prediction. We propose a new deep learning model, named DeepPFP-CO, which uses Graph Convolutional Network (GCN) to explore and capture the co-occurrence of GO terms to improve the protein function prediction performance. In this way, we can further deduce the protein functions by fusing the predicted propensity of the center function and its co-occurrence functions. We use Fmax and AUPR to evaluate the performance of DeepPFP-CO and compare DeepPFP-CO with state-of-the-art methods such as DeepGOPlus and DeepGOA. The computational results show that DeepPFP-CO outperforms DeepGOPlus and other methods. Moreover, we further analyze our model at the protein level. The results have demonstrated that DeepPFP-CO improves the performance of protein function prediction. DeepPFP-CO is available at https://csuligroup.com/DeepPFP/. Min Li 0007, Fuhao Zhang, Min Zeng 0004, Yaohang Li |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2023 | The Impact of Computational Drug Discovery on SocietyabstractGreetings and welcome to the fifth issue of IEEE Transactions on Computational Social Systems (TCSS) for 2023. This edition presents a collection of 55 diverse regular articles that illuminate various facets of the interaction between computer technology and society. Jianxin Wang 0001, Min Li 0007, Edwin Wang, Jing Tang 0002, Bin Hu 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | Fusion-Based Deep Learning Architecture for Detecting Drug-Target Binding Affinity Using Target and Drug Sequence and StructureabstractAccurately predicting drug-target binding affinity plays a vital role in accelerating drug discovery. Many computational approaches have been proposed due to costly and time-consuming of wet laboratory experiments. In the input representation, most methods only focus on the target sequence properties or target structure properties while ignore the overall contribution. Therefore, we develop a novel fusion protocol based on multiscale convolutional neural networks and graph neural networks, named CGraphDTA, to predict drug-target binding affinity using target sequence and structure. Unlike existing methods, CGraphDTA is the first model constructed with target sequence and structure as input. Concretely, the multiscale convolutional neural networks are utilized to extract target and drug presentation from sequence, graph neural networks are employed to extract graph presentation from target and drug molecular structure. We compare CGraphDTA with the state-of-the-art methods, the results show that our model outperforms the current methods on the test sets. Furthermore, we conduct ablation studies, biological interpretation examination and drug selectivity evaluation, all results suggest that CGraphDTA is a useful tool to predict drug-target binding affinity and accelerate drug discovery. Min Li 0007 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | CACO: A Core-Attachment Method With Cross-Species Functional Ortholog Information to Detect Human Protein ComplexesabstractProtein complexes play an essential role in living cells. Detecting protein complexes is crucial to understand protein functions and treat complex diseases. Due to high time and resource consumption of experiment approaches, many computational approaches have been proposed to detect protein complexes. However, most of them are only based on protein-protein interaction (PPI) networks, which heavily suffer from the noise in PPI networks. Therefore, we propose a novel core-attachment method, named CACO, to detect human protein complexes, by integrating the functional information from other species via protein ortholog relations. First, CACO constructs a cross-species ortholog relation matrix and transfers GO terms from other species as a reference to evaluate the confidence of PPIs. Then, a PPI filter strategy is adopted to clean the PPI network and thus a weighted clean PPI network is constructed. Finally, a new effective core-attachment algorithm is proposed to detect protein complexes from the weighted PPI network. Compared to other thirteen state-of-the-art methods, CACO outperforms all of them in terms of F-measure and Composite Score, showing that integrating ortholog information and the proposed core-attachment algorithm are effective in detecting protein complexes. Wenkang Wang, Xiangmao Meng, Ju Xiang, Yunyan Shuai, Hayat Dino Bedru, Min Li 0007 |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | LDAGSO: Predicting 1ncRNA-Disease Associations from Graph Sequences and Disease Ontology via Deep Learning techniquesabstractRecent studies have confirmed the significant effects of long non-coding RNAs (IncRNAs) in understanding the mechanism of diseases. Because of the relatively small number of validated associations between IncRNAs and diseases, and previous computational methods have limited performance without capturing important features of sequences and ontology information, we developed LDAGSO, a novel deep learning framework to predict IncRNA and disease associations from IncRNA sequences and disease ontology. For IncRNA sequences, we converted them into graph structure based on k-mer technique and de Bruijn graph, and captured high-level features of the graph using graph convolutional networks. For diseases, we extracted ontology term paths from the disease ontology tree, and treated them as sentences to obtain their feature representation using Bidirectional Encoder Representations from Transformers (BERT) technique. Finally, these two kinds of features were fed into a fully connected layer to perform the task of association prediction between 1ncRNAs and diseases. According to the results, our approach provides state-of-the-art results when evaluated by leave-one-out cross-validation. Norah Saeed Awn, Baoying Zhao, Min Zeng 0004, Min Li 0007 |
BIBM | 5 |
| 2022 | BayesImpute: a Bayesian imputation method for single-cell RNA-seq dataabstractSingle-cell RNA-sequencing (scRNA-seq) data suffer from a large number of zeros. Such dropout events hinder the downstream data analyses. We propose BayesImpute, a statistical algorithm to impute dropouts in the scRNA-seq data. BayesImpute first identifies likely dropouts based on expression rate and coefficient of variation of genes within cell subpopulation, and then constructs the posterior distribution for each gene and utilizes the posterior mean to impute dropout values. With several simulated and real scRNA-seq datasets, we demonstrate that BayesImpute is capable of effectively identifying dropouts. In addition, BayesImpute successfully recovers the true expression levels of missing values, improves the clustering and visualization of cell subpopulations, and enhances the identification of differential expression genes. We also show that BayesImpute is scalable and fast with minimal memory usage compared with other statistical-based imputation methods. Ruiqing Zheng, Luyi Tian, Fang-Xiang Wu, Min Li 0007 |
BIBM | 5 |
| 2022 | MERAS: A method for entity recognition, alignment, and structuration from Chinese Electronic Medical RecordsabstractThe description of entity information in Chinese Electronic Medical Records (EMRs) is not always isolated, and it often has rich attribute constraints, especially the symptom entities. Moreover, these entity descriptions are unstructured and may be colloquial, non-terminological, and misspelled. This paper proposed a new method, which is called MERAS for entity recognition, alignment, and structuration from Chinese EMRs. Different from the traditional entity extraction tasks, a series of complex entities with attribute constraints are annotated and recognized by MERAS, and then the entities are further converted to the standardized and structured entity sequences in JSON format based on the end-to-end translation idea. BiLSTM-CRF is applied to recognize the complex entities, and Transformer is applied to realize the entity alignment and structuration in this paper. With the manually annotated data set as the reference, 9 types of entities are extracted from the EMRs, and the corresponding corpus-feature-based enhanced strategies are proposed to support the model training. The experimental results show that MERAS can effectively extract standardized and structured entities from Chinese EMRs, and it can have great potential in practical applications. (Abstract) Yifan Wu 0008, Min Li 0007 |
BIBM | 3 |
| 2022 | A single cell potency inference method based on the local cell-specific network entropyabstractAt present, some methods have been proposed to solve the problem from the perspective of the chaos degree in gene functions or interaction network. However, errors in differentiation potency estimates arise if the scRNA-seq profile and underlying interaction network are disturbed by technique-induced or biological-induced noise. Thus, we proposed SPIDE, a single cell potency inference method based on local cell-specific network entropy. SPIDE constructs the weighted cell-specific network for each cell to preserve the heterogeneity of PPI network during differentiation, then estimates the entropy based on each network. The results show that SPIDE reveals better decreasing trends of cells’ differentiation potency than other state-of-the-art methods on most datasets. To conclude, our study provides a universal framework for cell entropy estimation with higher prediction accuracy and universal applicability. Ruiqing Zheng, Edwin Wang, Min Li 0007 |
BIBM | 5 |
| 2022 | Coded Residual Transform for Generalizable Deep Metric LearningabstractA fundamental challenge in deep metric learning is the generalization capability of the feature embedding network model since the embedding network learned on training classes need to be evaluated on new test classes. To address this challenge, in this paper, we introduce a new method called coded residual transform (CRT) for deep metric learning to significantly improve its generalization capability. Specifically, we learn a set of diversified prototype features, project the feature map onto each prototype, and then encode its features using their projection residuals weighted by their correlation coefficients with each prototype. The proposed CRT method has the following two unique characteristics. First, it represents and encodes the feature map from a set of complimentary perspectives based on projections onto diversified prototypes. Second, unlike existing transformer-based feature representation approaches which encode the original values of features based on global correlation analysis, the proposed coded residual transform encodes the relative differences between the original features and their projected prototypes. Embedding space density and spectral decay analysis show that this multi perspective projection onto diversified prototypes and coded residual representation are able to achieve significantly improved generalization capability in metric learning. Finally, to further enhance the generalization performance, we propose to enforce the consistency on their feature similarity matrices between coded residual transforms with different sizes of projection prototypes and embedding dimensions. Our extensive experimental results and ablation studies demonstrate that the proposed CRT method outperform the state-of-the-art deep metric learning methods by large margins and improving upon the current best method by up to 4.28% on the CUB dataset. Shichao Kan, Yixiong Liang, Min Li 0007, Yi-Gang Cen, Jianxin Wang 0001, Zhihai He |
NeurIPS | 3 |
| 2022 | HyMM: hybrid method for disease-gene prediction by integrating multiscale module structureabstractMOTIVATION: Identifying disease-related genes is an important issue in computational biology. Module structure widely exists in biomolecule networks, and complex diseases are usually thought to be caused by perturbations of local neighborhoods in the networks, which can provide useful insights for the study of disease-related genes. However, the mining and effective utilization of the module structure is still challenging in such issues as a disease gene prediction. RESULTS: We propose a hybrid disease-gene prediction method integrating multiscale module structure (HyMM), which can utilize multiscale information from local to global structure to more effectively predict disease-related genes. HyMM extracts module partitions from local to global scales by multiscale modularity optimization with exponential sampling, and estimates the disease relatedness of genes in partitions by the abundance of disease-related genes within modules. Then, a probabilistic model for integration of gene rankings is designed in order to integrate multiple predictions derived from multiscale module partitions and network propagation, and a parameter estimation strategy based on functional information is proposed to further enhance HyMM's predictive power. By a series of experiments, we reveal the importance of module partitions at different scales, and verify the stable and good performance of HyMM compared with eight other state-of-the-arts and its further performance improvement derived from the parameter estimation. CONCLUSIONS: The results confirm that HyMM is an effective framework for integrating multiscale module structure to enhance the ability to predict disease-related genes, which may provide useful insights for the study of the multiscale module structure and its application in such issues as a disease-gene prediction. Ju Xiang, Xiangmao Meng, Yi-chao Zhao, Fang-Xiang Wu, Min Li 0007 |
Briefings Bioinform. | 5 |
| 2022 | Biomedical data, computational methods and tools for evaluating disease-disease associationsabstractIn recent decades, exploring potential relationships between diseases has been an active research field. With the rapid accumulation of disease-related biomedical data, a lot of computational methods and tools/platforms have been developed to reveal intrinsic relationship between diseases, which can provide useful insights to the study of complex diseases, e.g. understanding molecular mechanisms of diseases and discovering new treatment of diseases. Human complex diseases involve both external phenotypic abnormalities and complex internal molecular mechanisms in organisms. Computational methods with different types of biomedical data from phenotype to genotype can evaluate disease-disease associations at different levels, providing a comprehensive perspective for understanding diseases. In this review, available biomedical data and databases for evaluating disease-disease associations are first summarized. Then, existing computational methods for disease-disease associations are reviewed and classified into five groups in terms of the usages of biomedical data, including disease semantic-based, phenotype-based, function-based, representation learning-based and text mining-based methods. Further, we summarize software tools/platforms for computation and analysis of disease-disease associations. Finally, we give a discussion and summary on the research of disease-disease associations. This review provides a systematic overview for current disease association research, which could promote the development and applications of computational methods and tools/platforms for disease-disease associations. Ju Xiang, Jiashuai Zhang, Yi-chao Zhao, Fang-Xiang Wu, Min Li 0007 |
Briefings Bioinform. | 5 |
| 2022 | GLOBE: a contrastive learning-based framework for integrating single-cell transcriptome datasetsabstractIntegration of single-cell transcriptome datasets from multiple sources plays an important role in investigating complex biological systems. The key to integration of transcriptome datasets is batch effect removal. Recent methods attempt to apply a contrastive learning strategy to correct batch effects. Despite their encouraging performance, the optimal contrastive learning framework for batch effect removal is still under exploration. We develop an improved contrastive learning-based batch correction framework, GLOBE. GLOBE defines adaptive translation transformations for each cell to guarantee the stability of approximating batch effects. To enhance the consistency of representations alignment, GLOBE utilizes a loss function that is both hardness-aware and consistency-aware to learn batch effect-invariant representations. Moreover, GLOBE computes batch-corrected gene matrix in a transparent approach to support diverse downstream analysis. Benchmarking results on a wide spectrum of datasets show that GLOBE outperforms other state-of-the-art methods in terms of robust batch mixing and superior conservation of biological signals. We further apply GLOBE to integrate two developing mouse neocortex datasets and show GLOBE succeeds in removing batch effects while preserving the contiguous structure of cells in raw data. Finally, a comprehensive study is conducted to validate the effectiveness of GLOBE. Xuhua Yan, Ruiqing Zheng, Min Li 0007 |
Briefings Bioinform. | 3 |
| 2022 | A framework for predicting variable-length epitopes of human-adapted viruses using machine learning methodsabstractThe coronavirus disease 2019 pandemic has alerted people of the threat caused by viruses. Vaccine is the most effective way to prevent the disease from spreading. The interaction between antibodies and antigens will clear the infectious organisms from the host. Identifying B-cell epitopes is critical in vaccine design, development of disease diagnostics and antibody production. However, traditional experimental methods to determine epitopes are time-consuming and expensive, and the predictive performance using the existing in silico methods is not satisfactory. This paper develops a general framework to predict variable-length linear B-cell epitopes specific for human-adapted viruses with machine learning approaches based on Protvec representation of peptides and physicochemical properties of amino acids. QR decomposition is incorporated during the embedding process that enables our models to handle variable-length sequences. Experimental results on large immune epitope datasets validate that our proposed model's performance is superior to the state-of-the-art methods in terms of AUROC (0.827) and AUPR (0.831) on the testing set. Moreover, sequence analysis also provides the results of the viral category for the corresponding predicted epitopes with high precision. Therefore, this framework is shown to reliably identify linear B-cell epitopes of human-adapted viruses given protein sequences and could provide assistance for potential future pandemics and epidemics. Rui Yin 0002, Xianghe Zhu, Min Zeng 0004, Min Li 0007, Chee Keong Kwoh 0001 |
Briefings Bioinform. | 5 |
| 2022 | DeepLncLoc: a deep learning framework for long non-coding RNA subcellular localization prediction based on subsequence embeddingabstractLong non-coding RNAs (lncRNAs) are a class of RNA molecules with more than 200 nucleotides. A growing amount of evidence reveals that subcellular localization of lncRNAs can provide valuable insights into their biological functions. Existing computational methods for predicting lncRNA subcellular localization use k-mer features to encode lncRNA sequences. However, the sequence order information is lost by using only k-mer features. We proposed a deep learning framework, DeepLncLoc, to predict lncRNA subcellular localization. In DeepLncLoc, we introduced a new subsequence embedding method that keeps the order information of lncRNA sequences. The subsequence embedding method first divides a sequence into some consecutive subsequences and then extracts the patterns of each subsequence, last combines these patterns to obtain a complete representation of the lncRNA sequence. After that, a text convolutional neural network is employed to learn high-level features and perform the prediction task. Compared with traditional machine learning models, popular representation methods and existing predictors, DeepLncLoc achieved better performance, which shows that DeepLncLoc could effectively predict lncRNA subcellular localization. Our study not only presented a novel computational model for predicting lncRNA subcellular localization but also introduced a new subsequence embedding method which is expected to be applied in other sequence-based prediction tasks. The DeepLncLoc web server is freely accessible at http://bioinformatics.csu.edu.cn/DeepLncLoc/, and source code and datasets can be downloaded from https://github.com/CSUBioGroup/DeepLncLoc. Min Zeng 0004, Yifan Wu 0008, Chengqian Lu, Fuhao Zhang, Fang-Xiang Wu, Min Li 0007 |
Briefings Bioinform. | 6 |
| 2022 | DeepDISOBind: accurate prediction of RNA-, DNA- and protein-binding intrinsically disordered residues with deep multi-task learningabstractProteins with intrinsically disordered regions (IDRs) are common among eukaryotes. Many IDRs interact with nucleic acids and proteins. Annotation of these interactions is supported by computational predictors, but to date, only one tool that predicts interactions with nucleic acids was released, and recent assessments demonstrate that current predictors offer modest levels of accuracy. We have developed DeepDISOBind, an innovative deep multi-task architecture that accurately predicts deoxyribonucleic acid (DNA)-, ribonucleic acid (RNA)- and protein-binding IDRs from protein sequences. DeepDISOBind relies on an information-rich sequence profile that is processed by an innovative multi-task deep neural network, where subsequent layers are gradually specialized to predict interactions with specific partner types. The common input layer links to a layer that differentiates protein- and nucleic acid-binding, which further links to layers that discriminate between DNA and RNA interactions. Empirical tests show that this multi-task design provides statistically significant gains in predictive quality across the three partner types when compared to a single-task design and a representative selection of the existing methods that cover both disorder- and structure-trained tools. Analysis of the predictions on the human proteome reveals that DeepDISOBind predictions can be encoded into protein-level propensities that accurately predict DNA- and RNA-binding proteins and protein hubs. DeepDISOBind is available at https://www.csuligroup.com/DeepDISOBind/. Fuhao Zhang, Bi Zhao, Min Li 0007, Lukasz A. Kurgan |
Briefings Bioinform. | 4 |
| 2022 | SEPA: signaling entropy-based algorithm to evaluate personalized pathway activation for survival analysis on pan-cancer dataabstractMOTIVATION: Biomarkers with prognostic ability and biological interpretability can be used to support decision-making in the survival analysis. Genes usually form functional modules to play synergistic roles, such as pathways. Predicting significant features from the functional level can effectively reduce the adverse effects of heterogeneity and obtain more reproducible and interpretable biomarkers. Personalized pathway activation inference can quantify the dysregulation of essential pathways involved in the initiation and progression of cancers, and can contribute to the development of personalized medical treatments. RESULTS: In this study, we propose a novel method to evaluate personalized pathway activation based on signaling entropy for survival analysis (SEPA), which is a new attempt to introduce the information-theoretic entropy in generating pathway representation for each patient. SEPA effectively integrates pathway-level information into gene expression data, converting the high-dimensional gene expression data into the low-dimensional biological pathway activation scores. SEPA shows its classification power on the prognostic pan-cancer genomic data, and the potential pathway markers identified based on SEPA have statistical significance in the discrimination of high- and low-risk cohorts and are likely to be associated with the initiation and progress of cancers. The results show that SEPA scores can be used as an indicator to precisely distinguish cancer patients with different clinical outcomes, and identify important pathway features with strong discriminative power and biological interpretability. AVAILABILITY AND IMPLEMENTATION: The MATLAB-package for SEPA is freely available from https://github.com/xingyili/SEPA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xingyi Li 0003, Min Li 0007, Ju Xiang, Zhelin Zhao, Xuequn Shang 0001 |
Bioinform. | 2 |
| 2022 | BACPI: a bi-directional attention neural network for compound-protein interaction and binding affinity predictionabstractMOTIVATION: The identification of compound-protein interactions (CPIs) is an essential step in the process of drug discovery. The experimental determination of CPIs is known for a large amount of funds and time it consumes. Computational model has therefore become a promising and efficient alternative for predicting novel interactions between compounds and proteins on a large scale. Most supervised machine learning prediction models are approached as a binary classification problem, which aim to predict whether there is an interaction between the compound and the protein or not. However, CPI is not a simple binary on-off relationship, but a continuous value reflects how tightly the compound binds to a particular target protein, also called binding affinity. RESULTS: In this study, we propose an end-to-end neural network model, called BACPI, to predict CPI and binding affinity. We employ graph attention network and convolutional neural network (CNN) to learn the representations of compounds and proteins and develop a bi-directional attention neural network model to integrate the representations. To evaluate the performance of BACPI, we use three CPI datasets and four binding affinity datasets in our experiments. The results show that, when predicting CPIs, BACPI significantly outperforms other available machine learning methods on both balanced and unbalanced datasets. This suggests that the end-to-end neural network model that predicts CPIs directly from low-level representations is more robust than traditional machine learning-based methods. And when predicting binding affinities, BACPI achieves higher performance on large datasets compared to other state-of-the-art deep learning methods. This comparison result suggests that the proposed method with bi-directional attention neural network can capture the important regions of compounds and proteins for binding affinity prediction. AVAILABILITY AND IMPLEMENTATION: Data and source codes are available at https://github.com/CSUBioGroup/BACPI. Min Li 0007, Zhangli Lu, Yifan Wu 0008, Yaohang Li |
Bioinform. | 1 |
| 2022 | BridgeDPI: a novel Graph Neural Network for predicting drug-protein interactionsabstractMOTIVATION: Exploring drug-protein interactions (DPIs) provides a rapid and precise approach to assist in laboratory experiments for discovering new drugs. Network-based methods usually utilize a drug-protein association network and predict DPIs by the information of its associated proteins or drugs, called 'guilt-by-association' principle. However, the 'guilt-by-association' principle is not always true because sometimes similar proteins cannot interact with similar drugs. Recently, learning-based methods learn molecule properties underlying DPIs by utilizing existing databases of characterized interactions but neglect the network-level information. RESULTS: We propose a novel method, namely BridgeDPI. We devise a class of virtual nodes to bridge the gap between drugs and proteins and construct a learnable drug-protein association network. The network is optimized based on the supervised signals from the downstream task-the DPI prediction. Through information passing on this drug-protein association network, a Graph Neural Network can capture the network-level information among diverse drugs and proteins. By combining the network-level information and the learning-based method, BridgeDPI achieves significant improvement in three real-world DPI datasets. Moreover, the case study further verifies the effectiveness and reliability of BridgeDPI. AVAILABILITY AND IMPLEMENTATION: The source code of BridgeDPI can be accessed at https://github.com/SenseTime-Knowledge-Mining/BridgeDPI. The source data used in this study is available on the https://github.com/IBM/InterpretableDTIP (for the BindingDB dataset), https://github.com/masashitsubaki/CPI_prediction (for the C.ELEGANS and HUMAN) datasets, http://dude.docking.org/ (for the DUD-E dataset), repectively. Yifan Wu 0008, Min Zeng 0004, Jie Zhang 0122, Min Li 0007 |
Bioinform. | 5 |
| 2022 | Guest editorial: Deep neural networks for precision medicine
Fang-Xiang Wu, Min Li 0007, Lukasz A. Kurgan, Luis Rueda 0001 |
Neurocomputing | 2 |
| 2022 | KAICD: A knowledge attention-based deep learning framework for automatic ICD coding
Yifan Wu 0008, Min Zeng 0004, Zhihui Fei, Fang-Xiang Wu, Min Li 0007 |
Neurocomputing | 6 |
| 2022 | Guest Editors' Introduction to the Special Section on Bioinformatics Research and ApplicationsabstractThe papers in this special section were presented at the 15th International Symposium on Bioinformatics Research and Applications (ISBRA 2019), which was held at Technical University of Catalonia, Barcelona, Spain on June 3-6, 2019. Zhipeng Cai 0001, Min Li 0007, Pavel Skums |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | In Silico Prediction of New Mutations That Can Improve the Binding Abilities Between 2019-nCoV Coronavirus and Human ACE2abstractThe Coronavirus Disease 2019 (COVID-19) has become an international public health emergency, posing a serious threat to human health and safety around the world. The 2019-nCoV coronavirus spike protein was confirmed to be highly susceptible to various mutations, which can trigger apparent changes of virus transmission capacity and the pathogenic mechanism. In this article, the binding interface was obtained by analyzing the interaction modes between 2019-nCoV coronavirus and the human ACE2. Based on the "SIFT server" and the "bubble" identification mechanism, 9 amino acid sites were selected as potential mutation-sites from the 2019-nCoV-S1-ACE2 binding interface. Subsequently, a total number of 171 mutant systems for 9 mutation-sites were optimized for binding-pattern comparsion analysis, and 14 mutations that may improve the binding capacity of 2019-nCoV-S1 to ACE2 were selected. The Molecular Dynamic Simulations were conducted to calculate the binding free energies of all the 14 mutant systems. Finally, we found that most of the 14 mutations on the 2019-nCoV-S1 protein could enhance the binding ability between 2019-nCoV coronavirus and human ACE2. Among which, the binding capacities for G446R, Y449R and F486Y mutations could be increased by 20 percent, and that for S494R mutant increased even by 38.98 percent. We hope this research could provide significant help for the future epidemic detection, drug and vaccine development. Senbiao Fang, Ruoqian Zheng, Chuqi Lei, Jianxin Wang 0001, Renyi Zhou, Min Li 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2022 | NIMCE: A Gene Regulatory Network Inference Approach Based on Multi Time Delays Causal EntropyabstractGene regulatory networks (GRNs)are involved in various biological processes, such as cell cycle, differentiation and apoptosis. The existing large amount of expression data, especially the time-series expression data, provide a chance to infer GRNs by computational methods. These data can reveal the dynamics of gene expression and imply the regulatory relationships among genes. However, identify the indirect regulatory links is still a big challenge as most studies treat time points as independent observations, while ignoring the influences of time delays. In this study, we propose a GRN inference method based on information-theory measure, called NIMCE. NIMCE incorporates the transfer entropy to measure the regulatory links between each pair of genes, then applies the causation entropy to filter indirect relationships. In addition, NIMCE applies multi time delays to identify indirect regulatory relationships from candidate genes. Experiments on simulated and colorectal cancer data show NIMCE outperforms than other competing methods. All data and codes used in this study are publicly available at https://github.com/CSUBioGroup/NIMCE. Haonan Feng, Ruiqing Zheng, Jianxin Wang 0001, Fang-Xiang Wu, Min Li 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | A Dual Ranking Algorithm Based on the Multiplex Network for Heterogeneous Complex Disease AnalysisabstractIdentifying biomarkers of heterogeneous complex diseases has always been one of the focuses in medical research. In previous studies, the powerful network propagation methods have been applied to finding marker genes related to specific diseases, but existing methods are mostly based on a single network, which may be greatly affected by the incompleteness of the network and the ignorance of a large amount of information about physical and functional interactions between biological components. Other methods that directly integrate multiple types of interactions into an aggregate network have the risks that different types of data may conflict with each other and the characteristics and topologies of each individual network are lost. Meanwhile, biomarkers used in clinical trials should have the characteristics of small quantity and strong discriminate ability. In this study, we developed a multiplex network-based dual ranking framework (DualRank) for heterogeneous complex disease analysis. We applied the proposed method to heterogeneous complex diseases for diagnosis, prognosis, and classification. The results showed that DualRank outperformed competing methods and could identify biomarkers with the small quantity, great prediction performance (average AUC = 0.818) and biological interpretability. Xingyi Li 0003, Ju Xiang, Fang-Xiang Wu, Min Li 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | Accurate Prediction of Human Essential Proteins Using Ensemble Deep LearningabstractEssential proteins are considered the foundation of life as they are indispensable for the survival of living organisms. Computational methods for essential protein discovery provide a fast way to identify essential proteins. But most of them heavily rely on various biological information, especially protein-protein interaction networks, which limits their practical applications. With the rapid development of high-throughput sequencing technology, sequencing data has become the most accessible biological data. However, using only protein sequence information to predict essential proteins has limited accuracy. In this paper, we propose EP-EDL, an ensemble deep learning model using only protein sequence information to predict human essential proteins. EP-EDL integrates multiple classifiers to alleviate the class imbalance problem and to improve prediction accuracy and robustness. In each base classifier, we employ multi-scale text convolutional neural networks to extract useful features from protein sequence feature matrices with evolutionary information. Our computational results show that EP-EDL outperforms the state-of-the-art sequence-based methods. Furthermore, EP-EDL provides a more practical and flexible way for biologists to accurately predict essential proteins. The source code and datasets can be downloaded from https://github.com/CSUBioGroup/EP-EDL. Min Zeng 0004, Yifan Wu 0008, Yaohang Li, Min Li 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | DPCMNE: Detecting Protein Complexes From Protein-Protein Interaction Networks Via Multi-Level Network EmbeddingabstractBiological functions of a cell are typically carried out through protein complexes. The detection of protein complexes is therefore of great significance for understanding the cellular organizations and protein functions. In the past decades, many computational methods have been proposed to detect protein complexes. However, most of the existing methods just search the local topological information to mine dense subgraphs as protein complexes, ignoring the global topological information. To tackle this issue, we propose the DPCMNE method to detect protein complexes via multi-level network embedding. It can preserve both the local and global topological information of biological networks. First, DPCMNE employs a hierarchical compressing strategy to recursively compress the input protein-protein interaction (PPI) network into multi-level smaller PPI networks. Then, a network embedding method is applied on these smaller PPI networks to learn protein embeddings of different levels of granularity. The embeddings learned from all the compressed PPI networks are concatenated to represent the final protein embeddings of the original input PPI network. Finally, a core-attachment based strategy is adopted to detect protein complexes in the weighted PPI network constructed by the pairwise similarity of protein embeddings. To assess the efficiency of our proposed method, DPCMNE is compared with other eight clustering algorithms on two yeast datasets. The experimental results show that the performance of DPCMNE outperforms those state-of-the-art complex detection methods in terms of F1 and F1+Acc. Furthermore, the results of functional enrichment analysis indicate that protein complexes detected by DPCMNE are more biologically significant in terms of P-score. Xiangmao Meng, Ju Xiang, Ruiqing Zheng, Fang-Xiang Wu, Min Li 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | Partner-Specific Drug Repositioning Approach Based on Graph Convolutional NetworkabstractDrug repositioning identifies novel therapeutic potentials for existing drugs and is considered an attractive approach due to the opportunity for reduced development timelines and overall costs. Prior computational methods usually learned a drug's representation from an entire graph of drug-disease associations. Therefore, the representation of learned drugs representation are static and agnostic to various diseases. However, for different diseases, a drug's mechanism of actions (MoAs) are different. The relevant context information should be differentiated for the same drug to target different diseases. Computational methods are thus required to learn different representations corresponding to different drug-disease associations for the given drug. In view of this, we propose an end-to-end partner-specific drug repositioning approach based on graph convolutional network, named PSGCN. PSGCN firstly extracts specific context information around drug-disease pairs from an entire graph of drug-disease associations. Then, it implements a graph convolutional network on the extracted graph to learn partner-specific graph representation. As the different layers of graph convolutional network contribute differently to the representation of the partner-specific graph, we design a layer self-attention mechanism to capture multi-scale layer information. Finally, PSGCN utilizes sortpool strategy to obtain the partner-specific graph embedding and formulates a drug-disease association prediction as a graph classification task. A fully-connected module is established to classify the partner-specific graph representations. The experiments on three benchmark datasets prove that the representation learning of partner-specific graph can lead to superior performances over state-of-the-art methods. In particular, case studies on small cell lung cancer and breast carcinoma confirmed that PSGCN is able to retrieve more actual drug-disease associations in the top prediction results. Moreover, in comparison with other static approaches, PSGCN can partly distinguish the different disease context information for the given drug. Xinliang Sun, Bei Wang 0004, Jie Zhang 0122, Min Li 0007 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | A Pseudo Label-Wise Attention Network for Automatic ICD CodingabstractAutomatic International Classification of Diseases (ICD) coding is defined as a kind of text multi-label classification problem, which is difficult because the number of labels is very large and the distribution of labels is unbalanced. The label-wise attention mechanism is widely used in automatic ICD coding because it can assign weights to every word in full Electronic Medical Records (EMR) for different ICD codes. However, the label-wise attention mechanism is redundant and costly in computing. In this paper, we propose a pseudo label-wise attention mechanism to tackle the problem. Instead of computing different attention modes for different ICD codes, the pseudo label-wise attention mechanism automatically merges similar ICD codes and computes only one attention mode for the similar ICD codes, which greatly compresses the number of attention modes and improves the predicted accuracy. In addition, we apply a more convenient and effective way to obtain the ICD vectors, and thus our model can predict new ICD codes by calculating the similarities between EMR vectors and ICD vectors. Our model demonstrates effectiveness in extensive computational experiments. On the public MIMIC-III dataset and private Xiangya dataset, our model achieves the best performance on micro F1 (0.583 and 0.806), micro AUC (0.986 and 0.994), P@8 (0.756 and 0.413), and costs much smaller GPU memory (about 26.1% of the models with label-wise attention). Furthermore, we verify the ability of our model in predicting new ICD codes. The interpretablility analysis and case study show the effectiveness and reliability of the patterns obtained by the pseudo label-wise attention mechanism. Yifan Wu 0008, Min Zeng 0004, Yaohang Li, Min Li 0007 |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | A Hybrid Pooling Based Deep Learning Framework For Automated ICD CodingabstractICD coding is the practice of allocating diagnostic and procedure codes to the clinical records following the International Classification of Diseases. The manual allocation of ICD codes to clinical notes is a very tedious job which has become costly, time-consuming and error-prone. Up to now, various methods for automated ICD coding have been devised, ranging from machine learning to deep learning methodologies. Earlier cutting-edge models relied on CNN’s with one or several fixed window widths. However, the length and dependency of text fragments linked to ICD labels in clinical literature differ considerably, posing a difficulty in determining the optimal window size. Apart from that, in prior models that utilized CNN architecture, the average features of the clinical notes have been ignored, resulting in a lack of preparation of rich features for the classifier. In this research, we present a Deep Recurrent Convolutional Neural Network with Hybrid Pooling (DRCNN-HP), which addresses all of the above mentioned issues. DRCNN-HP takes into account the different lengths as well as the dependency of the ICD code-related text chunks. Furthermore, we applied a powerful hybrid pooling layer in DRCNN-HP to capture the rich feature representation (i.e., by concatenating maximum and average features of text) for the classifier, which resulted in giving state-of-the-art results on the MIMIC-III top 50 dataset as compared to the prior competitive models. Sajida Raz Bhutto, Yifan Wu 0008, Akhtar Hussain 0001, Min Li 0007 |
BIBM | 5 |
| 2021 | DeepCI: a deep learning based clustering method for single cell RNA-seq dataabstractSingle cell RNA sequencing enables researchers to analyze cellular heterogeneity at high resolution. In the cellular heterogeneity analysis, unsupervised clustering has been a common and powerful way to identify cell types. Nevertheless, the high dropout rate and high dimension of scRNA-seq data make it still a challenging task. In this study, we proposed DeepCI, a deep neural network based single cell clustering method, which simultaneously accomplishes low-dimensional representation learning and clustering with implicit imputation of scRNA-seq data. Tested on real datasets, DeepCI obtained overall better clustering and visualization performance than several state-of-the-art approaches. Zhenlan Liang, Ruiqing Zheng, Xuhua Yan, Min Li 0007 |
BIBM | 5 |
| 2021 | nPTAS: A Novel Platform for Text Annotation and ServiceabstractNatural Language Processing (NLP) is a critical research area in artificial intelligence, which has spawned a variety of useful applications. However, creating an online NLP service from scratch is still challenging for non-expert users, which has to go through complex steps such as corpus annotation, model training and deployment. Existing tools mainly focus on corpus annotation, which rarely support collaborative annotation or online model training and deployment. To meet the requirement of rapid deployment of NLP services, we develop a novel full process platform nPTAS, which supports collaborative annotation, online model training and model deploying with RESTful interfaces. To validate the effectiveness of the platform, we build a clinical entities recognition service from Chinese Electronic Medical Records on it. The experimental results show that nPTAS improves the efficiency and quality of corpus annotation and greatly reduces the effort to build online NLP services. The platform is available at http://nptas.c2cloud.cn. Junwen Duan, Min Li 0007 |
BIBM | 4 |
| 2021 | Improving human essential protein prediction using only protein sequences via ensemble learningabstractAccurate prediction of essential proteins by using computational methods can effectively reduce the cost of wet-lab experiments. Existing computational methods usually rely on constructed protein-protein interaction (PPI) networks with different kinds of biological data. However, high-quality PPI networks and other biological data are not available for all proteins. Thus, it is very necessary and valuable to develop accurate methods for fast and effective prediction of essential proteins by using only protein sequences. We propose EPGBDT, a machine learning ensemble model, to improve the performance of essential protein prediction by using only protein sequences. EP-GBDT has an ensemble structure that combines multiple Gradient Boosting Decision Tree (GBDT) base classifiers. In addition, to reduce the effects of imbalanced dataset, EP-GBDT uses a sampling technique. The results show that EP-GBDT outperforms state-of-the-art sequence-based methods and network-based centrality measures. The source code and datasets can be downloaded from https://github.com/CSUBioGroup/EP-GBDT. Min Zeng 0004, Yifan Wu 0008, Fang-Xiang Wu, Min Li 0007 |
BIBM | 6 |
| 2021 | MKG: a mutual information based method to infer single cell gene regulatory networkabstractGRN is the core of all living organisms that can explain how genes and their products interact at different levels. To infer the potential GRNs from gene expression data remains a great challenge in bioinformatics. Recently, with the development of single cell RNA sequencing technology, inferring cell specific GRNs involving in cell differentiation or cell function becomes a hot topic. Although there are some methods proposed to accomplish the task, it is still less than ideal because of the additional noises of pseudo time and high dropouts in datasets. Therefore, we propose a time-delayed mutual information based method, named MKG. MKG handles the above problems by partitioning the whole trajectory into several time windows then takes the average expression value of cells in each window as the representative cell. To further reduce the impact of dropouts, the mixed KSG estimator is applied to quantify the high-order time-delayed mutual information between pairs of genes. According to the experimental results on multiple simulated and real datasets, MKG has better performance and stability compared with other state-of-the-art algorithms. Yanping Zeng, Xuhua Yan, Zhenlan Liang, Ruiqing Zheng, Min Li 0007 |
BIBM | 5 |
| 2021 | Overlapping Protein Complexes Detection Based on Multi-level Topological Similarities
Wenkang Wang, Xiangmao Meng, Ju Xiang, Min Li 0007 |
ISBRA | 4 |
| 2021 | Key residues influencing binding affinities of 2019-nCoV with ACE2 in different speciesabstractThe Novel Coronavirus Disease 2019 (COVID-19) has become an international public health emergency, which poses the most serious threat to the human health around the world. Accumulating evidences have shown that the new coronavirus could not only infect human beings, but also can infect other species which might result in the cross-species infections. In this research, 1056 ACE2 protein sequences are collected from the NCBI database, and 173 species with >60% sequence identity compared with that of human beings are selected for further analysis. We find 14 polar residues forming the binding interface of ACE2/2019-nCoV-Spike complex play an important role in maintaining protein-protein stability. Among them, 8 polar residues at the same positions with that of human ACE2 are highly conserved, which ensure its basic binding affinity with the novel coronavirus. 5 of other 6 unconserved polar residues (positions at human ACE2: Q24, D30, K31, H34 and E35) are proved to have an effect on the binding patterns among species. We select 21 species keeping close contacts with human beings, construct their ACE2 three-dimensional structures by Homology Modeling method and calculate the binding free energies of their ACE2/2019-nCoV-Spike complexes. We find the ACE2 from all the 21 species possess the capabilities to bind with the novel coronavirus. Compared with the human beings, 8 species (cow, deer, cynomys, chimpanzee, monkey, sheep, dolphin and whale) present almost the same binding abilities, and 3 species (bat, pig and dog) show significant improvements in binding affinities. We hope this research could provide significant help for the future epidemic detection, drug and vaccine development and even the global eco-system protections. Senbiao Fang, Ruoqian Zheng, Chuqi Lei, Jianxin Wang 0001, Ruiqing Zheng, Min Li 0007 |
Briefings Bioinform. | 6 |
| 2021 | Biomedical data and computational models for drug repositioning: a comprehensive reviewabstractDrug repositioning can drastically decrease the cost and duration taken by traditional drug research and development while avoiding the occurrence of unforeseen adverse events. With the rapid advancement of high-throughput technologies and the explosion of various biological data and medical data, computational drug repositioning methods have been appealing and powerful techniques to systematically identify potential drug-target interactions and drug-disease interactions. In this review, we first summarize the available biomedical data and public databases related to drugs, diseases and targets. Then, we discuss existing drug repositioning approaches and group them based on their underlying computational models consisting of classical machine learning, network propagation, matrix factorization and completion, and deep learning based models. We also comprehensively analyze common standard data sets and evaluation metrics used in drug repositioning, and give a brief comparison of various prediction methods on the gold standard data sets. Finally, we conclude our review with a brief discussion on challenges in computational drug repositioning, which includes the problem of reducing the noise and incompleteness of biomedical data, the ensemble of various computation drug repositioning methods, the importance of designing reliable negative samples selection methods, new techniques dealing with the data sparseness problem, the construction of large-scale and comprehensive benchmark data sets and the analysis and explanation of the underlying mechanisms of predicted interactions. Huimin Luo, Min Li 0007, Mengyun Yang, Fang-Xiang Wu, Yaohang Li, Jianxin Wang 0001 |
Briefings Bioinform. | 2 |
| 2021 | DeepDTAF: a deep learning method to predict protein-ligand binding affinityabstractBiomolecular recognition between ligand and protein plays an essential role in drug discovery and development. However, it is extremely time and resource consuming to determine the protein-ligand binding affinity by experiments. At present, many computational methods have been proposed to predict binding affinity, most of which usually require protein 3D structures that are not often available. Therefore, new methods that can fully take advantage of sequence-level features are greatly needed to predict protein-ligand binding affinity and accelerate the drug discovery process. We developed a novel deep learning approach, named DeepDTAF, to predict the protein-ligand binding affinity. DeepDTAF was constructed by integrating local and global contextual features. More specifically, the protein-binding pocket, which possesses some special properties for directly binding the ligand, was firstly used as the local input feature for protein-ligand binding affinity prediction. Furthermore, dilated convolution was used to capture multiscale long-range interactions. We compared DeepDTAF with the recent state-of-art methods and analyzed the effectiveness of different parts of our model, the significant accuracy improvement showed that DeepDTAF was a reliable tool for affinity prediction. The resource codes and data are available at https: //github.com/KailiWang1/DeepDTAF. Renyi Zhou, Yaohang Li, Min Li 0007 |
Briefings Bioinform. | 4 |
| 2021 | NIDM: network impulsive dynamics on multiplex biological network for disease-gene predictionabstractThe prediction of genes related to diseases is important to the study of the diseases due to high cost and time consumption of biological experiments. Network propagation is a popular strategy for disease-gene prediction. However, existing methods focus on the stable solution of dynamics while ignoring the useful information hidden in the dynamical process, and it is still a challenge to make use of multiple types of physical/functional relationships between proteins/genes to effectively predict disease-related genes. Therefore, we proposed a framework of network impulsive dynamics on multiplex biological network (NIDM) to predict disease-related genes, along with four variants of NIDM models and four kinds of impulsive dynamical signatures (IDSs). NIDM is to identify disease-related genes by mining the dynamical responses of nodes to impulsive signals being exerted at specific nodes. By a series of experimental evaluations in various types of biological networks, we confirmed the advantage of multiplex network and the important roles of functional associations in disease-gene prediction, demonstrated superior performance of NIDM compared with four types of network-based algorithms and then gave the effective recommendations of NIDM models and IDS signatures. To facilitate the prioritization and analysis of (candidate) genes associated to specific diseases, we developed a user-friendly web server, which provides three kinds of filtering patterns for genes, network visualization, enrichment analysis and a wealth of external links (http://bioinformatics.csu.edu.cn/DGP/NID.jsp). NIDM is a protocol for disease-gene prediction integrating different types of biological networks, which may become a very useful computational tool for the study of disease-related genes. Ju Xiang, Jiashuai Zhang, Ruiqing Zheng, Xingyi Li 0003, Min Li 0007 |
Briefings Bioinform. | 5 |
| 2021 | Improving circRNA-disease association prediction by sequence and ontology representations with convolutional and recurrent neural networksabstractMOTIVATION: Emerging studies indicate that circular RNAs (circRNAs) are widely involved in the progression of human diseases. Due to its special structure which is stable, circRNAs are promising diagnostic and prognostic biomarkers for diseases. However, the experimental verification of circRNA-disease associations is expensive and limited to small-scale. Effective computational methods for predicting potential circRNA-disease associations are regarded as a matter of urgency. Although several models have been proposed, over-reliance on known associations and the absence of characteristics of biological functions make precise predictions are still challenging. RESULTS: In this study, we propose a method for predicting CircRNA-disease associations based on sequence and ontology representations, named CDASOR, with convolutional and recurrent neural networks. For sequences of circRNAs, we encode them with continuous k-mers, get low-dimensional vectors of k-mers, extract their local feature vectors with 1D CNN and learn their long-term dependencies with bi-directional long short-term memory. For diseases, we serialize disease ontology into sentences containing the hierarchy of ontology, obtain low-dimensional vectors for disease ontology terms and get terms' dependencies. Furthermore, we get association patterns of circRNAs and diseases from known circRNA-disease associations with neural networks. After the above steps, we get circRNAs' and diseases' high-level representations, which are informative to improve the prediction. The experimental results show that CDASOR provides an accurate prediction. Importing the characteristics of biological functions, CDASOR achieves impressive predictions in the de novo test. In addition, 6 of the top-10 predicted results are verified by the published literature in the case studies. AVAILABILITY AND IMPLEMENTATION: The code and data of CDASOR are freely available at https://github.com/BioinformaticsCSU/CDASOR. Chengqian Lu, Min Zeng 0004, Fang-Xiang Wu, Min Li 0007, Jianxin Wang 0001 |
Bioinform. | 4 |
| 2021 | Protein interaction networks: centrality, modularity, dynamics, and applications
Xiangmao Meng, Xiaoqing Peng, Yaohang Li, Min Li 0007 |
Frontiers Comput. Sci. | 5 |
| 2021 | DeepPPF: A deep learning framework for predicting protein family
Shehu Mohammed Yusuf, Fuhao Zhang, Min Zeng 0004, Min Li 0007 |
Neurocomputing | 4 |
| 2021 | DeepDSC: A Deep Learning Method to Predict Drug Sensitivity of Cancer Cell LinesabstractHigh-throughput screening technologies have provided a large amount of drug sensitivity data for a panel of cancer cell lines and hundreds of compounds. Computational approaches to analyzing these data can benefit anticancer therapeutics by identifying molecular genomic determinants of drug sensitivity and developing new anticancer drugs. In this study, we have developed a deep learning architecture to improve the performance of drug sensitivity prediction based on these data. We integrated both genomic features of cell lines and chemical information of compounds to predict the half maximal inhibitory concentrations [Formula: see text] on the Cancer Cell Line Encyclopedia (CCLE) and the Genomics of Drug Sensitivity in Cancer (GDSC) datasets using a deep neural network, which we called DeepDSC. Specifically, we first applied a stacked deep autoencoder to extract genomic features of cell lines from gene expression data, and then combined the compounds' chemical features to these genomic features to produce final response data. We conducted 10-fold cross-validation to demonstrate the performance of our deep model in terms of root-mean-square error (RMSE) and coefficient of determination [Formula: see text]. We show that our model outperforms the previous approaches with RMSE of 0.23 and [Formula: see text] of 0.78 on CCLE dataset, and RMSE of 0.52 and [Formula: see text] of 0.78 on GDSC dataset, respectively. Moreover, to demonstrate the prediction ability of our models on novel cell lines or novel compounds, we left cell lines originating from the same tissue and each compound out as the test sets, respectively, and the rest as training sets. The performance was comparable to other methods. Min Li 0007, Yake Wang, Ruiqing Zheng, Xinghua Shi, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | FUNMarker: Fusion Network-Based Method to Identify Prognostic and Heterogeneous Breast Cancer BiomarkersabstractBreast cancer is a heterogeneous disease with many clinically distinguishable molecular subtypes each corresponding to a cluster of patients. Identification of prognostic and heterogeneous biomarkers for breast cancer is to detect cluster-specific gene biomarkers which can be used for accurate survival prediction of breast cancer outcomes. In this study, we proposed a FUsion Network-based method (FUNMarker) to identify prognostic and heterogeneous breast cancer biomarkers by considering the heterogeneity of patient samples and biological information from multiple sources. To reduce the affect of heterogeneity of patients, samples were first clustered using the K-means algorithm based on the principal components of gene expression. For each cluster, to comprehensively evaluate the influence of genes on breast cancer, genes were weighted from three aspects: biological function, prognostic ability and correlation with known disease genes. Then they were ranked via a label propagation model on a fusion network that combined physical protein interactions from seven types of networks and thus could reduce the impact of incompleteness of interactome. We compared FUNMarker with three state-of-the-art methods and the results showed that biomarkers identified by FUNMarker were biological interpretable and had stronger discriminative power than the existing methods in differentiating patients with different prognostic outcomes. Xingyi Li 0003, Ju Xiang, Jianxin Wang 0001, Jinyan Li 0001, Fang-Xiang Wu, Min Li 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2021 | EPGA-SC : A Framework for de novo Assembly of Single-Cell Sequencing ReadsabstractAssembling genomes from single-cell sequencing data is essential for single-cell studies. However, single-cell assemblies are challenging due to (i) the highly non-uniform read coverage and (ii) the elevated levels of sequencing errors and chimeric reads. Although several assemblers for single-cell data have been proposed in recent years, most of them fail to construct correct long contigs. In this study, we present a new framework called EPGA-SC for de novo assembly of single-cell sequencing reads. The EPGA assembler has designed strategies to solve the problems caused by sequencing errors, sequencing biases, and repetitive regions. However, the extremely unbalanced and richer error types prevent EPGA to achieve high performance in single-cell sequencing data. In this study, we designed EPGA-SC based on EPGA. The main innovations of EPGA-SC are as follows: (i) classifying reads to reduce the proportion of false reads; (ii) using multiple sets of high precision paired-end reads generated from the high precision assemblies produced by other assembler such as SPAdes to overcome the impact of sequencing biases and repetitive regions; and (iii) developing novel algorithms for removing chimeric errors and extending contigs. We test EPGA-SC with seven datasets. The experimental results show that EPGA-SC can generate better assemblies than most current tools in most time in term of MAX contig, N50, NG50, NA50, and NGA50. Xingyu Liao, Min Li 0007, You Zou, Fang-Xiang Wu, Yi Pan 0001, Feng Luo 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | A Novel Drug Repositioning Approach Based on Collaborative Metric LearningabstractComputational drug repositioning, which is an efficient approach to find potential indications for drugs, has been used to increase the efficiency of drug development. The drug repositioning problem essentially is a top-K recommendation task that recommends most likely diseases to drugs based on drug and disease related information. Therefore, many recommendation methods can be adopted to drug repositioning. Collaborative metric learning (CML) algorithm can produce distance metrics that capture the important relationships among objects, and has been widely used in recommendation domains. By applying CML in drug repositioning, a joint metric space is learned to encode drug's relationships with different diseases. In this study, we propose a novel drug repositioning computational method using Collaborative Metric Learning to predict novel drug-disease associations based on known drug and disease related information. Specifically, the proposed method learns latent vectors of drugs and diseases by applying metric learning, and then predicts the association probability of one drug-disease pair based on the learned vectors. The comprehensive experimental results show that CMLDR outperforms the other state-of-the-art drug repositioning algorithms in terms of precision, recall, and AUPR. Huimin Luo, Jianxin Wang 0001, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | A Deep Learning Framework for Identifying Essential Proteins by Integrating Multiple Types of Biological InformationabstractComputational methods including centrality and machine learning-based methods have been proposed to identify essential proteins for understanding the minimum requirements of the survival and evolution of a cell. In centrality methods, researchers are required to design a score function which is based on prior knowledge, yet is usually not sufficient to capture the complexity of biological information. In machine learning-based methods, some selected biological features cannot represent the complete properties of biological information as they lack a computational framework to automatically select features. To tackle these problems, we propose a deep learning framework to automatically learn biological features without prior knowledge. We use node2vec technique to automatically learn a richer representation of protein-protein interaction (PPI) network topologies than a score function. Bidirectional long short term memory cells are applied to capture non-local relationships in gene expression data. For subcellular localization information, we exploit a high dimensional indicator vector to characterize their feature. To evaluate the performance of our method, we tested it on PPI network of S. cerevisiae. Our experimental results demonstrate that the performance of our method is better than traditional centrality methods and is superior to existing machine learning-based methods. To explore which of the three types of biological information is the most vital element, we conduct an ablation study by removing each component in turn. Our results show that the PPI network embedding contributes most to the improvement. In addition, gene expression profiles and subcellular localization information are also helpful to improve the performance in identification of essential proteins. Min Zeng 0004, Min Li 0007, Zhihui Fei, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | DMFLDA: A Deep Learning Framework for Predicting lncRNA-Disease AssociationsabstractA growing amount of evidence suggests that long non-coding RNAs (lncRNAs) play important roles in the regulation of biological processes in many human diseases. However, the number of experimentally verified lncRNA-disease associations is very limited. Thus, various computational approaches are proposed to predict lncRNA-disease associations. Current matrix factorization-based methods cannot capture the complex non-linear relationship between lncRNAs and diseases, and traditional machine learning-based methods are not sufficiently powerful to learn the representation of lncRNAs and diseases. Considering these limitations in existing computational methods, we propose a deep matrix factorization model to predict lncRNA-disease associations (DMFLDA in short). DMFLDA uses a cascade of non-linear hidden layers to learn latent representation to represent lncRNAs and diseases. By using non-linear hidden layers, DMFLDA captures the more complex non-linear relationship between lncRNAs and diseases than traditional matrix factorization-based methods. In addition, DMFLDA learns features directly from the lncRNA-disease interaction matrix and thus can obtain more accurate representation learning for lncRNAs and diseases than traditional machine learning methods. The low dimensional representations of the lncRNAs and diseases are fused to estimate the new interaction value. To evaluate the performance of DMFLDA, we perform leave-one-out cross-validation and 5-fold cross-validation on known experimentally verified lncRNA-disease associations. The experimental results show that DMFLDA performs better than the existing methods. The case studies show that many predicted interactions of colorectal cancer, prostate cancer, and renal cancer have been verified by recent biomedical literature. The source code and datasets can be obtained from https://github.com/CSUBioGroup/DMFLDA. Min Zeng 0004, Chengqian Lu, Zhihui Fei, Fang-Xiang Wu, Yaohang Li, Jianxin Wang 0001, Min Li 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2021 | Deletion Detection Method Using the Distribution of Insert Size and a Precise Alignment StrategyabstractHomozygous and heterozygous deletions commonly exist in the human genome. For current structural variation detection tools, it is significant to determine whether a deletion is homozygous or heterozygous. However, the problems of sequencing errors, micro-homologies, and micro-insertions prohibit common alignment tools from identifying accurate breakpoint locations, and often result in detecting false structural variations. In this study, we present a novel deletion detection tool called Sprites2. Comparing with Sprites, Sprites2 makes the following modifications: (1) The distribution of insert size is used in Sprites2, which can identify the type of deletions and improve the accuracy of deletion calls. (2) A precise alignment method based on AGE (one algorithm simultaneously aligning 5' and 3' ends between two sequences) is adopted in Sprites2 to identify breakpoints, which is helpful to resolve the problems introduced by sequencing errors, micro-homologies, and micro-insertions. In order to test and verify the performance of Sprites2, some simulated and real datasets are adopted in our experiments, and Sprites2 is compared with five popular tools. The experimental results show that Sprites2 can improve the performance of deletion detection. Sprites2 can be downloaded from https://github.com/zhangzhen/sprites2. Zhen Zhang 0024, Juan Shang, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | A Deep Learning Framework for Gene Ontology Annotations With Sequence- and Network-Based InformationabstractKnowledge of protein functions plays an important role in biology and medicine. With the rapid development of high-throughput technologies, a huge number of proteins have been discovered. However, there are a great number of proteins without functional annotations. A protein usually has multiple functions and some functions or biological processes require interactions of a plurality of proteins. Additionally, Gene Ontology provides a useful classification for protein functions and contains more than 40,000 terms. We propose a deep learning framework called DeepGOA to predict protein functions with protein sequences and protein-protein interaction (PPI) networks. For protein sequences, we extract two types of information: sequence semantic information and subsequence-based features. We use the word2vec technique to numerically represent protein sequences, and utilize a Bi-directional Long and Short Time Memory (Bi-LSTM) and multi-scale convolutional neural network (multi-scale CNN) to obtain the global and local semantic features of protein sequences, respectively. Additionally, we use the InterPro tool to scan protein sequences for extracting subsequence-based information, such as domains and motifs. Then, the information is plugged into a neural network to generate high-quality features. For the PPI network, the Deepwalk algorithm is applied to generate its embedding information of PPI. Then the two types of features are concatenated together to predict protein functions. To evaluate the performance of DeepGOA, several different evaluation methods and metrics are utilized. The experimental results show that DeepGOA outperforms DeepGO and BLAST. Fuhao Zhang, Hong Song 0004, Min Zeng 0004, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001, Min Li 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2021 | An Ensemble Method to Reconstruct Gene Regulatory Networks Based on Multivariate Adaptive Regression SplinesabstractGene regulatory networks (GRNs) play a key role in biological processes. However, GRNs are diverse under different biological conditions. Reconstructing gene regulatory networks (GRNs) from gene expression has become an important opportunity and challenge in the past decades. Although there are a lot of existing methods to infer the topology of GRNs, such as mutual information, random forest, and partial least squares, the accuracy is still low due to the noise and high dimension of the expression data. In this paper, we introduce an ensemble Multivariate Adaptive Regression Splines (MARS) based method to reconstruct the directed GRNs from multifactorial gene expression data, called PBMarsNet. PBMarsNet incorporates part mutual information (PMI) to pre-weight the candidate regulatory genes and then uses MARS to detect the nonlinear regulatory links. Moreover, we apply bootstrap to run the MARS multiple times and average the outputs of each MARS as the final score of regulatory links. The results on DREAM4 challenge and DREAM5 challenge datasets show PBMarsNet has a superior performance and generalization over other state-of-the-art methods. Ruiqing Zheng, Min Li 0007, Xiang Chen 0029, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | Deep Matrix Factorization Improves Prediction of Human CircRNA-Disease AssociationsabstractIn recent years, more and more evidence indicates that circular RNAs (circRNAs) with covalently closed loop play various roles in biological processes. Dysregulation and mutation of circRNAs may be implicated in diseases. Due to its stable structure and resistance to degradation, circRNAs provide great potential to be diagnostic biomarkers. Therefore, predicting circRNA-disease associations is helpful in disease diagnosis. However, there are few experimentally validated associations between circRNAs and diseases. Although several computational methods have been proposed, precisely representing underlying features and grasping the complex structures of data are still challenging. In this paper, we design a new method, called DMFCDA (Deep Matrix Factorization CircRNA-Disease Association), to infer potential circRNA-disease associations. DMFCDA takes both explicit and implicit feedback into account. Then, it uses a projection layer to automatically learn latent representations of circRNAs and diseases. With multi-layer neural networks, DMFCDA can model the non-linear associations to grasp the complex structure of data. We assess the performance of DMFCDA using leave-one cross-validation and 5-fold cross-validation on two datasets. Computational results show that DMFCDA efficiently infers circRNA-disease associations according to AUC values, the percentage of precisely retrieved associations in various top ranks, and statistical comparison. We also conduct case studies to evaluate DMFCDA. All results show that DMFCDA provides accurate predictions. Chengqian Lu, Min Zeng 0004, Fuhao Zhang, Fang-Xiang Wu, Min Li 0007, Jianxin Wang 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | A topological AUC-based biomarker ensemble method for the complex disease analysisabstractComplex diseases are affected by many factors, and their pathogenic mechanism is complicated, which brings difficulties to the analysis and treatment of diseases. AUC, the area under the ROC curve, is often used as a gold standard to evaluate the performance of a binary classifier. The existing methods of constructing classifier by optimizing AUC are easy to fall into local optimum, and have high time complexity, which is not suitable for real-time analysis of high-dimensional gene expression data. With the rapid development of high-throughput sequencing technology, feature selection and model estimation become the necessary means to reduce the dimension and complexity of data, and the selected important features have the potential as biomarkers to reveal the pathogenesis of diseases. In this paper, we proposed a topological AUC-based biomarker ensemble method for the complex disease analysis, which uses gene expression data and the topological information derived from the protein-protein interaction network to identify biomarkers. The main contribution is to optimize two objectives simultaneously: maximizing the AUC score and minimizing the number of selected features. We applied the proposed method to analyze two types of problems: 1) prognosis of breast cancer, 2) classification of similar diseases. The results show that our method can effectively identify a small set of biomarkers with the powerful classification ability and the biological interpretability. Xingyi Li 0003, Ju Xiang, Fang-Xiang Wu, Min Li 0007 |
BIBM | 4 |
| 2020 | A robust single cell clustering method based on subspace learning and partial imputationabstractCell heterogeneity analysis is an important and urgent task in single cell data research. Numerous cell type identification methods have been proposed to address the issue. Due to the high rate of dropout and complex biological background, it is still a challenging task to obtain the accurate clusters of cells. In this study, we propose a robust single cell clustering method based on subspace learning and partial imputation, called RCSLI. RCSLI incorporates a modified variable genes selection method and utilizes the self-expression of scRNA-seq data to learn sparse cell-to-cell similarity and impute part of missing expression values. To evaluate the clustering performance of RCSLI, we compare it with nine state-of-the-art single cell clustering methods on eight scRNA-seq datasets. The experimental results show that RCSLI gets more accurate and robust clustering results. The imputation impact on the specific gene markers is evaluated on PBMC data. The classification results by taking these marker genes as predictors show RCSLI recovers the real dropouts, meanwhile, introduces less noise. Ruiqing Zheng, Zhenlan Liang, Xiangmao Meng, Yu Tian 0015, Min Li 0007 |
BIBM | 5 |
| 2020 | SPOC: Identification of Drug Targets in Biological Networks via Set Preference Output Control
Min Li 0007, Fang-Xiang Wu |
ISBRA | 2 |
| 2020 | Ess-NEXG: Predict Essential Proteins by Constructing a Weighted Protein Interaction Network Based on Node Embedding and XGBoost
Min Zeng 0004, Jiashuai Zhang, Min Li 0007 |
ISBRA | 5 |
| 2020 | Network-based methods for predicting essential genes or proteins: a surveyabstractGenes that are thought to be critical for the survival of organisms or cells are called essential genes. The prediction of essential genes and their products (essential proteins) is of great value in exploring the mechanism of complex diseases, the study of the minimal required genome for living cells and the development of new drug targets. As laboratory methods are often complicated, costly and time-consuming, a great many of computational methods have been proposed to identify essential genes/proteins from the perspective of the network level with the in-depth understanding of network biology and the rapid development of biotechnologies. Through analyzing the topological characteristics of essential genes/proteins in protein-protein interaction networks (PINs), integrating biological information and considering the dynamic features of PINs, network-based methods have been proved to be effective in the identification of essential genes/proteins. In this paper, we survey the advanced methods for network-based prediction of essential genes/proteins and present the challenges and directions for future research. Xingyi Li 0003, Min Zeng 0004, Ruiqing Zheng, Min Li 0007 |
Briefings Bioinform. | 5 |
| 2020 | Protein-protein interaction site prediction through combining local and global features with deep neural networksabstractMOTIVATION: Protein-protein interactions (PPIs) play important roles in many biological processes. Conventional biological experiments for identifying PPI sites are costly and time-consuming. Thus, many computational approaches have been proposed to predict PPI sites. Existing computational methods usually use local contextual features to predict PPI sites. Actually, global features of protein sequences are critical for PPI site prediction. RESULTS: A new end-to-end deep learning framework, named DeepPPISP, through combining local contextual and global sequence features, is proposed for PPI site prediction. For local contextual features, we use a sliding window to capture features of neighbors of a target amino acid as in previous studies. For global sequence features, a text convolutional neural network is applied to extract features from the whole protein sequence. Then the local contextual and global sequence features are combined to predict PPI sites. By integrating local contextual and global sequence features, DeepPPISP achieves the state-of-the-art performance, which is better than the other competing methods. In order to investigate if global sequence features are helpful in our deep learning model, we remove or change some components in DeepPPISP. Detailed analyses show that global sequence features play important roles in DeepPPISP. AVAILABILITY AND IMPLEMENTATION: The DeepPPISP web server is available at http://bioinformatics.csu.edu.cn/PPISP/. The source code can be obtained from https://github.com/CSUBioGroup/DeepPPISP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Min Zeng 0004, Fuhao Zhang, Fang-Xiang Wu, Yaohang Li, Jianxin Wang 0001, Min Li 0007 |
Bioinform. | 6 |
| 2020 | PROBselect: accurate prediction of protein-binding residues from proteins sequences via dynamic predictor selectionabstractMOTIVATION: Knowledge of protein-binding residues (PBRs) improves our understanding of protein-protein interactions, contributes to the prediction of protein functions and facilitates protein-protein docking calculations. While many sequence-based predictors of PBRs were published, they offer modest levels of predictive performance and most of them cross-predict residues that interact with other partners. One unexplored option to improve the predictive quality is to design consensus predictors that combine results produced by multiple methods. RESULTS: We empirically investigate predictive performance of a representative set of nine predictors of PBRs. We report substantial differences in predictive quality when these methods are used to predict individual proteins, which contrast with the dataset-level benchmarks that are currently used to assess and compare these methods. Our analysis provides new insights for the cross-prediction concern, dissects complementarity between predictors and demonstrates that predictive performance of the top methods depends on unique characteristics of the input protein sequence. Using these insights, we developed PROBselect, first-of-its-kind consensus predictor of PBRs. Our design is based on the dynamic predictor selection at the protein level, where the selection relies on regression-based models that accurately estimate predictive performance of selected predictors directly from the sequence. Empirical assessment using a low-similarity test dataset shows that PROBselect provides significantly improved predictive quality when compared with the current predictors and conventional consensuses that combine residue-level predictions. Moreover, PROBselect informs the users about the expected predictive quality for the prediction generated from a given input protein. AVAILABILITY AND IMPLEMENTATION: PROBselect is available at http://bioinformatics.csu.edu.cn/PROBselect/home/index. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fuhao Zhang, Jian Zhang 0020, Min Zeng 0004, Min Li 0007, Lukasz A. Kurgan |
Bioinform. | 5 |
| 2020 | DeepFrag-k: a fragment-based deep learning approach for protein fold recognitionabstractBACKGROUND: One of the most essential problems in structural bioinformatics is protein fold recognition. In this paper, we design a novel deep learning architecture, so-called DeepFrag-k, which identifies fold discriminative features at fragment level to improve the accuracy of protein fold recognition. DeepFrag-k is composed of two stages: the first stage employs a multi-modal Deep Belief Network (DBN) to predict the potential structural fragments given a sequence, represented as a fragment vector, and then the second stage uses a deep convolutional neural network (CNN) to classify the fragment vector into the corresponding fold. RESULTS: Our results show that DeepFrag-k yields 92.98% accuracy in predicting the top-100 most popular fragments, which can be used to generate discriminative fragment feature vectors to improve protein fold recognition. CONCLUSIONS: There is a set of fragments that can serve as structural "keywords" distinguishing between major protein folds. The deep learning architecture in DeepFrag-k is able to accurately identify these fragments as structure features to improve protein fold recognition. Wessam Elhefnawy, Min Li 0007, Jianxin Wang 0001, Yaohang Li |
BMC Bioinform. | 2 |
| 2020 | MADA: a web service for analysing DNA methylation array dataabstractBACKGROUND: DNA methylation in the human genome is acknowledged to be widely associated with biological processes and complex diseases. The Illumina Infinium methylation arrays have been approved as one of the most efficient and universal technologies to investigate the whole genome changes of methylation patterns. As methylation arrays may still be the dominant method for detecting methylation in the anticipated future, it is crucial to develop a reliable workflow to analysis methylation array data. RESULTS: In this study, we develop a web service MADA for the whole process of methylation arrays data analysis, which includes the steps of a comprehensive differential methylation analysis pipeline: pre-processing (data loading, quality control, data filtering, and normalization), batch effect correction, differential methylation analysis, and downstream analysis. In addition, we provide the visualization of pre-processing, differentially methylated probes or regions, gene ontology, pathway and cluster analysis results. Moreover, a customization function for users to define their own workflow is also provided in MADA. CONCLUSIONS: With the analysis of two case studies, we have shown that MADA can complete the whole procedure of methylation array data analysis. MADA provides a graphical user interface and enables users with no computational skills and limited bioinformatics background to carry on complicated methylation array data analysis. The web server is available at: http://120.24.94.89:8080/MADA. Linconghua Wang, Fang-Xiang Wu, Min Li 0007 |
BMC Bioinform. | 5 |
| 2020 | NEDD: a network embedding based method for predicting drug-disease associationsabstractBACKGROUND: Drug discovery is known for the large amount of money and time it consumes and the high risk it takes. Drug repositioning has, therefore, become a popular approach to save time and cost by finding novel indications for approved drugs. In order to distinguish these novel indications accurately in a great many of latent associations between drugs and diseases, it is necessary to exploit abundant heterogeneous information about drugs and diseases. RESULTS: In this article, we propose a meta-path-based computational method called NEDD to predict novel associations between drugs and diseases using heterogeneous information. First, we construct a heterogeneous network as an undirected graph by integrating drug-drug similarity, disease-disease similarity, and known drug-disease associations. NEDD uses meta paths of different lengths to explicitly capture the indirect relationships, or high order proximity, within drugs and diseases, by which the low dimensional representation vectors of drugs and diseases are obtained. NEDD then uses a random forest classifier to predict novel associations between drugs and diseases. CONCLUSIONS: The experiments on a gold standard dataset which contains 1933 validated drug-disease associations show that NEDD produces superior prediction results compared with the state-of-the-art approaches. Renyi Zhou, Zhangli Lu, Huimin Luo, Ju Xiang, Min Zeng 0004, Min Li 0007 |
BMC Bioinform. | 6 |
| 2020 | miRTRS: A Recommendation Algorithm for Predicting miRNA TargetsabstractmicroRNAs (miRNAs) are small and important non-coding RNAs that regulate gene expression in transcriptional and post-transcriptional level by combining with their targets (genes). Predicting miRNA targets is an important problem in biological research. It is expensive and time-consuming to identify miRNA targets by using biological experiments. Many computational methods have been proposed to predict miRNA targets. In this study, we develop a novel method, named miRTRS, for predicting miRNA targets based on a recommendation algorithm. miRTRS can predict targets for an isolated (new) miRNA with miRNA sequence similarity, as well as isolated (new) targets for a miRNA with gene sequence similarity. Furthermore, when compared to supervised machine learning methods, miRTRS does not need to select negative samples. We use 10-fold cross validation and independent datasets to evaluate the performance of our method. We compared miRTRS with two most recently published methods for miRNA target prediction. The experimental results have shown that our method miRTRS outperforms competing prediction methods in terms of AUC and other evaluation metrics. Hui Jiang 0008, Jianxin Wang 0001, Min Li 0007, Wei Lan 0001, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | United Neighborhood Closeness Centrality and Orthology for Predicting Essential ProteinsabstractIdentifying essential proteins plays an important role in disease study, drug design, and understanding the minimal requirement for cellular life. Computational methods for essential proteins discovery overcome the disadvantages of biological experimental methods that are often time-consuming, expensive, and inefficient. The topological features of protein-protein interaction (PPI) networks are often used to design computational prediction methods, such as Degree Centrality (DC), Betweenness Centrality (BC), Closeness Centrality (CC), Subgraph Centrality (SC), Eigenvector Centrality (EC), Information Centrality (IC), and Neighborhood Centrality (NC). However, the prediction accuracies of these individual methods still have space to be improved. Studies show that additional information, such as orthologous relations, helps discover essential proteins. Many researchers have proposed different methods by combining multiple information sources to gain improvement of prediction accuracy. In this study, we find that essential proteins appear in triangular structure in PPI network significantly more often than nonessential ones. Based on this phenomenon, we propose a novel pure centrality measure, so-called Neighborhood Closeness Centrality (NCC). Accordingly, we develop a new combination model, Extended Pareto Optimality Consensus model, named EPOC, to fuse NCC and Orthology information and a novel essential proteins identification method, NCCO, is fully proposed. Compared with seven existing classic centrality methods (DC, BC, IC, CC, SC, EC, and NC) and three consensus methods (PeC, ION, and CSC), our results on S.cerevisiae and E.coli datasets show that NCCO has clear advantages. As a consensus method, EPOC also yields better performance than the random walk model. Gaoshi Li, Min Li 0007, Jianxin Wang 0001, Yaohang Li, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | Identification of Protein Complexes by Using a Spatial and Temporal Active Protein Interaction NetworkabstractThe rapid development of proteomics and high-throughput technologies has produced a large amount of Protein-Protein Interaction (PPI) data, which makes it possible for considering dynamic properties of protein interaction networks (PINs) instead of static properties. Identification of protein complexes from dynamic PINs becomes a vital scientific problem for understanding cellular life in the post genome era. Up to now, plenty of models or methods have been proposed for the construction of dynamic PINs to identify protein complexes. However, most of the constructed dynamic PINs just focus on the temporal dynamic information and thus overlook the spatial dynamic information of the complex biological systems. To address the limitation of the existing dynamic PIN analysis approaches, in this paper, we propose a new model-based scheme for the construction of the Spatial and Temporal Active Protein Interaction Network (ST-APIN) by integrating time-course gene expression data and subcellular location information. To evaluate the efficiency of ST-APIN, the commonly used classical clustering algorithm MCL is adopted to identify protein complexes from ST-APIN and the other three dynamic PINs, NF-APIN, DPIN, and TC-PIN. The experimental results show that, the performance of MCL on ST-APIN outperforms those on the other three dynamic PINs in terms of matching with known complexes, sensitivity, specificity, and f-measure. Furthermore, we evaluate the identified protein complexes by Gene Ontology (GO) function enrichment analysis. The validation shows that the identified protein complexes from ST-APIN are more biologically significant. This study provides a general paradigm for constructing the ST-APINs, which is essential for further understanding of molecular systems and the biomedical mechanism of complex diseases. Min Li 0007, Xiangmao Meng, Ruiqing Zheng, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Improving de novo Assembly Based on Read ClassificationabstractDue to sequencing bias, sequencing error, and repeat problems, the genome assemblies usually contain misarrangements and gaps. When tackling these problems, current assemblers commonly consider the read libraries as a whole and adopt the same strategy to deal with them. However, if we can divide reads into different categories and take different assembly strategies for different read categories, we expect to reduce the mutual effects on problems in genome assembly and facilitate to produce satisfactory assemblies. In this paper, we present a new pipeline for genome assembly based on read classification (ARC). ARC classifies reads into three categories according to the frequencies of k-mers they contain. The three categories refer to (1) low depth reads, which contain a certain low frequency k-mers and are often caused by sequencing errors or bias; (2) high depth reads, which contain a certain high frequency k-mers and usually come from repetitive regions; and (3) normal depth reads, which are the rest of reads. After read classification, an existing assembler is used to assemble different read categories separately, which is beneficial to resolve problems in the genome assembly. ARC adopts loose assembly parameters for low depth reads, and strict assembly parameters for normal depth and high depth reads. We test ARC using five datasets. The experimental results show that, assemblers combining with ARC can generate better assemblies in terms of NA50, NGA50, and genome fraction. Xingyu Liao, Min Li 0007, You Zou, Fang-Xiang Wu, Yi Pan 0001, Feng Luo 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | An Efficient Trimming Algorithm based on Multi-Feature Fusion Scoring Model for NGS DataabstractNext-generation sequencing (NGS) has enabled an exponential growth rate of sequencing data. However, several sequence artifacts, including error reads (base calling errors and small insertions or deletions) and poor quality reads, which can impose significant impact on the downstream sequence processing and analysis. Here, we present PE-Trimmer, a sensitive and special trimming algorithm for NGS sequence. First, PE-Trimmer removes technical sequences in paired-end reads based on the characteristics of low quality reads in NGS data. Second, PE-Trimmer determines the range of reads that need to be trimmed according to the quality score statistics histogram of reads in the library. To improve the accuracy of this algorithm, we design a light-weight and easy-to-explain scoring model to evaluate candidates in the pattern of trimming step. Finally, PE-Trimmer selects the appropriate trimming strategy to process the low quality reads based on the location determined by the scoring model. PE-Trimmer is able to locate and remove adapter residues from the paired-end reads. It is easily configurable and offers superior throughput in the multi-threaded mode. We test PE-Trimmer on five datasets, and compare it with the current five latest methods. The experimental results demonstrate that PE-Trimmer produces more superior results, compared with other trimmers. Xingyu Liao, Min Li 0007, You Zou, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | GapReduce: A Gap Filling Algorithm Based on Partitioned Read SetsabstractWith the advances in technologies of sequencing and assembly, draft sequences of more and more genomes are available. However, there commonly exist gaps in these draft sequences which influence various downstream analysis of biological studies. Gap filling methods can shorten the length of gaps and improve the completion of these draft sequences of genomes. Although some gap filling tools have been developed, their effectiveness and accuracy need to be improved. In this study, we develop a novel tool, called GapReduce, which can fill the gaps using the paired reads. For a gap, GapReduce selects the reads whose mate reads are aligned on the left or the right flanking region, and partitions the reads to two sets. Then GapReduce adopts different $k$k values and $k$k-$mer$mer frequency thresholds to iteratively construct De Bruijn graphs, which are used for finding the correct path to fill the gap. For overcoming the branching problems caused by repetitive regions and sequencing errors in the procedure of path selection, GapReduce designs a novel approach that simultaneously considers $k$k-$mer$mer frequency and distribution of paired reads based on the partitioned read sets. We compare the performance of GapReduce with current popular gap filling tools. The experimental results demonstrate that GapReduce can produce satisfactory gap filling results, especially for long insert size datasets. GapReduce is publicly available for downloading at https://github.com/bioinfomaticsCSU/GapReduce. Jianxin Wang 0001, Juan Shang, Huimin Luo, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2020 | MEC: Misassembly Error Correction in Contigs based on Distribution of Paired-End Reads and Statistics of GC-contentsabstractThe de novo assembly tools aim at reconstructing genomes from next-generation sequencing (NGS) data. However, the assembly tools usually generate a large amount of contigs containing many misassemblies, which are caused by problems of repetitive regions, chimeric reads, and sequencing errors. As they can improve the accuracy of assembly results, detecting and correcting the misassemblies in contigs are appealing, yet challenging. In this study, a novel method, called MEC, is proposed to identify and correct misassemblies in contigs. Based on the insert size distribution of paired-end reads and the statistical analysis of GC-contents, MEC can identify more misassemblies accurately. We evaluate our MEC with the metrics (NA50, NGA50) on four datasets, compared it with the most available misassembly correction tools, and carry out experiments to analyze the influence of MEC on scaffolding results, which shows that MEC can reduce misassemblies effectively and result in quantitative improvements in scaffolding quality. MEC is publicly available at https://github.com/bioinfomaticsCSU/MEC. Binbin Wu, Min Li 0007, Xingyu Liao, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | miRTMC: A miRNA Target Prediction Method Based on Matrix Completion AlgorithmabstractmicroRNAs (miRNAs) are small non-coding RNAs which modulate the stability of gene targets and their rates of translation into proteins at transcriptional level and post-transcriptional level. miRNA dysfunctions can lead to human diseases because of dysregulation of their targets. Correct miRNA target prediction will lead to better understanding of the mechanisms of human diseases and provide hints on curing them. In recent years, computational miRNA target prediction methods have been proposed according to the interaction rules between miRNAs and targets. However, these methods suffer from high false positive rates due to the complicated relationship between miRNAs and their targets. The rapidly growing number of experimentally validated miRNA targets enables predicting miRNA targets with high precision via accurate data analysis. Taking advantage of these known miRNA targets, a novel recommendation system model (miRTMC) for miRNA target prediction is established using a new matrix completion algorithm. In miRTMC, a heterogeneous network is constructed by integrating the miRNA similarity network, the gene similarity network, and the miRNA-gene interaction network. Our assumption is that the latent factors determining whether a gene is the target of miRNA or not are highly correlated, i.e., the adjacency matrix of the heterogeneous network is low-rank, which is then completed by using a nuclear norm regularized linear least squares model under non-negative constraints. Alternating direction method of multipliers (ADMM) is adopted to numerically solve the matrix completion problem. Our results show that miRTMC outperforms the competing methods in terms of various evaluation metrics. Our software package is available at https://github.com/hjiangcsu/miRTMC. Hui Jiang 0008, Mengyun Yang, Xiang Chen 0029, Min Li 0007, Yaohang Li, Jianxin Wang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Predicting Human lncRNA-Disease Associations Based on Geometric Matrix CompletionabstractRecently, increasing evidences reveal that dysregulations of long non-coding RNAs (lncRNAs) are relevant to diverse diseases. However, the number of experimentally verified lncRNA-disease associations is limited. Prioritizing potential associations is beneficial not only for disease diagnosis, but also disease treatment, more important apprehending disease mechanisms at lncRNA level. Various computational methods have been proposed, but precise prediction and full use of data's intrinsic structure are still challenging. In this work, we design a new method, denominated GMCLDA (Geometric Matrix Completion lncRNA-Disease Association), to infer underlying associations based on geometric matrix completion. Utilizing association patterns among functionally similar lncRNAs and phenotypically similar diseases, GMCLCA makes use of the intrinsic structure embedded in the association matrix. Besides, limiting the scope of the predicted values gives rise to a certain sparsity in computation and enhances the robustness of GMCLDA. GMCLDA computes disease semantic similarity according to the Disease Ontology (DO) hierarchy and lncRNA Gaussian interaction profile kernel similarity according to known interaction profiles. Then, GMCLDA measures lncRNA sequence similarity using Needleman-Wunsch algorithm. For a new lncRNA, GMCLDA prefills interaction profile on account of its K-nearest neighbors defined by sequence similarity. Finally, GMCLDA estimates the missing entries of the association matrix based on geometric matrix completion model. Compared with state-of-the-art methods, GMCLDA can provide more accurate lncRNA-disease prediction. Further case studies prove that GMCLDA is able to correctly infer possible lncRNAs for renal cancer. Chengqian Lu, Mengyun Yang, Min Li 0007, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2019 | DoRC: Discovery of rare cells from ultra-large scRNA-seq dataabstractThe advent of droplet-based transcriptomics platforms has enabled parallel screening over thousands or millions of cells. One of the challenging issues is to identify the rare cells from the ultra-large scRNA-seq data. Existing algorithms to find rare cells are time consuming or memory-exhausting. We propose an efficient and accurate method, Discovery of Rare Cells (DoRC). The rareness scores generated by DoRC can help biologists focus the downstream analyses only on a fraction of expression profiles within ultra-large scRNA-seq data. We also demonstrate the efficacy of DoRC in delineating human blood dendritic cell sub-types using ~68k single-cell expression profiles of human blood cells. DoRC can recover artificially planted rare cells and is sensitive to cell type identities as well. Xiang Chen 0029, Fang-Xiang Wu, Jin Chen 0004, Min Li 0007 |
BIBM | 4 |
| 2019 | Classification of Schizophrenia by Iterative Random Forest Feature Selection Based on DNA Methylation Array DataabstractChanges in DNA methylation are widely thought to be involved in the evolution of the disease, and most studies suggest that whole genome hypo-methylation levels are widespread in patients with schizophrenia. Since the exploration of DNA methylation changes in the etiology and pathogenesis of schizophrenia will be important for the prevention and early intervention, it is crucial to identify differentially methylated sites with high specificity and sensitivity for disease classification. In this study, we present a comprehensive approach MethIRF for the DNA methylation-based classification of schizophrenia by iterative random forest feature selection. The results show that MethIRF has a powerful discrimination ability compared to other four criteria methods for detecting differentially methylation sites. Moreover, the proposed method can discover significant CpG sites associated with schizophrenia and explore changes in the biological mechanisms of diseases. Min Li 0007, Linconghua Wang, Xingyi Li 0003, Fang-Xiang Wu, Jianxin Wang 0001 |
BIBM | 2 |
| 2019 | DualRank: multiplex network-based dual ranking for heterogeneous complex disease analysisabstractAnalysis of heterogeneous complex diseases based on the expression of biomarkers has always been the focus of medical research. In the past studies, the powerful network propagation has been applied in finding marker genes related to specific diseases. However, the network propagation model largely depends on the reliability and integrity of the network data, current networks may cause some problems due to the incompleteness of the networks. In this study, we developed a multiplex network-based dual ranking framework (DualRank) for heterogeneous complex disease analysis. We applied the proposed method to heterogeneous complex diseases for disease diagnosis, cancer prognosis, and similar disease classification. The results showed that DualRank outperformed current methods and could identify biomarkers with small quantity, strong prediction accuracy and biological interpretability. Xingyi Li 0003, Ju Xiang, Fang-Xiang Wu, Min Li 0007 |
BIBM | 5 |
| 2019 | HNEDTI: Prediction of drug-target interaction based on heterogeneous network embeddingabstractIdentifying drug-target interactions (DTIs) is an important task in drug discovery. Various computational models have been proposed to predict potential association between drugs and targets. However, it is still a great challenge to accurately predict the potential drug-target interactions with rare known drug-target interactions. In this work, we propose a heterogeneous network embedding model to predict drug-target interactions, called HNEDTI. Based on the assumption that similar drugs share similar patterns of relationships with target proteins, we integrate the drug-drug similarity network, target-target similarity network and known drug-target interactions into a heterogeneous network. HNEDTI can learn more accurate feature representation of drugs and targets by extract both local and global information of the heterogeneous network from different lengths of meta-paths. The low dimensional feature representation vectors of drugs and targets are applied to random forest model to predict whether the given drug-target pair has an interaction. The evaluation on four benchmark datasets (Enzyme, Ion Channel, GPCR and Nuclear Receptor) shows that our method HNEDTI outperforms the previous methods. Zhangli Lu, Yake Wang, Min Zeng 0004, Min Li 0007 |
BIBM | 4 |
| 2019 | Detecting protein complex based on hierarchical compressing network embeddingabstractDetecting protein complexes from protein-protein interaction (PPI) networks provides biologists an opportunity to efficiently understand the cellular organizations and functions. Existing computational methods just focus on mining high-density regions as the protein complexes by searching the local topological information of a PPI network and ignore the global topological information. To address this limitation, in this study, we present a novel protein complex detection method based on hierarchical compressing network embedding, named DPC-HCNE. The proposed method can preserve both the local topological information and global topological information of a PPI network. To evaluate the performance of our method, DPC-HCNE is compared with other eight typical clustering algorithms to detect protein complexes on two yeast datasets. The experimental results show that DPC-HCNE outperforms those state-of-the-art complex detection methods. Xiangmao Meng, Xiaoqing Peng, Fang-Xiang Wu, Min Li 0007 |
BIBM | 4 |
| 2019 | Tentative diagnosis prediction via deep understanding of patient narrativesabstractA tentative diagnosis is a preliminary suspicion of patient status, which is usually made by physicians according to patient narrative right at admission. It largely depends on the experiences and professional knowledge of physicians. We explored a combination model for automatic tentative diagnosis prediction based on clinical narratives. Text features are extracted in two ways. Firstly, the context semantic features are extracted by attention-based bidirectional long-short term memory (BiLSTM) network. Secondly, the symptom concepts recognized from input texts by Metamap and are vectorized by TF-IDF. Two combination strategies are proposed to utilize both two features for one candidate international classification of diseases (ICD) code recommendation: feature vectors combination and prediction results combination. The experiments performed on MIMIC III dataset. Both of the two combination strategies achieved better performance, comparing with either of the model based on single type feature. Min Li 0007, Liangliang Liu 0001, Fang-Xiang Wu, Jianxin Wang 0001 |
BIBM | 2 |
| 2019 | LncRNA-disease association prediction through combining linear and non-linear features with matrix factorization and deep learning techniquesabstractLong non-coding RNAs (lncRNAs) are the foundation for understanding mechanisms of many human diseases. Considering the limited number of known experimentally verified associations between lncRNAs and diseases, it is appealing to develop accurate and effective computational methods to identify lncRNA-disease associations. Conventional matrix factorization-based methods cannot model complicated associations between lncRNAs and diseases. In this study, we propose a novel computational framework, through combining linear and non-linear features, which is used for lncRNA-disease association prediction. In our model, a conventional matrix factorization method is applied to extract linear features between lncRNAs and diseases. Deep learning techniques (fully connected layers) are applied to extract nonlinear features between lncRNAs and diseases. Finally, linear and non-linear features are fused to improve predictive performance. Compared to previous studies, our model can take advantages of the combination of linear and non-linear features between lncRNAs and diseases, and thus can effectively identify potential lncRNA-disease associations. The results show that our method achieves state-of-the-art performance in the leave-one-out cross-validation. The source codes of our method can be found at https://github.com/CSUBioGroup/DMFLDA2. Min Zeng 0004, Chengqian Lu, Fuhao Zhang, Zhangli Lu, Fang-Xiang Wu, Yaohang Li, Min Li 0007 |
BIBM | 7 |
| 2019 | Identification of Prognostic and Heterogeneous Breast Cancer Biomarkers Based on Fusion Network and Multiple Scoring Strategies
Xingyi Li 0003, Ju Xiang, Jianxin Wang 0001, Fang-Xiang Wu, Min Li 0007 |
ICIC (2) | 5 |
| 2019 | Control principles for complex biological networksabstractNetworks have been widely used to model the structure of various biological systems. Currently, a series of approaches have been developed to construct reliable biological networks. However, the ultimate understanding of a biological system is to steer its states to the desired ones by imposing signals. The control process is dominated by the intrinsic structure and the dynamic propagation. To understand the underlying mechanisms behind the life process, the control theory can be applied to biological networks with specific target requirements. In this article, we first introduce the structural controllability of complex networks and discuss its advantages and disadvantages. Then, we review the effective control to meet the specific requirements for complex biological networks. Moreover, we summarize the existing methods for finding the unique minimum set of driver nodes via the optimal control for complex networks. Finally, we discuss the relationships between biological networks and structural controllability, effective control and optimal control. Moreover, potential applications of general control principles are pointed out. Min Li 0007, Jianxin Wang 0001, Fang-Xiang Wu |
Briefings Bioinform. | 1 |
| 2019 | SCOP: a novel scaffolding algorithm based on contig classification and optimizationabstractMOTIVATION: Scaffolding is an essential step during the de novo sequence assembly process to infer the direction and order relationships between the contigs and make the sequence assembly results more continuous and complete. However, scaffolding still faces the challenges of repetitive regions in genome, sequencing errors and uneven sequencing depth. Moreover, the accuracy of scaffolding greatly depends on the quality of contigs. Generally, the existing scaffolding methods construct a scaffold graph, and then optimize the graph by deleting spurious edges. Nevertheless, due to the wrong joints between contigs, some correct edges connecting contigs may be deleted. RESULTS: In this study, we present a novel scaffolding method SCOP, which is the first method to classify the contigs and utilize the vertices and edges to optimize the scaffold graph. Specially, SCOP employs alignment features and GC-content of paired reads to evaluate the quality of contigs (vertices), and divide the contigs into three types (True, Uncertain and Misassembled), and then optimizes the scaffold graph based on the classification of contigs together with the alignment of edges. The experiment results on the datasets of GAGE-A and GAGE-B demonstrate that SCOP performs better than 12 other competing scaffolders. AVAILABILITY AND IMPLEMENTATION: SCOP is publicly available for download at https://github.com/bioinfomaticsCSU/SCOP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Min Li 0007, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
Bioinform. | 1 |
| 2019 | BiXGBoost: a scalable, flexible boosting-based method for reconstructing gene regulatory networksabstractMOTIVATION: Reconstructing gene regulatory networks (GRNs) based on gene expression profiles is still an enormous challenge in systems biology. Random forest-based methods have been proved a kind of efficient methods to evaluate the importance of gene regulations. Nevertheless, the accuracy of traditional methods can be further improved. With time-series gene expression data, exploiting inherent time information and high order time lag are promising strategies to improve the power and accuracy of GRNs inference. RESULTS: In this study, we propose a scalable, flexible approach called BiXGBoost to reconstruct GRNs. BiXGBoost is a bidirectional-based method by considering both candidate regulatory genes and target genes for a specific gene. Moreover, BiXGBoost utilizes time information efficiently and integrates XGBoost to evaluate the feature importance. Randomization and regularization are also applied in BiXGBoost to address the over-fitting problem. The results on DREAM4 and Escherichia coli datasets show the good performance of BiXGBoost on different scale of networks. AVAILABILITY AND IMPLEMENTATION: Our Python implementation of BiXGBoost is available at https://github.com/zrq0123/BiXGBoost. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ruiqing Zheng, Min Li 0007, Xiang Chen 0029, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
Bioinform. | 2 |
| 2019 | SinNLRR: a robust subspace clustering method for cell type detection by non-negative and low-rank representationabstractMOTIVATION: The development of single-cell RNA-sequencing (scRNA-seq) provides a new perspective to study biological problems at the single-cell level. One of the key issues in scRNA-seq analysis is to resolve the heterogeneity and diversity of cells, which is to cluster the cells into several groups. However, many existing clustering methods are designed to analyze bulk RNA-seq data, it is urgent to develop the new scRNA-seq clustering methods. Moreover, the high noise in scRNA-seq data also brings a lot of challenges to computational methods. RESULTS: In this study, we propose a novel scRNA-seq cell type detection method based on similarity learning, called SinNLRR. The method is motivated by the self-expression of the cells with the same group. Specifically, we impose the non-negative and low rank structure on the similarity matrix. We apply alternating direction method of multipliers to solve the optimization problem and propose an adaptive penalty selection method to avoid the sensitivity to the parameters. The learned similarity matrix could be incorporated with spectral clustering, t-distributed stochastic neighbor embedding for visualization and Laplace score for prioritizing gene markers. In contrast to other scRNA-seq clustering methods, our method achieves more robust and accurate results on different datasets. AVAILABILITY AND IMPLEMENTATION: Our MATLAB implementation of SinNLRR is available at, https://github.com/zrq0123/SinNLRR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ruiqing Zheng, Min Li 0007, Zhenlan Liang, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
Bioinform. | 2 |
| 2019 | CSA: a web service for the complete process of ChIP-Seq analysisabstractBACKGROUND: Chromatin immunoprecipitation sequencing (ChIP-seq) is a technology that combines chromatin immunoprecipitation (ChIP) with next generation of sequencing technology (NGS) to analyze protein interactions with DNA. At present, most ChIP-seq analysis tools adopt the command line, which lacks user-friendly interfaces. Although some web services with graphical interfaces have been developed for ChIP-seq analysis, these sites cannot provide a comprehensive analysis of ChIP-seq from raw data to downstream analysis. RESULTS: In this study, we develop a web service for the whole process of ChIP-Seq Analysis (CSA), which covers mapping, quality control, peak calling, and downstream analysis. In addition, CSA provides a customization function for users to define their own workflows. And the visualization of mapping, peak calling, motif finding, and pathway analysis results are also provided in CSA. For the different types of ChIP-seq datasets, CSA can provide the corresponding tool to perform the analysis. Moreover, CSA can detect differences in ChIP signals between ChIP samples and controls to identify absolute binding sites. CONCLUSIONS: The two case studies demonstrate the effectiveness of CSA, which can complete the whole procedure of ChIP-seq analysis. CSA provides a web interface for users, and implements the visualization of every analysis step. The website of CSA is available at http://CompuBio.csu.edu.cn. Min Li 0007, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
BMC Bioinform. | 1 |
| 2019 | DeepEP: a deep learning framework for identifying essential proteinsabstractBACKGROUND: Essential proteins are crucial for cellular life and thus, identification of essential proteins is an important topic and a challenging problem for researchers. Recently lots of computational approaches have been proposed to handle this problem. However, traditional centrality methods cannot fully represent the topological features of biological networks. In addition, identifying essential proteins is an imbalanced learning problem; but few current shallow machine learning-based methods are designed to handle the imbalanced characteristics. RESULTS: We develop DeepEP based on a deep learning framework that uses the node2vec technique, multi-scale convolutional neural networks and a sampling technique to identify essential proteins. In DeepEP, the node2vec technique is applied to automatically learn topological and semantic features for each protein in protein-protein interaction (PPI) network. Gene expression profiles are treated as images and multi-scale convolutional neural networks are applied to extract their patterns. In addition, DeepEP uses a sampling method to alleviate the imbalanced characteristics. The sampling method samples the same number of the majority and minority samples in a training epoch, which is not biased to any class in training process. The experimental results show that DeepEP outperforms traditional centrality methods. Moreover, DeepEP is better than shallow machine learning-based methods. Detailed analyses show that the dense vectors which are generated by node2vec technique contribute a lot to the improved performance. It is clear that the node2vec technique effectively captures the topological and semantic properties of PPI network. The sampling method also improves the performance of identifying essential proteins. CONCLUSION: We demonstrate that DeepEP improves the prediction performance by integrating multiple deep learning techniques and a sampling method. DeepEP is more effective than existing methods. Min Zeng 0004, Min Li 0007, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001 |
BMC Bioinform. | 2 |
| 2019 | Deep learning for biological/clinical data
Fang-Xiang Wu, Min Li 0007 |
Neurocomputing | 2 |
| 2019 | Automatic ICD-9 coding via deep transfer learning
Min Zeng 0004, Min Li 0007, Zhihui Fei, Yi Pan 0001, Jianxin Wang 0001 |
Neurocomputing | 2 |
| 2019 | Automatic ICD code assignment of Chinese clinical notes based on multilayer attention BiRNN
Min Li 0007, Liangliang Liu 0001, Zhihui Fei, Fang-Xiang Wu, Jianxin Wang 0001 |
J. Biomed. Informatics | 2 |
| 2019 | Decoding the Structural Keywords in Protein Structure Universe
Wessam Elhefnawy, Min Li 0007, Jianxin Wang 0001, Yaohang Li |
J. Comput. Sci. Technol. | 2 |
| 2019 | Controllability and Its Applications to Biological Networks
Min Li 0007, Jianxin Wang 0001, Fang-Xiang Wu |
J. Comput. Sci. Technol. | 2 |
| 2019 | Automated ICD-9 Coding via A Deep Learning ApproachabstractICD-9 (the Ninth Revision of International Classification of Diseases) is widely used to describe a patient's diagnosis. Accurate automated ICD-9 coding is important because manual coding is expensive, time-consuming, and inefficient. Inspired by the recent successes of deep learning, in this study, we present a deep learning framework called DeepLabeler to automatically assign ICD-9 codes. DeepLabeler combines the convolutional neural network with the 'Document to Vector' technique to extract and encode local and global features. Our proposed DeepLabeler demonstrates its effectiveness by achieving state-of-the-art performance, i.e., 0.335 micro F-measure on MIMIC-II dataset and 0.408 micro F-measure on MIMIC-III dataset. It outperforms classical hierarchy-based SVM and flat-SVM both on these two datasets by at least 14 percent. Furthermore, we analyze the deep neural network structure to discover the vital elements in the success of DeepLabeler. We find that the convolutional neural network is the most effective component in our network and the 'Document to Vector' technique is also necessary for enhancing classification performance since it extracts well-recognized global features. Extensive experimental results demonstrate that the great promise of deep learning techniques in the field of text multi-label classification and automated medical coding. Min Li 0007, Zhihui Fei, Min Zeng 0004, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | Construction of Refined Protein Interaction Network for Predicting Essential ProteinsabstractIdentification of essential proteins based on protein interaction network (PIN) is a very important and hot topic in the post genome era. Up to now, a number of network-based essential protein discovery methods have been proposed. Generally, a static protein interaction network was constructed by using the protein-protein interactions obtained from different experiments or databases. Unfortunately, most of the network-based essential protein discovery methods are sensitive to the reliability of the constructed PIN. In this paper, we propose a new method for constructing refined PIN by using gene expression profiles and subcellular location information. The basic idea behind refining the PIN is that two proteins should have higher possibility to physically interact with each other if they appear together at the same subcellular location and are active together at least at a time point in the cell cycle. The original static PIN is denoted by S-PIN while the final PIN refined by our method is denoted by TS-PIN. To evaluate whether the constructed TS-PIN is more suitable to be used in the identification of essential proteins, 10 network-based essential protein discovery methods (DC, EC, SC, BC, CC, IC, LAC, NC, BN, and DMNC) are applied on it to identify essential proteins. A comparison of TS-PIN and two other networks: S-PIN and NF-APIN (a noise-filtered active PIN constructed by using gene expression data and S-PIN) is implemented on the prediction of essential proteins by using these ten network-based methods. The comparison results show that all of the 10 network-based methods achieve better results when being applied on TS-PIN than that being applied on S-PIN and NF-APIN. Min Li 0007, Xiaopei Chen, Jianxin Wang 0001, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | A Novel Scaffolding Algorithm Based on Contig Error Correction and Path ExtensionabstractThe sequence assembly process can be divided into three stages: contigs extension, scaffolding, and gap filling. The scaffolding method is an essential step during the process to infer the direction and sequence relationships between the contigs. However, scaffolding still faces the challenges of uneven sequencing depth, genome repetitive regions, and sequencing errors, which often leads to many false relationships between contigs. The performance of scaffolding can be improved by removing potential false conjunctions between contigs. In this study, a novel scaffolding algorithm which is on the basis of path extension Loose-Strict-Loose strategy and contig error correction, called iLSLS. iLSLS helps reduce the false relationships between contigs, and improve the accuracy of subsequent steps. iLSLS utilizes a scoring function, which estimates the correctness of candidate paths by the distribution of paired reads, and try to conduction the extension with the path which is scored the highest. What's more, iLSLS can precisely estimate the gap size. We conduct experiments on two real datasets, and the results show that LSLS strategy is efficient to increase the correctness of scaffolds, and iLSLS performs better than other scaffolding methods. Min Li 0007, Zhongxiang Liao, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | MGT-SM: A Method for Constructing Cellular Signal Transduction NetworksabstractA cellular signal transduction network is an important means to describe biological responses to environmental stimuli and exchange of biological signals. Constructing the cellular signal transduction network provides an important basis for the study of the biological activities, the mechanism of the diseases, drug targets and so on. The statistical approaches to network inference are popular in literature. Granger test has been used as an effective method for causality inference. Compared with bivariate granger tests, multivariate granger tests reduce the indirect causality and were used widely for the construction of cellular signal transduction networks. A multivariate Granger test requires that the number of time points in the time-series data is more than the number of nodes involved in the network. However, there are many real datasets with a few time points which are much less than the number of nodes in the network. In this study, we propose a new multivariate Granger test-based framework to construct cellular signal transduction network, called MGT-SM. Our MGT-SM uses SVD to compute the coefficient matrix from gene expression data and adopts Monte Carlo simulation to estimate the significance of directed edges in the constructed networks. We apply the proposed MGT-SM to Yeast Synthetic Network and MDA-MB-468, and evaluate its performance in terms of the recall and the AUC. The results show that MGT-SM achieves better results, compared with other popular methods (CGC2SPR, PGC, and DBN). Min Li 0007, Ruiqing Zheng, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | Computational Drug Repositioning with Random Walk on a Heterogeneous NetworkabstractDrug repositioning is an efficient and promising strategy to identify new indications for existing drugs, which can improve the productivity of traditional drug discovery and development. Rapid advances in high-throughput technologies have generated various types of biomedical data over the past decades, which lay the foundations for furthering the development of computational drug repositioning approaches. Although many researches have tried to improve the repositioning accuracy by integrating information from multiple sources and different levels, it is still appealing to further investigate how to efficiently exploit valuable data for drug repositioning. In this study, we propose an efficient approach, Random Walk on a Heterogeneous Network for Drug Repositioning (RWHNDR), to prioritize candidate drugs for diseases. First, an integrated heterogeneous network is constructed by combining multiple sources including drugs, drug targets, diseases and disease genes data. Then, a random walk model is developed to capture the global information of the heterogeneous network. RWHNDR takes advantage of drug targets and disease genes data more comprehensively for drug repositioning. The experiment results show that our approach can achieve better performance, compared with other state-of-the-art approaches which prioritized candidate drugs based on multi-source data. Huimin Luo, Jianxin Wang 0001, Min Li 0007, Kaijie Zhao, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | Disease Inference with Symptom Extraction and Bidirectional Recurrent Neural Network
Donglin Guo, Min Li 0007, Yaohang Li, Guihua Duan, Fang-Xiang Wu, Jianxin Wang 0001 |
BIBM | 2 |
| 2018 | A Deep Learning Framework for Identifying Essential Proteins Based on Protein-Protein Interaction Network and Gene Expression Data
Min Zeng 0004, Min Li 0007, Zhihui Fei, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001 |
BIBM | 2 |
| 2018 | Using Deep Neural Network to Predict Drug Sensitivity of Cancer Cell Lines
Yake Wang, Min Li 0007, Ruiqing Zheng, Xinghua Shi, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001 |
ICIC (2) | 2 |
| 2018 | Sprites2: Detection of Deletions Based on an Accurate Alignment Strategy
Zhen Zhang 0024, Jianxin Wang 0001, Juan Shang, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001 |
ISBRA | 5 |
| 2018 | PBMarsNet: A Multivariate Adaptive Regression Splines Based Method to Reconstruct Gene Regulatory Networks
Ruiqing Zheng, Xiang Chen 0029, Yaohang Li, Fang-Xiang Wu, Min Li 0007 |
ISBRA | 6 |
| 2018 | DyNetViewer: a Cytoscape app for dynamic network construction, analysis and visualizationabstractSummary: The molecular interactions in a cell are varying with time and surrounded environmental cues. The construction and analysis of dynamic molecular networks can elucidate dynamic cellular mechanisms of different biological functions and provide a chance to understand complex diseases at the systems level. Here, we develop DyNetViewer, a Cytoscape application that provides a range of functionalities for the construction, analysis and visualization of dynamic protein-protein interaction networks. The current version of DyNetViewer consists of four different dynamic network construction methods, twelve topological variation analysis methods and four clustering algorithms. Moreover, visualization of different topological variation of nodes and clusters over time enables users to quickly identify the most variations across many network states. Availability and implementation: DyNetViewer is freely available with tutorials at the Cytoscape (3.4+) App Store (http://apps.cytoscape.org/apps/dynetviewer). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Min Li 0007, Jie Yang 0057, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
Bioinform. | 1 |
| 2018 | Prediction of lncRNA-disease associations based on inductive matrix completionabstractMotivation: Accumulating evidences indicate that long non-coding RNAs (lncRNAs) play pivotal roles in various biological processes. Mutations and dysregulations of lncRNAs are implicated in miscellaneous human diseases. Predicting lncRNA-disease associations is beneficial to disease diagnosis as well as treatment. Although many computational methods have been developed, precisely identifying lncRNA-disease associations, especially for novel lncRNAs, remains challenging. Results: In this study, we propose a method (named SIMCLDA) for predicting potential lncRNA-disease associations based on inductive matrix completion. We compute Gaussian interaction profile kernel of lncRNAs from known lncRNA-disease interactions and functional similarity of diseases based on disease-gene and gene-gene onotology associations. Then, we extract primary feature vectors from Gaussian interaction profile kernel of lncRNAs and functional similarity of diseases by principal component analysis, respectively. For a new lncRNA, we calculate the interaction profile according to the interaction profiles of its neighbors. At last, we complete the association matrix based on the inductive matrix completion framework using the primary feature vectors from the constructed feature matrices. Computational results show that SIMCLDA can effectively predict lncRNA-disease associations with higher accuracy compared with previous methods. Furthermore, case studies show that SIMCLDA can effectively predict candidate lncRNAs for renal cancer, gastric cancer and prostate cancer. Availability and implementation: https://github.com//bioinfomaticsCSU/SIMCLDA. Supplementary information: Supplementary data are available at Bioinformatics online. Chengqian Lu, Mengyun Yang, Feng Luo 0001, Fang-Xiang Wu, Min Li 0007, Yi Pan 0001, Yaohang Li, Jianxin Wang 0001 |
Bioinform. | 5 |
| 2018 | Computational drug repositioning using low-rank matrix approximation and randomized algorithmsabstractMotivation: Computational drug repositioning is an important and efficient approach towards identifying novel treatments for diseases in drug discovery. The emergence of large-scale, heterogeneous biological and biomedical datasets has provided an unprecedented opportunity for developing computational drug repositioning methods. The drug repositioning problem can be modeled as a recommendation system that recommends novel treatments based on known drug-disease associations. The formulation under this recommendation system is matrix completion, assuming that the hidden factors contributing to drug-disease associations are highly correlated and thus the corresponding data matrix is low-rank. Under this assumption, the matrix completion algorithm fills out the unknown entries in the drug-disease matrix by constructing a low-rank matrix approximation, where new drug-disease associations having not been validated can be screened. Results: In this work, we propose a drug repositioning recommendation system (DRRS) to predict novel drug indications by integrating related data sources and validated information of drugs and diseases. Firstly, we construct a heterogeneous drug-disease interaction network by integrating drug-drug, disease-disease and drug-disease networks. The heterogeneous network is represented by a large drug-disease adjacency matrix, whose entries include drug pairs, disease pairs, known drug-disease interaction pairs and unknown drug-disease pairs. Then, we adopt a fast Singular Value Thresholding (SVT) algorithm to complete the drug-disease adjacency matrix with predicted scores for unknown drug-disease pairs. The comprehensive experimental results show that DRRS improves the prediction accuracy compared with the other state-of-the-art approaches. In addition, case studies for several selected drugs further demonstrate the practical usefulness of the proposed method. Availability and implementation: http://bioinformatics.csu.edu.cn/resources/softs/DrugRepositioning/DRRS/index.html. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Huimin Luo, Min Li 0007, Shaokai Wang, Yaohang Li, Jianxin Wang 0001 |
Bioinform. | 2 |
| 2018 | CytoCtrlAnalyser: a Cytoscape app for biomolecular network controllability analysisabstractSummary: Studying the controllability of biomolecular networks can result in profound knowledge about molecular biological systems. However, there is no comprehensive and easy-to-use platform for analyzing controllability of biomolecular networks although various algorithms for analyzing complex network controllability have been proposed recently. In this application note, we develop the CytoCtrlAnalyser which is a Cytoscape app to provide a comprehensive platform for analyzing controllability of biomolecular networks. Nine algorithms have been integrated in CytoCtrlAnalyser. With network topologies and customized control settings imported into CytoCtrlAnalyser, users can identify the steering nodes which should be actuated by input control signals for achieving different control objectives as well as investigate the importance of nodes from different perspectives in the controllability of networks. CytoCtrlAnalyser offers a tool for many promising applications, such as identification of potential drug targets or biologically important nodes in biomolecular networks. Availability and implementation: Freely available for downloading at http://apps.cytoscape.org/apps/cytoctrlanalyser. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Min Li 0007, Jianxin Wang 0001, Fang-Xiang Wu |
Bioinform. | 2 |
| 2018 | Predicting MicroRNA-Disease Associations Based on Improved MicroRNA and Disease SimilaritiesabstractMicroRNAs (miRNAs) are a type of non-coding RNAs with about ∼22nt nucleotides. Increasing evidences have shown that miRNAs play critical roles in many human diseases. The identification of human disease-related miRNAs is helpful to explore the underlying pathogenesis of diseases. More and more experimental validated associations between miRNAs and diseases have been reported in the recent studies, which provide useful information for new miRNA-disease association discovery. In this study, we propose a computational framework, KBMF-MDI, to predict the associations between miRNAs and diseases based on their similarities. The sequence and function information of miRNAs are used to measure similarity among miRNAs while the semantic and function information of disease are used to measure similarity among diseases, respectively. In addition, the kernelized Bayesian matrix factorization method is employed to infer potential miRNA-disease associations by integrating these data sources. We applied this method to 6,084 known miRNA-disease associations and utilized 5-fold cross validation to evaluate the performance. The experimental results demonstrate that our method can effectively predict unknown miRNA-disease associations. Wei Lan 0001, Jianxin Wang 0001, Min Li 0007, Jin Liu 0012, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | Classification of Alzheimer's Disease Using Whole Brain Hierarchical NetworkabstractRegions of interest (ROIs) based classification has been widely investigated for analysis of brain magnetic resonance imaging (MRI) images to assist the diagnosis of Alzheimer's disease (AD) including its early warning and developing stages, e.g., mild cognitive impairment (MCI) including MCI converted to AD (MCIc) and MCI not converted to AD (MCInc). Since an ROI representation of brain structures is obtained either by pre-definition or by adaptive parcellation, the corresponding ROI in different brains can be measured. However, due to noise and small sample size of MRI images, representations generated from single or multiple ROIs may not be sufficient to reveal the underlying anatomical differences between the groups of disease-affected patients and health controls (HC). In this paper, we employ a whole brain hierarchical network (WBHN) to represent each subject. The whole brain of each subject is divided into 90, 54, 14, and 1 regions based on Automated Anatomical Labeling (AAL) atlas. The connectivity between each pair of regions is computed in terms of Pearson's correlation coefficient and used as classification feature. Then, to reduce the dimensionality of features, we select the features with higher scores. Finally, we use multiple kernel boosting (MKBoost) algorithm to perform the classification. Our proposed method is evaluated on MRI images of 710 subjects (200 AD, 120 MCIc, 160 MCInc, and 230 HC) from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database. The experimental results show that our proposed method achieves an accuracy of 94.65 percent and an area under the receiver operating characteristic (ROC) curve (AUC) of 0.954 for AD/HC classification, an accuracy of 89.63 percent and an AUC of 0.907 for AD/MCI classification, an accuracy of 85.79 percent and an AUC of 0.826 for MCI/HC classification, and an accuracy of 72.08 percent and an AUC of 0.716 for MCIc/MCInc classification, respectively. Our results demonstrate that our proposed method is efficient and promising for clinical applications for the diagnosis of AD via MRI images. Jin Liu 0012, Min Li 0007, Wei Lan 0001, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | An interpretable model for predicting side effects of analgesics for osteoarthritisabstractOsteoarthritis (OA) is the most common type of arthritis. Analgesics are widely used in the process of the treatment of arthritis. Analgesics are particularly used by OA patients which may increase the risk of cardiovascular disease by 20% to 50% overall. In this study, we proposed an interpretable model to predict side effects of analgesics on cardiovascular disease for OA patients. One task of our study is to predict whether OA patients can use analgesics. We weighed accuracy and interpretability among state-of-the-art methods, and constructed a non-linear model by the Gradient Boosting Decision Tree technique. The AUC of the prediction model was 0.96. Another task was to select informative risk features (RFs) by our proposed model. We sought to identify risk features in literature from the biomedical. Most of the selected RFs are validated by the medical literature and some new RFs could attract the interest across the medical research. The performance of the proposed model, showed its superiority compared with well-known machine learning algorithms in terms of AUC. Liangliang Liu 0001, Jianxin Wang 0001, Min Li 0007, Fang-Xiang Wu, Hong-Dong Li, Zhihui Fei |
BIBM | 3 |
| 2017 | MEC: Misassembly error correction in contigs using a combination of paired-end reads and GC-contentsabstractThe de novo assembly aims to reconstruct the genome of the unknown species. Many algorithms have been proposed for de novo assemblies. Due to problems of repetitive regions and sequencing errors, contigs usually contain a large amount of misassemblies. Consequently, the misassembly correction of contigs is a challenging and significant work, which receives considerable attentions from researchers. In this study, we propose a novel method, called MEC, to identify and correct misassemblies in contigs. Firstly, MEC takes fragment coverage as the feature to detect the candidate misassemblies. Then, it can distinguish a large number of false positives from the candidate misassemblies based on the distribution of paired-end reads and the statistical analysis of GC-contents. We apply MEC to four real contig datasets, and carry out experiments to analyze the influence of MEC on scaffolding results, which shows that MEC can reduce misassemblies effectively and result in quantitative improvements in scaffolding quality. MEC is publicly available for download at https://github.com/bioinfomaticsCSU/MEC. Binbin Wu, Jianxin Wang 0001, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001 |
BIBM | 4 |
| 2017 | LSLS: A Novel Scaffolding Method Based on Path Extension
Min Li 0007, Zhongxiang Liao, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
ICIC (2) | 1 |
| 2017 | Construction of Protein Backbone Fragments Libraries on Large Protein Sets Using a Randomized Spectral Clustering Algorithm
Wessam Elhefnawy, Min Li 0007, Jianxin Wang 0001, Yaohang Li |
ISBRA | 2 |
| 2017 | Relating Diseases Based on Disease Module Theory
Min Li 0007, Ping Zhong 0002, Guihua Duan, Jianxin Wang 0001, Yaohang Li, Fang-Xiang Wu |
ISBRA | 2 |
| 2017 | LDAP: a web server for lncRNA-disease association predictionabstractMotivation: Increasing evidences have demonstrated that long noncoding RNAs (lncRNAs) play important roles in many human diseases. Therefore, predicting novel lncRNA-disease associations would contribute to dissect the complex mechanisms of disease pathogenesis. Some computational methods have been developed to infer lncRNA-disease associations. However, most of these methods infer lncRNA-disease associations only based on single data resource. Results: In this paper, we propose a new computational method to predict lncRNA-disease associations by integrating multiple biological data resources. Then, we implement this method as a web server for lncRNA-disease association prediction (LDAP). The input of the LDAP server is the lncRNA sequence. The LDAP predicts potential lncRNA-disease associations by using a bagging SVM classifier based on lncRNA similarity and disease similarity. Availability and Implementation: The web server is available at http://bioinformatics.csu.edu.cn/ldap Contact: [email protected]. Supplimentary Information: Supplementary data are available at Bioinformatics online. Wei Lan 0001, Min Li 0007, Kaijie Zhao, Jin Liu 0012, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001 |
Bioinform. | 2 |
| 2017 | BOSS: a novel scaffolding algorithm based on an optimized scaffold graphabstractMOTIVATION: While aiming to determine orientations and orders of fragmented contigs, scaffolding is an essential step of assembly pipelines and can make assembly results more complete. Most existing scaffolding tools adopt scaffold graph approaches. However, due to repetitive regions in genome, sequencing errors and uneven sequencing depth, constructing an accurate scaffold graph is still a challenge task. RESULTS: In this paper, we present a novel algorithm (called BOSS), which employs paired reads for scaffolding. To construct a scaffold graph, BOSS utilizes the distribution of insert size to decide whether an edge between two vertices (contigs) should be added and how an edge should be weighed. Moreover, BOSS adopts an iterative strategy to detect spurious edges whose removal can guarantee no contradictions in the scaffold graph. Based on the scaffold graph constructed, BOSS employs a heuristic algorithm to sort vertices (contigs) and then generates scaffolds. The experimental results demonstrate that BOSS produces more satisfactory scaffolds, compared with other popular scaffolding tools on real sequencing data of four genomes. AVAILABILITY AND IMPLEMENTATION: BOSS is publicly available for download at https://github.com/bioinfomaticsCSU/BOSS CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Jianxin Wang 0001, Zhen Zhang 0024, Min Li 0007, Fang-Xiang Wu |
Bioinform. | 4 |
| 2017 | VAliBS: a visual aligner for bisulfite sequencesabstractBACKGROUND: Methylation is a common modification of DNA. It has been a very important and hot topic to study the correlation between methylation and diseases in medical science. Because of the special process with bisulfite treatment, traditional mapping tools do not work well with such methylation experimental reads. Traditional aligners are not designed for mapping bisulfite-treated reads, where the un-methylated 'C's are converted to 'T's. RESULTS: In this paper, we develop a reliable and visual tool, named VAliBS, for mapping bisulfate sequences to a genome reference. VAliBS works well even on large scale data or high noise data. By comparing with other state-of-the-art tools (BisMark, BSMAP, BS-Seeker2), VAliBS can improve the accuracy of bisulfite mapping. Moreover, VAliBS is a visual tool which makes its operations more easily and the alignment results are shown with colored marks which makes it easier to be read. VAliBS provides fast and accurate mapping of bisulfite-converted reads, and a friendly window system to visualize the detail of mapping of each read. CONCLUSIONS: VAliBS works well on both simulated data and real data. It can be useful in DNA methylation research. VALiBS implements an X-Window user interface where the methylation positions are visual and the operations are friendly. Min Li 0007, Jianxin Wang 0001, Yi Pan 0001, Fang-Xiang Wu |
BMC Bioinform. | 1 |
| 2017 | ISEA: Iterative Seed-Extension Algorithm for De Novo Assembly Using Paired-End Information and Insert Size DistributionabstractThe purpose of de novo assembly is to report more contiguous, complete, and less error prone contigs. Thanks to the advent of the next generation sequencing (NGS) technologies, the cost of producing high depth reads is reduced greatly. However, due to the disadvantages of NGS, de novo assembly has to face the difficulties brought by repeat regions, error rate, and low sequencing coverage in some regions. Although many de novo algorithms have been proposed to solve these problems, the de novo assembly still remains a challenge. In this article, we developed an iterative seed-extension algorithm for de novo assembly, called ISEA. To avoid the negative impact induced by error rate, ISEA utilizes reads overlap and paired-end information to correct error reads before assemblying. During extending seeds in a De Bruijn graph, ISEA uses an elaborately designed score function based on paired-end information and the distribution of insert size to solve the repeat region problem. By employing the distribution of insert size, the score function can also reduce the influence of error reads. In scaffolding, ISEA adopts a relaxed strategy to join contigs that were terminated for low coverage during the extension. The performance of ISEA was compared with six previous popular assemblers on four real datasets. The experimental results demonstrate that ISEA can effectively obtain longer and more accurate scaffolds. Min Li 0007, Zhongxiang Liao, Jianxin Wang 0001, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | United Complex Centrality for Identification of Essential Proteins from PPI NetworksabstractEssential proteins are indispensable for the survival or reproduction of an organism. Identification of essential proteins is not only necessary for the understanding of the minimal requirements for cellular life, but also important for the disease study and drug design. With the development of high-throughput techniques, a large number of protein-protein interaction data are available, which promotes the studies of essential proteins from the network level. Up to now, though a series of computational methods have been proposed, the prediction precision still needs to be improved. In this paper, we propose a new method, United complex Centrality (UC), to identify essential proteins by integrating the protein complexes with the topological features of protein-protein interaction (PPI) networks. By analyzing the relationship between the essential proteins and the known protein complexes of S. cerevisiae and human, we find that the proteins in complexes are more likely to be essential compared with the proteins not included in any complexes and the proteins appeared in multiple complexes are more inclined to be essential compared to those only appeared in a single complex. Considering that some protein complexes generated by computational methods are inaccurate, we also provide a modified version of UC with parameter alpha, named UC-P. The experimental results show that protein complex information can help identify the essential proteins more accurate both for the PPI network of S. cerevisiae and that of human. The proposed method UC performs obviously better than the eight previously proposed methods (DC, IC, EC, SC, BC, CC, NC, and LAC) for identifying essential proteins. Min Li 0007, Zhibei Niu, Fang-Xiang Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | Predicting Protein Functions by Using Unbalanced Random Walk Algorithm on Three Biological NetworksabstractWith the gap between the sequence data and their functional annotations becomes increasing wider, many computational methods have been proposed to annotate functions for unknown proteins. However, designing effective methods to make good use of various biological resources is still a big challenge for researchers due to function diversity of proteins. In this work, we propose a new method named ThrRW, which takes several steps of random walking on three different biological networks: protein interaction network (PIN), domain co-occurrence network (DCN), and functional interrelationship network (FIN), respectively, so as to infer functional information from neighbors in the corresponding networks. With respect to the topological and structural differences of the three networks, the number of walking steps in the three networks will be different. In the course of working, the functional information will be transferred from one network to another according to the associations between the nodes in different networks. The results of experiment on S. cerevisiae data show that our method achieves better prediction performance not only than the methods that consider both PIN data and GO term similarities, but also than the methods using both PIN data and protein domain information, which verifies the effectiveness of our method on integrating multiple biological data sources. Wei Peng 0004, Min Li 0007, Lusheng Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | Protein Inference from the Integration of Tandem MS Data and Interactome NetworksabstractSince proteins are digested into a mixture of peptides in the preprocessing step of tandem mass spectrometry (MS), it is difficult to determine which specific protein a shared peptide belongs to. In recent studies, besides tandem MS data and peptide identification information, some other information is exploited to infer proteins. Different from the methods which first use only tandem MS data to infer proteins and then use network information to refine them, this study proposes a protein inference method named TMSIN, which uses interactome networks directly. As two interacting proteins should co-exist, it is reasonable to assume that if one of the interacting proteins is confidently inferred in a sample, its interacting partners should have a high probability in the same sample, too. Therefore, we can use the neighborhood information of a protein in an interactome network to adjust the probability that the shared peptide belongs to the protein. In TMSIN, a multi-weighted graph is constructed by incorporating the bipartite graph with interactome network information, where the bipartite graph is built with the peptide identification information. Based on multi-weighted graphs, TMSIN adopts an iterative workflow to infer proteins. At each iterative step, the probability that a shared peptide belongs to a specific protein is calculated by using the Bayes' law based on the neighbor protein support scores of each protein which are mapped by the shared peptides. We carried out experiments on yeast data and human data to evaluate the performance of TMSIN in terms of ROC, q-value, and accuracy. The experimental results show that AUC scores yielded by TMSIN are 0.742 and 0.874 in yeast dataset and human dataset, respectively, and TMSIN yields the maximum number of true positives when q-value less than or equal to 0.05. The overlap analysis shows that TMSIN is an effective complementary approach for protein inference. Jiancheng Zhong, Jianxin Wang 0001, Zhen Zhang 0024, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2016 | Predicting microRNA-environmental factor interactions based on bi-random walk and multi-label learningabstractIncreasing evidences have shown that microRNAs (miRNAs) play important roles in many diseases. The environmental factors (EFs) can regulate the expression level of miRNAs in human tissues. Therefore, identifying potential miRNA-environmental factor interactions is helpful not only for understanding the pathogenesis of diseases, but also for disease diagnosis, prognosis and treatment. In this paper, we propose a computational framework, MEI-BRWMLL (MiRNA-EF Interaction prediction based on Bi-Random walk and Multi-Label Learning), to identify interactions between miRNAs and environmental factors. The sequence and topology information of miRNA and structure, anatomical therapeutic chemical and topology information of environmental factor are employed to measure similarity of miRNAs and environmental factors, respectively. In addition, we use similarity network fusion method to integrate biological information of miRNAs and environmental factors, respectively. In the last, the bi-random walk and multi-label learning method are utilized to identify potential miRNA-environmental factor interactions. In order to evaluate the performance of MEI-BRWMLL, we implement the ten-fold cross validation in the experiment. The MEI-BRWMLL achieves an AUC of 0.8208. It has been shown that MEI-BRWMLL is able to identify known miRNA-environmental factor interactions. Wei Lan 0001, Jianxin Wang 0001, Min Li 0007, Chengqian Lu, Fang-Xiang Wu, Yi Pan 0001 |
BIBM | 3 |
| 2016 | Construction of the spatial and temporal active protein interaction network for identifying protein complexesabstractWith the advances in high-throughput technology, a large number of protein interactions data have been burgeoning in recent years, which makes it possible for considering dynamic properties of protein interaction networks(PINs) instead of static properties. To address the limitation of the existing dynamic PIN analysis approaches, in this paper, we proposed a new model-based scheme for the construction of the Spatial and Temporal Active Protein Interaction Network (ST-APIN) by integrating time-course gene expression data and subcellular location information. To evaluate the efficiency of ST-APIN, the commonly used classical clustering algorithm MCL was adopted to identify protein complexes from ST-APIN and other three dynamic PINs, NF-APIN, DPIN, TC-PIN. The experimental results showed that, the performance of MCL on ST-APIN outperforms those on the three other dynamic networks in terms of matching with known complexes, sensitivity, specificity and f-measure. Furthermore, we evaluated the identified protein complexes by GO (Gene Ontology) function enrichment analysis. The validation showed that the identified protein complexes from ST-APIN were more biologically significance. This study provided a general paradigm for constructing the ST-APINs, which can be used for theoretical studies and clinic applications. Xiangmao Meng, Min Li 0007, Jianxin Wang 0001, Fang-Xiang Wu, Yi Pan 0001 |
BIBM | 2 |
| 2016 | The MSS of complex networks with centrality based preference and its application to biomolecular networksabstractNetworks are employed to represent many real world complex systems. For biological systems, biomolecules interact with each other to form so-called biomolecular networks. The explorations on the connections between structural control theory and biological networks have uncovered some interesting biological phenomena. Recently, some studies have paid attentions to the structural controllability of networks in notion of the minimum steering sets (MSSs). However, the MSSs for a complex network are not unique. Therefore, it is meaningful to find out the most special one with some centrality-based preference. The MSS of a network which has the maximum (minimum) average value of a certain centrality among all possible MSSs of the network can be identified by our method. Then we apply the method to the human liver metabolic network and find that centralities of steering nodes in different MSSs can be remarkably different. In addition, we observe that, for some centralities, the liver cancer reactions are significantly enriched in the MSSs with the minimum average centrality value. This result suggests that when investigating the controllability of biomolecular networks, the centralities, which could provide more meaningful biological information, can be taken into consideration. Lingkai Tang, Min Li 0007, Jianxin Wang 0001, Fang-Xiang Wu |
BIBM | 3 |
| 2016 | Identifying Essential Proteins by Purifying Protein Interaction Networks
Min Li 0007, Xiaopei Chen, Jianxin Wang 0001, Yi Pan 0001 |
ISBRA | 1 |
| 2016 | Drug repositioning based on comprehensive similarity measures and Bi-Random walk algorithmabstractMOTIVATION: Drug repositioning, which aims to identify new indications for existing drugs, offers a promising alternative to reduce the total time and cost of traditional drug development. Many computational strategies for drug repositioning have been proposed, which are based on similarities among drugs and diseases. Current studies typically use either only drug-related properties (e.g. chemical structures) or only disease-related properties (e.g. phenotypes) to calculate drug or disease similarity, respectively, while not taking into account the influence of known drug-disease association information on the similarity measures. RESULTS: In this article, based on the assumption that similar drugs are normally associated with similar diseases and vice versa, we propose a novel computational method named MBiRW, which utilizes some comprehensive similarity measures and Bi-Random walk (BiRW) algorithm to identify potential novel indications for a given drug. By integrating drug or disease features information with known drug-disease associations, the comprehensive similarity measures are firstly developed to calculate similarity for drugs and diseases. Then drug similarity network and disease similarity network are constructed, and they are incorporated into a heterogeneous network with known drug-disease interactions. Based on the drug-disease heterogeneous network, BiRW algorithm is adopted to predict novel potential drug-disease associations. Computational experiment results from various datasets demonstrate that the proposed approach has reliable prediction performance and outperforms several recent computational drug repositioning approaches. Moreover, case studies of five selected drugs further confirm the superior performance of our method to discover potential indications for drugs practically. AVAILABILITY AND IMPLEMENTATION: http://github.com//bioinfomaticsCSU/MBiRW CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huimin Luo, Jianxin Wang 0001, Min Li 0007, Xiaoqing Peng, Fang-Xiang Wu, Yi Pan 0001 |
Bioinform. | 3 |
| 2016 | Predicting essential proteins based on subcellular localization, orthology and PPI networksabstractBACKGROUND: Essential proteins play an indispensable role in the cellular survival and development. There have been a series of biological experimental methods for finding essential proteins; however they are time-consuming, expensive and inefficient. In order to overcome the shortcomings of biological experimental methods, many computational methods have been proposed to predict essential proteins. The computational methods can be roughly divided into two categories, the topology-based methods and the sequence-based ones. The former use the topological features of protein-protein interaction (PPI) networks while the latter use the sequence features of proteins to predict essential proteins. Nevertheless, it is still challenging to improve the prediction accuracy of the computational methods. RESULTS: Comparing with nonessential proteins, essential proteins appear more frequently in certain subcellular locations and their evolution more conservative. By integrating the information of subcellular localization, orthologous proteins and PPI networks, we propose a novel essential protein prediction method, named SON, in this study. The experimental results on S.cerevisiae data show that the prediction accuracy of SON clearly exceeds that of nine competing methods: DC, BC, IC, CC, SC, EC, NC, PeC and ION. CONCLUSIONS: We demonstrate that, by integrating the information of subcellular localization, orthologous proteins with PPI networks, the accuracy of predicting essential proteins can be improved. Our proposed method SON is effective for predicting essential proteins. Gaoshi Li, Min Li 0007, Jianxin Wang 0001, Jingli Wu, Fang-Xiang Wu, Yi Pan 0001 |
BMC Bioinform. | 2 |
| 2016 | FLEXc: protein flexibility prediction using context-based statistics, predicted structural features, and sequence informationabstractBACKGROUND: The fluctuation of atoms around their average positions in protein structures provides important information regarding protein dynamics. This flexibility of protein structures is associated with various biological processes. Predicting flexibility of residues from protein sequences is significant for analyzing the dynamic properties of proteins which will be helpful in predicting their functions. RESULTS: In this paper, an approach of improving the accuracy of protein flexibility prediction is introduced. A neural network method for predicting flexibility in 3 states is implemented. The method incorporates sequence and evolutionary information, context-based scores, predicted secondary structures and solvent accessibility, and amino acid properties. Context-based statistical scores are derived, using the mean-field potentials approach, for describing the different preferences of protein residues in flexibility states taking into consideration their amino acid context. The 7-fold cross validated accuracy reached 61 % when context-based scores and predicted structural states are incorporated in the training process of the flexibility predictor. CONCLUSIONS: Incorporating context-based statistical scores with predicted structural states are important features to improve the performance of predicting protein flexibility, as shown by our computational results. Our prediction method is implemented as web service called "FLEXc" and available online at: http://hpcr.cs.odu.edu/flexc . Ashraf Yaseen, Mais Nijim, Brandon Williams, Min Li 0007, Jianxin Wang 0001, Yaohang Li |
BMC Bioinform. | 5 |
| 2016 | Predicting drug-target interaction using positive-unlabeled learning
Wei Lan 0001, Jianxin Wang 0001, Min Li 0007, Jin Liu 0012, Yaohang Li, Fang-Xiang Wu, Yi Pan 0001 |
Neurocomputing | 3 |
| 2015 | A two-step logistic regression algorithm for identifying individual-cancer-related genesabstractThe identification of cancer-related genes is important towards the understanding of complex genetic diseases. Although many machine learning algorithms are proposed to identify disease-related genes, they often either have poor performance to identify locus heterogeneity cancer-related genes or are not applicable to predict individual-disease-related genes due to the lack of positive instances (imbalanced classification). To overcome these two issues, a two-step logistic regression (LR) based algorithm is proposed in this study for identifying individual-cancer-related genes. A set of high potential cancer-class-related genes is first generated in step 1, followed by a second round of LR-based algorithm conducted on this smaller dataset for identifying individual-cancer-related genes. Numerical experiments show that the proposed two-step LR-based algorithm not only works well for locus heterogeneity data, but also has good performance to handle the imbalanced classification problem. The individual-cancer-related gene identification experiments achieve AUC values of around 0.85 when the threshold of posterior probability is chosen between 0.3 and 0.6. All evaluations are conducted by using the leave-one-out cross validation method. Xuequn Shang 0001, Min Li 0007, Jianxin Wang 0001, Fang-Xiang Wu |
BIBM | 3 |
| 2015 | Predicting microRNA-disease associations by integrating multiple biological informationabstractMicroRNAs (miRNAs) are a set of small non-coding RNAs that play critical roles in many human diseases. Identifying potential miRNA-disease association is helpful to explore the underlying molecular mechanisms of disease. Currently, it is expensive and time-consuming to detect miRNA-disease associations with experimental methods. On the other hand, many known associations between miRNAs and diseases provide useful information for new miRNA-disease interaction discovery. In this study, we propose a computational framework to infer the relationship between miRNA and disease by integrating multiple data resources. We use sequence and function information of miRNA and semantic and function information of disease to measure similarity of miRNA and disease, respectively. In addition, kernelized Bayesian matrix factorization method is employed to infer potential miRNA-disease association by integrating these data resources. The experimental results demonstrate that our method can effectively predict unknown miRNA-disease association. Wei Lan 0001, Jianxin Wang 0001, Min Li 0007, Jin Liu 0012, Yi Pan 0001 |
BIBM | 3 |
| 2015 | EPGA2: memory-efficient de novo assemblerabstractMOTIVATION: In genome assembly, as coverage of sequencing and genome size growing, most current softwares require a large memory for handling a great deal of sequence data. However, most researchers usually cannot meet the requirements of computing resources which prevent most current softwares from practical applications. RESULTS: In this article, we present an update algorithm called EPGA2, which applies some new modules and can bring about improved assembly results in small memory. For reducing peak memory in genome assembly, EPGA2 adopts memory-efficient DSK to count K-mers and revised BCALM to construct De Bruijn Graph. Moreover, EPGA2 parallels the step of Contigs Merging and adds Errors Correction in its pipeline. Our experiments demonstrate that all these changes in EPGA2 are more useful for genome assembly. AVAILABILITY AND IMPLEMENTATION: EPGA2 is publicly available for download at https://github.com/bioinfomaticsCSU/EPGA2. Jianxin Wang 0001, Zhen Zhang 0024, Fang-Xiang Wu, Min Li 0007, Yi Pan 0001 |
Bioinform. | 6 |
| 2015 | EPGA: de novo assembly using the distributions of reads and insert sizeabstractMOTIVATION: In genome assembly, the primary issue is how to determine upstream and downstream sequence regions of sequence seeds for constructing long contigs or scaffolds. When extending one sequence seed, repetitive regions in the genome always cause multiple feasible extension candidates which increase the difficulty of genome assembly. The universally accepted solution is choosing one based on read overlaps and paired-end (mate-pair) reads. However, this solution faces difficulties with regard to some complex repetitive regions. In addition, sequencing errors may produce false repetitive regions and uneven sequencing depth leads some sequence regions to have too few or too many reads. All the aforementioned problems prohibit existing assemblers from getting satisfactory assembly results. RESULTS: In this article, we develop an algorithm, called extract paths for genome assembly (EPGA), which extracts paths from De Bruijn graph for genome assembly. EPGA uses a new score function to evaluate extension candidates based on the distributions of reads and insert size. The distribution of reads can solve problems caused by sequencing errors and short repetitive regions. Through assessing the variation of the distribution of insert size, EPGA can solve problems introduced by some complex repetitive regions. For solving uneven sequencing depth, EPGA uses relative mapping to evaluate extension candidates. On real datasets, we compare the performance of EPGA and other popular assemblers. The experimental results demonstrate that EPGA can effectively obtain longer and more accurate contigs and scaffolds. Jianxin Wang 0001, Zhen Zhang 0024, Fang-Xiang Wu, Min Li 0007, Yi Pan 0001 |
Bioinform. | 5 |
| 2015 | Re-alignment of the unmapped reads with base quality scoreabstractMOTIVATION: Based on the next generation genome sequencing technologies, a variety of biological applications are developed, while alignment is the first step once the sequencing reads are obtained. In recent years, many software tools have been developed to efficiently and accurately align short reads to the reference genome. However, there are still many reads that can't be mapped to the reference genome, due to the exceeding of allowable mismatches. Moreover, besides the unmapped reads, the reads with low mapping qualities are also excluded from the downstream analysis, such as variance calling. If we can take advantages of the confident segments of these reads, not only can the alignment rates be improved, but also more information will be provided for the downstream analysis. RESULTS: This paper proposes a method, called RAUR (Re-align the Unmapped Reads), to re-align the reads that can not be mapped by alignment tools. Firstly, it takes advantages of the base quality scores (reported by the sequencer) to figure out the most confident and informative segments of the unmapped reads by controlling the number of possible mismatches in the alignment. Then, combined with an alignment tool, RAUR re-align these segments of the reads. We run RAUR on both simulated data and real data with different read lengths. The results show that many reads which fail to be aligned by the most popular alignment tools (BWA and Bowtie2) can be correctly re-aligned by RAUR, with a similar Precision. Even compared with the BWA-MEM and the local mode of Bowtie2, which perform local alignment for long reads to improve the alignment rate, RAUR also shows advantages on the Alignment rate and Precision in some cases. Therefore, the trimming strategy used in RAUR is useful to improve the Alignment rate of alignment tools for the next-generation genome sequencing. AVAILABILITY: All source code are available at http://netlab.csu.edu.cn/bioinformatics/RAUR.html. Xiaoqing Peng, Jianxin Wang 0001, Zhen Zhang 0024, Qianghua Xiao, Min Li 0007, Yi Pan 0001 |
BMC Bioinform. | 5 |
| 2015 | A Topology Potential-Based Method for Identifying Essential Proteins from PPI NetworksabstractEssential proteins are indispensable for cellular life. It is of great significance to identify essential proteins that can help us understand the minimal requirements for cellular life and is also very important for drug design. However, identification of essential proteins based on experimental approaches are typically time-consuming and expensive. With the development of high-throughput technology in the post-genomic era, more and more protein-protein interaction data can be obtained, which make it possible to study essential proteins from the network level. There have been a series of computational approaches proposed for predicting essential proteins based on network topologies. Most of these topology based essential protein discovery methods were to use network centralities. In this paper, we investigate the essential proteins' topological characters from a completely new perspective. To our knowledge it is the first time that topology potential is used to identify essential proteins from a protein-protein interaction (PPI) network. The basic idea is that each protein in the network can be viewed as a material particle which creates a potential field around itself and the interaction of all proteins forms a topological field over the network. By defining and computing the value of each protein's topology potential, we can obtain a more precise ranking which reflects the importance of proteins from the PPI network. The experimental results show that topology potential-based methods TP and TP-NC outperform traditional topology measures: degree centrality (DC), betweenness centrality (BC), closeness centrality (CC), subgraph centrality (SC), eigenvector centrality (EC), information centrality (IC), and network centrality (NC) for predicting essential proteins. In addition, these centrality measures are improved on their performance for identifying essential proteins in biological network when controlled by topology potential. Min Li 0007, Jianxin Wang 0001, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2015 | ClusterViz: A Cytoscape APP for Cluster Analysis of Biological NetworkabstractCluster analysis of biological networks is one of the most important approaches for identifying functional modules and predicting protein functions. Furthermore, visualization of clustering results is crucial to uncover the structure of biological networks. In this paper, ClusterViz, an APP of Cytoscape 3 for cluster analysis and visualization, has been developed. In order to reduce complexity and enable extendibility for ClusterViz, we designed the architecture of ClusterViz based on the framework of Open Services Gateway Initiative. According to the architecture, the implementation of ClusterViz is partitioned into three modules including interface of ClusterViz, clustering algorithms and visualization and export. ClusterViz fascinates the comparison of the results of different algorithms to do further related analysis. Three commonly used clustering algorithms, FAG-EC, EAGLE and MCODE, are included in the current version. Due to adopting the abstract interface of algorithms in module of the clustering algorithms, more clustering algorithms can be included for the future use. To illustrate usability of ClusterViz, we provided three examples with detailed steps from the important scientific articles, which show that our tool has helped several research teams do their research work on the mechanism of the biological networks. Jianxin Wang 0001, Jiancheng Zhong, Gang Chen 0010, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2014 | A logistic regression based algorithm for identifying human disease genesabstractThe identification of disease genes is the first step towards the understanding of genetic disease mechanisms. Although many computational algorithms are proposed to identify disease genes, they either have poor performance in terms of AUC scores or are very time consuming. To overcome these two problems, a logistic regression based algorithm is proposed in this study for identifying disease genes. The issue of disease gene identification is formulated as a two-class classification problem, where one class represents those disease genes, while the other class represents non-disease genes. A binary logistic regression is employed to predict the posterior probability of a gene associated with disease by taking prior labels as the categorical dependent variables and label related feature vectors as predictor variables. Numerical experiments show that the proposed logistic regression based algorithm not only have a very good performance, but also significantly reduce the computing time. The AUC score is 0.737 when no prior information is used and it increases to 0.766 when protein complex data are integrated. Averagely, the proposed algorithm only takes 1.31% and 37.35% running time of the existing MRF method and RWR algorithm, respectively, when generating one prediction in the leave-one-out cross validation method. Min Li 0007, Jianxin Wang 0001, Fang-Xiang Wu |
BIBM | 2 |
| 2014 | Searching SNP Combinations Related to Evolutionary Information of Human Populations on HapMap Data
Haihua Gu, Zhen Zhang 0024, Min Li 0007, Fang-Xiang Wu |
ISBRA | 4 |
| 2014 | Identification of Essential Proteins by Using Complexes and Interaction Network
Min Li 0007, Zhibei Niu, Fang-Xiang Wu, Yi Pan 0001 |
ISBRA | 1 |
| 2014 | Drug Target Identification Based on Structural Output Controllability of Complex Networks
Min Li 0007, Fang-Xiang Wu |
ISBRA | 3 |
| 2014 | Detecting Protein Complexes Basedon Uncertain Graph ModelabstractAdvanced biological technologies are producing large-scale protein-protein interaction (PPI) data at an ever increasing pace, which enable us to identify protein complexes from PPI networks. Pair-wise protein interactions can be modeled as a graph, where vertices represent proteins and edges represent PPIs. However most of current algorithms detect protein complexes based on deterministic graphs, whose edges are either present or absent. Neighboring information is neglected in these methods. Based on the uncertain graph model, we propose the concept of expected density to assess the density degree of a subgraph, the concept of relative degree to describe the relationship between a protein and a subgraph in a PPI network. We develop an algorithm called DCU (detecting complex based on uncertain graph model) to detect complexes from PPI networks. In our method, the expected density combined with the relative degree is used to determine whether a subgraph represents a complex with high cohesion and low coupling. We apply our method and the existing competing algorithms to two yeast PPI networks. Experimental results indicate that our method performs significantly better than the state-of-the-art methods and the proposed model can provide more insights for future study in PPI networks. Jianxin Wang 0001, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2013 | Prioritization of candidate genes based on disease similarity and protein's proximity in PPI networksabstractIdentifying the genes causing a genetic disease is a key challenge in human health. Recently molecular interaction data has been used to prioritize candidate genes with respect to a particular disease. As a result, different methods have been implemented to rank genes which cause a given disease. However it has been suggested in literature that, to prioritize candidate genes it is necessary to consider disease similarity along with the protein's proximity to disease genes in a protein-protein interaction (PPI) network. This paper proposes a new algorithm called proximity disease similarity algorithm (ProSim) which considers both properties simultaneously. Prostate cancer, Alzheimer disease and diabetes mellitus type 2 case studies are then used to test the proposed method. Results in terms of leave-one-out cross validation and ROC curves indicate that the proposed approach outperforms existing methods. Gamage Upeksha Ganegoda, Jianxin Wang 0001, Fang-Xiang Wu, Min Li 0007 |
BIBM | 4 |
| 2013 | Identifying dynamic protein complexes based on gene expression profiles and PPI networksabstractSummary form only given. Identification of protein complexes from protein-protein interaction network has become a key problem for understanding cellular life in post-genomic era. Many computational methods have been proposed for identifying protein complexes. Up to now, the existing computational methods are mostly applied on static PPI networks. However, proteins and their interactions are dynamic in reality. Identifying dynamic protein complexes is more meaningful and challenging. In this paper, a novel algorithm, named DPC, is proposed to identify dynamic protein complexes by integrating PPI data and gene expression profiles. Not only is the topological characters but also dynamic meaning considered in DPC. The protein complexes produced by our algorithm DPC contain two parts: static core expressed in all the molecular cycle and dynamic attachments short-lived. According to core-attachment assumption, these proteins which are always active in the molecular cycle are regarded as core proteins. The protein-complex cores are identified from these always active proteins by detecting dense sub-graphs. All possible protein complexes are extended from the protein-complex cores by adding attachments based on a topological character of “closeness”. Others which not belong to always active proteins are considered as potential attachments. On a certain time course, an attachment protein can only participate in one protein complex. Based on this idea, we first find a best protein-complex core for each potential attachment. It means that if a protein would be active at the some time, it would be added into the best protein-complex core for forming protein complexes. According to the formation and function of a protein complex, it should be active in two or more continual time courses. Based on the above analysis, we use the following rules to filter false positive complexes: 1) A protein complex should include at least two proteins; 2) The attachment proteins should be active in the same time course or in different but adjacent time courses; 3) If the attachments of a possible protein complex do not satisfy the second rule and the protein-complex core involves at least two proteins, the core will be kept as a final protein complex. So final protein complexes are extended from the protein-complex cores by adding attachments based on a topological character of “closeness” and dynamic meaning. The protein complexes produced by our algorithm DPC contain two parts: static core expressed in all the molecular cycle and dynamic attachments short-lived. The proposed algorithm DPC was applied on the data of Scaccharomves cerevisiae and the experimental results show that DPC outperforms CMC, MCL, SPICi, HC-PIN, COACH and Core-Attachment based on the validation of matching with known complexes and hF-measures. Min Li 0007, Jianxin Wang 0001, Fang-Xiang Wu, Yi Pan 0001 |
BIBM | 1 |
| 2013 | A new method for predicting essential proteins based on topology potentialabstractEssential proteins are indispensable for cellular life. It is of great significance to identify essential proteins that can help us understand the minimal requirements for cellular life and is also very important for drug design. However, identification of essential proteins based on experimental approaches are always time-consuming and expensive. With the development of high-throughput technology in the post-genomic era, more and more protein-protein interaction data can be obtained, which make us study essential proteins from the network level become possible. There have been a series of computational approaches proposed for predicting essential proteins based on network topologies. Most of these topology based essential protein discovery methods were to use network centrality. In this paper, we investigate the essential proteins' topological characters from a completely new perspective. To our knowledge it is the first time that topology potential is used to identify essential proteins from protein-protein interaction network. The basic idea is that each protein in the network can be viewed as a material particle which creates a potential field around itself and the interaction of all proteins forms a topological field over the network. By defining and computing the value of each protein's topology potential, we can obtain a more precise ranking which reflects the importance of proteins from the protein-protein interaction network. The experiment results show that topology potential outperforms traditional topology measures: Degree Centrality (DC), Betweenness Centrality (BC), Closeness Centrality (CC), Subgraph Centrality(SC), Eigenvector Centrality(EC), Information Centrality(IC), and Sum of ECC (NC) for predicting essential proteins. In addition, these centrality measures are improved on their performance for identifying essential proteins in biological network when controlled by topology potential. Min Li 0007, Yi Pan 0001, Jianxin Wang 0001 |
BIBM | 2 |
| 2013 | A clustering algorithm for identifying hierarchical and overlapping protein complexes in large PPI networksabstractWith the development of high-throughput technique, protein-protein interactions (PPIs) are increasing fast and available conveniently, which make it possible to identify protein complexes in PPI network[1-7]. Many evidences have demonstrated that protein complexes are overlapping and hierarchically organized in PPI networks[8-9], which requires protein complex detection methods can identify both overlapping and hierarchical protein complexes in a PPI network. Meanwhile, the large size of PPI network requires protein complex detection methods based on PPI network running fast. Up to now, few methods can achieve all above requirements. Jianxin Wang 0001, Min Li 0007 |
BIBM | 3 |
| 2013 | A novel algorithm for mining protein complex from the weighted networkabstractThe vast amount of genes and proteins that participate in biological networks imposes the need for determination of protein complexes within the network in order to reduce the complexity, while these complexes will be the first step in deciphering the composite genetic or cellular interactions of the overall network. Xiwei Tang, Jianxin Wang 0001, Min Li 0007, Yi Pan 0001 |
BIBM | 3 |
| 2012 | Identifying protein complexes based on local fitness methodabstractIdentifying protein complexes from a PPI network is crucial to understand principles of cellular organization and functional mechanisms. However, it is still a difficult task because protein complexes have various topologies in PPI networks. In the paper, a novel protein complex identifying method, named LF-PIN, is proposed based on local fitness method. Firstly, LF-PIN calculates each PPI's weight based on its clustering value in the PPI network and selects seed edges by the edge weight. Then, protein complexes are extended from seed edges based on the evaluation of their neighbors' fitness values until their fitness reach the local maximum value. We apply the proposed algorithm LF-PIN and other nine previous algorithms, including HC-PIN, NFC, MCODE, DPClus, IPCA, CPM, MCL, CMC and Core-Attachment, to the PPI network of S.cerevisiae and compare their performances. Experimental results show that LF-PIN outperforms other competing algorithms in terms of matching with known complexes and functional enrichment. Jianxin Wang 0001, Min Li 0007 |
BIBM | 3 |
| 2012 | Towards the identification of protein complexes and functional modules by integrating PPI network and gene expression dataabstractBACKGROUND: Identification of protein complexes and functional modules from protein-protein interaction (PPI) networks is crucial to understanding the principles of cellular organization and predicting protein functions. In the past few years, many computational methods have been proposed. However, most of them considered the PPI networks as static graphs and overlooked the dynamics inherent within these networks. Moreover, few of them can distinguish between protein complexes and functional modules. RESULTS: In this paper, a new framework is proposed to distinguish between protein complexes and functional modules by integrating gene expression data into protein-protein interaction (PPI) data. A series of time-sequenced subnetworks (TSNs) is constructed according to the time that the interactions were activated. The algorithm TSN-PCD was then developed to identify protein complexes from these TSNs. As protein complexes are significantly related to functional modules, a new algorithm DFM-CIN is proposed to discover functional modules based on the identified complexes. The experimental results show that the combination of temporal gene expression data with PPI data contributes to identifying protein complexes more precisely. A quantitative comparison based on f-measure reveals that our algorithm TSN-PCD outperforms the other previous protein complex discovery algorithms. Furthermore, we evaluate the identified functional modules by using "Biological Process" annotated in GO (Gene Ontology). The validation shows that the identified functional modules are statistically significant in terms of "Biological Process". More importantly, the relationship between protein complexes and functional modules are studied. CONCLUSIONS: The proposed framework based on the integration of PPI data and gene expression data makes it possible to identify protein complexes and functional modules more effectively. Moveover, the proposed new framework and algorithms can distinguish between protein complexes and functional modules. Our findings suggest that functional modules are closely related to protein complexes and a functional module may consist of one or multiple protein complexes. The program is available at http://netlab.csu.edu.cn/bioinfomatics/limin/DFM-CIN/index.html. Min Li 0007, Jianxin Wang 0001, Yi Pan 0001 |
BMC Bioinform. | 1 |
| 2012 | Identification of Essential Proteins Based on Edge Clustering CoefficientabstractIdentification of essential proteins is key to understanding the minimal requirements for cellular life and important for drug design. The rapid increase of available protein-protein interaction (PPI) data has made it possible to detect protein essentiality on network level. A series of centrality measures have been proposed to discover essential proteins based on network topology. However, most of them tended to focus only on the location of single protein, but ignored the relevance between interactions and protein essentiality. In this paper, a new centrality measure for identifying essential proteins based on edge clustering coefficient, named as NC, is proposed. Different from previous centrality measures, NC considers both the centrality of a node and the relationship between it and its neighbors. For each interaction in the network, we calculate its edge clustering coefficient. A node’s essentiality is determined by the sum of the edge clustering coefficients of interactions connecting it and its neighbors. The new centrality measure NC takes into account the modular nature of protein essentiality. NC is applied to three different types of yeast protein-protein interaction networks, which are obtained from the DIP database, the MIPS database and the BioGRID database, respectively. The experimental results on the three different networks show that the number of essential proteins discovered by NC universally exceeds that discovered by the six other centrality measures: DC, BC, CC, SC, EC, and IC. Moreover, the essential proteins discovered by NC show significant cluster effect. Jianxin Wang 0001, Min Li 0007, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2011 | Essential Protein Discovery Based on Network Motif and Gene OntologyabstractEssential proteins are indispensable to support cellular life and constitute a minimal set required for a living cell. Fast progress in high-throughput technologies and large amount of data enable to discover essential proteins in system level by analyzing protein-protein interaction networks. A number of centrality algorithms are suggested to detect essential proteins, but they focus only on network structures. In this paper, we develop a new centrality algorithm, named MCGO which uses network motifs for centrality measure in the graph pruned by EDGEGO. EDGEGO algorithm utilizes Gene Ontology(GO) to trim a number of uninformative edges from the network. We compare the performance of our algorithm with DC (degree centrality) and SoECC (sum of edge clustering coefficient) against various evaluation measures. Experimental results applied to an yeast protein-protein interaction network downloaded from DIP database show that MCGO performs significantly better than DC and SoECC. We also show that DC and SoECC improve greatly when EDGEGO is applied to them. Min Li 0007, Jianxin Wang 0001, Yi Pan 0001 |
BIBM | 2 |
| 2011 | A New Measurement for Evaluating Clusters in Protein Interaction NetworksabstractClustering of protein-protein interaction networks is one of the most prevalent methods for identifying protein complexes and functional modules, which is crucial to understanding the principles of cellular organization and prediction of protein functions. In the past few years, many computational methods have been proposed. However, it is always a challenging task to evaluate how well the clusters are identified. Even for the most popular measurements, F-measure and Pvalue, bias exists for evaluating the identified clusters. In this paper, we propose a new measurement, named hF-measure, to evaluate clusters more finely and distinctly. First, we defined the hierarchical consistency and the hierarchical similarity. Then, we propose a new hierarchical measurement of hF-measure by taking into account the hierarchical organization of functional annotations and the functional similarities among proteins. The new measurement hF-measure can discriminate between different types of errors which cannot be distinguished by F-measure. The experimental results based on Gene Ontology (GO) and yeast functional modules show that hF-measure evaluates clusters more accurately when compared to F-measure. Min Li 0007, Jianxin Wang 0001, Yi Pan 0001 |
BIBM | 1 |
| 2011 | Active Protein Interaction Network and Its Application on Protein Complex DetectionabstractIn recent years, more and more attentions are focused on modelling and analyzing dynamic network. Some researchers attempted to extract dynamic network by combining the dynamic information from gene expression data or subcellular localization data with protein network. However, the dynamics of proteins' presence does not guarantee the dynamics of interactions, since the presence of a protein does not indicate the protein's activity. The activity of a protein is closely connected with its function. Thus only the dynamics of proteins activity ensure the dynamics of interaction. The gene expression of a cellular process or cycle carries more information than only the dynamics of proteins' presence. We assume that a protein is active when its expression values are near its maximum expression value, since the expression quantity will decrease after it has performed its function that leads a feedback for controlling the expression quantity. In this paper, we proposed a method to identify active time points for each protein in a cellular process or cycle by using a 3-sigma principle to compute an active threshold for each gene according to the characteristics of its expression curve. Combined the activity information and protein interaction network, we can construct an active protein interaction network (APPI). To demonstrate the efficiency of APPI network model, we applied it on complex detection. Compared with single threshold time series networks, APPI network achieves a better performance on protein complex prediction. Jianxin Wang 0001, Xiaoqing Peng, Min Li 0007, Yi Pan 0001 |
BIBM | 3 |
| 2011 | Prediction of Essential Proteins by Integration of PPI Network Topology and Protein Complexes Information
Jianxin Wang 0001, Min Li 0007 |
ISBRA | 3 |
| 2011 | A New Method for Identifying Essential Proteins Based on Edge Clustering Coefficient
Min Li 0007, Jianxin Wang 0001, Yi Pan 0001 |
ISBRA | 2 |
| 2011 | A comparison of the functional modules identified from time course and static PPI network dataabstractBACKGROUND: Cellular systems are highly dynamic and responsive to cues from the environment. Cellular function and response patterns to external stimuli are regulated by biological networks. A protein-protein interaction (PPI) network with static connectivity is dynamic in the sense that the nodes implement so-called functional activities that evolve in time. The shift from static to dynamic network analysis is essential for further understanding of molecular systems. RESULTS: In this paper, Time Course Protein Interaction Networks (TC-PINs) are reconstructed by incorporating time series gene expression into PPI networks. Then, a clustering algorithm is used to create functional modules from three kinds of networks: the TC-PINs, a static PPI network and a pseudorandom network. For the functional modules from the TC-PINs, repetitive modules and modules contained within bigger modules are removed. Finally, matching and GO enrichment analyses are performed to compare the functional modules detected from those networks. CONCLUSIONS: The comparative analyses show that the functional modules from the TC-PINs have much more significant biological meaning than those from static PPI networks. Moreover, it implies that many studies on static PPI networks can be done on the TC-PINs and accordingly, the experimental results are much more satisfactory. The 36 PPI networks corresponding to 36 time points, identified as part of this study, and other materials are available at http://bioinfo.csu.edu.cn/txw/TC-PINs. Xiwei Tang, Jianxin Wang 0001, Min Li 0007, Gang Chen 0010, Yi Pan 0001 |
BMC Bioinform. | 4 |
| 2011 | A Fast Hierarchical Clustering Algorithm for Functional Modules Discovery in Protein Interaction NetworksabstractAs advances in the technologies of predicting protein interactions, huge data sets portrayed as networks have been available. Identification of functional modules from such networks is crucial for understanding principles of cellular organization and functions. However, protein interaction data produced by high-throughput experiments are generally associated with high false positives, which makes it difficult to identify functional modules accurately. In this paper, we propose a fast hierarchical clustering algorithm HC-PIN based on the local metric of edge clustering value which can be used both in the unweighted network and in the weighted network. The proposed algorithm HC-PIN is applied to the yeast protein interaction network, and the identified modules are validated by all the three types of Gene Ontology (GO) Terms: Biological Process, Molecular Function, and Cellular Component. The experimental results show that HC-PIN is not only robust to false positives, but also can discover the functional modules with low density. The identified modules are statistically significant in terms of three types of GO annotations. Moreover, HC-PIN can uncover the hierarchical organization of functional modules with the variation of its parameter's value, which is approximatively corresponding to the hierarchical structure of GO annotations. Compared to other previous competing algorithms, our algorithm HC-PIN is faster and more accurate. Jianxin Wang 0001, Min Li 0007, Jianer Chen, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2010 | Essential Proteins Discovery from Weighted Protein Interaction Networks
Min Li 0007, Jianxin Wang 0001, Yi Pan 0001 |
ISBRA | 1 |
| 2010 | An Agglomerate Algorithm for Mining Overlapping and Hierarchical Functional Modules in Protein Interaction Networks
Jianxin Wang 0001, Jianer Chen, Min Li 0007, Gang Chen 0010 |
ISBRA | 4 |
| 2009 | Hierarchical Organization of Functional Modules in Weighted Protein Interaction Networks Using Clustering Coefficient
Min Li 0007, Jianxin Wang 0001, Jianer Chen, Yi Pan 0001 |
ISBRA | 1 |
| 2008 | A Graph-Theoretic Method for Mining Overlapping Functional Modules in Protein Interaction Networks
Min Li 0007, Jianxin Wang 0001, Jianer Chen |
ISBRA | 1 |
| 2008 | Modifying the DPClus algorithm for identifying protein complexes based on new topological structuresabstractBACKGROUND: Identification of protein complexes is crucial for understanding principles of cellular organization and functions. As the size of protein-protein interaction set increases, a general trend is to represent the interactions as a network and to develop effective algorithms to detect significant complexes in such networks. RESULTS: Based on the study of known complexes in protein networks, this paper proposes a new topological structure for protein complexes, which is a combination of subgraph diameter (or average vertex distance) and subgraph density. Following the approach of that of the previously proposed clustering algorithm DPClus which expands clusters starting from seeded vertices, we present a clustering algorithm IPCA based on the new topological structure for identifying complexes in large protein interaction networks. The algorithm IPCA is applied to the protein interaction network of Sacchromyces cerevisiae and identifies many well known complexes. Experimental results show that the algorithm IPCA recalls more known complexes than previously proposed clustering algorithms, including DPClus, CFinder, LCMA, MCODE, RNSC and STM. CONCLUSION: The proposed algorithm based on the new topological structure makes it possible to identify dense subgraphs in protein interaction networks, many of which correspond to known protein complexes. The algorithm is robust to the known high rate of false positives and false negatives in data from high-throughout interaction techniques. The program is available at http://netlab.csu.edu.cn/bioinformatics/limin/IPCA. Min Li 0007, Jianer Chen, Jianxin Wang 0001, Bin Hu 0001, Gang Chen 0010 |
BMC Bioinform. | 1 |