EDBT 2026 Demo / reviewers in the wild / expert
Fei Guo 0001
dblp:85/3639-1
· DBLP profile ↗
138ranked-venue papers
7as first author
115since 2021 · last 2026
0000-0001-8346-0798ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 85 · 7 first-author · 71 since 2021Artificial intelligence and machine learning · 38 · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 15 since 2021Databases, data management, data science and information retrieval · 7 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PharmaQA: Prompt-Based Molecular Representation Learning via Pharmacophore-Oriented Question AnsweringabstractMolecular representation plays a central role in computational drug discovery. Pharmacophores, functional groups responsible for molecular bioactivity, have been widely studied in cheminformatics. However, their incorporation into molecular representation learning, particularly in a context reasoning or generalization, remains relatively limited. To address this gap, we propose PharmaQA, a pharmacophore oriented question answering framework that formulates tailored prompts to extract context-aware molecular semantics. Rather than encoding pharmacophore features, PharmaQA learns to answer pharmacophore related queries. This design enables flexible reasoning across diverse tasks, including molecular property prediction, compound-target interaction prediction, and binding affinity estimation. Experimental results on benchmark datasets demonstrate that PharmaQA achieves competitive performance. In a ligand discovery case study using FDA-approved compounds, the framework identified potential inhibitors for three therapeutic targets, with strong docking performance. As a generalizable and modular solution, PharmaQA incorporates pharmacophoric knowledge into molecular embeddings, enhancing both predictive accuracy and interpretability in drug discovery applications. Chengwei Ai, Qiaozhen Meng, Mengwei Sun, Ruihan Dong, Hongpeng Yang, Shiqiang Ma, Cheng Liang 0001, Fei Guo 0001 |
AAAI | 9 |
| 2026 | Geometry-Aware Variational Information Maximization for Deep Incomplete Multi-view ClusteringabstractIncomplete multi-view clustering (IMVC) aims to group data into meaningful clusters when each sample is only partially observed across multiple views. Most existing methods either rely on imputation strategies that may introduce noise and distort the underlying data distribution, or adopt cross-view alignment techniques that focus on pairwise relationships, often resulting in suboptimal representations and unstable clustering performance. In this paper, we propose Geometry-Aware Variational Information Maximization for Deep Incomplete Multi-view Clustering (GAVIM), a novel imputation-free variational framework that enables robust and coherent incomplete multi-view clustering. Specifically, GAVIM leverages mutual information maximization to preserve the high mutual information between the available multi-view data and the shared embedding. Moreover, we explicitly retain local geometric consistency within each view-specific latent space under the guidance of an adaptive global supervision signal. Lastly, GAVIM aligns all views simultaneously using a Gramian representation alignment measure, ensuring coherent structure across modalities and promoting unified, semantically meaningful representations. Extensive experiments on five benchmark IMVC datasets with varying levels of view incompleteness demonstrate that GAVIM consistently outperforms state-of-the-art methods in clustering accuracy and representation quality. Wenlan Chen, Daoyuan Wang, Fei Guo 0001, Cheng Liang 0001 |
AAAI | 4 |
| 2026 | Piercing the Fog: Disentangling Key Features for Vision Models in Multi-Degradation ScenariosabstractIn natural scenarios, vision models often encounter the challenge of complex degradation scenarios(e.g., rain, snow, fog, or motion blur). These degradations severely corrupt image features, causing existing models to treat rarely seen or unseen degraded images as “unfamiliar”, thereby losing their inherent recognition and perception capabilities. To address this challenge, we propose a novel degradation disentanglement model (DDM) aimed at precisely disentangling degraded features from the image. The model enhances its perception of various degradations by controlling the matching of features across different degradation types and further strengthens the cross-correlation of target features by introducing a degradation suppression module. This enables the model to re-identify and re-localize targets while removing degradations. We validated the effectiveness of our method on more challenging few-shot segmentation datasets Degraded-Pascal and Degraded-COCO. Results on them outperform SOTA with 3.71% and 3.69% improvement respectively. The experimental results show that our method significantly improves the performance of vision models in various degradation scenarios and provides new ideas and solutions for visual understanding tasks in complex environments. Shiqiang Ma, Fei Guo 0001 |
AAAI | 3 |
| 2026 | Closer to Biological Mechanism: Drug-Drug Interaction Prediction from the Perspective of PharmacophoreabstractDrug combinations are widely used in modern medicine but may cause severe adverse drug reactions. Therefore, making effective drug-drug interactions (DDI) prediction is crucial for pharmacovigilance. Existing DDI prediction models are typically built from a structural perspective, assuming that drugs with similar molecular structures may exhibit similar interactions. However, such approaches overlook the biological mechanisms underlying DDI in the human body. This not only weakens the generalization ability of the model, but also makes its interpretability less convincing. Inspired by this, we propose a new method called PC-DDI. Unlike structure-based models, PC-DDI utilizes pharmacophores as basic unit, and designs a complete pharmacophore feature processing framework. It further constructs a pharmacophore-based bipartite graph to model interactions between pharmacophores. This approach allows us to explore the underlying mechanisms of DDI from a functional perspective. We also design a spatial attention weight graph convolution module to optimize the message passing process by integrating pharmacophore position features with node features. Furthermore, we apply causal inference to identify key pharmacophores in pharmacophore bipartite graph, enhancing the interpretability. Compared with the SOTA, PC-DDI achieves an accuracy improvement of 1.84% under the transductive setting and consistently outperforms others in all other experiments. Mingliang Dou, Linfeng Wen 0005, Jinyang Xie, Jijun Tang, Shiqiang Ma, Fei Guo 0001 |
AAAI | 6 |
| 2026 | Make Foundation Models Trustworthy Again: Causal Fine-Adaptation for Medical Image SegmentationabstractVision foundation models (e.g., SAM2, CLIP) show strong generalization in natural image analysis but degrade significantly in specialized domains like medical imaging. This is critical for tasks such as brain tumor segmentation, where errors directly affect surgical planning and patient outcomes. In such contexts, segmentation must be highly reliable and structurally precise, underscoring the need for adaptable methods with low error tolerance. While fine-tuning is the dominant strategy, it is computationally expensive and prone to forgetting. To address this, we propose CausalBridgeNet, a causality-guided correction framework for medical image segmentation. Inspired by predictive coding theories of the Bayesian brain, our method introduces a Predictive Causal Reasoning Unit (PCRU) that estimates structured error maps and delivers targeted feedback to iteratively refine predictions. This forms a closed-loop, error-aware correction mechanism without modifying the foundation model. By keeping the backbone frozen, CausalBridgeNet preserves general visual priors while enhancing task-specific accuracy. On the BraTS 2025 benchmark, it achieves an average Dice score of 84.48 and HD95 of 5.48 across tumor subregions, demonstrating its effectiveness for high-precision medical segmentation. Hongpeng Yang, Yingxin Chen 0001, Shiqiang Ma, Fei Guo 0001 |
AAAI | 4 |
| 2026 | BioLemons: Latent Conditional Diffusion Model with VAE Embedding for Enhancing Spatial Transcriptomics
Haolu Zhou, Wenying He, Yude Bai, Fei Guo 0001 |
DASFAA (3) | 4 |
| 2026 | Enhancing Sample Discrimination: Drug-Drug Interaction Prediction Based on Bidirectional Event Semantics Guidance
Shiqiang Ma, Mingliang Dou, Fei Guo 0001, Jijun Tang |
ICIC (15) | 4 |
| 2026 | Motif-Aware Graph Attention Networks for Hemolytic Peptide Prediction
Qiaozhen Meng, Bangguo Tan, Fei Guo 0001 |
ISBRA (1) | 4 |
| 2026 | DualDis: A Dual Disentanglement Network for Vehicle Re-identification
Wenying He, Guangquan Xu, Yude Bai, Fei Guo 0001 |
WWW | 5 |
| 2026 | Refprogen: a reference-guided molecular generation model with protein-ligand joint representation for property-aware drug design
Chengwei Ai, Jijun Tang, Fei Guo 0001 |
Expert Syst. Appl. | 5 |
| 2026 | EssLM-MoE: A mixture-of-experts-enhanced framework for protein essentiality prediction using fused protein language models
Min Zeng 0004, Qianpei Liu, Wenkang Wang, Fuhao Zhang, Fei Guo 0001, Min Li 0007 |
Neurocomputing | 7 |
| 2026 | High-order correlation and consistency-aware multi-view clustering via anchor graph learning
Cheng Liang 0001, Wenchao Zang, Daoyuan Wang, Fei Guo 0001 |
Neural Networks | 4 |
| 2026 | Graph-Embedded Deep Generative Clustering for Single-Cell Multi-Omics Data IntegrationabstractThe advancement of sequencing technologies has generated an unprecedented volume of single-cell multi-omics data, providing new opportunities for biological discovery and medical research. However, due to the high heterogeneity across different omics types, effective integration of single-cell multi-omics data remains a critical challenge. Existing methods generally ignore the graph structure information among cells or resort to additional knowledge to construct the cell graphs, leading to suboptimal performance and potentially limited practical utility. In this study, we propose a novel Graph-embedded Deep Generative Clustering model (GeDGC) for single-cell multi-omics data integration. Specifically, GeDGC simultaneously learns the shared latent representations and cluster factors across multiple omics by leveraging Gaussian mixture models. Moreover, we impose the graph embedding constraint on both the latent representations and the cluster assignments to ensure the preservation of intrinsic local data structure among cells. As a result, our model captures complex correlations across omics and obtains informative shared latent embeddings for downstream tasks. Extensive experimental results with seventeen competing methods on ten datasets confirm the superiority of GeDGC in single-cell multi-omics data integration. Cheng Liang 0001, Wenlan Chen, Chang-Dong Wang 0001, Shichao Zhang 0001, Fei Guo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | MultiPert: An adversarial alignment and dual attention framework for single-cell multi-omics perturbation predictionabstractPrecise prediction of perturbation responses is essential in systems biology research, as it plays a pivotal role in characterizing cellular identities and elucidating the regulatory mechanisms of biological pathways. Existing perturbation-responses prediction approaches are predominantly confined to single-modality transcriptomic data, limiting their capacity to capture cross-layer molecular effects. Here, we present MultiPert, a deep learning framework specifically designed for predicting perturbation responses in single-cell multi-omics data. MultiPert employs modality-specific encoders with dedicated pretraining, integrates perturbation through a dual-attention mechanism, and achieves cross-modal alignment via adversarial training. Benchmarking on human THP-1 and kidney multi-omics datasets demonstrates that MultiPert reliably predicts both perturbed gene expression and protein abundance profiles, achieving superior accuracy and stability compared to state-of-the-art strategies. MultiPert generalizes to unseen perturbations and uncovers regulatory mechanisms of immune checkpoint molecules based on perturbed proteomic predictions. In addition, enrichment analyses of perturbed transcriptomic predictions reveal immune-related pathways. By providing an integrated and interpretable framework, MultiPert expands the scope of perturbation modeling at the multi-omics level, thereby offering a robust methodological foundation for comprehensive research into pathogenesis and drug discovery. Xinyue Tang, Jiawei Li 0018, Cheng Liang 0001, Jijun Tang, Fei Guo 0001 |
PLoS Comput. Biol. | 6 |
| 2026 | EdgeCLIP: Injecting Edge-Awareness Into Visual-Language Models for Zero-Shot Semantic SegmentationabstractEffective segmentation of unseen categories in zero-shot semantic segmentation is hindered by models’ limited ability to interpret edges in unfamiliar contexts. In this paper, we propose EdgeCLIP, which addresses this by integrating CLIP with explicit edge-awareness. Based on the premise that edge variation patterns are similar across both seen and unseen class objects, EdgeCLIP introduces the Contextual Edge Sensing module. This module accurately discerns and utilizes edge information, which is crucial in complex border areas where conventional models struggle. Further, our Text-Guided Dense Feature Matching strategy precisely aligns text encodings with corresponding visual edge features, effectively distinguishing them from background edges. This strategy not only optimizes the training of CLIP’s image and text encoders but also leverages the intrinsic completeness of objects, enhancing the model’s ability to generalize and accurately segment objects in unseen classes. EdgeCLIP significantly outperforms the current state-of-the-art method, achieving a deep impressive margin of 17.5% on COCO-20i datasets. Our code is available at github.com/aqingaqinghh/EdgeCLIP. Jiaxiang Fang, Shiqiang Ma, Guihua Duan, Fei Guo 0001, Shengfeng He |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Incomplete Multi-View Clustering via Robust Representation Learning and Tensor-Based Co-RegularizationabstractAs incompleteness is common in real-world data, incomplete multi-view clustering is of great significance in the unsupervised learning field because it allows the partitioning of multi-view data with missing information into distinct groups. In this paper, we propose a novel generalized framework for incomplete multi-view clustering based on robust representation learning and tensor-based co-regularization (RRLTCR). Specifically, a robust principal component analysis is first used to learn a robust representation for each view. To explore high-order relationships among views, the view-specific spectral embeddings are stacked into a third-order tensor with a Schattenp-norm constraint. By spreading the complementary information of the high-quality available data from each view on a global scale, our model is able to alleviate the adverse effects of data noise and uncover the underlying common cluster structure. An effective iterative optimization strategy is developed to efficiently solve our model. According to the experimental results on seven datasets, our proposed framework has the potential to improve the clustering performance for a variety of incomplete multi-view clustering problems. Our research work brings a generalized framework for incomplete multi-view clustering, which can also assist in exploring the large cohort of existing incomplete multimodality datasets for other downstream tasks. Cheng Liang 0001, Daoyuan Wang, Fei Guo 0001, Shichao Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | MF-DocDDI: Drug Entity Multi-Feature Fusion for Document-Level Drug-Drug Interaction Relation ExtractionabstractDrug-drug interactions (DDIs) are crucial in clinical medicine, as they can lead to adverse events. Existing DDI extraction methods focus on sentence-level tasks, limiting their ability to identify cross-sentence DDIs. Moreover, the only document-level method available considers only internal drug features, leading to suboptimal performance. To address this, we propose MF-DocDDI, a document-level DDI extraction model using drug entity multi-feature fusion. We first construct a document-level dataset based on DDI Extraction 2013. Then, we introduce document-entity embeddings to capture internal drug features and employ a simplified U-shaped network to extract external features. Finally, we integrate these features to enhance interaction modeling. Experimental results show MF-DocDDI outperforms existing methods, improving the F1 score by 5.33 %. Case studies confirm its ability to identify cross-sentence DDIs, such as (naloxone, morphine) and (HEXALEN, cisplatin). Beyond DDI extraction, MF-DocDDI can be applied to other biomedical tasks like protein-protein interaction (PPI) extraction. Mingliang Dou, Jijun Tang, Fei Guo 0001 |
BIBM | 4 |
| 2025 | LGATFormer: A Dual-Path Model Combining Line Graph Attention and Transformer for Gene Regulatory Network InferenceabstractReconstructing high-precision gene regulatory networks (GRNs) from single-cell RNA sequencing (scRNA-seq) data presents significant challenges, including high noise levels, data sparsity, and structural complexity. We introduce LGATFormer, a novel deep learning model that combines local structural modeling with global dependency extraction to address these challenges effectively. LGATFormer extracts enclosing subgraphs centered on target gene pairs, transforms them into line graphs, and uses Graph Attention Networks (GATs) to learn edge-level representations, capturing high-order regulatory structures. For global modeling, the model employs a Transformer encoder to process the entire gene expression matrix, utilizing self-attention mechanisms to model long-range dependencies between genes. Across multiple benchmarks, LGATFormer outperforms existing methods by achieving state-of-the-art AUROC and AUPRC on 89.29% of STRING and Non-Specific ground-truth networks, while exhibiting superior generalization, robustness, and interpretability. This model offers an effective and reliable solution for GRN inference, advancing both theoretical and practical applications in systems biology. Wenying He, Yaowei Zhu, Rentao Zhang, Haolu Zhou, Yude Bai, Fei Guo 0001 |
BIBM | 6 |
| 2025 | MolInterAct: Multiscale Cross-Modal Interaction for Robust Molecular Representation LearningabstractMolecular representation learning, which captures the fundamental characteristics of chemical compounds, is crucial for AI-driven drug discovery. Existing methods integrate various modalities (e.g., 2D topology and 3D geometry) to develop robust representations. However, current multi-modal fusion strategies either align embedding space through independent models separately, thereby overlooking complementary information, or bridge modalities at a coarse-grained level, failing to capture inherent correlations. We present MolInterAct, an innovative pretraining framework designed to promote multiscale interactions between 2D and 3D modalities at both atomic-level and moleculelevel. Specifically, we propose a fine-grained fusion module, coupled with a customized complementary masking strategy, to seamlessly integrate information at the atomic-level, mitigating overlap and similarity between 2D and 3D representations. In addition, we introduce a fusion contrastive module, which operates at the molecule level, to further strengthen the fusion of 2D and 3D representations while preserving modality-specific features. Finally, we incorporate an intra-modal reconstruction module to reconstruct the original information, further refining the model's understanding of individual modality. Extensive experiments demonstrate that our model outperforms existing molecular pretraining methods across both 2D and 3D benchmarks, highlighting the effectiveness of multiscale fusion between modalities. Mengwei Sun, Chengwei Ai, Diya Zhang, Qiaozhen Meng, Shiqiang Ma, Fei Guo 0001 |
BIBM | 7 |
| 2025 | MultiPepDec: Decoupled Prompt Learning for Multi-Activity Therapeutic PeptidesabstractTherapeutic peptides demonstrate significant potential in anti-infection, antitumor, and immunomodulation therapies owing to their high specificity and low toxicity. However, existing computational methods are predominantly limited to single-activity design, restricting their clinical applicability. Here, we present MultiPepDec, a novel decoupled prompt learning framework based on protein language model for concurrent generation of multifunctional peptides, which includes antimicrobial, anticancer, toxic, and metabolic activities. Our approach employs: i) Shared-prompts capturing universal therapeutic patterns via adversarial purification; ii) Private-prompts encoding activity-specific knowledge through contrastive learning, ensuring functional decoupling between four activities. Experimental results demonstrate that generated antimicrobial peptides achieve 80.38% predicted efficacy against E. coli, with comparable performance against most clinically relevant pathogens. This confirms robust broad-spectrum capabilities without requiring pathogen-specific training, while maintaining low computational costs. For other therapeutic activities, the designed sequences not only exhibit the intended biological functions but also show significantly improved diversity. This work establishes a new paradigm for efficient multi-activity peptide design, with potential extensions to other biomolecular engineering domains. Xingdan Wang, Diya Zhang, Chengwei Ai, Shiqiang Ma, Qiaozhen Meng, Junwen Duan, Fei Guo 0001 |
BIBM | 7 |
| 2025 | Unlocking Multimodal Potential for Few-Shot Semantic Segmentation with Vision-Enriched Text
Jiaxiang Fang, Shiqiang Ma, Fei Guo 0001 |
DASFAA (1) | 4 |
| 2025 | Catching mRNA's Hiddens Marks: A Dual-Path Network by Contrastive Learning for N4-acetylcytidine PredictionabstractN4-acetylcytidine (ac4C) is a crucial RNA modification associated with mRNA stability and translational efficiency. Accurate identification of ac4C sites is essential for understanding their regulatory functions. However, experimental detection remains expensive and labor-intensive. At the same time, existing computational models suffer from limited generalization and insufficient feature discrimination, especially in distinguishing subtle nucleotide patterns. In this work, we propose a deep learning model named SNN-ac4C, which is based on a contrastive learning-based neural network. The model integrates a dual-path structure that combines BiLSTM and Multi-Head Self-Attention (MHSA) for capturing long-range dependencies and global context, while using CNN to extract local biological sequence features. The contrastive learning module further enhances the discriminative ability of ac4C and Non-ac4C sites by increasing the separation between positive and negative samples. Experiments on the test set confirm the effectiveness of SNN-ac4C, which achieves an accuracy (ACC) of 84.60% and a Matthews Correlation Coefficient (MCC) of 0.6934. Compared with NBCR-ac4C, the current state-of-the-art model, SNN-ac4C improves ACC and MCC by 1.09% and 0.0228, respectively. The source code and relevant supplementary are publicly available at https://github.com/2103374200/SNN. Wenying He, Haolu Zhou, Yun Zuo 0001, Yude Bai, Fei Guo 0001 |
ECAI | 6 |
| 2025 | Beyond Slice-by-Slice: 3D Lesion Segmentation via Cross-Frame PredictionabstractThree-dimensional medical image segmentation plays a significant role in clinical diagnosis, treatment planning, and disease research, as it provides doctors with precise anatomical and lesion information and improves the accuracy and efficiency of medical decision-making. However, most existing 3D segmentation approaches rely heavily on densely volumetric data and often fail to perform segmentation properly for incomplete 3D volume acquisition, i.e., missing slices. In this work, we present InterFrameNet, a framework designed to predict intermediate lesion structures by modeling spatial relationships across frames, enabling robust segmentation performance under sparse acquisition conditions, without requiring full-volume information. Our method explicitly models cross-frame spatial continuity and leverages structural relationships between available frames to accurately infer missing lesion regions. This design significantly reduces the dependence on consecutive frames while fully exploiting contextual anatomical information. Extensive experiments on brain lesion datasets demonstrate that our approach achieves robust segmentation performance under sparse acquisition settings, offering a practical solution to maximize usability of incomplete clinical imaging data. Hongpeng Yang, Yingxin Chen 0001, Xiangyu Hu 0005, Srihari Nelakuditi, Shiqiang Ma, Fei Guo 0001 |
ECAI | 7 |
| 2025 | Self-Support Prototype-Aware For Few-Shot Semantic SegmentationabstractIn recent years, significant progress has been made in prototype-based learning methods for few-shot semantic segmentation. However, prototype features originating from the support images are interfered with by intra-class diversity and thus cannot be aligned with the query foreground, resulting in poor segmentation accuracy. Therefore, we propose a novel self-support prototype-aware (SSPA) network to obtain highly confident query foreground pixel points and their corresponding query features. We design Cycle Consistency Collection module and Self-Support Collection module to address the interference of invalid support prototypes. Experimental results demonstrate that our SSPA significantly improves the quality of prototypes and achieves state-of-the-art segmentation results on multiple datasets. In particular, SSPA achieves mIoU scores of 69.7% and 76.4% for 1-shot and 5-shot segmentation, respectively, on PASCAL-5i. Jiaxiang Fang, Shiqiang Ma, Shengfeng He, Fei Guo 0001 |
ICASSP | 4 |
| 2025 | RetroInText: A Multimodal Large Language Model Enhanced Framework for Retrosynthetic Planning via In-Context Representation LearningabstractDevelopment of robust and effective strategies for retrosynthetic planning requires a deep understanding of the synthesis process. A critical step in achieving this goal is accurately identifying synthetic intermediates. Current machine learning-based methods often overlook the valuable context from the overall route, focusing only on predicting reactants from the product, requiring cost annotations for every reaction step, and ignoring the multi-faced nature of molecular, resulting in inaccurate synthetic route predictions. Therefore, we introduce RetroInText, an advanced end-to-end framework based on a multimodal Large Language Model (LLM), featuring in-context learning with TEXT descriptions of synthetic routes. First, RetroInText including ChatGPT presents detailed descriptions of the reaction procedure. It learns the distinct compound representations in parallel with corresponding molecule encoders to extract multi-modal representations including 3D features. Subsequently, we propose an attention-based mechanism that offers a fusion module to complement these multi-modal representations with in-context learning and a fine-tuned language model for a single-step model. As a result, RetroInText accurately represents and effectively captures the complex relationship between molecules and the synthetic route. In experiments on the USPTO pathways dataset RetroBench, RetroInText outperforms state-of-the-art methods, achieving up to a 5% improvement in Top-1 test accuracy, particularly for long synthetic routes. These results demonstrate the superiority of RetroInText by integrating with context information over routes. They also demonstrate its potential for advancing pathway design and facilitating the development of organic chemistry. Code is available at https://github.com/guofei-tju/RetroInText. Chenglong Kang, Fei Guo 0001 |
ICLR | 3 |
| 2025 | Image-Enhanced Hybrid Encoding with Reinforced Contrastive Learning for Spatial Domain Identification in Spatial TranscriptomicsabstractSpatial transcriptomics integrates spatial, gene expression, and multichannel immunohistochemistry image data, enabling advanced insights into cellular organization. However, existing methods often struggle to effectively fuse these multimodal data, limiting their potential for accurate spatial domain identification. Here, we propose IE-HERCL (Image-Enhanced Hybrid Encoding with Reinforced Contrastive Learning), a novel framework designed to address this challenge. Specifically, IE-HERCL employs hybrid encoding to capture both the non-spatial features and spatial dependencies for both gene and image modalities via autoencoders and GraphSAGE, respectively. These features are then fused using cross-view attention mechanisms to generate the unified informative embedding. To enhance the representation learning capability, we introduce a reinforced contrastive learning strategy to mitigate the influences of false negative samples, where we detect potential positive counterparts with high-order random walks. In addition, the cluster alignment is dynamically refined through optimal transport, which ensures that the fused consensus representation is coherent and robust, enabling accurate spatial domain identification. Our approach achieves state-of-the-art performance on five image-enhanced spatial transcriptomics datasets, demonstrating its robustness and effectiveness in multimodal integration and spatial domain identification. IE-HERCL offers a powerful and innovative solution for advancing spatial transcriptomics analysis. The code is released on https://github.com/wdyi701/IE-HERCL. Daoyuan Wang, Wenlan Chen, Cheng Liang 0001, Fei Guo 0001 |
IJCAI | 5 |
| 2025 | Unlocking Dark Vision Potential for Medical Image SegmentationabstractAccurate segmentation of lesions is crucial for disease diagnosis and treatment planning. However, blurring and low contrast in the imaging process can affect segmentation results. We have observed that noninvasive medical imaging shares considerable similarities with natural images under low light conditions and that nocturnal animals possess extremely strong night vision capabilities. Inspired by the dark vision of these nocturnal animals, we proposed a novel plug-and-play dark vision network (DVNet) to enhance the model's perception for low-contrast medical images. Specifically, by employing the wavelet transform, we decompose medical images into subbands of varying frequencies, mimicking the sensitivity of photoreceptor cells to different light intensities. To simulate the antagonistic receptive fields of horizontal cells and bipolar cells, we design a Mamba-Enhanced Fusion Module to achieve global information correlation and enhance contrast between lesions and surrounding healthy tissues. Extensive experiments demonstrate that the DVNet achieves SOTA performance in various medical image segmentation tasks. Hongpeng Yang, Xiangyu Hu 0005, Yingxin Chen 0001, Srihari Nelakuditi, Shiqiang Ma, Fei Guo 0001 |
IJCAI | 8 |
| 2025 | Deep Variational Incomplete Multi-View Clustering with Information-Theoretic Guidance
Wenlan Chen, Cheng Liang 0001, Fei Guo 0001 |
ACM Multimedia | 4 |
| 2025 | Dual-Level Distribution Alignment for Deep Incomplete Multi-View ClusteringabstractIncomplete Multi-view Clustering (IMvC) aims to perform effective clustering in the presence of missing views by exploiting the available information. While many existing approaches demonstrate satisfactory performance, their failure to adequately optimize the recovered data often limits the quality of learned representations and thus hampers clustering performance. To address this challenge, we propose a novel method, Dual-Level Distribution Alignment for Deep Incomplete Multi-View Clustering (DDAIMVC). To effectively address missing data, DDAIMVC employs a fusion-fill strategy to recover incomplete views. The recovered data from each view are then concatenated and processed through an attention mechanism to generate a unified high-level representation. To ensure consistent information across views, the framework performs distribution alignment at both the instance and cluster levels. Specifically, instance-level distribution alignment is conducted by minimizing the maximum mean discrepancy among views, while cluster-level distribution alignment is enhanced via prototypical contrastive learning, which encourages coherent cluster assignments across different modalities. Through the co-optimization of dual-level distribution alignment, the common representation reveals a clear clustering structure. Experimental results on benchmark multi-view datasets demonstrate that DDAIMVC consistently achieves state-of-the-art clustering performance. Fujian Ren, Wenlan Chen, Fei Guo 0001, Cheng Liang 0001 |
ACM Multimedia | 4 |
| 2025 | Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionabstractCross-modal representation learning aims to extract semantically aligned representations from heterogeneous modalities such as images and text. Existing multimodal VAE-based models often suffer from limited capability to align heterogeneous modalities or lack sufficient structural constraints to clearly separate the modality-specific and shared factors. In this work, we propose a novel framework, termed **D**isentangled **C**ross-**M**odal Representation Learning with **E**nhanced **M**utual Supervision (DCMEM). Specifically, our model disentangles the common and distinct information across modalities and regularizes the shared representation learned from each modality in a mutually supervised manner. Moreover, we incorporate the information bottleneck principle into our model to ensure that the shared and modality-specific factors encode exclusive yet complementary information. Notably, our model is designed to be trainable on both complete and partial multimodal datasets with a valid Evidence Lower Bound. Extensive experimental results demonstrate significant improvements of our model over existing methods on various tasks including cross-modal generation, clustering, and classification. Wenlan Chen, Daoyuan Wang, Fei Guo 0001, Cheng Liang 0001 |
NeurIPS | 4 |
| 2025 | DynaPhArM: Adaptive and Physics-Constrained Modeling for Target-Drug Complexes with Drug-Specific AdaptationsabstractAccurately modeling the target-drug complex at atom level presents a significant challenge in the computer-aided drug design. Traditional methods that rely solely on rigid transformations often fail to capture the adaptive interactions between targets and drugs, particularly during substantial conformational changes in targets upon ligand binding, which becomes especially critical when learning target-drug interactions in drug design. Accurately modeling these changes is crucial for understanding target-drug interactions and improving drug efficacy. To address these challenges, we introduce DynaPhArM, an SE(3)-Equivariant Transformer model specifically designed to capture adaptive alterations occurring within target-drug interactions. DynaPhArM utilizes the cooperative scalar-vector representation, drug-specific embeddings, and a diffusion process to effectively model the evolving dynamics of interactions between targets and drugs. Furthermore, we integrate physical information and energetic principles that maintain essential geometric constraints, such as bond lengths, bond angles, van der Waals forces (vdW), within a multi-task learning (MTL) framework to enhance accuracy. Experimental results demonstrate that DynaPhArM achieves state-of-the-art performance with an overall root mean square deviation (RMSD) of 2.01 Å and a sc-RMSD of 0.29 Å while exhibiting higher success rates compared to existing methodologies. Additionally, DynaPhArM shows promise in enhancing drug specificity, thereby simulating how targets adapt to various drugs through precise modeling of atomic-level interactions and conformational flexibility. Diya Zhang, Mengwei Sun, Xingdan Wang, Cheng Liang 0001, Qiaozhen Meng, Shiqiang Ma, Fei Guo 0001 |
NeurIPS | 7 |
| 2025 | deepTAD: an approach for identifying topologically associated domains based on convolutional neural network and transformer modelabstractMOTIVATION: Topologically associated domains (TADs) play a key role in the 3D organization and function of genomes, and accurate detection of TADs is essential for revealing the relationship between genomic structure and function. Most current methods are developed to extract features in Hi-C interaction matrix to identify TADs. However, due to complexities in Hi-C contact matrices, it is difficult to directly extract features associated with TADs, which prevents current methods from identifying accurate TADs. RESULTS: In this paper, a novel method is proposed, deepTAD, which is developed based on a convolutional neural network (CNN) and transformer model. First, based on Hi-C contact matrix, deepTAD utilizes CNN to directly extract features associated with TAD boundaries. Next, deepTAD takes advantage of the transformer model to analyze the variation features around TAD boundaries and determines the TAD boundaries. Second, deepTAD uses the Wilcoxon rank-sum test to further identify false-positive boundaries. Finally, deepTAD computes cosine similarity among identified TAD boundaries and assembles TAD boundaries to obtain hierarchical TADs. The experimental results show that TAD boundaries identified by deepTAD have a significant enrichment of biological features, including structural proteins, histone modifications, and transcription start site loci. Additionally, when evaluating the completeness and accuracy of identified TADs, deepTAD has a good performance compared with other methods. The source code of deepTAD is available at https://github.com/xiaoyan-wang99/deepTAD. Huimin Luo, Fei Guo 0001 |
Briefings Bioinform. | 5 |
| 2025 | RNALoc-LM: RNA subcellular localization prediction using pre-trained RNA language modelabstractMOTIVATION: Accurately predicting RNA subcellular localization is crucial for understanding the cellular functions and regulatory mechanisms of RNAs. Although many computational methods have been developed to predict the subcellular localization of lncRNAs, miRNAs, and circRNAs, very few of them are designed to simultaneously predict the subcellular localization of multiple types of RNAs. In addition, the emergence of pre-trained RNA language model has shown remarkable performance in various bioinformatics tasks, such as structure prediction and functional annotation. Despite these advancements, there remains a significant gap in applying pre-trained RNA language models specifically for predicting RNA subcellular localization. RESULTS: In this study, we proposed RNALoc-LM, the first interpretable deep-learning framework that leverages a pre-trained RNA language model for predicting RNA subcellular localization. RNALoc-LM uses a pre-trained RNA language model to encode RNA sequences, then captures local patterns and long-range dependencies through TextCNN and BiLSTM modules. A multi-head attention mechanism is used to focus on important regions within the RNA sequences. The results demonstrate that RNALoc-LM significantly outperforms both deep-learning baselines and existing state-of-the-art predictors. Additionally, motif analysis highlights RNALoc-LM's potential for discovering important motifs, while an ablation study confirms the effectiveness of the RNA sequence embeddings generated by the pre-trained RNA language model. AVAILABILITY AND IMPLEMENTATION: The RNALoc-LM web server is available at http://csuligroup.com:8000/RNALoc-LM. The source code can be obtained from https://github.com/CSUBioGroup/RNALoc-LM. Min Zeng 0004, Chengqian Lu, Rui Yin 0002, Fei Guo 0001, Min Li 0007 |
Bioinform. | 6 |
| 2025 | Leveraging protein language models for cross-variant CRISPR/Cas9 sgRNA activity predictionabstractMOTIVATION: Accurate prediction of single-guide RNA (sgRNA) activity is crucial for optimizing the CRISPR/Cas9 gene-editing system, as it directly influences the efficiency and accuracy of genome modifications. However, existing prediction methods mainly rely on large-scale experimental data of a single Cas9 variant to construct Cas9 protein (variants)-specific sgRNA activity prediction models, which limits their generalization ability and prediction performance across different Cas9 protein (variants), as well as their scalability to the continuously discovered new variants. RESULTS: In this study, we proposed PLM-CRISPR, a novel deep learning-based model that leverages protein language models to capture Cas9 protein (variants) representations for cross-variant sgRNA activity prediction. PLM-CRISPR uses tailored feature extraction modules for both sgRNA and protein sequences, incorporating a cross-variant training strategy and a dynamic feature fusion mechanism to effectively model their interactions. Extensive experiments demonstrate that PLM-CRISPR outperforms existing methods across datasets spanning seven Cas9 protein (variants) in three real-world scenarios, demonstrating its superior performance in handling data-scarce situations, including cases with few or no samples for novel variants. Comparative analyses with traditional machine learning and deep learning models further confirm the effectiveness of PLM-CRISPR. Additionally, motif analysis reveals that PLM-CRISPR accurately identifies high-activity sgRNA sequence patterns across diverse Cas9 protein (variants). Overall, PLM-CRISPR provides a robust, scalable, and generalizable solution for sgRNA activity prediction across diverse Cas9 protein (variants). AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/CSUBioGroup/PLM-CRISPR. Yalin Hou, Ruiqing Zheng, Fuhao Zhang, Fei Guo 0001, Min Li 0007, Min Zeng 0004 |
Bioinform. | 5 |
| 2025 | Cancer survival prediction based on soft-label guided contrastive learning and global feature fusionabstractMOTIVATION: The high complexity and heterogeneity of cancer pose significant challenges to personalized treatment, making the improvement of cancer survival prediction accuracy crucial for clinical decision-making. The integration of multi-omics data enables a more comprehensive capture of multi-layered information in complex biological processes. However, existing survival analysis models still face limitations in accurately extracting and effectively integrating the unique and shared information from multi-omics data. RESULTS: In this article, we propose a novel prediction model for cancer survival based on soft-label guided contrastive learning and global feature fusion, namely SLCGF. Our model first extracts paired feature representations for each omics using Siamese encoders. We then perform intra-view and inter-view contrastive learning simultaneously, employing a neighborhood-based paradigm to enhance feature discrimination and alignment across omics. To ensure reliable neighbor retention and improve model robustness, we treat the affinities between samples and their high-order neighbors as soft labels to guide the contrastive learning process at both levels. In addition, we adopt a global self-attention mechanism to obtain the unified representation for cancer survival prediction, where the cross-omics connections are fully exploited and complementary information is adaptively integrated. We comprehensively evaluate the performance of our model on 13 cancer multi-omics datasets, and the experimental results demonstrate its superiority over existing approaches. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/LiangSDNULab/SLCGF. Huiying Jiang, Wenlan Chen, Fei Guo 0001, Cheng Liang 0001 |
Bioinform. | 3 |
| 2025 | 2OMe-LM: predicting 2′-O-methylation sites in human RNA using a pre-trained RNA language modelabstractMOTIVATION: 2'-O-methylation (2OMe) is a common post-transcriptional modification in RNA that plays a crucial role in regulating gene expression and is implicated in various biological processes and diseases. Computational methods offer an efficient alternative to the time-consuming and costly experimental identification of 2OMe sites. Recent advancements in RNA pre-trained language models have revolutionized RNA bioinformatics. However, there remains a gap in their application specifically for predicting 2OMe sites. RESULTS: In the study, we propose a novel deep learning framework, 2OMe-LM, for predicting 2OMe sites in RNA. 2OMe-LM integrates RNA sequence features derived from RNA pre-trained language models with those obtained from the word2vec technique. Then, 2OMe-LM employs fully connected layers and a bidirectional long short-term memory network to process the two types of features separately, followed by a feature fusion module for the final prediction. Additionally, an attention block is incorporated to provide the interpretability of the prediction results. The results demonstrate that 2OMe-LM significantly outperforms existing state-of-the-art predictors, with features from RNA pre-trained language models proving to be critical. Motif analysis further demonstrates 2OMe-LM's potential for discovering 2OMe-related motifs. AVAILABILITY AND IMPLEMENTATION: The 2OMe-LM web server is available at https://csuligroup.com:9200/2OMe-LM. The source code can be obtained from https://github.com/CSUBioGroup/2OMe-LM. Qianpei Liu, Min Zeng 0004, Chengqian Lu, Shichao Kan, Fei Guo 0001, Min Li 0007 |
Bioinform. | 6 |
| 2025 | HyperPhS: a pharmacophore-guided multimodal representation framework for metabolic stability prediction through contrastive hypergraph learningabstractMOTIVATION: Metabolic stability is crucial in the early stage of drug discovery and development. Drug candidate screening and optimization can be streamlined through the accurate prediction of stability. Functional groups within drug molecules are known as pharmacophores, which bind directly to receptors or biological macromolecules to produce biological effects, thereby affecting metabolic stability. Therefore, determining metabolic stability via the pharmacophore groups remains a significant challenge. RESULTS: To address these issues, we propose a Pharmacophore-guided Hypergraph representation framework for predicting metabolic Stability (HyperPhS). In this study, we introduce a hypergraph-based method to extract features from metabolic pharmacophores with multi-view representation and contrastive learning. In particular, we introduce a pharmacophore-based contrastive learning encoder that captures the consistency between functional and nonfunctional structures. Our method applies ChatGPT simultaneously to metabolites and heterogeneous encoders and integrates multimodal representations by using attention-driven fusion modules coupled with fully connected neural networks. On the HLM dataset, HyperPhS achieves outstanding performance with 87.6% in AUC and 62.6% in MCC, alongside an external test AUC of 88.3%. In addition, pharmacophore groups studied by HyperPhS are validated for their interpretability through case studies. Overall, HyperPhS is an effective and interpretable tool for determining metabolic stability, identifying critical functional groups, and optimizing compounds. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://github.com/xiaoyiliu-usc/HyperPhS. Chenglong Kang, Chengwei Ai, Hongpeng Yang, Jijun Tang, Fei Guo 0001 |
Bioinform. | 7 |
| 2025 | PEFN: A Patches Enhancement and Hierarchical Fusion Network for Robust Vehicle ReidentificationabstractVehicle Re-Identification (Re-ID), which is a significant application in the Internet of Things, aims to accurately retrieve the remaining images of a given vehicle across different cameras views. The improvement in vehicle Re-ID performance largely stems from better addressing the issues of inter-class similarity and intra-class variance. Existing methods, relying solely on max or average pooling after using attention modules, fail to obtain significantly complete and pure global and local features, and neglect the false guidance that some unique individual information on images bring to re-identification. Moreover, models combining global and local features have shown good results in vehicle Re-ID, but these successes neglect the interaction between features across different convolutional layers, resulting in the loss of crucial details for vehicle Re-ID. To tackle these issues, we introduce a Patches Enhancement and hierarchical Fusion Network (PEFN) based on a multi-branch architecture, divided into a Global and Local Attention Supplement (GLAS) branch, and an Enhanced Hierarchical feature fusion (EnHi) branch. The GLAS branch, through the Identity-related Feature Remodeling (IDFR) module’s staged supplementation of spatial and channel features, has achieved the enhancement of both global and local features and effectively mitigated the negative impacts of individual information. The EnHi branch enhances the robustness of feature representation by interacting hierarchical features. Extensive experiments on two large-scale vehicle re-identification datasets demonstrate that our PEFN method outperforms state-of-the-art vehicle re-identification approaches. Specifically, without utilizing extra data and re-ranking, our model achieves 85.15% mAP on the VeRi776 dataset. Code is available at https://github.com/711L/PEFN. Wenying He, Yude Bai, Naixue Xiong, Guangquan Xu, Fei Guo 0001 |
IEEE Internet Things J. | 6 |
| 2025 | Block sparse Bayes-based fuzzy system for RNA N6-methyladenosine sites predictionabstractN6-methyladenosine (m6A) can significantly affect RNA expression, gene regulation, and determination of cell fate. As a common and abundant post-transcriptional modification (PTM) of RNA, m6A is also closely associated with the occurrence of numerous diseases. Thus, identifying the m6A modification site in the RNA sequence is a prerequisite for related research. High-throughput sequencing technology has high requirements and low cost performance. Computational methods have made encouraging progress in site prediction. However, most models only consider the effects of different species, ignoring the simultaneous exploration of RNA modifications in different tissues within the same species. We develop and validate a fuzzy system based on Block Sparse Bayesian Learning (BSBL), named BSBL-TSK-FS, which is a powerful sequence-level m6A prediction model. We introduce a Bayesian method that provides a posterior probability output to produce more sparse solutions so that the model has higher accuracy. The model classifies the m6A sites in several tissues of mouse, human, and rat. Under the five-fold cross-validation method (5-CV), the precision of the BSBL-TSK-FS model is 0.84∼0.95. The accuracy of our model improves by 9.4% over the existing SOTA predictors. BSBL-TSK-FS achieves superior performance over current SOTA methods. Finally, in order to verify the generalizability of the model, we carry out cross-species tests, and the results prove the robustness and adaptability of the model. An accurate and reliable sequence modification prediction model is developed to better understand the complex landscape of methylation modification. Yuqing Qian, Wenhuan Lu, Yijie Ding, Fei Guo 0001 |
PLoS Comput. Biol. | 7 |
| 2025 | Prediction of ncRNA-Disease Association Based on Correntropy Induced Loss Matrix Factorization ModelabstractIn recent years, numerous studies have demonstrated a close connection between human diseases and the regulation of non-coding RNAs (ncRNAs). Predicting potential ncRNAs associated with disease can help provide critical information for diagnosis and treatment of disease, leading to better disease analysis and prevention. Building good algorithms for predicting associations between ncRNAs and disease is critical. Many current algorithms have poor performance in identifying the association between ncRNAs and diseases. As a method for predicting the association between ncRNAs and diseases, we develop a Matrix Factorization method based on the Correntropy Induced Loss (C-loss) function (C-lossMF). In our model, we first construct ncRNA similarity matrix and disease similarity matrix by considering some important similarity information, and extract effective information of ncRNA and disease from them. Next, we perform matrix decomposition of ncRNA-disease association matrix and apply $L2$ loss and C-loss. Then we add collaborative regularization of RNA similarity matrix and the collaborative regularization of disease similarity matrix to take full advantage of the information in the similarity matrix. In particular, we propose a method that combines semi-quadratic optimization and gradient descent to optimize the model. In the experiments, we utilize the five-fold cross validation method on four datasets to evaluate the performance of C-lossMF. Comparing this model with other advanced models, the results show that it performs better. Yuqing Qian, Junhai Xu, Yijie Ding, Fei Guo 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | scMID: A Deep Multi-Omics Integration Framework for Comprehensive Single-Cell Data AnalysisabstractBiological research on single cells has witnessed remarkable progress in recent years, with downstream analyses playing a crucial role in uncovering cellular functions and mechanisms. Traditional single-cell analyses, which predominantly rely on single-omics data such as single-cell RNA sequencing, are inherently limited. These methods can only capture one aspect of cellular information, overlooking the complex interplay between different molecular layers, and thus are prone to introducing biases in results. The advent of single-cell multi-omics sequencing technologies has revolutionized this landscape. By enabling the integration of diverse molecular profiles, including transcriptomics, epigenomics, and proteomics, these technologies offer a more holistic view of cellular functions. However, existing integration methods often lack the ability to handle the complexity and heterogeneity of multi-omics data, limiting their application in in-depth single-cell studies. In this study, we propose an analysis method based on single-cell multi-omics data integration and dropout pattern (scMID). Specifically, scMID utilizes omics-independent deep autoencoders for the alignment of multi-omics data, employs GCN algorithm for data integration, and calculates the gene importance by combining the gene similarity obtained from the binarized dropout pattern. Meanwhile, scMID proposes a dual-strategy for feature gene screening, aiming to identify genes with high biological significance that best match the structural characteristics of reference data. Experimental results demonstrate that scMID significantly improves the accuracy of single-cell clustering in downstream analyses, breaking through the limitations of traditional feature selection methods and providing a superior analytical framework for decoding complex biological information. Qiu Xiao, Wanwan Shi, Ying Zuo, Fei Guo 0001, Jiawei Luo 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2025 | Structured Sparse Regularization-Based Deep Fuzzy Networks for RNA N6-Methyladenosine Sites PredictionabstractIn many biological processes, N6-methyladenosine (m6A) plays a critical role. Experimental methods for identifying m6A sites have proven to be costly, and existing computational methods still require improvement. To address these challenges, we develop a novel computational method called structured sparse regularization-based fuzzy hierarchical echo state network to identify m6A sites in mammals. We apply fuzzy systems to deep learning. Compared with traditional fuzzy inference systems, this deep fuzzy network has the ability to generate feature representations. Echo state network (ESN) is a special type of recurrent neural network, which consists of an input layer, a randomly generated large fixed hidden layer (called a reservoir), and an adaptive output layer. The advantages of our method over ESNs are that it is capable of mining and capturing hidden features layer-by-layer within reservoirs and has better approximation performance. In order to remove redundancy, the output layer weights are trained by structured sparse learning, which enhances the generalizability and robustness of the method. Evaluation of our method by testing it on tissue-specific datasets shows that it outperforms existing tools. Yuqing Qian, Hao Xie 0003, Yijie Ding, Fei Guo 0001 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2025 | CellCircLoc: Deep Neural Network for Predicting and Explaining Cell Line-Specific CircRNA Subcellular LocalizationabstractThe subcellular localization of circular RNAs (circRNAs) is crucial for understanding their functional relevance and regulatory mechanisms. CircRNA subcellular localization exhibits variations across different cell lines, demonstrating the diversity and complexity of circRNA regulation within distinct cellular contexts. However, existing computational methods for predicting circRNA subcellular localization often ignore the importance of cell line specificity and instead train a general model on aggregated data from all cell lines. Considering the diversity and context-dependent behavior of circRNAs across different cell lines, it is imperative to develop cell line-specific models to accurately predict circRNA subcellular localization. In the study, we proposed CellCircLoc, a sequence-based deep learning model for circRNA subcellular localization prediction, which is trained for different cell lines. CellCircLoc utilizes a combination of convolutional neural networks, Transformer blocks, and bidirectional long short-term memory to capture both sequence local features and long-range dependencies within the sequences. In the Transformer blocks, CellCircLoc uses an attentive convolution mechanism to capture the importance of individual nucleotides. Extensive experiments demonstrate the effectiveness of CellCircLoc in accurately predicting circRNA subcellular localization across different cell lines, outperforming other computational models that do not consider cell line specificity. Moreover, the interpretability of CellCircLoc facilitates the discovery of important motifs associated with circRNA subcellular localization. Min Zeng 0004, Jingwei Lu, Chengqian Lu, Shichao Kan, Fei Guo 0001, Min Li 0007 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | DP-BERT: a pre-trained deep language model for depression prediction using microarray dataabstractIn recent years, the increasing number of individuals diagnosed with depression and the growing awareness of its impact on modern society have highlighted the significance of accurate depression diagnosis. Microarray data has played a crucial role in uncovering the genetic mechanisms underlying depression. However, existing methods for depression prediction using microarray data often rely on the selection of differentially expressed genes. This approach disregards important information from other genes and is susceptible to batch effects, thereby limiting generalizability and model stability. To address these limitations, we propose DP-BERT, a depression prediction model based on Bidirectional Encoder Representations from Transformers (BERT). DP-BERT follows a pre-training and fine-tuning paradigm, leveraging a large amount of unlabeled microarray data from diverse sequencing platforms for pretraining to extract comprehensive genetic-level representations of psychiatric disorders. Subsequently, supervised fine-tuning is performed for depression prediction. Experimental results demonstrate that the pre-trained model achieves superior performance in depression prediction. The source code can be obtained from https://github.com/CSUBioGroup/DP-BERT. Junyu Gao 0004, Min Zeng 0004, Fang Wang 0028, Ruiqing Zheng, Jin Liu 0012, Fei Guo 0001, Min Li 0007 |
BIBM | 7 |
| 2024 | Multi-Task Driven Multi-Level Dynamical Fusion for Single-Cell Multi-Omics Cell Type AnnotationabstractThe emergence of single-cell multi-omics sequencing technology has enabled the simultaneous profiling of diverse omics data within individual cells. It offers a more comprehensive perspective on cellular phenotypes and heterogeneity. However, single-cell multi-omics data are inherently high-dimensional and heterogeneous. Due to technical limitations and scarce starting materials, the data are often affected by noise and dropout effects. To address these challenges, we propose a novel multitask driven multi-level dynamical fusion algorithm for single-cell multi-omics cell type annotation, named scMMDyn. Our approach incorporates reconstruction and classification auxiliary tasks to guide the training of trustworthy modules at both the feature and modality levels. It executes dynamical fusion during these stages and finally achieves cross-modality fusion via an attention mechanism. This method effectively mitigates data quality issues through reconstruction tasks and feature-level dynamical fusion while providing interpretability at both feature and modality levels. Experimental results across diverse single-cell multi-omics datasets show that our method surpasses existing approaches in cell type annotation. Jiawei Li 0018, Shizhan Chen, Zongbo Han, Jijun Tang, Fei Guo 0001 |
BIBM | 6 |
| 2024 | ComLMEss: Combining multiple protein language models enables accurate essential protein predictionabstractAccurately predicting essential proteins is vital for comprehending organism survival, aiding in drug discovery, and informing strategies for treating diseases. While previous computational methods for essential protein prediction have predominantly focused on network-based approaches, recent advancements have seen rapid development in sequence-based prediction methods. However, existing sequence-based prediction methods tend to focus only on sequence-level features, ignoring other biological information at diverse levels. To make use of the diverse information across various biological levels, in this study, we introduce ComLMEss, a novel deep learning framework that combines three protein language models. ComLMEss integrates ProtTrans, ESMFold and OntoProtein, which contain different levels of biological information include protein sequence, conservation, structural, and functional information. ComLMEss employs convolutional neural networks and transformer structure to refine and contextualize the representations from three language models, enabling accurate and robust predictions. Experimental results demonstrate that ComLMEss consistently outperforms existing methods. Ablation studies confirm that the effectiveness of combining different language models focus on different biological information. All results underscore the potential of ComLMEss in essential protein prediction. The source code can be obtained at https://github.com/CSUBioGroup/ComLMEss. Fuhao Zhang, Ruiqing Zheng, Fei Guo 0001, Min Li 0007, Min Zeng 0004 |
BIBM | 5 |
| 2024 | SiamSegNet: A multimodal Segmentation Method Based on Cross-modal Generation for Medical Image Segmentation
Shiqiang Ma, Fei Guo 0001, Jijun Tang |
DASFAA (3) | 2 |
| 2024 | RetroCaptioner: beyond attention in end-to-end retrosynthesis transformer via contrastively captioned learnable graph representationabstractMOTIVATION: Retrosynthesis identifies available precursor molecules for various and novel compounds. With the advancements and practicality of language models, Transformer-based models have increasingly been used to automate this process. However, many existing methods struggle to efficiently capture reaction transformation information, limiting the accuracy and applicability of their predictions. RESULTS: We introduce RetroCaptioner, an advanced end-to-end, Transformer-based framework featuring a Contrastive Reaction Center Captioner. This captioner guides the training of dual-view attention models using a contrastive learning approach. It leverages learned molecular graph representations to capture chemically plausible constraints within a single-step learning process. We integrate the single-encoder, dual-encoder, and encoder-decoder paradigms to effectively fuse information from the sequence and graph representations of molecules. This involves modifying the Transformer encoder into a uni-view sequence encoder and a dual-view module. Furthermore, we enhance the captioning of atomic correspondence between SMILES and graphs. Our proposed method, RetroCaptioner, achieved outstanding performance with 67.2% in top-1 and 93.4% in top-10 exact matched accuracy on the USPTO-50k dataset, alongside an exceptional SMILES validity score of 99.4%. In addition, RetroCaptioner has demonstrated its reliability in generating synthetic routes for the drug protokylol. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://github.com/guofei-tju/RetroCaptioner. Chengwei Ai, Hongpeng Yang, Ruihan Dong, Jijun Tang, Shuangjia Zheng, Fei Guo 0001 |
Bioinform. | 7 |
| 2024 | TranSiam: Aggregating multi-modal visual features with locality for medical image segmentation
Shiqiang Ma, Junhai Xu, Jijun Tang, Shengfeng He, Fei Guo 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Structured Sparse Regularization based Random Vector Functional Link Networks for DNA N4-methylcytosine sites predictionabstractAs an epigenetic modification that plays an important role in modifying gene function and controlling gene expression during cell development, DNA N4-methylcytosine (4mC) is still lack of researching. It is therefore necessary to accurately predict the 4mC sites to make fully aware of its mechanism and function. In this paper, we propose a novel model which is called Structural Sparse Regularized Random Vector Functional Link Network (SSR-RVFL) for predicting 4mC sites. Compared with other state-of-the-art methods, SSR-RVFL performs better and achieves higher prediction accuracy. There are total six benchmark datasets used in the experiments, namely C.elegans, D.elanogaster, E.coli, A.thaliana G.subterraneus and G.pickeringii. Our model improves the accuracy by 0.42%, 0.45%, 0.48%, 0.91%, 0.66% and 0.7% on these six benchmark datasets respectively, so it can be regarded as a more effective prediction tool. Hao Xie 0003, Yijie Ding, Yuqing Qian, Prayag Tiwari, Fei Guo 0001 |
Expert Syst. Appl. | 5 |
| 2024 | Rapid screening of multi-point mutations for enzyme thermostability modification by utilizing computational tools
Jia Jin, Qiaozhen Meng, Min Zeng 0004, Guihua Duan, Ercheng Wang, Fei Guo 0001 |
Future Gener. Comput. Syst. | 6 |
| 2024 | MTMol-GPT: De novo multi-target molecular generation with transformer-based generative adversarial imitation learningabstractDe novo drug design is crucial in advancing drug discovery, which aims to generate new drugs with specific pharmacological properties. Recently, deep generative models have achieved inspiring progress in generating drug-like compounds. However, the models prioritize a single target drug generation for pharmacological intervention, neglecting the complicated inherent mechanisms of diseases, and influenced by multiple factors. Consequently, developing novel multi-target drugs that simultaneously target specific targets can enhance anti-tumor efficacy and address issues related to resistance mechanisms. To address this issue and inspired by Generative Pre-trained Transformers (GPT) models, we propose an upgraded GPT model with generative adversarial imitation learning for multi-target molecular generation called MTMol-GPT. The multi-target molecular generator employs a dual discriminator model using the Inverse Reinforcement Learning (IRL) method for a concurrently multi-target molecular generation. Extensive results show that MTMol-GPT generates various valid, novel, and effective multi-target molecules for various complex diseases, demonstrating robustness and generalization capability. In addition, molecular docking and pharmacophore mapping experiments demonstrate the drug-likeness properties and effectiveness of generated molecules potentially improve neuropsychiatric interventions. Furthermore, our model's generalizability is exemplified by a case study focusing on the multi-targeted drug design for breast cancer. As a broadly applicable solution for multiple targets, MTMol-GPT provides new insight into future directions to enhance potential complex disease therapeutics by generating high-quality multi-target molecules in drug discovery. Chengwei Ai, Hongpeng Yang, Ruihan Dong, Yijie Ding, Fei Guo 0001 |
PLoS Comput. Biol. | 6 |
| 2024 | scRNMF: An imputation method for single-cell RNA-seq data by robust and non-negative matrix factorizationabstractSingle-cell RNA sequencing (scRNA-seq) has emerged as a powerful tool in genomics research, enabling the analysis of gene expression at the individual cell level. However, scRNA-seq data often suffer from a high rate of dropouts, where certain genes fail to be detected in specific cells due to technical limitations. This missing data can introduce biases and hinder downstream analysis. To overcome this challenge, the development of effective imputation methods has become crucial in the field of scRNA-seq data analysis. Here, we propose an imputation method based on robust and non-negative matrix factorization (scRNMF). Instead of other matrix factorization algorithms, scRNMF integrates two loss functions: L2 loss and C-loss. The L2 loss function is highly sensitive to outliers, which can introduce substantial errors. We utilize the C-loss function when dealing with zero values in the raw data. The primary advantage of the C-loss function is that it imposes a smaller punishment for larger errors, which results in more robust factorization when handling outliers. Various datasets of different sizes and zero rates are used to evaluate the performance of scRNMF against other state-of-the-art methods. Our method demonstrates its power and stability as a tool for imputation of scRNA-seq data. Yuqing Qian, Quan Zou 0001, Yi Liu 0112, Fei Guo 0001, Yijie Ding |
PLoS Comput. Biol. | 5 |
| 2024 | Boundary-Aware Dual Biaffine Model for Sequential Sentence Classification in Biomedical DocumentsabstractAssigning appropriate rhetorical roles, such as "background," "intervention," and "outcome," to sentences in biomedical documents can streamline the process for physicians to locate evidence and resources for medical treatment and decision-making. While sequence labeling and span-based methods are frequently employed for this task, the former disregards a document's semantic structure, resulting in a lack of semantic coherence across continuous sentences. Span-based approaches, on the other hand, either necessitate the enumeration of all potential spans, which can be time-consuming, or may lead to the misclassification of sentences over extended spans. Consequently, an approach is required that models the semantic structure of documents explicitly and captures boundary information to achieve precise and effective sentence labeling in biomedical documents. To address these challenges, we propose a new approach, the boundary-aware dual biaffine model, which explicitly models the semantic structure of documents and incorporates boundary information via a dual biaffine layer. We introduce a dynamic programming algorithm to minimize missing labels and overlapping predictions, and achieve globally optimal decoding results. We evaluate our approach on three benchmark datasets, namely PubMed 20 k RCT, PubMed-PICO and NICTA-PIBOSO. The experimental results demonstrate that our approach outperforms strong baselines and achieves state-of-the-art performance on PubMed 20 k RCT and PubMed-PICO. Additionally, our method also achieves competitive results on NICTA-PIBOSO. Junwen Duan, Huai Guo, Fei Guo 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Graph-Based Fusion of Imaging, Genetic and Clinical Data for Degenerative Disease DiagnosisabstractGraph learning methods have achieved noteworthy performance in disease diagnosis due to their ability to represent unstructured information such as inter-subject relationships. While it has been shown that imaging, genetic and clinical data are crucial for degenerative disease diagnosis, existing methods rarely consider how best to use their relationships. How best to utilize information from imaging, genetic and clinical data remains a challenging problem. This study proposes a novel graph-based fusion (GBF) approach to meet this challenge. To extract effective imaging-genetic features, we propose an imaging-genetic fusion module which uses an attention mechanism to obtain modality-specific and joint representations within and between imaging and genetic data. Then, considering the effectiveness of clinical information for diagnosing degenerative diseases, we propose a multi-graph fusion module to further fuse imaging-genetic and clinical features, which adopts a learnable graph construction strategy and a graph ensemble method. Experimental results on two benchmarks for degenerative disease diagnosis (Alzheimers Disease Neuroimaging Initiative and Parkinson's Progression Markers Initiative) demonstrate its effectiveness compared to state-of-the-art graph-based methods. Our findings should help guide further development of graph-based models for dealing with imaging, genetic and clinical data. Rui Guo 0009, Hanhe Lin, Stephen J. McKenna, Hong-Dong Li, Fei Guo 0001, Jin Liu 0012 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | PPRTGI: A Personalized PageRank Graph Neural Network for TF-Target Gene Interaction DetectionabstractTranscription factors (TFs) regulation is required for the vast majority of biological processes in living organisms. Some diseases may be caused by improper transcriptional regulation. Identifying the target genes of TFs is thus critical for understanding cellular processes and analyzing disease molecular mechanisms. Computational approaches can be challenging to employ when attempting to predict potential interactions between TFs and target genes. In this paper, we present a novel graph model (PPRTGI) for detecting TF-target gene interactions using DNA sequence features. Feature representations of TFs and target genes are extracted from sequence embeddings and biological associations. Then, by combining the aggregated node feature with graph structure, PPRTGI uses a graph neural network with personalized PageRank to learn interaction patterns. Finally, a bilinear decoder is applied to predict interaction scores between TF and target gene nodes. We designed experiments on six datasets from different species. The experimental results show that PPRTGI is effective in regulatory interaction inference, with our proposed model achieving an area under receiver operating characteristic score of 93.87% and an area under precision-recall curves score of 88.79% on the human dataset. This paper proposes a new method for predicting TF-target gene interactions, which provides new insights into modeling molecular networks and can thus be used to gain a better understanding of complex biological systems. Jiawei Li 0018, Ibrahim Zamit, Fei Guo 0001, Jijun Tang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | DMAMP: A Deep-Learning Model for Detecting Antimicrobial Peptides and Their Multi-ActivitiesabstractDue to the broad-spectrum and high-efficiency antibacterial activity, antimicrobial peptides (AMPs) and their functions have been studied in the field of drug discovery. Using biological experiments to detect the AMPs and corresponding activities require a high cost, whereas computational technologies do so for much less. Currently, most computational methods solve the identification of AMPs and their activities as two independent tasks, which ignore the relationship between them. Therefore, the combination and sharing of patterns for two tasks is a crucial problem that needs to be addressed. In this study, we propose a deep learning model, called DMAMP, for detecting AMPs and activities simultaneously, which is benefited from multi-task learning. The first stage is to utilize convolutional neural network models and residual blocks to extract the sharing hidden features from two related tasks. The next stage is to use two fully connected layers to learn the distinct information of two tasks. Meanwhile, the original evolutionary features from the peptide sequence are also fed to the predictor of the second task to complement the forgotten information. The experiments on the independent test dataset demonstrate that our method performs better than the single-task model with 4.28% of Matthews Correlation Coefficient (MCC) on the first task, and achieves 0.2627 of an average MCC which is higher than the single-task model and two existing methods for five activities on the second task. To understand whether features derived from the convolutional layers of models capture the differences between target classes, we visualize these high-dimensional features by projecting into 3D space. In addition, we show that our predictor has the ability to identify peptides that achieve activity against Severe Acute Respiratory Syndrome Coronavirus-2 (SARS-CoV-2). We hope that our proposed method can give new insights into the discovery of novel antiviral peptide drugs. Qiaozhen Meng, Genlang Chen, Shixin Zheng, Yulai Lin, Jijun Tang, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2024 | RGCNPPIS: A Residual Graph Convolutional Network for Protein-Protein Interaction Site PredictionabstractAccurate identification of protein-protein interaction (PPI) sites is crucial for understanding the mechanisms of biological processes, developing PPI networks, and detecting protein functions. Currently, most computational methods primarily concentrate on sequence context features and rarely consider the spatial neighborhood features. To address this limitation, we propose a novel residual graph convolutional network for structure-based PPI site prediction (RGCNPPIS). Specifically, we use a GCN module to extract the global structural features from all spatial neighborhoods, and utilize the GraphSage module to extract local structural features from local spatial neighborhoods. To the best of our knowledge, this is the first work utilizing local structural features for PPI site prediction. We also propose an enhanced residual graph connection to combine the initial node representation, local structural features, and the previous GCN layer's node representation, which enables information transfer between layers and alleviates the over-smoothing problem. Evaluation results demonstrate that RGCNPPIS outperforms state-of-the-art methods on three independent test sets. In addition, the results of ablation experiments and case studies confirm that RGCNPPIS is an effective tool for PPI site prediction. Qichang Zhao, Ruikang Zhou, Lishen Zhang, Fei Guo 0001, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | Fuzzy Neural Tangent Kernel Model for Identifying DNA N4-Methylcytosine SitesabstractDNA N4-methylcytosine (4mC) site identification is a crucial field in bioinformatics, where machine learning methods have been effectively utilized. Due to the presence of noise, the existing deep learning methods for detecting 4mC have consistently low recognition rates in positive samples. With fuzzy rules and membership functions, fuzzy systems can achieve good results in processing noisy signals. In contrast to traditional fuzzy systems that lack deep feature representation and sample measurement, we introduce novel techniques to enhance generalization and feature representation. By incorporating the neural tangent kernel (NTK) and kernel learning algorithm into the fuzzy system, we propose the fuzzy NTK (FNTK) model and the radius-based FNTK (R-FNTK) model to predict DNA 4mC sites. To achieve better generalization performance than traditional kernel functions, we first train the NTK for feature representation learning and sample measurement. Based on the membership function and NTK matrix, different fuzzy kernel matrices are constructed for each fuzzy subset of the fuzzy system. Finally, we utilize two types of iterative kernel optimization algorithms to effectively fuse multiple NTK-based fuzzy kernels and obtain the final prediction model. Rigorous testing using six benchmark datasets demonstrates the superiority of our approach, yielding significant improvements in the experiment's performance. Yijie Ding, Prayag Tiwari, Fei Guo 0001, Quan Zou 0001, Weiping Ding 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2024 | Robust Tensor Subspace Learning for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering has represented a significant role in grouping real images. In this study, a novel robust tensor subspace learning (RTSL) is proposed for incomplete multi-view clustering. Specifically, the missing samples within views are first recovered by matrix factorization. The recovered information is utilized for latent representations learning. And then, the obtained latent representations are organized from all views into a third-order tensor and the intrinsic sample relations are captured with tensor linear representation. Moreover, a low-rank sample coefficient tensor is sought to capture high-order connections among views by imposing the tensor nuclear norm. Compared with traditional learning paradigms in the vector space, the sample relations within each view as well as across views could be preserved with the aid of robust tensor subspace learning. As a result, our model can simultaneously handle the missing samples and exploit the intrinsic correlations, leading to enhanced representation capability and better quality of the recovered data. We design an efficient iterative optimization strategy to solve the proposed method. Experimental results on eight datasets show that our model outperforms other competing approaches. Cheng Liang 0001, Daoyuan Wang, Huaxiang Zhang 0001, Shichao Zhang 0001, Fei Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | MixUNet: Mix the 2D and 3D Models for Robust Medical Image SegmentationabstractBrain tumor segmentation is pivotal in the diagnosis and treatment of brain tumors. As functional imaging technologies like CT and MR advance, analyzing 3D medical image data becomes more time-consuming. Several challenges exist in 3D medical image segmentation: 1) 2D networks, when applied to 3D segmentation tasks, suffer from a lack of 3D structural information. 2) Pure 3D networks, due to their vast parameter count and smaller training sample, are susceptible to overfitting. 3) Current 2.5D networks do not fully leverage the available 3D structural information. In this study, we introduce the Mix-UNet, a multi-branch network that synergizes 2D and 3D networks. This design preserves essential 3D structural details for precise segmentation while ensuring computational efficiency. Our model comprises two main branches and a fusion module: a 2D branch for coarse segmentation without 3D structural information, a 3D branch to capture comprehensive 3D structural details, and a fusion module for pixel-level integration to produce the final segmentation. Experimental results demonstrate the model’s ability to reduce parameter count, increase robustness, and maintain high precision. When tested on the BraTS 2020 validation dataset, our model achieved mean dice coefficients of 90.4%, 80.7%, and 71.2% for the whole tumor, tumor core, and enhancing tumor, respectively, with only 2.2M parameters. Jiawei Li 0018, Shizhan Chen, Shiqiang Ma, Fei Guo 0001, Jijun Tang |
BIBM | 4 |
| 2023 | IK-DDI: a novel framework based on instance position embedding and key external text for DDI extractionabstractDetermining drug-drug interactions (DDIs) is an important part of pharmacovigilance and has a vital impact on public health. Compared with drug trials, obtaining DDI information from scientific articles is a faster and lower cost but still a highly credible approach. However, current DDI text extraction methods consider the instances generated from articles to be independent and ignore the potential connections between different instances in the same article or sentence. Effective use of external text data could improve prediction accuracy, but existing methods cannot extract key information from external data accurately and reasonably, resulting in low utilization of external data. In this study, we propose a DDI extraction framework, instance position embedding and key external text for DDI (IK-DDI), which adopts instance position embedding and key external text to extract DDI information. The proposed framework integrates the article-level and sentence-level position information of the instances into the model to strengthen the connections between instances generated from the same article or sentence. Moreover, we introduce a comprehensive similarity-matching method that uses string and word sense similarity to improve the matching accuracy between the target drug and external text. Furthermore, the key sentence search method is used to obtain key information from external data. Therefore, IK-DDI can make full use of the connection between instances and the information contained in external text data to improve the efficiency of DDI extraction. Experimental results show that IK-DDI outperforms existing methods on both macro-averaged and micro-averaged metrics, which suggests our method provides complete framework that can be used to extract relationships between biomedical entities and process external text data. Mingliang Dou, Jiaqi Ding, Genlang Chen, Junwen Duan, Fei Guo 0001, Jijun Tang |
Briefings Bioinform. | 5 |
| 2023 | GraphLncLoc: long non-coding RNA subcellular localization prediction using graph convolutional networks based on sequence to graph transformationabstractThe subcellular localization of long non-coding RNAs (lncRNAs) is crucial for understanding lncRNA functions. Most of existing lncRNA subcellular localization prediction methods use k-mer frequency features to encode lncRNA sequences. However, k-mer frequency features lose sequence order information and fail to capture sequence patterns and motifs of different lengths. In this paper, we proposed GraphLncLoc, a graph convolutional network-based deep learning model, for predicting lncRNA subcellular localization. Unlike previous studies encoding lncRNA sequences by using k-mer frequency features, GraphLncLoc transforms lncRNA sequences into de Bruijn graphs, which transforms the sequence classification problem into a graph classification problem. To extract the high-level features from the de Bruijn graph, GraphLncLoc employs graph convolutional networks to learn latent representations. Then, the high-level feature vectors derived from de Bruijn graph are fed into a fully connected layer to perform the prediction task. Extensive experiments show that GraphLncLoc achieves better performance than traditional machine learning models and existing predictors. In addition, our analyses show that transforming sequences into graphs has more distinguishable features and is more robust than k-mer frequency features. The case study shows that GraphLncLoc can uncover important motifs for nucleus subcellular localization. GraphLncLoc web server is available at http://csuligroup.com:8000/GraphLncLoc/. Min Li 0007, Baoying Zhao, Rui Yin 0002, Chengqian Lu, Fei Guo 0001, Min Zeng 0004 |
Briefings Bioinform. | 5 |
| 2023 | MVML-MPI: Multi-View Multi-Label Learning for Metabolic Pathway InferenceabstractDevelopment of robust and effective strategies for synthesizing new compounds, drug targeting and constructing GEnome-scale Metabolic models (GEMs) requires a deep understanding of the underlying biological processes. A critical step in achieving this goal is accurately identifying the categories of pathways in which a compound participated. However, current machine learning-based methods often overlook the multifaceted nature of compounds, resulting in inaccurate pathway predictions. Therefore, we present a novel framework on Multi-View Multi-Label Learning for Metabolic Pathway Inference, hereby named MVML-MPI. First, MVML-MPI learns the distinct compound representations in parallel with corresponding compound encoders to fully extract features. Subsequently, we propose an attention-based mechanism that offers a fusion module to complement these multi-view representations. As a result, MVML-MPI accurately represents and effectively captures the complex relationship between compounds and metabolic pathways and distinguishes itself from current machine learning-based methods. In experiments conducted on the Kyoto Encyclopedia of Genes and Genomes pathways dataset, MVML-MPI outperformed state-of-the-art methods, demonstrating the superiority of MVML-MPI and its potential to utilize the field of metabolic pathway design, which can aid in optimizing drug-like compounds and facilitating the development of GEMs. The code and data underlying this article are freely available at https://github.com/guofei-tju/MVML-MPI. Contact: [email protected], [email protected] or [email protected]. Hongpeng Yang, Chengwei Ai, Yijie Ding, Fei Guo 0001, Jijun Tang |
Briefings Bioinform. | 5 |
| 2023 | Improved structure-related prediction for insufficient homologous proteins using MSA enhancement and pre-trained language modelabstractIn recent years, protein structure problems have become a hotspot for understanding protein folding and function mechanisms. It has been observed that most of the protein structure works rely on and benefit from co-evolutionary information obtained by multiple sequence alignment (MSA). As an example, AlphaFold2 (AF2) is a typical MSA-based protein structure tool which is famous for its high accuracy. As a consequence, these MSA-based methods are limited by the quality of the MSAs. Especially for orphan proteins that have no homologous sequence, AlphaFold2 performs unsatisfactorily as MSA depth decreases, which may pose a barrier to its widespread application in protein mutation and design problems in which there are no rich homologous sequences and rapid prediction is needed. In this paper, we constructed two standard datasets for orphan and de novo proteins which have insufficient/none homology information, called Orphan62 and Design204, respectively, to fairly evaluate the performance of the various methods in this case. Then, depending on whether or not utilizing scarce MSA information, we summarized two approaches, MSA-enhanced and MSA-free methods, to effectively solve the issue without sufficient MSAs. MSA-enhanced model aims to improve poor MSA quality from the data source by knowledge distillation and generation models. MSA-free model directly learns the relationship between residues on enormous protein sequences from pre-trained models, bypassing the step of extracting the residue pair representation from MSA. Next, we evaluated the performance of four MSA-free methods (trRosettaX-Single, TRFold, ESMFold and ProtT5) and MSA-enhanced (Bagging MSA) method compared with a traditional MSA-based method AlphaFold2, in two protein structure-related prediction tasks, respectively. Comparison analyses show that trRosettaX-Single and ESMFold which belong to MSA-free method can achieve fast prediction ($\sim\! 40$s) and comparable performance compared with AF2 in tertiary structure prediction, especially for short peptides, $\alpha $-helical segments and targets with few homologous sequences. Bagging MSA utilizing MSA enhancement improves the accuracy of our trained base model which is an MSA-based method when poor homology information exists in secondary structure prediction. Our study provides biologists an insight of how to select rapid and appropriate prediction tools for enzyme engineering and peptide drug development. CONTACT: [email protected], [email protected]. Qiaozhen Meng, Fei Guo 0001, Jijun Tang |
Briefings Bioinform. | 2 |
| 2023 | A multi-scale multi-model deep neural network via ensemble strategy on high-throughput microscopy image for protein subcellular localization
Jiaqi Ding, Junhai Xu, Jianguo Wei, Jijun Tang, Fei Guo 0001 |
Expert Syst. Appl. | 5 |
| 2023 | A deep multiple kernel learning-based higher-order fuzzy inference system for identifying DNA N4-methylcytosine sites
Yijie Ding, Prayag Tiwari, Junhai Xu, Wenhuan Lu, Khan Muhammad 0001, Victor Hugo C. de Albuquerque, Fei Guo 0001 |
Inf. Sci. | 8 |
| 2023 | Multi-view unsupervised feature selection with tensor robust principal component analysis and consensus graph learningabstractRecently, multi-view unsupervised feature selection has attracted much attention due to its efficiency and better interpretability in processing high-dimensional multi-view datasets. Most existing methods rely on the constructed similarity matrices to obtain reliable pseudo labels to guide the feature selection. However, the considerable adverse noise in the raw data inevitably impedes the exploration of true underlying similarity structures. Besides, the inter-view correlations are often ignored during the common representation learning , which limits the effective fusion of the essential information from multiple views. To solve these issues, we design a novel robust multi-view unsupervised feature selection framework. Specifically, our method seeks a set of noise-free view-specific similarity matrices by leveraging tensor robust principal component analysis , where the high-order connections among different views are well exploited through the constructed low-rank tensor. Meanwhile, a high-quality consensus similarity matrix is adaptively learned from the view-specific representations within the same unified framework to capture the shared local structures. To enhance the discriminative ability of the feature selection matrix, we further impose a rank constraint on the consensus similarity matrix to obtain reliable pseudo cluster indicators. We present an efficient optimization algorithm ground on the alternating direction method of multipliers to solve the proposed model. Experimental results on six multi-view datasets confirm the superiority of our method. Cheng Liang 0001, Lianzhi Wang, Li Liu 0031, Huaxiang Zhang 0001, Fei Guo 0001 |
Pattern Recognit. | 5 |
| 2023 | Low Rank Matrix Factorization Algorithm Based on Multi-Graph Regularization for Detecting Drug-Disease AssociationabstractDetecting potential associations between drugs and diseases plays an indispensable role in drug development, which has also become a research hotspot in recent years. Compared with traditional methods, some computational approaches have the advantages of fast speed and low cost, which greatly accelerate the progress of predicting the drug-disease association. In this study, we propose a novel similarity-based method of low-rank matrix decomposition based on multi-graph regularization. On the basis of low-rank matrix factorization with$L_{2}$regularization, the multi-graph regularization constraint is constructed by combining a variety of similarity matrices from drugs and diseases respectively. In the experiments, we analyze the difference in the combination of different similarities, resulting that combining all the similarity information on drug space is unnecessary, and only a part of the similarity information can achieve the desired performance. Then our method is compared with other existing models on three data sets (Fdataset, Cdataset and LRSSLdataset) and have a good advantage in the evaluation measurement of AUPR. Besides, a case study experiment is conducted and showing that the superior ability for predicting the potential disease-related drugs of our model. Finally, we compare our model with some methods on six real world datasets, and our model has a good performance in detecting real world data. Chengwei Ai, Hongpeng Yang, Yijie Ding, Jijun Tang, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | Laplacian Regularized Sparse Representation Based Classifier for Identifying DNA N4-Methylcytosine Sites via $L_{2,1/2}$L2,1/2-Matrix NormabstractN4-methylcytosine (4mC) is one of important epigenetic modifications in DNA sequences. Detecting 4mC sites is time-consuming. The computational method based on machine learning has provided effective help for identifying 4mC. To further improve the performance of prediction, we propose a Laplacian Regularized Sparse Representation based Classifier with L2,1/2-matrix norm (LapRSRC). We also utilize kernal trick to derive the kernel LapRSRC for nonlinear modeling. Matrix factorization technology is employed to solve the sparse representation coefficients of all test samples in the training set. And an efficient iterative algorithm is proposed to solve the objective function. We implement our model on six benchmark datasets of 4mC and eight UCI datasets to test evaluate performance. The results show that the performance of our method is better or comparable. Yijie Ding, Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | Multi-View Kernel Sparse Representation for Identification of Membrane Protein TypesabstractMembrane proteins are the main undertaker of biomembrane functions and play a vital role in many biological activities of organisms. Prediction of membrane protein types has a great help in determining the function of proteins and understanding the interactions of membrane proteins. However, the biochemical experiment is expensive and not suitable for the large-scale identification of membrane protein types. Therefore, computational methods were used to improve the efficiency of biological experiments. Most existing computational methods only use a single feature of protein, or use multiple features but do not integrate these well. In our study, the protein sequence is described via three different views (features), including amino acid composition, evolutionary information and physicochemical properties of amino acids. To exploit information among all views (features), we introduce a coupling strategy for Kernel Sparse Representation based Classification (KSRC) and construct a new model called Multi-view KSRC (MvKSRC). We implement our method on 4 benchmark data sets of membrane proteins. The comparison results indicate that our method is much superior to all existing methods. Yuqing Qian, Yijie Ding, Quan Zou 0001, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | BP-DDI: Drug-drug interaction prediction based on biological information and pharmacological textabstractIn the treatment of many diseases, combination drug therapy has been widely used and achieved good clinical efficacy. However, drug-drug interaction (DDI) may occur between multiple drugs and pose a huge threat to the health of patients. Therefore, predicting the presence or absence of DDI among multiple drugs is an important part of pharmacovigilance. Currently, various computational methods for DDI prediction usually use biological information such as molecular structures, targets and enzymes of drugs, or construct heterogeneous networks about drugs, diseases, and genes, so as to obtain abundant information related to drugs. In addition to biological data, pharmacology texts also contain a wealth of information about drug properties, but these texts have not yet been applied to DDI predictions. In this study, we first collect six types of pharmacology texts from DrugBank that can reflect properties of drugs, and propose a novel method named BP-DDI which can combine biological information and pharmacological text to realize DDI event prediction. BP-DDI first extracts biological features (chemical substructure features and target features) from biological data, and then extracts specific types of text features from the collected pharmacology text data. Finally, the biological features are fused with different types of pharmacological text features in order to predict DDI events. Our experiments demonstrate that BP-DDI outperforms existing methods on all three types of prediction tasks. BP-DDI achieves 0.9052 on ACC, and achieves 0.9612 on AUPR. Mingliang Dou, Genlang Chen, Fei Guo 0001, Jijun Tang |
BIBM | 4 |
| 2022 | Integrating Prior Knowledge with Graph Encoder for Gene Regulatory Inference from Single-cell RNA-Seq DataabstractInferring gene regulatory networks based on single-cell transcriptomes is critical for systematically understanding cell-specific regulatory networks and discovering drug targets in tumor cells. Here we show that existing methods mainly perform co-expression analysis and apply the image-based model to deal with the non-euclidean scRNA-seq data, which may not reasonably handle the dropout problem and not fully take advantage of the validated gene regulatory topology. We propose a graph-based end-to-end deep learning model for GRN inference (GRNInfer) with the help of known regulatory relations through transductive learning. The robustness and superiority of the model are demonstrated by comparative experiments. Jiawei Li 0018, Fan Yang 0081, Fang Wang 0028, Yu Rong 0001, Peilin Zhao, Shizhan Chen, Jianhua Yao 0001, Jijun Tang, Fei Guo 0001 |
BIBM | 9 |
| 2022 | Multi-scale Neighborhood Attention Transformer on U-Net for Medical Image SegmentationabstractU-shaped network structures with skip connections played an irreplaceable role in medical image analysis, but the limitation of convolution makes it unable to learn long-distance semantic information well. The recent success of Transformer in natural language processing and image classification shows that it can benefit from global information modeling by using self-attention mechanisms. However, both local and global features are equally important for dense prediction tasks. Transformer ignores local semantic information to a certain extent. In this study, we propose a Unet-like Transformer for medical image segmentation, named MN-Unet, which can simultaneously extract local and global features. MN-Unet consists of encoder, decoder, and skip connections. Specially, we design an encoder based on the Neighborhood Attention Transformer, which fuse three neighborhood sizes of different dimensions to simultaneously extract local and global features. In the decoder, we use bilinear interpolation to restore the image to its original size. Skip connection is added to alleviate the distortion of low resolution to high resolution. MN-Unet can achieve accurate segmentation of medical images without any pre-training. Extensive experimental results on two medical image datasets (LiTS 2017 and BraTS 2020) show that we achieve relatively better performance than state-of-the-art methods. The codes and trained models will be publicly available a https://github.com/hutchinsonian/MN_Unet Nanxing Zhang, Shiqiang Ma, Jijun Tang, Fei Guo 0001 |
BIBM | 6 |
| 2022 | Identification of protein-nucleotide binding residues via graph regularized k-local hyperplane distance nearest neighbor model
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Appl. Intell. | 4 |
| 2022 | Identification of drug-target interactions via multiple kernel-based triple collaborative matrix factorizationabstractTargeted drugs have been applied to the treatment of cancer on a large scale, and some patients have certain therapeutic effects. It is a time-consuming task to detect drug-target interactions (DTIs) through biochemical experiments. At present, machine learning (ML) has been widely applied in large-scale drug screening. However, there are few methods for multiple information fusion. We propose a multiple kernel-based triple collaborative matrix factorization (MK-TCMF) method to predict DTIs. The multiple kernel matrices (contain chemical, biological and clinical information) are integrated via multi-kernel learning (MKL) algorithm. And the original adjacency matrix of DTIs could be decomposed into three matrices, including the latent feature matrix of the drug space, latent feature matrix of the target space and the bi-projection matrix (used to join the two feature spaces). To obtain better prediction performance, MKL algorithm can regulate the weight of each kernel matrix according to the prediction error. The weights of drug side-effects and target sequence are the highest. Compared with other computational methods, our model has better performance on four test data sets. Yijie Ding, Jijun Tang, Fei Guo 0001, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2022 | Microbe-bridged disease-metabolite associations identification by heterogeneous graph fusionabstractMOTIVATION: Metabolomics has developed rapidly in recent years, and metabolism-related databases are also gradually constructed. Nowadays, more and more studies are being carried out on diverse microbes, metabolites and diseases. However, the logics of various associations among microbes, metabolites and diseases are limited understanding in the biomedicine of gut microbial system. The collection and analysis of relevant microbial bioinformation play an important role in the revelation of microbe-metabolite-disease associations. Therefore, the dataset that integrates multiple relationships and the method based on complex heterogeneous graphs need to be developed. RESULTS: In this study, we integrated some databases and extracted a variety of associations data among microbes, metabolites and diseases. After obtaining the three interconnected bilateral association data (microbe-metabolite, metabolite-disease and disease-microbe), we considered building a heterogeneous graph to describe the association data. In our model, microbes were used as a bridge between diseases and metabolites. In order to fuse the information of disease-microbe-metabolite graph, we used the bipartite graph attention network on the disease-microbe and metabolite-microbe bipartite graph. The experimental results show that our model has good performance in the prediction of various disease-metabolite associations. Through the case study of type 2 diabetes mellitus, Parkinson's disease, inflammatory bowel disease and liver cirrhosis, it is noted that our proposed methodology are valuable for the mining of other associations and the prediction of biomarkers for different human diseases.Availability and implementation: https://github.com/Selenefreeze/DiMiMe.git. Jitong Feng, Shengbo Wu, Hongpeng Yang, Chengwei Ai, Jianjun Qiao, Junhai Xu, Fei Guo 0001 |
Briefings Bioinform. | 7 |
| 2022 | Two-stage-vote ensemble framework based on integration of mutation data and gene interaction network for uncovering driver genesabstractIdentifying driver genes, exactly from massive genes with mutations, promotes accurate diagnosis and treatment of cancer. In recent years, a lot of works about uncovering driver genes based on integration of mutation data and gene interaction networks is gaining more attention. However, it is in suspense if it is more effective for prioritizing driver genes when integrating various types of mutation information (frequency and functional impact) and gene networks. Hence, we build a two-stage-vote ensemble framework based on somatic mutations and mutual interactions. Specifically, we first represent and combine various kinds of mutation information, which are propagated through networks by an improved iterative framework. The first vote is conducted on iteration results by voting methods, and the second vote is performed to get ensemble results of the first poll for the final driver gene list. Compared with four excellent previous approaches, our method has better performance in identifying driver genes on $33$ types of cancer from The Cancer Genome Atlas. Meanwhile, we also conduct a comparative analysis about two kinds of mutation information, five gene interaction networks and four voting strategies. Our framework offers a new view for data integration and promotes more latent cancer genes to be admitted. Yingxin Kan, Limin Jiang, Jijun Tang, Fei Guo 0001 |
Briefings Bioinform. | 5 |
| 2022 | Identification of drug-side effect association via restricted Boltzmann machines with penalized termabstractIn the entire life cycle of drug development, the side effect is one of the major failure factors. Severe side effects of drugs that go undetected until the post-marketing stage leads to around two million patient morbidities every year in the United States. Therefore, there is an urgent need for a method to predict side effects of approved drugs and new drugs. Following this need, we present a new predictor for finding side effects of drugs. Firstly, multiple similarity matrices are constructed based on the association profile feature and drug chemical structure information. Secondly, these similarity matrices are integrated by Centered Kernel Alignment-based Multiple Kernel Learning algorithm. Then, Weighted K nearest known neighbors is utilized to complement the adjacency matrix. Next, we construct Restricted Boltzmann machines (RBM) in drug space and side effect space, respectively, and apply a penalized maximum likelihood approach to train model. At last, the average decision rule was adopted to integrate predictions from RBMs. Comparison results and case studies demonstrate, with four benchmark datasets, that our method can give a more accurate and reliable prediction result. Yuqing Qian, Yijie Ding, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2022 | HDIContact: a novel predictor of residue-residue contacts on hetero-dimer interfaces via sequential information and transfer learning strategyabstractProteins maintain the functional order of cell in life by interacting with other proteins. Determination of protein complex structural information gives biological insights for the research of diseases and drugs. Recently, a breakthrough has been made in protein monomer structure prediction. However, due to the limited number of the known protein structure and homologous sequences of complexes, the prediction of residue-residue contacts on hetero-dimer interfaces is still a challenge. In this study, we have developed a deep learning framework for inferring inter-protein residue contacts from sequential information, called HDIContact. We utilized transfer learning strategy to produce Multiple Sequence Alignment (MSA) two-dimensional (2D) embedding based on patterns of concatenated MSA, which could reduce the influence of noise on MSA caused by mismatched sequences or less homology. For MSA 2D embedding, HDIContact took advantage of Bi-directional Long Short-Term Memory (BiLSTM) with two-channel to capture 2D context of residue pairs. Our comprehensive assessment on the Escherichia coli (E. coli) test dataset showed that HDIContact outperformed other state-of-the-art methods, with top precision of 65.96%, the Area Under the Receiver Operating Characteristic curve (AUROC) of 83.08% and the Area Under the Precision Recall curve (AUPR) of 25.02%. In addition, we analyzed the potential of HDIContact for human-virus protein-protein complexes, by achieving top five precision of 80% on O75475-P04584 related to Human Immunodeficiency Virus. All experiments indicated that our method was a valuable technical tool for predicting inter-protein residue contacts, which would be helpful for understanding protein-protein interaction mechanisms. Qiaozhen Meng, Jianxin Wang 0001, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2022 | A hybrid deep learning framework for gene regulatory network inference from single-cell transcriptomic dataabstractInferring gene regulatory networks (GRNs) based on gene expression profiles is able to provide an insight into a number of cellular phenotypes from the genomic level and reveal the essential laws underlying various life phenomena. Different from the bulk expression data, single-cell transcriptomic data embody cell-to-cell variance and diverse biological information, such as tissue characteristics, transformation of cell types, etc. Inferring GRNs based on such data offers unprecedented advantages for making a profound study of cell phenotypes, revealing gene functions and exploring potential interactions. However, the high sparsity, noise and dropout events of single-cell transcriptomic data pose new challenges for regulation identification. We develop a hybrid deep learning framework for GRN inference from single-cell transcriptomic data, DGRNS, which encodes the raw data and fuses recurrent neural network and convolutional neural network (CNN) to train a model capable of distinguishing related gene pairs from unrelated gene pairs. To overcome the limitations of such datasets, it applies sliding windows to extract valuable features while preserving the direction of regulation. DGRNS is constructed as a deep learning model containing gated recurrent unit network for exploring time-dependent information and CNN for learning spatially related information. Our comprehensive and detailed comparative analysis on the dataset of mouse hematopoietic stem cells illustrates that DGRNS outperforms state-of-the-art methods. The networks inferred by DGRNS are about 16% higher than the area under the receiver operating characteristic curve of other unsupervised methods and 10% higher than the area under the precision recall curve of other supervised methods. Experiments on human datasets show the strong robustness and excellent generalization of DGRNS. By comparing the predictions with standard network, we discover a series of novel interactions which are proved to be true in some specific cell types. Importantly, DGRNS identifies a series of regulatory relationships with high confidence and functional consistency, which have not yet been experimentally confirmed and merit further research. Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 5 |
| 2022 | Inferring gene regulatory network via fusing gene expression image and RNA-seq dataabstractMOTIVATION: Recently, with the development of high-throughput experimental technology, reconstruction of gene regulatory network (GRN) has ushered in new opportunities and challenges. Some previous methods mainly extract gene expression information based on RNA-seq data, but the associated information is very limited. With the establishment of gene expression image database, it is possible to infer GRN from image data with rich spatial information. RESULTS: First, we propose a new convolutional neural network (called SDINet), which can extract gene expression information from images and identify the interaction between genes. SDINet can obtain the detailed information and high-level semantic information from the images well. And it can achieve satisfying performance on image data (Acc: 0.7196, F1: 0.7374). Second, we apply the idea of our SDINet to build an RNA-model, which also achieves good results on RNA-seq data (Acc: 0.8962, F1: 0.8950). Finally, we combine image data and RNA-seq data, and design a new fusion network to explore the potential relationship between them. Experiments show that our proposed network fusing two modalities can obtain satisfying performance (Acc: 0.9116, F1: 0.9118) than any single data. AVAILABILITY AND IMPLEMENTATION: Data and code are available from https://github.com/guofei-tju/Combine-Gene-Expression-images-and-RNA-seq-data-For-infering-GRN. Shiqiang Ma, Jin Liu 0012, Jijun Tang, Fei Guo 0001 |
Bioinform. | 5 |
| 2022 | String kernels construction and fusion: a survey with bioinformatics application
Ren Qi, Fei Guo 0001, Quan Zou 0001 |
Frontiers Comput. Sci. | 2 |
| 2022 | A multi-layer multi-kernel neural network for determining associations between non-coding RNAs and diseases
Chengwei Ai, Hongpeng Yang, Yijie Ding, Jijun Tang, Fei Guo 0001 |
Neurocomputing | 5 |
| 2022 | Sparse regularized joint projection model for identifying associations of non-coding RNAs and human diseasesabstractCurrent human biomedical research shows that human diseases are closely related to non-coding RNAs, so it is of great significance for human medicine to study the relationship between diseases and non-coding RNAs. Current research has found associations between non-coding RNAs and human diseases through a variety of effective methods, but most of the methods are complex and targeted at a single RNA or disease. Therefore, we urgently need an effective and simple method to discover the associations between non-coding RNAs and human diseases. In this paper, we propose a sparse regularized joint projection model (SRJP) to identify the associations between non-coding RNAs and diseases. First, we extract information through a series of ncRNA similarity matrices and disease similarity matrices and assign average weights to the similarity matrices of the two sides. Then we decompose the similarity matrices of the two spaces into low-rank matrices and put them into SRJP. In SRJP, we innovatively use the projection matrix to combine the ncRNA side and the disease side to identify the associations between ncRNAs and diseases. Finally, the regularization term in SRJP effectively improves the robustness and generalization ability of the model. We test our model on different datasets involving three types of ncRNAs: circRNA, microRNA and long non-coding RNA. The experimental results show that SRJP has superior ability to identify and predict the associations between ncRNAs and diseases. Prayag Tiwari, Junhai Xu, Yuqing Qian, Chengwei Ai, Yijie Ding, Fei Guo 0001 |
Knowl. Based Syst. | 7 |
| 2022 | Inferring human microbe-drug associations via multiple kernel fusion on graph neural network
Hongpeng Yang, Yijie Ding, Jijun Tang, Fei Guo 0001 |
Knowl. Based Syst. | 4 |
| 2022 | Res2Unet: A multi-scale channel attention network for retinal vessel segmentation
Jiaqi Ding, Jijun Tang, Fei Guo 0001 |
Neural Comput. Appl. | 4 |
| 2022 | Shared subspace-based radial basis function neural network for identifying ncRNAs subcellular localizationabstractNon-coding RNAs (ncRNAs) play an important role in revealing the mechanism of human disease for anti-tumor and anti-virus substances. Detecting subcellular locations of ncRNAs is a necessary way to study ncRNA. Traditional biochemical methods are time-consuming and labor-intensive, and computational-based methods can help detect the location of ncRNAs on a large scale. However, many models did not consider the correlation information among multiple subcellular localizations of ncRNAs. This study proposes a radial basis function neural network based on shared subspace learning (RBFNN-SSL), which extract shared structures in multi-labels. To evaluate performance, our classifier is tested on three ncRNA datasets. Our model achieves better performance in experimental results. Yijie Ding, Prayag Tiwari, Fei Guo 0001, Quan Zou 0001 |
Neural Networks | 3 |
| 2022 | DeepFusionDTA: Drug-Target Binding Affinity Prediction With Information Fusion and Hybrid Deep-Learning Ensemble ModelabstractIdentification of drug-target interaction (DTI) is the most important issue in the broad field of drug discovery. Using purely biological experiments to verify drug-target binding profiles takes lots of time and effort, so computational technologies for this task obviously have great benefits in reducing the drug search space. Most of computational methods to predict DTI are proposed to solve a binary classification problem, which ignore the influence of binding strength. Therefore, drug-target binding affinity prediction is still a challenging issue. Currently, lots of studies only extract sequence information that lacks feature-rich representation, but we consider more spatial features in order to merge various data in drug and target spaces. In this study, we propose a two-stage deep neural network ensemble model for detecting drug-target binding affinity, called DeepFusionDTA, via various information analysis modules. First stage is to utilize sequence and structure information to generate fusion feature map of candidate protein and drug pair through various analysis modules based deep learning. Second stage is to apply bagging-based ensemble learning strategy for regression prediction, and we obtain outstanding results by combining the advantages of various algorithms in efficient feature abstraction and regression calculation. Importantly, we evaluate our novel method, DeepFusionDTA, which delivers 1.5 percent CI increase on KIBA dataset and 1.0 percent increase on Davis dataset, by comparing with existing prediction tools, DeepDTA. Furthermore, the ideas we have offered can be applied to in-silico screening of the interaction space, to provide novel DTIs which can be experimentally pursued. The codes and data are available from https://github.com/guofei-tju/DeepFusionDTA. Yuqian Pu, Jiawei Li 0018, Jijun Tang, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | Identify ncRNA Subcellular Localization via Graph Regularized $k$k-Local Hyperplane Distance Nearest Neighbor Model on Multi-Kernel LearningabstractNon-coding RNAs (ncRNAs) are a type of RNAs which are not used to encode protein sequences. Emerging evidence shows that lots of ncRNAs may participate in many biological processes and must be widely involved in many types of cancers. Therefore, understanding their functionality is of great importance. Similar to proteins, various functions of ncRNAs relies on their subcellular localizations. Traditional high-throughput methods in wet-lab to identify subcellular localization is time-consuming and costly. In this paper, we propose a novel computational method based on multi-kernel learning to identify multi-label ncRNA subcellular localizations, via graph regularized k-local hyperplane distance nearest neighbor algorithm. First, we construct six types of sequence-based feature descriptors and select important feature vectors. Then, we build a multi-kernel learning model with Hilbert-Schmidt independence criterion (HSIC) to obtain optimal weights for vairous features. Furthermore, we propose the graph regularized k-local hyperplane distance nearest neighbor algorithm (GHKNN) as a binary classification model for detecting one kind of non-coding RNA subcellular localization. Finally, we apply One-vs-Rest strategy to decompose multi-label problem of non-coding RNA subcellular localizations. Our method achieves excellent performance on three ncRNA datasets and three human ncRNA datasets, and out-performs other outstanding machine learning methods. Comparing to existing method, our model also performs well especially on small datasets. We expect that this model will be useful for the prediction of subcellular localization and the study of important functional mechanisms of ncRNAs. Furthermore, we establish user-friendly web server (http://ncrna.lbci.net/) with the implementation of our method, which can be easily used by most experimental scientists. Haohao Zhou, Jijun Tang, Yijie Ding, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | C-Loss Based Higher Order Fuzzy Inference Systems for Identifying DNA N4-Methylcytosine SitesabstractDNA methylation is an epigenetic marker that plays an important role in the biological processes of regulating gene expression, maintaining chromatin structure, imprinting genes, inactivating X chromosomes, and developing embryos. The traditional detection method is time-consuming. Currently, researchers have used effective computational methods to improve the efficiency of methylation detection. This study proposes a fuzzy model with correntropy induced loss (C-loss) function to identify DNA N4-methylcytosine (4 mC) sites. To improve the robustness and performance of the model, we use kernel method and the C-loss function to build a higher order fuzzy inference systems. To test performance, our model is implemented on six 4 mC and eight University of California Irvine (UCI) datasets. The experimental results show that our model achieves better prediction performance. Yijie Ding, Prayag Tiwari, Quan Zou 0001, Fei Guo 0001, Hari Mohan Pandey |
IEEE Trans. Fuzzy Syst. | 4 |
| 2021 | Document-level DDI relation extraction with document-entity embeddingabstractDDI is an important part of drug-related research and pharmacovigilance. Extracting DDI information from scientific literature has become a low-cost and highly reliable way. Currently, existing works are all sentence-level DDI relation extraction. In fact, the entity relationship is often expressed by multiple sentences. Moreover, the sentence-level DDI relation extraction also causes a large amount of redundancy in the whole dataset with increasing in negative instance data. In this study, we propose a document-level DDI relation extraction method based on document-entity embedding. Our method performs special processing on the DDI Extraction 2013 for the first time, in order to calculate document-level relation extraction. For obtaining document-level entity information, we propose a document-entity embedding method to integrate the information of all same drugs in the same article. The experimental results show that the processing of DDI Extraction 2013 dataset is reasonable. In addition, the proposed method has achieved good performance on document-level DDI dataset, and the best F1 score is 62.51%. This is the first time that DDI Extraction 2013 has been processed into a document-level dataset, and document-level DDI relation extraction has been realized. Mingliang Dou, Jijun Tang, Fei Guo 0001 |
BIBM | 3 |
| 2021 | MIASNet: A medical image segmentation method predicting future based on past and current casesabstractFast and accurate segmentation of medical images is essential for the diagnosis and treatment of diseases. The automatic segmentation technology based on deep learning has achieved encouraging performance in segmentation accuracy. However, the improvement of segmentation accuracy usually requires a larger network structure, which also leads to a decrease in segmentation speed. In this study, we propose a medical image anticipation segmentation net (MIASNet), in order to further improve the segmentation speed under the premise of excellent segmentation accuracy. For 3D medical images, we use the spatial association of the previous frame and the current frame as input data to predict the segmentation results of the next frame. Our approach consists of three lightweight sub-networks, which are used to learn the mapping relationship of the spatial domain. In order to make full use of the generating ability of the deep learning network, we use group convolution to obtain diversified prediction results. On the multimodal brain tumor image segmentation (BraTS) 2020 dataset, MIASNet achieves excellent segmentation accuracy without using the target frame that need to be segmented as the network input. Therefore, our proposed segmentation network can be used in a wider range of real-time medical applications. Shiqiang Ma, Jijun Tang, Fei Guo 0001 |
BIBM | 4 |
| 2021 | GEU-Net: Rethinking the information transmission in the skip connection of U-Net architectureabstractWith the wide application of deep learning technology in medical image processing, the performance of medical image segmentation has been improved in a breakthrough. U-Net architecture has excellent performance in medical image segmentation tasks. In order to solve the problem of image signal loss caused by the autoencoder structure, U-Net has added skip connections to its network to transfer the low-level features of the encoder path to the decoder path. Although this method can roughly solve the problem of image information loss, while it introduces a new problem, that is, the simple feature fusion method causes the high-level semantic information to be diluted. In order to solve the problem that the simple fusion of low-level edge information and high-level semantic information creates the semantic gap and dilutes high-level semantic information, we propose a novel U-shaped architecture, namely GEU-Net. GEU-Net utilizes ensemble learning methods to obtain better segmentation performance with a small computational cost. In addition, We propose a multi-scale group convolution block namely Group Residual (GR) module to reduce the semantic gap between encoder and decoder. We have evaluated our model on the BraTS 2020 Challenge, and have achieved competitive segmentation results. Shiqiang Ma, Jijun Tang, Fei Guo 0001 |
BIBM | 5 |
| 2021 | Multi-AMP: detecting the antimicrobial peptides and their activities using the multi-task learningabstractRecently due to the broad-spectrum and high-efficiency antibacterial activity, antimicrobial peptides (AMPs) have become the best alternative to antibiotics. With the rapid increase of the antibacterial peptides, many computational methods have been developed to identify the AMPs and their specific antibacterial activities. However, most existing methods regard these two problems as independent sub-problems and ignore the correlation between tasks. In this paper, we propose a method, Multi-AMP, which utilizes multi-task learning and solves two tasks simultaneously: 1) whether a given peptide is AMP, 2) which activities it performs. The two tasks share the parameters at the bottom layers of the model and learn the specific information at the top layers. Experiments indicate that our multi-task model performs better than single-task models and two existing predictors, which can give insights to the drug discovery process. Qiaozhen Meng, Jijun Tang, Fei Guo 0001 |
BIBM | 3 |
| 2021 | A Zero-Shot Method for 3D Medical Image SegmentationabstractAccurate automatic medical image segmentation technology plays an important role for the diagnosis and treatment of brain tumor. However, existing methods based on outstanding 2.5D and 3D segmentation strategies are time-consumption and hardware-consumption while ensuring high accuracy. In order to reduce the high demand for automatic segmentation of tumor images and avoid the noise interference in a single input image, we propose an end-to-end zero-shot CNN segmentation method. Our method only utilizes two adjacent images, instead of the target image, as the input data of deep neural network to predict the brain tumor area in the target image. Avoiding noise interference in the target image, this method makes full use of the spatial context feature between adjacent slices in order to obtain accurate zero-shot segmentation results. We compare with the state-of-the-art segmentation frameworks on the same benchmark and notice that our method has strong competitiveness. Shiqiang Ma, Jijun Tang, Fei Guo 0001 |
ICME | 4 |
| 2021 | MMFGRN: a multi-source multi-model fusion method for gene regulatory network reconstructionabstractLots of biological processes are controlled by gene regulatory networks (GRNs), such as growth and differentiation of cells, occurrence and development of the diseases. Therefore, it is important to persistently concentrate on the research of GRN. The determination of the gene-gene relationships from gene expression data is a complex issue. Since it is difficult to efficiently obtain the regularity behind the gene-gene relationship by only relying on biochemical experimental methods, thus various computational methods have been used to construct GRNs, and some achievements have been made. In this paper, we propose a novel method MMFGRN (for "Multi-source Multi-model Fusion for Gene Regulatory Network reconstruction") to reconstruct the GRN. In order to make full use of the limited datasets and explore the potential regulatory relationships contained in different data types, we construct the MMFGRN model from three perspectives: single time series data model, single steady-data model and time series and steady-data joint model. And, we utilize the weighted fusion strategy to get the final global regulatory link ranking. Finally, MMFGRN model yields the best performance on the DREAM4 InSilico_Size10 data, outperforming other popular inference algorithms, with an overall area under receiver operating characteristic score of 0.909 and area under precision-recall (AUPR) curves score of 0.770 on the 10-gene network. Additionally, as the network scale increases, our method also has certain advantages with an overall AUPR score of 0.335 on the DREAM4 InSilico_Size100 data. These results demonstrate the good robustness of MMFGRN on different scales of networks. At the same time, the integration strategy proposed in this paper provides a new idea for the reconstruction of the biological network model without prior knowledge, which can help researchers to decipher the elusive mechanism of life. Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2021 | Predicting MHC class I binder: existing approaches and a novel recurrent neural network solutionabstractMajor histocompatibility complex (MHC) possesses important research value in the treatment of complex human diseases. A plethora of computational tools has been developed to predict MHC class I binders. Here, we comprehensively reviewed 27 up-to-date MHC I binding prediction tools developed over the last decade, thoroughly evaluating feature representation methods, prediction algorithms and model training strategies on a benchmark dataset from Immune Epitope Database. A common limitation was identified during the review that all existing tools can only handle a fixed peptide sequence length. To overcome this limitation, we developed a bilateral and variable long short-term memory (BVLSTM)-based approach, named BVLSTM-MHC. It is the first variable-length MHC class I binding predictor. In comparison to the 10 mainstream prediction tools on an independent validation dataset, BVLSTM-MHC achieved the best performance in six out of eight evaluated metrics. A web server based on the BVLSTM-MHC model was developed to enable accurate and efficient MHC class I binder prediction in human, mouse, macaque and chimpanzee. Limin Jiang, Jiawei Li 0018, Jijun Tang, Fei Guo 0001 |
Briefings Bioinform. | 6 |
| 2021 | DeepATT: a hybrid category attention neural network for identifying functional effects of DNA sequencesabstractQuantifying DNA properties is a challenging task in the broad field of human genomics. Since the vast majority of non-coding DNA is still poorly understood in terms of function, this task is particularly important to have enormous benefit for biology research. Various DNA sequences should have a great variety of representations, and specific functions may focus on corresponding features in the front part of learning model. Currently, however, for multi-class prediction of non-coding DNA regulatory functions, most powerful predictive models do not have appropriate feature extraction and selection approaches for specific functional effects, so that it is difficult to gain a better insight into their internal correlations. Hence, we design a category attention layer and category dense layer in order to select efficient features and distinguish different DNA functions. In this study, we propose a hybrid deep neural network method, called DeepATT, for identifying $919$ regulatory functions on nearly $5$ million DNA sequences. Our model has four built-in neural network constructions: convolution layer captures regulatory motifs, recurrent layer captures a regulatory grammar, category attention layer selects corresponding valid features for different functions and category dense layer classifies predictive labels with selected features of regulatory functions. Importantly, we compare our novel method, DeepATT, with existing outstanding prediction tools, DeepSEA and DanQ. DeepATT performs significantly better than other existing tools for identifying DNA functions, at least increasing $1.6\%$ area under precision recall. Furthermore, we can mine the important correlation among different DNA functions according to the category attention module. Moreover, our novel model can greatly reduce the number of parameters by the mechanism of attention and locally connected, on the basis of ensuring accuracy. Jiawei Li 0018, Yuqian Pu, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 5 |
| 2021 | EP3: an ensemble predictor that accurately identifies type III secreted effectorsabstractType III secretion systems (T3SS) can be found in many pathogenic bacteria, such as Dysentery bacillus, Salmonella typhimurium, Vibrio cholera and pathogenic Escherichia coli. The routes of infection of these bacteria include the T3SS transferring a large number of type III secreted effectors (T3SE) into host cells, thereby blocking or adjusting the communication channels of the host cells. Therefore, the accurate identification of T3SEs is the precondition for the further study of pathogenic bacteria. In this article, a new T3SEs ensemble predictor was developed, which can accurately distinguish T3SEs from any unknown protein. In the course of the experiment, methods and models are strictly trained and tested. Compared with other methods, EP3 demonstrates better performance, including the absence of overfitting, strong robustness and powerful predictive ability. EP3 (an ensemble predictor that accurately identifies T3SEs) is designed to simplify the user's (especially nonprofessional users) access to T3SEs for further investigation, which will have a significant impact on understanding the progression of pathogenic bacterial infections. Based on the integrated model that we proposed, a web server had been established to distinguish T3SEs from non-T3SEs, where have EP3_1 and EP3_2. The users can choose the model according to the species of the samples to be tested. Our related tools and data can be accessed through the link http://lab.malab.cn/∼lijing/EP3.html. Leyi Wei, Fei Guo 0001, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2021 | SubLocEP: a novel ensemble predictor of subcellular localization of eukaryotic mRNA based on machine learningabstractMOTIVATION: mRNA location corresponds to the location of protein translation and contributes to precise spatial and temporal management of the protein function. However, current assignment of subcellular localization of eukaryotic mRNA reveals important limitations: (1) turning multiple classifications into multiple dichotomies makes the training process tedious; (2) the majority of the models trained by classical algorithm are based on the extraction of single sequence information; (3) the existing state-of-the-art models have not reached an ideal level in terms of prediction and generalization ability. To achieve better assignment of subcellular localization of eukaryotic mRNA, a better and more comprehensive model must be developed. RESULTS: In this paper, SubLocEP is proposed as a two-layer integrated prediction model for accurate prediction of the location of sequence samples. Unlike the existing models based on limited features, SubLocEP comprehensively considers additional feature attributes and is combined with LightGBM to generated single feature classifiers. The initial integration model (single-layer model) is generated according to the categories of a feature. Subsequently, two single-layer integration models are weighted (sequence-based: physicochemical properties = 3:2) to produce the final two-layer model. The performance of SubLocEP on independent datasets is sufficient to indicate that SubLocEP is an accurate and stable prediction model with strong generalization ability. Additionally, an online tool has been developed that contains experimental data and can maximize the user convenience for estimation of subcellular localization of eukaryotic mRNA. Shida He, Fei Guo 0001, Quan Zou 0001 |
Briefings Bioinform. | 4 |
| 2021 | A spectral clustering with self-weighted multiple kernel learning method for single-cell RNA-seq dataabstractSingle-cell RNA-sequencing (scRNA-seq) data widely exist in bioinformatics. It is crucial to devise a distance metric for scRNA-seq data. Almost all existing clustering methods based on spectral clustering algorithms work in three separate steps: similarity graph construction; continuous labels learning; discretization of the learned labels by k-means clustering. However, this common practice has potential flaws that may lead to severe information loss and degradation of performance. Furthermore, the performance of a kernel method is largely determined by the selected kernel; a self-weighted multiple kernel learning model can help choose the most suitable kernel for scRNA-seq data. To this end, we propose to automatically learn similarity information from data. We present a new clustering method in the form of a multiple kernel combination that can directly discover groupings in scRNA-seq data. The main proposition is that automatically learned similarity information from scRNA-seq data is used to transform the candidate solution into a new solution that better approximates the discrete one. The proposed model can be efficiently solved by the standard support vector machine (SVM) solvers. Experiments on benchmark scRNA-Seq data validate the superior performance of the proposed model. Spectral clustering with multiple kernels is implemented in Matlab, licensed under Massachusetts Institute of Technology (MIT) and freely available from the Github website, https://github.com/Cuteu/SMSC/. Ren Qi, Jin Wu 0002, Fei Guo 0001, Lei Xu 0047, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2021 | Exploring associations of non-coding RNAs in human diseases via three-matrix factorization with hypergraph-regular terms on center kernel alignmentabstractRelationship of accurate associations between non-coding RNAs and diseases could be of great help in the treatment of human biomedical research. However, the traditional technology is only applied on one type of non-coding RNA or a specific disease, and the experimental method is time-consuming and expensive. More computational tools have been proposed to detect new associations based on known ncRNA and disease information. Due to the ncRNAs (circRNAs, miRNAs and lncRNAs) having a close relationship with the progression of various human diseases, it is critical for developing effective computational predictors for ncRNA-disease association prediction. In this paper, we propose a new computational method of three-matrix factorization with hypergraph regularization terms (HGRTMF) based on central kernel alignment (CKA), for identifying general ncRNA-disease associations. In the process of constructing the similarity matrix, various types of similarity matrices are applicable to circRNAs, miRNAs and lncRNAs. Our method achieves excellent performance on five datasets, involving three types of ncRNAs. In the test, we obtain best area under the curve scores of $0.9832$, $0.9775$, $0.9023$, $0.8809$ and $0.9185$ via 5-fold cross-validation and $0.9832$, $0.9836$, $0.9198$, $0.9459$ and $0.9275$ via leave-one-out cross-validation on five datasets. Furthermore, our novel method (CKA-HGRTMF) is also able to discover new associations between ncRNAs and diseases accurately. Availability: Codes and data are available: https://github.com/hzwh6910/ncRNA2Disease.git. Contact:[email protected]. Jijun Tang, Yijie Ding, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2021 | Exploring effectiveness of ab-initio protein-protein docking methods on a novel antibacterial protein complex datasetabstractDiseases caused by bacterial infections become a critical problem in public heath. Antibiotic, the traditional treatment, gradually loses their effectiveness due to the resistance. Meanwhile, antibacterial proteins attract more attention because of broad spectrum and little harm to host cells. Therefore, exploring new effective antibacterial proteins is urgent and necessary. In this paper, we are committed to evaluating the effectiveness of ab-initio docking methods in antibacterial protein-protein docking. For this purpose, we constructed a three-dimensional (3D) structure dataset of antibacterial protein complex, called APCset, which contained $19$ protein complexes whose receptors or ligands are homologous to antibacterial peptides from Antimicrobial Peptide Database. Then we selected five representative ab-initio protein-protein docking tools including ZDOCK3.0.2, FRODOCK3.0, ATTRACT, PatchDock and Rosetta to identify these complexes' structure, whose performance differences were obtained by analyzing from five aspects, including top/best pose, first hit, success rate, average hit count and running time. Finally, according to different requirements, we assessed and recommended relatively efficient protein-protein docking tools. In terms of computational efficiency and performance, ZDOCK was more suitable as preferred computational tool, with average running time of $6.144$ minutes, average Fnat of best pose of $0.953$ and average rank of best pose of $4.158$. Meanwhile, ZDOCK still yielded better performance on Benchmark 5.0, which proved ZDOCK was effective in performing docking on large-scale dataset. Our survey can offer insights into the research on the treatment of bacterial infections by utilizing the appropriate docking methods. Qiaozhen Meng, Jijun Tang, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2021 | A comprehensive overview and critical evaluation of gene regulatory network inference technologiesabstractGene regulatory network (GRN) is the important mechanism of maintaining life process, controlling biochemical reaction and regulating compound level, which plays an important role in various organisms and systems. Reconstructing GRN can help us to understand the molecular mechanism of organisms and to reveal the essential rules of a large number of biological processes and reactions in organisms. Various outstanding network reconstruction algorithms use specific assumptions that affect prediction accuracy, in order to deal with the uncertainty of processing. In order to study why a certain method is more suitable for specific research problem or experimental data, we conduct research from model-based, information-based and machine learning-based method classifications. There are obviously different types of computational tools that can be generated to distinguish GRNs. Furthermore, we discuss several classical, representative and latest methods in each category to analyze core ideas, general steps, characteristics, etc. We compare the performance of state-of-the-art GRN reconstruction technologies on simulated networks and real networks under different scaling conditions. Through standardized performance metrics and common benchmarks, we quantitatively evaluate the stability of various methods and the sensitivity of the same algorithm applying to different scaling networks. The aim of this study is to explore the most appropriate method for a specific GRN, which helps biologists and medical scientists in discovering potential drug targets and identifying cancer biomarkers. Wenying He, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 5 |
| 2021 | A sequence-based multiple kernel model for identifying DNA-binding proteinsabstractBACKGROUND: DNA-Binding Proteins (DBP) plays a pivotal role in biological system. A mounting number of researchers are studying the mechanism and detection methods. To detect DBP, the tradition experimental method is time-consuming and resource-consuming. In recent years, Machine Learning methods have been used to detect DBP. However, it is difficult to adequately describe the information of proteins in predicting DNA-binding proteins. In this study, we extract six features from protein sequence and use Multiple Kernel Learning-based on Centered Kernel Alignment to integrate these features. The integrated feature is fed into Support Vector Machine to build predictive model and detect new DBP. RESULTS: In our work, date sets of PDB1075 and PDB186 are employed to test our method. From the results, our model obtains better results (accuracy) than other existing methods on PDB1075 ([Formula: see text]) and PDB186 ([Formula: see text]), respectively. CONCLUSION: Multiple kernel learning could fuse the complementary information between different features. Compared with existing methods, our method achieves comparable and best results on benchmark data sets. Yuqing Qian, Limin Jiang, Yijie Ding, Jijun Tang, Fei Guo 0001 |
BMC Bioinform. | 5 |
| 2021 | Identification of drug-target interactions via multi-view graph regularized link propagation model
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Neurocomputing | 3 |
| 2021 | Granular multiple kernel learning for identifying RNA-binding protein residues via integrating sequence and structure information
Yijie Ding, Qiaozhen Meng, Jijun Tang, Fei Guo 0001 |
Neural Comput. Appl. | 5 |
| 2021 | Protein Crystallization Identification via Fuzzy Model on Linear Neighborhood RepresentationabstractX-ray crystallography is the most popular approach for analyzing protein 3D structure. However, the success rate of protein crystallization is very low (2-10 percent). To reduce the cost of time and resources, lots of computation-based methods are developed to detect the protein crystallization. Improving the accuracy of predicting protein crystallization is very important for the determination of protein structure by X-ray crystallography. At present, many machine learning methods are used to predict protein crystallization. In this article, we propose a Fuzzy Support Vector Machine based on Linear Neighborhood Representation (FSVM-LNR) to predict the crystallization propensity of proteins. Proteins are represented by three types of features (PsePSSM, PSSM-DWT, MMI-PS), and these features are serially combined and fed into FSVM-LNR. FSVM-LNR can filter outliers by membership score, which is calculated via reconstruction residuals of k nearest samples. To evaluate the performance of our predictive model, we test FSVM-LNR on the datasets of TRAIN3587, TEST3585 and TEST500. Our method achieves better Mathew's correlation coefficient (MCC) on TRAIN3587 (MCC: 0.56) and TEST3585 (MCC: 0.58). Although the performance of independent test is not the best on TEST500, FSVM-LNR also has a certain predictability (MCC: 0.70) in the identification of protein crystallization. The good performance on the datasets proves the effectiveness of our method and the better performance on large datasets further demonstrates the stability and superiority of our method. Yijie Ding, Jijun Tang, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | iEnhancer-KL: A Novel Two-Layer Predictor for Identifying Enhancers by Position Specific of Nucleotide CompositionabstractAn enhancer is a short region of DNA with the ability to recruit transcription factors and their complexes, increasing the likelihood of the transcription of a particular gene. Considering the importance of enhancers, enhancer identification is a prevailing problem in computational biology. In this paper, we propose a novel two-layer enhancer predictor called iEnhancer-KL, using computational biology algorithms to identify enhancers and then classify these enhancers into strong or weak types. Kullback-Leibler (KL) divergence is creatively taken into consideration to improve the feature extraction method PSTNPss. Then, LASSO is used to reduce the dimension of features and finally helps to get better prediction performance. Furthermore, the selected features are tested on several machine learning models, and the SVM algorithm achieves the best performance. The rigorous cross-validation indicates that our predictor is remarkably superior to the existing state-of-the-art methods with an Acc of 84.23 percent and the MCC of 0.6849 for identifying enhancers. Our code and results can be freely downloaded from https://github.com/Not-so-middle/iEnhancer-KL.git. Yinuo Lyu, Jiawei Li 0018, Wenying He, Yijie Ding, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2021 | CrystalM: A Multi-View Fusion Approach for Protein Crystallization PredictionabstractImproving the accuracy of predicting protein crystallization is very important for protein crystallization projects, which is a critical step for the determination of protein structure by X-ray crystallography. At present, many machine learning methods are used to predict protein crystallization. Here, we use a novel feature combination to construct a SVM model in the prediction of protein crystallization, called as CrystalM. In this work, we extract six features to represent protein sequences, namely Average Block-Position specific scoring matrix (AVBlock-PSSM), Average Block-Secondary Structure (AVBlock-SS), Global Encoding (GE), Pseudo-Position specific scoring matrix (PsePSSM), Protscale, and Discrete Wavelet Transform-Position specific scoring matrix (DWT-PSSM). Moreover, we employ two training datasets (TRAIN3587 and TRAIN1500) and their corresponding independent test datasets (TEST3585 and TEST500) to evaluate CrystalM by feeding multi-view features into Support Vector Machine (SVM) classifier. Two training datasets are employed for five-fold cross validation, and two test datasets are separately used to test the corresponding datasets. Finally, we compare CrystalM with other existing methods in the performance. For the datasets of TRAIN3587 and TEST3585, CrystalM achieves best Accuracy (ACC), best Specificity (SP), and the same Mathew's correlation coefficient (MCC) as the previous outperforming methods in the five-fold cross validation. In particular, ACC, SP, and MCC have surpassed the existing methods in independent test, which proves the effectiveness of CrystalM. Meanwhile, ACC, SP, and MCC are higher than existing methods in the five-fold cross validation for TRAIN1500. Although the performance of independent test for TEST500 is not the best, CrystalM also has a certain predictability in the prediction of protein crystallization. In addition, we find that only choosing the first four features can improve the performance of prediction for TRAIN1500 and TEST500, not only in independent tests but also in five-fold cross validation. This phenomenon indicates that the latter two features can not effectively represent proteins of TRAIN1500 and TEST500. CrystalM is a sequence-based protein crystallization prediction method. The good performance on the datasets proves the effectiveness of CrystalM and the better performance on large datasets further demonstrates the stability and superiority of CrystalM. Yijie Ding, Jijun Tang, Yu Dai 0005, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | AIEpred: An Ensemble Predictive Model of Classifier Chain to Identify Anti-Inflammatory PeptidesabstractAnti-inflammatory peptides (AIEs) have recently emerged as promising therapeutic agent for treatment of various inflammatory diseases, such as rheumatoid arthritis and Alzheimer's disease. Therefore, detecting the correlation between amino acid sequence and its anti-inflammatory property is of great importance for the discovery of new AIEs. To address this issue, we propose a novel prediction tool for accurate identification of peptides as anti-inflammatory epitopes or non anti-inflammatory epitopes. Most of all, we encode the original peptide sequence for better mining and exploring the information and patterns, based on the three feature representations as amino acid contact, position specific scoring matrix, physicochemical property. At the same time, we exploit several feature extraction models and utilize one feature selection model, in order to construct many base classifiers from various feature representations. More specifically, we develop an effective classification model, with which we can extract and learn a set of informative features from the ensemble classifier chain model with different group of base classifiers. Furthermore, in order to test the predictive power of our model, we conduct the comparative experiments on the leave-one-out cross-validation and the independent test. It shows that our novel predictor performs great accurate for identification of AIEs as well as existing outstanding prediction tools. Source codes are available at https://github.com/guofei-tju/Ensemble-classifier-chain-model. Lianrong Pu, Jijun Tang, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | CEPZ: A Novel Predictor for Identification of DNase I Hypersensitive SitesabstractDNase I hypersensitive sites (DHSs) have proven to be tightly associated with cis-regulatory elements, commonly indicating specific function on the chromatin structure. Thus, identifying DHSs plays a fundamental role in decoding gene regulatory behavior. While traditional experimental methods turn to be time-consuming and expensive, computational techniques promise to be practical to discovering and analyzing regulatory factors. In this study, we applied an efficient model that considered composition information and physicochemical properties and effectively selected features with a boosting algorithm. CEPZ, our predictor, greatly improved a Matthews correlation coefficient and accuracy of 0.7740 and 0.9113 respectively, more competitive than any predictor before. This result suggests that it may become a useful tool for DHSs research in the human and other complex genomes. Our research was anchored on the properties of dinucleotides and we identified several dinucleotides with significant differences in the distribution of DHS and non-DHS samples, which are likely to have a special meaning in the chromatin structure. The datasets, feature sets and the relevant algorithm are available at https://github.com/YanZheng-16/CEPZ_DHS/. Yijie Ding, Fei Guo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | iPro2L-PSTKNC: A Two-Layer Predictor for Discovering Various Types of Promoters by Position Specific of Nucleotide CompositionabstractPromoters are DNA regulatory elements located proximal to the transcription start site, which are in charge of the initiation of specific gene transcription. In Escherichia coli, promoters can be recognized by σ factors that have multiple families based on distinct function and structure, such as σ24, σ28, σ32, σ38, σ54and σ70. At present, biological methods are mainly used to identify these promoters. However, because it is time-consuming and material-consuming to do biological experiments, computational biology algorithm has emerged as a more effective way to predict the classification. In this study, we develop a novel two-layer seamless predictor called iPro2L-PSTKNC to identify the promoters of the E. coli genome, which based on the feature extraction model we newly proposed that is named as the position specific tendencies of k-mer nucleotide composition (PSTKNC). On the first layer, it is a binary classification predicting whether a sequence is promoter or not. And the second layer is a multiple classification identifying which type the identified promoter belongs to. The ensemble classification SVM performsbest comparing with other algorithms, which gets a promising accuracy and the Matthews correlation coefficient (MCC) at 90.05% and 80.13%. Our data and code are available at https://github.com/lyuyinuo/iPro2L-PSTKNC. Yinuo Lyu, Wenying He, Quan Zou 0001, Fei Guo 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Multi-Scale Time-Series Kernel-Based Learning Method for Brain Disease DiagnosisabstractThe functional magnetic resonance imaging (fMRI) is a noninvasive technique for studying brain activity, such as brain network analysis, neural disease automated diagnosis and so on. However, many existing methods have some drawbacks, such as limitations of graph theory, lack of global topology characteristic, local sensitivity of functional connectivity, and absence of temporal or context information. In addition to many numerical features, fMRI time series data also cover specific contextual knowledge and global fluctuation information. Here, we propose multi-scale time-series kernel-based learning model for brain disease diagnosis, based on Jensen-Shannon divergence. First, we calculate correlation value within and between brain regions over time. In addition, we extract multi-scale synergy expression probability distribution (interactional relation) between brain regions. Also, we produce state transition probability distribution (sequential relation) on single brain regions. Then, we build time-series kernel-based learning model based on Jensen-Shannon divergence to measure similarity of brain functional connectivity. Finally, we provide an efficient system to deal with brain network analysis and neural disease automated diagnosis. On Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset, our proposed method achieves accuracy of 0.8994 and AUC of 0.8623. On Major Depressive Disorder (MDD) dataset, our proposed method achieves accuracy of 0.9166 and AUC of 0.9263. Experiments show that our proposed method outperforms other existing excellent neural disease automated diagnosis approaches. It shows that our novel prediction method performs great accurate for identification of brain diseases as well as existing outstanding prediction tools. Jiaqi Ding, Junhai Xu, Jijun Tang, Fei Guo 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Melanoma Classification in Dermoscopy Images via Ensemble Learning on Deep Neural NetworkabstractAuotmatic melanoma classification in dermoscopy images is a very important task, which can help improve diagnostic accuracy and reduce mortality. Deep convolutional neural network (DCNN) has developed rapidly in recent years, but it is still a challenging task due to the intra-class variation and inter-class similarity of melanoma. We proposed a novel neural network integration model, which is composed of three parts: First, we use U-net segmentation network to generate masks and use the masks to crop original images; Second, we use five state-of-the-art DCNNs to extract features of cropped images, and add the squeeze-excitation block (SE block) to emphasize useful features; Finally, we construct a new neural network with local connection to integrate the classification results, extract features of different class of results, and integrate the results of each class separately. Local connection can integrate each class separately, maximizing the advantages of different networks in various classes. We evaluate our model on ISIC 2017 challenge dataset, and the result shows that our method has better performance compared with the existing methods. Jiawei Li 0018, Shiqiang Ma, Jijun Tang, Fei Guo 0001 |
BIBM | 5 |
| 2020 | An two-layer predictive model of ensemble classifier chain for detecting antimicrobial peptidesabstractAntimicrobial peptides (AMPs) are innate immune molecules that exhibit activities against a range of microbes. According to their special functions, AMPs are generally classified into several categories. Over the last decade, a number of AMP prediction tools have been designed and made freely available online, which show potential to discriminate AMPs from non-AMPs. However, the relative quality of existing AMP predictions produced by various tools is difficult to quantify. In fact, a comprehensive benchmark dataset used to train the prediction model is one of key points to solving the problem. Also, how to address the multi-label character of new synthetic instance is obviously very important to both basic research and drug development. In view of this, AMPs prediction should be a task of two-level multi-label classification, in which the first step is to identify whether a query peptide is AMP, and the second step is to identify which functional type(s) the peptide belongs to. To establish a really useful prediction method, we construct a valid benchmark dataset to train the predictor, and develop a powerful algorithm to operate the prediction. In this paper, we propose a novel two-layer prediction model for identifying AMP and its functional types, using ADASYN oversampling technology to solve imbalance multi-label classification problem. First, we construct a novel benchmark AMPs dataset with seven different AMP functional types. Then, we encode AMPs within three different feature representations, and use various feature extraction models to convert the variable length coding matrix into some equidimensional features. Furthermore, we use modelbased feature selection method for filtering effective and sparse features. Finally, we apply ensemble classifier chain model to identify whether a query peptide is an AMPs or non-AMPs. In the second layer prediction, we use ADASYN to oversample different functional types of AMPs, and build a multi-label multi-class prediction model to identify which functional type(s) it belongs to. To be specific, our novel method outperforms outstanding rather than other tools in most respects on our novel benchmark datasets. Our novel benchmark dataset and source codes are available at https://github.com/guofei-tju/Two_Level_Ensemble-classifier-chain. Yijie Ding, Jijun Tang, Fei Guo 0001 |
BIBM | 5 |
| 2020 | Critical evaluation of web-based prediction tools for human protein subcellular localizationabstractHuman protein subcellular localization has an important research value in biological processes, also in elucidating protein functions and identifying drug targets. Over the past decade, a number of protein subcellular localization prediction tools have been designed and made freely available online. The purpose of this paper is to summarize the progress of research on the subcellular localization of human proteins in recent years, including commonly used data sets proposed by the predecessors and the performance of all selected prediction tools against the same benchmark data set. We carry out a systematic evaluation of several publicly available subcellular localization prediction methods on various benchmark data sets. Among them, we find that mLASSO-Hum and pLoc-mHum provide a statistically significant improvement in performance, as measured by the value of accuracy, relative to the other methods. Meanwhile, we build a new data set using the latest version of Uniprot database and construct a new GO-based prediction method HumLoc-LBCI in this paper. Then, we test all selected prediction tools on the new data set. Finally, we discuss the possible development directions of human protein subcellular localization. Availability: The codes and data are available from http://www.lbci.cn/syn/. Yinan Shen, Yijie Ding, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
Briefings Bioinform. | 5 |
| 2020 | Identification of membrane protein types via multivariate information fusion with Hilbert-Schmidt Independence Criterion
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Neurocomputing | 4 |
| 2020 | Identification of Drug-Target Interactions via Dual Laplacian Regularized Least Squares with Multiple Kernel Fusion
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Knowl. Based Syst. | 3 |
| 2020 | Identification of drug-target interactions via fuzzy bipartite local model
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Neural Comput. Appl. | 3 |
| 2020 | DeepAVP: A Dual-Channel Deep Neural Network for Identifying Variable-Length Antiviral PeptidesabstractAntiviral peptides (AVPs) have been experimentally verified to block virus into host cells, which have antiviral activity with decapeptide amide. Therefore, utilization of experimentally validated antiviral peptides is a potential alternative strategy for targeting medically important viruses. In this article, we propose a dual-channel deep neural network ensemble method for analyzing variable-length antiviral peptides. The LSTM channel can capture long-term dependencies for effectively studying original variable-length sequence data. The CONV channel can build dynamic neural network for analyzing the local evolution information. Also, our model can fine-tune the substitution matrix for specifically functional peptides. Applying it to a novel experimentally verified dataset, our AVPs predictor, DeepAVP, demonstrates state-of-the-art performance of [Formula: see text] accuracy and 0.85 MCC, which is far better than existing prediction methods for identifying antiviral peptides. Therefore, DeepAVP, web server for predicting the effective AVPs, would make significantly contributions to peptide-based antiviral research. Jiawei Li 0018, Yuqian Pu, Jijun Tang, Quan Zou 0001, Fei Guo 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2019 | Identification of DNA-Binding Proteins via Fuzzy Multiple Kernel Model and Sequence Information
Yijie Ding, Jijun Tang, Fei Guo 0001 |
ICIC (2) | 3 |
| 2019 | Identifying protein-protein interface via a novel multi-scale local sequence and structural representationabstractAbstract Background Protein-protein interaction plays a key role in a multitude of biological processes, such as signal transduction, de novo drug design, immune responses, and enzymatic activities. Gaining insights of various binding abilities can deepen our understanding of the interaction. It is of great interest to understand how proteins in a complex interact with each other. Many efficient methods have been developed for identifying protein-protein interface. Results In this paper, we obtain the local information on protein-protein interface, through multi-scale local average block and hexagon structure construction. Given a pair of proteins, we use a trained support vector regression (SVR) model to select best configurations. On Benchmark v4.0, our method achieves average Irmsd value of 3.28Å and overall Fnat value of 63%, which improves upon Irmsd of 3.89Å and Fnat of 49% for ZRANK, and Irmsd of 3.99Å and Fnat of 46% for ClusPro. On CAPRI targets, our method achieves average Irmsd value of 3.45Å and overall Fnat value of 46%, which improves upon Irmsd of 4.18Å and Fnat of 40% for ZRANK, and Irmsd of 5.12Å and Fnat of 32% for ClusPro. The success rates by our method, FRODOCK 2.0, InterEvDock and SnapDock on Benchmark v4.0 are 41.5%, 29.0%, 29.4% and 37.0%, respectively. Conclusion Experiments show that our method performs better than some state-of-the-art methods, based on the prediction quality improved in terms of CAPRI evaluation criteria. All these results demonstrate that our method is a valuable technological tool for identifying protein-protein interface. Fei Guo 0001, Quan Zou 0001, Jijun Tang, Junhai Xu |
BMC Bioinform. | 1 |
| 2019 | Identification of drug-side effect association via multiple information integration with centered kernel alignment
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Neurocomputing | 3 |
| 2019 | Identification of Drug-Side Effect Association via Semisupervised Model and Multiple Kernel LearningabstractDrug-side effect association contains the information on marketed medicines and their recorded adverse drug reactions. Traditional experimental method is time consuming and expensive. All associations of drugs and side-effects are seen as a bipartite network. Therefore, many computational approaches have been developed to deal with this problem, which are used to predict new potential associations. However, lots of methods did not consider multiple kernel learning (MKL) algorithm, which can integrate multiple sources of information and further improve prediction performance. In this study, we develop a novel predictor of drug-side effect association. First, we build multiple kernels from drug space and side-effect space. What is more, these corresponding kernels are linear weighted by MKL algorithm in drug space and side-effect space, respectively. Finally, a graph-based semisupervised learning is employed to construct drug-side effect predictor. Compared with existing methods, our method achieves better results on three benchmark data sets. The values of area under the precision recall curve are 0.668, 0.673, and 0.670 on three benchmark data sets, respectively. Our method is a useful tool for the side-effects prediction of drugs. Yijie Ding, Jijun Tang, Fei Guo 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2018 | Progressive approach for SNP calling and haplotype assembly using single molecular sequencing dataabstractMotivation: Haplotype information is essential to the complete description and interpretation of genomes, genetic diversity and genetic ancestry. The new technologies can provide Single Molecular Sequencing (SMS) data that cover about 90% of positions over chromosomes. However, the SMS data has a higher error rate comparing to 1% error rate for short reads. Thus, it becomes very difficult for SNP calling and haplotype assembly using SMS reads. Most existing technologies do not work properly for the SMS data. Results: In this paper, we develop a progressive approach for SNP calling and haplotype assembly that works very well for the SMS data. Our method can handle more than 200 million non-N bases on Chromosome 1 with millions of reads, more than 100 blocks, each of which contains more than 2 million bases and more than 3K SNP sites on average. Experiment results show that the false discovery rate and false negative rate for our method are 15.7 and 11.0% on NA12878, and 16.5 and 11.0% on NA24385. Moreover, the overall switch errors for our method are 7.26 and 5.21 with average 3378 and 5736 SNP sites per block on NA12878 and NA24385, respectively. Here, we demonstrate that SMS reads alone can generate a high quality solution for both SNP calling and haplotype assembly. Availability and implementation: Source codes and results are available at https://github.com/guofeieileen/SMRT/wiki/Software. Fei Guo 0001, Lusheng Wang 0001 |
Bioinform. | 1 |
| 2018 | Approximation algorithms for the scaffolding problem and its generalizations
Zhi-Zhong Chen, Youta Harada, Fei Guo 0001, Lusheng Wang 0001 |
Theor. Comput. Sci. | 3 |
| 2017 | Improved prediction of protein-protein interactions using novel negative samples, features, and an ensemble classifier
Leyi Wei, Pengwei Xing, Jian-Cang Zeng, Jin-Xiu Chen, Ran Su, Fei Guo 0001 |
Artif. Intell. Medicine | 6 |
| 2017 | Identification of drug-target interactions via multiple information integration
Yijie Ding, Jijun Tang, Fei Guo 0001 |
Inf. Sci. | 3 |
| 2016 | Predicting protein-protein interactions via multivariate mutual information of protein sequencesabstractBACKGROUND: Protein-protein interactions (PPIs) are central to a lot of biological processes. Many algorithms and methods have been developed to predict PPIs and protein interaction networks. However, the application of most existing methods is limited since they are difficult to compute and rely on a large number of homologous proteins and interaction marks of protein partners. In this paper, we propose a novel sequence-based approach with multivariate mutual information (MMI) of protein feature representation, for predicting PPIs via Random Forest (RF). METHODS: Our method constructs a 638-dimentional vector to represent each pair of proteins. First, we cluster twenty standard amino acids into seven function groups and transform protein sequences into encoding sequences. Then, we use a novel multivariate mutual information feature representation scheme, combined with normalized Moreau-Broto Autocorrelation, to extract features from protein sequence information. Finally, we feed the feature vectors into a Random Forest model to distinguish interaction pairs from non-interaction pairs. RESULTS: To evaluate the performance of our new method, we conduct several comprehensive tests for predicting PPIs. Experiments show that our method achieves better results than other outstanding methods for sequence-based PPIs prediction. Our method is applied to the S.cerevisiae PPIs dataset, and achieves 95.01 % accuracy and 92.67 % sensitivity repectively. For the H.pylori PPIs dataset, our method achieves 87.59 % accuracy and 86.81 % sensitivity respectively. In addition, we test our method on other three important PPIs networks: the one-core network, the multiple-core network, and the crossover network. CONCLUSIONS: Compared to the Conjoint Triad method, accuracies of our method are increased by 6.25,2.06 and 18.75 %, respectively. Our proposed method is a useful tool for future proteomics studies. Yijie Ding, Jijun Tang, Fei Guo 0001 |
BMC Bioinform. | 3 |
| 2016 | Learning from real imbalanced data of 14-3-3 proteins binding specificity
Jijun Tang, Fei Guo 0001 |
Neurocomputing | 3 |
| 2015 | A novel multivariate performance optimization method based on sparse coding and hyper-predictor learning
Zhiyong Ding, Fei Guo 0001, Huogen Wang |
Neural Networks | 3 |
| 2013 | Detecting Protein Conformational Changes in Interactions via Scaling Known Structures
Fei Guo 0001, Shuaicheng Li 0001, Wenji Ma, Lusheng Wang 0001 |
RECOMB | 1 |
| 2012 | P-Binder: A System for the Protein-Protein Binding Sites Identification
Fei Guo 0001, Shuaicheng Li 0001, Lusheng Wang 0001 |
ISBRA | 1 |
| 2012 | Protein-protein binding site identification by enumerating the configurationsabstractBACKGROUND: The ability to predict protein-protein binding sites has a wide range of applications, including signal transduction studies, de novo drug design, structure identification and comparison of functional sites. The interface in a complex involves two structurally matched protein subunits, and the binding sites can be predicted by identifying structural matches at protein surfaces. RESULTS: We propose a method which enumerates "all" the configurations (or poses) between two proteins (3D coordinates of the two subunits in a complex) and evaluates each configuration by the interaction between its components using the Atomic Contact Energy function. The enumeration is achieved efficiently by exploring a set of rigid transformations. Our approach incorporates a surface identification technique and a method for avoiding clashes of two subunits when computing rigid transformations. When the optimal transformations according to the Atomic Contact Energy function are identified, the corresponding binding sites are given as predictions. Our results show that this approach consistently performs better than other methods in binding site identification. CONCLUSIONS: Our method achieved a success rate higher than other methods, with the prediction quality improved in terms of both accuracy and coverage. Moreover, our method is being able to predict the configurations of two binding proteins, where most of other methods predict only the binding sites. The software package is available at http://sites.google.com/site/guofeics/dobi for non-commercial use. Fei Guo 0001, Shuaicheng Li 0001, Lusheng Wang 0001, Daming Zhu |
BMC Bioinform. | 1 |
| 2012 | Computing the protein binding sitesabstractBACKGROUND: Identifying the location of binding sites on proteins is of fundamental importance for a wide range of applications including molecular docking, de novo drug design, structure identification and comparison of functional sites. Structural genomic projects are beginning to produce protein structures with unknown functions. Therefore, efficient methods are required if all these structures are to be properly annotated. Lots of methods for finding binding sites involve 3D structure comparison. Here we design a method to find protein binding sites by direct comparison of protein 3D structures. RESULTS: We have developed an efficient heuristic approach for finding similar binding sites from the surface of given proteins. Our approach consists of three steps: local sequence alignment, protein surface detection, and 3D structures comparison. We implement the algorithm and produce a software package that works well in practice. When comparing a complete protein with all complete protein structures in the PDB database, experiments show that the average recall value of our approach is 82% and the average precision value of our approach is also significantly better than the existing approaches. CONCLUSIONS: Our program has much higher recall values than those existing programs. Experiments show that all the existing approaches have recall values less than 50%. This implies that more than 50% of real binding sites cannot be reported by those existing approaches. The software package is available at http://sites.google.com/site/guofeics/bsfinder. Fei Guo 0001, Lusheng Wang 0001 |
BMC Bioinform. | 1 |
| 2011 | Computing the Protein Binding Sites
Fei Guo 0001, Lusheng Wang 0001 |
ISBRA | 1 |