EDBT 2026 Demo / reviewers in the wild / expert
Ran Su
dblp:115/7315
· DBLP profile ↗
64ranked-venue papers
18as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 43 · 13 first-author · 29 since 2021Artificial intelligence and machine learning · 18 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Multi-Modal Contrastive Learning Framework for Cyclic Peptide Permeability PredictionabstractCyclic peptides represent a rapidly growing class of therapeutics, yet their development is often hindered by the challenge of predicting cell membrane permeability, a critical determinant of drug efficacy. Existing computational methods often struggle to integrate the diverse structural information inherent in these complex molecules, resulting in suboptimal predictive accuracy. Here, we introduce MCPerm, a multi-modal deep learning framework that synergistically integrates 1D SMILES, 2D topological, and 3D geometric information through a novel modality share and contrastive learning strategy to accurately predict cyclic peptide permeability. MCPerm fine-tunes a pretrained peptide language model for SMILES encoding and uses a parameter-sharing graph transformer for structural representation, while a dual contrastive learning mechanism enforces representational consistency both within and between modalities. On the benchmark PAMPA dataset, MCPerm achieves state-of-the-art performance, significantly outperforming leading methods. We further demonstrate its robustness and competitive transferability across three independent assays (Caco-2, MDCK, and RRCK). Our work presents a robust in silico framework that holds potential to accelerate the rational design and discovery of cell-permeable cyclic peptide drugs. Furthermore, to move beyond predictive accuracy, we introduced an attention-based visualization analysis. The results demonstrate that our model is not a "black box"; it has learned key chemical principles governing cyclic peptide permeability. Shuwen Xiong, Feifei Cui, Rao Zeng, Ran Su, Leyi Wei |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2026 | KG-CMI: Knowledge Graph Enhanced Cross-Mamba Interaction for Medical Visual Question Answering
Xianyao Zheng, Hui Cui 0002, Changming Sun, Xiangyu Li 0004, Ran Su, Leyi Wei, Qiangguo Jin |
IEEE Trans. Ind. Informatics | 6 |
| 2026 | PGST: A Prototype-Guided Parameter-Efficient Network for Spatial Transcriptomics PredictionabstractSpatial transcriptomics (ST) aims to decode spatially resolved gene expression patterns while preserving tissue morphology. Current methods tend to use lower-cost deep learning approaches for gene expression prediction, yet face severe challenges. First, existing methods fail to give sufficient consideration to the spatial specificity of positional encoding inherent in ST; second, they neglect to leverage spatially coherent co-expression patterns across different domains; third, their reliance on linearly weighted aggregation induces vulnerability to noise and distribution shifts; and finally, these architectures exhibit limited parameter efficiency. To address these issues, we introduce prototype-guided network for spatial transcriptomics (PGST), which includes four parts: (1) oriented signal propagation through polar embedding strategy for spatial transcriptomics (PEST); (2) prototype-guided aggregation for global co-feature preservation; (3) global consistency enforcement via shared decoder with reconstruction loss; and (4) lightweight architectural design. Our framework integrates contrastive learning with graph neural networks to balance local-global spatial dependencies and cross-modal consistency. Experimental results on multiple datasets from ST demonstrate the superior performance of our PGST model than existing methods. Yuan He 0016, Kaimiao Hu, Changming Sun, Leyi Wei, Ran Su |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | ADSA-Net: Addressing Intra- and Inter-Class Variabilities for Severity Assessment of Atopic DermatitisabstractAtopic dermatitis (AD) is a chronic inflammatory skin disorder characterized by recurrent itching, erythema, dryness, and eczematous lesions. Automated AD severity assessment is crucial for cost-effective and precision clinical decision-making but remains challenging. This is due to the subtle contrast variations between key dermatological signs and significant variations in lesion sizes across patients and disease stages. To address these issues, we propose ADSA-Net, which is designed to handle both intra- and inter-class variabilities. ADSA-Net first extracts multi-scale texture-aware features to effectively model variations in lesion size and texture. It then leverages contrastive learning to enhance intra- and inter-class differentiation, strengthening model's discriminatory ability for samples that are difficult to distinguish. Finally, ADSA-Net refines the learning process by leveraging a dynamic feature pool of correctly classified samples to guide the calibration of misclassified instances, enhancing overall accuracy. We further establish a dataset for AD severity assessment. Comprehensive experiments on this dataset show that ADSA-Net significantly outperforms existing state-of-the-art methods. Qiangguo Jin, Xurong Chen, Hui Cui 0002, Changming Sun, Youpeng Deng, Cong Cong 0001, Yuqi Fang, Ran Su, Leyi Wei |
BIBM | 8 |
| 2025 | M3ST-DTI: A Multi-Task Learning Model for Drug-Target Interactions Based on Multi-Modal Features and Multi-Stage AlignmentabstractAccurate prediction of drug-target interactions (DTI) is pivotal in drug discovery. However, existing approaches often fail to capture deep intra-modal feature in-teractions or achieve effective cross-modal alignment, limiting predictive performance and generalization. To address these challenges, we propose M3ST-DTI, a multi-task learning model that enables multi-stage integration and alignment of multi-modal features for DTI prediction. M3ST-DTI incorporates three types of features-textual, structural, and functional and enhances intra-modal representations using self-attention mechanisms and a hybrid pooling graph attention module. For early-stage feature alignment and fusion, the model integrates MCA with Gram loss as a structural constraint. In the later stage, a BCA module captures fine-grained interactions between drugs and targets within each modality, while a deep orthogonal fusion module mitigates feature redundancy. Extensive evaluations on benchmark datasets demonstrate that M3ST-DTI consistently outperforms state-of-the-art methods across diverse metrics. Ran Su, Liangliang Liu 0001 |
BIBM | 2 |
| 2025 | Iterative clustering algorithm G-DESC-E and pan-cancer key gene analysis based on single-cell sequencing dataabstractSingle-cell sequencing technology has profoundly revolutionized the field of cancer genomics, enabling researchers to explore gene expression profiles at the resolution of individual cells. Despite its extensive applications in the study of cancer gene states, pan-cancer analyses remain relatively underexplored. In this study, we propose the G-DESC-E algorithm, which effectively distinguishes dimensionality-reduced data through a grid-based approach, filters out outliers during the preprocessing phase, and employs the Louvain algorithm for prescreening cluster centroids as initial clusters. We construct an objective function by integrating label entropy with the Kullback-Leibler divergence formula, achieving final clustering results through iterative optimization. Our findings demonstrate the effectiveness of the G-DESC-E algorithm in enhancing clustering accuracy. By applying our methodology to real-world datasets, we illustrate its capability to identify critical transcriptional features associated with distinct cancer subtypes. Coupled with clustering visualization and gene ontology analysis, we identify over thirty genes potentially related to cancer occurrence and progression. The algorithm and research framework presented in this study pave the way for new directions in clinical research by applying single-cell sequencing technology to the analysis of key genes within the realm of pan-cancer analysis for the first time. This approach offers valuable insights that can inform further clinical investigations. Ke Wu 0023, Changming Sun, Jie Geng 0001, Leyi Wei, Ran Su |
Briefings Bioinform. | 7 |
| 2025 | Synergizing multimodal data and fingerprint space exploration for mechanism of action predictionabstractMOTIVATION: Effective computational methods for predicting the mechanism of action (MoA) of compounds are essential in drug discovery. Current MoA prediction models mainly utilize the structural information of compounds. However, high-throughput screening technologies have generated more targeted cell perturbation data for MoA prediction, a factor frequently disregarded by the majority of current approaches. Moreover, exploring the commonalities and specificities among different fingerprint representations remains challenging. RESULTS: In this paper, we propose IFMoAP, a model integrating cell perturbation image and fingerprint data for MoA prediction. Firstly, we modify the Res-Net to accommodate the feature extraction of five-channel cell perturbation images and establish a granularity-level attention mechanism to combine coarse- and fine-grained features. To learn both common and specific fingerprint features, we introduce an FP-CS module, projecting four fingerprint embeddings into distinct spaces and incorporating two loss functions for effective learning. Finally, we construct two independent classifiers based on image and fingerprint features for prediction and for weighting the two prediction scores. Experimental results demonstrate that our model achieves highest accuracy of 0.941 when using multimodal data. The comparison with other methods and explorations further highlights the superiority of our proposed model and the complementary characteristics of multimodal data. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/ s1mplehu/IFMoAP. The raw image data of Cell Painting can be accessed from Figshare (https://doi.org/10.17044/scilifelab.21378906). Kaimiao Hu, Jianguo Wei, Changming Sun, Jie Geng 0001, Leyi Wei, Ran Su |
Bioinform. | 7 |
| 2025 | PKDF-Net: Anticancer peptide prediction via a prior-knowledge-aware dual-path feature-entangled network
Qiangguo Jin, Ankang Wu, Leyi Wei, Hui Cui 0002, Ping Xuan, Xikang Feng, Ran Su |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | Iterative pseudo-labeling based adaptive copy-paste supervision for semi-supervised tumor segmentation
Qiangguo Jin, Hui Cui 0002, Changming Sun, Yimiao He, Ping Xuan, Cong Cong 0001, Leyi Wei, Ran Su |
Knowl. Based Syst. | 10 |
| 2025 | Drug-induced liver injury prediction based on graph convolutional networks and toxicogenomicsabstractDrug-induced liver injury is a leading cause of high attrition rates for both candidate drugs and marketed medications. Previous in silico models may not effectively utilize biological drug property information and often lack robust model validation. In this study, we developed a graph convolutional network embedded with a biological graph learning (BioGL) module-named BioGL-GCN(Biological Graph Learning-Graph Convolutional Network)-for drug-induced liver injury prediction using toxicogenomic profiles. The BioGL module learned the optimal graph representations of gene interactions by utilizing the constructed protein-protein interaction network, which represents initial gene relationships, and gene frequency information obtained from gene enrichment analysis. Finally, the graph convolutional network was used to identify drug hepatotoxicity. Our method pays more attention to gene-gene relationships compared to previous approaches, thereby achieving more accurate predictive performance. We applied BioGL-GCN to predict DILI risk for active components in the integrated traditional Chinese medicine (ITCM) database and validated these predictions through hepatotoxicity experiments using a 3D primary human hepatocyte (PHH) model. The results showed that our model achieved a prediction accuracy of 79%, thus further validating the reliability of the constructed model. Tong Xiao 0018, Kaimiao Hu, Kaimin Guo, Weihua Lei, Shuiping Zhou, Yunhui Hu, Ran Su |
PLoS Comput. Biol. | 11 |
| 2025 | TPNET: A Time-Sensitive Small Sample Multimodal Network for Cardiotoxicity Risk PredictionabstractCancer therapy-related cardiac dysfunction (CTRCD) is a potential complication associated with cancer treatment, particularly in patients with breast cancer, requiring monitoring of cardiac health during the treatment process. Tissue Doppler imaging (TDI) is a remarkable technique that can provide a comprehensive reflection of the left ventricle's physiological status. We hypothesized that the combination of TDI features with deep learning techniques could be utilized to predict CTRCD. To evaluate the hypothesis, we developed a temporal-multimodal pattern network for efficient training (TPNET) model to predict the incidence of CTRCD over a 24-month period based on TDI, function, and clinical data from 270 patients. Our model achieved an area under curve (AUC) of 0.83 and sensitivity of 0.88, demonstrating greater robustness compared to other existing visual models. To further translate our model's findings into practical applications, we utilized the integrated gradients (IG) attribution to perform a detailed evaluation of all the features. This analysis has identified key pathogenic signs that may have remained unnoticed, providing a viable option for implementing our model in preoperative breast cancer patients. Additionally, our findings demonstrate the potential of TPNET in discovering new causative agents for CTRCD. Yuan He 0016, Fengyun Zhang, Kaimiao Hu, Changming Sun, Jie Geng 0001, Ning Ren, Ran Su |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | BFGTP: A BERT-Guided Two-Stage Molecular Representation Learning Framework for Toxicity PredictionabstractAccurate prediction of molecular toxicity is vital for drug development. Most mainstream methods rely on fingerprints or graph-based feature extraction, the emergence of large language models (LLMs) offers new prospects for molecular representation learning in toxicity prediction. Although several studies attempt to leverage LLMs to integrate molecular sequence data for pretraining molecular representations, certain limitations remain. Current LLM-based approaches usually utilize solely on class embedding features, overlooking the rich information in sequence embedding. Moreover, integrating pre-trained molecular representations with multi-modal molecular data may further enhance performance in toxicity prediction. To address these challenges, we propose BFGTP, a BERT-guided two-stage molecular representation learning framework for toxicity prediction. Firstly, we design independent encoders for molecular descriptions of three modalities, where the fingerprint encoder with dual level attention mechanisms effectively integrates multi-category fingerprints. Then, the two-stage guide strategy is introduced to fully utilize the prior knowledge of LLMs, employing contrastive learning to align and fuse the tri-modal representations and knowledge distillation to align predicted value distributions. BFGTP ultimately combines fingerprint and graph representations to predict molecular toxicity. Experiments on seven toxicity datasets show that BFGTP outperforms baselines, achieving the highest AUC on five datasets and the best average performance across five evaluation metrics. Ablation studies, t-SNE visualization and case study confirm the effectiveness of BFGTP's components and its ability to capture meaningful molecular representations. Kaimiao Hu, Yuan He 0016, Jianguo Wei, Changming Sun, Jie Geng 0001, Leyi Wei, Ran Su |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Multi-Modal Deep Representation Learning Accurately Identifies and Interprets Drug-Target InteractionsabstractDeep learning offers efficient solutions for drug-target interaction prediction, but current methods often fail to capture the full complexity of multi-modal data (i.e., sequence, graphs, and three-dimensional structures), limiting both performance and generalization. Here, we present UnitedDTA, a novel explainable deep learning framework capable of integrating multi-modal biomolecule data to improve the binding affinity prediction, especially for novel (unseen) drugs and targets. UnitedDTA enables automatic learning unified discriminative representations from multi-modality data via contrastive learning and cross-attention mechanisms for cross-modality alignment and integration. Comparative results on multiple benchmark datasets show that UnitedDTA significantly outperforms the state-of-the-art drug-target affinity prediction methods and exhibits better generalization ability in predicting unseen drug-target pairs. More importantly, unlike most "black-box" deep learning methods, our well-established model offers better interpretability which enables us to directly infer the important substructures of the drug-target complexes that influence the binding activity, thus providing the insights in unveiling the binding preferences. Moreover, by extending UnitedDTA to other downstream tasks (e.g., molecular property prediction), we showcase the proposed multi-modal representation learning is capable of capturing the latent molecular representations that are closely associated with the molecular property, demonstrating the broad application potential for advancing the drug discovery process. Jiayue Hu, Xiangxiang Zeng, Quan Zou 0001, Ran Su, Leyi Wei |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Multiview Deep Learning-Based Molecule Design and Structural Optimization Accelerates Inhibitor DiscoverabstractIn this work, we propose MEDICO, a multiview deep generative model for molecule generation, structural optimization, and the SARS-CoV-2 inhibitor discovery. To the best of our knowledge, MEDICO is the first-of-this-kind graph generative model that can generate molecular graphs similar to the structure of targeted molecules, with a multiview representation learning framework to sufficiently and adaptively learn comprehensive structural semantics from targeted molecular topology and geometry. We show that our MEDICO significantly outperforms the state-of-the-art methods in generating valid, novel, and unique molecules under benchmarking comparisons, particularly achieving $\tilde {8}5 \%$ improvement compared with the state-of-the-art methods in terms of validity. Importantly, we showcase that the multiview deep learning model enables us to generate not only the molecules structurally similar to the targeted molecules but also the molecules with desired chemical properties. Moreover, case study results on targeted molecule generation for the SARS-CoV-2 main protease (Mpro) show that we successfully generate new small molecules with desired drug-like properties for the Mpro by integrating molecular docking into our model as a chemical priori, potentially accelerating the de novo design of COVID-19 drugs. Furthermore, we apply MEDICO to the structural optimization of three well-known Mpro inhibitors (N3, 11a, and GC376) and achieve $\tilde {8}8 \%$ improvement compared with the origin inhibitors in their binding affinity to Mpro, demonstrating the application value of our model for the development of therapeutics for SARS-CoV-2 infection. Ruheng Wang, Quan Zou 0001, Xiangxiang Zeng, Ran Su, Leyi Wei |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | TSEML: A task-specific embedding-based method for few-shot classification of cancer molecular subtypesabstractMolecular subtyping of cancer is recognized as a critical and challenging upstream task for personalized therapy. Existing deep learning methods have achieved significant performance in this domain when abundant data samples are available. However, the acquisition of densely labeled samples for cancer molecular subtypes remains a significant challenge for conventional data-intensive deep learning approaches. In this work, we focus on the few-shot molecular subtype prediction problem in heterogeneous and small cancer datasets, aiming to enhance precise diagnosis and personalized treatment. We first construct a new few-shot dataset for cancer molecular subtype classification and auxiliary cancer classification, named TCGA Few-Shot, from existing publicly available datasets. To effectively leverage the relevant knowledge from both tasks, we introduce a task-specific embedding-based meta-learning framework (TSEML). TSEML leverages the synergistic strengths of a model-agnostic meta-learning (MAML) approach and a prototypical network (ProtoNet) to capture diverse and fine-grained features. Comparative experiments conducted on the TCGA FewShot dataset demonstrate that our TSEML framework achieves superior performance in addressing the problem of few-shot molecular subtype classification. Ran Su, Hui Cui 0002, Ping Xuan, Chengyan Fang, Xikang Feng, Qiangguo Jin |
BIBM | 1 |
| 2024 | MSKI-Net: Towards modality-specific knowledge interaction for glioma survival predictionabstractGliomas hold a prominent position in neurooncology due to their high malignancy and poor survival rates. Accurately predicting the prognosis and survival risk of glioma patients is crucial for clinical treatment. Recent advances in survival prediction methods emphasize the importance of integrating complementary information from diverse modalities while neglecting the significant modality gap between pathological images and genomic data. To address this issue, we propose a modality-specific knowledge interaction network (MSKI-Net), which integrates whole slide images (WSI), RNA-Seq gene expression data, and copy number variation (CNV) data for glioma survival analysis. The MSKI-Net consists of a modality-specific feature enhancement (MSFE) module, a modality-interactive cross-attention (MICA) module, and a modality-specific knowledge-guided representation learning (MSKR) module. The three modules collaborate by complementing modality-specific features with modality-agnostic knowledge to improve the learning capability of MSKI-Net. Furthermore, we construct a dataset named TCGAmm, which combines WSI, RNA-Seq, and CNV data from The Cancer Genome Atlas (TCGA) to address the issue of data scarcity. Extensive experiments demonstrate that MSKI-Net achieves superior performance in predicting the survival risk of glioma cancer. Ran Su, Hui Cui 0002, Ping Xuan, Xikang Feng, Leyi Wei, Qiangguo Jin |
BIBM | 1 |
| 2024 | Location Embedding Based Pairwise Distance Learning for Fine-Grained Diagnosis of Urinary Stones
Qiangguo Jin, Jiapeng Huang, Changming Sun, Hui Cui 0002, Ping Xuan, Ran Su, Leyi Wei, Yu-Jie Wu, Chia-An Wu, Henry Been-Lirn Duh, Yueh-Hsun Lu |
MICCAI (11) | 6 |
| 2024 | StructuralDPPIV: a novel deep learning model based on atom structure for predicting dipeptidyl peptidase-IV inhibitory peptidesabstractMOTIVATION: Diabetes is a chronic metabolic disorder that has been a major cause of blindness, kidney failure, heart attacks, stroke, and lower limb amputation across the world. To alleviate the impact of diabetes, researchers have developed the next generation of anti-diabetic drugs, known as dipeptidyl peptidase IV inhibitory peptides (DPP-IV-IPs). However, the discovery of these promising drugs has been restricted due to the lack of effective peptide-mining tools. RESULTS: Here, we presented StructuralDPPIV, a deep learning model designed for DPP-IV-IP identification, which takes advantage of both molecular graph features in amino acid and sequence information. Experimental results on the independent test dataset and two wet experiment datasets show that our model outperforms the other state-of-art methods. Moreover, to better study what StructuralDPPIV learns, we used CAM technology and perturbation experiment to analyze our model, which yielded interpretable insights into the reasoning behind prediction results. AVAILABILITY AND IMPLEMENTATION: The project code is available at https://github.com/WeiLab-BioChem/Structural-DPP-IV. Junru Jin, Zhongshen Li, Mushuang Fan, Sirui Liang, Ran Su, Leyi Wei |
Bioinform. | 7 |
| 2024 | Inter- and intra-uncertainty based feature aggregation model for semi-supervised histopathology image segmentation
Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiangbin Zheng 0001, Leilei Cao, Leyi Wei, Ran Su |
Expert Syst. Appl. | 8 |
| 2024 | Self-Attention Factor Graph Neural Network for Multiagent Collaborative Target TrackingabstractCollaborative target tracking is an essential task in positioning systems, particularly in environments characterized by high dynamics, multi-source heterogeneous data, and interactive multi-agent scenarios. The challenge in such networks lies in the direct utilization of multi-source heterogeneous data as feature input for models. Additionally, the presence of high-dynamic time series data complicates the extraction of dependencies by the models. To address these issues, we introduce a novel approach that integrates a factor graph-based data fusion method with a graph neural network. This combination is designed to uncover potential dependencies between time series data and positional information within dynamic networks. Furthermore, we employ a self-attention mechanism, enabling distance-agnostic autonomous selection of complex network features. This innovation allows the model to achieve enhanced accuracy performance while simultaneously reducing computational costs. We validated our approach through simulation experiments. The results demonstrated the method’s effectiveness in fusing and selecting multi-source heterogeneous information within collaborative networks. It also excelled in identifying potential relationships between feature information and positional data, showcasing the robustness and applicability of our proposed solution in challenging collaborative target tracking environments. Cheng Xu 0003, Ran Su, Ran Wang 0014, Shihong Duan |
IEEE Internet Things J. | 2 |
| 2023 | Shape-aware contrastive deep supervision for esophageal tumor segmentation from CT scansabstractAccurate tumor segmentation is crucial for esophageal cancer radiotherapy treatment planning. The low contrast among the esophagus, tumors, and surrounding tissues, and irregular tumor shapes limit the performance of automatic segmentation methods. In this paper, we aim to exploit the irregular shapes of tumors to facilitate accurate segmentation. We propose a simple and pluggable shape-aware contrastive deep supervision network (SCDSNet) with shape-aware regularization and voxel-to-voxel contrastive deep supervision. Specifically, the shape-aware regularization with an uncertainty minimization strategy encourages the precise predictions of an additional shape-aware head. The voxel-to-voxel contrastive deep supervision enhances the multi-scale shape-tumor contrast for better voxel-to-voxel prediction of shapes. The proposed method is simple and highly pluggable, which can easily be extended to other frameworks. Further, we establish a large in-house dataset on esophageal cancer to validate the effectiveness of our proposed method. The quantitative and qualitative experimental results demonstrate the effectiveness of SCDSNet on the esophageal cancer dataset. Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiapeng Huang, Ping Xuan, Yiyue Xu, Leilei Cao, Leyi Wei, Ran Su |
BIBM | 10 |
| 2023 | OPE-SR: Orthogonal Position Encoding for Designing a Parameter-free Upsampling Module in Arbitrary-scale Image Super-ResolutionabstractArbitrary-scale image super-resolution (SR) is often tackled using the implicit neural representation (INR) approach, which relies on a position encoding scheme to im-prove its representation ability. In this paper, we introduce orthogonal position encoding (OPE), an extension of po-sition encoding, and an OPE-Upscale module to replace the INR-based upsampling module for arbitrary-scale im-age super-resolution. Our OPE-Upscale module takes 2D coordinates and latent code as inputs, just like INR, but does not require any training parameters. This parameter-free feature allows the OPE-Upscale module to directly perform linear combination operations, resulting in con-tinuous image reconstruction and achieving arbitrary-scale image reconstruction. As a concise SR framework, our method is computationally efficient and consumes less mem-ory than state-of-the-art methods, as confirmed by exten-sive experiments and evaluations. In addition, our method achieves comparable results with state-of-the-art methods in arbitrary-scale image super-resolution. Lastly, we show that OPE corresponds to a set of orthogonal basis, validating our design principle.11Project page: https://github.com/gaochao-s/ope-sr Gaochao Song, Qian Sun 0003, Luo Zhang 0002, Ran Su, Jianfeng Shi 0001, Ying He 0001 |
CVPR | 4 |
| 2023 | Multi-modality Contrastive Learning for Sarcopenia Screening from Hip X-rays and Clinical Information
Qiangguo Jin, Changjiang Zou, Hui Cui 0002, Changming Sun, Shu-Wei Huang, Yi-Jie Kuo, Ping Xuan, Leilei Cao, Ran Su, Leyi Wei, Henry Been-Lirn Duh, Yu-Pin Chen |
MICCAI (6) | 9 |
| 2023 | An effective negotiating agent framework based on deep offline reinforcement learningabstractLearning is crucial for automated negotiation, and recent years have witnessed a remarkable achievement in application of reinforcement learning (RL) for various negotiation tasks. Conventional RL methods focus generally on learning from active interactions with opposing negotiators. However, collecting online data is expensive in many realistic negotiation scenarios. While previous studies partially mitigate this problem through the use of opponent simulators (i.e., agents following known strategies), in reality it is usually hard to fully capture an opponent’s negotiation strategy. Moreover, a further challenge lies in an agent’s capability of adapting to dynamic variations of an opponent’s preferences or strategies, which may happen from time to time for different reasons in subsequent negotiations. In response to these challenges, this article proposes a novel Deep Offline Reinforcement learning Negotiating Agent framework that allows to learn an effective strategy using previously collected negotiation datasets without requiring interaction with an opponent. This is in contrast to existing RL-based negotiation approaches that all rely on active interaction with opponents. Furthermore, the strategy fine-tuning mechanism is included to adjust the learned strategy in response to the preferences or strategy changes of the opponent. The performance of the proposed framework is evaluated based on a diverse set of state-of-the-art baselines under different settings. Experimental results show that the framework allows to learn effective strategies exclusively with offline datasets, and is also capable of effectively adapting to changes of an opponent’s preferences or strategy. Siqi Chen 0001, Gerhard Weiss 0001, Ran Su, Kaiyou Lei |
UAI | 4 |
| 2022 | Semi-supervised Histological Image Segmentation via Hierarchical Consistency Enforcement
Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiangbin Zheng 0001, Leyi Wei, Zhenyu Fang, Zhaopeng Meng, Ran Su |
MICCAI (2) | 8 |
| 2022 | Accelerating bioactive peptide discovery via mutual information-based meta-learningabstractRecently, machine learning methods have been developed to identify various peptide bio-activities. However, due to the lack of experimentally validated peptides, machine learning methods cannot provide a sufficiently trained model, easily resulting in poor generalizability. Furthermore, there is no generic computational framework to predict the bioactivities of different peptides. Thus, a natural question is whether we can use limited samples to build an effective predictive model for different kinds of peptides. To address this question, we propose Mutual Information Maximization Meta-Learning (MIMML), a novel meta-learning-based predictive model for bioactive peptide discovery. Using few samples from various functional peptides, MIMML can sufficiently learn the discriminative information amongst various functions and characterize functional differences. Experimental results show excellent performance of MIMML though using far fewer training samples as compared to the state-of-the-art methods. We also decipher the latent relationships among different kinds of functions to understand what meta-model learned to improve a specific task. In summary, this study is a pioneering work in the field of functional peptide mining and provides the first-of-its-kind solution for few-sample learning problems in biological sequence analysis, accelerating the new functional peptide discovery. The source codes and datasets are available on https://github.com/TearsWaiting/MIMML. Junru Jin, Zhongshen Li, Jiaojiao Zhao, Balachandran Manavalan, Ran Su, Xin Gao 0001, Leyi Wei |
Briefings Bioinform. | 7 |
| 2022 | SRDFM: Siamese Response Deep Factorization Machine to improve anti-cancer drug recommendationabstractPredicting the response of cancer patients to a particular treatment is a major goal of modern oncology and an important step toward personalized treatment. In the practical clinics, the clinicians prefer to obtain the most-suited drugs for a particular patient instead of knowing the exact values of drug sensitivity. Instead of predicting the exact value of drug response, we proposed a deep learning-based method, named Siamese Response Deep Factorization Machines (SRDFM) Network, for personalized anti-cancer drug recommendation, which directly ranks the drugs and provides the most effective drugs. A Siamese network (SN), a type of deep learning network that is composed of identical subnetworks that share the same architecture, parameters and weights, was used to measure the relative position (RP) between drugs for each cell line. Through minimizing the difference between the real RP and the predicted RP, an optimal SN model was established to provide the rank for all the candidate drugs. Specifically, the subnetwork in each side of the SN consists of a feature generation level and a predictor construction level. On the feature generation level, both drug property and gene expression, were adopted to build a concatenated feature vector, which even enables the recommendation for newly designed drugs with only chemical property known. Particularly, we developed a response unit here to generate weighted genetic feature vector to simulate the biological interaction mechanism between a specific drug and the genes. For the predictor construction level, we built this level integrating a factorization machine (FM) component with a deep neural network component. The FM can well handle the discrete chemical information and both low-order and high-order feature interactions could be sufficiently learned. Impressively, the SRDFM works well on both single-drug recommendation and synergic drug combination. Experiment result on both single-drug and synergetic drug data sets have shown the efficiency of the SRDFM. The Python implementation for the proposed SRDFM is available at at https://github.com/RanSuLab/SRDFM Contact: [email protected], [email protected] and [email protected]. Ran Su, Guobao Xiao, Leyi Wei |
Briefings Bioinform. | 1 |
| 2022 | Distant metastasis identification based on optimized graph representation of gene interaction patternsabstractMetastasis is a major cause of cancer morbidity and mortality, and most cancer deaths are caused by cancer metastasis rather than by the primary tumor. The prediction of metastasis based on computational methods has not been explored much in the previous research. In this study, we proposed a graph convolutional network embedded with a graph learning (GL) module, named glmGCN, to predict the distant metastasis of cancer. Both the mRNA and lncRNA expressions were used to provide more genetic information than using the mRNA alone and we used them to construct gene interaction graph representation to consider the effect of genetic interaction. Then, the prediction of the cancer metastasis was performed under a GCN framework, which extracted informative and advanced features from the built non-regular graph structures. Particularly, a GL module was embedded in the proposed glmGCN to learn an optimal graph representation of the gene interaction. We firstly constructed the protein-protein interaction network to represent the initial gene(node) relationship graph. Then, through the GL module, a new graph representation was built which optimally learned the gene interaction strength. Finally, the GCN was adopted to identify the distant metastasis cases. It is worth mentioning that the proposed method pays more attentions on the gene-gene relation than the previous GCN-based method, so more accurate prediction performance can be obtained. The glmGCN was trained based on two types of cancer and was further validated using two other cancer types. A series of experiments have shown that the effectiveness of the proposed method. The implementation for the proposed method is available at https://github.com/RanSuLab/Metastasis-glmGCN. Ran Su, Quan Zou 0001, Leyi Wei |
Briefings Bioinform. | 1 |
| 2022 | EOCSA: Predicting prognosis of Epithelial ovarian cancer with whole slide histopathological images
Tianling Liu, Ran Su, Changming Sun, Xiu-Ting Li, Leyi Wei |
Expert Syst. Appl. | 2 |
| 2022 | A multi-label learning model for predicting drug-induced pathology in multi-organ based on toxicogenomics dataabstractDrug-induced toxicity damages the health and is one of the key factors causing drug withdrawal from the market. It is of great significance to identify drug-induced target-organ toxicity, especially the detailed pathological findings, which are crucial for toxicity assessment, in the early stage of drug development process. A large variety of studies have devoted to identify drug toxicity. However, most of them are limited to single organ or only binary toxicity. Here we proposed a novel multi-label learning model named Att-RethinkNet, for predicting drug-induced pathological findings targeted on liver and kidney based on toxicogenomics data. The Att-RethinkNet is equipped with a memory structure and can effectively use the label association information. Besides, attention mechanism is embedded to focus on the important features and obtain better feature presentation. Our Att-RethinkNet is applicable in multiple organs and takes account the compound type, dose, and administration time, so it is more comprehensive and generalized. And more importantly, it predicts multiple pathological findings at the same time, instead of predicting each pathology separately as the previous model did. To demonstrate the effectiveness of the proposed model, we compared the proposed method with a series of state-of-the-arts methods. Our model shows competitive performance and can predict potential hepatotoxicity and nephrotoxicity in a more accurate and reliable way. The implementation of the proposed method is available at https://github.com/RanSuLab/Drug-Toxicity-Prediction-MultiLabel. Ran Su, Haitang Yang, Leyi Wei, Siqi Chen 0001, Quan Zou 0001 |
PLoS Comput. Biol. | 1 |
| 2021 | PSSP-MVIRT: peptide secondary structure prediction based on a multi-view deep learning architectureabstractThe prediction of peptide secondary structures is fundamentally important to reveal the functional mechanisms of peptides with potential applications as therapeutic molecules. In this study, we propose a multi-view deep learning method named Peptide Secondary Structure Prediction based on Multi-View Information, Restriction and Transfer learning (PSSP-MVIRT) for peptide secondary structure prediction. To sufficiently exploit discriminative information, we introduce a multi-view fusion strategy to integrate different information from multiple perspectives, including sequential information, evolutionary information and hidden state information, respectively, and generate a unified feature space. Moreover, we construct a hybrid network architecture of Convolutional Neural Network and Bi-directional Gated Recurrent Unit to extract global and local features of peptides. Furthermore, we utilize transfer learning to effectively alleviate the lack of training samples (peptides with experimentally validated structures). Comparative results on independent tests demonstrate that our proposed method significantly outperforms state-of-the-art methods. In particular, our method exhibits better performance at the segment level, suggesting the strong ability of our model in capturing local discriminative information. The case study also shows that our PSSP-MVIRT achieves promising and robust performance in the prediction of new peptide secondary structures. Importantly, we establish a webserver to implement the proposed method, which is currently accessible via http://server.malab.cn/PSSP-MVIRT. We expect it can be a useful tool for the researchers of interest, facilitating the wide use of our method. Xiao Cao, Zitan Chen, Lesong Wei, Li-Zhen Cui 0001, Ran Su, Leyi Wei |
Briefings Bioinform. | 9 |
| 2021 | Iterative feature representation algorithm to improve the predictive performance of N7-methylguanosine sitesabstractMOTIVATION: N7-methylguanosine (m7G) is an important epigenetic modification, playing an essential role in gene expression regulation. Therefore, accurate identification of m7G modifications will facilitate revealing and in-depth understanding their potential functional mechanisms. Although high-throughput experimental methods are capable of precisely locating m7G sites, they are still cost ineffective. Therefore, it's necessary to develop new methods to identify m7G sites. RESULTS: In this work, by using the iterative feature representation algorithm, we developed a machine learning based method, namely m7G-IFL, to identify m7G sites. To demonstrate its superiority, m7G-IFL was evaluated and compared with existing predictors. The results demonstrate that our predictor outperforms existing predictors in terms of accuracy for identifying m7G sites. By analyzing and comparing the features used in the predictors, we found that the positive and negative samples in our feature space were more separated than in existing feature space. This result demonstrates that our features extracted more discriminative information via the iterative feature learning process, and thus contributed to the predictive performance improvement. Chichi Dai, Pengmian Feng, Li-Zhen Cui 0001, Ran Su, Wei Chen 0064, Leyi Wei |
Briefings Bioinform. | 4 |
| 2021 | Classification and gene selection of triple-negative breast cancer subtype embedding gene connectivity matrix in deep neural networkabstractTriple-negative breast cancer (TNBC) has been a challenging breast cancer subtype for oncological therapy. Normally, it can be classified into different molecular subtypes. Accurate and stable classification of the six subtypes is essential for personalized treatment of TNBC. In this study, we proposed a new framework to distinguish the six subtypes of TNBC, and this is one of the handful studies that completed the classification based on mRNA and long noncoding RNA expression data. Particularly, we developed a gene selection approach named DGGA, which takes correlation information between genes into account in the process of measuring gene importance and then effectively removes redundant genes. A gene scoring approach that combined GeneRank scores with gene importance generated by deep neural network (DNN), taking inter-subtype discrimination and inner-gene correlations into account, was came up to improve gene selection performance. More importantly, we embedded a gene connectivity matrix in the DNN for sparse learning, which takes additional consideration with weight changes during training when obtaining the measurement of the relative importance of each gene. Finally, Genetic Algorithm was used to simulate the natural evolutionary process to search for the optimal subset of TNBC subtype classification. We validated the proposed method through cross-validation, and the results demonstrate that it can use fewer genes to obtain more accurate classification results. The implementation for the proposed method is available at https://github.com/RanSuLab/TNBC. Ran Su, Leyi Wei |
Briefings Bioinform. | 2 |
| 2021 | Protein subcellular localization based on deep image features and criterion learning strategyabstractThe spatial distribution of proteome at subcellular levels provides clues for protein functions, thus is important to human biology and medicine. Imaging-based methods are one of the most important approaches for predicting protein subcellular location. Although deep neural networks have shown impressive performance in a number of imaging tasks, its application to protein subcellular localization has not been sufficiently explored. In this study, we developed a deep imaging-based approach to localize the proteins at subcellular levels. Based on deep image features extracted from convolutional neural networks (CNNs), both single-label and multi-label locations can be accurately predicted. Particularly, the multi-label prediction is quite a challenging task. Here we developed a criterion learning strategy to exploit the label-attribute relevancy and label-label relevancy. A criterion that was used to determine the final label set was automatically obtained during the learning procedure. We concluded an optimal CNN architecture that could give the best results. Besides, experiments show that compared with the hand-crafted features, the deep features present more accurate prediction with less features. The implementation for the proposed method is available at https://github.com/RanSuLab/ProteinSubcellularLocation. Ran Su, Linlin He, Tianling Liu, Xiaofeng Liu 0004, Leyi Wei |
Briefings Bioinform. | 1 |
| 2021 | Predicting drug-induced hepatotoxicity based on biological feature maps and diverse classification strategiesabstractIdentifying hepatotoxicity as early as possible is significant in drug development. In this study, we developed a drug-induced hepatotoxicity prediction model taking account of both the biological context and the computational efficacy based on toxicogenomics data. Specifically, we proposed a novel gene selection algorithm considering gene's participation, named BioCB, to choose the discriminative genes and make more efficient prediction. Then instead of using the raw gene expression levels to characterize each drug, we developed a two-dimensional biological process feature pattern map to represent each drug. Then we employed two strategies to handle the maps and identify the hepatotoxicity, the direct use of maps, named Two-dim branch, and vectorization of maps, named One-dim branch. The two strategies subsequently used the deep convolutional neural networks and LightGBM as predictors, respectively. Additionally, we here for the first time proposed a stacked vectorized gene matrix, which was more predictive than the raw gene matrix. Results validated on both in vivo and in vitro data from two public data sets, the TG-GATES and DrugMatrix, show that the proposed One-dim branch outperforms the deep framework, the Two-dim branch, and has achieved high accuracy and efficiency. The implementation of the proposed method is available at https://github.com/RanSuLab/Hepatotoxicity. Ran Su, Huichen Wu, Leyi Wei |
Briefings Bioinform. | 1 |
| 2021 | Computational prediction and interpretation of cell-specific replication origin sites from multiple eukaryotes by exploiting stacking frameworkabstractOrigins of replication sites (ORIs), which refers to the initiative locations of genomic DNA replication, play essential roles in DNA replication process. Detection of ORIs' distribution in genome scale is one of key steps to in-depth understanding their regulation mechanisms. In this study, we presented a novel machine learning-based approach called Stack-ORI encompassing 10 cell-specific prediction models for identifying ORIs from four different eukaryotic species (Homo sapiens, Mus musculus, Drosophila melanogaster and Arabidopsis thaliana). For each cell-specific model, we employed 12 feature encoding schemes that cover nucleic acid composition, position-specific and physicochemical properties information. The optimal feature set was identified from each encoding individually and developed their respective baseline models using the eXtreme Gradient Boosting (XGBoost) classifier. Subsequently, the predicted scores of 12 baseline models are integrated as a novel feature vector to train XGBoost and develop the final model. Extensive experimental results show that Stack-ORI achieves significantly better performance as compared with their baseline models on both training and independent datasets. Interestingly, Stack-ORI consistently outperforms existing predictor in all cell-specific models, not only on training but also on independent test. Moreover, our novel approach provides necessary interpretations that help understanding model success by leveraging the powerful SHapley Additive exPlanation algorithm, thus underlining the most important feature encoding schemes significant for predicting cell-specific ORIs. Leyi Wei, Adeel Malik, Ran Su, Li-Zhen Cui 0001, Balachandran Manavalan |
Briefings Bioinform. | 4 |
| 2021 | Learning embedding features based on multisense-scaled attention architecture to improve the predictive performance of anticancer peptidesabstractMOTIVATION: Anticancer peptides (ACPs) have recently emerged as effective anticancer drugs in cancer therapy. Machine learning-based predictors have been developed to identify ACPs and achieve satisfactory performance. However, existing methods suffer from experience-based feature engineering, which not only restricts the representation ability of the models to a certain extent but also lacks adaptivity for different data, limiting the further improvement of the predictive performance and impacting the robustness of the predictive models. To alleviate the above problems, we propose a novel deep-learning-based predictor named ACPred-LAF, in which we propose a novel multisense and multiscaled embedding algorithm to automatically learn and extract context sequential characteristics of ACPs. RESULTS: Through the feature comparative analysis, we demonstrate that our learnable and self-adaptive embedding features are better than hand-crafted features in capturing discriminative information, which can effectively benefit the performance improvement for ACP prediction. In addition, benchmarking comparison results demonstrate that our ACPred-LAF outperforms the state-of-the-art methods both on existing benchmark datasets and our newly constructed dataset. Furthermore, we also prove and validate the robustness of the model via the data interference experiment. To avoid potential evaluation bias, here, we construct a new ACP benchmark dataset named ACP-Mixed by integrating existing datasets. We expect our newly constructed dataset to be a golden standard benchmark dataset in this field. To facilitate the use of our model, we develop a web server as the implementation of ACPred-LAF. AVAILABILITY AND IMPLEMENTATION: Our proposed ACPred-LAF, newly constructed benchmark dataset ACP-Mixed are open source collaborative initiatives available in the GitHub repository (https://github.com/TearsWaiting/ACPred-LAF). Besides, a webserver as the implementation of ACPred-LAF that can be accessed via: http://server.malab.cn/ACPred-LAF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Li-Zhen Cui 0001, Ran Su, Leyi Wei |
Bioinform. | 4 |
| 2021 | Domain adaptation based self-correction model for COVID-19 infection segmentation in CT images
Qiangguo Jin, Hui Cui 0002, Changming Sun, Zhaopeng Meng, Leyi Wei, Ran Su |
Expert Syst. Appl. | 6 |
| 2021 | Free-form tumor synthesis in computed tomography images via richer generative adversarial network
Qiangguo Jin, Hui Cui 0002, Changming Sun, Zhaopeng Meng, Ran Su |
Knowl. Based Syst. | 5 |
| 2021 | Identification of glioblastoma molecular subtype and prognosis based on deep MRI features
Ran Su, Qiangguo Jin, Xiaofeng Liu 0004, Leyi Wei |
Knowl. Based Syst. | 1 |
| 2020 | Towards Human Motion Tracking: An Open-source Platform based on Multi-sensory Fusion MethodsabstractHuman motion tracking (HMT) has been a research focus in the last decades. In this paper, we propose an IMU/TOA-fusion-based platform to solve this problem. Firstly, Time-of-arrival (TOA)-based distance ranging method is considered to compensate for the drifting errors and accumulation introduced by inertial sensors. Secondly, a geometrical kinematic model and maximum correntropy criterion (MCC)-based Kalman filter method are proposed to fuse the multiple information. The open-source hardware and software are detailed in this paper for real-time human motion capture and reconstruction applications. Experiment results show that our proposed hardware can be easily equipped for total body motion reconstruction with a considerable enhancement of the wear-ability and comfort. Furthermore, the main achievements have been presented with a performance comparison between the proposed platform and state-of-the-art commercial ones. Above all, our proposed platform can significantly suppress the accumulative error and drifting problem of conventional inertial systems. More importantly, it realizes the open-source software and hardware, thus it has promising prospects for wearable human motion tracking applications. Cheng Xu 0003, Ran Su, Shihong Duan |
SMC | 2 |
| 2020 | CPPred-FL: a sequence-based predictor for large-scale identification of cell-penetrating peptides by feature representation learningabstractCell-penetrating peptides (CPPs) have been shown to be a transport vehicle for delivering cargoes into live cells, offering great potential as future therapeutics. It is essential to identify CPPs for better understanding of their functional mechanisms. Machine learning-based methods have recently emerged as a main approach for computational identification of CPPs. However, one of the main challenges and difficulties is to propose an effective feature representation model that sufficiently exploits the inner difference and relevance between CPPs and non-CPPs, in order to improve the predictive performance. In this paper, we have developed CPPred-FL, a powerful bioinformatics tool for fast, accurate and large-scale identification of CPPs. In our predictor, we introduce a new feature representation learning scheme that enables one to learn feature representations from totally 45 well-trained random forest models with multiple feature descriptors from different perspectives, such as compositional information, position-specific information and physicochemical properties, etc. We integrate class and probabilistic information into our feature representations. To improve the feature representation ability, we further remove redundant and irrelevant features by feature space optimization. Benchmarking experiments showed that CPPred-FL, using 19 informative features only, is able to achieve better performance than the state-of-the-art predictors. We anticipate that CPPred-FL will be a powerful tool for large-scale identification of CPPs, facilitating the characterization of their functional mechanisms and accelerating their applications in clinical therapy. Xiaoli Qiang, Xiucai Ye, Pufeng Du, Ran Su, Leyi Wei |
Briefings Bioinform. | 5 |
| 2020 | ACPred-Fuse: fusing multi-view information improves the prediction of anticancer peptidesabstractFast and accurate identification of the peptides with anticancer activity potential from large-scale proteins is currently a challenging task. In this study, we propose a new machine learning predictor, namely, ACPred-Fuse, that can automatically and accurately predict protein sequences with or without anticancer activity in peptide form. Specifically, we establish a feature representation learning model that can explore class and probabilistic information embedded in anticancer peptides (ACPs) by integrating a total of 29 different sequence-based feature descriptors. In order to make full use of various multiview information, we further fused the class and probabilistic features with handcrafted sequential features and then optimized the representation ability of the multiview features, which are ultimately used as input for training our prediction model. By comparing the multiview features and existing feature descriptors, we demonstrate that the fused multiview features have more discriminative ability to capture the characteristics of ACPs. In addition, the information from different views is complementary for the performance improvement. Finally, our benchmarking comparison results showed that the proposed ACPred-Fuse is more precise and promising in the identification of ACPs than existing predictors. To facilitate the use of the proposed predictor, we built a web server, which is now freely available via http://server.malab.cn/ACPred-Fuse. Bing Rao, Guoying Zhang, Ran Su, Leyi Wei |
Briefings Bioinform. | 4 |
| 2020 | Empirical comparison and analysis of web-based cell-penetrating peptide prediction toolsabstractCell-penetrating peptides (CPPs) facilitate the delivery of therapeutically relevant molecules, including DNA, proteins and oligonucleotides, into cells both in vitro and in vivo. This unique ability explores the possibility of CPPs as therapeutic delivery and its potential applications in clinical therapy. Over the last few decades, a number of machine learning (ML)-based prediction tools have been developed, and some of them are freely available as web portals. However, the predictions produced by various tools are difficult to quantify and compare. In particular, there is no systematic comparison of the web-based prediction tools in performance, especially in practical applications. In this work, we provide a comprehensive review on the biological importance of CPPs, CPP database and existing ML-based methods for CPP prediction. To evaluate current prediction tools, we conducted a comparative study and analyzed a total of 12 models from 6 publicly available CPP prediction tools on 2 benchmark validation sets of CPPs and non-CPPs. Our benchmarking results demonstrated that a model from the KELM-CPPpred, namely KELM-hybrid-AAC, showed a significant improvement in overall performance, when compared to the other 11 prediction models. Moreover, through a length-dependency analysis, we find that existing prediction tools tend to more accurately predict CPPs and non-CPPs with the length of 20-25 residues long than peptides in other length ranges. Ran Su, Quan Zou 0001, Balachandran Manavalan, Leyi Wei |
Briefings Bioinform. | 1 |
| 2020 | MinE-RFE: determine the optimal subset from RFE by minimizing the subset-accuracy-defined energyabstractRecursive feature elimination (RFE), as one of the most popular feature selection algorithms, has been extensively applied to bioinformatics. During the training, a group of candidate subsets are generated by iteratively eliminating the least important features from the original features. However, how to determine the optimal subset from them still remains ambiguous. Among most current studies, either overall accuracy or subset size (SS) is used to select the most predictive features. Using which one or both and how they affect the prediction performance are still open questions. In this study, we proposed MinE-RFE, a novel RFE-based feature selection approach by sufficiently considering the effect of both factors. Subset decision problem was reflected into subset-accuracy space and became an energy-minimization problem. We also provided a mathematical description of the relationship between the overall accuracy and SS using Gaussian Mixture Models together with spline fitting. Besides, we comprehensively reviewed a variety of state-of-the-art applications in bioinformatics using RFE. We compared their approaches of deciding the final subset from all the candidate subsets with MinE-RFE on diverse bioinformatics data sets. Additionally, we also compared MinE-RFE with some well-used feature selection algorithms. The comparative results demonstrate that the proposed approach exhibits the best performance among all the approaches. To facilitate the use of MinE-RFE, we further established a user-friendly web server with the implementation of the proposed approach, which is accessible at http://qgking.wicp.net/MinE/. We expect this web server will be a useful tool for research community. Ran Su, Leyi Wei |
Briefings Bioinform. | 1 |
| 2020 | Meta-GDBP: a high-level stacked regression model to improve anticancer drug response predictionabstractAnticancer drug response prediction plays an important role in personalized medicine. In particular, precisely predicting drug response in specific cancer types and patients is still a challenge problem. Here we propose Meta-GDBP, a novel anticancer drug-response model, which involves two levels. At the first level of Meta-GDBP, we build four optimized base models (BMs) using genetic information, chemical properties and biological context with an ensemble optimization strategy, while at the second level, we construct a weighted model to integrate the four BMs. Notably, the weights of the models are learned upstream, thus the parameter cost is significantly reduced compared to previous methods. We evaluate the Meta-GDBP on Genomics of Drug Sensitivity in Cancer (GDSC) and the Cancer Cell Line Encyclopedia (CCLE) data sets. Benchmarking results demonstrate that compared to other methods, the Meta-GDBP achieves a much higher correlation between the predicted and the observed responses for almost all the drugs. Moreover, we apply the Meta-GDBP to predict the GDSC-missing drug response and use the CCLE-known data to validate the performance. The results show quite a similar tendency between these two response sets. Particularly, we here for the first time introduce a biological context-based frequency matrix (BCFM) to associate the biological context with the drug response. It is encouraging that the proposed BCFM is biologically meaningful and consistent with the reported biological mechanism, further demonstrating its efficacy for predicting drug response. The R implementation for the proposed Meta-GDBP is available at https://github.com/RanSuLab/Meta-GDBP. Ran Su, Guobao Xiao, Leyi Wei |
Briefings Bioinform. | 1 |
| 2020 | Comparative analysis and prediction of quorum-sensing peptides using feature representation learning and machine learning algorithmsabstractQuorum-sensing peptides (QSPs) are the signal molecules that are closely associated with diverse cellular processes, such as cell-cell communication, and gene expression regulation in Gram-positive bacteria. It is therefore of great importance to identify QSPs for better understanding and in-depth revealing of their functional mechanisms in physiological processes. Machine learning algorithms have been developed for this purpose, showing the great potential for the reliable prediction of QSPs. In this study, several sequence-based feature descriptors for peptide representation and machine learning algorithms are comprehensively reviewed, evaluated and compared. To effectively use existing feature descriptors, we used a feature representation learning strategy that automatically learns the most discriminative features from existing feature descriptors in a supervised way. Our results demonstrate that this strategy is capable of effectively capturing the sequence determinants to represent the characteristics of QSPs, thereby contributing to the improved predictive performance. Furthermore, wrapping this feature representation learning strategy, we developed a powerful predictor named QSPred-FL for the detection of QSPs in large-scale proteomic data. Benchmarking results with 10-fold cross validation showed that QSPred-FL is able to achieve better performance as compared to the state-of-the-art predictors. In addition, we have established a user-friendly webserver that implements QSPred-FL, which is currently available at http://server.malab.cn/QSPred-FL. We expect that this tool will be useful for the high-throughput prediction of QSPs and the discovery of important functional mechanisms of QSPs. Leyi Wei, Fuyi Li, Jiangning Song, Ran Su, Quan Zou 0001 |
Briefings Bioinform. | 5 |
| 2020 | Identification of expression signatures for non-small-cell lung carcinoma subtype classificationabstractMOTIVATION: Non-small-cell lung carcinoma (NSCLC) mainly consists of two subtypes: lung squamous cell carcinoma (LUSC) and lung adenocarcinoma (LUAD). It has been reported that the genetic and epigenetic profiles vary strikingly between LUAD and LUSC in the process of tumorigenesis and development. Efficient and precise treatment can be made if subtypes can be identified correctly. Identification of discriminative expression signatures has been explored recently to aid the classification of NSCLC subtypes. RESULTS: In this study, we designed a classification model integrating both mRNA and long non-coding RNA (lncRNA) expression data to effectively classify the subtypes of NSCLC. A gene selection algorithm, named WGRFE, was proposed to identify the most discriminative gene signatures within the recursive feature elimination (RFE) framework. GeneRank scores considering both expression level and correlation, together with the importance generated by classifiers were all taken into account to improve the selection performance. Moreover, a module-based initial filtering of the genes was performed to reduce the computation cost of RFE. We validated the proposed algorithm on The Cancer Genome Atlas (TCGA) dataset. The results demonstrate that the developed approach identified a small number of expression signatures for accurate subtype classification and particularly, we here for the first time show the potential role of LncRNA in building computational NSCLC subtype classification models. AVAILABILITY AND IMPLEMENTATION: The R implementation for the proposed approach is available at https://github.com/RanSuLab/NSCLC-subtype-classification. Ran Su, Xiaofeng Liu 0004, Leyi Wei |
Bioinform. | 1 |
| 2020 | Fusing convolutional neural network features with hand-crafted features for osteoporosis diagnoses
Ran Su, Tianling Liu, Changming Sun, Qiangguo Jin, Rachid Jennane, Leyi Wei |
Neurocomputing | 1 |
| 2020 | Construction of Retinal Vessel Segmentation Models Based on Convolutional Neural Network
Qiangguo Jin, Zhaopeng Meng, Ran Su |
Neural Process. Lett. | 5 |
| 2020 | Ultra-Wideband Radio Channel Characteristics for Near-Ground Swarm Robots CommunicationabstractUltra-Wideband (UWB) technology has great potential for the cooperation and navigation among near ground mobile robots in GPS-denied environments. In this paper, an efficient two-segment UWB radio channel model is proposed with considering the multi-path condition in very near-ground environments and different surface roughness. We conducted field measurements to collect channel information, with both transmitter and receiver antennas placed at different heights above the ground: 0cm-20cm. Signal frequency was chosen at 4.3GHz with bandwidth of 1GHz. Three ground coverings were tested in common scenarios: brick, grass and robber fields. The proposed model has enhanced accuracy achieved by careful assessment of dominant propagation mechanisms in each segment, such as diffraction loss due to obstruction of the first Fresnel zone and higher-order waves produced by ground roughness. It is realized that antenna height and distance are the most influential geometric parameters to affect the path loss model. Once the antenna height is known, there exists a breakpoint distance in UWB propagation, which separates two segmentation using the different path-loss mechanism. Different surface types can cause different signal attenuation. Monte Carlo simulations are used to investigate the effects of antenna height, distance, ground surface type on mobile robots swarm communication to find out the antenna height is also a dominant factor on connectivity and the average number of neighbors. Within a certain range, the higher the antenna height and the closer the communication distance, the better the communication performance will usually be. Cramér-Rao lower bound(CRLB) of path loss estimator based on the proposed model is derived to show the relationship between CRLB with height and distance. Shihong Duan, Ran Su, Cheng Xu 0003, Jie He 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2019 | Exploring sequence-based features for the improved prediction of DNA N4-methylcytosine sites in multiple speciesabstractMOTIVATION: As one of important epigenetic modifications, DNA N4-methylcytosine (4mC) is recently shown to play crucial roles in restriction-modification systems. For better understanding of their functional mechanisms, it is fundamentally important to identify 4mC modification. Machine learning methods have recently emerged as an effective and efficient approach for the high-throughput identification of 4mC sites, although high predictive error rates are still challenging for existing methods. Therefore, it is highly desirable to develop a computational method to more accurately identify m4C sites. RESULTS: In this study, we propose a machine learning based predictor, namely 4mcPred-SVM, for the genome-wide detection of DNA 4mC sites. In this predictor, we present a new feature representation algorithm that sufficiently exploits sequence-based information. To improve the feature representation ability, we use a two-step feature optimization strategy, thereby obtaining the most representative features. Using the resulting features and Support Vector Machine (SVM), we adaptively train the optimal models for different species. Comparative results on benchmark datasets from six species indicate that our predictor is able to achieve generally better performance in predicting 4mC sites as compared to the state-of-the-art predictors. Importantly, the sequence-based features can reliably and robust predict 4mC sites, facilitating the discovery of potentially important sequence characteristics for the prediction of 4mC sites. AVAILABILITY AND IMPLEMENTATION: The user-friendly webserver that implements the proposed 4mcPred-SVM is well established, and is freely accessible at http://server.malab.cn/4mcPred-SVM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leyi Wei, Shasha Luan, Luis Augusto Eijy Nagai, Ran Su, Quan Zou 0001 |
Bioinform. | 4 |
| 2019 | Iterative feature representations improve N4-methylcytosine site predictionabstractMOTIVATION: Accurate identification of N4-methylcytosine (4mC) modifications in a genome wide can provide insights into their biological functions and mechanisms. Machine learning recently have become effective approaches for computational identification of 4mC sites in genome. Unfortunately, existing methods cannot achieve satisfactory performance, owing to the lack of effective DNA feature representations that are capable to capture the characteristics of 4mC modifications. RESULTS: In this work, we developed a new predictor named 4mcPred-IFL, aiming to identify 4mC sites. To represent and capture discriminative features, we proposed an iterative feature representation algorithm that enables to learn informative features from several sequential models in a supervised iterative mode. Our analysis results showed that the feature representations learnt by our algorithm can capture the discriminative distribution characteristics between 4mC sites and non-4mC sites, enlarging the decision margin between the positives and negatives in feature space. Additionally, by evaluating and comparing our predictor with the state-of-the-art predictors on benchmark datasets, we demonstrate that our predictor can identify 4mC sites more accurately. AVAILABILITY AND IMPLEMENTATION: The user-friendly webserver that implements the proposed 4mcPred-IFL is well established, and is freely accessible at http://server.malab.cn/4mcPred-IFL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leyi Wei, Ran Su, Shasha Luan, Zhijun Liao, Balachandran Manavalan, Quan Zou 0001 |
Bioinform. | 2 |
| 2019 | PEPred-Suite: improved and robust prediction of therapeutic peptides using adaptive feature representation learningabstractMOTIVATION: Prediction of therapeutic peptides is critical for the discovery of novel and efficient peptide-based therapeutics. Computational methods, especially machine learning based methods, have been developed for addressing this need. However, most of existing methods are peptide-specific; currently, there is no generic predictor for multiple peptide types. Moreover, it is still challenging to extract informative feature representations from the perspective of primary sequences. RESULTS: In this study, we have developed PEPred-Suite, a bioinformatics tool for the generic prediction of therapeutic peptides. In PEPred-Suite, we introduce an adaptive feature representation strategy that can learn the most representative features for different peptide types. To be specific, we train diverse sequence-based feature descriptors, integrate the learnt class information into our features, and utilize a two-step feature optimization strategy based on the area under receiver operating characteristic curve to extract the most discriminative features. Using the learnt representative features, we trained eight random forest models for eight different types of functional peptides, respectively. Benchmarking results showed that as compared with existing predictors, PEPred-Suite achieves better and robust performance for different peptides. As far as we know, PEPred-Suite is currently the first tool that is capable of predicting so many peptide types simultaneously. In addition, our work demonstrates that the learnt features can reliably predict different peptides. AVAILABILITY AND IMPLEMENTATION: The user-friendly webserver implementing the proposed PEPred-Suite is freely accessible at http://server.malab.cn/PEPred-Suite. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leyi Wei, Ran Su, Quan Zou 0001 |
Bioinform. | 3 |
| 2019 | Integration of deep feature representations and handcrafted features to improve the prediction of N6-methyladenosine sites
Leyi Wei, Ran Su, Xiu-Ting Li, Quan Zou 0001, Xing Gao 0004 |
Neurocomputing | 2 |
| 2019 | DUNet: A deformable network for retinal vessel segmentation
Qiangguo Jin, Zhaopeng Meng, Tuan D. Pham, Leyi Wei, Ran Su |
Knowl. Based Syst. | 6 |
| 2019 | Developing a Multi-Dose Computational Model for Drug-Induced Hepatotoxicity Prediction Based on Toxicogenomics DataabstractDrug-induced hepatotoxicity may cause acute and chronic liver disease, leading to great concern for patient safety. It is also one of the main reasons for drug withdrawal from the market. Toxicogenomics data has been widely used in hepatotoxicity prediction. In our study, we proposed a multi-dose computational model to predict the drug-induced hepatotoxicity based on gene expression and toxicity data. The dose/concentration information after drug treatment is fully utilized in our study based on the dose-response curve, thus a more informative representative of the dose-response relationship is considered. We also proposed a new feature selection method, named MEMO, which is also one important aspect of our multi-dose model in our study, to deal with the high-dimensional toxicogenomics data. We validated the proposed model using the TG-GATEs, which is a large database recording toxicogenomics data from multiple views. The experimental results show that the drug-induced hepatotoxicity can be predicted with high accuracy and efficiency using the proposed predictive model. Ran Su, Huichen Wu, Xiaofeng Liu 0004, Leyi Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | Encoded Texture Features to Characterize Bone Radiograph ImagesabstractOsteoporosis is the most common reason that causes the fracture among the elderly. For the purpose of convenience and safety, 2D texture analysis has been used to diagnose osteoporosis. In this study, a supervised method using proposed texture features to identify osteoporotic cases from healthy was proposed. We designed two groups of new features, Encoded GLCM and Encoded LBP, each of which contains two subgroups through encoding the Gabor and Hessian information into the Gray Level Co-Occurrence Matrix (GLCM) features and Local Binary Patterns (LBP) features respectively. These two groups of features, together with the raw feature group containing the GLCM and LBP features, totally 560 features, were categorized into various groups and used to train the Random Forest classifier. Classification performances using these features were compared inter-and intra-groups/subgroups. And the performance using each individual feature was also provided. We conducted feature selection based on Recursive Feature Elimination (RFE) inside a voting scheme to further increase the efficiency. The inter-and intra-groups/subgroups results indicate that the Encoded GLCM and Encoded LBP, are more discriminative than the raw GLCM and LBP features for the identification of the osteoporosis; The best individual feature is from the Encoded LBP group and can achieve 70% of balanced accuracy; Furthermore, using only ten of the proposed features through feature selection, the balanced accuracy can even be improved from 60% to 71%. This shows that the proposed method is promising to assist the early diagnosis of osteoporosis. Ran Su, Leyi Wei, Xiu-Ting Li, Qiangguo Jin, Wenyuan Tao |
ICPR | 1 |
| 2018 | ACPred-FL: a sequence-based predictor using effective feature representation to improve the prediction of anti-cancer peptidesabstractMotivation: Anti-cancer peptides (ACPs) have recently emerged as promising therapeutic agents for cancer treatment. Due to the avalanche of protein sequence data in the post-genomic era, there is an urgent need to develop automated computational methods to enable fast and accurate identification of novel ACPs within the vast number of candidate proteins and peptides. Results: To address this, we propose a novel predictor named Anti-Cancer peptide Predictor with Feature representation Learning (ACPred-FL) for accurate prediction of ACPs based on sequence information. More specifically, we develop an effective feature representation learning model, with which we can extract and learn a set of informative features from a pool of support vector machine-based models trained using sequence-based feature descriptors. By doing so, the class label information of data samples is fully utilized. To improve the feature representation, we further employ a two-step feature selection technique, resulting in a most informative five-dimensional feature vector for the final peptide representation. Experimental results show that such five features provide the most discriminative power for identifying ACPs than currently available feature descriptors, highlighting the effectiveness of the proposed feature representation learning approach. The developed ACPred-FL method significantly outperforms state-of-the-art methods. Availability and implementation: The web-server of ACPred-FL is available at http://server.malab.cn/ACPred-FL. Supplementary information: Supplementary data are available at Bioinformatics online. Leyi Wei, Huangrong Chen, Jiangning Song, Ran Su |
Bioinform. | 5 |
| 2018 | Prediction of human protein subcellular localization using deep learning
Leyi Wei, Yijie Ding, Ran Su, Jijun Tang, Quan Zou 0001 |
J. Parallel Distributed Comput. | 3 |
| 2017 | Improved prediction of protein-protein interactions using novel negative samples, features, and an ensemble classifier
Leyi Wei, Pengwei Xing, Jian-Cang Zeng, Jin-Xiu Chen, Ran Su, Fei Guo 0001 |
Artif. Intell. Medicine | 5 |
| 2014 | Supervised prediction of drug-induced nephrotoxicity based on interleukin-6 and -8 expression levelsabstractBACKGROUND: Drug-induced nephrotoxicity causes acute kidney injury and chronic kidney diseases, and is a major reason for late-stage failures in the clinical trials of new drugs. Therefore, early, pre-clinical prediction of nephrotoxicity could help to prioritize drug candidates for further evaluations, and increase the success rates of clinical trials. Recently, an in vitro model for predicting renal-proximal-tubular-cell (PTC) toxicity based on the expression levels of two inflammatory markers, interleukin (IL)-6 and -8, has been described. However, this and other existing models usually use linear and manually determined thresholds to predict nephrotoxicity. Automated machine learning algorithms may improve these models, and produce more accurate and unbiased predictions. RESULTS: Here, we report a systematic comparison of the performances of four supervised classifiers, namely random forest, support vector machine, k-nearest-neighbor and naive Bayes classifiers, in predicting PTC toxicity based on IL-6 and -8 expression levels. Using a dataset of human primary PTCs treated with 41 well-characterized compounds that are toxic or not toxic to PTC, we found that random forest classifiers have the highest cross-validated classification performance (mean balanced accuracy = 87.8%, sensitivity = 89.4%, and specificity = 85.9%). Furthermore, we also found that IL-8 is more predictive than IL-6, but a combination of both markers gives higher classification accuracy. Finally, we also show that random forest classifiers trained automatically on the whole dataset have higher mean balanced accuracy than a previous threshold-based classifier constructed for the same dataset (99.3% vs. 80.7%). CONCLUSIONS: Our results suggest that a random forest classifier can be used to automatically predict drug-induced PTC toxicity based on the expression levels of IL-6 and -8. Ran Su, Daniele Zink, Lit-Hsin Loo |
BMC Bioinform. | 1 |
| 2014 | A new method for linear feature and junction enhancement in 2D images based on morphological operation, oriented anisotropic Gaussian function and Hessian information
Ran Su, Changming Sun, Chao Zhang 0011, Tuan D. Pham |
Pattern Recognit. | 1 |
| 2012 | Junction detection for linear structures based on Hessian, correlation and shape information
Ran Su, Changming Sun, Tuan D. Pham |
Pattern Recognit. | 1 |