EDBT 2026 Demo / reviewers in the wild / expert
Jianyu Shi
dblp:05/2421 · also Jian-Yu Shi
· DBLP profile ↗
39ranked-venue papers
10as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 9 first-author · 19 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unlocking molecular optimization potential by leveraging active learning for scaffold-preserving optimization with ALMO
Zhen-Yi Wu, Peng-Cheng Zhao, Jia-Ning Li, Jianyu Shi, Bing-Xue Du |
Expert Syst. Appl. | 4 |
| 2026 | An Explainable Molecular Token Estimation Method for Knowledge-Aware Drug-Drug Interaction PredictionabstractIn molecular representation learning(MRL), tokens (e.g., atoms, motifs, and fingerprints) are the basic elements to represent molecules. It is a common practice by using various tokens to enhance the expressive power of Graph Neural Networks (GNNs) on molecular graphs. Although prior GNNs-based methods employing tokens achieve promising performances in drug-drug interaction (DDI) prediction, the influence of the token on the expressiveness of molecular embedding models remains underexplored. To bridge the gap, we provide an axiomatic definition of MRL from a frequency domain perspective, revealing that the model's performance is closely related to the number of tokens and deriving a theoretical upper bound of likelihood-based model convergency. Building on these insights, we propose SimMotifPro, a simple yet efficient motif-based method, for DDI prediction. Specifically, SimMotifPro uses a variant of DeeperGCN encoder and builds a motif-motif knowledge graph to capture motif interconnections. A Motif Ranker module is also introduced to decouple learned representations and differentiate the contributions of selected motifs. Empirically, we demonstrate that SimMotifPro adheres to the properties demonstrated in our theoretical upper bound and validate the general applicability of our theory across different methods. Furthermore, our approach achieves state-of-the-art performance on various benchmarks for DDI prediction. Hui Yu 0011, Xinkun Li, Jianyu Shi |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | Small Target Insect Detection Based on Improved YOLOv8nabstractInsect pests greatly affect the growth and harvest of crops, and accurate identification of insect species is very important to the agricultural industry. The insect dataset has the problem that the target is too small to locate, detect, and identify. To solve this problem, We designed a modified YOLOv8n for detecting small target insects. First, we add a small object detection head to the network structure of the YOLOv8n (You Only Look Once) model. Then, we improve C2f and Conv in YOLOv8n, construct C2f-RFAConv by adding Receptive-Field Attention Convolutional Operation (RFAConv) into C2f for the first time, and replace the original convolution with RFAConv for feature extraction. Finally, we improve Wise-IoU (WIoU) loss function to solve the problem of small object detection difficulty. When the number of parameters is reduced, the performance of our model YOLOv8n-Improved on the Yellow Sticky Traps dataset and self-built dataset is greatly improved compared with YOLOv8n, and it also has great advantages compared with other models in the YOLO series. This work plays an important role in our efforts to achieve intelligent pest management in agriculture. Jianyu Shi, Yuan Jia, Zhenhong Jia |
ICASSP | 1 |
| 2025 | EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone GenerationabstractDesigning enzyme backbones with substrate-specific functionality is a critical challenge in computational protein engineering. Current generative models excel in protein design but face limitations in binding data, substrate-specific control, and flexibility for de novo enzyme backbone generation. To address this, we introduce **EnzyBind**, a dataset with 11,100 experimentally validated enzyme-substrate pairs specifically curated from PDBbind. Building on this, we propose **EnzyControl**, a method that enables functional and substrate-specific control in enzyme backbone generation. Our approach generates enzyme backbones conditioned on MSA-annotated catalytic sites and their corresponding substrates, which are automatically extracted from curated enzyme-substrate data. At the core of EnzyControl is **EnzyAdapter**, a lightweight, modular component integrated into a pretrained motif-scaffolding model, allowing it to become substrate-aware. A two-stage training paradigm further refines the model's ability to generate accurate and functional enzyme structures. Experiments show that our EnzyControl achieves the best performance across structural and functional metrics on EnzyBind and EnzyBench benchmarks, with particularly notable improvements of 13% in designability and 13% in catalytic efficiency compared to the baseline models. The code is released at https://github.com/Vecteur-libre/EnzyControl. Jianyu Shi, Hui Yu 0011, Yihang Zhou |
NeurIPS | 6 |
| 2025 | Few-shot drug synergy prediction via rapid cross-tier adaptation meta-optimizationabstractDrug combination therapy offers key advantages over monotherapy in personalized oncology by reducing drug resistance and toxicity. However, predicting synergistic effects for rare cell lines remains challenging, as existing methods suffer from poor generalizability in data-scarce scenarios owing to their reliance on large training datasets and inability to effectively transfer knowledge across distinct cellular contexts. Here, we present MetaSynergy, a Rapid Cross-tier Adaptation Meta-Optimization (R-CAMO)-based framework for few-shot drug synergy prediction through cross-domain knowledge transfer and meta-optimized adaptation. We first designed a multimodal feature learning architecture integrating drug molecular graphs with cell line omics profiles, then implemented a stage-wise training strategy based on R-CAMO for few-shot drug synergy prediction: (i) cross-domain pretraining establishes meta-initialized representations by transferring knowledge from data-rich cell lines to scarce target domains, enhancing feature representation capability in data-scarce scenarios. (ii) Cross-tier meta-optimization enables rapid adaptation to data-scarce scenarios: the inner-tier refines task-specific parameters of the prediction network on the target domain, while the outer-tier meta-learns task-shared, generalizable parameters by minimizing the cross-cell line prediction loss. (iii) Fine-tuning further refines task-specific parameters, improving generalizability to novel drug combinations within the same cellular context. Experimental results demonstrate that MetaSynergy achieves excellent performance in few-shot, zero-shot and low-similarity tasks, surpassing most baseline methods and highlighting its robustness and generalizability. Ablation studies confirmed the pivotal role of R-CAMO strategy in data-scarce cell lines. Furthermore, MetaSynergy successfully identified novel synergistic drug combinations in several understudied malignancies, underscoring its potential in precision oncology. Yue-Hua Feng, Ze-Lin Feng, Xiao-Ying Yan, Shaowu Zhang 0001, Jianyu Shi |
Briefings Bioinform. | 5 |
| 2024 | Recognition of cyanobacteria promoters via Siamese network-based contrastive learning under novel non-promoter generationabstractIt is a vital step to recognize cyanobacteria promoters on a genome-wide scale. Computational methods are promising to assist in difficult biological identification. When building recognition models, these methods rely on non-promoter generation to cope with the lack of real non-promoters. Nevertheless, the factitious significant difference between promoters and non-promoters causes over-optimistic prediction. Moreover, designed for E. coli or B. subtilis, existing methods cannot uncover novel, distinct motifs among cyanobacterial promoters. To address these issues, this work first proposes a novel non-promoter generation strategy called phantom sampling, which can eliminate the factitious difference between promoters and generated non-promoters. Furthermore, it elaborates a novel promoter prediction model based on the Siamese network (SiamProm), which can amplify the hidden difference between promoters and non-promoters through a joint characterization of global associations, upstream and downstream contexts, and neighboring associations w.r.t. k-mer tokens. The comparison with state-of-the-art methods demonstrates the superiority of our phantom sampling and SiamProm. Both comprehensive ablation studies and feature space illustrations also validate the effectiveness of the Siamese network and its components. More importantly, SiamProm, upon our phantom sampling, finds a novel cyanobacterial promoter motif ('GCGATCGC'), which is palindrome-patterned, content-conserved, but position-shifted. Guang Yang 0043, Jinlu Hu, Jianyu Shi |
Briefings Bioinform. | 4 |
| 2024 | GGI-DDI: Identification for key molecular substructures by granule learning to interpret predicted drug-drug interactions
Hui Yu 0011, Omayo Silver, Zun Liu, JingTao Yao 0001, Jianyu Shi |
Expert Syst. Appl. | 7 |
| 2024 | Multi-view clustering with semantic fusion and contrastive learning
Hui Yu 0011, Hui-Xiang Bian, Zi-Ling Chong, Zun Liu, Jianyu Shi |
Neurocomputing | 5 |
| 2024 | Identifying the reaction centers of molecule based on dual-view representation
Hui Yu 0011, Jianyu Shi |
Knowl. Based Syst. | 4 |
| 2023 | GELKcat: An Integration Learning of Substrate Graph with Enzyme Embedding for Kcat predictionabstractComputational modeling and identification of the enzyme turnover number kcatare crucial for synthetic biology and early-stage lead optimization. Therefore, the accurate assessment of the kcatfor enzyme-substrate pairs is essential. Considering wet-lab experiment is time-consuming, laborious, and expensive, in silico prediction of kcatis an alternative choice. However, few computational methods have been developed to address this task and other enzyme kinetics predictions. To address this, we develop a novel end-to-end dual-representation framework GELKcat by harnessing graph transformers for substrate molecular encoding and CNNs for enzyme word2vec embeddings. We further integrate substrate and enzyme features using the adaptive gate network, which assigns optimal weights to capture the most suitable feature combinations. The comparison with several state-of-the-art methods exhibits the superiority of our GELKcat. The Ablation studies further illuminate the invaluable roles of the word2vec embeddings of enzymes. It is anticipated that this work can bridge current gaps in enzyme-substrate representation, which can give some guidance for drug discovery and synthetic biology. Bing-Xue Du, Bei Zhu, Yahui Long, Min Wu 0008, Jianyu Shi |
BIBM | 6 |
| 2023 | MTGL-ADMET: A Novel Multi-task Graph Learning Framework for ADMET Prediction Enhanced by Status-Theory and Maximum Flow
Bing-Xue Du, Siu-Ming Yiu, Hui Yu 0011, Jianyu Shi |
RECOMB | 5 |
| 2023 | Bamboo Agents: Exploring the Potentiality of Digital Craft by Decoding and Recoding ProcessabstractAs an emerging field in HCI, Digital Craft is often involved in debate on its concept that leads to distinctive practices. In this paper, the authors argue that the hybridization of digital power as computing and fabrication and human skills of ideation and hand-making sheds light on an important research direction for future inquiry of digital craft. In particular, the reported project focuses on bamboo craft making to explore the potentiality of digital craft through constructive design research methods. Through the collaborative making with craftspeople, methods of hybridizing digital power and human skills are explored by decoding and recoding the bamboo making process with inventive digital intervention, which enriches bamboo artifacts’ forms and integrates digital fabrication, such as 3D printing. Meanwhile, digital toolkits for bamboo weaving and a digital platform that can make computational design and craft more compatible are created and reported. The authors conclude with reported hybrid craft cases that the hybridization has the potential to stimulate a new generation of artisans. Peizhong Gao, Tanhao Gao, Yanbin Yang, Zhenyuan Liu 0004, Jianyu Shi, Jin Li 0039 |
TEI | 5 |
| 2023 | A social theory-enhanced graph representation learning framework for multitask prediction of drug-drug interactionsabstractCurrent machine learning-based methods have achieved inspiring predictions in the scenarios of mono-type and multi-type drug-drug interactions (DDIs), but they all ignore enhancive and depressive pharmacological changes triggered by DDIs. In addition, these pharmacological changes are asymmetric since the roles of two drugs in an interaction are different. More importantly, these pharmacological changes imply significant topological patterns among DDIs. To address the above issues, we first leverage Balance theory and Status theory in social networks to reveal the topological patterns among directed pharmacological DDIs, which are modeled as a signed and directed network. Then, we design a novel graph representation learning model named SGRL-DDI (social theory-enhanced graph representation learning for DDI) to realize the multitask prediction of DDIs. SGRL-DDI model can capture the task-joint information by integrating relation graph convolutional networks with Balance and Status patterns. Moreover, we utilize task-specific deep neural networks to perform two tasks, including the prediction of enhancive/depressive DDIs and the prediction of directed DDIs. Based on DDI entries collected from DrugBank, the superiority of our model is demonstrated by the comparison with other state-of-the-art methods. Furthermore, the ablation study verifies that Balance and Status patterns help characterize directed pharmacological DDIs, and that the joint of two tasks provides better DDI representations than individual tasks. Last, we demonstrate the practical effectiveness of our model by a version-dependent test, where 88.47 and 81.38% DDI out of newly added entries provided by the latest release of DrugBank are validated in two predicting tasks respectively. Yue-Hua Feng, Shaowu Zhang 0001, Yi-Yang Feng, Qing-Qing Zhang, Ming-Hui Shi, Jianyu Shi |
Briefings Bioinform. | 6 |
| 2023 | Comprehensive evaluation of deep and graph learning on drug-drug interactions predictionabstractRecent advances and achievements of artificial intelligence (AI) as well as deep and graph learning models have established their usefulness in biomedical applications, especially in drug-drug interactions (DDIs). DDIs refer to a change in the effect of one drug to the presence of another drug in the human body, which plays an essential role in drug discovery and clinical research. DDIs prediction through traditional clinical trials and experiments is an expensive and time-consuming process. To correctly apply the advanced AI and deep learning, the developer and user meet various challenges such as the availability and encoding of data resources, and the design of computational methods. This review summarizes chemical structure based, network based, natural language processing based and hybrid methods, providing an updated and accessible guide to the broad researchers and development community with different domain knowledge. We introduce widely used molecular representation and describe the theoretical frameworks of graph neural network models for representing molecular structures. We present the advantages and disadvantages of deep and graph learning methods by performing comparative experiments. We discuss the potential technical challenges and highlight future directions of deep and graph learning models for accelerating DDIs prediction. Xuan Lin, Lichang Dai, Yafang Zhou, Jianyu Shi, Dong-Sheng Cao 0001, Bosheng Song, Philip S. Yu, Xiangxiang Zeng |
Briefings Bioinform. | 6 |
| 2023 | Attention-based cross domain graph neural network for prediction of drug-drug interactionsabstractDrug-drug interactions (DDI) may lead to adverse reactions in human body and accurate prediction of DDI can mitigate the medical risk. Currently, most of computer-aided DDI prediction methods construct models based on drug-associated features or DDI network, ignoring the potential information contained in drug-related biological entities such as targets and genes. Besides, existing DDI network-based models could not make effective predictions for drugs without any known DDI records. To address the above limitations, we propose an attention-based cross domain graph neural network (ACDGNN) for DDI prediction, which considers the drug-related different entities and propagate information through cross domain operation. Different from the existing methods, ACDGNN not only considers rich information contained in drug-related biomedical entities in biological heterogeneous network, but also adopts cross-domain transformation to eliminate heterogeneity between different types of entities. ACDGNN can be used in the prediction of DDIs in both transductive and inductive setting. By conducting experiments on real-world dataset, we compare the performance of ACDGNN with several state-of-the-art methods. The experimental results show that ACDGNN can effectively predict DDIs and outperform the comparison models. Hui Yu 0011, Wenmin Dong, Shuanghong Song, Jianyu Shi |
Briefings Bioinform. | 6 |
| 2023 | CMMS-GCL: cross-modality metabolic stability prediction with graph contrastive learningabstractMOTIVATION: Metabolic stability plays a crucial role in the early stages of drug discovery and development. Accurately modeling and predicting molecular metabolic stability has great potential for the efficient screening of drug candidates as well as the optimization of lead compounds. Considering wet-lab experiment is time-consuming, laborious, and expensive, in silico prediction of metabolic stability is an alternative choice. However, few computational methods have been developed to address this task. In addition, it remains a significant challenge to explain key functional groups determining metabolic stability. RESULTS: To address these issues, we develop a novel cross-modality graph contrastive learning model named CMMS-GCL for predicting the metabolic stability of drug candidates. In our framework, we design deep learning methods to extract features for molecules from two modality data, i.e. SMILES sequence and molecule graph. In particular, for the sequence data, we design a multihead attention BiGRU-based encoder to preserve the context of symbols to learn sequence representations of molecules. For the graph data, we propose a graph contrastive learning-based encoder to learn structure representations by effectively capturing the consistencies between local and global structures. We further exploit fully connected neural networks to combine the sequence and structure representations for model training. Extensive experimental results on two datasets demonstrate that our CMMS-GCL consistently outperforms seven state-of-the-art methods. Furthermore, a collection of case studies on sequence data and statistical analyses of the graph structure module strengthens the validation of the interpretability of crucial functional groups recognized by CMMS-GCL. Overall, CMMS-GCL can serve as an effective and interpretable tool for predicting metabolic stability, identifying critical functional groups, and thus facilitating the drug discovery process and lead compound optimization. AVAILABILITY AND IMPLEMENTATION: The code and data underlying this article are freely available at https://github.com/dubingxue/CMMS-GCL. Bing-Xue Du, Yahui Long, Xiaoli Li 0001, Min Wu 0008, Jianyu Shi |
Bioinform. | 5 |
| 2023 | CProMG: controllable protein-oriented molecule generation with desired binding affinity and drug-like propertiesabstractMOTIVATION: Deep learning-based molecule generation becomes a new paradigm of de novo molecule design since it enables fast and directional exploration in the vast chemical space. However, it is still an open issue to generate molecules, which bind to specific proteins with high-binding affinities while owning desired drug-like physicochemical properties. RESULTS: To address these issues, we elaborate a novel framework for controllable protein-oriented molecule generation, named CProMG, which contains a 3D protein embedding module, a dual-view protein encoder, a molecule embedding module, and a novel drug-like molecule decoder. Based on fusing the hierarchical views of proteins, it enhances the representation of protein binding pockets significantly by associating amino acid residues with their comprising atoms. Through jointly embedding molecule sequences, their drug-like properties, and binding affinities w.r.t. proteins, it autoregressively generates novel molecules having specific properties in a controllable manner by measuring the proximity of molecule tokens to protein residues and atoms. The comparison with state-of-the-art deep generative methods demonstrates the superiority of our CProMG. Furthermore, the progressive control of properties demonstrates the effectiveness of CProMG when controlling binding affinity and drug-like properties. After that, the ablation studies reveal how its crucial components contribute to the model respectively, including hierarchical protein views, Laplacian position encoding as well as property control. Last, a case study w.r.t. protein illustrates the novelty of CProMG and the ability to capture crucial interactions between protein pockets and molecules. It's anticipated that this work can boost de novo molecule design. AVAILABILITY AND IMPLEMENTATION: The code and data underlying this article are freely available at https://github.com/lijianing0902/CProMG. Jia-Ning Li, Guang Yang 0043, Peng-Cheng Zhao, Xue-Xin Wei, Jianyu Shi |
Bioinform. | 5 |
| 2023 | Data Augmentation Generated by Generative Adversarial Network for Small Sample Datasets Clustering
Hui Yu 0011, Qiao Feng Wang, Jianyu Shi |
Neural Process. Lett. | 3 |
| 2023 | Few-Shot Drug Synergy Prediction With a Prior-Guided Hypernetwork ArchitectureabstractPredicting drug synergy is critical to tailoring feasible drug combination treatment regimens for cancer patients. However, most of the existing computational methods only focus on data-rich cell lines, and hardly work on data-poor cell lines. To this end, here we proposed a novel few-shot drug synergy prediction method (called HyperSynergy) for data-poor cell lines by designing a prior-guided Hypernetwork architecture, in which the meta-generative network based on the task embedding of each cell line generates cell line dependent parameters for the drug synergy prediction network. In HyperSynergy model, we designed a deep Bayesian variational inference model to infer the prior distribution over the task embedding to quickly update the task embedding with a few labeled drug synergy samples, and presented a three-stage learning strategy to train HyperSynergy for quickly updating the prior distribution by a few labeled drug synergy samples of each data-poor cell line. Moreover, we proved theoretically that HyperSynergy aims to maximize the lower bound of log-likelihood of the marginal distribution over each data-poor cell line. The experimental results show that our HyperSynergy outperforms other state-of-the-art methods not only on data-poor cell lines with a few samples (e.g., 10, 5, 0), but also on data-rich cell lines. Qing-Qing Zhang, Shaowu Zhang 0001, Yue-Hua Feng, Jianyu Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | DGANDDI: Double Generative Adversarial Networks for Drug-Drug Interaction PredictionabstractCo-administration of multiple drugs may cause adverse drug interactions and side effects that damage the body. Therefore, accurate prediction of drug-drug interaction (DDI) events is of great importance. Recently, many computational methods have been proposed for predicting DDI associated events. However, most existing methods merely considered drug associated attribute information or topological information in DDI network, ignoring the complementary knowledge between them. Therefore, to effectively explore the complementarity of drug attribute and topological information of DDI network, we propose a deep learning model based adversarial learning strategy, which is named as DGANDDI. In DGANDDI, we design a two-GAN architecture to deeply capture the complementary knowledge between drug attribute and topological information of DDI network, thus more comprehensive drug representations can be learned. We conduct extensive experiments on real world dataset. The experimental results show that DGANDDI can effectively predict DDI occurrence and outperforms the comparison of the state-of-the-art models. We also perform ablation studies that demonstrate that DGANDDI is effective and that it is robust in DDI prediction tasks, even in the case of a scarcity of labeled DDIs. Hui Yu 0011, Jianyu Shi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Enhanced CT Image Generation by GAN for Improving Thyroid Anatomy DetectionabstractComputed tomography (CT) is one of the most imaging methods widely used to locate lesions such as nodules, tumors, and cysts, and make primary diagnosis. For clearer imaging of anatomical or lesions, contrast-enhanced CT (CECT) scans are imaging with injecting a contrast agent into a patient during examination. But there are limits to iodine contrast injections so that CECT scans are not convenient like non-contrast enhanced CT (NECT). Recently, deep learning models bring impressive results in computer vision, including image translation. So, we would like to apply image translation methods to generate CECT images from the more accessible NECT images, and evaluate the effects of generated images on image detection tasks. In this study, we propose a method called cross-modal enhancement training strategy for thyroid anatomy detection, which employs CycleGAN to translate non-constrast enhanced CT images to enhanced CT style images with content reserved. The experiments are conducted on thyroid CT images with anatomy object annotation. The experimental results show that by adding translated images into the training dataset, the performance of thyroid anatomy detection can be effectively improved. We achieve the best mAP of 82.5% compared to 73.2% in the along non-contrast enhanced CT training. Jianyu Shi, Xiaohong Liu 0007, Guoxing Yang |
BIBM | 1 |
| 2022 | AIAT: Adaptive Iteration Adversarial Training for Robust Pulmonary Nodule DetectionabstractLung cancer is one of the leading causes of death worldwide. Early diagnosis through cancer screening can significantly improve lung cancer patients’ survival. Recently, deep learning based diagnostic systems for nodule detection have shown great potential in assisting radiologists to screen cancer more efficiently. However, studies have found that deep learning models lack robustness against imperceptible crafted adversarial attacks and few studied improving the robustness of pulmonary nodule detection. Therefore, making pulmonary nodule detection models robust remains challenges. Moreover, traditional adversarial training methods either hurt the natural generalization or need expensive computational cost. To address these challenges, here we propose a novel adversarial training method called, Adaptive Iteration Adversarial Training (AIAT). AIAT generates adversarial samples by adding adversarial noise with an adaptive iteration strategy, so that it can stably and fast train models with improving robustness. Extensive experiments on the LUNA 16 dataset show that AIAT improves robustness for pulmonary nodule detection without compromising the natural generalization, and largely reduces training time. Guoxing Yang, Xiaohong Liu 0007, Jianyu Shi, Xianchao Zhang 0002 |
BIBM | 3 |
| 2022 | Directed graph attention networks for predicting asymmetric drug-drug interactionsabstractIt is tough to detect unexpected drug-drug interactions (DDIs) in poly-drug treatments because of high costs and clinical limitations. Computational approaches, such as deep learning-based approaches, are promising to screen potential DDIs among numerous drug pairs. Nevertheless, existing approaches neglect the asymmetric roles of two drugs in interaction. Such an asymmetry is crucial to poly-drug treatments since it determines drug priority in co-prescription. This paper designs a directed graph attention network (DGAT-DDI) to predict asymmetric DDIs. First, its encoder learns the embeddings of the source role, the target role and the self-roles of a drug. The source role embedding represents how a drug influences other drugs in DDIs. In contrast, the target role embedding represents how it is influenced by others. The self-role embedding encodes its chemical structure in a role-specific manner. Besides, two role-specific items, aggressiveness and impressionability, capture how the number of interaction partners of a drug affects its interaction tendency. Furthermore, the predictor of DGAT-DDI discriminates direction-specific interactions by the combination between two proximities and the above two role-specific items. The proximities measure the similarity between source/target embeddings and self-role embeddings. In the designated experiments, the comparison with state-of-the-art deep learning models demonstrates the superiority of DGAT-DDI across a direction-specific predicting task and a direction-blinded predicting task. An ablation study reveals how well each component of DGAT-DDI contributes to its ability. Moreover, a case study of finding novel DDIs confirms its practical ability, where 7 out of the top 10 candidates are validated in DrugBank. Yi-Yang Feng, Hui Yu 0011, Yue-Hua Feng, Jianyu Shi |
Briefings Bioinform. | 4 |
| 2022 | Drug-drug interaction prediction with learnable size-adaptive molecular substructuresabstractDrug-drug interactions (DDIs) are interactions with adverse effects on the body, manifested when two or more incompatible drugs are taken together. They can be caused by the chemical compositions of the drugs involved. We introduce gated message passing neural network (GMPNN), a message passing neural network which learns chemical substructures with different sizes and shapes from the molecular graph representations of drugs for DDI prediction between a pair of drugs. In GMPNN, edges are considered as gates which control the flow of message passing, and therefore delimiting the substructures in a learnable way. The final DDI prediction between a drug pair is based on the interactions between pairs of their (learned) substructures, each pair weighted by a relevance score to the final DDI prediction output. Our proposed method GMPNN-CS (i.e. GMPNN + prediction module) is evaluated on two real-world datasets, with competitive results on one, and improved performance on the other compared with previous methods. Source code is freely available at https://github.com/kanz76/GMPNN-CS. Arnold K. Nyamabo, Hui Yu 0011, Zun Liu, Jianyu Shi |
Briefings Bioinform. | 4 |
| 2022 | STNN-DDI: a Substructure-aware Tensor Neural Network to predict Drug-Drug InteractionsabstractComputational prediction of multiple-type drug-drug interaction (DDI) helps reduce unexpected side effects in poly-drug treatments. Although existing computational approaches achieve inspiring results, they ignore to study which local structures of drugs cause DDIs, and their interpretability is still weak. In this paper, by supposing that the interactions between two given drugs are caused by their local chemical structures (substructures) and their DDI types are determined by the linkages between different substructure sets, we design a novel Substructure-aware Tensor Neural Network model for DDI prediction (STNN-DDI). The proposed model learns a 3-D tensor of $\langle $ substructure, substructure, interaction type $\rangle $ triplets, which characterizes a substructure-substructure interaction (SSI) space. According to a list of predefined substructures with specific chemical meanings, the mapping of drugs into this SSI space enables STNN-DDI to perform the multiple-type DDI prediction in both transductive and inductive scenarios in a unified form with an explicable manner. The comparison with deep learning-based state-of-the-art baselines demonstrates the superiority of STNN-DDI with the significant improvement of AUC, AUPR, Accuracy and Precision. More importantly, case studies illustrate its interpretability by both revealing an important substructure pair across drugs regarding a DDI type of interest and uncovering interaction type-specific substructure pairs in a given DDI. In summary, STNN-DDI provides an effective approach to predicting DDIs as well as explaining the interaction mechanisms among drugs. Source code is freely available at https://github.com/zsy-9/STNN-DDI. Hui Yu 0011, Jianyu Shi |
Briefings Bioinform. | 3 |
| 2022 | MLGL-MP: a Multi-Label Graph Learning framework enhanced by pathway interdependence for Metabolic Pathway predictionabstractMOTIVATION: During lead compound optimization, it is crucial to identify pathways where a drug-like compound is metabolized. Recently, machine learning-based methods have achieved inspiring progress to predict potential metabolic pathways for drug-like compounds. However, they neglect the knowledge that metabolic pathways are dependent on each other. Moreover, they are inadequate to elucidate why compounds participate in specific pathways. RESULTS: To address these issues, we propose a novel Multi-Label Graph Learning framework of Metabolic Pathway prediction boosted by pathway interdependence, called MLGL-MP, which contains a compound encoder, a pathway encoder and a multi-label predictor. The compound encoder learns compound embedding representations by graph neural networks. After constructing a pathway dependence graph by re-trained word embeddings and pathway co-occurrences, the pathway encoder learns pathway embeddings by graph convolutional networks. Moreover, after adapting the compound embedding space into the pathway embedding space, the multi-label predictor measures the proximity of two spaces to discriminate which pathways a compound participates in. The comparison with state-of-the-art methods on KEGG pathways demonstrates the superiority of our MLGL-MP. Also, the ablation studies reveal how its three components contribute to the model, including the pathway dependence, the adapter between compound embeddings and pathway embeddings, as well as the pre-training strategy. Furthermore, a case study illustrates the interpretability of MLGL-MP by indicating crucial substructures in a compound, which are significantly associated with the attending metabolic pathways. It is anticipated that this work can boost metabolic pathway predictions in drug discovery. AVAILABILITY AND IMPLEMENTATION: The code and data underlying this article are freely available at https://github.com/dubingxue/MLGL-MP. Bing-Xue Du, Peng-Cheng Zhao, Bei Zhu, Siu-Ming Yiu, Arnold K. Nyamabo, Hui Yu 0011, Jianyu Shi |
Bioinform. | 7 |
| 2022 | Predict multi-type drug-drug interactions in cold start scenarioabstractBACKGROUND: Prediction of drug-drug interactions (DDIs) can reveal potential adverse pharmacological reactions between drugs in co-medication. Various methods have been proposed to address this issue. Most of them focus on the traditional link prediction between drugs, however, they ignore the cold-start scenario, which requires the prediction between known drugs having approved DDIs and new drugs having no DDI. Moreover, they're restricted to infer whether DDIs occur, but are not able to deduce diverse DDI types, which are important in clinics. RESULTS: In this paper, we propose a cold start prediction model for both single-type and multiple-type drug-drug interactions, referred to as CSMDDI. CSMDDI predict not only whether two drugs trigger pharmacological reactions but also what reaction types they induce in the cold start scenario. We implement several embedding methods in CSMDDI, including SVD, GAE, TransE, RESCAL and compare it with the state-of-the-art multi-type DDI prediction method DeepDDI and DDIMDL to verify the performance. The comparison shows that CSMDDI achieves a good performance of DDI prediction in the case of both the occurrence prediction and the multi-type reaction prediction in cold start scenario. CONCLUSIONS: Our approach is able to predict not only conventional binary DDIs but also what reaction types they induce in the cold start scenario. More importantly, it learns a mapping function who can bridge the drugs attributes to their network embeddings to predict DDIs. The main contribution of CSMDDI contains the development of a generalized framework to predict the single-type and multi-type of DDIs in the cold start scenario, as well as the implementations of several embedding models for both single-type and multi-type of DDIs. The dataset and source code can be accessed at https://github.com/itsosy/csmddi . Zun Liu, Xing-Nan Wang, Hui Yu 0011, Jianyu Shi, Wenmin Dong |
BMC Bioinform. | 4 |
| 2022 | RANEDDI: Relation-aware network embedding for drug-drug interaction prediction
Hui Yu 0011, Wenmin Dong, Jianyu Shi |
Inf. Sci. | 3 |
| 2021 | SSI-DDI: substructure-substructure interactions for drug-drug interaction predictionabstractA major concern with co-administration of different drugs is the high risk of interference between their mechanisms of action, known as adverse drug-drug interactions (DDIs), which can cause serious injuries to the organism. Although several computational methods have been proposed for identifying potential adverse DDIs, there is still room for improvement. Existing methods are not explicitly based on the knowledge that DDIs are fundamentally caused by chemical substructure interactions instead of whole drugs' chemical structures. Furthermore, most of existing methods rely on manually engineered molecular representation, which is limited by the domain expert's knowledge.We propose substructure-substructure interaction-drug-drug interaction (SSI-DDI), a deep learning framework, which operates directly on the raw molecular graph representations of drugs for richer feature extraction; and, most importantly, breaks the DDI prediction task between two drugs down to identifying pairwise interactions between their respective substructures. SSI-DDI is evaluated on real-world data and improves DDI prediction performance compared to state-of-the-art methods. Source code is freely available at https://github.com/kanz76/SSI-DDI. Arnold K. Nyamabo, Hui Yu 0011, Jianyu Shi |
Briefings Bioinform. | 3 |
| 2020 | DPDDI: a deep predictor for drug-drug interactionsabstractBACKGROUND: The treatment of complex diseases by taking multiple drugs becomes increasingly popular. However, drug-drug interactions (DDIs) may give rise to the risk of unanticipated adverse effects and even unknown toxicity. DDI detection in the wet lab is expensive and time-consuming. Thus, it is highly desired to develop the computational methods for predicting DDIs. Generally, most of the existing computational methods predict DDIs by extracting the chemical and biological features of drugs from diverse drug-related properties, however some drug properties are costly to obtain and not available in many cases. RESULTS: In this work, we presented a novel method (namely DPDDI) to predict DDIs by extracting the network structure features of drugs from DDI network with graph convolution network (GCN), and the deep neural network (DNN) model as a predictor. GCN learns the low-dimensional feature representations of drugs by capturing the topological relationship of drugs in DDI network. DNN predictor concatenates the latent feature vectors of any two drugs as the feature vector of the corresponding drug pairs to train a DNN for predicting the potential drug-drug interactions. Experiment results show that, the newly proposed DPDDI method outperforms four other state-of-the-art methods; the GCN-derived latent features include more DDI information than other features derived from chemical, biological or anatomical properties of drugs; and the concatenation feature aggregation operator is better than two other feature aggregation operators (i.e., inner product and summation). The results in case studies confirm that DPDDI achieves reasonable performance in predicting new DDIs. CONCLUSION: We proposed an effective and robust method DPDDI to predict the potential DDIs by utilizing the DDI network information without considering the drug properties (i.e., drug chemical and biological properties). The method should also be useful in other DDI-related scenarios, such as the detection of unexpected side effects, and the guidance of drug combination. Yue-Hua Feng, Shaowu Zhang 0001, Jianyu Shi |
BMC Bioinform. | 3 |
| 2018 | TMFUF: a triple matrix factorization-based unified framework for predicting comprehensive drug-drug interactions of new drugsabstractBACKGROUND: A significant number of adverse drug reactions is caused by unexpected Drug-drug interactions (DDIs). The identification of DDIs becomes crucial before the co-prescription of multiple drugs is made. Such a task in clinics or in drug discovery usually requires high costs and numerous limitations, while computational approaches are able to predict potential DDIs effectively by utilizing diverse drug attributes (e.g. side effects). Nevertheless, they're incapable when required to predict enhancive and degressive DDIs, which change increasingly and decreasingly the pharmacological behavior of interacting drugs respectively. The pharmacological change of DDIs is one of the most important factors when making a multi-drug prescription. RESULTS: In this work, we design a Triple Matrix Factorization-based Unified Framework (TMFUF) to address the above issue. By leveraging a group of side effect entries of drugs, TMFUF achieves the inspiring result (AUC = 0.842 and AUPR = 0.526) in the case of conventional DDI prediction under the traditional screening task. In the comparison with two state-of-the-art approaches, TMFUF demonstrates it superiority by ~ 7% and ~ 20% improvement in terms of AUC and AUPR respectively. More importantly, TMFUF shows its ability in the comprehensive DDI prediction under different screening tasks. Finally, a utilization TMFUF reveals the significant pairs of side effects, which contribute to form enhancive and degressive DDIs, for further clinical validation. CONCLUSIONS: The proposed TMFUF is first capable to predict both conventional binary DDIs and comprehensive DDIs such that it captures the pharmacological changes caused by DDIs. Furthermore, it provides a unified solution of DDI prediction for two screening scenarios, which involves newly given drugs having no prior interaction. Another advantage is its ability to indicate how significantly the pairs of drug features contribute to form DDIs. Jianyu Shi, Yanning Zhang 0001, Siu-Ming Yiu |
BMC Bioinform. | 1 |
| 2018 | BMCMDA: a novel model for predicting human microbe-disease associations via binary matrix completionabstractBACKGROUND: Human Microbiome Project reveals the significant mutualistic influence between human body and microbes living in it. Such an influence lead to an interesting phenomenon that many noninfectious diseases are closely associated with diverse microbes. However, the identification of microbe-noninfectious disease associations (MDAs) is still a challenging task, because of both the high cost and the limitation of microbe cultivation. Thus, there is a need to develop fast approaches to screen potential MDAs. The growing number of validated MDAs enables us to meet the demand in a new insight. Computational approaches, especially machine learning, are promising to predict MDA candidates rapidly among a large number of microbe-disease pairs with the advantage of no limitation on microbe cultivation. Nevertheless, a few computational efforts at predicting MDAs are made so far. RESULTS: In this paper, grouping a set of MDAs into a binary MDA matrix, we propose a novel predictive approach (BMCMDA) based on Binary Matrix Completion to predict potential MDAs. The proposed BMCMDA assumes that the incomplete observed MDA matrix is the summation of a latent parameterizing matrix and a noising matrix. It also assumes that the independently occurring subscripts of observed entries in the MDA matrix follows a binomial model. Adopting a standard mean-zero Gaussian distribution for the nosing matrix, we model the relationship between the parameterizing matrix and the MDA matrix under the observed microbe-disease pairs as a probit regression. With the recovered parameterizing matrix, BMCMDA deduces how likely a microbe would be associated with a particular disease. In the experiment under leave-one-out cross-validation, it exhibits the inspiring performance (AUC = 0.906, AUPR =0.526) and demonstrates its superiority by ~ 7% and ~ 5% improvements in terms of AUC and AUPR respectively in the comparison with the pioneering approach KATZHMDA. CONCLUSIONS: Our BMCMDA provides an effective approach for predicting MDAs and can be also extended to other similar predicting tasks of binary relationship (e.g. protein-protein interaction, drug-target interaction). Jianyu Shi, Yanning Zhang 0001, Jiang-Bo Cao, Siu-Ming Yiu |
BMC Bioinform. | 1 |
| 2017 | Predicting combinative drug pairs towards realistic screening via integrating heterogeneous featuresabstractBACKGROUND: Drug Combination is one of the effective approaches for treating complex diseases. However, determining combinative drug pairs in clinical trials is still costly. Thus, computational approaches are used to identify potential drug pairs in advance. Existing computational approaches have the following shortcomings: (i) the lack of an effective integration of heterogeneous features leads to a time-consuming training and even results in an over-fitted classifier; and (ii) the narrow consideration of predicting potential drug combinations only among known drugs having known combinations cannot meet the demand of realistic screenings, which pay more attention to potential combinative pairs among newly-coming drugs that have no approved combination with other drugs at all. RESULTS: In this paper, to tackle the above two problems, we propose a novel drug-driven approach for predicting potential combinative pairs on a large scale. We define four new features based on heterogeneous data and design an efficient fusion scheme to integrate these feature. Moreover importantly, we elaborate appropriate cross-validations towards realistic screening scenarios of drug combinations involving both known drugs and new drugs. In addition, we perform an extra investigation to show how each kind of heterogeneous features is related to combinative drug pairs. The investigation inspires the design of our approach. Experiments on real data demonstrate the effectiveness of our fusion scheme for integrating heterogeneous features and its predicting power in three scenarios of realistic screening. In terms of both AUC and AUPR, the prediction among known drugs achieves 0.954 and 0.821, that between known drugs and new drugs achieves 0.909 and 0.635, and that among new drugs achieves 0.809 and 0.592 respectively. CONCLUSIONS: Our approach provides not only an effective tool to integrate heterogeneous features but also the first tool to predict potential combinative pairs among new drugs. Jianyu Shi, Siu-Ming Yiu |
BMC Bioinform. | 1 |
| 2016 | The need of accelerators in analyzing biological networksabstractSummary form only given. As the development of high-throughput techniques in both biology and its related disciplines (chemistry or medicine), the huge number of biological entries are available. The discovered relationship between them (e.g. interactions or associations) reveals important biological facts, which are never found in individual-based biological experiments. A biological network is an appropriate tool to systematically analyze and uncover such facts. The relationship between biological molecules is usually modeled as a monopartite network, such as protein-protein interactions, while that between biological molecules and other objects is modeled as a bipartite network, such as chemical compound-protein interactions, gene-disease associations and ncRNA-disease associations. A biological network may contain a large number of nodes, of which each owns many heterogeneous attributes, including binary, real-valued and semantic forms. Current algorithms for systematical analysis based on large-scale biological networks have always a need of either using much memory or taking much time, because of their high computational complexity. Take the compound-protein interaction network as an example. Over 90 million compounds are available in PubChem and each compound is characterized as a high-dimensional vector (e.g. 881-d PubChem fingerprint or 4860-d Klekota-Roth fingerprint). Meanwhile, a protein can be characterized as a 20K-demensional vector if the K-mer descriptor is adopted. However, involving intensive matrix manipulation (e.g. matrix factorization, inverse and tensor product), current algorithms cannot be directly applied to predict compound-protein interactions on a large scale. For example, having the complexity O(n3), singular value decomposition (SVD) runs for a 6,000□6,000 matrix in MATLAB 2013b (64 bits) under Windows 7(64bits) with Intel Corei7-4700MQ (2.40G) and GeForce GTX 765M. SVD spends 81.9, 77.9, and 51.4 seconds when using CPU only, CPU with four workers and CPU plus GPU respectively. Consequently, there is an urge need to turn them into accelerator-enabled parallel algorithms or develop novel accelerators to speed up the knowledge-mining in biological networks. Jianyu Shi |
BIBM | 1 |
| 2016 | LCM-DS: A novel approach of predicting drug-drug interactions for new drugs via Dempster-Shafer theory of evidenceabstractThere is an urgent need to discover or predict DDIs, which would cause serious adverse drug reactions. However, preclinical detection of DDIs bear high cost. Similarity-based computational approaches can be the assistance of experimental approaches. Utilizing pre-market drug similarities, they are able to predict DDIs on a large scale. However, they neglect the topological structure among DDIs and non-DDIs and have a burden of slow training and much memory. Or, they bear the bias that the pairs between a newly-given drug and the drugs having many DDIs tend to obtain high ranks. More importantly, they lack an effective combination of multiple predictions. To address these issues, we develop a local classification-based model (LCM), which has the advantages of faster training, less memory requirement as well as no that bias. We further design a novel supervised algorithm of fusion based on Dempster-Shafer (DS) theory of evidence for combine multiple predictions. Finally, the experiments demonstrate that our LCM-DS is significantly superior to three state-of-the-art approaches and outperforms both individual LCMs and classical fusion algorithms. Jianyu Shi, Xuequn Shang 0001, Siu-Ming Yiu |
BIBM | 1 |
| 2016 | Predicting existing targets for new drugs base on strategies for missing interactionsabstractBACKGROUND: There has been paid more and more attention to supervised classification models in the area of predicting drug-target interactions (DTIs). However, in terms of classification, unavoidable missing DTIs in data would cause three issues which have not yet been addressed appropriately by former approaches. Directly labeled as negatives (non-DTIs), missing DTIs increase the confusion of positives (DTIs) and negatives, aggravate the imbalance between few positives and many negatives, and are usually discriminated as highly-scored false positives, which influence the existing measures sharply. RESULTS: Under the framework of local classification model (LCM), this work focuses on the scenario of predicting how possibly a new drug interacts with known targets. To address the first two issues, two strategies, Spy and Super-target, are introduced accordingly and further integrated to form a two-layer LCM. In the bottom layer, Spy-based local classifiers for protein targets are built by positives, as well as reliable negatives identified among unlabeled drug-target pairs. In the top layer, regular local classifiers specific to super-targets are built with more positives generated by grouping similar targets and their interactions. Furthermore, to handle the third issue, an additional performance measure, Coverage, is presented for assessing DTI prediction. The experiments based on benchmark datasets are finally performed under five-fold cross validation of drugs to evaluate this approach. The main findings are concluded as follows. (1) Both two individual strategies and their combination are effective to missing DTIs, and the combination wins the best. (2) Having the advantages of less confusing decision boundary at the bottom layer and less biased decision boundary at the top layer, our two-layer LCM outperforms two former approaches. (3) Coverage is more robust to missing interactions than other measures and is able to evaluate how far one needs to go down the list of targets to cover all the proper targets of a drug. CONCLUSIONS: Proposing two strategies and one performance measure, this work has addressed the issues derived from missing interactions, which cause confusing and biased decision boundaries in classifiers, as well as the inappropriate measure of predicting performance, in the scenario of predicting interactions between new drugs and known targets. Jianyu Shi, Hui-Meng Lu |
BMC Bioinform. | 1 |
| 2015 | SRP: A concise non-parametric similarity-rank-based model for predicting drug-target interactionsabstractThe identification of drug-target interactions in web lab is costly and time-consuming. Computational approaches become important to help identifying potential candidates for laboratory experiments. However, they usually involve solving optimization problems or assuming statistical distribution based on prior knowledge, and may require estimating tunable parameters. This paper is motivated by the concepts behind “follow-on” drugs. They are the drugs developed by drug companies to substitute the pioneering drug which was firstly discovered and patented for a specific target and determined a new therapeutic class. There are three observations from “follow-on” drugs. The first observation has been used by many existing methods: drugs interacting with a common target usually have higher similar scores (e.g. the similarity score in terms of chemical structure). The second one is that a drug candidate for a specific target gains more attention if it is more similar to those drugs interacting with the target than other known drugs, even though the similarity score is low. Lastly, people intuitively tend to design a “follow-on” drug for the targets already having more drugs because of less cost and less risk. In our approach, the above observations are translated into more evidences for predicted drug-target interaction. Designing an interaction tendency index to characterize these observations, we propose the similarity-rank-based predictor (SRP). Unlike other models, SRP is a non-parametric model and requires neither solving an optimization problem nor prior statistical knowledge. Based on real benchmark datasets, we show that our model is able to achieve higher accuracy than the two most recent models and our approach is able to cope with two real predicting scenario of missing interactions. Jianyu Shi, Siu-Ming Yiu |
BIBM | 1 |
| 2014 | Predicting drug-target interaction for new drugs using enhanced similarity measures and super-target clustering1abstractPredicting drug-target interaction using computational approaches is an important step in drug discovery and repositioning. To predict whether there will be an interaction between a drug and a target, most existing methods identify similar drugs and targets in the database. The prediction is then made based on the known interactions of these drugs and targets. This idea is promising. However, there are two shortcomings that have not yet been addressed appropriately. Firstly, most of the methods only use 2D chemical structures and protein sequences to measure the similarity of drugs and targets respectively. However, this information may not fully capture the characteristics determining whether a drug will interact with a target. Secondly, there are very few known interactions, i.e. many interactions are “missing” in the database. Existing approaches are biased towards known interactions and have no good solutions to handle possibly missing interactions which affect the accuracy of the prediction. In this paper, we enhance the similarity measures to include non-structural (and non-sequence-based) information and introduce the concept of a “super-target” to handle the problem of possibly missing interactions. Based on evaluations on real data, we show that our similarity measure is better than the existing measures and our approach is able to achieve higher accuracy than the two best existing algorithms, WNN-GIP and KBMF2K. Jianyu Shi, Siu-Ming Yiu, Henry C. M. Leung, Francis Y. L. Chin |
BIBM | 1 |
| 2008 | A new system for computer-aided intraoperative simulation and postoperative facial appearance prediction of orthognathic surgeryabstractThis paper presents a new system for computer-aided orthognathic surgery which aims at correcting dento-facial defects and keeping patient away from much radiation and high cost of CT or MRI imaging. The system allows simulating the surgical correction on virtual 3-D model of the skull reconstructed with the acquisition of radiographs of three orthogonal views and facial mesh data generated by 3-D laser scanner and the use of build-in skull template mesh. Surgery simulation includes maxilla Le Fort I and mandible bilateral SSRO osteotomies by interactive cutting of jaw bone. With the help of a special registration procedure, we are able to align the reconstructed skull model and a facial skin model. Upon fulfilment of the registration and virtual osteotomies, the system provides further assistance with predicting postoperative appearance of patient. Finally, we apply the system on one patient in clinical practice and achieve satisfied result. Yanning Zhang 0001, Pei-Fang Zhai, Jianyu Shi, Jiangbin Zheng 0001, Jiang-Bo Li |
MMSP | 3 |