Seyedeh Shaghayegh Sadeghi

dblp:290/6020 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0003-2346-1558ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2024 Can large language models understand molecules?
abstract
PURPOSE: Large Language Models (LLMs) like Generative Pre-trained Transformer (GPT) from OpenAI and LLaMA (Large Language Model Meta AI) from Meta AI are increasingly recognized for their potential in the field of cheminformatics, particularly in understanding Simplified Molecular Input Line Entry System (SMILES), a standard method for representing chemical structures. These LLMs also have the ability to decode SMILES strings into vector representations. METHOD: We investigate the performance of GPT and LLaMA compared to pre-trained models on SMILES in embedding SMILES strings on downstream tasks, focusing on two key applications: molecular property prediction and drug-drug interaction prediction. RESULTS: We find that SMILES embeddings generated using LLaMA outperform those from GPT in both molecular property and DDI prediction tasks. Notably, LLaMA-based SMILES embeddings show results comparable to pre-trained models on SMILES in molecular prediction tasks and outperform the pre-trained models for the DDI prediction tasks. CONCLUSION: The performance of LLMs in generating SMILES embeddings shows great potential for further investigation of these models for molecular embedding. We hope our study bridges the gap between LLMs and molecular embedding, motivating additional research into the potential of LLMs in the molecular representation field. GitHub: https://github.com/sshaghayeghs/LLaMA-VS-GPT .
Seyedeh Shaghayegh Sadeghi, Alan Bui, Ali Forooghi, Jianguo Lu, Alioune Ngom
BMC Bioinform.1
2022 DeePSLiM: A Deep Learning Approach to Identify Predictive Short-linear Motifs for Protein Sequence Classification
abstract
SLiMs (Short Linear Motifs) are patterns of three to 20 amino acids within proteins that are sufficient to fulfill certain functions. SLiMs play a critical role in many biological processes. Hence, with the increasing quantity of biological data, it is important to develop algorithms that can quickly find patterns in large databases of DNA, RNA and protein sequences. Previous research has been very successful at applying deep learning methods to the problems of motif detection as well as classification of biological sequences. There are, however, limitations to these approaches. Most are limited to finding motifs of a single length. In addition, most research has focused on DNA and RNA, both of which use a four-letter alphabet. A few of these have attempted to apply deep learning methods on the larger, twenty letter, alphabet of proteins. We present an enhanced deep learning model, called DeePSLiM, capable of detecting predictive SLiMs in protein sequences. The model is a shallow network that can be trained quickly on large amounts of data. The SLiMs are predictive because they can be used to classify the sequences into their respective families. In this study, first, we propose a new deep learning approach for finding predictive SLiMs in protein sequences. Then, we use these predictive SLiMs for the classification task of protein sequences to evaluate our proposed method. The model was able to reach scores of 94.5% on accuracy, precision, recall, F1-Score and Matthews-correlation coefficient, as well as 99.9% area under the receiver operator characteristic curve (AUROC). Availability: The source code, sample data, and supplementary material are available via a Github project at https://github.com/sshaghayeghs/DeePSLiM.
Alexandru Filip, Seyedeh Shaghayegh Sadeghi, Alioune Ngom, Luis Rueda 0001
CIBCB2
2022 DDIPred: Graph Convolutional Network-Based Drug-drug Interactions Prediction Using Drug Chemical Structure Embedding
abstract
A drug-drug interaction (DDI) describes a circumstance in which drugs affect the activity of each other. Drugs may interact with each other to cause side effects that are unexpected or more severe than anticipated. Drugs may also interact and oppose the results of one another, leading to one (or both) medications not having their intended effect. Most drug interactions are negligible, but some can be significantly harmful if not discovered and appropriately overseen. DDI data can be helpful for other drug-related research topics such as drug repurposing and drug-target interaction, which leads to improving the drug development process. This paper presents a new method for DDI prediction named DDIPred. It is based on drug chemical structure embedding and graph convolutional networks for predicting new DDIs. DDIPred First extracts a representation of Simplified Molecular Input Line Entry System (SMILES) strings using SELF-referencIng Embedded Strings (SELFIES) and Doc2Vec. Then the representation, along with the DDI network structure, is used to predict new DDIs. Our method achieved acceptable performance when tested on the BIOSNAP DDI network dataset while outperforming other existing methods. Availability: The source code and sample data are available via a Github project at https://github.com/sshaghayeghs/DDIPred.
Seyedeh Shaghayegh Sadeghi, Alioune Ngom
CIBCB1
2022 A network-based drug repurposing method via non-negative matrix factorization
abstract
MOTIVATION: Drug repurposing is a potential alternative to the traditional drug discovery process. Drug repurposing can be formulated as a recommender system that recommends novel indications for available drugs based on known drug-disease associations. This article presents a method based on non-negative matrix factorization (NMF-DR) to predict the drug-related candidate disease indications. This work proposes a recommender system-based method for drug repurposing to predict novel drug indications by integrating drug and diseases related data sources. For this purpose, this framework first integrates two types of disease similarities, the associations between drugs and diseases, and the various similarities between drugs from different views to make a heterogeneous drug-disease interaction network. Then, an improved non-negative matrix factorization-based method is proposed to complete the drug-disease adjacency matrix with predicted scores for unknown drug-disease pairs. RESULTS: The comprehensive experimental results show that NMF-DR achieves superior prediction performance when compared with several existing methods for drug-disease association prediction. AVAILABILITY AND IMPLEMENTATION: The program is available at https://github.com/sshaghayeghs/NMF-DR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Seyedeh Shaghayegh Sadeghi, Jianguo Lu, Alioune Ngom
Bioinform.1
2021 An Analytical Review of Computational Drug Repurposing
abstract
Drug repurposing is a vital function in pharmaceutical fields and has gained popularity in recent years in both the pharmaceutical industry and research community. It refers to the process of discovering new uses and indications for existing or failed drugs. It is cost-effective and reliable in contrast to experimental drug discovery, which is a costly, time-consuming, and risky process and limited to a relatively small number of targets. Accordingly, a plethora of computational methodologies have been propounded to repurpose drugs on a large scale by utilizing available high throughput data. The available literature, however, lacks a contemporary and comprehensive analysis of the current computational drug repurposing methodologies. In this paper, we presented a systematic analysis of computational drug repurposing which consists of three main sections: Initially, we categorize the computational drug repurposing methods based on their technical approach and artificial intelligence perspective and discuss the strengths and weaknesses of various methods. Secondly, some general criteria are recommended to analyze our proposed categorization. In the third and final section, a qualitative comparison is made between each approach which is a guide to understanding their preference to one another. Further, this systematic analysis can help in the efficient selection and improvement of drug repurposing techniques based on the nature of computational methods implemented on biological resources.
Seyedeh Shaghayegh Sadeghi, Mohammad Reza Keyvanpour
IEEE ACM Trans. Comput. Biol. Bioinform.1