VLDB 2026 Research / reviewers in the wild / expert
Joel Arrais
dblp:02/5399 · also Joel P. Arrais, Joel Perdiz Arrais
· DBLP profile ↗
33ranked-venue papers
1as first author
19since 2021 · last 2025
0000-0003-4937-2334ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | xLSTM-AmPEP60: An xLSTM Model for the Minimum Inhibitory Concentrations Prediction of Antimicrobial PeptidesabstractAntimicrobial peptides (AMPs) represent a promising alternative to combat escalating bacterial drug resistance. While pretrained protein language models (PLMs) have demonstrated success in identifying AMPs from sequence data, overfitting remains a major challenge. This issue is mainly caused by the high complexity of PLMs and limited availability of data. To address this, we proposed xLSTM-AmPEP60, a deep learning method based on an extended Long Short-Term Memory (xLSTM) architecture. This model used lightweight parameters to derive embedding features from peptide sequences, enabling the quantitative prediction of minimum inhibitory concentrations (MICs) of the common bacterial species Escherichia coli (E. coli). In five independent experiments, with 10% leave-out sequences as test sets, xLSTM-AmPEP60 outperformed state-of-the-art regression methods, achieving a mean squared error of 0.2349$(\log \mu \mathbf{M})$. These results demonstrate that lightweight xLSTM architectures can accurately assess AMP activity against E. coli. Jianxiu Cai, Xinpo Lou, Jielu Yan, Joel Arrais, Shirley W. I. Siu |
BIBM | 4 |
| 2025 | Rethinking transformers with convolution and graph embeddings for few-shot molecular property discoveryabstractThe prediction of molecular properties is a critical step in drug discovery campaigns. Computational methods such as graph neural networks (GNNs) and Transformers have effectively leveraged the small-range and long-range dependencies in molecules to preserve the local and global patterns for multiple molecular property prediction tasks. However, the dependence of these models on large amounts of experimental data poses a challenge, particularly on smaller biological datasets prevalent across the drug discovery pipeline. This paper introduces FS-GCvTR, a few-shot graph-based convolutional Transformer architecture designed to predict chemical properties with a small amount of labeled compounds. The convolutional Transformer is presented as a crucial component, effectively integrating both local and global dependencies of molecular graph embeddings by propagating a set of convolutional tokens across Transformer attention layers for molecular property prediction. Furthermore, a few-shot meta-learning approach is introduced to iteratively adapt model parameters across multiple few-shot tasks while generalizing to new chemical properties with limited available data. Experiments including few-shot evaluations on multi-property datasets show that the FS-GCvTR model outperformed other few-shot graph-based baselines in specific molecular property prediction tasks. • A few-shot GNN-Transformer, FS-GCvTR is proposed for molecular property discovery. • A Convolutional Transformer learns local and global information in graph embeddings. • A meta-learning approach adapts FS-GCvTR across tasks to predict molecular properties. • Experiments show that FS-GCvTR outperforms standard graph-based methods. Luis H. M. Torres, Joel Arrais, Bernardete Ribeiro |
Pattern Recognit. | 2 |
| 2024 | Predicting drug activity against cancer through genomic profiles and SMILESabstractDue to the constant increase in cancer rates, the disease has become a leading cause of death worldwide, enhancing the need for its detection and treatment. In the era of personalized medicine, the main goal is to incorporate individual variability in order to choose more precisely which therapy and prevention strategies suit each person. However, predicting the sensitivity of tumors to anticancer treatments remains a challenge. In this work, we propose two deep neural network models to predict the impact of anticancer drugs in tumors through the half-maximal inhibitory concentration (IC50). These models join biological and chemical data to apprehend relevant features of the genetic profile and the drug compounds, respectively. In order to predict the drug response in cancer cell lines, this study employed different DL methods, resorting to Recurrent Neural Networks (RNNs) and Convolutional Neural Networks (CNNs). In the first stage, two autoencoders were pre-trained with high-dimensional gene expression and mutation data of tumors. Afterward, this genetic background is transferred to the prediction models that return the IC50 value that portrays the potency of a substance in inhibiting a cancer cell line. When comparing RSEM Expected counts and TPM as methods for displaying gene expression data, RSEM has been shown to perform better in deep models and CNNs model can obtain better insight in these types of data. Moreover, the obtained results reflect the effectiveness of the extracted deep representations in the prediction of the IC50 value that portrays the potency of a substance in inhibiting a tumor, achieving a performance of a mean squared error of 1.06 and surpassing previous state-of-the-art models. Maryam Abbasi, Filipa G. Carvalho, Bernardete Ribeiro, Joel Arrais |
Artif. Intell. Medicine | 4 |
| 2024 | TAG-DTA: Binding-region-guided strategy to predict drug-target affinity using transformersabstractThe proper assessment of target-specific compound selectivity is paramount in the drug discovery context, promoting the identification of drug-target interactions (DTIs) and the discovery of potential leads. On that account, the accurate prediction of an unbiased drug-target binding affinity (DTA) metric is pivotal to understanding the binding process. Most in silico computational approaches, however, neglect the inter-dependency of the proteomics, chemical, and pharmacological spaces and the explainability during the model construction. Furthermore, these methods have yet to actively include information associated with binding pockets during the learning process, which is essential to DTA prediction performance and model explainability. In this study, we propose an end-to-end binding-region-guided Transformer-based architecture that simultaneously predicts the 1D binding pocket and the binding affinity of DTI pairs, where the prediction of the 1D binding pocket guides and conditions the prediction of DTA. This architecture uses 1D raw sequential and structural data to represent the proteins and compounds, respectively, and combines multiple Transformer-Encoder blocks to capture and learn the proteomics, chemical, and pharmacological contexts. The predicted 1D binding pocket conditions the attention mechanism of the Transformer-Encoder used to learn the pharmacological space in order to model the inter-dependency amongst binding-related positions. The results show that the proposed architecture, TAG-DTA, achieved the best performance in DTA prediction compared to state-of-the-art benchmarks, including in unknown subsets of the proteomics and chemical representation spaces. Moreover, the 1D binding pocket prediction increases the discriminative power and robustness of the aggregate representation of the pharmacological space and improves the DTA prediction performance. Overall, this research study validates the applicability of an end-to-end Transformer-based architecture in the context of drug discovery, and that combining computationally different yet contextually related tasks is critical to new findings in the DTI domain. Additionally, it shows that TAG-DTA is capable of providing increasing DTI and prediction understanding due to the nature of the attention blocks and prediction of the 1D binding pocket. The data and source code used in this study are available at: https://github.com/larngroup/TAG-DTA. Nelson R. C. Monteiro, José Luís Oliveira, Joel Arrais |
Expert Syst. Appl. | 3 |
| 2023 | Convolutional Transformer via Graph Embeddings for Few-shot Toxicity and Side Effect PredictionabstractThe prediction of chemical toxicity and adverse side effects is a crucial task in drug discovery.Graph neural networks (GNNs) have accelerated the discovery of compounds with improved molecular profiles for effective drug development.Recently, Transformer networks have also managed to capture the long-range dependence in molecules to preserve the global aspects of molecular embeddings for molecular property prediction.In this paper, we propose a few-shot GNN-Transformer, FS-GNNCvTR to face the challenge of low-data toxicity and side effect prediction.Specifically, we introduce a convolutional Transformer to model the local spatial context of molecular graph embeddings while preserving the global information of deep representations.Furthermore, a two-module meta-learning framework is proposed to iteratively update model parameters across fewshot tasks with limited available data.Experiments on small-sized biological datasets for toxicity and side effect prediction, Tox21 and SIDER, demonstrate a superior performance of FS-GNNCvTR compared to standard graph-based methods.The code and data underlying this article are available in the repository, https://github.com/larngroup/FS-GNNCvTR. Luis H. M. Torres, Bernardete Ribeiro, Joel Arrais |
ESANN | 3 |
| 2023 | On the Quantization of Recurrent Neural Networks for Smiles GenerationabstractThis paper focuses on the effects of applying quantization during training to Recurrent Neural Networks (RNNs) used in Simplified Molecular-Input Line-Entry System (SMILES) generation, a form of line notation for molecular information used in the development of pharmaceutical drugs, from the PubChem database. It offers the flexibility to choose the precision used by the model, by defining the number of bits at each layer. The RNNs are the focus of the current study, by comparing the performance of three of the most used algorithms, Simple RNN, Long Short-Term Memory (LSTM), and Gated Recurrent Unit (GRU). The models were trained on a selection of SMILES. By exploiting the QK-eras library, quantization performance was compared to their floating-point equivalent for several combinations of parameters. The goal of the testing program developed is to generate a large number of novel SMILES, facilitating the process of Drug Discovery which is traditionally long, and thus very expensive and difficult. By understanding how the behavior of quantized networks deviates from the regular model, in relation to the parameters used, we are able to control the process of choosing whether to quantize a model and to which degree it becomes more or less efficient. In this study, we observed good performance even for 4-bit models making use of LSTM and GRU layers, the same way we concluded that Simple RNN quantization does not compensate the effort. Adriano Durao, Joel Arrais, Bernardete Ribeiro, Gabriel Falcao |
ICASSP | 2 |
| 2023 | Enhancing reinforcement learning for de novo molecular design applying self-attention mechanismsabstractThe drug discovery process can be significantly improved by applying deep reinforcement learning (RL) methods that learn to generate compounds with desired pharmacological properties. Nevertheless, RL-based methods typically condense the evaluation of sampled compounds into a single scalar value, making it difficult for the generative agent to learn the optimal policy. This work combines self-attention mechanisms and RL to generate promising molecules. The idea is to evaluate the relative significance of each atom and functional group in their interaction with the target, and to utilize this information for optimizing the Generator. Therefore, the framework for de novo drug design is composed of a Generator that samples new compounds combined with a Transformer-encoder and a biological affinity Predictor that evaluate the generated structures. Moreover, it takes the advantage of the knowledge encapsulated in the Transformer's attention weights to evaluate each token individually. We compared the performance of two output prediction strategies for the Transformer: standard and masked language model (MLM). The results show that the MLM Transformer is more effective in optimizing the Generator compared with the state-of-the-art works. Additionally, the evaluation models identified the most important regions of each molecule for the biological interaction with the target. As a case study, we generated synthesizable hit compounds that can be putative inhibitors of the enzyme ubiquitin-specific protein 7 (USP7). Tiago Pereira 0001, Maryam Abbasi, Joel Arrais |
Briefings Bioinform. | 3 |
| 2023 | Few-shot learning with transformers via graph embeddings for molecular property predictionabstractMolecular property prediction is an essential task in drug discovery. Recently, deep neural networks have accelerated the discovery of compounds with improved molecular profiles for effective drug development. In particular, graph neural networks (GNNs) have played a pivotal role in identifying promising drug candidates with desirable molecular properties. However, it is common for only a few molecules to share the same set of properties, which presents a low-data problem unanswered by regular machine learning (ML) approaches. Transformer networks have also emerged as a promising solution to model the long-range dependence in molecular embeddings and achieve encouraging results across a wide range of molecular property prediction tasks. Nonetheless, these methods still require a large number of data points per task to achieve acceptable performance. In this study, we propose a few-shot GNN-Transformer architecture, FS-GNNTR to face the challenge of low-data in molecular property prediction. The proposed model accepts molecules in the form of molecular graphs to model the local spatial context of molecular graph embeddings while preserving the global information of deep representations. Furthermore, we introduce a two-module meta-learning framework to iteratively update model parameters across few-shot tasks and predict new molecular properties with limited available data. Finally, we conduct multiple experiments on small-sized biological datasets for molecular property prediction, Tox21 and SIDER, and our results demonstrate the superior performance of FS-GNNTR compared to simpler graph-based baselines. The code and data underlying this article are available in the repository, https://github.com/ltorres97/FS-GNNTR. Luis H. M. Torres, Bernardete Ribeiro, Joel Arrais |
Expert Syst. Appl. | 3 |
| 2023 | Few-shot learning via graph embeddings with convolutional networks for low-data molecular property predictionabstractAbstract Graph neural networks and convolutional architectures have proven to be pivotal in improving the prediction of molecular properties in drug discovery. However, this is fundamentally a low data problem that is incompatible with regular deep learning approaches. Contemporary deep networks require large amounts of training data, which severely limits the prediction of new molecular entities from limited available data. In this paper, we address the challenge of low data in molecular property prediction by: (1) defining a set of deep learning architectures that accept compound chemical structures in the form of molecular graphs, (2) creating a few-shot learning strategy across graph neural networks and convolutional neural networks to leverage the rich information of graph embeddings, and (3) proposing a two-module meta-learning framework to learn from task-transferable knowledge and predict molecular properties on few-shot data. Furthermore, we conduct multiple experiments on two benchmark multiproperty datasets to demonstrate a superior performance over conventional graph-based baselines. ROC-AUC results for 10-shot experiments show an average improvement of $$+11.37\%$$ + 11.37 % on Tox21 and $$+0.53\%$$ + 0.53 % on SIDER, which are representative small-sized biological datasets for molecular property prediction. Luis H. M. Torres, Joel Arrais, Bernardete Ribeiro |
Neural Comput. Appl. | 2 |
| 2022 | Deep Model for Anticancer Drug Response through Genomic Profiles and Compound StructuresabstractCancer is among the deadliest diseases, enhancing the need for its detection and treatment. In the era of precision medicine, the main goal is to take into account individual vari-ability in order to choose more accurately which treatment and prevention strategies suit each person. However, drug response prediction for cancer therapy remains a challenge. In this work, we propose a deep neural network model to predict the effect of anticancer drugs in tumors through the half-maximal inhibitory concentration (IC50). The model can be seen as two-fold: first, we pre-trained two autoencoders with high-dimensional gene expression and mutation data to capture the crucial features from tumors; then, this genetic background is translated to cancer cell lines to predict the impact of the genetic variants on a given drug. Moreover, SMILES structures were introduced so that the model can apprehend relevant features regarding the drug compound. Finally, we use drug sensitivity data correlated to the genomic and drugs data to identify features that predict the IC50 value for each pair of drug-cell line. The obtained results demonstrate the effectiveness of the extracted deep representations in the prediction of drug-target interactions, achieving a performance of a mean squared error of 1.07 and surpassing previous state-of-the-art models. Filipa G. Carvalho, Maryam Abbasi, Bernardete Ribeiro, Joel Arrais |
CBMS | 4 |
| 2022 | Deep generative model for therapeutic targets using transcriptomic disease-associated data - USP7 case studyabstractThe generation of candidate hit molecules with the potential to be used in cancer treatment is a challenging task. In this context, computational methods based on deep learning have been employed to improve in silico drug design methodologies. Nonetheless, the applied strategies have focused solely on the chemical aspect of the generation of compounds, disregarding the likely biological consequences for the organism's dynamics. Herein, we propose a method to implement targeted molecular generation that employs biological information, namely, disease-associated gene expression data, to conduct the process of identifying interesting hits. When applied to the generation of USP7 putative inhibitors, the framework managed to generate promising compounds, with more than 90% of them containing drug-like properties and essential active groups for the interaction with the target. Hence, this work provides a novel and reliable method for generating new promising compounds focused on the biological context of the disease. Tiago Pereira 0001, Maryam Abbasi, Rita I. Oliveira, Romina A. Guedes, Jorge A. R. Salvador, Joel Arrais |
Briefings Bioinform. | 6 |
| 2022 | Explainable deep drug-target representations for binding affinity predictionabstractBACKGROUND: Several computational advances have been achieved in the drug discovery field, promoting the identification of novel drug-target interactions and new leads. However, most of these methodologies have been overlooking the importance of providing explanations to the decision-making process of deep learning architectures. In this research study, we explore the reliability of convolutional neural networks (CNNs) at identifying relevant regions for binding, specifically binding sites and motifs, and the significance of the deep representations extracted by providing explanations to the model's decisions based on the identification of the input regions that contributed the most to the prediction. We make use of an end-to-end deep learning architecture to predict binding affinity, where CNNs are exploited in their capacity to automatically identify and extract discriminating deep representations from 1D sequential and structural data. RESULTS: The results demonstrate the effectiveness of the deep representations extracted from CNNs in the prediction of drug-target interactions. CNNs were found to identify and extract features from regions relevant for the interaction, where the weight associated with these spots was in the range of those with the highest positive influence given by the CNNs in the prediction. The end-to-end deep learning model achieved the highest performance both in the prediction of the binding affinity and on the ability to correctly distinguish the interaction strength rank order when compared to baseline approaches. CONCLUSIONS: This research study validates the potential applicability of an end-to-end deep learning architecture in the context of drug discovery beyond the confined space of proteins and ligands with determined 3D structure. Furthermore, it shows the reliability of the deep representations extracted from the CNNs by providing explainability to the decision-making process. Nelson R. C. Monteiro, Carlos J. V. Simões, Henrique V. Ávila, Maryam Abbasi, José Luís Oliveira, Joel Arrais |
BMC Bioinform. | 6 |
| 2021 | Optimizing Recurrent Neural Network Architectures for De Novo Drug DesignabstractIn drug discovery, Deep Learning algorithms are emerging as a potential method to generate novel chemical structures since they can speed up the traditional process and decrease expenditure. Recurrent architectures are amongst the most promising methods for computational de novo drug design. One current challenge consists in finding the optimal architecture and parameters for the recurrent network that assures the generation of valid molecules that span the chemical space. In this work we perform an evaluation on Recurrent Neural Networks which can learn the syntax of molecular representation in terms of SMILES notation. We optimize the computational framework based on the recurrent architecture and its hyper-parameters. Moreover, we evaluate the performance of two types of encoding and spatial arrangement of molecules: Embedding and One-hot Encoding, and datasets with and without stereo-chemical information, respectively. The proposed model showed improved performance when compared to the current literature, both in terms of percentage of valid generated SMILES and diversity with 98.7% and 0.88, for the ChEMBL dataset, respectively. Even when considering the ZINC biogenic library, with stereochemical information, the values were 94.5% and 0.90. The obtained results reveal the potential of the recurrent architectures in learning the SMILES syntax and adding novelty to generate promising compounds. Beatriz P. Santos, Maryam Abbasi, Tiago Pereira 0001, Bernardete Ribeiro, Joel Arrais |
CBMS | 5 |
| 2021 | Analysing Games for Health through Users' Opinion MiningabstractSerious games are a category of games which purpose extends beyond entertainment. Among these, we find a specific type of games, exergames, which aim to promote physical activity. Despite the positive influence of exergames on their users, players often stop playing them after a short period of time, losing the positive benefits of gameplay. It is in this context, that the need to better understand the experience and the opinions of players emerges. To grasp what users feel and (dis)like is key so that games can be redesigned and improved to fit users preferences. This work proposes to analyse users' comments from YouTube, using Natural Language Processing techniques, to extract knowledge on usability, user experience and perceived impacts on health, in particular on quality of life, that could inform the redesign of the game thereafter. This paper is a work in progress that reports on preliminary work that explores users opinions about the Just Dance game. In mining users' opinions, the process of annotation of usability, user experience and quality of life dimensions is based on a pre-established vocabulary. Each extracted opinion is annotated with the concepts present in the opinion, using an approach that is based on the English dictionary lexicon in conjunction with sentiment analysis. The results obtained are then displayed on a dashboard, where the data extracted from the previously collected user comments can be viewed, analysed, and explored. Renato Santos, Joel Arrais, Paula Alexandra Silva |
CBMS | 2 |
| 2021 | Multiobjective Reinforcement Learning in Optimized Drug DesignabstractMachine learning has been increasingly applied with success in generating synthetically reasonable molecules.However, a complete system capable of both producing valid molecules and optimizing multiple traits has remained elusive.This paper employs multiobjective reinforcement learning to draw a framework to design compounds.Different multiobjective techniques have been evaluated, such as weighted sum and Chebyshev.The results show that the implemented model can be effectively optimized towards different and competing molecular properties.Nonetheless, the model implemented with the weighted sum scalarization technique with a weight of 0.55 for biological affinity is the one with the most appropriate trade-off for the different evaluated properties. Maryam Abbasi, Tiago Pereira 0001, Beatriz P. Santos, Bernardete Ribeiro, Joel Arrais |
ESANN | 5 |
| 2021 | Decay Momentum for Improving Federated LearningabstractWe propose two novel Federated Learning (FL) algorithms based on decaying momentum (Demon): Federated Demon (FedDemon) and Federated Demon Adam (FedDemonAdam).In particular, we apply Demon to Momentum Stochastic Gradient Descent (SGD) and Adam in a Federated setting, which has shown to improve results in a centralized environment.We empirically show that FedDemon and FedDemonAdam have a faster convergence rate and performance improvements compared to state-of-the-art algorithms including FedAvg, FedAvgM and FedAdam.17 Miguel Fernandes, Catarina Silva 0001, Joel Arrais, Alberto Cardoso, Bernardete Ribeiro |
ESANN | 3 |
| 2021 | Improvement on Generative Adversarial Network for Targeted Drug DesignabstractThis paper provides a generative network framework that can replicate the molecular space distribution to satisfy a set of desirable features.The approach incorporates two effective machine learning techniques: an Encoder-Decoder architecture that converts the string notations of molecules into latent space and a generative adversarial network to learn the data distribution and generate new compounds.We train this joint model on a dataset that includes stereo-chemical information.The results show an improvement in the Encoder-Decoder performance, reaching 89% of correctly reconstructed molecules.The framework can generate a wide variety of compounds biased towards specific molecular properties using Transfer Learning. Beatriz P. Santos, Maryam Abbasi, Tiago Pereira 0001, Bernardete Ribeiro, Joel Arrais |
ESANN | 5 |
| 2021 | Optimizing blood-brain barrier permeation through deep reinforcement learning for de novo drug designabstractMOTIVATION: The process of placing new drugs into the market is time-consuming, expensive and complex. The application of computational methods for designing molecules with bespoke properties can contribute to saving resources throughout this process. However, the fundamental properties to be optimized are often not considered or conflicting with each other. In this work, we propose a novel approach to consider both the biological property and the bioavailability of compounds through a deep reinforcement learning framework for the targeted generation of compounds. We aim to obtain a promising set of selective compounds for the adenosine A2A receptor and, simultaneously, that have the necessary properties in terms of solubility and permeability across the blood-brain barrier to reach the site of action. The cornerstone of the framework is based on a recurrent neural network architecture, the Generator. It seeks to learn the building rules of valid molecules to sample new compounds further. Also, two Predictors are trained to estimate the properties of interest of the new molecules. Finally, the fine-tuning of the Generator was performed with reinforcement learning, integrated with multi-objective optimization and exploratory techniques to ensure that the Generator is adequately biased. RESULTS: The biased Generator can generate an interesting set of molecules, with approximately 85% having the two fundamental properties biased as desired. Thus, this approach has transformed a general molecule generator into a model focused on optimizing specific objectives. Furthermore, the molecules' synthesizability and drug-likeness demonstrate the potential applicability of the de novo drug design in medicinal chemistry. AVAILABILITY AND IMPLEMENTATION: All code is publicly available in the https://github.com/larngroup/De-Novo-Drug-Design. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tiago Pereira 0001, Maryam Abbasi, José Luís Oliveira, Bernardete Ribeiro, Joel Arrais |
Bioinform. | 5 |
| 2021 | Drug-Target Interaction Prediction: End-to-End Deep Learning ApproachabstractThe discovery of potential Drug-Target Interactions (DTIs) is a determining step in the drug discovery and repositioning process, as the effectiveness of the currently available antibiotic treatment is declining. Although putting efforts on the traditional in vivo or in vitro methods, pharmaceutical financial investment has been reduced over the years. Therefore, establishing effective computational methods is decisive to find new leads in a reasonable amount of time. Successful approaches have been presented to solve this problem but seldom protein sequences and structured data are used together. In this paper, we present a deep learning architecture model, which exploits the particular ability of Convolutional Neural Networks (CNNs) to obtain 1D representations from protein sequences (amino acid sequence) and compounds SMILES (Simplified Molecular Input Line Entry System) strings. These representations can be interpreted as features that express local dependencies or patterns that can then be used in a Fully Connected Neural Network (FCNN), acting as a binary classifier. The results achieved demonstrate that using CNNs to obtain representations of the data, instead of the traditional descriptors, lead to improved performance. The proposed end-to-end deep learning method outperformed traditional machine learning approaches in the correct classification of both positive and negative interactions. Nelson R. C. Monteiro, Bernardete Ribeiro, Joel Arrais |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Exploring a Siamese Neural Network Architecture for One-Shot Drug DiscoveryabstractThe application of deep neural networks in drug discovery is mainly due to their enormous potential to significantly increase the predictive power when inferring the properties and activities of small-molecules. However, in the traditional drug discovery process, where supervised data is scarce, the lead-optimization step is a low-data problem, making it difficult to find molecules with the desired therapeutic activity and obtain accurate predictions for candidate compounds. One major requirement to ensure the validity of the obtained neural network models is the need for a large number of training examples per class, which is not always feasible in drug discovery applications. This invalidates the use of instances whose classes were not considered in the training phase or in data where the number of classes is high and oscillates dynamically. The main objective of the study is to optimize the discovery of novel compounds based on a reduced set of candidate drugs. We propose a Siamese neural network architecture for one-shot classification, based on Convolutional Neural Networks (CNNs), that learns from a similarity score between two input molecules according to a given similarity function. Using a one-shot learning strategy, few instances per class are needed for training, and a small amount of data and computational resources are required to build an accurate model. The results achieved demonstrate that using a Siamese Deep Neural Network for one-shot classification leads to overall improved performance when compared to other state-of the-art models. The proposed architecture provides an accurate and reliable prediction of novel compounds considering the lack of biological data available for drug discovery tasks. Luis H. M. Torres, Nelson R. C. Monteiro, José Luís Oliveira, Joel Arrais, Bernardete Ribeiro |
BIBE | 4 |
| 2020 | Using a Novel Unbiased Dataset and Deep Learning Architectures to Predict Protein-Protein InteractionsabstractProteins are indispensable to the living organisms and are the backbone of almost all cellular processes. However, these macromolecules rarely act alone, forming the protein-protein interactions. Given their biological significance it should come as no surprise that their deregulation is one of the main causes to several disease states.The sudden surge of interest in this field of study motivated the development of innovative in silico methods. Despite the obvious advances in recent years, the effectiveness of these computational methods remains questionable. There is still not enough evidence to support the exclusive use of in silico techniques to predict protein-protein interactions not yet experimentally determined. It is proved that one of the primary reasons leading to this situation is the non-existence of a ”gold-standard” negative interactions dataset. Contrary to the high abundance of publicly available positive interactions, the negative examples are often artificially generated, culminating in biased samples.In this paper a new unbiased dataset that does not overly constraint the negative interactions distribution is presented. Beyond the novel dataset, also distinct deep learning models are proposed as a tool to predict whether two individual proteins are capable of interacting with each other, using exclusively the complete raw amino acid sequences. The obtained results firmly indicate that the proposed models are actually a valuable tool to predict protein-protein interactions, mainly when compared with the existing approaches, while also highlighting that there is still some room for improvement when implemented in unbiased datasets. Luís Silva, Carlos Pereira, Joel Arrais |
BIBM | 3 |
| 2020 | Exploring Time-Series Through Force-Directed TimelinesabstractTemporal datasets are a product of many scientific disciplines and analyzing the events that they describe may help provide valuable insight into their respective research subjects and help move towards solutions to existing problems. Time-series analysis is still an open problem which prompts new solutions, particularly the discovery of patterns across complex temporal networks. Visualization has proven to be a valuable tool in the analysis of such datasets, with the emergence of new models such as Time Curves, which distorts timelines to position time points based on their similarity, creating visualizations that highlight behavior patterns. In this paper, we further explore time-series functionally and aesthetically by revising the dynamic Time Curves models in CroP, a visualization tool with coordinated multiple views. Firstly, we propose the additional of new visual elements and interactive functions, coordinated with a network visualization to help discover and understand temporal patterns across complex datasets. Secondly, we visually explore time-series through Time Paths, a parameter-based force-directed layout that can dynamically transform the original model to either highlight small data variations or reduce visual noise in favor of overall patterns. António Cruz, Joel Arrais, Penousal Machado |
IV | 2 |
| 2020 | CroP - Coordinated Panel visualization for biological networks analysisabstractSUMMARY: CroP is a data visualization application that focuses on the analysis of relational data that changes over time. While it was specifically designed for addressing the preeminent need to interpret large scale time series from gene expression studies, CroP is prepared to analyze datasets from multiple contexts. Multiple datasets can be uploaded simultaneously and viewed through dynamic visualization models, which are contained within flexible panels that allow users to adapt the workspace to their data. Through clustering and the time curve visualization it is possible to quickly identify groups of data points with similar proprieties or behaviors, as well as temporal patterns across all points, such as periodic waves of expression. Additionally, it integrates a public biomedical database for gene annotation. CroP will be of major interest to biologists who seek to extract relations from complex sets of data. AVAILABILITY AND IMPLEMENTATION: CroP is freely available for download as an executable jar at https://cdv.dei.uc.pt/crop/. António Cruz, Penousal Machado, Joel Arrais |
Bioinform. | 3 |
| 2019 | Interactive and coordinated visualization approaches for biological data analysisabstractThe field of computational biology has become largely dependent on data visualization tools to analyze the increasing quantities of data gathered through the use of new and growing technologies. Aside from the volume, which often results in large amounts of noise and complex relationships with no clear structure, the visualization of biological data sets is hindered by their heterogeneity, as data are obtained from different sources and contain a wide variety of attributes, including spatial and temporal information. This requires visualization approaches that are able to not only represent various data structures simultaneously but also provide exploratory methods that allow the identification of meaningful relationships that would not be perceptible through data analysis algorithms alone. In this article, we present a survey of visualization approaches applied to the analysis of biological data. We focus on graph-based visualizations and tools that use coordinated multiple views to represent high-dimensional multivariate data, in particular time series gene expression, protein-protein interaction networks and biological pathways. We then discuss how these methods can be used to help solve the current challenges surrounding the visualization of complex biological data sets. António Cruz, Joel Arrais, Penousal Machado |
Briefings Bioinform. | 2 |
| 2018 | Interactive Network Visualization of Gene Expression Time-Series DataabstractVisualization models have shown to be remarkably important in the interpretation of datasets across many fields of study. In the field of Biology, data visualization is used to better understand processes that range from phylogenetic trees to multiple layers of molecular networks. The latter is especially challenging due to the large quantities of varying elements and complex relationships, often with no perceptible structure. Although various tools have been proposed to improve the visualization of molecular networks, many challenges still persist. In this paper, we propose a tool that uses interactive visualization models to represent the dynamic behaviors of molecular networks. The tool employs various methods to explore and organize the data, including clustering, force-directed layouts, and a timeline for navigating through time-series data. To further analyze temporal attributes, the timeline can be distorted through a force-directed layout to spatially position time points according to their similarity. Additionally, gene expression can be annotated through an integrated biological database. The visualization model was validated with the use of time-series gene expression RNA-Seq data from the HIV-1 infection. António Cruz, Joel Arrais, Penousal Machado |
IV | 2 |
| 2016 | Ensemble-Based Methodology for the Prediction of Drug-Target InteractionsabstractAntibacterial resistance has been progressively increasing mostly due to selective antibiotic pressure, forcing pathogens to either adapt or die. The development of antibacterial resistance to last-line antibiotics urges the formulation of alternative strategies for drug discovery. Recently, attention has been devoted to the development of computational methods to predict drug-target interactions (DTIs). Here we present a computational strategy to predict proteome-scale DTIs based on the combination of the drugs' chemical features and substructural fingerprints, and on the structural information and physicochemical properties of the proteins. We propose an ensemble learning combination of Support-Vector Machine and Random Forest to deal with the complexity of DTI classification. Two distinct classification models were developed to ascertain whether taking the type of protein target (i.e., enzymes, g-protein-coupled receptors, ion channels and nuclear receptors) into account improves classification performance. External validation analysis was consistent with internal five-fold cross-validation, with an AUC of 0.87. This strategy was applied to the proteome of methicillin-resistant Staphylococcus aureus COL (MRSA COL, taxonomy id: 93062), a major nosocomial pathogen worldwide whose antimicrobial resistance and incidence rate keeps steadily increasing. Our predictive framework is available at http://bioinformatics.ua.pt/software/dtipred. Edgar D. Coelho, Joel Arrais, José Luís Oliveira |
CBMS | 2 |
| 2016 | Reconstructing the temporal progression of HIV-1 immune response pathwaysabstractMOTIVATION: Most methods for reconstructing response networks from high throughput data generate static models which cannot distinguish between early and late response stages. RESULTS: We present TimePath, a new method that integrates time series and static datasets to reconstruct dynamic models of host response to stimulus. TimePath uses an Integer Programming formulation to select a subset of pathways that, together, explain the observed dynamic responses. Applying TimePath to study human response to HIV-1 led to accurate reconstruction of several known regulatory and signaling pathways and to novel mechanistic insights. We experimentally validated several of TimePaths' predictions highlighting the usefulness of temporal models. AVAILABILITY AND IMPLEMENTATION: Data, Supplementary text and the TimePath software are available from http://sb.cs.cmu.edu/timepath CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Siddhartha Jain 0001, Joel Arrais, Narasimhan J. Venkatachari, Velpandi Ayyavoo, Ziv Bar-Joseph |
Bioinform. | 2 |
| 2016 | Computational Discovery of Putative Leads for Drug Repositioning through Drug-Target Interaction PredictionabstractDe novo experimental drug discovery is an expensive and time-consuming task. It requires the identification of drug-target interactions (DTIs) towards targets of biological interest, either to inhibit or enhance a specific molecular function. Dedicated computational models for protein simulation and DTI prediction are crucial for speed and to reduce the costs associated with DTI identification. In this paper we present a computational pipeline that enables the discovery of putative leads for drug repositioning that can be applied to any microbial proteome, as long as the interactome of interest is at least partially known. Network metrics calculated for the interactome of the bacterial organism of interest were used to identify putative drug-targets. Then, a random forest classification model for DTI prediction was constructed using known DTI data from publicly available databases, resulting in an area under the ROC curve of 0.91 for classification of out-of-sampling data. A drug-target network was created by combining 3,081 unique ligands and the expected ten best drug targets. This network was used to predict new DTIs and to calculate the probability of the positive class, allowing the scoring of the predicted instances. Molecular docking experiments were performed on the best scoring DTI pairs and the results were compared with those of the same ligands with their original targets. The results obtained suggest that the proposed pipeline can be used in the identification of new leads for drug repositioning. The proposed classification model is available at http://bioinformatics.ua.pt/software/dtipred/. Edgar D. Coelho, Joel Arrais, José Luís Oliveira |
PLoS Comput. Biol. | 2 |
| 2014 | geneCommittee: a web-based tool for extensively testing the discriminatory power of biologically relevant gene sets in microarray data classificationabstractBACKGROUND: The diagnosis and prognosis of several diseases can be shortened through the use of different large-scale genome experiments. In this context, microarrays can generate expression data for a huge set of genes. However, to obtain solid statistical evidence from the resulting data, it is necessary to train and to validate many classification techniques in order to find the best discriminative method. This is a time-consuming process that normally depends on intricate statistical tools. RESULTS: geneCommittee is a web-based interactive tool for routinely evaluating the discriminative classification power of custom hypothesis in the form of biologically relevant gene sets. While the user can work with different gene set collections and several microarray data files to configure specific classification experiments, the tool is able to run several tests in parallel. Provided with a straightforward and intuitive interface, geneCommittee is able to render valuable information for diagnostic analyses and clinical management decisions based on systematically evaluating custom hypothesis over different data sets using complementary classifiers, a key aspect in clinical research. CONCLUSIONS: geneCommittee allows the enrichment of microarrays raw data with gene functional annotations, producing integrated datasets that simplify the construction of better discriminative hypothesis, and allows the creation of a set of complementary classifiers. The trained committees can then be used for clinical research and diagnosis. Full documentation including common use cases and guided analysis workflows is freely available at http://sing.ei.uvigo.es/GC/. Miguel Reboiro-Jato, Joel Arrais, José Luís Oliveira, Florentino Fernández Riverola |
BMC Bioinform. | 2 |
| 2010 | GeneBrowser 2: an application to explore and identify common biological traits in a set of genesabstractBACKGROUND: The development of high-throughput laboratory techniques created a demand for computer-assisted result analysis tools. Many of these techniques return lists of genes whose interpretation requires finding relevant biological roles for the problem at hand. The required information is typically available in public databases, and usually, this information must be manually retrieved to complement the analysis. This process is a very time-consuming task that should be automated as much as possible. RESULTS: GeneBrowser is a web-based tool that, for a given list of genes, combines data from several public databases with visualisation and analysis methods to help identify the most relevant and common biological characteristics. The functionalities provided include the following: a central point with the most relevant biological information for each inserted gene; a list of the most related papers in PubMed and gene expression studies in ArrayExpress; and an extended approach to functional analysis applied to Gene Ontology, homologies, gene chromosomal localisation and pathways. CONCLUSIONS: GeneBrowser provides a unique entry point to several visualisation and analysis methods, providing fast and easy analysis of a set of genes. GeneBrowser fills the gap between Web portals that analyse one gene at a time and functional analysis tools that are limited in scope and usually desktop-based. Joel Arrais, João E. Pereira, José Luís Oliveira |
BMC Bioinform. | 1 |
| 2010 | Concept-based query expansion for retrieving gene related publications from MEDLINEabstractBACKGROUND: Advances in biotechnology and in high-throughput methods for gene analysis have contributed to an exponential increase in the number of scientific publications in these fields of study. While much of the data and results described in these articles are entered and annotated in the various existing biomedical databases, the scientific literature is still the major source of information. There is, therefore, a growing need for text mining and information retrieval tools to help researchers find the relevant articles for their study. To tackle this, several tools have been proposed to provide alternative solutions for specific user requests. RESULTS: This paper presents QuExT, a new PubMed-based document retrieval and prioritization tool that, from a given list of genes, searches for the most relevant results from the literature. QuExT follows a concept-oriented query expansion methodology to find documents containing concepts related to the genes in the user input, such as protein and pathway names. The retrieved documents are ranked according to user-definable weights assigned to each concept class. By changing these weights, users can modify the ranking of the results in order to focus on documents dealing with a specific concept. The method's performance was evaluated using data from the 2004 TREC genomics track, producing a mean average precision of 0.425, with an average of 4.8 and 31.3 relevant documents within the top 10 and 100 retrieved abstracts, respectively. CONCLUSIONS: QuExT implements a concept-based query expansion scheme that leverages gene-related information available on a variety of biological resources. The main advantage of the system is to give the user control over the ranking of the results by means of a simple weighting scheme. Using this approach, researchers can effortlessly explore the literature regarding a group of genes and focus on the different aspects relating to these genes. Sérgio Matos, Joel Arrais, João Maia-Rodrigues, José Luís Oliveira |
BMC Bioinform. | 2 |
| 2008 | Dynamic service integration using web-based workflowsabstractWeb services have been the main leverage to the development of Service Oriented Architecture (SOA), essentially a collection of interacting software agents with a loosely coupling organization. Despite several composition solutions already exist, the integration of services in a seamless and user-friendly way is not yet a complete solved problem. Most of the times, this integration is performed in hard-coded monolithic desktop applications. Pedro Lopes 0002, Joel Arrais, José Luís Oliveira |
iiWAS | 2 |
| 2004 | Hardware/Software Implementation of FPGA-Targeted Matrix-Oriented SAT Solvers
Valery Sklyarov, Iouliia Skliarova, Bruno Figueiredo Pimentel, Joel Arrais |
FPL | 4 |