Oscar Dias

dblp:71/835 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-1765-7178ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Correction: A diel multi-tissue genome-scale metabolic model of Vitis vinifera
abstract
[This corrects the article DOI: 10.1371/journal.pcbi.1012506.].
Marta Sampaio, Miguel Rocha 0001, Oscar Dias
PLoS Comput. Biol.3
2025 Comparative Assessment of Protein Large Language Models for Enzyme Commission Number Prediction
abstract
BACKGROUND: Protein large language models (LLM) have been used to extract representations of enzyme sequences to predict their function, which is encoded by enzyme commission (EC) numbers. However, a comprehensive comparison of different LLMs for this task is still lacking, leaving questions about their relative performance. Moreover, protein sequence alignments (e.g. BLASTp or DIAMOND) are often combined with machine learning models to assign EC numbers from homologous enzymes, thus compensating for the shortcomings of these models' predictions. In this context, LLMs and sequence alignment methods have not been extensively compared as individual predictors, raising unaddressed questions about LLMs' performance and limitations relative to the alignment methods. In this study, we set out to assess the performance of ESM2, ESM1b, and ProtBERT language models in their ability to predict EC numbers, comparing them with BLASTp, against each other and against models that rely on one-hot encodings of amino acid sequences. RESULTS: Our findings reveal that combining these LLMs with fully connected neural networks surpasses the performance of deep learning models that rely on one-hot encodings. Moreover, although BLASTp provided marginally better results overall, DL models provide results that complement BLASTp's, revealing that LLMs better predict certain EC numbers while BLASTp excels in predicting others. The ESM2 stood out as the best model among the LLMs tested, providing more accurate predictions on difficult annotation tasks and for enzymes without homologs. CONCLUSIONS: Crucially, this study demonstrates that LLMs still have to be improved to become the gold standard tool over BLASTp in mainstream enzyme annotation routines. On the other hand, LLMs can provide good predictions for more difficult-to-annotate enzymes, particularly when the identity between the query sequence and the reference database falls below 25%. Our results reinforce the claim that BLASTp and LLM models complement each other and can be more effective when used together.
João Capela, Maria Zimmermann-Kogadeeva, Aalt D. J. van Dijk, Dick de Ridder, Oscar Dias, Miguel Rocha 0001
BMC Bioinform.5
2024 A diel multi-tissue genome-scale metabolic model of Vitis vinifera
abstract
Vitis vinifera, also known as grapevine, is widely cultivated and commercialized, particularly to produce wine. As wine quality is directly linked to fruit quality, studying grapevine metabolism is important to understand the processes underlying grape composition. Genome-scale metabolic models (GSMMs) have been used for the study of plant metabolism and advances have been made, allowing the integration of omics datasets with GSMMs. On the other hand, Machine learning (ML) has been used to analyze and integrate omics data, and while the combination of ML with GSMMs has shown promising results, it is still scarcely used to study plants. Here, the first GSSM of V. vinifera was reconstructed and validated, comprising 7199 genes, 5399 reactions, and 5141 metabolites across 8 compartments. Tissue-specific models for the stem, leaf, and berry of the Cabernet Sauvignon cultivar were generated from the original model, through the integration of RNA-Seq data. These models have been merged into diel multi-tissue models to study the interactions between tissues at light and dark phases. The potential of combining ML with GSMMs was explored by using ML to analyze the fluxomics data generated by green and mature grape GSMMs and provide insights regarding the metabolism of grapes at different developmental stages. Therefore, the models developed in this work are useful tools to explore different aspects of grapevine metabolism and understand the factors influencing grape quality.
Marta Sampaio, Miguel Rocha 0001, Oscar Dias
PLoS Comput. Biol.3
2024 BioISO: An Objective-Oriented Application for Assisting the Curation of Genome-Scale Metabolic Models
abstract
As the reconstruction of Genome-Scale Metabolic Models (GEMs) becomes standard practice in systems biology, the number of organisms having at least one metabolic model is peaking at an unprecedented scale. The automation of laborious tasks, such as gap-finding and gap-filling, allowed the development of GEMs for poorly described organisms. However, the quality of these models can be compromised by the automation of several steps, which may lead to erroneous phenotype simulations. Biological networks constraint-based In Silico Optimisation (BioISO) is a computational tool aimed at accelerating the reconstruction of GEMs. This tool facilitates manual curation steps by reducing the large search spaces often met when debugging in silico biological models. BioISO uses a recursive relation-like algorithm and Flux Balance Analysis (FBA) to evaluate and guide debugging of in silico phenotype simulations. The potential of BioISO to guide the debugging of model reconstructions was showcased and compared with the results of two other state-of-the-art gap-filling tools (Meneco and fastGapFill). In this assessment, BioISO is better suited to reducing the search space for errors and gaps in metabolic networks by identifying smaller ratios of dead-end metabolites. Furthermore, BioISO was used as Meneco's gap-finding algorithm to reduce the number of proposed solutions for filling the gaps.
Fernando Cruz, João Capela, Eugénio C. Ferreira, Miguel Rocha 0001, Oscar Dias
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 TranSyT, an innovative framework for identifying transport systems
abstract
MOTIVATION: The importance and rate of development of genome-scale metabolic models have been growing for the last few years, increasing the demand for software solutions that automate several steps of this process. However, since TRIAGE's release, software development for the automatic integration of transport reactions into models has stalled. RESULTS: Here, we present the Transport Systems Tracker (TranSyT). Unlike other transport systems annotation software, TranSyT does not rely on manual curation to expand its internal database, which is derived from highly curated records retrieved from the Transporters Classification Database and complemented with information from other data sources. TranSyT compiles information regarding transporter families and proteins, and derives reactions into its internal database, making it available for rapid annotation of complete genomes. All transport reactions have GPR associations and can be exported with identifiers from four different metabolite databases. TranSyT is currently available as a plugin for merlin v4.0 and an app for KBase. AVAILABILITY AND IMPLEMENTATION: TranSyT web service: https://transyt.bio.di.uminho.pt/; GitHub for the tool: https://github.com/BioSystemsUM/transyt; GitHub with examples and instructions to run TranSyT: https://github.com/ecunha1996/transyt_paper.
Emanuel Cunha, Davide Lagoa, José P. Faria, Filipe Liu, Christopher S. Henry, Oscar Dias
Bioinform.6
2023 The first multi-tissue genome-scale metabolic model of a woody plant highlights suberin biosynthesis pathways in Quercus suber
abstract
Over the last decade, genome-scale metabolic models have been increasingly used to study plant metabolic behaviour at the tissue and multi-tissue level under different environmental conditions. Quercus suber, also known as the cork oak tree, is one of the most important forest communities of the Mediterranean/Iberian region. In this work, we present the genome-scale metabolic model of the Q. suber (iEC7871). The metabolic model comprises 7871 genes, 6231 reactions, and 6481 metabolites across eight compartments. Transcriptomics data was integrated into the model to obtain tissue-specific models for the leaf, inner bark, and phellogen, with specific biomass compositions. The tissue-specific models were merged into a diel multi-tissue metabolic model to predict interactions among the three tissues at the light and dark phases. The metabolic models were also used to analyse the pathways associated with the synthesis of suberin monomers, namely the acyl-lipids, phenylpropanoids, isoprenoids, and flavonoids production. The models developed in this work provide a systematic overview of the metabolism of Q. suber, including its secondary metabolism pathways and cork formation.
Emanuel Cunha, Inês Chaves, Hüseyin Demirci, Davide Lagoa, Miguel Rocha 0001, Isabel Rocha, Oscar Dias
PLoS Comput. Biol.9
2021 A kinetic model of the central carbon metabolism for acrylic acid production in Escherichia coli
abstract
Acrylic acid is a value-added chemical used in industry to produce diapers, coatings, paints, and adhesives, among many others. Due to its economic importance, there is currently a need for new and sustainable ways to synthesise it. Recently, the focus has been laid in the use of Escherichia coli to express the full bio-based pathway using 3-hydroxypropionate as an intermediary through three distinct pathways (glycerol, malonyl-CoA, and β-alanine). Hence, the goals of this work were to use COPASI software to assess which of the three pathways has a higher potential for industrial-scale production, from either glucose or glycerol, and identify potential targets to improve the biosynthetic pathways yields. When compared to the available literature, the models developed during this work successfully predict the production of 3-hydroxypropionate, using glycerol as carbon source in the glycerol pathway, and using glucose as a carbon source in the malonyl-CoA and β-alanine pathways. Finally, this work allowed to identify four potential over-expression targets (glycerol-3-phosphate dehydrogenase (G3pD), acetyl-CoA carboxylase (AccC), aspartate aminotransferase (AspAT), and aspartate carboxylase (AspC)) that should, theoretically, result in higher AA yields.
Alexandre Oliveira, Joana Rodrigues, Eugénio C. Ferreira, Lígia R. Rodrigues, Oscar Dias
PLoS Comput. Biol.5
2019 Predicting promoters in phage genomes using PhagePromoter
abstract
SUMMARY: The growing interest in phages as antibacterial agents has led to an increase in the number of sequenced phage genomes, increasing the need for intuitive bioinformatics tools for performing genome annotation. The identification of phage promoters is indeed the most difficult step of this process. Due to the lack of online tools for phage promoter prediction, we developed PhagePromoter, a tool for locating promoters in phage genomes, using machine learning methods. This is the first online tool for predicting promoters that uses phage promoter data and the first to identify both host and phage promoters with different motifs. AVAILABILITY AND IMPLEMENTATION: This tool was integrated in the Galaxy framework and it is available online at: https://bit.ly/2Dfebfv. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marta Sampaio, Miguel Rocha 0001, Hugo Oliveira, Oscar Dias
Bioinform.4
2019 SamPler - a novel method for selecting parameters for gene functional annotation routines
abstract
BACKGROUND: As genome sequencing projects grow rapidly, the diversity of organisms with recently assembled genome sequences peaks at an unprecedented scale, thereby highlighting the need to make gene functional annotations fast and efficient. However, the (high) quality of such annotations must be guaranteed, as this is the first indicator of the genomic potential of every organism. Automatic procedures help accelerating the annotation process, though decreasing the confidence and reliability of the outcomes. Manually curating a genome-wide annotation of genes, enzymes and transporter proteins function is a highly time-consuming, tedious and impractical task, even for the most proficient curator. Hence, a semi-automated procedure, which balances the two approaches, will increase the reliability of the annotation, while speeding up the process. In fact, a prior analysis of the annotation algorithm may leverage its performance, by manipulating its parameters, hastening the downstream processing and the manual curation of assigning functions to genes encoding proteins. RESULTS: Here SamPler, a novel strategy to select parameters for gene functional annotation routines is presented. This semi-automated method is based on the manual curation of a randomly selected set of genes/proteins. Then, in a multi-dimensional array, this sample is used to assess the automatic annotations for all possible combinations of the algorithm's parameters. These assessments allow creating an array of confusion matrices, for which several metrics are calculated (accuracy, precision and negative predictive value) and used to reach optimal values for the parameters. CONCLUSIONS: The potential of this methodology is demonstrated with four genome functional annotations performed in merlin, an in-house user-friendly computational framework for genome-scale metabolic annotation and model reconstruction. For that, SamPler was implemented as a new plugin for the merlin tool.
Fernando Cruz, Davide Lagoa, Isabel Rocha, Eugénio C. Ferreira, Miguel Rocha 0001, Oscar Dias
BMC Bioinform.7
2017 Genome-Wide Semi-Automated Annotation of Transporter Systems
abstract
Usually, transport reactions are added to genome-scale metabolic models (GSMMs) based on experimental data and literature. This approach does not allow associating specific genes with transport reactions, which impairs the ability of the model to predict effects of gene deletions. Novel methods for systematic genome-wide transporter functional annotation and their integration into GSMMs are therefore necessary. In this work, an automatic system to detect and classify all potential membrane transport proteins for a given genome and integrate the related reactions into GSMMs is proposed, based on the identification and classification of genes that encode transmembrane proteins. The Transport Reactions Annotation and Generation (TRIAGE) tool identifies the metabolites transported by each transmembrane protein and its transporter family. The localization of the carriers is also predicted and, consequently, their action is confined to a given membrane. The integration of the data provided by TRIAGE with highly curated models allowed the identification of new transport reactions. TRIAGE is included in the new release of merlin, a software tool previously developed by the authors, which expedites the GSMM reconstruction processes.
Oscar Dias, Daniel Gomes, Paulo Vilaça, João G. R. Cardoso, Miguel Rocha 0001, Eugénio C. Ferreira, Isabel Rocha
IEEE ACM Trans. Comput. Biol. Bioinform.1