Miguel Rocha 0001

dblp:17/2067 · also Miguel Francisco de Almeida Pereira da Rocha, Miguel P. Rocha 0001 · DBLP profile ↗
← Back
71ranked-venue papers
14as first author
20since 2021 · last 2026
0000-0001-8439-8172ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 11 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 32 · 2 first-author · 11 since 2021Computer networks · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 NeuroScaler: Towards Energy-Optimal Autoscaling for Container-Based Services
Alisson O. Chaves, Rodrigo Moreira, Larissa F. Rodrigues Moreira, Joao Correia, David Santos, Tiago Barros, Daniel Corujo, Miguel Rocha 0001, Flávio Oliveira Silva 0001
ICC9
2026 Correction: A diel multi-tissue genome-scale metabolic model of Vitis vinifera
abstract
[This corrects the article DOI: 10.1371/journal.pcbi.1012506.].
Marta Sampaio, Miguel Rocha 0001, Oscar Dias
PLoS Comput. Biol.2
2025 Evolutionary Algorithms for Metabolic Transformation through Multi-gene Knockout Optimization
abstract
The Metabolic Transformation Algorithm (MTA) leverages constraint-based modeling to identify metabolic interventions capable of shifting a biological system from an undesired to a desired state. Its robust extension (rMTA) strengthens predictive accuracy through worst-case scenario analyses and the integration of Minimization Of Metabolic Adjustment (MOMA) algorithm.
Bruno Sá, Alexandre Oliveira, Miguel Rocha 0001
GECCO3
2025 Comparative Assessment of Protein Large Language Models for Enzyme Commission Number Prediction
abstract
BACKGROUND: Protein large language models (LLM) have been used to extract representations of enzyme sequences to predict their function, which is encoded by enzyme commission (EC) numbers. However, a comprehensive comparison of different LLMs for this task is still lacking, leaving questions about their relative performance. Moreover, protein sequence alignments (e.g. BLASTp or DIAMOND) are often combined with machine learning models to assign EC numbers from homologous enzymes, thus compensating for the shortcomings of these models' predictions. In this context, LLMs and sequence alignment methods have not been extensively compared as individual predictors, raising unaddressed questions about LLMs' performance and limitations relative to the alignment methods. In this study, we set out to assess the performance of ESM2, ESM1b, and ProtBERT language models in their ability to predict EC numbers, comparing them with BLASTp, against each other and against models that rely on one-hot encodings of amino acid sequences. RESULTS: Our findings reveal that combining these LLMs with fully connected neural networks surpasses the performance of deep learning models that rely on one-hot encodings. Moreover, although BLASTp provided marginally better results overall, DL models provide results that complement BLASTp's, revealing that LLMs better predict certain EC numbers while BLASTp excels in predicting others. The ESM2 stood out as the best model among the LLMs tested, providing more accurate predictions on difficult annotation tasks and for enzymes without homologs. CONCLUSIONS: Crucially, this study demonstrates that LLMs still have to be improved to become the gold standard tool over BLASTp in mainstream enzyme annotation routines. On the other hand, LLMs can provide good predictions for more difficult-to-annotate enzymes, particularly when the identity between the query sequence and the reference database falls below 25%. Our results reinforce the claim that BLASTp and LLM models complement each other and can be more effective when used together.
João Capela, Maria Zimmermann-Kogadeeva, Aalt D. J. van Dijk, Dick de Ridder, Oscar Dias, Miguel Rocha 0001
BMC Bioinform.6
2024 A diel multi-tissue genome-scale metabolic model of Vitis vinifera
abstract
Vitis vinifera, also known as grapevine, is widely cultivated and commercialized, particularly to produce wine. As wine quality is directly linked to fruit quality, studying grapevine metabolism is important to understand the processes underlying grape composition. Genome-scale metabolic models (GSMMs) have been used for the study of plant metabolism and advances have been made, allowing the integration of omics datasets with GSMMs. On the other hand, Machine learning (ML) has been used to analyze and integrate omics data, and while the combination of ML with GSMMs has shown promising results, it is still scarcely used to study plants. Here, the first GSSM of V. vinifera was reconstructed and validated, comprising 7199 genes, 5399 reactions, and 5141 metabolites across 8 compartments. Tissue-specific models for the stem, leaf, and berry of the Cabernet Sauvignon cultivar were generated from the original model, through the integration of RNA-Seq data. These models have been merged into diel multi-tissue models to study the interactions between tissues at light and dark phases. The potential of combining ML with GSMMs was explored by using ML to analyze the fluxomics data generated by green and mature grape GSMMs and provide insights regarding the metabolism of grapes at different developmental stages. Therefore, the models developed in this work are useful tools to explore different aspects of grapevine metabolism and understand the factors influencing grape quality.
Marta Sampaio, Miguel Rocha 0001, Oscar Dias
PLoS Comput. Biol.2
2024 BioISO: An Objective-Oriented Application for Assisting the Curation of Genome-Scale Metabolic Models
abstract
As the reconstruction of Genome-Scale Metabolic Models (GEMs) becomes standard practice in systems biology, the number of organisms having at least one metabolic model is peaking at an unprecedented scale. The automation of laborious tasks, such as gap-finding and gap-filling, allowed the development of GEMs for poorly described organisms. However, the quality of these models can be compromised by the automation of several steps, which may lead to erroneous phenotype simulations. Biological networks constraint-based In Silico Optimisation (BioISO) is a computational tool aimed at accelerating the reconstruction of GEMs. This tool facilitates manual curation steps by reducing the large search spaces often met when debugging in silico biological models. BioISO uses a recursive relation-like algorithm and Flux Balance Analysis (FBA) to evaluate and guide debugging of in silico phenotype simulations. The potential of BioISO to guide the debugging of model reconstructions was showcased and compared with the results of two other state-of-the-art gap-filling tools (Meneco and fastGapFill). In this assessment, BioISO is better suited to reducing the search space for errors and gaps in metabolic networks by identifying smaller ratios of dead-end metabolites. Furthermore, BioISO was used as Meneco's gap-finding algorithm to reduce the number of proposed solutions for filling the gaps.
Fernando Cruz, João Capela, Eugénio C. Ferreira, Miguel Rocha 0001, Oscar Dias
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 Word embeddings for protein sequence analysis
abstract
Nowadays, the ability to predict protein functions directly from amino acid sequences alone remains a major biological challenge. The understanding of protein properties and functions is extremely important and can have a wide range of biotechnological and medical applications. High-level representations from the field of deep learning can provide new alternatives to address these problems, particularly Natural Language Processes methods, such as word embeddings (WE), which have shown particular success when applied to protein sequence analysis. Here, a software tool that eases the implementation of WE models toward protein representation and classification is presented. This tool was validated using two protein classification problems namely, the identification of plant ubiquitylation sites and lysine crotonylation sites, and further used to explore enzyme functional annotation. Several WE were tested and fed to different Machine and Deep Learning networks. Overall, WE achieved good results being even competitive with state-of-the-art, reinforcing the idea that language-based methods can be applied with success to a wide range of protein classification problems.
Ana Marta Sequeira, Ivan Gomes, Miguel Rocha 0001
CIBCB3
2023 Combining Evolutionary Algorithms with Reaction Rules Towards Focused Molecular Design
abstract
Designing novel small molecules with desirable properties and feasible synthesis continues to pose a significant challenge in drug discovery, particularly in the realm of natural products. Reaction-based gradient-free methods are promising approaches for designing new molecules as they ensure synthetic feasibility and provide potential synthesis paths. However, it is important to note that the novelty and diversity of the generated molecules highly depend on the availability of comprehensive reaction templates. To address this challenge, we introduce ReactEA, a new open-source evolutionary framework for computer-aided drug discovery that solely utilizes biochemical reaction rules. ReactEA optimizes molecular properties using a comprehensive set of 22,949 reaction rules, ensuring chemical validity and synthetic feasibility. ReactEA is versatile, as it can virtually optimize any objective function and track potential synthetic routes during the optimization process. To demonstrate its effectiveness, we apply ReactEA to various case studies, including the design of novel drug-like molecules and the optimization of pre-existing ligands. The results show that ReactEA consistently generates novel molecules with improved properties and reasonable synthetic routes, even for complex tasks such as improving binding affinity against the PARP1 enzyme when compared to existing inhibitors.
João Correia 0002, Vítor Pereira 0001, Miguel Rocha 0001
GECCO3
2023 A systematic evaluation of deep learning methods for the prediction of drug synergy in cancer
abstract
One of the main obstacles to the successful treatment of cancer is the phenomenon of drug resistance. A common strategy to overcome resistance is the use of combination therapies. However, the space of possibilities is huge and efficient search strategies are required. Machine Learning (ML) can be a useful tool for the discovery of novel, clinically relevant anti-cancer drug combinations. In particular, deep learning (DL) has become a popular choice for modeling drug combination effects. Here, we set out to examine the impact of different methodological choices on the performance of multimodal DL-based drug synergy prediction methods, including the use of different input data types, preprocessing steps and model architectures. Focusing on the NCI ALMANAC dataset, we found that feature selection based on prior biological knowledge has a positive impact-limiting gene expression data to cancer or drug response-specific genes improved performance. Drug features appeared to be more predictive of drug response, with a 41% increase in coefficient of determination (R2) and 26% increase in Spearman correlation relative to a baseline model that used only cell line and drug identifiers. Molecular fingerprint-based drug representations performed slightly better than learned representations-ECFP4 fingerprints increased R2 by 5.3% and Spearman correlation by 2.8% w.r.t the best learned representations. In general, fully connected feature-encoding subnetworks outperformed other architectures. DL outperformed other ML methods by more than 35% (R2) and 14% (Spearman). Additionally, an ensemble combining the top DL and ML models improved performance by about 6.5% (R2) and 4% (Spearman). Using a state-of-the-art interpretability method, we showed that DL models can learn to associate drug and cell line features with drug response in a biologically meaningful way. The strategies explored in this study will help to improve the development of computational methods for the rational design of effective drug combinations for cancer therapy.
Delora Baptista, Pedro G. Ferreira, Miguel Rocha 0001
PLoS Comput. Biol.3
2023 The first multi-tissue genome-scale metabolic model of a woody plant highlights suberin biosynthesis pathways in Quercus suber
abstract
Over the last decade, genome-scale metabolic models have been increasingly used to study plant metabolic behaviour at the tissue and multi-tissue level under different environmental conditions. Quercus suber, also known as the cork oak tree, is one of the most important forest communities of the Mediterranean/Iberian region. In this work, we present the genome-scale metabolic model of the Q. suber (iEC7871). The metabolic model comprises 7871 genes, 6231 reactions, and 6481 metabolites across eight compartments. Transcriptomics data was integrated into the model to obtain tissue-specific models for the leaf, inner bark, and phellogen, with specific biomass compositions. The tissue-specific models were merged into a diel multi-tissue metabolic model to predict interactions among the three tissues at the light and dark phases. The metabolic models were also used to analyse the pathways associated with the synthesis of suberin monomers, namely the acyl-lipids, phenylpropanoids, isoprenoids, and flavonoids production. The models developed in this work provide a systematic overview of the metabolism of Q. suber, including its secondary metabolism pathways and cork formation.
Emanuel Cunha, Inês Chaves, Hüseyin Demirci, Davide Lagoa, Miguel Rocha 0001, Isabel Rocha, Oscar Dias
PLoS Comput. Biol.7
2022 Variational Autoencoders and Evolutionary Algorithms for Targeted Novel Enzyme Design
abstract
Recent developments in Generative Deep Learning have fostered new engineering methods for protein design. Although deep generative models trained on protein sequence can learn biologically meaningful representations, the design of proteins with optimised properties remains a challenge. We combined deep learning architectures with evolutionary computation to steer the protein generative process towards specific sets of properties to address this problem. The latent space of a Variational Autoencoder is explored by evolutionary algorithms to find the best candidates. A set of single-objective and multi-objective problems were conceived to evaluate the algorithms' capacity to optimise proteins. The optimisation tasks consider the average proteins' hydrophobicity, their solubility and the probability of being generated by a defined functional Hidden Markov Model profile. The results show that Evolutionary Algorithms can achieve good results while allowing for more variability in the design of the experiment, thus resulting in a much greater set of possibly functional novel proteins.
Miguel Martins, Miguel Rocha 0001, Vítor Pereira 0001
CEC2
2022 Development of Deep Learning approaches to predict relationships between chemical structures and sweetness
abstract
The non-caloric sweeteners market is catching up with the market of conventionally used sugars due to the benefits of preventing obesity, tooth decay and other health problems. Developing strategies for designing easier-to-produce novel molecules with a sweet taste and less toxicity are up-to-date motivations for the food industry. In this sense, Machine Learning (ML) approaches have been reported as cutting-edge technologies to guide the design of new molecules towards specific objectives, including sweet taste. The largest known dataset of sweet molecules is here provided. The dataset contains fully integrated 9541 sweeteners and 1141 bitterants from FooDB, FlavorDB and literature. This robust dataset allowed the development of standard Machine and Deep Learning pipelines towards conceiving Structure-Activity Relationships (SAR) between molecules and sweetness. In this work, we showcase that Textual Convolutional Neural Networks (TextCNN), Graph Convolutional Networks (GCN), and Deep Neural Networks (DNNs) outperformed most of traditional “shallow” learning approaches. These Deep Learning (DL) models produced platforms to guide the design of new sweeteners and repurposing existing compounds. Sixty million compounds from PubChem were evaluated using these models. Herein, we deliver a dataset of 67724 compounds that present high probabilities of being sweet. Quick searches in literature allowed us to find 13 molecules reported as potent sweetening agents, revealing that our approach is suitable for finding new sweeteners, valuable to expand food chemistry databases, repurposing existing chemicals and designing novel molecules with a sweet taste.
João Capela, João Correia 0002, Vítor Pereira 0001, Miguel Rocha 0001
IJCNN4
2022 Predicting the number of biochemical transformations needed to synthesize a compound
abstract
Exploiting the natural metabolic abilities of microorganisms for the production of bioactive compounds has been a research problem of great interest. The economical and environmental costs associated with petrochemical-derived industries have promoted the emergence of biochemical processes from renewable carbon sources. However, optimally rewiring microbial metabolism in a competitive and sustainable manner is still a challenge. Recently, some retrobiosynthesis tools for the design of de novo biosynthetic pathways have been proposed. These tools generate a large number of intermediate compounds that are beyond experimental feasibility. Thus, effective methods to reduce the number of compounds by selecting the most promising ones are still needed. Here, we propose the use of classification and regression deep learning models, such as fully-connected neural networks and 1D convolutional neural networks, to predict the number of biochemical transformations needed to produce a compound. The data to train and evaluate the models was generated using a set of 13055 reaction rules and 673 compounds from Escherichia coli metabolism as starting compounds. The data was generated up to 5 steps resulting in a dataset of over 2.6 million compounds. This approach can be effectively used in biochemical applications, including retrobiosyntesis, to prioritize compounds that can be produced using fewer biochemical transformations.
João Correia 0002, Rafael Carreira, Vítor Pereira 0001, Miguel Rocha 0001
IJCNN4
2022 ProPythia: A Python package for protein classification based on machine and deep learning
Ana Marta Sequeira, Diana Lousa, Miguel Rocha 0001
Neurocomputing3
2022 A comparison of multi-objective optimization algorithms for weight setting problems in traffic engineering
Vítor Pereira 0001, Pedro Sousa 0001, Miguel Rocha 0001
Nat. Comput.3
2022 A pipeline for the reconstruction and evaluation of context-specific human metabolic models at a large-scale
abstract
Constraint-based (CB) metabolic models provide a mathematical framework and scaffold for in silico cell metabolism analysis and manipulation. In the past decade, significant efforts have been done to model human metabolism, enabled by the increased availability of multi-omics datasets and curated genome-scale reconstructions, as well as the development of several algorithms for context-specific model (CSM) reconstruction. Although CSM reconstruction has revealed insights on the deregulated metabolism of several pathologies, the process of reconstructing representative models of human tissues still lacks benchmarks and appropriate integrated software frameworks, since many tools required for this process are still disperse across various software platforms, some of which are proprietary. In this work, we address this challenge by assembling a scalable CSM reconstruction pipeline capable of integrating transcriptomics data in CB models. We combined omics preprocessing methods inspired by previous efforts with in-house implementations of existing CSM algorithms and new model refinement and validation routines, all implemented in the Troppo Python-based open-source framework. The pipeline was validated with multi-omics datasets from the Cancer Cell Line Encyclopedia (CCLE), also including reference fluxomics measurements for the MCF7 cell line. We reconstructed over 6000 models based on the Human-GEM template model for 733 cell lines featured in the CCLE, using MCF7 models as reference to find the best parameter combinations. These reference models outperform earlier studies using the same template by comparing gene essentiality and fluxomics experiments. We also analysed the heterogeneity of breast cancer cell lines, identifying key changes in metabolism related to cancer aggressiveness. Despite the many challenges in CB modelling, we demonstrate using our pipeline that combining transcriptomics data in metabolic models can be used to investigate key metabolic shifts. Significant limitations were found on these models ability for reliable quantitative flux prediction, thus motivating further work in genome-wide phenotype prediction.
Vítor Vieira, Miguel Rocha 0001
PLoS Comput. Biol.3
2021 Combining Multi-objective Evolutionary Algorithms with Deep Generative Models Towards Focused Molecular Design
João Correia 0002, Vítor Pereira 0001, Miguel Rocha 0001
EvoApplications4
2021 R2L: Routing With Reinforcement Learning
abstract
In a packet network, the routes taken by traffic can be determined according to predefined objectives. Assuming that the network conditions remain static and the defined objectives do not change, mathematical tools such as linear programming could be used to solve this routing problem. However, networks can be dynamic or the routing requirements may change. In that context, Reinforcement Learning (RL), which can learn to adapt in dynamic conditions and offers flexibility of behavior through the reward function, presents as a suitable tool to find good routing strategies. In this work, we train an RL agent, which we call R2L, to address the routing problem. The policy function used in R2L is a neural network and we use an evolution strategy algorithm to determine its weights and biases. We tested R2L in two different scenarios: static and dynamic networks conditions. In the first, we used a 16-node network and experimented with different reward functions, observing that R2L was able to adapt its routing behavior accordingly. Finally, in the second experiment, we used a 5-node network topology where a given link's transmission rate changed during the simulation. In this scenario, we observed that R2L was able to deliver a competitive performance, compared to heuristic benchmarks, with changing network conditions.
Truong Khoa Phan, Morteza Kheirkhah, David Griffin 0001, Miguel Rocha 0001, Miguel Rio
IJCNN6
2021 Deep learning for drug response prediction in cancer
abstract
Predicting the sensitivity of tumors to specific anti-cancer treatments is a challenge of paramount importance for precision medicine. Machine learning(ML) algorithms can be trained on high-throughput screening data to develop models that are able to predict the response of cancer cell lines and patients to novel drugs or drug combinations. Deep learning (DL) refers to a distinct class of ML algorithms that have achieved top-level performance in a variety of fields, including drug discovery. These types of models have unique characteristics that may make them more suitable for the complex task of modeling drug response based on both biological and chemical data, but the application of DL to drug response prediction has been unexplored until very recently. The few studies that have been published have shown promising results, and the use of DL for drug response prediction is beginning to attract greater interest from researchers in the field. In this article, we critically review recently published studies that have employed DL methods to predict drug response in cancer cell lines. We also provide a brief description of DL and the main types of architectures that have been used in these studies. Additionally, we present a selection of publicly available drug screening data resources that can be used to develop drug response prediction models. Finally, we also address the limitations of these approaches and provide a discussion on possible paths for further improvement. Contact: [email protected].
Delora Baptista, Pedro G. Ferreira, Miguel Rocha 0001
Briefings Bioinform.3
2021 MEWpy: a computational strain optimization workbench in Python
abstract
SUMMARY: Metabolic Engineering aims to favour the overproduction of native, as well as non-native, metabolites by modifying or extending the cellular processes of a specific organism. In this context, Computational Strain Optimization (CSO) plays a relevant role by putting forward mathematical approaches able to identify potential metabolic modifications to achieve the defined production goals. We present MEWpy, a Python workbench for metabolic engineering, which covers a wide range of metabolic and regulatory modelling approaches, as well as phenotype simulation and CSO algorithms. AVAILABILITY AND IMPLEMENTATION: MEWpy can be installed from PyPi (pip install mewpy), the source code being available at https://github.com/BioSystemsUM/mewpy under the GPL license.
Vítor Pereira 0001, Fernando Cruz, Miguel Rocha 0001
Bioinform.3
2020 Traffic Engineering With Three-Segments Routing
abstract
IEEE Segment Routing (SR) is a new fertile ground for Traffic Engineering (TE). By decomposing forwarding paths into segments, which specify a list of intermediate delivery points that a packet must visit on its way to the final destination, SR improves TE tasks and enables new solutions for the optimization of network resource utilization. This work proposes an Evolutionary Computation approach that enables Path Computation Element (PCE), or Software-defined Network (SDN) controllers, to optimize SR configurations for improved traffic distribution. Furthermore, we present a robust semi-oblivious method to address the variability of traffic requirements as well as alternative approaches to ensure a good network performance after link failures. In all cases, the optimization of network resource utilization is achieved using at the most three segments to configure each SR path. Moreover, all proposed optimization methods are made publicly available in a optimization framework developed by the authors.
Vítor Pereira 0001, Miguel Rocha 0001, Pedro Sousa 0001
IEEE Trans. Netw. Serv. Manag.2
2019 Deep Neural Networks for Network Routing
abstract
In this work, we propose a Deep Learning (DL) based solution to the problem of routing traffic flows in computer networks. Routing decisions can be made in different ways depending on the desired objective and, based on that objective function, optimal solutions can be computed using a variety of techniques, e.g. with mixed integer linear programming. However, determining these solutions requires solving complex optimization problems and, thus, cannot be typically done at runtime. Instead, heuristics for these problems are often created but designing them is non-trivial in many cases. The routing framework proposed here presents an alternative to the design of heuristics, whilst still achieving good performance. This is done by building a DL model trained on the optimal decisions over flows from known traffic demands. To evaluate our solution, we focused on the problem of network congestion, even though a wide range of alternative objectives could be fitted into this framework. We ran experiments using two publicly available datasets of networks with real traffic demands and showed that our solution achieves close-to-optimal network congestion values.
Miguel Rocha 0001, Truong Khoa Phan, David Griffin 0001, Franck Le, Miguel Rio
IJCNN2
2019 Triptych: Multi-objective Optimisation of Service Deployment Costs, Application Delay and Bandwidth Usage
abstract
Advanced Internet services increasingly rely on many components to implement their functionality. These composite services have three important features: they are expensive to deploy, components need to be placed intelligently close to the users to improve quality of experience and they will potentially consume significant amounts of bandwidth. This paper presents Triptych, a multi-objective optimisation framework that tries to optimise according these three dimensions to help the three main stakeholders in the Internet ecosystem: users, application providers and network providers. Triptych implements evolutionary computation approaches for this complex problem, which simultaneously optimise service deployment costs, latency-based user utility and network congestion. These algorithms provide possible operating points, bringing important tools for network managements and resource allocation. A large set of simulations under different scenarios are provided to validate the algorithms.
Miguel Rocha 0001, Truong Khoa Phan, David Griffin 0001, Miguel Rio
Networking1
2019 Predicting promoters in phage genomes using PhagePromoter
abstract
SUMMARY: The growing interest in phages as antibacterial agents has led to an increase in the number of sequenced phage genomes, increasing the need for intuitive bioinformatics tools for performing genome annotation. The identification of phage promoters is indeed the most difficult step of this process. Due to the lack of online tools for phage promoter prediction, we developed PhagePromoter, a tool for locating promoters in phage genomes, using machine learning methods. This is the first online tool for predicting promoters that uses phage promoter data and the first to identify both host and phage promoters with different motifs. AVAILABILITY AND IMPLEMENTATION: This tool was integrated in the Galaxy framework and it is available online at: https://bit.ly/2Dfebfv. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marta Sampaio, Miguel Rocha 0001, Hugo Oliveira, Oscar Dias
Bioinform.2
2019 CoBAMP: a Python framework for metabolic pathway analysis in constraint-based models
abstract
SUMMARY: CoBAMP is a modular framework for the enumeration of pathway analysis concepts, such as elementary flux modes (EFM) and minimal cut sets in genome-scale constraint-based models (CBMs) of metabolism. It currently includes the K-shortest EFM algorithm and facilitates integration with other frameworks involving reading, manipulation and analysis of CBMs. AVAILABILITY AND IMPLEMENTATION: The software is implemented in Python 3, supported on most operating systems and requires a mixed-integer linear programming optimizer supported by the optlang framework. Source-code is available at https://github.com/BioSystemsUM/cobamp.
Vítor Vieira, Miguel Rocha 0001
Bioinform.2
2019 SamPler - a novel method for selecting parameters for gene functional annotation routines
abstract
BACKGROUND: As genome sequencing projects grow rapidly, the diversity of organisms with recently assembled genome sequences peaks at an unprecedented scale, thereby highlighting the need to make gene functional annotations fast and efficient. However, the (high) quality of such annotations must be guaranteed, as this is the first indicator of the genomic potential of every organism. Automatic procedures help accelerating the annotation process, though decreasing the confidence and reliability of the outcomes. Manually curating a genome-wide annotation of genes, enzymes and transporter proteins function is a highly time-consuming, tedious and impractical task, even for the most proficient curator. Hence, a semi-automated procedure, which balances the two approaches, will increase the reliability of the annotation, while speeding up the process. In fact, a prior analysis of the annotation algorithm may leverage its performance, by manipulating its parameters, hastening the downstream processing and the manual curation of assigning functions to genes encoding proteins. RESULTS: Here SamPler, a novel strategy to select parameters for gene functional annotation routines is presented. This semi-automated method is based on the manual curation of a randomly selected set of genes/proteins. Then, in a multi-dimensional array, this sample is used to assess the automatic annotations for all possible combinations of the algorithm's parameters. These assessments allow creating an array of confusion matrices, for which several metrics are calculated (accuracy, precision and negative predictive value) and used to reach optimal values for the parameters. CONCLUSIONS: The potential of this methodology is demonstrated with four genome functional annotations performed in merlin, an in-house user-friendly computational framework for genome-scale metabolic annotation and model reconstruction. For that, SamPler was implemented as a new plugin for the merlin tool.
Fernando Cruz, Davide Lagoa, Isabel Rocha, Eugénio C. Ferreira, Miguel Rocha 0001, Oscar Dias
BMC Bioinform.6
2019 Comparison of pathway analysis and constraint-based methods for cell factory design
abstract
BACKGROUND: Computational strain optimisation methods (CSOMs) have been successfully used to exploit genome-scale metabolic models, yielding strategies useful for allowing compound overproduction in metabolic cell factories. Minimal cut sets are particularly interesting since their definition allows searching for intervention strategies that impose strong growth-coupling phenotypes, and are not subject to optimality bias when compared with simulation-based CSOMs. However, since both types of methods have different underlying principles, they also imply different ways to formulate metabolic engineering problems, posing an obstacle when comparing their outputs. RESULTS: In this work, we perform an in-depth analysis of potential strategies that can be obtained with both methods, providing a critical comparison of performance, robustness, predicted phenotypes as well as strategy structure and size. To this end, we devised a pipeline including enumeration of strategies from evolutionary algorithms (EA) and minimal cut sets (MCS), filtering and flux analysis of predicted mutants to optimize the production of succinic acid in Saccharomyces cerevisiae. We additionally attempt to generalize problem formulations for MCS enumeration within the context of growth-coupled product synthesis. Strategies from evolutionary algorithms show the best compromise between acceptable growth rates and compound overproduction. However, constrained MCSs lead to a larger variety of phenotypes with several degrees of growth-coupling with production flux. The latter have proven useful in revealing the importance, in silico, of the gamma-aminobutyric acid shunt and manipulation of cofactor pools in growth-coupled designs for succinate production, mechanisms which have also been touted as potentially useful for metabolic engineering. CONCLUSIONS: The two main groups of CSOMs are valuable for finding growth-coupled mutants. Despite the limitations in maximum growth rates and large strategy sizes, MCSs help uncover novel mechanisms for compound overproduction and thus, analyzing outputs from both methods provides a richer overview on strategies that can be potentially carried over in vivo.
Vítor Vieira, Paulo Maia, Miguel Rocha 0001, Isabel Rocha
BMC Bioinform.3
2018 Utilitarian Placement of Composite Services
abstract
The emergence of distributed clouds opens up new research challenges for service deployment. Composite services consist of multiple components, potentially located in different geographical locations, which need to be interconnected and invoked in the correct order according to the overall service work-flow. The placement of composite services over distributed cloud node locations raises new challenges for efficient deployment and management. In this paper, we design exact models of the composite service placement problems using mixed integer linear program, and compare these to solutions based on genetic algorithms. We use a utility function, based initially on latency metrics, to evaluate the quality of service (QoS) of the deployed composite service. By maximizing the utility with respect to deployment cost, our approach can provide good QoS for users while satisfying budget constraints for service providers. Based on simulations using real data-center locations and traffic demand patterns, we show that our algorithms are scalable under a range of scenarios.
Truong Khoa Phan, Miguel Rocha 0001, David Griffin 0001, Miguel Rio
IEEE Trans. Netw. Serv. Manag.2
2017 Data-driven reverse engineering of signaling pathways using ensembles of dynamic models
abstract
Despite significant efforts and remarkable progress, the inference of signaling networks from experimental data remains very challenging. The problem is particularly difficult when the objective is to obtain a dynamic model capable of predicting the effect of novel perturbations not considered during model training. The problem is ill-posed due to the nonlinear nature of these systems, the fact that only a fraction of the involved proteins and their post-translational modifications can be measured, and limitations on the technologies used for growing cells in vitro, perturbing them, and measuring their variations. As a consequence, there is a pervasive lack of identifiability. To overcome these issues, we present a methodology called SELDOM (enSEmbLe of Dynamic lOgic-based Models), which builds an ensemble of logic-based dynamic models, trains them to experimental data, and combines their individual simulations into an ensemble prediction. It also includes a model reduction step to prune spurious interactions and mitigate overfitting. SELDOM is a data-driven method, in the sense that it does not require any prior knowledge of the system: the interaction networks that act as scaffolds for the dynamic models are inferred from data using mutual information. We have tested SELDOM on a number of experimental and in silico signal transduction case-studies, including the recent HPN-DREAM breast cancer challenge. We found that its performance is highly competitive compared to state-of-the-art methods for the purpose of recovering network topology. More importantly, the utility of SELDOM goes beyond basic network inference (i.e. uncovering static interaction networks): it builds dynamic (based on ordinary differential equation) models, which can be used for mechanistic interpretations and reliable dynamic predictions in new experimental conditions (i.e. not used in the training). For this task, SELDOM's ensemble prediction is not only consistently better than predictions from individual models, but also often outperforms the state of the art represented by the methods used in the HPN-DREAM challenge.
David Henriques, Alejandro Fernández Villaverde, Miguel Rocha 0001, Julio Saez-Rodriguez, Julio R. Banga
PLoS Comput. Biol.3
2017 Genome-Wide Semi-Automated Annotation of Transporter Systems
abstract
Usually, transport reactions are added to genome-scale metabolic models (GSMMs) based on experimental data and literature. This approach does not allow associating specific genes with transport reactions, which impairs the ability of the model to predict effects of gene deletions. Novel methods for systematic genome-wide transporter functional annotation and their integration into GSMMs are therefore necessary. In this work, an automatic system to detect and classify all potential membrane transport proteins for a given genome and integrate the related reactions into GSMMs is proposed, based on the identification and classification of genes that encode transmembrane proteins. The Transport Reactions Annotation and Generation (TRIAGE) tool identifies the metabolites transported by each transmembrane protein and its transporter family. The localization of the carriers is also predicted and, consequently, their action is confined to a given membrane. The integration of the data provided by TRIAGE with highly curated models allowed the identification of new transport reactions. TRIAGE is included in the new release of merlin, a software tool previously developed by the authors, which expedites the GSMM reconstruction processes.
Oscar Dias, Daniel Gomes, Paulo Vilaça, João G. R. Cardoso, Miguel Rocha 0001, Eugénio C. Ferreira, Isabel Rocha
IEEE ACM Trans. Comput. Biol. Bioinform.5
2015 Reconstructing transcriptional Regulatory Networks using data integration and Text Mining
abstract
Transcriptional Regulatory Networks (TRNs) are powerful tool for representing several interactions that occur within a cell. Recent studies have provided information to help researchers in the tasks of building and understanding these networks. One of the major sources of information to build TRNs is biomedical literature. However, due to the rapidly increasing number of scientific papers, it is quite difficult to analyse the large amount of papers that have been published about this subject. This fact has heightened the importance of Biomedical Text Mining approaches in this task. Also, owing to the lack of adequate standards, as the number of databases increases, several inconsistencies concerning gene and protein names and identifiers are common. In this work, we developed an integrated approach for the reconstruction of TRNs that retrieve the relevant information from important biological databases and insert it into a unique repository, named KREN. Also, we applied text mining techniques over this integrated repository to build TRNs. However, was necessary to create a dictionary of names and synonyms associated with these entities and also develop an approach that retrieves all the abstracts from the related scientific papers stored on PubMed, in order to create a corpora of data about genes. Furthermore, these tasks were integrated into @Note, a software system that allows to use some methods from the Biomedical Text Mining field, including an algorithms for Named Entity Recognition (NER), extraction of all relevant terms from publication abstracts, extraction relationships between biological entities (genes, proteins and transcription factors). And finally, extended this tool to allow the reconstruction Transcriptional Regulatory Networks through using scientific literature.
Rafael T. Pereira, Hugo Costa, Sónia Carneiro, Miguel Rocha 0001, Rui Mendes 0001
BIBM4
2015 Comparison of Single and Multi-objective Evolutionary Algorithms for Robust Link-State Routing
Vítor Pereira 0001, Pedro Sousa 0001, Paulo Cortez 0001, Miguel Rio, Miguel Rocha 0001
EMO (2)5
2015 Reverse engineering of logic-based differential equation models using a mixed-integer dynamic optimization approach
abstract
MOTIVATION: Systems biology models can be used to test new hypotheses formulated on the basis of previous knowledge or new experimental data, contradictory with a previously existing model. New hypotheses often come in the shape of a set of possible regulatory mechanisms. This search is usually not limited to finding a single regulation link, but rather a combination of links subject to great uncertainty or no information about the kinetic parameters. RESULTS: In this work, we combine a logic-based formalism, to describe all the possible regulatory structures for a given dynamic model of a pathway, with mixed-integer dynamic optimization (MIDO). This framework aims to simultaneously identify the regulatory structure (represented by binary parameters) and the real-valued parameters that are consistent with the available experimental data, resulting in a logic-based differential equation model. The alternative to this would be to perform real-valued parameter estimation for each possible model structure, which is not tractable for models of the size presented in this work. The performance of the method presented here is illustrated with several case studies: a synthetic pathway problem of signaling regulation, a two-component signal transduction pathway in bacterial homeostasis, and a signaling network in liver cancer cells. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected] or [email protected].
David Henriques, Miguel Rocha 0001, Julio Saez-Rodriguez, Julio R. Banga
Bioinform.2
2014 Genome-scale bacterial transcriptional regulatory networks: reconstruction and integrated analysis with metabolic models
abstract
Advances in sequencing technology are resulting in the rapid emergence of large numbers of complete genome sequences. High-throughput annotation and metabolic modeling of these genomes is now a reality. The high-throughput reconstruction and analysis of genome-scale transcriptional regulatory networks represent the next frontier in microbial bioinformatics. The fruition of this next frontier will depend on the integration of numerous data sources relating to mechanisms, components and behavior of the transcriptional regulatory machinery, as well as the integration of the regulatory machinery into genome-scale cellular models. Here, we review existing repositories for different types of transcriptional regulatory data, including expression data, transcription factor data and binding site locations and we explore how these data are being used for the reconstruction of new regulatory networks. From template network-based methods to de novo reverse engineering from expression data, we discuss how regulatory networks can be reconstructed and integrated with metabolic models to improve model predictions and performance. We also explore the impact these integrated models can have in simulating phenotypes, optimizing the production of compounds of interest or paving the way to a whole-cell model.
José P. Faria, Ross A. Overbeek, Fangfang Xia, Miguel Rocha 0001, Isabel Rocha, Christopher S. Henry
Briefings Bioinform.4
2014 D-Tailor: automated analysis and design of DNA sequences
abstract
MOTIVATION: Current advances in DNA synthesis, cloning and sequencing technologies afford high-throughput implementation of artificial sequences into living cells. However, flexible computational tools for multi-objective sequence design are lacking, limiting the potential of these technologies. RESULTS: We developed DNA-Tailor (D-Tailor), a fully extendable software framework, for property-based design of synthetic DNA sequences. D-Tailor permits the seamless integration of multiple sequence analysis tools into a generic Monte Carlo simulation that evolves sequences toward any combination of rationally defined properties. As proof of principle, we show that D-Tailor is capable of designing sequence libraries comprising all possible combinations among three different sequence properties influencing translation efficiency in Escherichia coli The capacity to design artificial sequences that systematically sample any given parameter space should support the implementation of more rigorous experimental designs. AVAILABILITY: Source code is available for download at https://sourceforge.net/projects/dtailor/ CONTACT: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online (D-Tailor Tutorial).
Joao C. Guimaraes, Miguel Rocha 0001, Adam P. Arkin, Guillaume Cambray
Bioinform.2
2014 An integrated network visualization framework towards metabolic engineering applications
abstract
BACKGROUND: Over the last years, several methods for the phenotype simulation of microorganisms, under specified genetic and environmental conditions have been proposed, in the context of Metabolic Engineering (ME). These methods provided insight on the functioning of microbial metabolism and played a key role in the design of genetic modifications that can lead to strains of industrial interest. On the other hand, in the context of Systems Biology research, biological network visualization has reinforced its role as a core tool in understanding biological processes. However, it has been scarcely used to foster ME related methods, in spite of the acknowledged potential. RESULTS: In this work, an open-source software that aims to fill the gap between ME and metabolic network visualization is proposed, in the form of a plugin to the OptFlux ME platform. The framework is based on an abstract layer, where the network is represented as a bipartite graph containing minimal information about the underlying entities and their desired relative placement. The framework provides input/output support for networks specified in standard formats, such as XGMML, SBGN or SBML, providing a connection to genome-scale metabolic models. An user-interface makes it possible to edit, manipulate and query nodes in the network, providing tools to visualize diverse effects, including visual filters and aspect changing (e.g. colors, shapes and sizes). These tools are particularly interesting for ME, since they allow overlaying phenotype simulation results or elementary flux modes over the networks. CONCLUSIONS: The framework and its source code are freely available, together with documentation and other resources, being illustrated with well documented case studies.
Alberto Noronha, Paulo Vilaça, Miguel Rocha 0001
BMC Bioinform.3
2014 Optimization of fed-batch fermentation processes with bio-inspired algorithms
Miguel Rocha 0001, Rui Mendes 0001, Orlando Rocha, Isabel Rocha, Eugénio C. Ferreira
Expert Syst. Appl.1
2013 Algorithms to infer metabolic flux ratios from fluxomics data
abstract
In silico cell simulation approaches based in the use of genome-scale metabolic models (GSMMs) and constraint-based methods such as Flux Balance Analysis are gaining importance, but methods to integrate these approaches with omics data are still greatly needed. In this work, the focus relies on fluxomics data that provide valuable information on the intracellular fluxes, although in many cases in an indirect, incomplete and noisy way. The proposed framework enables the integration of fluxomics data, in the form of13C labeling distribution for metabolite fragments, with GSMMs enriched with carbon atom transition maps. The algorithms implemented allow to infer labeling distributions for fragments/metabolites not measured and to build expressions for the relevant flux ratios that can be then used to enrich constraint-based methods for flux determination. This approach does not require any assumptions on the metabolic network and reaction reversibility, allowing to compute ratios originating from coupled joint points of the network. Also, when enough data do not exist, the system tries to infer ratio bounds from the measurements.
Rafael Carreira, Miguel Rocha 0001, Silas Granato Villas-Bôas, Isabel Rocha
BIBM2
2013 Evolutionary computation for predicting optimal reaction knockouts and enzyme modulation strategies
abstract
One of the main purposes of Metabolic Engineering is the quantitative prediction of cell behaviour under selected genetic modifications. These methods can then be used to support adequate strain optimization algorithms in a outer layer. The purpose of the present study is to explore methods in which dynamical models provide for phenotype simulation methods, that will be used as a basis for strain optimization algorithms to indicate enzyme under/over expression or deletion of a few reactions as to maximize the production of compounds with industrial interest. This work details the developed optimization algorithms, based on Evolutionary Computation approaches, to enhance the production of a target metabolite by finding an adequate set of reaction deletions or by changing the levels of expression of a set of enzymes. To properly evaluate the strains, the ratio of the flux value associated with the target metabolite divided by the wild-type counterpart was employed as a fitness function. The devised algorithms were applied to the maximization of Serine production by Escherichia coli, using a dynamic kinetic model of the central carbon metabolism. In this case study, the proposed algorithms reached a set of solutions with higher quality, as compared to the ones described in the literature using distinct optimization techniques.
Pedro Evangelista, Miguel Rocha 0001, Isabel Rocha
IEEE Congress on Evolutionary Computation2
2013 An integrated framework for strain optimization
abstract
The identification of genetic modifications leading to mutant strains able to overproduce compounds of industrial interest is a challenging task in Metabolic Engineering (ME). Several methods have been proposed but, to some extent, none of them is suitable for all the specificities of each particular strain optimization problem. This work proposes an integrated framework that allows its users to configure and fine tune all the various steps involved in a strain optimization strategy, including the loading of models in distinct formats, the definition of a suitable phenotype simulation method and the choice and configuration of the strain optimization engine. Moreover, it is designed to suit the needs of users skilled at programming, as well as less advanced users. The framework includes a GUI implemented as the strain optimization plug-in for the OptFlux workbench (version 3), a reference platform for ME (http://www.optflux.org). All the code is distributed under the GPLv3 licence and it is fully available (http://sourceforge.net/projects/optflux/).
Paulo Maia, Isabel Rocha, Miguel Rocha 0001
IEEE Congress on Evolutionary Computation3
2012 Evolutionary Symbiotic Feature Selection for Email Spam Detection
Paulo Cortez 0001, Rui Vaz, Miguel Rocha 0001, Miguel Rio, Pedro Sousa 0001
ICINCO (1)3
2012 Multi-scale Internet traffic forecasting using neural networks and time series methods
abstract
Abstract This article presents three methods to forecast accurately the amount of traffic in TCP/IP based networks: a novel neural network ensemble approach and two important adapted time series methods (ARIMA and Holt‐Winters). In order to assess their accuracy, several experiments were held using real‐world data from two large Internet service providers. In addition, different time scales (5 min, 1 h and 1 day) and distinct forecasting lookaheads were analysed. The experiments with the neural ensemble achieved the best results for 5 min and hourly data, while the Holt‐Winters is the best option for the daily forecasts. This research opens possibilities for the development of more efficient traffic engineering and anomaly detection tools, which will result in financial gains from better network resource management.
Paulo Cortez 0001, Miguel Rio, Miguel Rocha 0001, Pedro Sousa 0001
Expert Syst. J. Knowl. Eng.3
2011 Multiobjective Evolutionary Algorithms for intradomain routing optimization
abstract
Evolutionary Algorithms (EAs) have been used to develop methods for Traffic Engineering (TE) over IP-based networks in the last few years, being used to reach the best set of link weights in the configuration of intra-domain routing protocols, such as OSPF. In this work, the multiobjective nature of a class of optimization problems provided by TE with Quality of Service constraints is identified. Multiobjective EAs (MOEAs) are developed to tackle these tasks and their results are compared to previous approaches using single objective EAs. The effect of distinct genetic representations within the MOEAs is also explored. The results show that the MOEAs provide more flexible solutions for network management, but are in some cases unable to reach the level of quality obtained by single objective EAs. Furthermore, a freely available software application is described that allows the use of the mentioned optimization algorithms by network administrators, in an user-friendly way by providing adequate user interfaces for the main TE tasks.
Miguel Rocha 0001, Tiago Sá, Pedro Nuno Miranda de Sousa, Paulo Cortez 0001, Miguel Rio
IEEE Congress on Evolutionary Computation1
2011 Challenges in integrating Escherichia coli molecular biology data
abstract
One key challenge in Systems Biology is to provide mechanisms to collect and integrate the necessary data to be able to meet multiple analysis requirements. Typically, biological contents are scattered over multiple data sources and there is no easy way of comparing heterogeneous data contents. This work discusses ongoing standardisation and interoperability efforts and exposes integration challenges for the model organism Escherichia coli K-12. The goal is to analyse the major obstacles faced by integration processes, suggest ways to systematically identify them, and whenever possible, propose solutions or means to assist manual curation. Integration of gene, protein and compound data was evaluated by performing comparisons over EcoCyc, KEGG, BRENDA, ChEBI, Entrez Gene and UniProt contents. Cross-links, a number of standard nomenclatures and name information supported the comparisons. Except for the gene integration scenario, in no other scenario an element of integration performed well enough to support the process by itself. Indeed, both the integration of enzyme and compound records imply considerable curation. Results evidenced that, even for a well-studied model organism, source contents are still far from being as standardized as it would be desired and metadata varies considerably from source to source. Before designing any data integration pipeline, researchers should decide on the sources that best fit the purpose of analysis and be aware of existing conflicts/inconsistencies to be able to intervene in their resolution. Moreover, they should be aware of the limits of automatic integration such that they can define the extent of necessary manual curation for each application.
Anália Lourenço, Sónia Carneiro, Miguel Rocha 0001, Eugénio C. Ferreira, Isabel Rocha
Briefings Bioinform.3
2011 Semantic annotation of biological concepts interplaying microbial cellular responses
abstract
BACKGROUND: Automated extraction systems have become a time saving necessity in Systems Biology. Considerable human effort is needed to model, analyse and simulate biological networks. Thus, one of the challenges posed to Biomedical Text Mining tools is that of learning to recognise a wide variety of biological concepts with different functional roles to assist in these processes. RESULTS: Here, we present a novel corpus concerning the integrated cellular responses to nutrient starvation in the model-organism Escherichia coli. Our corpus is a unique resource in that it annotates biomedical concepts that play a functional role in expression, regulation and metabolism. Namely, it includes annotations for genetic information carriers (genes and DNA, RNA molecules), proteins (transcription factors, enzymes and transporters), small metabolites, physiological states and laboratory techniques. The corpus consists of 130 full-text papers with a total of 59043 annotations for 3649 different biomedical concepts; the two dominant classes are genes (highest number of unique concepts) and compounds (most frequently annotated concepts), whereas other important cellular concepts such as proteins account for no more than 10% of the annotated concepts. CONCLUSIONS: To the best of our knowledge, a corpus that details such a wide range of biological concepts has never been presented to the text mining community. The inter-annotator agreement statistics provide evidence of the importance of a consolidated background when dealing with such complex descriptions, the ambiguities naturally arising from the terminology and their impact for modelling purposes.Availability is granted for the full-text corpora of 130 freely accessible documents, the annotation scheme and the annotation guidelines. Also, we include a corpus of 340 abstracts.
Rafael Carreira, Sónia Carneiro, Rui Pereira, Miguel Rocha 0001, Isabel Rocha, Eugénio C. Ferreira, Anália Lourenço
BMC Bioinform.4
2011 Symbiotic filtering for spam email detection
Clotilde Lopes, Paulo Cortez 0001, Pedro Nuno Miranda de Sousa, Miguel Rocha 0001, Miguel Rio
Expert Syst. Appl.4
2010 Pluggable Parallelization of Evolutionary Algorithms Applied to the Optimization of Biological Processes
abstract
Current wide availability of multicore systems requires tools that can help scientists to smoothly update their applications to take advantage of the parallel processing capabilities of these systems. In this paper, we present an experience with aspect-oriented programming (AOP) techniques to perform this move. We describe the parallelization of a Java library that implements algorithms from the Evolutionary Computation field (JECoLi), applied to two case studies in Bioinformatics, namely the optimization of feeding profiles in fed-batch fermentations and in silico strain optimization in Metabolic Engineering. AOP allowed us to enable the library to take advantage of multicore systems with minimal impact on the original code and to simultaneously develop the parallelization and the original library. Moreover, we developed modules that extend the library’s behavior for a better usage of multicore resources. Performance results show that this approach boosts performance, does not compromise the quality of the final solutions and enables a more loosely coupled development.
Jorge Henrique Martins de Pinho, Miguel Rocha 0001, João Luís Ferreira Sobral
PDP2
2010 BioDR: Semantic indexing networks for biomedical document retrieval
abstract
In Biomedical research, retrieving documents that match an interesting query is a task performed quite frequently. Typically, the set of obtained results is extensive containing many non-interesting documents and consists in a flat list, i.e., not organized or indexed in any way. This work proposes BioDR, a novel approach that allows the semantic indexing of the results of a query, by identifying relevant terms in the documents. These terms emerge from a process of Named Entity Recognition that annotates occurrences of biological terms (e.g. genes or proteins) in abstracts or full-texts. The system is based on a learning process that builds an Enhanced Instance Retrieval Network (EIRN) from a set of manually classified documents, regarding their relevance to a given problem. The resulting EIRN implements the semantic indexing of documents and terms, allowing for enhanced navigation and visualization tools, as well as the assessment of relevance for new documents.
Anália Lourenço, Rafael Carreira, Daniel Glez-Peña, José Ramón Méndez 0001, Sónia Carneiro, Luis M. Rocha, Fernando Díaz 0001, Eugénio C. Ferreira, Isabel Rocha, Florentino Fernández Riverola, Miguel Rocha 0001
Expert Syst. Appl.11
2009 Implementing Metaheuristic Optimization Algorithms with JECoLi
abstract
This work proposes JECoLi-a novel Java-based library for the implementation of metaheuristic optimization algorithms with a focus on Genetic and Evolutionary Computation based methods. The library was developed based on the principles of flexibility, usability, adaptability, modularity, extensibility, transparency, scalability, robustness and computational efficiency. The project is open-source, so JECoLi is made available under the GPL license, together with extensive documentation and examples, all included in a community Wiki-based web site (http://darwin.di.uminho.pt/jecoli). JECoLi has been/is being used in several research projects that helped to shape its evolution, ranging application fields from Bioinformatics, to Data Mining and Computer Network optimization.
Pedro Evangelista, Paulo Maia, Miguel Rocha 0001
ISDA3
2009 Symbiotic Data Mining for Personalized Spam Filtering
abstract
Unsolicited e-mail (spam) is a severe problem due to intrusion of privacy, online fraud, viruses and time spent reading unwanted messages. To solve this issue, Collaborative Filtering (CF) and Content-Based Filtering (CBF) solutions have been adopted. We propose a new CBF-CF hybrid approach called Symbiotic Data Mining (SDM), which aims at aggregating distinct local filters in order to improve filtering at a personalized level using collaboration while preserving privacy. We apply SDM to spam e-mail detection and compare it with a local CBF filter (i.e. Naive Bayes). Several experiments were conducted by using a novel corpus based on the well known Enron datasets mixed with recent spam. The results show that the symbiotic strategy is competitive in performance when compared to CBF and also more robust to contamination attacks.
Paulo Cortez 0001, Clotilde Lopes, Pedro Nuno Miranda de Sousa, Miguel Rocha 0001, Miguel Rio
Web Intelligence4
2009 @Note: A workbench for Biomedical Text Mining
abstract
Biomedical Text Mining (BioTM) is providing valuable approaches to the automated curation of scientific literature. However, most efforts have addressed the benchmarking of new algorithms rather than user operational needs. Bridging the gap between BioTM researchers and biologists' needs is crucial to solve real-world problems and promote further research. We present @Note, a platform for BioTM that aims at the effective translation of the advances between three distinct classes of users: biologists, text miners and software developers. Its main functional contributions are the ability to process abstracts and full-texts; an information retrieval module enabling PubMed search and journal crawling; a pre-processing module with PDF-to-text conversion, tokenisation and stopword removal; a semantic annotation schema; a lexicon-based annotator; a user-friendly annotation view that allows to correct annotations and a Text Mining Module supporting dataset preparation and algorithm evaluation. @Note improves the interoperability, modularity and flexibility when integrating in-home and open-source third-party components. Its component-based architecture allows the rapid development of new applications, emphasizing the principles of transparency and simplicity of use. Although it is still on-going, it has already allowed the development of applications that are currently being used.
Anália Lourenço, Rafael Carreira, Sónia Carneiro, Paulo Maia, Daniel Glez-Peña, Florentino Fernández Riverola, Eugénio C. Ferreira, Isabel Rocha, Miguel Rocha 0001
J. Biomed. Informatics9
2008 A framework for the development of Biomedical Text Mining software tools
abstract
Over the last few years, a growing number of techniques has been successfully proposed to tackle diverse challenges in the biomedical text mining (BioTM) arena. However, the set of available software tools to researchers has not grown in a similar way. This work makes a contribution to close this gap, proposing a framework to ease the development of user-friendly and interoperable applications in this field, based on a set of available modular components. These modules can be connected in diverse ways to create applications that fit distinct user roles. Also, developers of new algorithms have a framework that allows them to easily integrate their implementations with state-of-the-art BioTM software for related tasks.
Anália Lourenço, Rafael Carreira, Sónia Carneiro, Paulo Maia, Daniel Glez-Peña, Florentino Fernández Riverola, Eugénio C. Ferreira, Isabel Rocha, Miguel Rocha 0001
BIBE9
2008 Evaluating evolutionary multiobjective algorithms for the in silico optimization of mutant strains
abstract
In Metabolic Engineering, the identification of genetic manipulations that lead to mutant strains able to produce a given compound of interest is a promising, while still complex process. Evolutionary Algorithms (EAs) have been a successful approach for tackling the underlying in silico optimization problems. The most common task is to solve a bi-level optimization problem, where the strain that maximizes the production of some compound is sought, while trying to keep the organism viable (maximizing biomass). In this work, this task is viewed as a multiobjective optimization problem and an approach based on multiobjective EAs is proposed. The algorithms are validated with a real world case study that uses E. coli to produce succinic acid. The results obtained are quite promising when compared to the available single objective algorithms.
Paulo Maia, Isabel Rocha, Eugénio C. Ferreira, Miguel Rocha 0001
BIBE4
2008 A framework for the integrated analysis of metabolic and regulatory networks
abstract
The analysis of cellular behavior and functionality is the most challenging aim of systems biology. The extensive analysis of the interactions between different classes of intra-cellular molecules reacting to genetic/environment changes can elucidate the mechanisms of regulation involved on different cellular processes. We propose a novel framework that enables the integrated analysis of metabolic and regulatory networks. The framework takes advantage on publicly available data repositories, sustaining the inference of knowledge from the integrated network. Since it is based on logic programming, it provides users with a powerful language to query information using both first and second order predicates. Also, it supports network topology analysis, motif finding and robustness evaluation. In this work, as an illustrative case study, our framework is used to build a model of the bacterium Escherichia coli K12.
Rui Mendes 0001, Anália Lourenço, Sónia Carneiro, Miguel Rocha 0001, Isabel Rocha, Eugénio C. Ferreira
BIBE4
2008 Natural computation meta-heuristics for the in silico optimization of microbial strains
abstract
BACKGROUND: One of the greatest challenges in Metabolic Engineering is to develop quantitative models and algorithms to identify a set of genetic manipulations that will result in a microbial strain with a desirable metabolic phenotype which typically means having a high yield/productivity. This challenge is not only due to the inherent complexity of the metabolic and regulatory networks, but also to the lack of appropriate modelling and optimization tools. To this end, Evolutionary Algorithms (EAs) have been proposed for in silico metabolic engineering, for example, to identify sets of gene deletions towards maximization of a desired physiological objective function. In this approach, each mutant strain is evaluated by resorting to the simulation of its phenotype using the Flux-Balance Analysis (FBA) approach, together with the premise that microorganisms have maximized their growth along natural evolution. RESULTS: This work reports on improved EAs, as well as novel Simulated Annealing (SA) algorithms to address the task of in silico metabolic engineering. Both approaches use a variable size set-based representation, thereby allowing the automatic finding of the best number of gene deletions necessary for achieving a given productivity goal. The work presents extensive computational experiments, involving four case studies that consider the production of succinic and lactic acid as the targets, by using S. cerevisiae and E. coli as model organisms. The proposed algorithms are able to reach optimal/near-optimal solutions regarding the production of the desired compounds and presenting low variability among the several runs. CONCLUSION: The results show that the proposed SA and EA both perform well in the optimization task. A comparison between them is favourable to the SA in terms of consistency in obtaining optimal solutions and faster convergence. In both cases, the use of variable size representations allows the automatic discovery of the approximate number of gene deletions, without compromising the optimality of the solutions.
Miguel Rocha 0001, Paulo Maia, Rui Mendes 0001, José P. Pinto, Eugénio C. Ferreira, Jens Nielsen, Kiran Raosaheb Patil, Isabel Rocha
BMC Bioinform.1
2007 Optimization of Bacterial Strains with Variable-Sized Evolutionary Algorithms
abstract
In metabolic engineering it is difficult to identify which set of genetic manipulations will result in a microbial strain that achieves a desired production goal, due to the complexity of the metabolic and regulatory cellular networks and to the lack of appropriate modeling and optimization tools. In this work, evolutionary algorithms (EAs) are proposed for the optimization of the set of gene deletions to apply to a microorganism, in order to maximize a given objective function. Each mutant strain is evaluated by resorting to the simulation of its phenotype using the flux-balance analysis approach, together with the premise that microorganisms have maximized their growth along natural evolution. A new set based representation is used in the EAs, using variable size chromosomes, allowing for the automatic discovery of the optimal number of gene deletions. This approach was compared with a traditional binary-based genetic algorithm. Two case studies are presented considering the production of succinic and lactic acid as the target, with the bacterium E. coli. The variable size EAs, outperformed the other approaches tested, allowing to reach good results regarding the production of the desired compounds, and additionally presenting low variability among the several runs
Miguel Rocha 0001, José P. Pinto, Isabel Rocha, Eugénio C. Ferreira
CIBCB1
2007 A platform for the selection of genes in DNA microarraydata using evolutionary algorithms
abstract
This paper presents a flexible framework to the task of featureselection in classification of DNA microarray data. Theuser can select a number of filter methods in the preprocessingstage and choose from a wide set of classifiers (models and algorithms from WEKA [17] are available) and accuracy estimation methods. This approach implements wrapper methods, where Evolutionary Algorithms, with variable sized set based representations are used to reduce the number of attributes. Two case studies were used to validate the approach, with three distinct classifiers (1-nearest neighbour, decision trees, SVMs), a filter method based on discriminant fuzzy patterns and k-fold cross-validation to estimate the generalization error.
Miguel Rocha 0001, Rui Mendes 0001, Paulo Maia, Daniel Glez-Peña, Florentino Fernández Riverola
GECCO1
2007 Topology Aware Internet Traffic Forecasting Using Neural Networks
Paulo Cortez 0001, Miguel Rio, Pedro Nuno Miranda de Sousa, Miguel Rocha 0001
ICANN (2)4
2007 Evolution of neural networks for classification and regression
Miguel Rocha 0001, Paulo Cortez 0001, José Neves 0001
Neurocomputing1
2006 A Comparison of Algorithms for the Optimization of Fermentation Processes
abstract
The optimization of biotechnological processes is a complex problem that has been intensively studied in the past few years due to the economic impact of the products obtained from fermentations. In fed-batch processes, the goal is to find the optimal feeding trajectory that maximizes the final productivity. Several methods, including evolutionary algorithms (EAs) have been applied to this task in a number of different fermentation processes. This paper performs an experimental comparison between particle swarm optimization, differential evolution and a real-valued EA in three distinct case studies, taken from previous work by the authors and literature, all considering the optimization of fed-batch fermentation processes.
Rui Mendes 0001, Isabel Rocha, Eugénio C. Ferreira, Miguel Rocha 0001
IEEE Congress on Evolutionary Computation4
2006 QoS Constrained Internet Routing with Evolutionary Algorithms
abstract
OSPFOSPF is the most common intra-domain routing protocol in Wide Area Networks. Thus, optimizing OSPF weights will produce tools for traffic engineering with Quality of Service constraints, without changing the network management model. Evolutionary Algorithms (EAs) provide a valuable tool to face this NP-hard problem, allowing flexible cost functions with several metrics of the network behavior. A novel framework is proposed that enriches current models for network congestion with delay constraints, setting the basis for EAs that allocate OSPF weights, guided by a bi-objective cost function. The results show that EAs make an efficient method, outperforming common heuristics and achieving effective network behavior under unfavorable scenarios.
Miguel Rocha 0001, Pedro Sousa 0001, Miguel Rio, Paulo Cortez 0001
IEEE Congress on Evolutionary Computation1
2006 Internet Traffic Forecasting using Neural Networks
abstract
The forecast of Internet traffic is an important issue that has received few attention from the computer networks field. By improving this task, efficient traffic engineering and anomaly detection tools can be created, resulting in economic gains from better resource management. This paper presents a neural network ensemble (NNE) for the prediction of TCP/IP traffic using a time series forecasting (TSF) point of view. Several experiments were devised by considering real-world data from two large Internet Service Providers. In addition, different time scales (e.g. every five minutes and hourly) and forecasting horizons were analyzed. Overall, the NNE approach is competitive when compared with other TSF methods (e.g. Holt-Winters and ARIMA).
Paulo Cortez 0001, Miguel Rio, Miguel Rocha 0001, Pedro Nuno Miranda de Sousa
IJCNN3
2005 A new representation in evolutionary algorithms for the optimization of bioprocesses
abstract
Evolutionary algorithms (EAs) have been used to achieve optimal feedforward control in a number of fed-batch fermentation processes. Typically, the optimization purpose is to set the optimal feeding trajectory, being the feeding profile over time given by a piecewise linear function, in order to reduce the number of parameters to the optimization algorithm. In this work, a novel representation scheme for the encoding of the feeding trajectory over time is proposed. Each gene in the variable sized chromosome has two components: a time label and the real value of the variable. The new approach is compared with a traditional real-valued EA, with chromosomes of constant size and fixed discretization steps. Three distinct case studies are presented, taken from previous work from the authors and literature, all considering the optimization of fed-batch fermentation processes. The experimental results show that the proposed approach is capable of results better or at the same level of quality of the best traditional EAs and is able to automatically evolve the best discretization steps for each case, thus simplifying the EA's setup
Miguel Rocha 0001, Isabel Rocha, Eugénio C. Ferreira
Congress on Evolutionary Computation1
2003 Adaptive Learning in Changing Environments
Miguel Rocha 0001, Paulo Cortez 0001, José Neves 0001
ESANN1
2002 A Lamarckian Approach for Neural Network Training
Paulo Cortez 0001, Miguel Rocha 0001, José Neves 0001
Neural Process. Lett.2
2001 Sitting guests at a wedding party: experiments on genetic and evolutionary constrained optimization
abstract
The complex task of giving out tables to guests, according to their preferences, at a wedding party, instantiates a broader class of clustering problems, whose purpose is to group a number of entities into a number of clusters, according to a set of hard constraints, and optimizing an objective function. In order to study the application of genetic and evolutionary algorithms (GEAs) to these class of problems, some experiments were conducted. These contemplated different approaches to constraint handling, namely the use of penalty functions and decoders. The encoding issue was also studied, being compared direct and indirect representations of the problem's solutions in the chromosomes. The development of hybrid genetic operators, that combine the synergies of the GEAs paradigm with those of problem dependent heuristics, were also taken into account. The overall result is a study on the performance of several approaches to constrained optimization by GEAs, that can be used to guide the application of the paradigm in real-world problems, in the combinatorial optimization arena.
Miguel Rocha 0001, Rui Mendes 0001, Paulo Cortez 0001, José Neves 0001
CEC1
2001 Lamarckian training of feedforward neural networks
Paulo Cortez 0001, Miguel Rocha 0001, José Neves 0001
ESANN2
2001 Genetic and Evolutionary Algorithms for Time Series Forecasting
Paulo Cortez 0001, Miguel Rocha 0001, José Neves 0001
IEA/AIE2
2001 A Genetic and Evolutionary Programming Environment with Spatially Structured Populations and Built-In Parallelism
Miguel Rocha 0001, Sónia Afonso, José Neves 0001
IEA/AIE1
2000 A Study of Order Based Genetic and Evolutionary Algorithms in Combinatorial Optimization Problems
Miguel Rocha 0001, Carla Vilela, José Neves 0001
IEA/AIE1
1999 Preventing Premature Convergence to Local Optima in Genetic Algorithms via Random Offspring Generation
Miguel Rocha 0001, José Neves 0001
IEA/AIE1