VLDB 2026 Research / reviewers in the wild / expert
Susana Vinga
dblp:87/4443
· DBLP profile ↗
26ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-1954-5487ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A technical review of multi-omics data integration methods: from classical statistical to deep generative approachesabstractThe rapid advancement of high-throughput sequencing and other assay technologies has resulted in the generation of large and complex multi-omics datasets, offering unprecedented opportunities for advancing precision medicine. However, multi-omics data integration remains challenging due to the high-dimensionality, heterogeneity, and frequency of missing values across data types. Computational methods leveraging statistical and machine learning approaches have been developed to address these issues and uncover complex biological patterns, improving our understanding of disease mechanisms. Here, we comprehensively review state-of-the-art multi-omics integration methods with a focus on deep generative models, particularly variational autoencoders (VAEs) that have been widely used for data imputation, augmentation, and batch effect correction. We explore the technical aspects of VAE loss functions and regularisation techniques, including adversarial training, disentanglement, and contrastive learning. Moreover, we highlight recent advancements in foundation models and multimodal data integration, outlining future directions in precision medicine research. Ana Rita Baião, Zhaoxiang Cai, Rebecca C. Poulos, Phillip J. Robinson, Roger R. Reddel, Qing Zhong 0002, Susana Vinga, Emanuel J. V. Gonçalves |
Briefings Bioinform. | 7 |
| 2023 | Causal Graph Discovery for Explainable Insights on Marine Biotoxin Shellfish Contamination
Filipe Ferraz, Marta B. Lopes, Susana Rodrigues, Pedro Reis Costa, Susana Vinga, Alexandra M. Carvalho |
IDEAL | 6 |
| 2023 | Identification of biomarkers predictive of metastasis development in early-stage colorectal cancer using network-based regularizationabstractColorectal cancer (CRC) is the third most common cancer and the second most deathly worldwide. It is a very heterogeneous disease that can develop via distinct pathways where metastasis is the primary cause of death. Therefore, it is crucial to understand the molecular mechanisms underlying metastasis. RNA-sequencing is an essential tool used for studying the transcriptional landscape. However, the high-dimensionality of gene expression data makes selecting novel metastatic biomarkers problematic. To distinguish early-stage CRC patients at risk of developing metastasis from those that are not, three types of binary classification approaches were used: (1) classification methods (decision trees, linear and radial kernel support vector machines, logistic regression, and random forest) using differentially expressed genes (DEGs) as input features; (2) regularized logistic regression based on the Elastic Net penalty and the proposed iTwiner-a network-based regularizer accounting for gene correlation information; and (3) classification methods based on the genes pre-selected using regularized logistic regression. Classifiers using the DEGs as features showed similar results, with random forest showing the highest accuracy. Using regularized logistic regression on the full dataset yielded no improvement in the methods' accuracy. Further classification using the pre-selected genes found by different penalty factors, instead of the DEGs, significantly improved the accuracy of the binary classifiers. Moreover, the use of network-based correlation information (iTwiner) for gene selection produced the best classification results and the identification of more stable and robust gene sets. Some are known to be tumor suppressor genes (OPCML-IT2), to be related to resistance to cancer therapies (RAC1P3), or to be involved in several cancer processes such as genome stability (XRCC6P2), tumor growth and metastasis (MIR602) and regulation of gene transcription (NME2P2). We show that the classification of CRC patients based on pre-selected features by regularized logistic regression is a valuable alternative to using DEGs, significantly increasing the models' predictive performance. Moreover, the use of correlation-based penalization for biomarker selection stands as a promising strategy for predicting patients' groups based on RNA-seq data. Carolina Peixoto, Marta B. Lopes, Marta Martins, Sandra Casimiro, Daniel Sobral, Ana Rita Grosso, Catarina Abreu, Daniela Macedo, Ana Lúcia Costa, Helena Pais, Cecília Alvim, André Mansinho, Pedro Filipe, Pedro Marques da Costa, Afonso Fernandes, Paula Borralho, Cristina Ferreira, João Malaquias, António Quintela, Shannon Kaplan, Mahdi Golkaram, Michael Salmans, Nafeesa Khan, Raakhee Vijayaraghavan, Shile Zhang, Traci Pawlowski, Jim Godsey, Alex So, Luís Costa, Susana Vinga |
BMC Bioinform. | 31 |
| 2023 | Using Markov chains and temporal alignment to identify clinical patterns in DementiaabstractIn the healthcare sector, resorting to big data and advanced analytics is a great advantage when dealing with complex groups of patients in terms of comorbidities, representing a significant step towards personalized targeting. In this work, we focus on understanding key features and clinical pathways of patients with multimorbidity suffering from Dementia. This disease can result from many heterogeneous factors, potentially becoming more prevalent as the population ages. We present a set of methods that allow us to identify medical appointment patterns within a cohort of 1924 patients followed from January 2007 to August 2021 in Hospital da Luz (Lisbon), and to stratify patients into subgroups that exhibit similar patterns of interaction. With Markov Chains, we are able to identify the most prevailing medical appointments attended by Dementia patients, as well as recurring transitions between these. To perform patient stratification, we applied AliClu, a temporal sequence alignment algorithm for clustering longitudinal clinical data, which allowed us to successfully identify patient subgroups with similar medical appointment activity. A feature analysis per cluster obtained allows the identification of distinct patterns and characteristics. This pipeline provides a tool to identify prevailing clinical pathways of medical appointments within the dataset, as well as the most common transitions between medical specialities within Dementia patients. This methodology, alongside demographic and clinical data, has the potential to provide early signalling of the most likely clinical pathways and serve as a support tool for health providers in deciding the best course of treatment, considering a patient as a whole. Luísa Marote Costa, João Pedro Colaço, Alexandra M. Carvalho, Susana Vinga, Andreia Sofia Teixeira |
J. Biomed. Informatics | 4 |
| 2022 | Tutorial - Machine Learning and Information Theoretic Methods for Molecular Biology and MedicineabstractA short introduction to the application of informationtheoretic and machine learning methods to biomolecular and medical data is provided as the motivating material that supports special session dedicated to this topic at ESANN 2022.In particular, we highlight current developments of foundation such as interpretability and model certainty.Further, we emphasize how theoretic models provide a natural framework to deal with heterogeneous and complex data structures as frequently occurring in biomedical research. Thomas Villmann, Jonas S. Almeida, Susana Vinga |
ESANN | 4 |
| 2021 | Structured sparsity regularization for analyzing high-dimensional omics dataabstractThe development of new molecular and cell technologies is having a significant impact on the quantity of data generated nowadays. The growth of omics databases is creating a considerable potential for knowledge discovery and, concomitantly, is bringing new challenges to statistical learning and computational biology for health applications. Indeed, the high dimensionality of these data may hamper the use of traditional regression methods and parameter estimation algorithms due to the intrinsic non-identifiability of the inherent optimization problem. Regularized optimization has been rising as a promising and useful strategy to solve these ill-posed problems by imposing additional constraints in the solution parameter space. In particular, the field of statistical learning with sparsity has been significantly contributing to building accurate models that also bring interpretability to biological observations and phenomena. Beyond the now-classic elastic net, one of the best-known methods that combine lasso with ridge penalizations, we briefly overview recent literature on structured regularizers and penalty functions that have been applied in biomedical data to build parsimonious models in a variety of underlying contexts, from survival to generalized linear models. These methods include functions of $\ell _k$-norms and network-based penalties that take into account the inherent relationships between the features. The successful application to omics data illustrates the potential of sparse structured regularization for identifying disease's molecular signatures and for creating high-performance clinical decision support systems towards more personalized healthcare. Supplementary information: Supplementary data are available at Briefings in Bioinformatics online. Susana Vinga |
Briefings Bioinform. | 1 |
| 2020 | MOMO - multi-objective metabolic mixed integer optimization: application to yeast strain engineeringabstractBACKGROUND: In this paper, we explore the concept of multi-objective optimization in the field of metabolic engineering when both continuous and integer decision variables are involved in the model. In particular, we propose a multi-objective model that may be used to suggest reaction deletions that maximize and/or minimize several functions simultaneously. The applications may include, among others, the concurrent maximization of a bioproduct and of biomass, or maximization of a bioproduct while minimizing the formation of a given by-product, two common requirements in microbial metabolic engineering. RESULTS: Production of ethanol by the widely used cell factory Saccharomyces cerevisiae was adopted as a case study to demonstrate the usefulness of the proposed approach in identifying genetic manipulations that improve productivity and yield of this economically highly relevant bioproduct. We did an in vivo validation and we could show that some of the predicted deletions exhibit increased ethanol levels in comparison with the wild-type strain. CONCLUSIONS: The multi-objective programming framework we developed, called MOMO, is open-source and uses POLYSCIP (Available at http://polyscip.zib.de/). as underlying multi-objective solver. MOMO is available at http://momo-sysbio.gforge.inria.fr. Ricardo Andrade, Mahdi Doostmohammadi, João L. Santos, Marie-France Sagot, Nuno P. Mira, Susana Vinga |
BMC Bioinform. | 6 |
| 2020 | Tracking intratumoral heterogeneity in glioblastoma via regularized classification of single-cell RNA-Seq dataabstractBACKGROUND: Understanding cellular and molecular heterogeneity in glioblastoma (GBM), the most common and aggressive primary brain malignancy, is a crucial step towards the development of effective therapies. Besides the inter-patient variability, the presence of multiple cell populations within tumors calls for the need to develop modeling strategies able to extract the molecular signatures driving tumor evolution and treatment failure. With the advances in single-cell RNA Sequencing (scRNA-Seq), tumors can now be dissected at the cell level, unveiling information from their life history to their clinical implications. RESULTS: We propose a classification setting based on GBM scRNA-Seq data, through sparse logistic regression, where different cell populations (neoplastic and normal cells) are taken as classes. The goal is to identify gene features discriminating between the classes, but also those shared by different neoplastic clones. The latter will be approached via the network-based twiner regularizer to identify gene signatures shared by neoplastic cells from the tumor core and infiltrating neoplastic cells originated from the tumor periphery, as putative disease biomarkers to target multiple neoplastic clones. Our analysis is supported by the literature through the identification of several known molecular players in GBM. Moreover, the relevance of the selected genes was confirmed by their significance in the survival outcomes in bulk GBM RNA-Seq data, as well as their association with several Gene Ontology (GO) biological process terms. CONCLUSIONS: We presented a methodology intended to identify genes discriminating between GBM clones, but also those playing a similar role in different GBM neoplastic clones (including migrating cells), therefore potential targets for therapy research. Our results contribute to a deeper understanding on the genetic features behind GBM, by disclosing novel therapeutic directions accounting for GBM heterogeneity. Marta B. Lopes, Susana Vinga |
BMC Bioinform. | 2 |
| 2019 | Twiner: correlation-based regularization for identifying common cancer gene signaturesabstractBACKGROUND: Breast and prostate cancers are typical examples of hormone-dependent cancers, showing remarkable similarities at the hormone-related signaling pathways level, and exhibiting a high tropism to bone. While the identification of genes playing a specific role in each cancer type brings invaluable insights for gene therapy research by targeting disease-specific cell functions not accounted so far, identifying a common gene signature to breast and prostate cancers could unravel new targets to tackle shared hormone-dependent disease features, like bone relapse. This would potentially allow the development of new targeted therapies directed to genes regulating both cancer types, with a consequent positive impact in cancer management and health economics. RESULTS: We address the challenge of extracting gene signatures from transcriptomic data of prostate adenocarcinoma (PRAD) and breast invasive carcinoma (BRCA) samples, particularly estrogen positive (ER+), and androgen positive (AR+) triple-negative breast cancer (TNBC), using sparse logistic regression. The introduction of gene network information based on the distances between BRCA and PRAD correlation matrices is investigated, through the proposed twin networks recovery (twiner) penalty, as a strategy to ensure similarly correlated gene features in two diseases to be less penalized during the feature selection procedure. CONCLUSIONS: Our analysis led to the identification of genes that show a similar correlation pattern in BRCA and PRAD transcriptomic data, and are selected as key players in the classification of breast and prostate samples into ER+ BRCA/AR+ TNBC/PRAD tumor and normal tissues, and also associated with survival time distributions. The results obtained are supported by the literature and are expected to unveil the similarities between the diseases, disclose common disease biomarkers, and help in the definition of new strategies for more effective therapies. Marta B. Lopes, Sandra Casimiro, Susana Vinga |
BMC Bioinform. | 3 |
| 2019 | Twiner: correlation-based regularization for identifying common cancer gene signaturesabstractBACKGROUND: Breast and prostate cancers are typical examples of hormone-dependent cancers, showing remarkable similarities at the hormone-related signaling pathways level, and exhibiting a high tropism to bone. While the identification of genes playing a specific role in each cancer type brings invaluable insights for gene therapy research by targeting disease-specific cell functions not accounted so far, identifying a common gene signature to breast and prostate cancers could unravel new targets to tackle shared hormone-dependent disease features, like bone relapse. This would potentially allow the development of new targeted therapies directed to genes regulating both cancer types, with a consequent positive impact in cancer management and health economics. RESULTS: We address the challenge of extracting gene signatures from transcriptomic data of prostate adenocarcinoma (PRAD) and breast invasive carcinoma (BRCA) samples, particularly estrogen positive (ER+), and androgen positive (AR+) triple-negative breast cancer (TNBC), using sparse logistic regression. The introduction of gene network information based on the distances between BRCA and PRAD correlation matrices is investigated, through the proposed twin networks recovery (twiner) penalty, as a strategy to ensure similarly correlated gene features in two diseases to be less penalized during the feature selection procedure. CONCLUSIONS: Our analysis led to the identification of genes that show a similar correlation pattern in BRCA and PRAD transcriptomic data, and are selected as key players in the classification of breast and prostate samples into ER+ BRCA/AR+ TNBC/PRAD tumor and normal tissues, and also associated with survival time distributions. The results obtained are supported by the literature and are expected to unveil the similarities between the diseases, disclose common disease biomarkers, and help in the definition of new strategies for more effective therapies. Marta B. Lopes, Sandra Casimiro, Susana Vinga |
BMC Bioinform. | 3 |
| 2018 | Ensemble outlier detection and gene selection in triple-negative breast cancer dataabstractBACKGROUND: Learning accurate models from 'omics data is bringing many challenges due to their inherent high-dimensionality, e.g. the number of gene expression variables, and comparatively lower sample sizes, which leads to ill-posed inverse problems. Furthermore, the presence of outliers, either experimental errors or interesting abnormal clinical cases, may severely hamper a correct classification of patients and the identification of reliable biomarkers for a particular disease. We propose to address this problem through an ensemble classification setting based on distinct feature selection and modeling strategies, including logistic regression with elastic net regularization, Sparse Partial Least Squares - Discriminant Analysis (SPLS-DA) and Sparse Generalized PLS (SGPLS), coupled with an evaluation of the individuals' outlierness based on the Cook's distance. The consensus is achieved with the Rank Product statistics corrected for multiple testing, which gives a final list of sorted observations by their outlierness level. RESULTS: We applied this strategy for the classification of Triple-Negative Breast Cancer (TNBC) RNA-Seq and clinical data from the Cancer Genome Atlas (TCGA). The detected 24 outliers were identified as putative mislabeled samples, corresponding to individuals with discrepant clinical labels for the HER2 receptor, but also individuals with abnormal expression values of ER, PR and HER2, contradictory with the corresponding clinical labels, which may invalidate the initial TNBC label. Moreover, the model consensus approach leads to the selection of a set of genes that may be linked to the disease. These results are robust to a resampling approach, either by selecting a subset of patients or a subset of genes, with a significant overlap of the outlier patients identified. CONCLUSIONS: The proposed ensemble outlier detection approach constitutes a robust procedure to identify abnormal cases and consensus covariates, which may improve biomarker selection for precision medicine applications. The method can also be easily extended to other regression models and datasets. Marta B. Lopes, André Veríssimo, Eunice Carrasquinha, Sandra Casimiro, Niko Beerenwinkel, Susana Vinga |
BMC Bioinform. | 6 |
| 2016 | DegreeCox - a network-based regularization method for survival analysisabstractBACKGROUND: Modeling survival oncological data has become a major challenge as the increase in the amount of molecular information nowadays available means that the number of features greatly exceeds the number of observations. One possible solution to cope with this dimensionality problem is the use of additional constraints in the cost function optimization. LASSO and other sparsity methods have thus already been successfully applied with such idea. Although this leads to more interpretable models, these methods still do not fully profit from the relations between the features, specially when these can be represented through graphs. We propose DEGREECOX, a method that applies network-based regularizers to infer Cox proportional hazard models, when the features are genes and the outcome is patient survival. In particular, we propose to use network centrality measures to constrain the model in terms of significant genes. RESULTS: We applied DEGREECOX to three datasets of ovarian cancer carcinoma and tested several centrality measures such as weighted degree, betweenness and closeness centrality. The a priori network information was retrieved from Gene Co-Expression Networks and Gene Functional Maps. When compared with RIDGE and LASSO, DEGREECOX shows an improvement in the classification of high and low risk patients in a par with NET-COX. The use of network information is especially relevant with datasets that are not easily separated. In terms of RMSE and C-index, DEGREECOX gives results that are similar to those of the best performing methods, in a few cases slightly better. CONCLUSIONS: Network-based regularization seems a promising framework to deal with the dimensionality problem. The centrality metrics proposed can be easily expanded to accommodate other topological properties of different biological networks. André Veríssimo, Arlindo L. Oliveira, Marie-France Sagot, Susana Vinga |
BMC Bioinform. | 4 |
| 2015 | Polynomial-time algorithm for learning optimal tree-augmented dynamic Bayesian networks
José L. Monteiro, Susana Vinga, Alexandra M. Carvalho |
UAI | 2 |
| 2014 | Editorial: Alignment-free methods in computational biologyabstractAlignment-free methods for biological sequence analysis and comparison have emerged as a natural framework to address the challenges of understanding the patterns and properties of biological sequences. These methods are based on mapping symbolic sequences describing DNA, RNA and proteins, onto vector spaces, in which many of the analysis can be performed more efficiently. The rational is to represent sequences as numerical real-valued vectors and to apply available tools to this domain, which range from filtering techniques, normalization, dissimilarity estimation and clustering. Broad frameworks such as probability, statistics and linear algebra are then at hand to provide a strong and extensive theoretical and computational background. This explains the high computational efficiency of alignment-free methods, which has led in the past decades to widening the range of successful applications. The main advantage of alignment-free methods, besides the key fact they are usually computational inexpensive, is the ability of effortlessly dealing with whole genomes, thus allowing the analysis of complete sequence information. They are robust to shuffling and recombination events and generally applicable when less conservation pushes beyond what alignment could handle. It should nevertheless be noted that for the majority of the alignment-free algorithms, the symbol order is lost, which might constitute a caveat if an alignment structure is preferred. The term alignment-free was coined in a review paper a decade ago [1] that reflected on research being performed by then, which avoided the caveats of dynamic programming. That survey tried to systematize, organize and categorize a diversity of disparate methods in a common framework. It was then observed that all the main categories identified shared, in their root, the same principle and rationale of not preprocessing the sequences being compared by aligning them. It was already clear that there was a central focus on vector-valued representations using L-tuple composition, sequence representation through iterative maps, compression and entropy estimation, although with contrasting notations and generally unaware of the possibility of a unifying perspective. The applications surveyed were surprisingly diverse, ranging from phylogenetic classification and motif analysis to genomic entropy estimation. This special issue on Alignment-free methods in computational biology reflects this growing trend and offers current reviews on several areas where alignment-free methods continue to provide relevant results in computational biology and bioinformatics. The papers included cover a wide area of this theme and constitute a tentative roadmap for future developments, which both go far beyond biological sequence analysis, and also are more in line with the new Big Data challenges brought about by next-generation sequencing. The statistical aspects of metrics designed for sequence comparison are described in New developments of alignment-free sequence comparison: measures, statistics and next-generation sequencing by Song, Ren, Reinert, Deng, Waterman and Sun. Previous definitions of adequate dissimilarly measures and metrics to apply in comparison tasks led to the development of statistical descriptions and strong theoretical work, addressed in this article. The authors review statistical properties of metrics based on L-tuple matches, in particular D2 statistic, illustrating its strength for clustering purposes of next-generation sequencing short read data in metagenomic studies. The close relation between alignment-free methods and pattern recognition algorithms is reviewed in Pattern recognition and probabilistic measures in alignment-free sequence analysis by Schwende and Pham, where the authors survey the general properties of alignment-free versus alignment-based methods and overview current metrics and software to perform the general analysis in the context of machine learning. The need for alternative sequence representation has led to the striking development of iterated maps, rooted in non-linear dynamics, reviewed in the paper Sequence analysis by iterated maps, a review by Almeida. The idea of mapping sequences onto vector spaces attained a high level of formal elegance in chaos game representation. These functions allowed to smoothly bridge concepts and applications from numerical representation to graphical structures. Iterated maps have also clear connections with stochastic processes and Markov chain models, besides owning strong links with information theory concepts. The developments in the past decade suggest a particularly intriguing angle on the scalable analysis of next-generation sequencing. The alignment-free methods covered have been particularly connected to information theory (IT) concepts. The statistical and linear algebra descriptions have clear associations with IT and both have been successfully applied in computational biology. The paper Information theory applications for biological sequence analysis by Vinga reviews and categorizes IT applications under an alignment-free framework. These include the global characterization of sequences, their local analysis and also methods that combine several levels of information into a unique integrated framework. In connection with information theory concepts, the algorithmic assessment of compression methods for sequences is addressed in the paper Compressive biological sequence analysis and archival in the era of high-throughput sequencing technologies by Giancarlo, Rombo and Utro. The authors exhaustively review current compression and storage techniques, with a focus on high-throughput sequencing (HTS) data. They further provide reference databases and available software tools. On the application to evolutionary research, the paper Alignment-free phylogenetics and population genetics by Haubold reviews alignment-free methods successfully applied to phylogenetics and population genetics. The author overviews the metrics more adequate to infer phylogenetic relationships and to estimate the distribution of mutations, and illustrates them in simulated sequences and in real genomes. Finally, the paper Applications of alignment-free methods in epigenomics by Pinello, Lo Bosco and Yuan illustrates the strength of this framework in the context of linking genome to epigenome. Several machine learning techniques are overviewed and their applications highlighted, namely for nucleosome positioning, DNA methylation and histone modifications. We hope that this special issue with current reviews of alignment-free methods will support algorithm advancements in the next decade as exciting as in the 11 years since the establishment of a unifying characterization. Susana Vinga |
Briefings Bioinform. | 1 |
| 2014 | Information theory applications for biological sequence analysisabstractInformation theory (IT) addresses the analysis of communication systems and has been widely applied in molecular biology. In particular, alignment-free sequence analysis and comparison greatly benefited from concepts derived from IT, such as entropy and mutual information. This review covers several aspects of IT applications, ranging from genome global analysis and comparison, including block-entropy estimation and resolution-free metrics based on iterative maps, to local analysis, comprising the classification of motifs, prediction of transcription factor binding sites and sequence characterization based on linguistic complexity and entropic profiles. IT has also been applied to high-level correlations that combine DNA, RNA or protein features with sequence-independent properties, such as gene mapping and phenotype analysis, and has also provided models based on communication systems theory to describe information transmission channels at the cell level and also during evolutionary processes. While not exhaustive, this review attempts to categorize existing methods and to indicate their relation with broader transversal topics such as genomic signatures, data compression and complexity, time series analysis and phylogenetic classification, providing a resource for future developments in this promising area. Susana Vinga |
Briefings Bioinform. | 1 |
| 2014 | Identifying IIR filter coefficients using particle swarm optimization with application to reconstruction of missing cardiovascular signals
András Hartmann, João Miranda Lemos, Rafael S. Costa, Susana Vinga |
Eng. Appl. Artif. Intell. | 4 |
| 2013 | Prediction of Forest Aboveground Biomass: An Exercise on Avoiding Overfitting
Sara Silva, Vijay Ingalalli, Susana Vinga, João Manuel de Brito Carreiras, Joana B. Melo, Mauro Castelli, Leonardo Vanneschi, Ivo Gonçalves, José Caldas |
EvoApplications | 3 |
| 2013 | BGFit: management and automated fitting of biological growth curvesabstractBACKGROUND: Existing tools to model cell growth curves do not offer a flexible integrative approach to manage large datasets and automatically estimate parameters. Due to the increase of experimental time-series from microbiology and oncology, the need for a software that allows researchers to easily organize experimental data and simultaneously extract relevant parameters in an efficient way is crucial. RESULTS: BGFit provides a web-based unified platform, where a rich set of dynamic models can be fitted to experimental time-series data, further allowing to efficiently manage the results in a structured and hierarchical way. The data managing system allows to organize projects, experiments and measurements data and also to define teams with different editing and viewing permission. Several dynamic and algebraic models are already implemented, such as polynomial regression, Gompertz, Baranyi, Logistic and Live Cell Fraction models and the user can add easily new models thus expanding current ones. CONCLUSIONS: BGFit allows users to easily manage their data and models in an integrated way, even if they are not familiar with databases or existing computational tools for parameter estimation. BGFit is designed with a flexible architecture that focus on extensibility and leverages free software with existing tools and methods, allowing to compare and evaluate different data modeling techniques. The application is described in the context of bacterial and tumor cells growth data fitting, but it is also applicable to any type of two-dimensional data, e.g. physical chemistry and macroeconomic time series, being fully scalable to high number of projects, data and model complexity. André Veríssimo, Laura Paixão, Ana Rita Neves, Susana Vinga |
BMC Bioinform. | 4 |
| 2011 | A Survey on Methods for Modeling and Analyzing Integrated Biological NetworksabstractUnderstanding how cellular systems build up integrated responses to their dynamically changing environment is one of the open questions in Systems Biology. Despite their intertwinement, signaling networks, gene regulation and metabolism have been frequently modeled independently in the context of well-defined subsystems. For this purpose, several mathematical formalisms have been developed according to the features of each particular network under study. Nonetheless, a deeper understanding of cellular behavior requires the integration of these various systems into a model capable of capturing how they operate as an ensemble. With the recent advances in the "omics" technologies, more data is becoming available and, thus, recent efforts have been driven toward this integrated modeling approach. We herein review and discuss methodological frameworks currently available for modeling and analyzing integrated biological networks, in particular metabolic, gene regulatory and signaling networks. These include network-based methods and Chemical Organization Theory, Flux-Balance Analysis and its extensions, logical discrete modeling, Petri Nets, traditional kinetic modeling, Hybrid Systems and stochastic models. Comparisons are also established regarding data requirements, scalability with network size and computational burden. The methods are illustrated with successful case studies in large-scale genome models and in particular subsystems of various organisms. Nuno Tenazinha, Susana Vinga |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2009 | Biological sequences as pictures - a generic two dimensional solution for iterated mapsabstractBACKGROUND: Representing symbolic sequences graphically using iterated maps has enjoyed an enduring popularity since it was first proposed in Jeffrey 1990 as chaos game representation (CGR). The usefulness of this representation goes beyond the convenience of a scale independent representation. It provides a variable memory length representation of transition. This includes the representation of succession with non-integer order, which comes with the promise of generalizing Markovian formalisms. The original proposal targeted genomic sequences only but since then several generalizations have been proposed, many specifically designed to handle protein data. RESULTS: The challenge of a general solution is that of deriving a bijective transformation of symbolic sequences into bi-dimensional planes. More specifically, it requires the regular fractal nesting of polygons. A first attempt at a general solution was proposed by Fiser 1994 by using non-overlapping circles that contain the polygons. This was used as a starting point to identify a more efficient solution where the encapsulating circles can overlap without the same happening for the sequence maps which are circumscribed to fractal polygon domains. CONCLUSION: We identified the optimal inscribed packing solution for iterated maps of any Biological sequence, indeed of any symbolic sequence. The new solution maintains the prized bijective mapping property and includes the Sierpinski triangle and the CGR square as particular solutions of the more encompassing formulation. Jonas S. Almeida, Susana Vinga |
BMC Bioinform. | 2 |
| 2008 | An analysis of the positional distribution of DNA motifs in promoter regions and its biological relevanceabstractBACKGROUND: Motif finding algorithms have developed in their ability to use computationally efficient methods to detect patterns in biological sequences. However the posterior classification of the output still suffers from some limitations, which makes it difficult to assess the biological significance of the motifs found. Previous work has highlighted the existence of positional bias of motifs in the DNA sequences, which might indicate not only that the pattern is important, but also provide hints of the positions where these patterns occur preferentially. RESULTS: We propose to integrate position uniformity tests and over-representation tests to improve the accuracy of the classification of motifs. Using artificial data, we have compared three different statistical tests (Chi-Square, Kolmogorov-Smirnov and a Chi-Square bootstrap) to assess whether a given motif occurs uniformly in the promoter region of a gene. Using the test that performed better in this dataset, we proceeded to study the positional distribution of several well known cis-regulatory elements, in the promoter sequences of different organisms (S. cerevisiae, H. sapiens, D. melanogaster, E. coli and several Dicotyledons plants). The results show that position conservation is relevant for the transcriptional machinery. CONCLUSION: We conclude that many biologically relevant motifs appear heterogeneously distributed in the promoter region of genes, and therefore, that non-uniformity is a good indicator of biological relevance and can be used to complement over-representation tests commonly used. In this article we present the results obtained for the S. cerevisiae data sets. Ana C. Casimiro, Susana Vinga, Ana T. Freitas, Arlindo L. Oliveira |
BMC Bioinform. | 2 |
| 2007 | Automated smoother for the numerical decoupling of dynamics modelsabstractBACKGROUND: Structure identification of dynamic models for complex biological systems is the cornerstone of their reverse engineering. Biochemical Systems Theory (BST) offers a particularly convenient solution because its parameters are kinetic-order coefficients which directly identify the topology of the underlying network of processes. We have previously proposed a numerical decoupling procedure that allows the identification of multivariate dynamic models of complex biological processes. While described here within the context of BST, this procedure has a general applicability to signal extraction. Our original implementation relied on artificial neural networks (ANN), which caused slight, undesirable bias during the smoothing of the time courses. As an alternative, we propose here an adaptation of the Whittaker's smoother and demonstrate its role within a robust, fully automated structure identification procedure. RESULTS: In this report we propose a robust, fully automated solution for signal extraction from time series, which is the prerequisite for the efficient reverse engineering of biological systems models. The Whittaker's smoother is reformulated within the context of information theory and extended by the development of adaptive signal segmentation to account for heterogeneous noise structures. The resulting procedure can be used on arbitrary time series with a nonstationary noise process; it is illustrated here with metabolic profiles obtained from in-vivo NMR experiments. The smoothed solution that is free of parametric bias permits differentiation, which is crucial for the numerical decoupling of systems of differential equations. CONCLUSION: The method is applicable in signal extraction from time series with nonstationary noise structure and can be applied in the numerical decoupling of system of differential equations into algebraic equations, and thus constitutes a rather general tool for the reverse engineering of mechanistic model descriptions from multivariate experimental time series. Marco Vilela, Carlos Cristiano H. Borges, Susana Vinga, Ana Tereza Ribeiro de Vasconcelos, Helena Santos, Eberhard O. Voit, Jonas S. Almeida |
BMC Bioinform. | 3 |
| 2007 | Local Renyi entropic profiles of DNA sequencesabstractBACKGROUND: In a recent report the authors presented a new measure of continuous entropy for DNA sequences, which allows the estimation of their randomness level. The definition therein explored was based on the Rényi entropy of probability density estimation (pdf) using the Parzen's window method and applied to Chaos Game Representation/Universal Sequence Maps (CGR/USM). Subsequent work proposed a fractal pdf kernel as a more exact solution for the iterated map representation. This report extends the concepts of continuous entropy by defining DNA sequence entropic profiles using the new pdf estimations to refine the density estimation of motifs. RESULTS: The new methodology enables two results. On the one hand it shows that the entropic profiles are directly related with the statistical significance of motifs, allowing the study of under and over-representation of segments. On the other hand, by spanning the parameters of the kernel function it is possible to extract important information about the scale of each conserved DNA region. The computational applications, developed in Matlab m-code, the corresponding binary executables and additional material and examples are made publicly available at http://kdbio.inesc-id.pt/~svinga/ep/. CONCLUSION: The ability to detect local conservation from a scale-independent representation of symbolic sequences is particularly relevant for biological applications where conserved motifs occur in multiple, overlapping scales, with significant future applications in the recognition of foreign genomic material and inference of motif structures. Susana Vinga, Jonas S. Almeida |
BMC Bioinform. | 1 |
| 2004 | Comparative evaluation of word composition distances for the recognition of SCOP relationshipsabstractMOTIVATION: Alignment-free metrics were recently reviewed by the authors, but have not until now been object of a comparative study. This paper compares the classification accuracy of word composition metrics therein reviewed. It also presents a new distance definition between protein sequences, the W-metric, which bridges between alignment metrics, such as scores produced by the Smith-Waterman algorithm, and methods based solely in L-tuple composition, such as Euclidean distance and Information content. RESULTS: The comparative study reported here used the SCOP/ASTRAL protein structure hierarchical database and accessed the discriminant value of alternative sequence dissimilarity measures by calculating areas under the Receiver Operating Characteristic curves. Although alignment methods resulted in very good classification accuracy at family and superfamily levels, alignment-free distances, in particular Standard Euclidean Distance, are as good as alignment algorithms when sequence similarity is smaller, such as for recognition of fold or class relationships. This observation justifies its advantageous use to pre-filter homologous proteins since word statistics techniques are computed much faster than the alignment methods. AVAILABILITY: All MATLAB code used to generate the data is available upon request to the authors. Additional material available at http://bioinformatics.musc.edu/wmetric Susana Vinga, Rodrigo Gouveia-Oliveira, Jonas S. Almeida |
Bioinform. | 1 |
| 2003 | Alignment-free sequence comparison-a reviewabstractMOTIVATION: Genetic recombination and, in particular, genetic shuffling are at odds with sequence comparison by alignment, which assumes conservation of contiguity between homologous segments. A variety of theoretical foundations are being used to derive alignment-free methods that overcome this limitation. The formulation of alternative metrics for dissimilarity between sequences and their algorithmic implementations are reviewed. RESULTS: The overwhelming majority of work on alignment-free sequence has taken place in the past two decades, with most reports published in the past 5 years. Two main categories of methods have been proposed-methods based on word (oligomer) frequency, and methods that do not require resolving the sequence with fixed word length segments. The first category is based on the statistics of word frequency, on the distances defined in a Cartesian space defined by the frequency vectors, and on the information content of frequency distribution. The second category includes the use of Kolmogorov complexity and Chaos Theory. Despite their low visibility, alignment-free metrics are in fact already widely used as pre-selection filters for alignment-based querying of large applications. Recent work is furthering their usage as a scale-independent methodology that is capable of recognizing homology when loss of contiguity is beyond the possibility of alignment. AVAILABILITY: Most of the alignment-free algorithms reviewed were implemented in MATLAB code and are available at http://bioinformatics.musc.edu/resources.html Susana Vinga, Jonas S. Almeida |
Bioinform. | 1 |
| 2002 | Universal sequence map (USM) of arbitrary discrete sequencesabstractBACKGROUND: For over a decade the idea of representing biological sequences in a continuous coordinate space has maintained its appeal but not been fully realized. The basic idea is that any sequence of symbols may define trajectories in the continuous space conserving all its statistical properties. Ideally, such a representation would allow scale independent sequence analysis--without the context of fixed memory length. A simple example would consist on being able to infer the homology between two sequences solely by comparing the coordinates of any two homologous units. RESULTS: We have successfully identified such an iterative function for bijective mapping psi of discrete sequences into objects of continuous state space that enable scale-independent sequence analysis. The technique, named Universal Sequence Mapping (USM), is applicable to sequences with an arbitrary length and arbitrary number of unique units and generates a representation where map distance estimates sequence similarity. The novel USM procedure is based on earlier work by these and other authors on the properties of Chaos Game Representation (CGR). The latter enables the representation of 4 unit type sequences (like DNA) as an order free Markov chain transition table. The properties of USM are illustrated with test data and can be verified for other data by using the accompanying web-based tool:http://bioinformatics.musc.edu/~jonas/usm/. CONCLUSIONS: USM is shown to enable a statistical mechanics approach to sequence analysis. The scale independent representation frees sequence analysis from the need to assume a memory length in the investigation of syntactic rules. Jonas S. Almeida, Susana Vinga |
BMC Bioinform. | 2 |