Sara C. Madeira

dblp:90/3540 · also Sara Cordeiro Madeira · DBLP profile ↗
← Back
43ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-1459-8096ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 9 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Software engineering, systems software and programming languages · 2Theory of computation · 1
YearPublicationVenuePosition
2026 PatientFlow: Learning to generate mixed-type longitudinal clinical data with flow matching
abstract
Synthetic longitudinal clinical data, with static and temporal mixed-type components, can help unlock large-scale deep learning models to tackle complex diseases. However, learning to generate realistic patients faces dual challenges: modeling the inherently complex structure of longitudinal data and protecting patient privacy. We introduce PatientFlow, a generative modeling method combining Variational Autoencoders for data representation with Flow Matching for patient generation. We extensively evaluated the generative model on a longitudinal cohort of patients with Amyotrophic Lateral Sclerosis (N = 1560) using both qualitative and quantitative methods. The ability of the method to generate realistic patient data, further validated by expert clinicians, shows its potential application to other diseases. Prognostic models trained on synthetic data across five clinically relevant endpoints matched and sometimes outperformed the models trained on real data. Our results demonstrate that PatientFlow can effectively model longitudinal clinical data with high fidelity, opening promising avenues for sharing and augmenting datasets for deep learning applications in healthcare without compromising privacy.
Ruben Branco, Marta Gromicho, Mamede de Carvalho, Piero Fariselli, Sara C. Madeira
Artif. Intell. Medicine5
2026 Cutting through the noise: Explaining residuals in multivariate time series with motif analysis
abstract
Modeling real-world system dynamics is challenging due to non-periodic patterns, such as stimuli-dependent physiological responses in health, event-driven traffic in mobility, and news-triggered interactions in societal systems. In the absence of contextual data, state-of-the-art methods—including advanced deep learning architectures—struggle to model these irregular behaviors. Furthermore, their predictive focus often limits their utility for descriptive analysis, hampering knowledge acquisition. This work addresses these challenges by proposing a methodology to decompose multivariate time series residuals into statistically significant, meaningful events, effectively filtering noise. We extend the motif discovery task to identify irregular patterns satisfying key properties: non-triviality, statistical significance, multi-dimensionality, and actionability. Our approach introduces principles to mitigate biases, evaluate statistical significance, place robust hyperparameterization, explore relationships in multivariate residuals, and incorporate domain knowledge through specialized masks. Real-world case studies validate the proposed methodology, uncovering explainable patterns relevant to human activity recognition, energy consumption, and urban planning, accounting for up to 50% of irregular components and revealing hidden system behaviors.
Miguel G. Silva, Sara C. Madeira, Rui Henriques
Pattern Recognit.2
2026 On why and how statistical significance criteria can guide multivariate time series motif analysis
abstract
The modeling of real-world system dynamics is often challenged by the presence of non-periodic patterns, including stimuli-dependent physiological responses in health systems, event-elicited traffic flows in mobility systems, or news-driven interactions in societal systems. While motif discovery has proven effective in revealing recurring events from time series data, the presence of spurious patterns hampers this process. Moreover, existing statistical significance stances are limited to univariate symbolic series. To address this critical gap, this study proposes a statistical frame to assess the likelihood of motifs to deviate from null expectations in multivariate time series with arbitrary order and variable types. It includes a principled discussion on the application of binomial testing as a strategy to guide motif discovery and reduce the incidence of false positives, considering variable dependencies, temporal associations, and multi-hypothesis corrections. Results from real-world case studies suggest that current motif discovery algorithms are vulnerable to spurious patterns, with up to 60% false positive discoveries under normative search conditions. The proposed assessment can enhance existing motif discovery algorithms by minimizing the occurrence of spurious patterns, prioritizing pattern importance, and refining the search space to decrease computational complexity.
Miguel G. Silva, Sara C. Madeira, Rui Henriques
Pattern Recognit. Lett.2
2025 TriHSPAM: Triclustering heterogeneous longitudinal clinical data using sequential patterns
abstract
Triclustering has become a well-established approach for handling the complexities of three-way data analysis, aiming to uncover patterns with unexpectedly strong coherence across subsets of observations, features, and time-points. In biomedical settings, this technique facilitates the analysis of multivariate physiological signals, omic datasets, and clinical records to discern coherent responses and distinct patient subgroups. Despite its significant potential, existing triclustering algorithms face limitations when dealing with heterogeneous data that include mixed-type features. Key challenges include establishing robust coherence criteria, managing noise and missing data, and effectively addressing temporal complexities. This paper introduces TriHSPAM, a novel triclustering algorithm designed to address these challenges and provide insights into heterogeneous clinical data. A novel merit function is proposed to evaluate heterogeneous triclusters and guide the search process. TriHSPAM demonstrates remarkable flexibility, accommodating diverse data intricacies and capturing complex relationships within data. This algorithm incorporates noise-tolerant techniques to enhance reliability and robustness, effectively handling noisy and sparse real-world data. TriHSPAM adeptly captures meaningful temporal dynamics in three-way heterogeneous data, addressing temporal specificities, such as temporal contiguity and misalignments, while providing statistical significance guarantees. Experimental validation on both synthetic and real datasets confirms TriHSPAM’s efficacy in identifying coherent temporal subspaces within heterogeneous data, thereby guiding knowledge discovery from longitudinal studies.
Diogo F. Soares, Rui Henriques, Sara C. Madeira
Pattern Recognit.3
2024 iDPP@CLEF 2024: The Intelligent Disease Progression Prediction Challenge
Helena Aidos, Roberto Bergamaschi, Paola Cavalla, Adriano Chiò, Arianna Dagliati, Barbara Di Camillo, Mamede de Carvalho, Nicola Ferro 0001, Piero Fariselli, Jose Manuel García Dominguez, Sara C. Madeira, Eleonora Tavazzi
ECIR (6)11
2024 Deep Temporal Consensus Clustering for Patient Stratification in Amyotrophic Lateral Sclerosis
abstract
Amyotrophic Lateral Sclerosis (ALS) is a fast-acting neurodegenerative disease, characterized by loss of muscle movement and heterogeneity in disease evolution.This poses a challenge in predicting the best time for therapy administration.Here, we propose Deep Temporal Consensus Clustering (DTCC), a stratification method to uncover patient groups with similar disease progression.Using only the initial 6-month follow-up period, DTCC uncovered five clusters that were evaluated in terms of disease evolution and time-to-event.For three critical events (non-invasive ventilation, gastrostomy and death) the attained groups show distinct 10year progressions, validating the approach.
Miguel Pego Roque, Andreia S. Martins, Marta Gromicho, Mamede de Carvalho, Sara C. Madeira, Pedro Tomás, Helena Aidos
ESANN5
2024 Biclustering data analysis: a comprehensive survey
abstract
Biclustering, the simultaneous clustering of rows and columns of a data matrix, has proved its effectiveness in bioinformatics due to its capacity to produce local instead of global models, evolving from a key technique used in gene expression data analysis into one of the most used approaches for pattern discovery and identification of biological modules, used in both descriptive and predictive learning tasks. This survey presents a comprehensive overview of biclustering. It proposes an updated taxonomy for its fundamental components (bicluster, biclustering solution, biclustering algorithms, and evaluation measures) and applications. We unify scattered concepts in the literature with new definitions to accommodate the diversity of data types (such as tabular, network, and time series data) and the specificities of biological and biomedical data domains. We further propose a pipeline for biclustering data analysis and discuss practical aspects of incorporating biclustering in real-world applications. We highlight prominent application domains, particularly in bioinformatics, and identify typical biclusters to illustrate the analysis output. Moreover, we discuss important aspects to consider when choosing, applying, and evaluating a biclustering algorithm. We also relate biclustering with other data mining tasks (clustering, pattern mining, classification, triclustering, N-way clustering, and graph mining). Thus, it provides theoretical and practical guidance on biclustering data analysis, demonstrating its potential to uncover actionable insights from complex datasets.
Eduardo N. Castanho, Helena Aidos, Sara C. Madeira
Briefings Bioinform.3
2024 Correction: G-bic: generating synthetic benchmarks for biclustering
Eduardo N. Castanho, João Lobo, Rui Henriques, Sara C. Madeira
BMC Bioinform.4
2024 Comprehensive assessment of triclustering algorithms for three-way temporal data analysis
abstract
The analysis of temporal data has gained increasing attention in recent years, aiming to identify patterns and trends that change over time. Temporal triclustering is a promising approach for this purpose, as it allows for the simultaneous clustering of three dimensions of data: objects, attributes, and time. In this work, we present a comparative study and experimental evaluation of state-of-the-art temporal triclustering algorithms. Our study provides a comprehensive quantitative assessment of several triclustering algorithms to unravel their strengths and limitations. To this end, we consider synthetic data with varying sizes and regularities, where true solutions are planted with different coherence and quality criteria, in order to assess the algorithms’ performance in datasets with diverse characteristics and assess their capacity to retrieve specific types of hidden patterns. This provides a more comprehensive evaluation of the algorithms and allows for a better understanding of their capabilities and limitations. This study is the first to compare state-of-the-art triclustering algorithms inherently prepared to deal with temporal data and provides new benchmark results for the Temporal Triclustering task. Our results on the algorithms’ performance can guide practitioners in selecting the most appropriate algorithm for their specific application.
Diogo F. Soares, Rui Henriques, Sara C. Madeira
Pattern Recognit.3
2023 iDPP@CLEF 2023: The Intelligent Disease Progression Prediction Challenge
Helena Aidos, Roberto Bergamaschi, Paola Cavalla, Adriano Chiò, Arianna Dagliati, Barbara Di Camillo, Mamede de Carvalho, Nicola Ferro 0001, Piero Fariselli, Jose Manuel García Dominguez, Sara C. Madeira, Eleonora Tavazzi
ECIR (3)11
2023 Artificial intelligence and statistical methods for stratification and prediction of progression in amyotrophic lateral sclerosis: A systematic review
abstract
BACKGROUND: Amyotrophic Lateral Sclerosis (ALS) is a fatal neurodegenerative disorder characterised by the progressive loss of motor neurons in the brain and spinal cord. The fact that ALS's disease course is highly heterogeneous, and its determinants not fully known, combined with ALS's relatively low prevalence, renders the successful application of artificial intelligence (AI) techniques particularly arduous. OBJECTIVE: This systematic review aims at identifying areas of agreement and unanswered questions regarding two notable applications of AI in ALS, namely the automatic, data-driven stratification of patients according to their phenotype, and the prediction of ALS progression. Differently from previous works, this review is focused on the methodological landscape of AI in ALS. METHODS: We conducted a systematic search of the Scopus and PubMed databases, looking for studies on data-driven stratification methods based on unsupervised techniques resulting in (A) automatic group discovery or (B) a transformation of the feature space allowing patient subgroups to be identified; and for studies on internally or externally validated methods for the prediction of ALS progression. We described the selected studies according to the following characteristics, when applicable: variables used, methodology, splitting criteria and number of groups, prediction outcomes, validation schemes, and metrics. RESULTS: Of the starting 1604 unique reports (2837 combined hits between Scopus and PubMed), 239 were selected for thorough screening, leading to the inclusion of 15 studies on patient stratification, 28 on prediction of ALS progression, and 6 on both stratification and prediction. In terms of variables used, most stratification and prediction studies included demographics and features derived from the ALSFRS or ALSFRS-R scores, which were also the main prediction targets. The most represented stratification methods were K-means, and hierarchical and expectation-maximisation clustering; while random forests, logistic regression, the Cox proportional hazard model, and various flavours of deep learning were the most widely used prediction methods. Predictive model validation was, albeit unexpectedly, quite rarely performed in absolute terms (leading to the exclusion of 78 eligible studies), with the overwhelming majority of included studies resorting to internal validation only. CONCLUSION: This systematic review highlighted a general agreement in terms of input variable selection for both stratification and prediction of ALS progression, and in terms of prediction targets. A striking lack of validated models emerged, as well as a general difficulty in reproducing many published studies, mainly due to the absence of the corresponding parameter lists. While deep learning seems promising for prediction applications, its superiority with respect to traditional methods has not been established; there is, instead, ample room for its application in the subfield of patient stratification. Finally, an open question remains on the role of new environmental and behavioural variables collected via novel, real-time sensors.
Erica Tavazzi, Enrico Longato, Martina Vettoretti, Helena Aidos, Isotta Trescato, Chiara Roversi, Andreia S. Martins, Eduardo N. Castanho, Ruben Branco, Diogo F. Soares, Alessandro Guazzo, Giovanni Birolo, Daniele Pala, Pietro Bosoni, Adriano Chiò, Umberto Manera, Mamede de Carvalho, Bruno Miranda, Marta Gromicho, Inês Alves, Riccardo Bellazzi, Arianna Dagliati, Piero Fariselli, Sara C. Madeira, Barbara Di Camillo
Artif. Intell. Medicine24
2023 G-bic: generating synthetic benchmarks for biclustering
abstract
BACKGROUND: Biclustering is increasingly used in biomedical data analysis, recommendation tasks, and text mining domains, with hundreds of biclustering algorithms proposed. When assessing the performance of these algorithms, more than real datasets are required as they do not offer a solid ground truth. Synthetic data surpass this limitation by producing reference solutions to be compared with the found patterns. However, generating synthetic datasets is challenging since the generated data must ensure reproducibility, pattern representativity, and real data resemblance. RESULTS: We propose G-Bic, a dataset generator conceived to produce synthetic benchmarks for the normative assessment of biclustering algorithms. Beyond expanding on aspects of pattern coherence, data quality, and positioning properties, it further handles specificities related to mixed-type datasets and time-series data.G-Bic has the flexibility to replicate real data regularities from diverse domains. We provide the default configurations to generate reproducible benchmarks to evaluate and compare diverse aspects of biclustering algorithms. Additionally, we discuss empirical strategies to simulate the properties of real data. CONCLUSION: G-Bic is a parametrizable generator for biclustering analysis, offering a solid means to assess biclustering solutions according to internal and external metrics robustly.
Eduardo N. Castanho, João Lobo, Rui Henriques, Sara C. Madeira
BMC Bioinform.4
2022 Biclustering fMRI time series: a comparative study
abstract
BACKGROUND: The effectiveness of biclustering, simultaneous clustering of rows and columns in a data matrix, was shown in gene expression data analysis. Several researchers recognize its potentialities in other research areas. Nevertheless, the last two decades have witnessed the development of a significant number of biclustering algorithms targeting gene expression data analysis and a lack of consistent studies exploring the capacities of biclustering outside this traditional application domain. RESULTS: This work evaluates the potential use of biclustering in fMRI time series data, targeting the Region × Time dimensions by comparing seven state-in-the-art biclustering and three traditional clustering algorithms on artificial and real data. It further proposes a methodology for biclustering evaluation beyond gene expression data analysis. The results discuss the use of different search strategies in both artificial and real fMRI time series showed the superiority of exhaustive biclustering approaches, obtaining the most homogeneous biclusters. However, their high computational costs are a challenge, and further work is needed for the efficient use of biclustering in fMRI data analysis. CONCLUSIONS: This work pinpoints avenues for the use of biclustering in spatio-temporal data analysis, in particular neurosciences applications. The proposed evaluation methodology showed evidence of the effectiveness of biclustering in finding local patterns in fMRI time series data. Further work is needed regarding scalability to promote the application in real scenarios.
Eduardo N. Castanho, Helena Aidos, Sara C. Madeira
BMC Bioinform.3
2022 Learning prognostic models using a mixture of biclustering and triclustering: Predicting the need for non-invasive ventilation in Amyotrophic Lateral Sclerosis
abstract
Longitudinal cohort studies to study disease progression generally combine temporal features produced under periodic assessments (clinical follow-up) with static features associated with single-time assessments, genetic, psychophysiological, and demographic profiles. Subspace clustering, including biclustering and triclustering stances, enables the discovery of local and discriminative patterns from such multidimensional cohort data. These patterns, highly interpretable, are relevant to identifying groups of patients with similar traits or progression patterns. Despite their potential, their use for improving predictive tasks in clinical domains remains unexplored. In this work, we propose to learn predictive models from static and temporal data using discriminative patterns, obtained via biclustering and triclustering, as features within a state-of-the-art classifier, thus enhancing model interpretation. triCluster is extended to find time-contiguous triclusters in temporal data (temporal patterns) and a biclustering algorithm to discover coherent patterns in static data. The transformed data space, composed of bicluster and tricluster features, capture local and cross-variable associations with discriminative power, yielding unique statistical properties of interest. As a case study, we applied our methodology to follow-up data from Portuguese patients with Amyotrophic Lateral Sclerosis (ALS) to predict the need for non-invasive ventilation (NIV) since the last appointment. The results showed that, in general, our methodology outperformed baseline results using the original features. Furthermore, the bicluster/tricluster-based patterns used by the classifier can be used by clinicians to understand the models by highlighting relevant prognostic patterns.
Diogo F. Soares, Rui Henriques, Marta Gromicho, Mamede de Carvalho, Sara C. Madeira
J. Biomed. Informatics5
2022 Impact of metrics on biclustering solution and quality: A review
Marta D. M. Noronha, Rui Henriques, Sara C. Madeira, Luis E. Zárate
Pattern Recognit.3
2022 Learning Prognostic Models Using Disease Progression Patterns: Predicting the Need for Non-Invasive Ventilation in Amyotrophic Lateral Sclerosis
abstract
Amyotrophic Lateral Sclerosis is a devastating neurodegenerative disease causing rapid degeneration of motor neurons and usually leading to death by respiratory failure. Since there is no cure, treatment's goal is to improve symptoms and prolong survival. Non-invasive Ventilation (NIV) is an effective treatment, leading to extended life expectancy and improved quality of life. In this scenario, it is paramount to predict its need in order to allow preventive or timely administration. In this work, we propose to use itemset mining together with sequential pattern mining to unravel disease presentation patterns together with disease progression patterns by analysing, respectively, static data collected at diagnosis and longitudinal data from patient follow-up. The goal is to use these static and temporal patterns as features in prognostic models, enabling to take disease progression into account in predictions and promoting model interpretability. As case study, we predict the need for NIV within 90, 180 and 365 days (short, mid and long-term predictions). The learnt prognostic models are promising. Pattern evaluation through growth rate suggests bulbar function and phrenic nerve response amplitude, additionally to respiratory function, are significant features towards determining patient evolution. This confirms clinical knowledge regarding relevant biomarkers of disease progression towards respiratory insufficiency.
Andreia S. Martins, Marta Gromicho, Susana Pinto, Mamede de Carvalho, Sara C. Madeira
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 Mining Actionable Patterns of Road Mobility From Heterogeneous Traffic Data Using Biclustering
abstract
The comprehensive access to road traffic patterns in the continuously growing urban areas is key to achieve a sustainable mobility. However, the inherent complexity of urban traffic poses many challenges to achieve this goal, including: i) the need to integrate heterogeneous views of road traffic (such as speed limits, jam size, delay, throughput) from available sources; ii) the complex spatiotemporal intricacies of geolocalized speed and loop counter data; iii) the need to mine congestion patterns robust to the inherent traffic variability and unexpected occurrence of events, taking also into consideration the varying degrees of congestion severity; and iv) the need to guarantee the statistical significance and interpretability of the target patterns. In the context of our work, a road traffic pattern is a recurrent congestion profile (w.r.t. speed limits, jam extent and flow) that can span multiple locations and time periods within a day. Biclustering, the discovery of coherent subspaces (local patterns) within real-valued data, has unique properties of interest, being positioned to unravel such traffic patterns, while satisfying the aforementioned challenges. Despite its relevance, the potentialities of applying biclustering in mobility domains remain unexplored. This work proposes a structured view on why, when and how to apply biclustering for mining traffic patterns of road mobility, a subject remaining largely unexplored up to date. Using the city of Lisbon as a guiding case, we illustrate the relevance of biclustering geolocalized speed data and loop counter data. The gathered results confirm the role of biclustering in comprehensively finding statistically significant and actionable spatiotemporal associations of road mobility.
Francisco Neves, Anna Carolina Finamore, Sara C. Madeira, Rui Henriques
IEEE Trans. Intell. Transp. Syst.3
2021 G-Tric: generating three-way synthetic datasets with triclustering solutions
abstract
BACKGROUND: Three-way data started to gain popularity due to their increasing capacity to describe inherently multivariate and temporal events, such as biological responses, social interactions along time, urban dynamics, or complex geophysical phenomena. Triclustering, subspace clustering of three-way data, enables the discovery of patterns corresponding to data subspaces (triclusters) with values correlated across the three dimensions (observations [Formula: see text] features [Formula: see text] contexts). With increasing number of algorithms being proposed, effectively comparing them with state-of-the-art algorithms is paramount. These comparisons are usually performed using real data, without a known ground-truth, thus limiting the assessments. In this context, we propose a synthetic data generator, G-Tric, allowing the creation of synthetic datasets with configurable properties and the possibility to plant triclusters. The generator is prepared to create datasets resembling real 3-way data from biomedical and social data domains, with the additional advantage of further providing the ground truth (triclustering solution) as output. RESULTS: G-Tric can replicate real-world datasets and create new ones that match researchers needs across several properties, including data type (numeric or symbolic), dimensions, and background distribution. Users can tune the patterns and structure that characterize the planted triclusters (subspaces) and how they interact (overlapping). Data quality can also be controlled, by defining the amount of missing, noise or errors. Furthermore, a benchmark of datasets resembling real data is made available, together with the corresponding triclustering solutions (planted triclusters) and generating parameters. CONCLUSIONS: Triclustering evaluation using G-Tric provides the possibility to combine both intrinsic and extrinsic metrics to compare solutions that produce more reliable analyses. A set of predefined datasets, mimicking widely used three-way data and exploring crucial properties was generated and made available, highlighting G-Tric's potential to advance triclustering state-of-the-art by easing the process of evaluating the quality of new triclustering approaches.
João Lobo, Rui Henriques, Sara C. Madeira
BMC Bioinform.3
2021 Learning dynamic Bayesian networks from time-dependent and time-independent data: Unraveling disease progression in Amyotrophic Lateral Sclerosis
abstract
Amyotrophic lateral sclerosis (ALS) is a neurodegenerative disease causing patients to quickly lose motor neurons. The disease is characterized by a fast functional impairment and ventilatory decline, leading most patients to die from respiratory failure. To estimate when patients should get ventilatory support, it is helpful to adequately profile the disease progression. For this purpose, we use dynamic Bayesian networks (DBNs), a machine learning model, that graphically represents the conditional dependencies among variables. However, the standard DBN framework only includes dynamic (time-dependent) variables, while most ALS datasets have dynamic and static (time-independent) observations. Therefore, we propose the sdtDBN framework, which learns optimal DBNs with static and dynamic variables. Besides learning DBNs from data, with polynomial-time complexity in the number of variables, the proposed framework enables the user to insert prior knowledge and to make inference in the learned DBNs. We use sdtDBNs to study the progression of 1214 patients from a Portuguese ALS dataset. First, we predict the values of every functional indicator in the patients' consultations, achieving results competitive with state-of-the-art studies. Then, we determine the influence of each variable in patients' decline before and after getting ventilatory support. This insightful information can lead clinicians to pay particular attention to specific variables when evaluating the patients, thus improving prognosis. The case study with ALS shows that sdtDBNs are a promising predictive and descriptive tool, which can also be applied to assess the progression of other diseases, given time-dependent and time-independent clinical observations.
Tiago Leão, Sara C. Madeira, Marta Gromicho, Mamede de Carvalho, Alexandra M. Carvalho
J. Biomed. Informatics2
2021 FleBiC: Learning classifiers from high-dimensional biomedical data using discriminative biclusters with non-constant patterns
Rui Henriques, Sara C. Madeira
Pattern Recognit.2
2020 Targeting the uncertainty of predictions at patient-level using an ensemble of classifiers coupled with calibration methods, Venn-ABERS, and Conformal Predictors: A case study in AD
Telma Pereira, Sandra Cardoso, Manuela Guerreiro, Alexandre de Mendonça, Sara C. Madeira
J. Biomed. Informatics5
2019 An Architecture Based on Fuzzy Systems for Personalized Medicine in ICUs
abstract
This paper proposes a decision support system based on fuzzy clustering, fuzzy modeling and fuzzy fingerprints, to provide personalized therapy for critically ill patients. It is hypothesized that the ‘collective experience’ from large clinical databases, where clinical decisions are linked with patient outcomes, can be used to identify specific patient sub-groups and build personalized therapy models towards a new era of personalized medicine, allowing the improvement of patient outcomes in the Intensive Care Unit (ICU). The validity of the proposed systems will be tested using the case study of patients admitted to the ICU who then develop acute kidney injury (AKI); Two-thirds of patients with AKI require renal support therapy. Generalized severity scoring systems have consistently performed poorly for patients with AKI.
João Miguel da Costa Sousa, Susana M. Vieira, João Paulo Carvalho 0001, Sara C. Madeira, Leo A. Celi, Stan N. Finkelstein
FUZZ-IEEE4
2018 BSig: evaluating the statistical significance of biclustering solutions
Rui Henriques, Sara C. Madeira
Data Min. Knowl. Discov.2
2017 BicPAMS: software for biological data analysis with pattern-based biclustering
abstract
BACKGROUND: Biclustering has been largely applied for the unsupervised analysis of biological data, being recognised today as a key technique to discover putative modules in both expression data (subsets of genes correlated in subsets of conditions) and network data (groups of coherently interconnected biological entities). However, given its computational complexity, only recent breakthroughs on pattern-based biclustering enabled efficient searches without the restrictions that state-of-the-art biclustering algorithms place on the structure and homogeneity of biclusters. As a result, pattern-based biclustering provides the unprecedented opportunity to discover non-trivial yet meaningful biological modules with putative functions, whose coherency and tolerance to noise can be tuned and made problem-specific. METHODS: To enable the effective use of pattern-based biclustering by the scientific community, we developed BicPAMS (Biclustering based on PAttern Mining Software), a software that: 1) makes available state-of-the-art pattern-based biclustering algorithms (BicPAM (Henriques and Madeira, Alg Mol Biol 9:27, 2014), BicNET (Henriques and Madeira, Alg Mol Biol 11:23, 2016), BicSPAM (Henriques and Madeira, BMC Bioinforma 15:130, 2014), BiC2PAM (Henriques and Madeira, Alg Mol Biol 11:1-30, 2016), BiP (Henriques and Madeira, IEEE/ACM Trans Comput Biol Bioinforma, 2015), DeBi (Serin and Vingron, AMB 6:1-12, 2011) and BiModule (Okada et al., IPSJ Trans Bioinf 48(SIG5):39-48, 2007)); 2) consistently integrates their dispersed contributions; 3) further explores additional accuracy and efficiency gains; and 4) makes available graphical and application programming interfaces. RESULTS: Results on both synthetic and real data confirm the relevance of BicPAMS for biological data analysis, highlighting its essential role for the discovery of putative modules with non-trivial yet biologically significant functions from expression and network data. CONCLUSIONS: BicPAMS is the first biclustering tool offering the possibility to: 1) parametrically customize the structure, coherency and quality of biclusters; 2) analyze large-scale biological networks; and 3) tackle the restrictive assumptions placed by state-of-the-art biclustering algorithms. These contributions are shown to be key for an adequate, complete and user-assisted unsupervised analysis of biological data. SOFTWARE: BicPAMS and its tutorial available in http://www.bicpams.com .
Rui Henriques, Francisco L. Ferreira, Sara C. Madeira
BMC Bioinform.3
2017 Erratum to: BicPAMS: software for biological data analysis with pattern-based biclustering
Rui Henriques, Francisco L. Ferreira, Sara C. Madeira
BMC Bioinform.3
2015 BicNET: Efficient Biclustering of Biological Networks to Unravel Non-Trivial Modules
Rui Henriques, Sara C. Madeira
WABI2
2015 Generative modeling of repositories of health records for predictive tasks
Rui Henriques, Cláudia Antunes, Sara C. Madeira
Data Min. Knowl. Discov.3
2015 Multi-period classification: learning sequent classes from temporal domains
Rui Henriques, Sara C. Madeira, Cláudia Antunes
Data Min. Knowl. Discov.2
2015 Prognostic models based on patient snapshots and time windows: Predicting disease progression to assisted ventilation in Amyotrophic Lateral Sclerosis
André V. Carreiro, Pedro M. T. Amaral, Susana Pinto, Pedro Tomás, Mamede de Carvalho, Sara C. Madeira
J. Biomed. Informatics6
2015 A structured view on pattern mining-based biclustering
Rui Henriques, Cláudia Antunes, Sara C. Madeira
Pattern Recognit.3
2015 Biclustering with Flexible Plaid Models to Unravel Interactions between Biological Processes
abstract
Genes can participate in multiple biological processes at a time and thus their expression can be seen as a composition of the contributions from the active processes. Biclustering under a plaid assumption allows the modeling of interactions between transcriptional modules or biclusters (subsets of genes with coherence across subsets of conditions) by assuming an additive composition of contributions in their overlapping areas. Despite the biological interest of plaid models, few biclustering algorithms consider plaid effects and, when they do, they place restrictions on the allowed types and structures of biclusters, and suffer from robustness problems by seizing exact additive matchings. We propose BiP (Biclustering using Plaid models), a biclustering algorithm with relaxations to allow expression levels to change in overlapping areas according to biologically meaningful assumptions (weighted and noise-tolerant composition of contributions). BiP can be used over existing biclustering solutions (seizing their benefits) as it is able to recover excluded areas due to unaccounted plaid effects and detect noisy areas non-explained by a plaid assumption, thus producing an explanatory model of overlapping transcriptional activity. Experiments on synthetic data support BiP's efficiency and effectiveness. The learned models from expression data unravel meaningful and non-trivial functional interactions between biological processes associated with putative regulatory modules.
Rui Henriques, Sara C. Madeira
IEEE ACM Trans. Comput. Biol. Bioinform.2
2014 Mining coherent evolution patterns in education through biclustering
André Vale, Sara C. Madeira, Cláudia Antunes
EDM2
2014 BicSPAM: flexible biclustering using sequential patterns
abstract
BACKGROUND: Biclustering is a critical task for biomedical applications. Order-preserving biclusters, submatrices where the values of rows induce the same linear ordering across columns, capture local regularities with constant, shifting, scaling and sequential assumptions. Additionally, biclustering approaches relying on pattern mining output deliver exhaustive solutions with an arbitrary number and positioning of biclusters. However, existing order-preserving approaches suffer from robustness, scalability and/or flexibility issues. Additionally, they are not able to discover biclusters with symmetries and parameterizable levels of noise. RESULTS: We propose new biclustering algorithms to perform flexible, exhaustive and noise-tolerant biclustering based on sequential patterns (BicSPAM). Strategies are proposed to allow for symmetries and to seize efficiency gains from item-indexable properties and/or from partitioning methods with conservative distance guarantees. Results show BicSPAM ability to capture symmetries, handle planted noise, and scale in terms of memory and time. BicSPAM also achieves the best match-scores for the recovery of hidden biclusters in synthetic datasets with varying noise distributions and levels of missing values. Finally, results on gene expression data lead to complete solutions, delivering new biclusters corresponding to putative modules with heightened biological relevance. CONCLUSIONS: BicSPAM provides an exhaustive way to discover flexible structures of order-preserving biclusters. To the best of our knowledge, BicSPAM is the first attempt to deal with order-preserving biclusters that allow for symmetries and that are robust to varying levels of noise.
Rui Henriques, Sara C. Madeira
BMC Bioinform.2
2014 LateBiclustering: Efficient Heuristic Algorithm for Time-Lagged Bicluster Identification
abstract
Identifying patterns in temporal data is key to uncover meaningful relationships in diverse domains, from stock trading to social interactions. Also of great interest are clinical and biological applications, namely monitoring patient response to treatment or characterizing activity at the molecular level. In biology, researchers seek to gain insight into gene functions and dynamics of biological processes, as well as potential perturbations of these leading to disease, through the study of patterns emerging from gene expression time series. Clustering can group genes exhibiting similar expression profiles, but focuses on global patterns denoting rather broad, unspecific responses. Biclustering reveals local patterns, which more naturally capture the intricate collaboration between biological players, particularly under a temporal setting. Despite the general biclustering formulation being NP-hard, considering specific properties of time series has led to efficient solutions for the discovery of temporally aligned patterns. Notably, the identification of biclusters with time-lagged patterns, suggestive of transcriptional cascades, remains a challenge due to the combinatorial explosion of delayed occurrences. Herein, we propose LateBiclustering, a sensible heuristic algorithm enabling a polynomial rather than exponential time solution for the problem. We show that it identifies meaningful time-lagged biclusters relevant to the response of Saccharomyces cerevisiae to heat stress.
Joana P. Gonçalves, Sara C. Madeira
IEEE ACM Trans. Comput. Biol. Bioinform.2
2011 TFRank: network-based prioritization of regulatory associations underlying transcriptional responses
abstract
MOTIVATION: Uncovering mechanisms underlying gene expression control is crucial to understand complex cellular responses. Studies in gene regulation often aim to identify regulatory players involved in a biological process of interest, either transcription factors coregulating a set of target genes or genes eventually controlled by a set of regulators. These are frequently prioritized with respect to a context-specific relevance score. Current approaches rely on relevance measures accounting exclusively for direct transcription factor-target interactions, namely overrepresentation of binding sites or target ratios. Gene regulation has, however, intricate behavior with overlapping, indirect effect that should not be neglected. In addition, the rapid accumulation of regulatory data already enables the prediction of large-scale networks suitable for higher level exploration by methods based on graph theory. A paradigm shift is thus emerging, where isolated and constrained analyses will likely be replaced by whole-network, systemic-aware strategies. RESULTS: We present TFRank, a graph-based framework to prioritize regulatory players involved in transcriptional responses within the regulatory network of an organism, whereby every regulatory path containing genes of interest is explored and incorporated into the analysis. TFRank selected important regulators of yeast adaptation to stress induced by quinine and acetic acid, which were missed by a direct effect approach. Notably, they reportedly confer resistance toward the chemicals. In a preliminary study in human, TFRank unveiled regulators involved in breast tumor growth and metastasis when applied to genes whose expression signatures correlated with short interval to metastasis.
Joana P. Gonçalves, Alexandre P. Francisco, Nuno P. Mira, Miguel C. Teixeira, Isabel Sá-Correia, Arlindo L. Oliveira, Sara C. Madeira
Bioinform.7
2010 Identification of Regulatory Modules in Time Series Gene Expression Data Using a Linear Time Biclustering Algorithm
abstract
Although most biclustering formulations are NP-hard, in time series expression data analysis, it is reasonable to restrict the problem to the identification of maximal biclusters with contiguous columns, which correspond to coherent expression patterns shared by a group of genes in consecutive time points. This restriction leads to a tractable problem. We propose an algorithm that finds and reports all maximal contiguous column coherent biclusters in time linear in the size of the expression matrix. The linear time complexity of CCC-Biclustering relies on the use of a discretized matrix and efficient string processing techniques based on suffix trees. We also propose a method for ranking biclusters based on their statistical significance and a methodology for filtering highly overlapping and, therefore, redundant biclusters. We report results in synthetic and real data showing the effectiveness of the approach and its relevance in the discovery of regulatory modules. Results obtained using the transcriptomic expression patterns occurring in Saccharomyces cerevisiae in response to heat stress show not only the ability of the proposed methodology to extract relevant information compatible with documented biological knowledge but also the utility of using this algorithm in the study of other environmental stresses and of regulatory modules in general.
Sara C. Madeira, Miguel C. Teixeira, Isabel Sá-Correia, Arlindo L. Oliveira
IEEE ACM Trans. Comput. Biol. Bioinform.1
2009 High Level Thread-Based Competitive Or-Parallelism in Logtalk
Paulo Moura, Ricardo Rocha 0001, Sara C. Madeira
PADL3
2008 Thread-Based Competitive Or-Parallelism
Paulo Moura, Ricardo Rocha 0001, Sara C. Madeira
ICLP3
2007 An Efficient Biclustering Algorithm for Finding Genes with Similar Patterns in Time-series Expression Data
Sara C. Madeira, Arlindo L. Oliveira
APBC1
2005 A Linear Time Biclustering Algorithm for Time Series Gene Expression Data
Sara C. Madeira, Arlindo L. Oliveira
WABI1
2004 Biclustering Algorithms for Biological Data Analysis: A Survey
abstract
A large number of clustering approaches have been proposed for the analysis of gene expression data obtained from microarray experiments. However, the results from the application of standard clustering methods to genes are limited. This limitation is imposed by the existence of a number of experimental conditions where the activity of genes is uncorrelated. A similar limitation exists when clustering of conditions is performed. For this reason, a number of algorithms that perform simultaneous clustering on the row and column dimensions of the data matrix has been proposed. The goal is to find submatrices, that is, subgroups of genes and subgroups of conditions, where the genes exhibit highly correlated activities for every condition. In this paper, we refer to this class of algorithms as biclustering. Biclustering is also referred in the literature as coclustering and direct clustering, among others names, and has also been used in fields such as information retrieval and data mining. In this comprehensive survey, we analyze a large number of existing approaches to biclustering, and classify them in accordance with the type of biclusters they can find, the patterns of biclusters that are discovered, the methods used to perform the search, the approaches used to evaluate the solution, and the target applications.
Sara C. Madeira, Arlindo L. Oliveira
IEEE ACM Trans. Comput. Biol. Bioinform.1
2003 Modeling charity donations using target selection for revenue maximization
abstract
This paper presents the results of one application of target selection in direct marketing: the mailing campaigns of a charity organization, where the clients are selected based on the expected amount of donation they are going to make. Target selection is an important data mining problem for which several modeling techniques have been used. Statistical regression, neural networks, decision trees, and clustering are the most utilized techniques. Fuzzy clustering can also be applied to target selection. In this paper, traditional and fuzzy techniques are compared by using cross-validation measures. The four techniques are applied based on recency, frequency and monetary value measures. The application to mailing campaigns of a charity organization, showed that fuzzy modeling obtains results similar to those of other classical target selection techniques.
João Miguel da Costa Sousa, Sara C. Madeira, Uzay Kaymak
FUZZ-IEEE2
2002 A comparative study of fuzzy target selection methods in direct marketing
abstract
Target selection in direct marketing is an important data mining problem for which fuzzy modeling can be used. The paper compares several fuzzy modeling techniques applied to target selection based on recency, frequency and monetary value measures. The comparison uses cross validation applied to mailing campaigns of a charity organization.
João Miguel da Costa Sousa, Uzay Kaymak, Sara C. Madeira
FUZZ-IEEE3