VLDB 2026 Research / reviewers in the wild / expert
Rui Henriques
dblp:55/9661
· DBLP profile ↗
41ranked-venue papers
13as first author
24since 2021 · last 2026
0000-0002-3993-0171ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Networked data science: a unified network modeling frameworkabstractAbstract This work introduces a data-centric framework for answering analytical questions using network models, transcending domain-specific modeling conventions. The unified network modeling framework (UNMF) provides an interdisciplinary strategy for principled modeling and analysis of complex networks, including multilayer and temporal networks that arise in sustainability-driven applications. UNMF connects dynamic analysis, pattern discovery, and network-grounded integration of heterogeneous sources (structured and unstructured). We present guided instantiations of UNMF in urban development, mobility, ecosystems, and social-network settings to show how explicit modeling choices can be documented, compared, and assessed within a shared evaluative framework. In doing so, we formalize and systematize Networked Data Science as a field at the intersection of network science and data science, and we define its scope and applications. This work contributes to network science by providing an auditable design-and-evaluation procedure for studying complex, evolving systems and for making representation choices more explicit, inspectable, and reusable across domains. João Tiago Aparício, Elisabete Arsenio, Rui Henriques |
Knowl. Inf. Syst. | 3 |
| 2026 | Cutting through the noise: Explaining residuals in multivariate time series with motif analysisabstractModeling real-world system dynamics is challenging due to non-periodic patterns, such as stimuli-dependent physiological responses in health, event-driven traffic in mobility, and news-triggered interactions in societal systems. In the absence of contextual data, state-of-the-art methods—including advanced deep learning architectures—struggle to model these irregular behaviors. Furthermore, their predictive focus often limits their utility for descriptive analysis, hampering knowledge acquisition. This work addresses these challenges by proposing a methodology to decompose multivariate time series residuals into statistically significant, meaningful events, effectively filtering noise. We extend the motif discovery task to identify irregular patterns satisfying key properties: non-triviality, statistical significance, multi-dimensionality, and actionability. Our approach introduces principles to mitigate biases, evaluate statistical significance, place robust hyperparameterization, explore relationships in multivariate residuals, and incorporate domain knowledge through specialized masks. Real-world case studies validate the proposed methodology, uncovering explainable patterns relevant to human activity recognition, energy consumption, and urban planning, accounting for up to 50% of irregular components and revealing hidden system behaviors. Miguel G. Silva, Sara C. Madeira, Rui Henriques |
Pattern Recognit. | 3 |
| 2026 | On why and how statistical significance criteria can guide multivariate time series motif analysisabstractThe modeling of real-world system dynamics is often challenged by the presence of non-periodic patterns, including stimuli-dependent physiological responses in health systems, event-elicited traffic flows in mobility systems, or news-driven interactions in societal systems. While motif discovery has proven effective in revealing recurring events from time series data, the presence of spurious patterns hampers this process. Moreover, existing statistical significance stances are limited to univariate symbolic series. To address this critical gap, this study proposes a statistical frame to assess the likelihood of motifs to deviate from null expectations in multivariate time series with arbitrary order and variable types. It includes a principled discussion on the application of binomial testing as a strategy to guide motif discovery and reduce the incidence of false positives, considering variable dependencies, temporal associations, and multi-hypothesis corrections. Results from real-world case studies suggest that current motif discovery algorithms are vulnerable to spurious patterns, with up to 60% false positive discoveries under normative search conditions. The proposed assessment can enhance existing motif discovery algorithms by minimizing the occurrence of spurious patterns, prioritizing pattern importance, and refining the search space to decrease computational complexity. Miguel G. Silva, Sara C. Madeira, Rui Henriques |
Pattern Recognit. Lett. | 3 |
| 2025 | Integrating statistical significance and discriminative power in pattern discoveryabstractPattern discovery plays a central role in knowledge acquisition across multiple domains. Actionable patterns must meet rigorous statistical significance criteria and, in the presence of target variables, further uphold discriminative power. This work addresses the underexplored area of guiding pattern discovery by integrating statistical significance and discriminative power criteria into state-of-the-art algorithms while preserving pattern quality. We also address how pattern quality thresholds, imposed by some algorithms, can be rectified to accommodate these additional criteria. To assess the proposed methodology, we select the triclustering task as the guiding pattern discovery case and extend well-known greedy and multi-objective optimization triclustering algorithms, δ -Trimax and TriGen, with various merit functions, such as Mean Squared Residual (MSR), Least Squared Lines (LSL), and Multi Slope Measure (MSL). Results from three case studies show the role of the proposed methodology in discovering patterns with pronouncedly higher discriminative power and statistical significance without quality deterioration, highlighting the relevance of combining these criteria to guide the underlying searches and aid knowledge discovery. Although the proposed methodology is motivated over multivariate time series data, it is straightforwardly extensible to pattern discovery tasks involving multivariate, N-way (N > 3), transactional, and sequential data structures. Availability: The code is freely available at https://github.com/JupitersMight/MOF_Triclustering • Integration of discriminative power views into pattern discovery. • Adaptable framework for pattern-centric knowledge acquisition. • Patterns with guarantees of quality, statistical significance, and discriminative power. • Extended triclustering views for pattern discovery in tensor data. • Mitigation of false positive and false negative risks during pattern discovery. Leonardo Alexandre, Rafael S. Costa, Rui Henriques |
Knowl. Based Syst. | 3 |
| 2025 | TriHSPAM: Triclustering heterogeneous longitudinal clinical data using sequential patternsabstractTriclustering has become a well-established approach for handling the complexities of three-way data analysis, aiming to uncover patterns with unexpectedly strong coherence across subsets of observations, features, and time-points. In biomedical settings, this technique facilitates the analysis of multivariate physiological signals, omic datasets, and clinical records to discern coherent responses and distinct patient subgroups. Despite its significant potential, existing triclustering algorithms face limitations when dealing with heterogeneous data that include mixed-type features. Key challenges include establishing robust coherence criteria, managing noise and missing data, and effectively addressing temporal complexities. This paper introduces TriHSPAM, a novel triclustering algorithm designed to address these challenges and provide insights into heterogeneous clinical data. A novel merit function is proposed to evaluate heterogeneous triclusters and guide the search process. TriHSPAM demonstrates remarkable flexibility, accommodating diverse data intricacies and capturing complex relationships within data. This algorithm incorporates noise-tolerant techniques to enhance reliability and robustness, effectively handling noisy and sparse real-world data. TriHSPAM adeptly captures meaningful temporal dynamics in three-way heterogeneous data, addressing temporal specificities, such as temporal contiguity and misalignments, while providing statistical significance guarantees. Experimental validation on both synthetic and real datasets confirms TriHSPAM’s efficacy in identifying coherent temporal subspaces within heterogeneous data, thereby guiding knowledge discovery from longitudinal studies. Diogo F. Soares, Rui Henriques, Sara C. Madeira |
Pattern Recognit. | 2 |
| 2024 | Comparing Deep Neural Networks for 1D Multi-Label Power Quality Disturbance Classification: ResNet, MobileNet, and DenseNetabstractPower quality (PQ) analysis has become increasingly challenging due to the widespread usage of power electronic devices and the variability of operational events in electric grids integrating renewable energy resources. Deep neural networks (DNNs) have emerged as the prevalent approach for identifying PQ disturbances (PQDs). The 1D DNN-based disturbance classification architectures, in contrast to 2D models, operate directly on original waveform measurements, offering sequential views to capture irregular PQDs. However, there remains a research gap on the impact of architectural choices of DNN frameworks in this context, as well as on understanding their robustness to noise and adequacy for multi-output identification tasks. This work delves into three advanced DNN architectures, namely ResNet, MobileNet, and DenseNet, to assess their suitability for 1D PQD classification tasks. The performance of different classification models is compared using synthetic datasets following the IEEE 1159-2019 standard. The results demonstrate that ResNet outperforms MobileNet and DenseNet in terms of average accuracy and noise tolerance. Rui Henriques, Hugo Morais, Elisabetta Tedeschi |
IECON | 2 |
| 2024 | Correction: G-bic: generating synthetic benchmarks for biclustering
Eduardo N. Castanho, João Lobo, Rui Henriques, Sara C. Madeira |
BMC Bioinform. | 3 |
| 2024 | Multiple-input neural networks for time series forecasting incorporating historical and prospective contextabstractAbstract Individual and societal systems are open systems continuously affected by their situational context. In recent years, context sources have been increasingly considered in different domains to aid short and long-term forecasts of systems’ behavior. Nevertheless, available research generally disregards the role of prospective context, such as calendrical planning or weather forecasts. This work proposes a multiple-input neural architecture consisting of a sequential composition of long short-term memory units or temporal convolutional networks able to incorporate both historical and prospective sources of situational context to aid time series forecasting tasks. Considering urban case studies, we further assess the impact that different sources of external context have on medical emergency and mobility forecasts. Results show that the incorporation of external context variables, including calendrical and weather variables, can significantly reduce forecasting errors against state-of-the-art forecasters. In particular, the incorporation of prospective context, generally neglected in related work, mitigates error increases along the forecasting horizon. João Palet, Vasco Manquinho, Rui Henriques |
Data Min. Knowl. Discov. | 3 |
| 2024 | Using dynamic knowledge graphs to detect emerging communities of knowledgeabstractKnowledge graphs represent relationships between entities. These graphs can take dynamic forms to trace changes along time through text models and further used by reasoning systems with the intent to answer queries. In this research we explore their applicability for extracting temporal patterns of knowledge in the form of communities. To this end, we propose a method for generating knowledge relationships over unconnected components of a knowledge graph, allowing for a targeted exploration of emerging contents in corpora. This analysis is applied to the corpora of the Conference on Knowledge Discovery and Data Mining (KDD) publications over the last decade. We find the key knowledge communities over time and rank the underlying concepts. Results show that the publication efforts increasingly focus on graph research and the creation of relationships instead of new concepts. The acquired results confirm the validity of the proposed knowledge discovery methodology for community-centered analysis of emerging changes in dynamic knowledge graphs. João Tiago Aparício, Elisabete Arsenio, Francisco Santos, Rui Henriques |
Knowl. Based Syst. | 4 |
| 2024 | TriSig: Evaluating the statistical significance of triclustersabstractTensor data analysis allows researchers to uncover novel patterns and relationships that cannot be obtained from tabular data alone. The information inferred from multi-way patterns can offer valuable insights into disease progression, bioproduction processes, behavioral responses, weather fluctuations, or social dynamics. However, spurious patterns often hamper this process. This work aims at proposing a statistical frame to assess the probability of patterns in tensor data to deviate from null expectations, extending well-established principles for assessing the statistical significance of patterns in tabular data. A principled discussion on binomial testing to mitigate false positive discoveries is entailed at the light of: variable dependencies, temporal associations and misalignments, and multi-hypothesis correction. Results gathered from the application of triclustering algorithms over distinct real-world case studies in biotechnological domains confer validity to the proposed statistical frame while revealing vulnerabilities of reference triclustering searches. The proposed assessment can be incorporated into existing triclustering algorithms to minimize spurious occurrences, rank patterns, and further prune the search space, reducing their computational complexity. Leonardo Alexandre, Rafael S. Costa, Rui Henriques |
Pattern Recognit. | 3 |
| 2024 | Comprehensive assessment of triclustering algorithms for three-way temporal data analysisabstractThe analysis of temporal data has gained increasing attention in recent years, aiming to identify patterns and trends that change over time. Temporal triclustering is a promising approach for this purpose, as it allows for the simultaneous clustering of three dimensions of data: objects, attributes, and time. In this work, we present a comparative study and experimental evaluation of state-of-the-art temporal triclustering algorithms. Our study provides a comprehensive quantitative assessment of several triclustering algorithms to unravel their strengths and limitations. To this end, we consider synthetic data with varying sizes and regularities, where true solutions are planted with different coherence and quality criteria, in order to assess the algorithms’ performance in datasets with diverse characteristics and assess their capacity to retrieve specific types of hidden patterns. This provides a more comprehensive evaluation of the algorithms and allows for a better understanding of their capabilities and limitations. This study is the first to compare state-of-the-art triclustering algorithms inherently prepared to deal with temporal data and provides new benchmark results for the Temporal Triclustering task. Our results on the algorithms’ performance can guide practitioners in selecting the most appropriate algorithm for their specific application. Diogo F. Soares, Rui Henriques, Sara C. Madeira |
Pattern Recognit. | 2 |
| 2024 | Detecting Fraudulent Student Communication in a Multiple Choice Online Test EnvironmentabstractOnline evaluation systems, pervasive nowadays, are known to be susceptible to higher fraud risks. This work proposes a novel and robust method to detect potential fraud acts in online multiple-choice question (MCQ) exams. For the first time, the communication probability between the examinees is statistically assessed based on the concordance of responses and answer time against null expectations and is subsequently used to identify potential fraud behavior. The model is sensitive to the direction of communication acts, distinguishing content consumption from production, as well as multiwise communication channels. Online remote tests from engineering courses at Técnico Lisboa are used as a case study. We show that the cumulative contribution of concordant responses between students, when recurrent, offers a way of signaling fraud behavior. Separating content production from consumption reveals the underlying student role played in potential fraud acts. Collusion behavior is assessed against null models of fraud and conformity, and therefore being statistically framed and offering a solid criterion to guide tutors in ascertaining fraud and discouraging communication. Mariana Carrasco, António Manuel Ferreira Rito da Silva, Rui Henriques |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | G-bic: generating synthetic benchmarks for biclusteringabstractBACKGROUND: Biclustering is increasingly used in biomedical data analysis, recommendation tasks, and text mining domains, with hundreds of biclustering algorithms proposed. When assessing the performance of these algorithms, more than real datasets are required as they do not offer a solid ground truth. Synthetic data surpass this limitation by producing reference solutions to be compared with the found patterns. However, generating synthetic datasets is challenging since the generated data must ensure reproducibility, pattern representativity, and real data resemblance. RESULTS: We propose G-Bic, a dataset generator conceived to produce synthetic benchmarks for the normative assessment of biclustering algorithms. Beyond expanding on aspects of pattern coherence, data quality, and positioning properties, it further handles specificities related to mixed-type datasets and time-series data.G-Bic has the flexibility to replicate real data regularities from diverse domains. We provide the default configurations to generate reproducible benchmarks to evaluate and compare diverse aspects of biclustering algorithms. Additionally, we discuss empirical strategies to simulate the properties of real data. CONCLUSION: G-Bic is a parametrizable generator for biclustering analysis, offering a solid means to assess biclustering solutions according to internal and external metrics robustly. Eduardo N. Castanho, João Lobo, Rui Henriques, Sara C. Madeira |
BMC Bioinform. | 3 |
| 2022 | Context-situated visualization of biclusters to aid decisions: going beyond subspaces with parallel coordinatesabstractPattern discovery and subspace clustering are pervasive tasks across biological, biotechnological, and biomedical domains. Parallel coordinates plots and heatmaps are reference visualizations for individual biclusters. Both have been object of improvements over time, with a special emphasis on heatmaps, commonly used in gene expression analysis. However, the emphasis is solely placed on the corresponding subspace, preventing an assessment of biclusters’ significance against global regularities. This work proposes an improvement on bicluster visualization by disruptively extending parallel coordinates representations with the means to compare the local bicluster against the remaining dataset instances helping in the contextualization of a pattern in the broader picture of an entire dataset. The proposed solution is the first able to deal with mixed data types and is independent from the underlying biclustering or pattern mining algorithm. Results in different data domains show the utility of the proposed visualization, especially in primary phases where visual inspection of biclusters is used. Rafael S. Costa, Rui Henriques |
AVI | 3 |
| 2022 | Order-preserving pattern matching indeterminate strings
Luís M. S. Russo, Diogo M. Costa, Rui Henriques, Hideo Bannai, Alexandre P. Francisco |
Inf. Comput. | 3 |
| 2022 | Learning prognostic models using a mixture of biclustering and triclustering: Predicting the need for non-invasive ventilation in Amyotrophic Lateral SclerosisabstractLongitudinal cohort studies to study disease progression generally combine temporal features produced under periodic assessments (clinical follow-up) with static features associated with single-time assessments, genetic, psychophysiological, and demographic profiles. Subspace clustering, including biclustering and triclustering stances, enables the discovery of local and discriminative patterns from such multidimensional cohort data. These patterns, highly interpretable, are relevant to identifying groups of patients with similar traits or progression patterns. Despite their potential, their use for improving predictive tasks in clinical domains remains unexplored. In this work, we propose to learn predictive models from static and temporal data using discriminative patterns, obtained via biclustering and triclustering, as features within a state-of-the-art classifier, thus enhancing model interpretation. triCluster is extended to find time-contiguous triclusters in temporal data (temporal patterns) and a biclustering algorithm to discover coherent patterns in static data. The transformed data space, composed of bicluster and tricluster features, capture local and cross-variable associations with discriminative power, yielding unique statistical properties of interest. As a case study, we applied our methodology to follow-up data from Portuguese patients with Amyotrophic Lateral Sclerosis (ALS) to predict the need for non-invasive ventilation (NIV) since the last appointment. The results showed that, in general, our methodology outperformed baseline results using the original features. Furthermore, the bicluster/tricluster-based patterns used by the classifier can be used by clinicians to understand the models by highlighting relevant prognostic patterns. Diogo F. Soares, Rui Henriques, Marta Gromicho, Mamede de Carvalho, Sara C. Madeira |
J. Biomed. Informatics | 2 |
| 2022 | Impact of metrics on biclustering solution and quality: A review
Marta D. M. Noronha, Rui Henriques, Sara C. Madeira, Luis E. Zárate |
Pattern Recognit. | 2 |
| 2022 | Mining Actionable Patterns of Road Mobility From Heterogeneous Traffic Data Using BiclusteringabstractThe comprehensive access to road traffic patterns in the continuously growing urban areas is key to achieve a sustainable mobility. However, the inherent complexity of urban traffic poses many challenges to achieve this goal, including: i) the need to integrate heterogeneous views of road traffic (such as speed limits, jam size, delay, throughput) from available sources; ii) the complex spatiotemporal intricacies of geolocalized speed and loop counter data; iii) the need to mine congestion patterns robust to the inherent traffic variability and unexpected occurrence of events, taking also into consideration the varying degrees of congestion severity; and iv) the need to guarantee the statistical significance and interpretability of the target patterns. In the context of our work, a road traffic pattern is a recurrent congestion profile (w.r.t. speed limits, jam extent and flow) that can span multiple locations and time periods within a day. Biclustering, the discovery of coherent subspaces (local patterns) within real-valued data, has unique properties of interest, being positioned to unravel such traffic patterns, while satisfying the aforementioned challenges. Despite its relevance, the potentialities of applying biclustering in mobility domains remain unexplored. This work proposes a structured view on why, when and how to apply biclustering for mining traffic patterns of road mobility, a subject remaining largely unexplored up to date. Using the city of Lisbon as a guiding case, we illustrate the relevance of biclustering geolocalized speed data and loop counter data. The gathered results confirm the role of biclustering in comprehensively finding statistically significant and actionable spatiotemporal associations of road mobility. Francisco Neves, Anna Carolina Finamore, Sara C. Madeira, Rui Henriques |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | fMRI Multiple Missing Values Imputation Regularized by a Recurrent Denoiser
David Calhas, Rui Henriques |
AIME | 2 |
| 2021 | UNIANO: robust and efficient anomaly consensus in time series sensitive to cross-correlated anomaly profilesabstractTime series anomaly detection is an active research area, combining dozens of state-of-the-art methods that place heterogeneous views on what is an anomaly.This diversity of views -local and global, point and segment, univariate and multivariate, context-free and context-aware anomalies -is associated with moderate-to-high output divergences between methods.As a result, the user is faced with the difficult and laborious task of selecting the most appropriate methods and identifying cross-method consensus in an attempt to optimize recall and precision.Despite the relevance of establishing agreement criteria, existing principles are scarce and suffer from major problems: 1) show biases towards methods with correlated/redundant anomaly profiles; 2) depend on anomaly score thresholding; 3) prevent online detection; and 4) offer consensus not subjected to sound statistical testing.This work proposes UNIANO (UNIfied ANOmaly), an approach that combines simple yet effective empirical multivariate distribution statistics to address these drawbacks, guaranteeing a parameter-free and statistically robust integration of heterogeneous anomaly views.In this context, anomalies detected by less prevalent and concordant anomaly profiles, such as context-aware profiles in the presence of complementary variables, are not undervalued.Given a n-length time series and m views, UNIANO is aided by adequate data structures to achieve O(n log m 2 n) training time and linear O(m) testing-and-updating time.The gathered results confirm the relevance of the proposed approach. Leonor Silva, Helena Galhardas, Vasco Manquinho, Rui Henriques |
SDM | 4 |
| 2021 | DI2: prior-free and multi-item discretization of biological data and its applicationsabstractBACKGROUND: A considerable number of data mining approaches for biomedical data analysis, including state-of-the-art associative models, require a form of data discretization. Although diverse discretization approaches have been proposed, they generally work under a strict set of statistical assumptions which are arguably insufficient to handle the diversity and heterogeneity of clinical and molecular variables within a given dataset. In addition, although an increasing number of symbolic approaches in bioinformatics are able to assign multiple items to values occurring near discretization boundaries for superior robustness, there are no reference principles on how to perform multi-item discretizations. RESULTS: In this study, an unsupervised discretization method, DI2, for variables with arbitrarily skewed distributions is proposed. Statistical tests applied to assess differences in performance confirm that DI2 generally outperforms well-established discretizations methods with statistical significance. Within classification tasks, DI2 displays either competitive or superior levels of predictive accuracy, particularly delineate for classifiers able to accommodate border values. CONCLUSIONS: This work proposes a new unsupervised method for data discretization, DI2, that takes into account the underlying data regularities, the presence of outlier values disrupting expected regularities, as well as the relevance of border values. DI2 is available at https://github.com/JupitersMight/DI2. Leonardo Alexandre, Rafael S. Costa, Rui Henriques |
BMC Bioinform. | 3 |
| 2021 | G-Tric: generating three-way synthetic datasets with triclustering solutionsabstractBACKGROUND: Three-way data started to gain popularity due to their increasing capacity to describe inherently multivariate and temporal events, such as biological responses, social interactions along time, urban dynamics, or complex geophysical phenomena. Triclustering, subspace clustering of three-way data, enables the discovery of patterns corresponding to data subspaces (triclusters) with values correlated across the three dimensions (observations [Formula: see text] features [Formula: see text] contexts). With increasing number of algorithms being proposed, effectively comparing them with state-of-the-art algorithms is paramount. These comparisons are usually performed using real data, without a known ground-truth, thus limiting the assessments. In this context, we propose a synthetic data generator, G-Tric, allowing the creation of synthetic datasets with configurable properties and the possibility to plant triclusters. The generator is prepared to create datasets resembling real 3-way data from biomedical and social data domains, with the additional advantage of further providing the ground truth (triclustering solution) as output. RESULTS: G-Tric can replicate real-world datasets and create new ones that match researchers needs across several properties, including data type (numeric or symbolic), dimensions, and background distribution. Users can tune the patterns and structure that characterize the planted triclusters (subspaces) and how they interact (overlapping). Data quality can also be controlled, by defining the amount of missing, noise or errors. Furthermore, a benchmark of datasets resembling real data is made available, together with the corresponding triclustering solutions (planted triclusters) and generating parameters. CONCLUSIONS: Triclustering evaluation using G-Tric provides the possibility to combine both intrinsic and extrinsic metrics to compare solutions that produce more reliable analyses. A set of predefined datasets, mimicking widely used three-way data and exploring crucial properties was generated and made available, highlighting G-Tric's potential to advance triclustering state-of-the-art by easing the process of evaluating the quality of new triclustering approaches. João Lobo, Rui Henriques, Sara C. Madeira |
BMC Bioinform. | 2 |
| 2021 | FleBiC: Learning classifiers from high-dimensional biomedical data using discriminative biclusters with non-constant patterns
Rui Henriques, Sara C. Madeira |
Pattern Recognit. | 1 |
| 2021 | Mining Pre-Surgical Patterns Able to Discriminate Post-Surgical Outcomes in the Oncological DomainabstractUnderstanding the individualized risks of undertaking surgical procedures is essential to personalize preparatory, intervention and post-care protocols for minimizing post-surgical complications. This knowledge is key in oncology given the nature of interventions, the fragile profile of patients with comorbidities and cytotoxic drug exposure, and the possible cancer recurrence. Despite its relevance, the discovery of discriminative patterns of post-surgical risk is hampered by major challenges: i) the unique physiological and demographic profile of individuals, as well as their differentiated post-surgical care; ii) the high-dimensionality and heterogeneous nature of available biomedical data, combining non-identically distributed risk factors, clinical and molecular variables; iii) the need to generalize tumors have significant histopathological differences and individuals undertake unique surgical procedures; iv) the need to focus on non-trivial patterns of post-surgical risk, while guaranteeing their statistical significance and discriminative power; and v) the lack of interpretability and actionability of current approaches. Biclustering, the discovery of groups of individuals correlated on subsets of variables, has unique properties of interest, being positioned to satisfy the aforementioned challenges. In this context, this work proposes a structured view on why, when and how to apply biclustering to mine discriminative patterns of post-surgical risk with guarantees of usability, a subject remaining unexplored up to date. These patterns offer a comprehensive view on how the patient profile, cancer histopathology and entailed surgical procedures determine: i) post-surgical complications, ii) survival, and iii) hospitalization needs. The gathered results confirm the role of biclustering in comprehensively finding interpretable, actionable and statistically significant patterns of post-surgical risk. The found patterns are already assisting healthcare professionals at IPO-Porto to establish specialized pre-habilitation protocols and bedside care. Leonardo Alexandre, Rafael S. Costa, Lucio Lara-Santos, Rui Henriques |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Moving from Formal Towards Coherent Concept Analysis: Why, When and How
Pavlo Kovalchuk, Diogo Proença, José Borbinha, Rui Henriques |
ECIR (1) | 4 |
| 2020 | Efficient discovery of emerging patternsin heterogeneous spatiotemporal data from mobile sensorsabstractHeterogeneous sensor networks, including traffic monitoring systems and telemetry systems, produce massive spatiotemporal data. Geolocated time series data and timestamped trajectory data are generally produced from fixed and mobile sensors in these systems, offering the possibility to detect events of interest. Events of interest generally comprise emerging and gradual changes in the behavior of those systems, including patterns of congestion in road, utility and communication networks. However, the comprehensive discovery of these actionable events is challenged by the: i) inherently spatiotemporal and heterogeneous nature of data produced by different sensors; ii) difficulty of detecting emerging patterns not yet markedly noticeable at early stages; and iii) massive data size. Francisco Neves, Anna Carolina Finamore, Rui Henriques |
MobiQuitous | 3 |
| 2020 | On the use of pairwise distance learning for brain signal classification with limited observations
David Calhas, Enrique Romero, Rui Henriques |
Artif. Intell. Medicine | 3 |
| 2019 | An Unsupervised Method for Concept Association Analysis in Text Collections
Pavlo Kovalchuk, Diogo Proença, José Borbinha, Rui Henriques |
TPDL | 4 |
| 2019 | On the Discovery of Educational Patterns using Biclustering
Rui Henriques, Anna Carolina Finamore, Marco A. Casanova |
ITS | 1 |
| 2018 | Order-Preserving Pattern Matching Indeterminate StringsabstractGiven an indeterminate string pattern $p$ and an indeterminate string text $t$, the problem of order-preserving pattern matching with character uncertainties ($μ$OPPM) is to find all substrings of $t$ that satisfy one of the possible orderings defined by $p$. When the text and pattern are determinate strings, we are in the presence of the well-studied exact order-preserving pattern matching (OPPM) problem with diverse applications on time series analysis. Despite its relevance, the exact OPPM problem suffers from two major drawbacks: 1) the inability to deal with indetermination in the text, thus preventing the analysis of noisy time series; and 2) the inability to deal with indetermination in the pattern, thus imposing the strict satisfaction of the orders among all pattern positions. This paper provides the first polynomial algorithm to answer the $μ$OPPM problem when indetermination is observed on the pattern or text. Given two strings with length $m$ and $O(r)$ uncertain characters per string position, we show that the $μ$OPPM problem can be solved in $O(mr\lg r)$ time when one string is indeterminate and $r\in\mathbb{N}^+$. Mappings into satisfiability problems are provided when indetermination is observed on both the pattern and the text, and results concerning the general problem complexity are presented as well, with $μ$OPPM problem proved to be NP-hard in general. Rui Henriques, Alexandre P. Francisco, Luís M. S. Russo, Hideo Bannai |
CPM | 1 |
| 2018 | BSig: evaluating the statistical significance of biclustering solutions
Rui Henriques, Sara C. Madeira |
Data Min. Knowl. Discov. | 1 |
| 2017 | BicPAMS: software for biological data analysis with pattern-based biclusteringabstractBACKGROUND: Biclustering has been largely applied for the unsupervised analysis of biological data, being recognised today as a key technique to discover putative modules in both expression data (subsets of genes correlated in subsets of conditions) and network data (groups of coherently interconnected biological entities). However, given its computational complexity, only recent breakthroughs on pattern-based biclustering enabled efficient searches without the restrictions that state-of-the-art biclustering algorithms place on the structure and homogeneity of biclusters. As a result, pattern-based biclustering provides the unprecedented opportunity to discover non-trivial yet meaningful biological modules with putative functions, whose coherency and tolerance to noise can be tuned and made problem-specific. METHODS: To enable the effective use of pattern-based biclustering by the scientific community, we developed BicPAMS (Biclustering based on PAttern Mining Software), a software that: 1) makes available state-of-the-art pattern-based biclustering algorithms (BicPAM (Henriques and Madeira, Alg Mol Biol 9:27, 2014), BicNET (Henriques and Madeira, Alg Mol Biol 11:23, 2016), BicSPAM (Henriques and Madeira, BMC Bioinforma 15:130, 2014), BiC2PAM (Henriques and Madeira, Alg Mol Biol 11:1-30, 2016), BiP (Henriques and Madeira, IEEE/ACM Trans Comput Biol Bioinforma, 2015), DeBi (Serin and Vingron, AMB 6:1-12, 2011) and BiModule (Okada et al., IPSJ Trans Bioinf 48(SIG5):39-48, 2007)); 2) consistently integrates their dispersed contributions; 3) further explores additional accuracy and efficiency gains; and 4) makes available graphical and application programming interfaces. RESULTS: Results on both synthetic and real data confirm the relevance of BicPAMS for biological data analysis, highlighting its essential role for the discovery of putative modules with non-trivial yet biologically significant functions from expression and network data. CONCLUSIONS: BicPAMS is the first biclustering tool offering the possibility to: 1) parametrically customize the structure, coherency and quality of biclusters; 2) analyze large-scale biological networks; and 3) tackle the restrictive assumptions placed by state-of-the-art biclustering algorithms. These contributions are shown to be key for an adequate, complete and user-assisted unsupervised analysis of biological data. SOFTWARE: BicPAMS and its tutorial available in http://www.bicpams.com . Rui Henriques, Francisco L. Ferreira, Sara C. Madeira |
BMC Bioinform. | 1 |
| 2017 | Erratum to: BicPAMS: software for biological data analysis with pattern-based biclustering
Rui Henriques, Francisco L. Ferreira, Sara C. Madeira |
BMC Bioinform. | 1 |
| 2015 | BicNET: Efficient Biclustering of Biological Networks to Unravel Non-Trivial Modules
Rui Henriques, Sara C. Madeira |
WABI | 1 |
| 2015 | Generative modeling of repositories of health records for predictive tasks
Rui Henriques, Cláudia Antunes, Sara C. Madeira |
Data Min. Knowl. Discov. | 1 |
| 2015 | Multi-period classification: learning sequent classes from temporal domains
Rui Henriques, Sara C. Madeira, Cláudia Antunes |
Data Min. Knowl. Discov. | 1 |
| 2015 | A structured view on pattern mining-based biclustering
Rui Henriques, Cláudia Antunes, Sara C. Madeira |
Pattern Recognit. | 1 |
| 2015 | Biclustering with Flexible Plaid Models to Unravel Interactions between Biological ProcessesabstractGenes can participate in multiple biological processes at a time and thus their expression can be seen as a composition of the contributions from the active processes. Biclustering under a plaid assumption allows the modeling of interactions between transcriptional modules or biclusters (subsets of genes with coherence across subsets of conditions) by assuming an additive composition of contributions in their overlapping areas. Despite the biological interest of plaid models, few biclustering algorithms consider plaid effects and, when they do, they place restrictions on the allowed types and structures of biclusters, and suffer from robustness problems by seizing exact additive matchings. We propose BiP (Biclustering using Plaid models), a biclustering algorithm with relaxations to allow expression levels to change in overlapping areas according to biologically meaningful assumptions (weighted and noise-tolerant composition of contributions). BiP can be used over existing biclustering solutions (seizing their benefits) as it is able to recover excluded areas due to unaccounted plaid effects and detect noisy areas non-explained by a plaid assumption, thus producing an explanatory model of overlapping transcriptional activity. Experiments on synthetic data support BiP's efficiency and effectiveness. The learned models from expression data unravel meaningful and non-trivial functional interactions between biological processes associated with putative regulatory modules. Rui Henriques, Sara C. Madeira |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2014 | BicSPAM: flexible biclustering using sequential patternsabstractBACKGROUND: Biclustering is a critical task for biomedical applications. Order-preserving biclusters, submatrices where the values of rows induce the same linear ordering across columns, capture local regularities with constant, shifting, scaling and sequential assumptions. Additionally, biclustering approaches relying on pattern mining output deliver exhaustive solutions with an arbitrary number and positioning of biclusters. However, existing order-preserving approaches suffer from robustness, scalability and/or flexibility issues. Additionally, they are not able to discover biclusters with symmetries and parameterizable levels of noise. RESULTS: We propose new biclustering algorithms to perform flexible, exhaustive and noise-tolerant biclustering based on sequential patterns (BicSPAM). Strategies are proposed to allow for symmetries and to seize efficiency gains from item-indexable properties and/or from partitioning methods with conservative distance guarantees. Results show BicSPAM ability to capture symmetries, handle planted noise, and scale in terms of memory and time. BicSPAM also achieves the best match-scores for the recovery of hidden biclusters in synthetic datasets with varying noise distributions and levels of missing values. Finally, results on gene expression data lead to complete solutions, delivering new biclusters corresponding to putative modules with heightened biological relevance. CONCLUSIONS: BicSPAM provides an exhaustive way to discover flexible structures of order-preserving biclusters. To the best of our knowledge, BicSPAM is the first attempt to deal with order-preserving biclusters that allow for symmetries and that are robust to varying levels of noise. Rui Henriques, Sara C. Madeira |
BMC Bioinform. | 1 |
| 2013 | Accessing Emotion Patterns from Affective Interactions Using Electrodermal ActivityabstractMeasuring evocative emotions in affective interactions has become a critical step for effective engagements with computers. Electro dermal activity is believed to accurately isolate sympathetic responses, revealing paths to excitement, attention and arousal, and to differentiate emotional states. However, the inability to deal with varying amplitude and length of responses across individuals has led to its use as a simple intensity barometer. Thus, two questions remain unanswered. To which extent can electro dermal activity be used to recognize emotions? How do electro dermal responses vary between human-to-human and human-to-robot interactions? To answer these questions, we propose a new method to mine the signal that surpasses the referred limitations, and conduct an extensive experiment to study the responses to emotion-evocative stimuli across different settings. Observations reveal emerging electro dermal patterns for each emotion and attractive accuracy levels for emotion recognition that increases when there is a link to the psychological traits of the subjects. Rui Henriques, Ana Paiva 0001, Cláudia Antunes |
ACII | 1 |
| 2013 | Sensors in the wild: exploring electrodermal activity in child-robot interaction
Iolanda Leite, Rui Henriques, Carlos Martinho, Ana Paiva 0001 |
HRI | 2 |