Pedro Pereira Rodrigues

dblp:r/PedroPereiraRodrigues · DBLP profile ↗
← Back
38ranked-venue papers
9as first author
5since 2021 · last 2024
0000-0001-7867-6682ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 15 · 2 first-authorDatabases, data management, data science and information retrieval · 7 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2024 Imputation of data Missing Not at Random: Artificial generation and benchmark analysis
abstract
Experimental assessment of different missing data imputation methods often compute error rates between the original values and the estimated ones. This experimental setup relies on complete datasets that are injected with missing values. The injection process is straightforward for the Missing Completely At Random and Missing At Random mechanisms; however, the Missing Not At Random mechanism poses a major challenge, since the available artificial generation strategies are limited. Furthermore, the studies focused on this latter mechanism tend to disregard a comprehensive baseline of state-of-the-art imputation methods. In this work, both challenges are addressed: four new Missing Not At Random generation strategies are introduced and a benchmark study is conducted to compare six imputation methods in an experimental setup that covers 10 datasets and five missingness levels (10% to 80%). The overall findings are that, for most missing rates and datasets, the best imputation method to deal with Missing Not At Random values is the Multiple Imputation by Chained Equations, whereas for higher missingness rates autoencoders show promising results.
Ricardo Cardoso Pereira, Pedro H. Abreu, Pedro Pereira Rodrigues, Mário A. T. Figueiredo
Expert Syst. Appl.3
2024 Call for Papers: Data Generation in Healthcare Environments
Ricardo Cardoso Pereira, Pedro Pereira Rodrigues, Irina S. Moreira, Pedro H. Abreu
J. Biomed. Informatics2
2022 Partial Multiple Imputation With Variational Autoencoders: Tackling Not at Randomness in Healthcare Data
abstract
Missing data can pose severe consequences in critical contexts, such as clinical research based on routinely collected healthcare data. This issue is usually handled with imputation strategies, but these tend to produce poor and biased results under the Missing Not At Random (MNAR) mechanism. A recent trend that has been showing promising results for MNAR is the use of generative models, particularly Variational Autoencoders. However, they have a limitation: the imputed values are the result of a single sample, which can be biased. To tackle it, an extension to the Variational Autoencoder that uses a partial multiple imputation procedure is introduced in this work. The proposed method was compared to 8 state-of-the-art imputation strategies, in an experimental setup with 34 datasets from the medical context, injected with the MNAR mechanism (10% to 80% rates). The results were evaluated through the Mean Absolute Error, with the new method being the overall best in 71% of the datasets, significantly outperforming the remaining ones, particularly for high missing rates. Finally, a case study of a classification task with heart failure data was also conducted, where this method induced improvements in 50% of the classifiers.
Ricardo Cardoso Pereira, Pedro H. Abreu, Pedro Pereira Rodrigues
IEEE J. Biomed. Health Informatics3
2021 GANs for Tabular Healthcare Data Generation: A Review on Utility and Privacy
João Coutinho-Almeida, Pedro Pereira Rodrigues, Ricardo João Cruz Correia
DS2
2021 CMIID: A comprehensive medical information identifier for clinical search harmonization in Data Safe Havens
Michael A. P. Domingues, Rui Camacho, Pedro Pereira Rodrigues
J. Biomed. Informatics3
2020 Missing Image Data Imputation using Variational Autoencoders with Weighted Loss
Ricardo Cardoso Pereira, Joana Cristo Santos, José Pereira Amorim, Pedro Pereira Rodrigues, Pedro H. Abreu
ESANN4
2020 VAE-BRIDGE: Variational Autoencoder Filter for Bayesian Ridge Imputation of Missing Data
abstract
The missing data issue is often found in real-world datasets and it is usually handled with imputation strategies that replace the missing values with new data. Recently, generative models such as Variational Autoencoders have been applied for this imputation task. However, they were always used to perform the entire imputation, which has presented limited results when comparing to other state-of-the-art methods. In this work, a new approach called Variational Autoencoder Filter for Bayesian Ridge Imputation is introduced. It uses a Variational Autoencoder at the beginning of the imputation pipeline to filter the instances that are later fitted to a Bayesian ridge regression used to predict the new values. The approach was compared to four state-of-the-art imputation methods using 10 datasets from the healthcare context covering clinical trials, all injected with missing values under different rates. The proposed approach significantly outperformed the remaining methods in all settings, achieving an overall improvement between 26% and 67%.
Ricardo Cardoso Pereira, Pedro H. Abreu, Pedro Pereira Rodrigues
IJCNN3
2020 Reviewing Autoencoders for Missing Data Imputation: Technical Trends, Applications and Outcomes
abstract
Missing data is a problem often found in real-world datasets and it can degrade the performance of most machine learning models. Several deep learning techniques have been used to address this issue, and one of them is the Autoencoder and its Denoising and Variational variants. These models are able to learn a representation of the data with missing values and generate plausible new ones to replace them. This study surveys the use of Autoencoders for the imputation of tabular data and considers 26 works published between 2014 and 2020. The analysis is mainly focused on discussing patterns and recommendations for the architecture, hyperparameters and training settings of the network, while providing a detailed discussion of the results obtained by Autoencoders when compared to other state-of-the-art methods, and of the data contexts where they have been applied. The conclusions include a set of recommendations for the technical settings of the network, and show that Denoising Autoencoders outperform their competitors, particularly the often used statistical methods.
Ricardo Cardoso Pereira, Miriam Seoane Santos, Pedro Pereira Rodrigues, Pedro H. Abreu
J. Artif. Intell. Res.3
2019 Guest Editorial Small Things and Big Data: Controversies and Challenges in Digital Healthcare
abstract
The papers in this special section focus on the challenges faced in the digital healthcare market. Recent advances in information and communication technologies (ICT), as well as biomedical engineering, sensor technology and data science, have acted as catalysts for significant developments in the sector of health care, strongly affecting medical diagnosis, patient and healthcare management, disease treatment and health education. In fact, small wearable, disposable sensors, implantable devices or medical devices, as well as elementary services are being featured as keys for monitoring health and facilitating well-being.
Panagiotis D. Bamidis, Stathis Th. Konstantinidis, Pedro Pereira Rodrigues, Sameer K. Antani, Daniela Giordano
IEEE J. Biomed. Health Informatics3
2018 Finding Groups in Obstructive Sleep Apnea Patients: A Categorical Cluster Analysis
abstract
Obstructive sleep apnea (OSA) is a significant sleep problem with various clinical presentations that have not been formally characterized. This poses critical challenges for its recognition, resulting in missed or delayed diagnosis. Recently, cluster analysis has been used in different clinical domains, particularly within numeric variables. We applied an extension of k-means to be used in categorical variables: k-modes, to identify groups of OSA patients. Demographic, physical examination, clinical history, and comorbidities characterization variables (n=46) were collected from 318 patients; missing values were all imputed with k-nearest neighbors (k-NN). Feature selection, through Chi-square test, was executed and 17 variables were inserted in cluster analysis, resulting in three clusters. Cluster 1 having an age between 65 and 90 years (54%), 78% of males, with the presence of diabetes and gastroesophageal reflux, and high OSA prevalence; Cluster 2 presented a lower percentage of OSA (46%), with middle-aged women without comorbidities, but with gastroesophageal reflux; and Cluster 3 was very similar to cluster 1, only differing in age (45-64) and comorbidities were not present. Our results suggest that there are different groups of OSA patients, creating the need to rethink the baseline characteristics of these patients before being sent to perform polysomnography (gold standard exam for diagnosis).
Daniela Ferreira Santos, Pedro Pereira Rodrigues
CBMS2
2018 Causality assessment of adverse drug reaction reports using an expert-defined Bayesian network
abstract
In pharmacovigilance, reported cases are considered suspected adverse drug reactions (ADR). Health authorities have thus adopted structured causality assessment methods, allowing the evaluation of the likelihood that a drug was the causal agent of an adverse reaction. The aim of this work was to develop and validate a new causality assessment support system used in a regional pharmacovigilance centre. A Bayesian network was developed, for which the structure was defined by experts while the parameters were learnt from 593 completely filled ADR reports evaluated by the Portuguese Northern Pharmacovigilance Centre medical expert between 2000 and 2012. Precision, recall and time to causality assessment (TTA) was evaluated, according to the WHO causality assessment guidelines, in a retrospective cohort of 466 reports (April-September 2014) and a prospective cohort of 1041 reports (January-December 2015). Additionally, a simplified assessment matrix was derived from the model, enabling its preliminary direct use by notifiers. Results show that the network was able to easily identify the higher levels of causality (recall above 80%), although struggling to assess reports with a lower level of causality. Nonetheless, the median (Q1:Q3) TTA was 4 (2:8) days using the network and 8 (5:14) days using global introspection, meaning the network allowed a faster time to assessment, which has a procedural deadline of 30 days, improving daily activities in the centre. The matrix expressed similar validity, allowing an immediate feedback to the notifiers, which may result in better future engagement of patients and health professionals in the pharmacovigilance system.
Pedro Pereira Rodrigues, Daniela Ferreira Santos, Jorge Polónia, Inês Ribeiro-Vaz
Artif. Intell. Medicine1
2017 Implementing Guidelines for Causality Assessment of Adverse Drug Reaction Reports: A Bayesian Network Approach
Pedro Pereira Rodrigues, Daniela Ferreira Santos, Jorge Polónia, Inês Ribeiro-Vaz
AIME1
2017 Anomaly Detection Through Temporal Abstractions on Intensive Care Data: Position Paper
abstract
A large amount of information is continuously generated in intensive health care. An analysis of these data streams can supply valuable insights to improve the monitoring of the patients. The volume, frequency and complexity of data, which come unlabeled, make their analysis a challenging task. Machine learning (ML) techniques have been successfully employed for mining data streams to extract useful knowledge for health care monitoring. It includes the detection of changes in the behavior of sensors, failures on machines or systems, and data anomalies. Anomaly (or outlier) detection is a ML task that aims to find exceptions or abnormalities in a dataset. These exceptions, in a medical context, can represent a new disease pattern, an event to be further investigated, behavior changes or potential health complications. Despite of its analysis in data streams is a challenging task, temporal abstractions techniques should help due to they deal with the management and abstraction of time based data, offering high level of visualization of each data object in its context. The aim of this paper is to review recent research in anomaly detection and temporal abstractions and discuss the application of their combination to intensive care data streams.
Giovana Jaskulski Gelatti, André C. P. L. F. de Carvalho, Pedro Pereira Rodrigues
CBMS3
2017 Bringing Bayesian Networks to Bedside: A Web-Based Framework
abstract
Bayesian networks are one of the most intuitive statistical models for both estimation, classification and prediction of patients outcomes. However, the availability of inference software in clinical settings is still limited. This work presents preliminary steps towards the creation of simple web-based forms that can access a powerful Bayesian network inference engine, making the derived models usable at bedside by both the clinicians and the patients themselves.
Raphael Oliveira, Joana Ferreira 0002, Diogo Libânio, Cláudia Camila Dias, Pedro Pereira Rodrigues
CBMS5
2017 Improving Diagnosis in Obstructive Sleep Apnea with Clinical Data: A Bayesian Network Approach
abstract
In obstructive sleep apnea, respiratory effort is maintained but ventilation decreases/disappears because of the partial/total occlusion in the upper airway. It affects about 4% of men and 2% of women in the world population. The aim was to define an auxiliary diagnostic method that can support the decision to perform polysomnography (standard test), based on risk and diagnostic factors. Our sample performed polysomnography between January and May 2015. Two Bayesian classifiers were used to build the models: Naïve Bayes (NB) and Tree augmented Naïve Bayes (TAN), using all 39 variables or just a selection of 13. Area under the ROC curve, sensitivity, specificity, predictive values were evaluated using cross-validation. From a collected total of 241 patients, only 194 fulfill the inclusion criteria. 123 (63%) were male, with a mean age of 58 years old. 66 (34%) patients had a normal result and 128 (66%) a diagnostic of obstructive sleep apnea. The AUCs for each model were: NB39 - 72%; TAN39 - 79%; NB13 - 75% and TAN13 - 75%. The high (34%) proportion of normal results confirm the need for a preevaluation prior to polysomnography. The constant seeking of a validated model to screen patients with suspicion of obstructive sleep apnea is essential, especially at the level of primary care.
Daniela Ferreira Santos, Pedro Pereira Rodrigues
CBMS2
2016 Disabling and Reoperation in Patients with Crohn's Disease Subject to Early Surgery or Immunosuppression: A Bayesian Network Prognostic Model
abstract
Crohn's disease is one type of inflammatory bowel disease whose incidence is currently increasing, subject to relapse and disabling, with unknown etiology, and usually diagnosed between the second and third decade of life. The aim of this work is to develop a Bayesian network tool to predict disabling and reoperation in patients with Crohn's disease subject to early surgery or immunosuppressors intake. Multi-centric study data from patients with surgery or immunosuppression in the first six months after diagnosis was used, focusing on the prognosis and the analysis of factors' interaction. Patients were grouped by the index episode: immunosuppressors intake, and surgery (stratified considering the use or not of immunosuppressors 6 months after surgery). Patient group was associated with disease behavior, upper gastrointestinal tract location (L4) and age at diagnosis, while disease extent was associated to perianal disease. For disabling, association between perianal disease and gender and location was also found. Association between gender and L4 was also found for reoperation. The cross-validated discriminative power of the models were high for both disabling (above 70%) and reoperation (above 80%). The generated models presented interesting insights on factor interaction and predictive ability for the prognosis, supporting their use in future clinical decision support systems.
Cláudia Camila Dias, Fernando Magro, Pedro Pereira Rodrigues
CBMS3
2015 Preliminary Study for a Bayesian Network Prognostic Model for Crohn's Disease
abstract
Crohn's disease is one type of inflammatory bowel disease whose incidence is currently increasing, and may affect any part of both the small and large intestine, possibly irritating deeper layers of the organs. Being a chronic disease, neither treatment nor surgery actually heals the patients. Thus, focus has been given to identifying good prognostic models based on clinical factors since they are more easily included in daily practice. The aim of this work is to provide an initial study on the adequacy of a Bayesian network model to enhance the prognosis prediction for patients with Crohn's disease. Multicentric study data of patients with surgery or immuno suppression in the six month after diagnosis was used to derive a Bayesian network, focusing on the prognosis and the analysis of factors interaction, including clinical features, disease course, treatment, follow-up plan, and adverse events. Two models were evaluated (naïve Bayes and Tree-Augmented Naïve Bayes) and also compared with logistic regression, using cross-validation and ROC curve analysis. Preliminary results showed competitive accuracy (above 75%) and discriminative power (above 70%). The generated models presented interesting insights on factor interaction and predictive ability for the prognosis, supporting their use in future clinical decision support systems.
Cláudia Camila Dias, Fernando Magro, Pedro Pereira Rodrigues
CBMS3
2015 Obstructive Sleep Apnea Diagnosis: The Bayesian Network Model Revisited
abstract
Obstructive Sleep Apnea (OSA) is a disease that affects approximately 4% of men and 2% of women worldwide but is still underestimated and underdiagnosed. The standard method for assessing this index, and therefore defining the OSA diagnosis, is polysomnography (PSG). Previous work developed relevant Bayesian network models but those were based only on variables univariatedly associated with the outcome, yielding a bias on the possible knowledge representation of the models. The aim of this work was to develop and validate new Bayesian network decision support models that could be used during sleep consult to assess the need for PSG. Bayesian models were developed using a) expert opinion, b) hill-climbing, c) naïve Bayes and d) TAN structures. Resulting models validity was assessed with in-sample AUC and stratified cross-validation, also comparing with previously published model. Overall, models achieved good discriminative power (AUC>70%) and validity (measures consistently above 70%). Main conclusions are a) the need to integrate a wider range of variables in the final models and b) the support of using Bayesian networks in the diagnosis of obstructive sleep apnea.
Pedro Pereira Rodrigues, Daniela Ferreira Santos, Liliana Leite
CBMS1
2015 Medical Mining: KDD 2015 Tutorial
abstract
In year 2015, we experience a proliferation of scientific publications, conferences and funding programs on KDD for medicine and healthcare. However, medical scholars and practitioners work differently from KDD researchers: their research is mostly hypothesis-driven, not data-driven. KDD researchers need to understand how medical researchers and practitioners work, what questions they have and what methods they use, and how mining methods can fit into their research frame and their everyday business. Purpose of this tutorial is to contribute to this learning process. We address medicine and healthcare; there the expertise of KDD scholars is needed and familiarity with medical research basics is a prerequisite. We aim to provide basics for (1) mining in epidemiology and (2) mining in the hospital. We also address, to a lesser extent, the subject of (3) preparing and annotating Electronic Health Records for mining.
Myra Spiliopoulou, Pedro Pereira Rodrigues, Ernestina Menasalvas Ruiz
KDD2
2015 Probabilistic change detection and visualization methods for the assessment of temporal stability in biomedical data quality
Carlos Sáez 0001, Pedro Pereira Rodrigues, João Gama 0001, Montserrat Robles, Juan Miguel García-Gómez
Data Min. Knowl. Discov.2
2014 Need and Requirements Elicitation for Electronic Access to Patient's Medication History in the Emergency Department
abstract
Electronic access to patient's medication history (PMH) in the emergency department (ED) in Portugal is not widely granted, nor has the importance of such access been clearly assessed. Given the known association between poor PMH and medication errors, the goal of this study was to gather requirements for such a system, assessing physicians' opinions regarding the importance of having access to PMH in the ED. A questionnaire was sent to all Portuguese public hospitals which approved the study, and forwarded by email by the internal services of each hospital to ED physicians. Fourteen hospitals authorized the study, from which 83 ED physicians answered the questionnaire. PMH-related information considered most important focused on medication name and posology (>90%) and date and dose of prescription (>80%), but also date of dispensing of medications (>40%). Other information such as allergies (99%) and adverse reactions (96%) were similarly considered important, and physicians agree with the inclusion of non-prescription medications (85%) as well as homeopathic medicines (64%). Overall, access to PMH in the ED appears to be important and present benefits to patients' care. Given this, electronic access to PHM should be settled in Portuguese ED.
Margarida David, Fernando Rosa, Pedro Pereira Rodrigues
CBMS3
2014 Using Probabilistic Graphical Models to Enhance the Prognosis of Health-Related Quality of Life in Adult Survivors of Critical Illness
abstract
Health-related quality of life (HR-QoL) is a subjective concept, reflecting the overall mental and physical state of the patient, and their own sense of well-being. Estimating current and future QoL has become a major outcome in the evaluation of critically ill patients. The aim of this study is to enhance the inference process of 6 weeks and 6 months prognosis of QoL after intensive care unit (ICU) stay, using the EQ-5D questionnaire. The main outcomes of the study were the EQ-5D five main dimensions: mobility, self-care, usual activities, pain and anxiety depression. For each outcome, three Bayesian classifiers were built and validated with 10-fold cross-validation. Sixty and 473 patients (6 weeks and 6 months, respectively) were included. Overall, 6 months QoL is higher than 6 weeks, with the probability of absence of problems ranging from 31% (6 weeks mobility) to 72% (6 months self-care). Bayesian models achieved prognosis accuracies of 56% (6 months, anxiety depression) up to 80% (6 weeks, mobility). The prognosis inference process for an individual patient was enhanced with the visual analysis of the models, showing that women, elderly, or people with longer ICU stay have higher risk of QoL problems at 6 weeks. Likewise, for the 6 months prognosis, a higher APACHE II severity score also leads to a higher risk of problems, except for anxiety depression where the youngest and active have increased risk. Bayesian networks are competitive with less descriptive strategies, improve the inference process by incorporating domain knowledge and present a more interpretable model. The relationships among different factors extracted by the Bayesian models are in accordance with those collected by previous state-of-the-art literature, hence showing their usability as inference model.
Cláudia Camila Dias, Cristina Granja, Altamiro da Costa Pereira, João Gama 0001, Pedro Pereira Rodrigues
CBMS5
2014 Can We Avoid Unnecessary Polysomnographies in the Diagnosis of Obstructive Sleep Apnea? A Bayesian Network Decision Support Tool
abstract
Obstructive Sleep Apnea (OSA) affects 2-4% of the population worldwide. The standard test for OSA diagnosis is polysomnography (PSG), an expensive exam limited to urban areas. Furthermore, nearly half of all PSG tests results are negative for OSA. This work aims to reduce these unnecessary exams, by defining an auxiliary diagnostic method that could be used to assess patient's need for PSG, according to their probability of OSA diagnosis. A prospective study was conducted on adult patients with OSA suspicion who performed PSG at our sleep laboratory in Portugal. The studied clinical variables were defined after literature review and collected during consultation. Two comparable cohorts were studied for derivation (n=86) and validation (n=33) of models. Three classifiers were analyzed - a multiple logistic regression classifier (AUC=80.0%) and two Bayesian networks classifiers - Naïve Bayes (AUC=81.3%) and Tree Augmented Naïve Bayes (TAN, AUC=81.4%) - aiming at the best possible specificity (identification of unnecessary exams). Overall, sensitivity-adjusted models could detect normal patients, preventing unnecessary PSG, while keeping sensitivity high. Furthermore, the graphical representation of TAN can be explored by the physician during consultation, making it a helpful tool to assess patients' need to perform PSG.
Liliana Leite, Cristina Costa-Santos, Pedro Pereira Rodrigues
CBMS3
2014 Enhancing data stream predictions with reliability estimators and explanation
Zoran Bosnic, Jaka Demsar, Grega Kespret, Pedro Pereira Rodrigues, João Gama 0001, Igor Kononenko 0001
Eng. Appl. Artif. Intell.4
2013 Telenursing in colorectal cancer patient follow-up and treatment assessment: A mixed methods evaluation study
abstract
The incidence of colorectal cancer cases in the Portuguese Institute of Oncology of Porto created the need of a telenursing program in the Gastro-Intestinal Cancer Unit. After staging, treatment may involve surgery radio and chemotherapy (either oral or IV). Patients with no treatment after surgery are scheduled for medical exams every 3 months in the first 2 years. Patients on chemotherapy need to be compliant and to have a close monitoring of adverse events. The GI Cancer Unit uses a telenursing information system to help assess colorectal cancer patients' follow-up after surgery, medical treatment compliance and adverse events. A mixed-methods evaluation was done to a) describe the target population, b) detect problems in the telenursing information system, and c) suggest changes to meet users' requirements. From 181 outbound phone calls, representing 67 patients (49 in treatment and 18 in follow-up), patients' main characteristics were extracted and system's problems were identified by the intervening nurses. Recommendations will be useful for a further development of the system.
Maria Jose Dias, Maria Fragoso, Lucio Lara-Santos, Pedro Pereira Rodrigues
CBMS4
2013 Learner's satisfaction within a breast imaging eLearning course for radiographers
abstract
Background: An asynchronous eLearning system was developed for radiographers in order to promote a better knowledge about senology and mammography. Objectives: to assess the learners' satisfaction. Methods: Target population included radiographers and radiography students, in order to assess eLearning satisfaction according to different experience levels in breast imaging. Satisfaction was measured through a questionnaire developed especially for eLearning systems, using a seven-point Likert scale. Main topics related are content, interface, personalization and learning community. Results: Overall, 85% of learners were satisfied with the course and 87,5% considered that the course is successful. Main areas that were evaluated by most learners in a positive way were interface and content (between six and seven-point); on the other hand, learning community presented a wider distribution of answers. Conclusions: The course provides an overall high degree of learner satisfaction, thus providing more effective knowledge gain on breast imaging for radiographers.
Inês C. Moreira, Sandra M. Rua Ventura, Pedro Pereira Rodrigues
CBMS4
2013 Predicting visualization of hospital clinical reports using survival analysis of access logs from a virtual patient record
abstract
The amount of data currently being produced, stored and used in hospital settings is stressing information technology infrastructure, making clinical reports to be stored in secondary memory devices. The aim of this work was to develop a model that predicts the probability of visualization, within a certain period after production, of each clinical report. We collected log data, from January 2013 till May 2011, from an existing virtual patient record, in a tertiary university hospital in Porto, Portugal, with information on report creation and report first-time visualization dates, along with contextual information. The main factors associated with visualization were defined using logistic regression. These factors were then used as explanatory variables for predicting the probability of a piece of information being accessed after production, using Kaplan-Meier analysis and the Weibull probability distribution. Clinical department, type of encounter and report type were found significantly associated with time-to-visualization and probability of visualization.
Pedro Pereira Rodrigues, Cláudia Camila Dias, Diana Rocha, Isabel Boldt, Armando Teixeira-Pinto, Ricardo João Cruz Correia
CBMS1
2013 An automatic clinical document importance estimator for an existing electronic patient record - Architecture and implementation
abstract
The goal of the OPTIM project is to optimize the graphical user interface of an electronic health record (EHR) by predicting clinical documents' relevance and provide a ranked list of relevant documents for the given user at a certain time. This paper describes the architecture of the relevance assignment and ranking prototype and some implementation issues. The prototype's design is based on two components: OPTIM Core, with logical representation, estimation server's integration and the webservice layer, and the OPTIM WebUI, with the user interface for presenting the results. The prototype was tested in integration with an EHR using a simulated environment. The results were encouraging but yet they revealed a certain lack of security (confidentiality). It has now the capacity of rating 10 documents per second. Nonetheless, the integration of features such as rating clinical relevance based on mathematical models can be included in existing EHR potentially improving their usability.
Pedro Pereira Rodrigues, Ricardo João Cruz Correia
CBMS2
2013 Contextual anomalies in medical data
abstract
Anomalies in data can cause a lot of problems in the data analysis processes. Thus, it is necessary to improve data quality by detecting and eliminating errors and inconsistencies in the data, known as the data cleaning process [1]. Since detection and correction of anomalies requires detailed domain knowledge, the involvement of experts in the field is essential to the success of the process of cleaning the data. However, considering the size of data to be processed, this process should be as automatic as possible so as to minimize the time spent [1].
Daniela Vasco, Pedro Pereira Rodrigues, João Gama 0001
CBMS2
2013 On evaluating stream learning algorithms
João Gama 0001, Raquel Sebastião, Pedro Pereira Rodrigues
Mach. Learn.3
2011 Clustering distributed sensor data streams using local processing and reduced communication
abstract
Nowadays applications produce infinite streams of data distributed across wide sensor networks. In this work we study the problem of continuously maintain a cluster structure over the data points generated by the entire network. Usual techniques ope
João Gama 0001, Pedro Pereira Rodrigues, Luís M. B. Lopes
Intell. Data Anal.2
2009 Issues in evaluation of stream learning algorithms
abstract
Learning from data streams is a research area of increasing importance. Nowadays, several stream learning algorithms have been developed. Most of them learn decision models that continuously evolve over time, run in resource-aware environments, detect and react to changes in the environment generating data. One important issue, not yet conveniently addressed, is the design of experimental work to evaluate and compare decision models that evolve over time. There are no golden standards for assessing performance in non-stationary environments. This paper proposes a general framework for assessing predictive stream learning algorithms. We defend the use of Predictive Sequential methods for error estimate - the prequential error. The prequential error allows us to monitor the evolution of the performance of models that evolve over time. Nevertheless, it is known to be a pessimistic estimator in comparison to holdout estimates. To obtain more reliable estimators we need some forgetting mechanism. Two viable alternatives are: sliding windows and fading factors. We observe that the prequential error converges to an holdout estimator when estimated over a sliding window or using fading factors. We present illustrative examples of the use of prequential error estimators, using fading factors, for the tasks of: i) assessing performance of a learning algorithm; ii) comparing learning algorithms; iii) hypothesis testing using McNemar test; and iv) change detection using Page-Hinkley test. In these tasks, the prequential error estimated using fading factors provide reliable estimators. In comparison to sliding windows, fading factors are faster and memory-less, a requirement for streaming applications. This paper is a contribution to a discussion in the good-practices on performance assessment when learning dynamic models that evolve over time.
João Gama 0001, Raquel Sebastião, Pedro Pereira Rodrigues
KDD3
2009 A system for analysis and prediction of electricity-load streams
abstract
Sensors distributed all around electrical-power distribution networks produce streams of data at high-speed. From a data mining perspective, this sensor network problem is characterized by a large number of variables (sensors), producing a continuous
Pedro Pereira Rodrigues, João Gama 0001
Intell. Data Anal.1
2008 Robust Division in Clustering of Streaming Time Series
abstract
Online learning algorithms which address fast data streams should process examples at the rate they arrive, using a single scan of data and fixed memory, maintaining a decision model at any time and being able to adapt the model to the most recent data. These features yield the necessity of using approximate models. One problem that usually arises with approximate models is the definition of a minimum number of observations necessary to assure convergence, which implies a high risk since the system may have to decide based only on a small subset of the entire data. One approach is to apply techniques based on the Hoeffding bound to enforce decisions with a confidence level. In divisive clustering of time series, the goal is to find clusters of similar time series over time. In online approaches there are two decisions to make: when to split and how to assign variables to new clusters. We can define a confidence level to both the decision of splitting and the assignment of data variables to new clusters. Previous works have already addressed confident decisions on the moment of split. Our proposal is to include a confidence level to the assignment process. When a split point is reported, creating two new clusters, we can directly assign points which are confidently closer to one cluster than the other, having a different strategy for those variables which do not satisfy the confidence level. In this paper we propose to assign the unsure variables to a third cluster. Experimental evaluation is presented in the context of a recently proposed hierarchical algorithm, assessing the advantages of the proposal, revealing also advantages on memory usage reduction and processing speed. Although this proposal is evaluated under the scope of an existent method, it can be generalized to any divisive procedure.
Pedro Pereira Rodrigues, João Gama 0001
ECAI1
2008 Clustering Distributed Sensor Data Streams
Pedro Pereira Rodrigues, João Gama 0001, Luís M. B. Lopes
ECML/PKDD (2)1
2008 Hierarchical Clustering of Time-Series Data Streams
abstract
This paper presents and analyzes an incremental system for clustering streaming time series. The Online Divisive-Agglomerative Clustering (ODAC) system continuously maintains a tree-like hierarchy of clusters that evolves with data, using a top-down strategy. The splitting criterion is a correlation-based dissimilarity measure among time series, splitting each node by the farthest pair of streams. The system also uses a merge operator that reaggregates a previously split node in order to react to changes in the correlation structure between time series. The split and merge operators are triggered in response to changes in the diameters of existing clusters, assuming that in stationary environments, expanding the structure leads to a decrease in the diameters of the clusters. The system is designed to process thousands of data streams that flow at a high rate. The main features of the system include update time and memory consumption that do not depend on the number of examples in the stream. Moreover, the time and memory required to process an example decreases whenever the cluster structure expands. Experimental results on artificial and real data assess the processing qualities of the system, suggesting a competitive performance on clustering streaming time series, exploring also its ability to deal with concept drift.
Pedro Pereira Rodrigues, João Gama 0001, João Pedro Pedroso
IEEE Trans. Knowl. Data Eng.1
2007 Stream-Based Electricity Load Forecast
João Gama 0001, Pedro Pereira Rodrigues
PKDD2
2006 ODAC: Hierarchical Clustering of Time Series Data Streams
abstract
This paper presents a time series whole clustering system that incrementally constructs a tree-like hierarchy of clusters, using a top-down strategy. The Online Divisive-Agglomerative Clustering (ODAC) system uses a correlation-based dissimilarity measure between time series over a data stream and possesses an agglomerative phase to enhance a dynamic behavior capable of concept drift detection. Main features include splitting and agglomerative criteria based on the diameters of existing clusters and supported by a significance level. At each new example, only the leaves are updated, reducing computation of unneeded dissimilarities and speeding up the process every time the structure grows. Experimental results on artificial and real data suggest competitive performance on clustering time series and show that the system is equivalent to a batch divisive clustering on stationary time series, being also capable of dealing with concept drift. With this work, we assure the possibility and importance of hierarchical incremental time series whole clustering in the data stream paradigm, presenting a valuable and usable option.
Pedro Pereira Rodrigues, João Gama 0001, João Pedro Pedroso
SDM1