EDBT 2026 Demo / reviewers in the wild / expert
João Gama 0001
dblp:g/JGama · also João Manuel Portela da Gama
· DBLP profile ↗
77ranked-venue papers in the field
13as first author
24since 2021 · last 2026
0000-0003-3357-1195ORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 58 (11 first)Database Systems & Data Management · 8Information Retrieval & Web Search · 4Knowledge Engineering, Semantic Web & Information Systems · 4 (2 first)Big Data, Cloud & Distributed Data Systems · 2Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conditional Motif-Based Graph Convolutional Network for Anomaly Detection in the Waste Management Network
Sara Oliveira, Shazia Tabassum, João Gama 0001, Ana Garcia, Pedro Santana |
IDA | 3 |
| 2026 | In-context Learning of Evolving Data Streams with Tabular Foundational ModelsabstractState-of-the-art data stream mining has long drawn from ensembles of the Very Fast Decision Tree, a seminal algorithm honored with the 2015 KDD Test-of-Time Award. However, the emergence of large tabular models, i.e., transformers designed for structured numerical data, marks a significant paradigm shift. These models move beyond traditional weight updates, instead employing in-context learning through prompt tuning. By using on-the-fly sketches to summarize unbounded streaming data, one can feed this information into a pre-trained model for efficient processing. This work bridges advancements from both areas, highlighting how transformers' implicit meta-learning abilities, pre-training on drifting natural data, and reliance on context optimization directly address the core challenges of adaptive learning in dynamic environments. Exploring real-time model adaptation, this research demonstrates that TabPFN, coupled with a simple sliding memory strategy, consistently outperforms ensembles of Hoeffding trees, such as Adaptive Random Forest, and Streaming Random Patches, across all non-stationary benchmarks. Afonso Lourenço, João Gama 0001, Eric P. Xing, Goreti Marreiros |
KDD (1) | 2 |
| 2026 | Fish swarm parameter self-tuning for data streams
Bruno M. Veloso, Hugo Deandrade Amorim Neto, Fernando B. Lima Neto, João Gama 0001 |
Data Min. Knowl. Discov. | 4 |
| 2026 | Unsupervised Concept Drift Detector for Data Streams With Varying Feature SpacesabstractData streams with varying feature spaces have received extensive attention recently, while the common concept drift in them remains underexplored. Unsupervised concept drift detectors can report potential drifts without class labels, making them suitable for practical scenarios where labeling is usually costly and difficult. However, existing unsupervised detectors usually operate under fixed feature spaces. To address this limitation, a Matching Degree Histogram-based unsupervised detector for data streams with Varying Feature Spaces (MDH-VFS) is proposed. Changes in input features are refined into four scenarios, specifying the sources of concept drifts in such data streams. Based on this, MDH-VFS monitors the distribution of each feature independently using the fix-slide windows model. A matching degree-based histogram (MD-Histogram) supporting online updating is proposed to model data distribution. MD-Histogram requires no prior distributions and captures data change more sensitively than traditional histograms. The dissimilarity between two MD-Histograms is measured by the Hellinger distance, and drift is detected using an adaptive thresholding strategy. Both the drift positions and drift features can be reported. Experimental results show that MDH-VFS can not only effectively detect drifts in data streams with varying feature spaces (achieving average F1-score/MCC above 77% and outperforming nine existing detectors with improvements of at least 43%), but also improve the classification performance of downstream learning algorithms (reaching a maximum average accuracy of 88% and yielding up to 7.23% improvement). Ruirui Zhao, Jiang Jiang 0001, João Gama 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Unveiling Fairness and Performance of Causal DiscoveryabstractData-driven decision models based on Artificial Intelligence (AI) are increasingly adopted across domains. However, these models are susceptible to bias that can result in unfair or discriminatory outcomes. Recent research has explored causal discovery methods as a promising way to understand and improve fairness in decision-making systems. In this work, we investigate how different conditional independence tests used in constraint-based causal discovery algorithms, specifically the PC algorithm, affect fairness and performance. We perform an empirical evaluation on several datasets, including Portuguese public contracts, COMPAS, and the German Credit dataset. Using seven conditional independence tests, we assess model behavior under fairness (demographic parity, accuracy parity, equalized odds and predictive rate parity) and performance (accuracy, F1-score, AUC) metrics. Our findings reveal that some tests, due to their statistical properties, fail to expose unfairness detectable via causal structures, even when performance metrics appear acceptable. Furthermore, we highlight significant differences in computational efficiency among the tests, with x2-adf, sp-mi, and sp-x2 being the least efficient. This study underscores the need for careful selection of conditional independence tests in causal discovery to ensure both fairness and reliability in data-driven decision systems. Sónia Teixeira, Ana Rita Nogueira, João Gama 0001 |
DSAA | 3 |
| 2025 | RMIDDM: an unsupervised and interpretable concept drift detection method for data streams
Ruivaldo Lobão-Neto, Brenno de Mello Alencar, Heitor Murilo Gomes, Albert Bifet, João Gama 0001, Guilherme Weigert Cassales, Ricardo Araújo Rios |
Data Min. Knowl. Discov. | 5 |
| 2025 | Online learning from drifting capricious data streams with flexible Hoeffding tree
Ruirui Zhao, Yaqian You, João Gama 0001, Jiang Jiang 0001 |
Inf. Process. Manag. | 4 |
| 2025 | Modelling Concept Drift in Dynamic Data Streams for Recommender SystemsabstractRecommendation systems play a crucial role in modern e-commerce and streaming services. However, the limited availability of public datasets hampers the rapid development of more efficient and accurate recommendation algorithms within the research community. This work introduces a stream-based data generator designed to generate user preferences for a set of items while accommodating progressive changes in user preferences. The underlying principle involves using user/item embeddings to derive preferences by exploring the proximity of these embeddings. Whether randomly generated or learned from a real finite data stream, these embeddings serve as the basis for generating new preferences. We investigate how this fundamental model can adapt to shifts in user behavior over time; in our framework, changes correspond to alterations in the structure of the tripartite graph, reflecting modifications in the underlying embeddings. Through an analysis of real-life data streams, we demonstrate that the proposed model is effective in capturing actual preferences and the changes that they can exhibit over time. Thus, we characterize these changes and develop a generalized method capable of simulating realistic data, thereby generating streams with similar yet controllable drift dynamics. Luciano Caroprese, Francesco Sergio Pisani, Bruno M. Veloso, Matthias König 0005, Giuseppe Manco 0001, Holger H. Hoos, João Gama 0001 |
Trans. Recomm. Syst. | 7 |
| 2024 | Recent Advances in Learning from Data Streams
João Gama 0001 |
IC3K | 1 |
| 2024 | Super-Resolution Analysis for Landfill Waste Classification
Matías Molina, Rita P. Ribeiro, Bruno M. Veloso, João Gama 0001 |
IDA (1) | 4 |
| 2024 | S+t-SNE - Bringing Dimensionality Reduction to Data Streams
Pedro C. Vieira, João P. Montrezol, João T. Vieira, João Gama 0001 |
IDA (2) | 4 |
| 2024 | Improving hyper-parameter self-tuning for data streams by adapting an evolutionary approach
Antonio R. Moya, Bruno M. Veloso, João Gama 0001, Sebastián Ventura |
Data Min. Knowl. Discov. | 3 |
| 2024 | Forecasting financial market structure from network features using machine learning
Douglas Castilho 0001, Thársis Tuani Pinto Souza, Soong Moon Kang, João Gama 0001, André C. P. L. F. de Carvalho |
Knowl. Inf. Syst. | 4 |
| 2024 | From fault detection to anomaly explanation: A case study on predictive maintenanceabstractPredictive Maintenance applications are increasingly complex, with interactions between many components. Black-box models are popular approaches based on deep-learning techniques due to their predictive accuracy. This paper proposes a neural-symbolic architecture that uses an online rule-learning algorithm to explain when the black-box model predicts failures. The proposed system solves two problems in parallel: (i) anomaly detection and (ii) explanation of the anomaly. For the first problem, we use an unsupervised state-of-the-art autoencoder. For the second problem, we train a rule learning system that learns a mapping from the input features to the autoencoder’s reconstruction error. Both systems run online and in parallel. The autoencoder signals an alarm for the examples with a reconstruction error that exceeds a threshold. The causes of the signal alarm are hard for humans to understand because they result from a non-linear combination of sensor data. The rule that triggers that example describes the relationship between the input features and the autoencoder’s reconstruction error. The rule explains the failure signal by indicating which sensors contribute to the alarm and allowing the identification of the component involved in the failure. The system can present global explanations for the black box model and local explanations for why the black box model predicts a failure. We evaluate the proposed system in a real-world case study of Metro do Porto and provide explanations that illustrate its benefits. João Gama 0001, Rita P. Ribeiro, Saulo Martiello Mastelini, Narjes Davari, Bruno M. Veloso |
J. Web Semant. | 1 |
| 2023 | Knowledge-driven Analytics and Systems Impacting Human Quality of Life- Neurosymbolic AI, Explainable AI and BeyondabstractThe management of knowledge-driven artificial intelligence technologies is essential in order to evaluate their impact on human life and society. Social networks and tech use can have a negative impact on us physically, emotionally, socially and mentally. On the other hand, intelligent systems can have a positive effect on people's lives. Currently, we are witnessing the power of large language models (LLMs) like chatGPT and its influence towards the society. The objective of the workshop is to contribute to the advancement of intelligent technologies designed to address the human condition. This could include precise and personalized medicine, better care for elderly people, reducing private data leaks, using AI to manage resources better, using AI to predict risks, augmenting human capabilities, and more. The workshop's objective is to present research findings and perspectives that demonstrate how knowledge-enabled technologies and applications improve human well-being. This workshop indeed focuses on the impacts at different granularity levels made by Artificial Intelligence (AI) research on the micro granular level, where the daily or regular functioning of human life is affected, and also the macro granulate level, where the long-term or far-future effects of artificial intelligence on people's lives and the human society could be pretty high. In conclusion, this workshop explores how AI research can potentially address the most pressing challenges facing modern societies, and how knowledge management can potentially contribute to these solutions. Arijit Ukil, João Gama 0001, Antonio J. Jara, Leandro Marín |
CIKM | 2 |
| 2023 | Online Influence Forest for Streaming Anomaly Detection
Inês Martins, João S. Resende, João Gama 0001 |
IDA | 3 |
| 2023 | XAI for Predictive MaintenanceabstractThe field of Explainable Predictive Maintenance (PM) is concerned with developing methods that can clarify how AI systems operate in the PM domain. One of the challenges of creating maintenance plans is integrating AI output with human decision-making pro- cesses and expertise. For AI to be helpful and trustworthy, fault predictions must be contextualized and easily comprehensible to humans. This involves providing tailored explanations to different actors depending on their roles and needs. For example, engineers can be connected to technical installation blueprints, while man- agers can evaluate system downtime costs, and lawyers can assess safety-threatening failures' potential liability. In many industries, black-box AI systems analyze sensor data to predict failures by detecting anomalies and deviations from typical behavior with impressive accuracy. However, PM is just one part of a broader context that aims to identify the most probable causes, develop a recovery plan, and estimate remaining useful life while providing alternative solutions. Achieving this requires complex interactions among various actors in industrial and decision-making processes. Our tutorial explores current trends, and promising research directions in Explainable AI (XAI) relevant to Explainable Predictive Maintenance (XPM), and future challenges and open issues on this topic. We will also present three case studies that highlight XPM's challenges in bus and train operations and steel factories. João Gama 0001, Slawomir Nowaczyk, Sepideh Pashami, Rita P. Ribeiro, Grzegorz J. Nalepa, Bruno M. Veloso |
KDD | 1 |
| 2022 | A Fault Detection Framework Based on LSTM Autoencoder: A Case Study for Volvo Bus Data Set
Narjes Davari, Sepideh Pashami, Bruno M. Veloso, Slawomir Nowaczyk, Yuantao Fan, Pedro Mota Pereira, Rita P. Ribeiro, João Gama 0001 |
IDA | 8 |
| 2022 | Bank Statements to Network Features: Extracting Features Out of Time Series Using Visibility Graph
Nirbhaya Shaji, João Gama 0001, Rita P. Ribeiro |
IDA | 2 |
| 2021 | Predictive maintenance based on anomaly detection using deep learning for air production unit in the railway industryabstractPredictive maintenance methods assist early detection of failures and errors in machinery before they reach critical stages. This study proposes a data-driven predictive maintenance framework for the air production unit (APU) system of a train of Metro do Porto by deep learning based on a sparse autoencoder (SAE) network that efficiently detects abnormal data and considerably reduces the false alarm rate. Several analog and digital sensors installed on the APU system allow the detection of behavioral changes and deviations from the normal pattern by analyzing the collected data. We implemented two versions of the SAE network in which we inputted analog sensors data and digital sensors data, and the experimental results show that the failures due to air leakage problems are predicted by analog sensors data while other types of failures are identified by digital sensors data. A low pass filter is applied to the output of the SAE network, and a sequence of abnormal data is used as an alarm for the APU system failure. Performance indicators of the SAE network with digital sensors data, in terms of F1 Score, Recall, and Precision, are respectively, about 33.6%, 42%, and 28% better than those of the SAE network with analog sensors data. For comparison purposes, we also implemented a variational autoencoder (VAE). The results show that SAE performance is better than that of VAE by 14%, 77%, and 37% respectively, for Recall, Precision and F1 Score. Narjes Davari, Bruno M. Veloso, Rita P. Ribeiro, Pedro Mota Pereira, João Gama 0001 |
DSAA | 5 |
| 2021 | Hyper-parameter Optimization for Latent Spaces
Bruno M. Veloso, Luciano Caroprese, Matthias König 0005, Sónia Teixeira, Giuseppe Manco 0001, Holger H. Hoos, João Gama 0001 |
ECML/PKDD (3) | 7 |
| 2021 | Chebyshev approaches for imbalanced data streams regression models
Ehsan Aminian, Rita P. Ribeiro, João Gama 0001 |
Data Min. Knowl. Discov. | 3 |
| 2021 | Multi-aspect renewable energy forecasting
Roberto Corizzo, Michelangelo Ceci, Hadi Fanaee-T, João Gama 0001 |
Inf. Sci. | 4 |
| 2021 | Statistically Robust Evaluation of Stream-Based Recommender SystemsabstractOnline incremental models for recommendation are nowadays pervasive in both the industry and the academia. However, there is not yet a standard evaluation methodology for the algorithms that maintain such models. Moreover, online evaluation methodologies available in the literature generally fall short on the statistical validation of results, since this validation is not trivially applicable to stream-based algorithms. We propose ak-fold validation framework for the pairwise comparison of recommendation algorithms that learn from user feedback streams, using prequential evaluation. Our proposal enables continuous statistical testing on adaptive-size sliding windows over the outcome of the prequential process, allowing practitioners and researchers to make decisions in real time based on solid statistical evidence. We present a set of experiments to gain insights on the sensitivity and robustness of two statistical tests-McNemar's and Wilcoxon signed rank-in a streaming data environment. Our results show that besides allowing a real-time, fine-grained online assessment, the online versions of the statistical tests are at least as robust as the batch versions, and definitely more robust than a simple prequential single-fold approach. João Vinagre, Alípio Mário Jorge, Conceição Rocha, João Gama 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | AutoML for Stream k-Nearest Neighbors ClassificationabstractThe last few decades have witnessed a significant evolution of technology in different domains, changing the way the world operates, which leads to an overwhelming amount of data generated in an open-ended way as streams. Over the past years, we observed the development of several machine learning algorithms to process big data streams. However, the accuracy of these algorithms is very sensitive to their hyper-parameters, which requires expertise and extensive trials to tune. Another relevant aspect is the high-dimensionality of data, which can causes degradation to computational performance. To cope with these issues, this paper proposes a stream k-nearest neighbors (kNN) algorithm that applies an internal dimension reduction to the stream in order to reduce the resource usage and uses an automatic monitoring system that tunes dynamically the configuration of the kNN algorithm and the output dimension size with big data streams. Experiments over a wide range of datasets show that the predictive and computational performances of the kNN algorithm are improved. Maroua Bahri, Bruno M. Veloso, Albert Bifet, João Gama 0001 |
IEEE BigData | 4 |
| 2020 | Using Network Features for Credit Scoring in MicroFinance: Extended AbstractabstractThis paper uses non-traditional data, from a MicroFinance Institution (MFI), in a Credit Scoring loan classification problem and addresses a common problem in emerging markets of the lack of a verifiable customers' credit history. We perform a set of experiments to define a baseline model and prove the relevance of node embedding features, in credit scoring models, using a real world dataset. Paulo Paraíso, Saulo Ruiz, Luís Rodrigues 0005, João Gama 0001 |
DSAA | 5 |
| 2020 | Improving Prediction with Causal Probabilistic VariablesabstractThe application of feature engineering in classification problems has been commonly used as a means to increase the classification algorithms performance. There are already many methods for constructing features, based on the combination of attributes but, to the best of our knowledge, none of these methods takes into account a particular characteristic found in many problems: causality. In many observational data sets, causal relationships can be found between the variables, meaning that it is possible to extract those relations from the data and use them to create new features. The main goal of this paper is to propose a framework for the creation of new supposed causal probabilistic features, that encode the inferred causal relationships between the target and the other variables. In this case, an improvement in the performance was achieved when applied to the Random Forest algorithm. Ana Rita Nogueira, João Gama 0001, Carlos Ferreira 0007 |
IDA | 2 |
| 2020 | A drift detection method based on dynamic classifier selection
Felipe Azevedo Pinage, Eulanda M. dos Santos, João Gama 0001 |
Data Min. Knowl. Discov. | 3 |
| 2020 | BRIGHT - Drift-Aware Demand Predictions for Taxi NetworksabstractMassive data broadcast by GPS-equipped vehicles provide unprecedented opportunities. One of the main tasks in order to optimize our transportation networks is to build data-driven real-time decision support systems. However, the dynamic environments where the networks operate disallow the traditional assumptions required to put in practice many off-the-shelf supervised learning algorithms, such as finite training sets or stationary distributions. In this paper, we propose BRIGHT: a drift-aware supervised learning framework to predict demand quantities. BRIGHT aims to provide accurate predictions for short-term horizons through a creative ensemble of time series analysis methods that handles distinct types of concept drift. By selecting neighborhoods dynamically, BRIGHT reduces the likelihood of overfitting. By ensuring diversity among the base learners, BRIGHT ensures a high reduction of variance while keeping bias stable. Experiments were conducted using three large-scale heterogeneous real-world transportation networks in Porto (Portugal), Shanghai (China), and Stockholm (Sweden), as well as with controlled experiments using synthetic data where multiple distinct drifts were artificially induced. The obtained results illustrate the advantages of BRIGHT in relation to state-of-the-art methods for this task. Amal Saadallah, Luís Moreira-Matias, Ricardo Teixeira Sousa, Jihed Khiari, Erik Jenelius, João Gama 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2019 | BRIGHT - Drift-Aware Demand Predictions for Taxi Networks (Extended Abstract)abstractThe dynamic behavior of urban mobility patterns makes matching taxi supply with demand as one of the biggest challenges in this industry. Recently, the increasing availability of massive broadcast GPS data has encouraged the exploration of this issue under different perspectives. One possible solution is to build a data-driven real-time taxi-dispatching recommender system. However, existing systems are based on strong assumptions such as stationary demand distributions and finite training sets, which make them inadequate for modeling the dynamic nature of the network. In this paper, we propose BRIGHT: a drift-aware supervised learning framework which aims to provide accurate predictions for short-term horizon taxi demand quantities through a creative ensemble of time series analysis methods that handle distinct types of concept drift. A large experimental set-up which includes three real-world transportation networks and a synthetic test-bed with artificially inserted concept drifts, was employed to illustrate the advantages of BRIGHT when compared to S.o.A methods for this problem. Amal Saadallah, Luís Moreira-Matias, Ricardo Teixeira Sousa, Jihed Khiari, Erik Jenelius, João Gama 0001 |
ICDE | 6 |
| 2019 | Special Issue of DASFAA 2019abstracttechniques to solve this problem, and developed an inverted list and sketch-based approach. Guoliang Li 0001, João Gama 0001, Jun Yang 0001 |
Data Sci. Eng. | 2 |
| 2019 | Learning under Concept Drift: A ReviewabstractConcept drift describes unforeseeable changes in the underlying distribution of streaming data overtime. Concept drift research involves the development of methodologies and techniques for drift detection, understanding, and adaptation. Data analysis has revealed that machine learning in a concept drift environment will result in poor learning results if the drift is not addressed. To help researchers identify which research topics are significant and how to apply related techniques in data analysis tasks, it is necessary that a high quality, instructive review of current research developments and trends in the concept drift field is conducted. In addition, due to the rapid development of concept drift in recent years, the methodologies of learning under concept drift have become noticeably systematic, unveiling a framework which has not been mentioned in literature. This paper reviews over 130 high quality publications in concept drift related research areas, analyzes up-to-date developments in methodologies and techniques, and establishes a framework of learning under concept drift including three main components: concept drift detection, concept drift understanding, and concept drift adaptation. This paper lists and discusses 10 popular synthetic datasets and 14 publicly available benchmark datasets used for evaluating the performance of learning algorithms aiming at handling concept drift. Also, concept drift related research directions are covered and discussed. By providing state-of-the-art knowledge, this survey will directly support researchers in their understanding of research developments in the field of learning under concept drift. Jie Lu 0001, Anjin Liu, João Gama 0001, Guangquan Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2018 | Clustering in the Presence of Concept Drift
Richard Hugh Moulton, Herna L. Viktor, Nathalie Japkowicz, João Gama 0001 |
ECML/PKDD (1) | 4 |
| 2018 | Dynamic graph summarization: a tensor decomposition approach
Sofia Fernandes 0001, Hadi Fanaee-T, João Gama 0001 |
Data Min. Knowl. Discov. | 3 |
| 2018 | Forgetting techniques for stream-based matrix factorization in recommender systems
Pawel Matuszyk, João Vinagre, Myra Spiliopoulou, Alípio Mário Jorge, João Gama 0001 |
Knowl. Inf. Syst. | 5 |
| 2017 | The Initialization and Parameter Setting Problem in Tensor Decomposition-Based Link PredictionabstractLink prediction is the task of social network analysis whose goal is to predict the links that will appear in the network in future instants. Among the link predictors exploiting the time evolution of the networks, we can find the tensor decomposition-based methods. A major limitation of these methods is the lack of appropriate approaches for estimating their parameters and initialization. In this paper, we address this problem by proposing a parameter setting method. Our proposed approach resorts to optimization techniques to drive the search for an adequate parameter and initialization choice. Sofia Fernandes 0001, Hadi Fanaee-T, João Gama 0001 |
DSAA | 3 |
| 2016 | Online Semi-supervised Learning for Multi-target Regression in Data Streams Using AMRules
Ricardo Teixeira Sousa, João Gama 0001 |
IDA | 2 |
| 2016 | IoT Big Data Stream MiningabstractThe challenge of deriving insights from the Internet of Things (IoT) has been recognized as one of the most exciting and key opportunities for both academia and industry. Advanced analysis of big data streams from sensors and devices is bound to become a key area of data mining research as the number of applications requiring such processing increases. Dealing with the evolution over time of such data streams, i.e., with concepts that drift or change completely, is one of the core issues in IoT stream mining. This tutorial is a gentle introduction to mining IoT big data streams. The first part introduces data stream learners for classification, regression, clustering, and frequent pattern mining. The second part deals with scalability issues inherent in IoT applications, and discusses how to mine data streams on distributed engines such as Spark, Flink, Storm, and Samza. Gianmarco De Francisci Morales, Albert Bifet, Latifur Khan, João Gama 0001, Wei Fan 0001 |
KDD | 4 |
| 2016 | Concept Neurons - Handling Drift Issues for Real-Time Industrial Data Mining
Luís Moreira-Matias, João Gama 0001, João Mendes-Moreira 0001 |
ECML/PKDD (3) | 2 |
| 2016 | MINAS: multiclass learning algorithm for novelty detection in data streams
Elaine Ribeiro de Faria, André C. P. L. F. de Carvalho, João Gama 0001 |
Data Min. Knowl. Discov. | 3 |
| 2016 | Adaptive Model Rules From High-Speed Data StreamsabstractDecision rules are one of the most expressive and interpretable models for machine learning. In this article, we present Adaptive Model Rules (AMRules), the first stream rule learning algorithm for regression problems. In AMRules, the antecedent of a rule is a conjunction of conditions on the attribute values, and the consequent is a linear combination of the attributes. In order to maintain a regression model compatible with the most recent state of the process generating data, each rule uses a Page-Hinkley test to detect changes in this process and react to changes by pruning the rule set. Online learning might be strongly affected by outliers. AMRules is also equipped with outliers detection mechanisms to avoid model adaption using anomalous examples. In the experimental section, we report the results of AMRules on benchmark regression problems, and compare the performance of our system with other streaming regression algorithms. João Duarte, João Gama 0001, Albert Bifet |
ACM Trans. Knowl. Discov. Data | 2 |
| 2015 | Multi-target regression from high-speed data streams with adaptive model rulesabstractMany real life prediction problems involve predicting a structured output. Multi-target regression is an instance of structured output prediction whose task is to predict for multiple target variables. Structured output algorithms are usually computationally and memory demanding, hence are not suited for dealing with massive amounts of data. Most of these algorithms can be categorized as local or global methods. Local methods produce individual models for each output component and combine them to produce the structured prediction. Global methods adapt traditional learning algorithms to predict the output structure as a whole. We propose the first rule-based algorithm for solving multi-target regression problems from data streams. The algorithm builds on the adaptive model rules framework. In contrast to the majority of the structured output predictors, this particular algorithm does not fall into the local and global categories. Instead, each rule specializes on related subsets of the output attributes. To evaluate the performance of the proposed algorithm, two other rule-based algorithms were developed, one using the local strategy and the other using the global strategy. These methods were compared considering their prediction error, memory usage, computational time, and model complexity. Experimental results on synthetic and real data show that the local-strategy algorithm usually obtains the lowest error. However, the proposed and the global-strategy algorithms use much less memory and run significantly much faster at the cost of a slightly increase in the error, which make them very attractive when computation resources are an important factor. Also, the models produced by the latter approaches are much easier to understand since considerably less rules are produced. João Duarte, João Gama 0001 |
DSAA | 2 |
| 2015 | Data Stream Classification Guided by Clustering on Nonstationary Environments and Extreme Verification LatencyabstractData stream classification algorithms for nonstationary environments frequently assume the availability of class labels, instantly or with some lag after the classification. However, certain applications, mainly those related to sensors and robotics, involve high costs to obtain new labels during the classification phase. Such a scenario in which the actual labels of processed data are never available is called extreme verification latency. Extreme verification latency requires new classification methods capable of adapting to possible changes over time without external supervision. This paper presents a fast, simple, intuitive and accurate algorithm to classify nonstationary data streams in an extreme verification latency scenario, namely Stream Classification Algorithm Guided by Clustering – SCARGC. Our method consists of a clustering followed by a classification step applied repeatedly in a closed loop fashion. We show in several classification tasks evaluated in synthetic and real data that our method is faster and more accurate than the state-of-the-art. Vinícius M. A. de Souza, Diego Furtado Silva, João Gama 0001, Gustavo Batista |
SDM | 3 |
| 2015 | Guest editors introduction: special issue of the ECMLPKDD 2015 journal track
Concha Bielza, João Gama 0001, Alípio Mário Jorge, Indre Zliobaite |
Data Min. Knowl. Discov. | 2 |
| 2015 | Very fast decision rules for classification in data streams
Petr Kosina, João Gama 0001 |
Data Min. Knowl. Discov. | 2 |
| 2015 | Probabilistic change detection and visualization methods for the assessment of temporal stability in biomedical data quality
Carlos Sáez 0001, Pedro Pereira Rodrigues, João Gama 0001, Montserrat Robles, Juan Miguel García-Gómez |
Data Min. Knowl. Discov. | 3 |
| 2015 | Validating the coverage of bus schedules: A Machine Learning approach
João Mendes-Moreira 0001, Luís Moreira-Matias, João Gama 0001, Jorge Freire de Sousa |
Inf. Sci. | 3 |
| 2015 | Evaluation of Multiclass Novelty Detection Algorithms for Data StreamsabstractData stream mining is an emergent research area that investigates knowledge extraction from large amounts of continuously generated data, produced by non-stationary distribution. Novelty detection, the ability to identify new or previously unknown situations, is a useful ability for learning systems, especially when dealing with data streams, where concepts may appear, disappear, or evolve overtime. There are several studies currently investigating the application of novelty detection techniques in data streams. However, there is no consensus regarding how to evaluate the performance of these techniques. In this study, we propose a new evaluation methodology for multiclass novelty detection in data streams able to deal with: i) unsupervised learning, which generates novelty patterns without an association with the true classes, where one class may be composed of a novelty set, ii) confusion matrix that increases overtime, iii) confusion matrix with a column representing unknown examples, i.e., those not explained by the model, and iv) representation of the evaluation measures overtime. We propose a new methodology to associate the novelty patterns detected by the algorithm, in an unsupervised fashion, with the true classes. Finally, we evaluate the performance of the proposed methodology through the use of known novelty detection algorithms with artificial and real data sets. Elaine Ribeiro de Faria, Isabel Ribeiro Gonçalves, João Gama 0001, André C. P. L. F. de Carvalho |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2014 | Distributed Adaptive Model Rules for mining big data streamsabstractDecision rules are among the most expressive data mining models. We propose the first distributed streaming algorithm to learn decision rules for regression tasks. The algorithm is available in SAMOA (Scalable Advanced Massive Online Analysis), an open-source platform for mining big data streams. It uses a hybrid of vertical and horizontal parallelism to distribute Adaptive Model Rules (AMRules) on a cluster. The decision rules built by AMRules are comprehensible models, where the antecedent of a rule is a conjunction of conditions on the attribute values, and the consequent is a linear combination of the attributes. Our evaluation shows that this implementation is scalable in relation to CPU and memory consumption. On a small commodity Samza cluster of 9 nodes, it can handle a rate of more than 30000 instances per second, and achieve a speedup of up to 4.7x over the sequential version. Anh Thu Vu, Gianmarco De Francisci Morales, João Gama 0001, Albert Bifet |
IEEE BigData | 3 |
| 2014 | An Incremental Probabilistic Model to Predict Bus Bunching in Real-Time
Luís Moreira-Matias, João Gama 0001, João Mendes-Moreira 0001, Jorge Freire de Sousa |
IDA | 2 |
| 2014 | Recurrent concepts in data streams classification
João Gama 0001, Petr Kosina |
Knowl. Inf. Syst. | 1 |
| 2013 | Adaptive Model Rules from Data Streams
Ezilda Almeida, Carlos Ferreira 0007, João Gama 0001 |
ECML/PKDD (1) | 3 |
| 2012 | Online Predictive Model for Taxi Services
Luís Moreira-Matias, João Gama 0001, Michel Ferreira, João Mendes-Moreira 0001, Luís Damas |
IDA | 2 |
| 2012 | Where Are We Going? Predicting the Evolution of Individuals
Zaigham Faraz Siddiqui, Márcia D. B. Oliveira, João Gama 0001, Myra Spiliopoulou |
IDA | 3 |
| 2012 | Mobile Data Stream Mining: From Algorithms to ApplicationsabstractThis paper presents an overview of the current state-of-the-art in mobile data stream mining. This area of mobile data stream mining is significant for a number of new application domains such as mobile crowd sensing and mobile activity recognition. The paper presents the strategies and techniques for adaptation that are essential in order to perform real-time, continuous data mining on mobile devices. We present an overview of the algorithms research in this area. Finally, we discuss the key toolkits, systems and applications of mobile data stream mining. Shonali Krishnaswamy, João Gama 0001, Mohamed Medhat Gaber |
MDM | 2 |
| 2012 | Handling Time Changing Data with Adaptive Very Fast Decision Rules
Petr Kosina, João Gama 0001 |
ECML/PKDD (1) | 2 |
| 2011 | Advances in data stream mining for mobile and ubiquitous environmentsabstractThe tutorial presents the state-of-the-art in mobile and ubiquitous data stream mining and discusses open research problems, issues, and challenges in this area. Shonali Krishnaswamy, João Gama 0001, Mohamed Medhat Gaber |
CIKM | 2 |
| 2011 | Online Evaluation of Email Streaming Classifiers Using GNUsmail
José M. Carmona-Cejudo, Manuel Baena-García, José del Campo-Ávila, Albert Bifet, João Gama 0001, Rafael Morales Bueno |
IDA | 5 |
| 2011 | Learning about the Learning Process
João Gama 0001, Petr Kosina |
IDA | 1 |
| 2011 | Learning model trees from evolving data streams
Elena Ikonomovska, João Gama 0001, Saso Dzeroski |
Data Min. Knowl. Discov. | 2 |
| 2011 | Best papers from the Fifth International Conference on Advanced Data Mining and Applications (ADMA 2009)
Jian Pei 0001, João Gama 0001, Qiang Yang 0001, Ronghuai Huang, Xue Li 0001 |
Knowl. Inf. Syst. | 2 |
| 2010 | Bipartite Graphs for Monitoring Clusters Transitions
Márcia D. B. Oliveira, João Gama 0001 |
IDA | 2 |
| 2010 | The next generation of transportation systems, greenhouse emissions, and data miningabstractControling Greenhouse gas (GHG) emissions for minimizing the impact on the environment is one of the major challenges in front of the human civilization. Although future concentrations, damages and costs are unknown, it Hillol Kargupta, João Gama 0001, Wei Fan 0001 |
KDD | 2 |
| 2009 | Issues in evaluation of stream learning algorithmsabstractLearning from data streams is a research area of increasing importance. Nowadays, several stream learning algorithms have been developed. Most of them learn decision models that continuously evolve over time, run in resource-aware environments, detect and react to changes in the environment generating data. One important issue, not yet conveniently addressed, is the design of experimental work to evaluate and compare decision models that evolve over time. There are no golden standards for assessing performance in non-stationary environments. This paper proposes a general framework for assessing predictive stream learning algorithms. We defend the use of Predictive Sequential methods for error estimate - the prequential error. The prequential error allows us to monitor the evolution of the performance of models that evolve over time. Nevertheless, it is known to be a pessimistic estimator in comparison to holdout estimates. To obtain more reliable estimators we need some forgetting mechanism. Two viable alternatives are: sliding windows and fading factors. We observe that the prequential error converges to an holdout estimator when estimated over a sliding window or using fading factors. We present illustrative examples of the use of prequential error estimators, using fading factors, for the tasks of: i) assessing performance of a learning algorithm; ii) comparing learning algorithms; iii) hypothesis testing using McNemar test; and iv) change detection using Page-Hinkley test. In these tasks, the prequential error estimated using fading factors provide reliable estimators. In comparison to sliding windows, fading factors are faster and memory-less, a requirement for streaming applications. This paper is a contribution to a discussion in the good-practices on performance assessment when learning dynamic models that evolve over time. João Gama 0001, Raquel Sebastião, Pedro Pereira Rodrigues |
KDD | 1 |
| 2008 | Clustering Distributed Sensor Data Streams
Pedro Pereira Rodrigues, João Gama 0001, Luís M. B. Lopes |
ECML/PKDD (2) | 2 |
| 2008 | Hierarchical Clustering of Time-Series Data StreamsabstractThis paper presents and analyzes an incremental system for clustering streaming time series. The Online Divisive-Agglomerative Clustering (ODAC) system continuously maintains a tree-like hierarchy of clusters that evolves with data, using a top-down strategy. The splitting criterion is a correlation-based dissimilarity measure among time series, splitting each node by the farthest pair of streams. The system also uses a merge operator that reaggregates a previously split node in order to react to changes in the correlation structure between time series. The split and merge operators are triggered in response to changes in the diameters of existing clusters, assuming that in stationary environments, expanding the structure leads to a decrease in the diameters of the clusters. The system is designed to process thousands of data streams that flow at a high rate. The main features of the system include update time and memory consumption that do not depend on the number of examples in the stream. Moreover, the time and memory required to process an example decreases whenever the cluster structure expands. Experimental results on artificial and real data assess the processing qualities of the system, suggesting a competitive performance on clustering streaming time series, exploring also its ability to deal with concept drift. Pedro Pereira Rodrigues, João Gama 0001, João Pedro Pedroso |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2007 | Stream-Based Electricity Load Forecast
João Gama 0001, Pedro Pereira Rodrigues |
PKDD | 1 |
| 2006 | Learning with Local Drift Detection
João Gama 0001, Gladys Castillo |
ADMA | 1 |
| 2006 | An Adaptive Prequential Learning Framework for Bayesian Network Classifiers
Gladys Castillo, João Gama 0001 |
PKDD | 2 |
| 2006 | ODAC: Hierarchical Clustering of Time Series Data StreamsabstractThis paper presents a time series whole clustering system that incrementally constructs a tree-like hierarchy of clusters, using a top-down strategy. The Online Divisive-Agglomerative Clustering (ODAC) system uses a correlation-based dissimilarity measure between time series over a data stream and possesses an agglomerative phase to enhance a dynamic behavior capable of concept drift detection. Main features include splitting and agglomerative criteria based on the diameters of existing clusters and supported by a significance level. At each new example, only the leaves are updated, reducing computation of unneeded dissimilarities and speeding up the process every time the structure grows. Experimental results on artificial and real data suggest competitive performance on clustering time series and show that the system is equivalent to a batch divisive clustering on stationary time series, being also capable of dealing with concept drift. With this work, we assure the possibility and importance of hierarchical incremental time series whole clustering in the data stream paradigm, presenting a valuable and usable option. Pedro Pereira Rodrigues, João Gama 0001, João Pedro Pedroso |
SDM | 2 |
| 2003 | Accurate decision trees for mining high-speed data streamsabstractIn this paper we study the problem of constructing accurate decision tree models from data streams. Data streams are incremental tasks that require incremental, online, and any-time learning algorithms. One of the most successful algorithms for mining data streams is VFDT. In this paper we extend the VFDT system in two directions: the ability to deal with continuous data and the use of more powerful classification techniques at tree leaves. The proposed system, VFDTc, can incorporate and classify new information online, with a single scan of the data, in time constant per example. The most relevant property of our system is the ability to obtain a performance similar to a standard decision tree algorithm even for medium size datasets. This is relevant due to the any-time property. We study the behaviour of VFDTc in different problems and demonstrate its utility in large and medium data sets. Under a bias-variance analysis we observe that VFDTc in comparison to C4.5 is able to reduce the variance component. João Gama 0001, Ricardo Rocha 0003, Pedro Medas |
KDD | 1 |
| 2001 | Functional Trees for ClassificationabstractThe design of algorithms that explore multiple representation languages and explore different search spaces has an intuitive appeal. In the context of classification problems, algorithms that generate multivariate trees are able to explore multiple representation languages by using decision tests based on a combination of attributes. The same applies to model-tree algorithms in regression domains, but using linear models at leaf nodes. In this paper, we study where to use combinations of attributes in decision tree learning. We present an algorithm for multivariate tree learning that combines a univariate decision tree with a discriminant function by means of constructive induction. This algorithm is able to use decision nodes with multivariate tests, and leaf nodes that predict a class using a discriminant function. Multivariate decision nodes are built when growing the tree, while functional leaves are built when pruning the tree. Functional trees can be seen as a generalization of multivariate trees. Our algorithm was compared against to its components and two simplified versions using 30 benchmark data sets. The experimental evaluation shows that our algorithm has clear advantages with respect to the generalization ability and model sizes at statistically significant confidence levels. João Gama 0001 |
ICDM | 1 |
| 2001 | Functional Trees for Regression
João Gama 0001 |
IDA | 1 |
| 1998 | Combining Classifiers by Constructive Induction
João Gama 0001 |
ECML | 1 |
| 1997 | Search-Based Class Discretization
Luís Torgo, João Gama 0001 |
ECML | 2 |
| 1997 | Oblique Linear Tree
João Gama 0001 |
IDA | 1 |
| 1994 | Characterizing the Applicability of Classification Algorithms Using Meta-Level Learning
Pavel Brazdil, João Gama 0001, Bob Henery |
ECML | 2 |