Luís Torgo

dblp:60/1153 · DBLP profile ↗
← Back
62ranked-venue papers
18as first author
11since 2021 · last 2025
0000-0002-6892-8871ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 11 first-author · 9 since 2021Databases, data management, data science and information retrieval · 31 · 11 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 1 since 2021Theory of computation · 8 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Cherry-Picking in Time Series Forecasting: How to Select Datasets to Make Your Model Shine
abstract
The importance of time series forecasting drives continuous research and the development of new approaches to tackle this problem. Typically, these methods are introduced through empirical studies that frequently claim superior accuracy for the proposed approaches. Nevertheless, concerns are rising about the reliability and generalizability of these results due to limitations in experimental setups. This paper addresses a critical limitation: the number and representativeness of the datasets used. We investigate the impact of dataset selection bias, particularly the practice of cherry-picking datasets, on the performance evaluation of forecasting methods. Through empirical analysis with a diverse set of benchmark datasets, our findings reveal that cherry-picking datasets can significantly distort the perceived performance of methods, often exaggerating their effectiveness. Furthermore, our results demonstrate that by selectively choosing just four datasets — what most studies report — 46% of methods could be deemed best in class, and 77% could rank within the top three. Additionally, recent deep learning-based approaches show high sensitivity to dataset selection, whereas classical methods exhibit greater robustness. Finally, our results indicate that, when empirically validating forecasting algorithms on a subset of the benchmarks, increasing the number of datasets tested from 3 to 6 reduces the risk of incorrectly identifying an algorithm as the best one by approximately 40%. Our study highlights the critical need for comprehensive evaluation frameworks that more accurately reflect real-world scenarios. Adopting such frameworks will ensure the development of robust and reliable forecasting methods.
Luis Roque, Vítor Cerqueira, Carlos Soares, Luís Torgo
AAAI4
2025 Meta Subspace Analysis: Understanding Model (Mis)behavior in the Metafeature Space
Carlos Soares, Paulo J. Azevedo, Vítor Cerqueira, Luís Torgo
DS4
2024 RHiOTS: A Framework for Evaluating Hierarchical Time Series Forecasting Algorithms
abstract
We introduce the Robustness of Hierarchically Organized Time Series (RHiOTS) framework, designed to assess the robustness of hierarchical time series forecasting models and algorithms on real-world datasets. Hierarchical time series, where lower-level forecasts must sum to upper-level ones, are prevalent in various contexts, such as retail sales across countries. Current empirical evaluations of forecasting methods are often limited to a small set of benchmark datasets, offering a narrow view of algorithm behavior. RHiOTS addresses this gap by systematically altering existing datasets and modifying the characteristics of individual series and their interrelations. It uses a set of parameterizable transformations to simulate those changes in the data distribution. Additionally, RHiOTS incorporates an innovative visualization component, turning complex, multidimensional robustness evaluation results into intuitive, easily interpretable visuals. This approach allows an in-depth analysis of algorithm and model behavior under diverse conditions. We illustrate the use of RHiOTS by analyzing the predictive performance of several algorithms. Our findings show that traditional statistical methods are more robust than state-of-the-art deep learning algorithms, except when the transformation effect is highly disruptive. Furthermore, we found no significant differences in the robustness of the algorithms when applying specific reconciliation methods, such as MinT. RHiOTS provides researchers with a comprehensive tool for understanding the nuanced behavior of forecasting algorithms, offering a more reliable basis for selecting the most appropriate method for a given problem.
Luis Roque, Carlos Soares, Luís Torgo
KDD3
2023 Subgroup mining for performance analysis of regression models
abstract
Abstract Machine learning algorithms have shown several advantages compared to humans, namely in terms of the scale of data that can be analysed, delivering high speed and precision. However, it is not always possible to understand how algorithms work. As a result of the complexity of some algorithms, users started to feel the need to ask for explanations, boosting the relevance of Explainable Artificial Intelligence. This field aims to explain and interpret models with the use of specific analytical methods that usually analyse how their predicted values and/or errors behave. While prediction analysis is widely studied, performance analysis has limitations for regression models. This paper proposes a rule‐based approach, Error Distribution Rules (EDRs), to uncover atypical error regions, while considering multivariate feature interactions without size restrictions. Extracting EDRs is a form of subgroup mining. EDRs are model agnostic and a drill‐down technique to evaluate regression models, which consider multivariate interactions between predictors. EDRs uncover regions of the input space with deviating performance providing an interpretable description of these regions. They can be regarded as a complementary tool to the standard reporting of the expected average predictive performance. Moreover, by providing interpretable descriptions of these specific regions, EDRs allow end users to understand the dangers of using regression tools for some specific cases that fall on these regions, thaṯ is, they improve the accountability of models. The performance of several models from different problems was studied, showing that our proposal allows the analysis of many situations and direct model comparison. In order to facilitate the examination of rules, two visualization tools based on boxplots and density plots were implemented. A network visualization tool is also provided to rapidly check interactions of every feature condition. An additional tool is provided by using a grid of boxplots, where comparison between quartiles of every distribution with a reference is performed. Based on this comparison, an extrapolation of counterfactual examples to regression was also implemented. A set of examples is described, including a setting where regression models performance is compared in detail using EDRs. Specifically, the error difference between two models in a dataset is studied by deriving rules highlighting regions of the input space where model performance difference is unexpected. The application of visual tools is illustrated using EDRs examples derived from public available datasets. Also, case studies illustrating the specialization of subgroups, identification of counter factual subgroups and detecting unanticipated complex models are presented. This paper extends the state of the art by providing a method to derive explanations for model performance instead of explanations for model predictions.
João Pimentel 0003, Paulo J. Azevedo, Luís Torgo
Expert Syst. J. Knowl. Eng.3
2023 STUDD: a student-teacher method for unsupervised concept drift detection
Vítor Cerqueira, Heitor Murilo Gomes, Albert Bifet, Luís Torgo
Mach. Learn.4
2023 Automated imbalanced classification via layered learning
Vítor Cerqueira, Luís Torgo, Paula Branco, Colin Bellinger
Mach. Learn.2
2023 Early anomaly detection in time series: a hierarchical approach for predicting critical health episodes
Vítor Cerqueira, Luís Torgo, Carlos Soares
Mach. Learn.2
2023 Model Selection for Time Series Forecasting An Empirical Analysis of Multiple Estimators
Vítor Cerqueira, Luís Torgo, Carlos Soares
Neural Process. Lett.2
2022 A case study comparing machine learning with statistical methods for time series forecasting: size matters
Vítor Cerqueira, Luís Torgo, Carlos Soares
J. Intell. Inf. Syst.2
2021 Active Learning for Imbalanced Domains: the ALOD and ALOD-RE Algorithms
abstract
Active learning strategies are used to acquire an enlarged labelled training set that allows the learner to achieve a better performance. To this end, unlabelled instances are carefully selected for labelling by an human expert in order to achieve the best performance with the smallest number of questions to this oracle. Several techniques exist to select the most informative samples within active learning strategies. However, the effectiveness of these methods when applied to problems with imbalanced classes was not studied before. In this paper, we focus on the improvement of instance selection strategies for active learning techniques in the presence of class imbalance. In an imbalanced setting, learning algorithms have difficulties to focus on the minority class due to its under-representation. Still, this is typically the class of interest for the end-user. This mismatch between the classes distribution and the goals of the end user is known as the class imbalance problem. In an active learning setting for class imbalance problems, this becomes an even more challenging issue. We propose two novel active learning algorithms, ALOD and ALOD-RE centered around selecting the most informative samples to be labelled, while considering the selection of possible minority class cases and the generation of synthetic minority class examples to improve the learner performance. To this end, our two proposed solutions combine: active learning procedures, resampling strategies, and anomaly detection methods. Through an extensive set of experiments we show that the incorporation of outlier detection and resampling techniques in the active learning procedure benefits the learners performance on imbalanced domains. The performance advantages of our proposed ALOD and ALOD-RE algorithms are clearly supported by our experimental results.
Mayukh Bhattacharjee, Hema Sri Kambhampati, Paula Branco, Luís Torgo
DSAA4
2021 SWS: an unsupervised trajectory segmentation algorithm based on change detection with interpolation kernels
Mohammad Etemad, Amílcar Soares Júnior 0001, Elham Etemad, Jordan Rose, Luís Torgo, Stan Matwin
GeoInformatica5
2020 Knowledge-based Reliability Metrics for Social Media Accounts
Nuno Guimarães, Álvaro Figueira, Luís Torgo
WEBIST3
2020 Visual interpretation of regression error
abstract
Abstract Several sophisticated machine learning tools (e.g., ensembles or deep networks) have shown outstanding performance in different regression forecasting tasks. In many real world application domains the numeric predictions of the models drive important and costly decisions. Nevertheless, decision makers frequently require more than a black box model to be able to “trust” the predictions up to the point that they base their decisions on them. In this context, understanding these black boxes has become one of the hot topics in Machine Learning research. This paper proposes a series of visualization tools that explain the relationship between the expected predictive performance of black box regression models and the values of the input variables of any given test case. This type of information thus allows end‐users to correctly assess the risks associated with the use of a model, by showing how concrete values of the predictors may affect the performance of the model. Our illustrations with different real world data sets and learning algorithms provide insights on the type of usage and information these tools bring to both the data analyst and the end‐user. Furthermore, a thorough evaluation of the proposed tools is performed to showcase the reliability of this approach.
Inês Areosa, Luís Torgo
Expert Syst. J. Knowl. Eng.2
2020 Evaluating time series forecasting models: an empirical study on performance estimation methods
Vítor Cerqueira, Luís Torgo, Igor Mozetic
Mach. Learn.2
2019 The CURE for Class Imbalance
Colin Bellinger, Paula Branco, Luís Torgo
DS3
2019 Layered Learning for Early Anomaly Detection: Predicting Critical Health Episodes
Vítor Cerqueira, Luís Torgo, Carlos Soares
DS2
2019 Biased Resampling Strategies for Imbalanced Spatio-Temporal Forecasting
abstract
Extreme and rare events, such as abnormal spikes in air pollution or weather conditions can have serious repercussions. Many of these sorts of events develop from spatio-temporal processes, and accurate predictions are a most valuable tool in addressing their impact, in a timely manner. In this paper, we propose a new set of resampling strategies for imbalanced spatio-temporal forecasting tasks, by introducing bias into formerly random processes. This spatio-temporal bias includes a hyper-parameter that regulates the relative importance of the temporal and spatial dimensions in the selection of observations during under-or over-sampling. We test and compare our proposals against standard versions of the strategies on 10 different geo-referenced numeric time series, using 3 distinct off-the-shelf learning algorithms. Experimental results show that our proposal provides an advantage over random resampling strategies in imbalanced spatio-temporal forecasting tasks. Additionally, we also find that valuing an observation's recency is more useful when over-sampling; while valuing its spatial distance to other cases with extreme values is more beneficial when under-sampling.
Mariana Oliveira 0001, Nuno Moniz, Luís Torgo, Vítor Santos Costa
DSAA3
2019 Explaining the Performance of Black Box Regression Models
abstract
The widespread usage of Machine Learning and Data Mining models in several key areas of our societies has raised serious concerns in terms of accountability and ability to justify and interpret the decisions of these models. This is even more relevant when models are too complex and often regarded as black boxes. In this paper we present several tools designed to help in understanding and explaining the reasons for the observed predictive performance of black box regression models. We describe, evaluate and propose several variants of Error Dependence Plots. These plots provide a visual display of the expected relationship between the prediction error of any model and the values of a predictor variable. They allow the end user to understand what to expect from the models given some concrete values of the predictor variables. These tools allow more accurate explanations on the conditions that may lead to some failures of the models. Moreover, our proposed extensions also provide a multivariate perspective of this analysis, and the ability to compare the behaviour of multiple models under different conditions. This comparative analysis empowers the end user with the ability to have a case-based analysis of the risks associated with different models, and thus select the model with lower expected risk for each test case, or even decide not to use any model because the expected error is unacceptable.
Inês Areosa, Luís Torgo
DSAA2
2019 A Study on the Impact of Data Characteristics in Imbalanced Regression Tasks
abstract
The class imbalance problem has been thoroughly studied over the past two decades. More recently, the research community realized that the problem of imbalanced distributions also occurred in other tasks beyond classification. Regression problems are among these newly studied tasks where the problem of imbalanced domains also poses important challenges. Imbalanced regression problems occur in a diversity of real world domains such as meteorological (predicting weather extreme values), financial (extreme stock returns forecasting) or medical (anticipate rare values). In imbalanced regression the end-user preferences are biased towards values of the target variable that are under-represented on the available data. Several pre-processing methods were proposed to address this problem. These methods change the training set to force the learner to focus on the rare cases. However, as far as we know, the relationship between the data intrinsic characteristics and the performance achieved by these methods has not yet been studied for imbalanced regression tasks. In this paper we describe a study of the impact certain data characteristics may have in the results of applying pre-processing methods to imbalanced regression problems. To achieve this goal, we define potentially interesting data characteristics of regression problems. We then conduct our study using a synthetic data repository build for this purpose. We show that all the different characteristics studied have a different behaviour that is related with the level at which the data characteristic is present and the learning algorithm used. The main contributions of our work are: i) to define interesting data characteristics for regression tasks; ii) to create the first repository of imbalanced regression tasks containing 6000 data sets with controlled data characteristics; and iii) to provide insights on the impact of intrinsic data characteristics in the results of pre-processing methods for handling imbalanced regression tasks.
Paula Branco, Luís Torgo
DSAA2
2019 Pre-processing approaches for imbalanced distributions in regression
Paula Branco, Luís Torgo, Rita P. Ribeiro
Neurocomputing2
2019 A Brief Overview on the Strategiesto Fight Back the Spreadof False Information
abstract
The proliferation of false information on social networks is one of the hardest challenges in today's society, with implications capable of changing users perception on what is a fact or rumor.Due to its complexity, there has been an overwhelming number of contributions from the research community like the analysis of specific events where rumors are spread, analysis of the propagation of false content on the network, or machine learning algorithms to distinguish what is a fact and what is "fake news".In this paper, we identify and summarize some of the most prevalent works on the different categories studied.Finally, we also discuss the methods applied to deceive users and what are the next main challenges of this area.
Álvaro Figueira, Nuno Guimarães, Luís Torgo
J. Web Eng.3
2019 Arbitrage of forecasting experts
Vítor Cerqueira, Luís Torgo, Fábio Pinto, Carlos Soares
Mach. Learn.2
2018 MetaUtil: Meta Learning for Utility Maximization in Regression
Paula Branco, Luís Torgo, Rita P. Ribeiro
DS2
2018 Analysis and Detection of Unreliable Users in Twitter: Two Case Studies
Nuno Guimarães, Álvaro Figueira, Luís Torgo
IC3K3
2018 Evaluation Procedures for Forecasting with Spatio-Temporal Data
Mariana Oliveira 0001, Luís Torgo, Vítor Santos Costa
ECML/PKDD (1)2
2018 Constructive Aggregation and Its Application to Forecasting with Dynamic Ensembles
Vítor Cerqueira, Fábio Pinto, Luís Torgo, Carlos Soares, Nuno Moniz
ECML/PKDD (1)3
2018 Current State of the Art to Detect Fake News in Social Media: Global Trendings and Next Challenges
Álvaro Figueira, Nuno Guimarães, Luís Torgo
WEBIST3
2018 Resampling with neighbourhood bias on imbalanced domains
abstract
Abstract Imbalanced domains are an important problem that arises in predictive tasks causing a loss in the performance on the most relevant cases for the user. This problem has been extensively studied for classification problems, where the target variable is nominal. Recently, it was recognized that imbalanced domains occur in several other contexts and for multiple tasks, such as regression tasks, where the target variable is continuous. This paper focuses on imbalanced domains in both classification and regression tasks. Resampling strategies are among the most successful approaches to address imbalanced domains. In this work, we propose variants of existing resampling strategies that are able to take into account the information regarding the neighbourhood of the examples. Instead of performing sampling uniformly, our proposals bias the strategies to reinforce some regions of the data sets. With an extensive set of experiments, we provide evidence of the advantage of introducing a neighbourhood bias in the resampling strategies for both classification and regression tasks with imbalanced data sets.
Paula Branco, Luís Torgo, Rita P. Ribeiro
Expert Syst. J. Knowl. Eng.2
2017 Learning Through Utility Optimization in Regression Tasks
abstract
Accounting for misclassification costs is important in many practical applications of machine learning, and cost-sensitive techniques for classification have been studied extensively. Utility-based learning provides a generalization of purely cost-based approaches that considers both costs and benefits, enabling application to domains with complex cost-benefit settings. However, there is little work on utility- or cost-based learning for regression. In this paper, we formally define the problem of utility-based regression and propose a strategy for maximizing the utility of regression models. We verify our findings in a large set of experiments that show the advantage of our proposal in a diverse set of domains, learning algorithms and cost/benefit settings.
Paula Branco, Luís Torgo, Rita P. Ribeiro, Eibe Frank, Bernhard Pfahringer, Markus Michael Rau
DSAA2
2017 Dynamic and Heterogeneous Ensembles for Time Series Forecasting
abstract
This paper addresses the issue of learning time series forecasting models in changing environments by leveraging the predictive power of ensemble methods. Concept drift adaptation is performed in an active manner, by dynamically combining base learners according to their recent performance using a non-linear function. Diversity in the ensembles is encouraged with several strategies that include heterogeneity among learners, sampling techniques and computation of summary statistics as extra predictors. Heterogeneity is used with the goal of better coping with different dynamic regimes of the time series. The driving hypotheses of this work are that (i) heterogeneous ensembles should better fit different dynamic regimes and (ii) dynamic aggregation should allow for fast detection and adaptation to regime changes. We extend some strategies typically used in classification tasks to time series forecasting. The proposed methods are validated using Monte Carlo simulations on 16 real-world univariate time series with numerical outcome as well as an artificial series with clear regime shifts. The results provide strong empirical evidence for our hypotheses. To encourage reproducibility the proposed method is publicly available as a software package.
Vítor Cerqueira, Luís Torgo, Mariana Oliveira 0001, Bernhard Pfahringer
DSAA2
2017 A Comparative Study of Performance Estimation Methods for Time Series Forecasting
abstract
Performance estimation denotes a task of estimating the loss that a predictive model will incur on unseen data. These procedures are part of the pipeline in every machine learning task and are used for assessing the overall generalisation ability of models. In this paper we address the application of these methods to time series forecasting tasks. For independent and identically distributed data the most common approach is cross-validation. However, the dependency among observations in time series raises some caveats about the most appropriate way to estimate performance in these datasets and currently there is no settled way to do so. We compare different variants of cross-validation and different variants of out-of-sample approaches using two case studies: One with 53 real-world time series and another with three synthetic time series. Results show noticeable differences in the performance estimation methods in the two scenarios. In particular, empirical experiments suggest that cross-validation approaches can be applied to stationary synthetic time series. However, in real-world scenarios the most accurate estimates are produced by the out-of-sample methods, which preserve the temporal order of observations.
Vítor Cerqueira, Luís Torgo, Jasmina Smailovic, Igor Mozetic
DSAA2
2017 Relevance-Based Evaluation Metrics for Multi-class Imbalanced Domains
Paula Branco, Luís Torgo, Rita P. Ribeiro
PAKDD (1)2
2017 Arbitrated Ensemble for Time Series Forecasting
Vítor Cerqueira, Luís Torgo, Fábio Pinto, Carlos Soares
ECML/PKDD (2)2
2017 A comparative study of approaches to forecast the correct trading actions
abstract
Abstract This paper addresses the problem of decision making in the context of financial markets, more specifically, the problem of forecasting the correct trading action for a certain future horizon. We study and compare two alternative ways of addressing these forecasting tasks: (a) using standard numeric prediction models to forecast the variation on the prices of the target asset and, on a second stage, transform these numeric predictions into a decision according to some predefined decision rules; and (b) use models that directly forecast the right decision thus ignoring the intermediate numeric forecasting task. The objective of our study is to determine if both strategies provide identical results or if there is any particular advantage worth being considered that may distinguish each alternative in the context of financial markets.
Luís Baía, Luís Torgo
Expert Syst. J. Knowl. Eng.2
2016 Predicting Wildfires - Propositional and Relational Spatio-Temporal Pre-processing Approaches
Mariana Oliveira 0001, Luís Torgo, Vítor Santos Costa
DS2
2016 Resampling Strategies for Imbalanced Time Series
abstract
Time series forecasting is a challenging task, where the non-stationary characteristics of the data portrays a hard setting for predictive tasks. A common issue is the imbalanced distribution of the target variable, where some intervals are very important to the user but severely underrepresented. Standard regression tools focus on the average behaviour of the data. However, the objective is the opposite in many forecasting tasks involving time series: predicting rare values. A common solution to forecasting tasks with imbalanced data is the use of resampling strategies, which operate on the learning data by changing its distribution in favor of a given bias. The objective of this paper is to provide solutions capable of significantly improving the predictive accuracy of rare cases in forecasting tasks using imbalanced time series data. We extend the application of resampling strategies to the time series context and introduce the concept of temporal and relevance bias in the case selection process of such strategies, presenting new proposals. We evaluate the results of standard regression tools and the use of resampling strategies, with and without bias over 24 time series data sets from 6 different sources. Results show a significant increase in predictive accuracy of rare cases associated with the use of resampling strategies, and the use of biased strategies further increases accuracy over the non-biased strategies.
Nuno Moniz, Paula Branco, Luís Torgo
DSAA3
2015 Class-Based Outlier Detection: Staying Zombies or Awaiting for Resurrection?
Leona Nezvalová, Lubos Popelínský, Luís Torgo, Karel Vaculík
IDA3
2015 Resampling strategies for regression
abstract
Abstract Several real world prediction problems involve forecasting rare values of a target variable. When this variable is nominal, we have a problem of class imbalance that was thoroughly studied within machine learning. For regression tasks, where the target variable is continuous, few works exist addressing this type of problem. Still, important applications involve forecasting rare extreme values of a continuous target variable. This paper describes a contribution to this type of tasks. Namely, we propose to address such tasks by resampling approaches that change the distribution of the given data set to decrease the problem of imbalance between the rare target cases and the most frequent ones. We present two modifications of well‐known resampling strategies for classification tasks: the under‐sampling and the synthetic minority over‐sampling technique (SMOTE) methods. These modifications allow the use of these strategies on regression tasks where the goal is to forecast rare extreme values of the target variable. In an extensive set of experiments, we provide empirical evidence for the superiority of our proposals for these particular regression tasks. The proposed resampling methods can be used with any existing regression algorithm, which means that they are general tools for addressing problems of forecasting rare extreme values of a continuous target variable.
Luís Torgo, Paula Branco, Rita P. Ribeiro, Bernhard Pfahringer
Expert Syst. J. Knowl. Eng.1
2014 Ensembles for Time Series Forecasting
Mariana Oliveira 0001, Luís Torgo
ACML2
2014 Resampling Approaches to Improve News Importance Prediction
Nuno Moniz, Luís Torgo, Fátima Rodrigues 0001
IDA2
2013 OpenML: A Collaborative Science Platform
Jan N. van Rijn, Bernd Bischl, Luís Torgo, Bo Gao 0002, Venkatesh Umaashankar, Simon Fischer 0001, Patrick Winter, Bernd Wiswedel, Michael R. Berthold, Joaquin Vanschoren
ECML/PKDD (3)3
2012 Spatial Interpolation Using Multiple Regression
abstract
Many real world data mining applications involve analyzing geo-referenced data. Frequently, this type of data sets are incomplete in the sense that not all geographical coordinates have measured values of the variable(s) of interest. This incompleteness may be caused by poor data collection, measurement errors, costs management and many other factors. These missing values may cause several difficulties in many applications. Spatial imputation/interpolation methods try to fill in these unknown values in geo-referenced data sets. In this paper we propose a new spatial imputation method based on machine learning algorithms and a series of data pre-processing steps. The key distinguishing factor of this method is allowing the use of data from faraway regions, contrary to the state of the art on spatial data mining. Images (e.g. from a satellite or video surveillance cameras) may also suffer from this incompleteness where some pixels are missing, which again may be caused by many factors. An image can be seen as a spatial data set in a Cartesian coordinates system, where each pixel (location) registers some value (e.g. degree of gray on a black and white image). Being able to recover the original image from a partial or incomplete version of the reality is a key application in many domains (e.g. surveillance, security, etc.). In this paper we evaluate our general methodology for spatial interpolation on this type of problems. Namely, we check the ability of our method to fill in unknown pixels on several images. We compare it to state of the art methods and provide strong experimental evidence of the advantages of our proposal.
Orlando Ohashi, Luís Torgo
ICDM2
2011 Utility-Based Fraud Detection
abstract
Fraud detection is a key activity with serious socioeconomical impact. Inspection activities associated with this task are usually constrained by limited available resources. Data analysis methods can provide help in the task of deciding where to allocate these limited resources in order to optimise the outcome of the inspection activities. This paper presents a multi-strategy learning method to address the question of which cases to inspect first. The proposed methodology is based on the utility theory and provides a ranking ordered by decreasing expected outcome of inspecting the candidate cases. This outcome is a function not only of the probability of the case being fraudulent but also of the inspection costs and expected payoff if the case is confirmed as a fraud. The proposed methodology is general and can be useful on fraud detection activities with limited inspection resources. We experimentally evaluate our proposal on both an artificial domain and on a real world task. 1
Luís Torgo, Elsa Lopes
IJCAI1
2011 2D-interval predictions for time series
abstract
Research on time series forecasting is mostly focused on point predictions - models are obtained to estimate the expected value of the target variable for a certain point in future. However, for several relevant applications this type of forecasts has limited utility (e.g. costumer wallet value estimation, wind and electricity power production, control of water quality, etc.). For these domains it is frequently more important to be able to forecast a range of plausible future values of the target variable. A typical example is wind power production, where it is of high relevance to predict the future wind variability in order to ensure that supply and demand are balanced. This type of predictions will allow timely actions to be taken in order to cope with the expected values of the target variable on a certain future time horizon. In this paper we study this type of predictions - the prediction of a range of expected values for a future time interval. We describe some possible approaches to this task and propose an alternative procedure that our extensive experiments on both artificial and real world domains show to have clear advantages.
Luís Torgo, Orlando Ohashi
KDD1
2010 Interval Forecast of Water Quality Parameters
Orlando Ohashi, Luís Torgo, Rita P. Ribeiro
ECAI2
2009 Precision and Recall for Regression
Luís Torgo, Rita P. Ribeiro
Discovery Science1
2007 Utility-Based Regression
Luís Torgo, Rita P. Ribeiro
PKDD1
2006 Rule-Based Prediction of Rare Extreme Values
Rita P. Ribeiro, Luís Torgo
Discovery Science2
2006 Predicting Rare Extreme Values
Luís Torgo, Rita P. Ribeiro
PAKDD1
2006 Design of an end-to-end method to extract information from tables
Ana Costa e Silva, Alípio Mário Jorge, Luís Torgo
Int. J. Document Anal. Recognit.3
2005 Regression error characteristic surfaces
abstract
This paper presents a generalization of Regression Error Characteristic (REC) curves. REC curves describe the cumulative distribution function of the prediction error of models and can be seen as a generalization of ROC curves to regression problems. REC curves provide useful information for analyzing the performance of models, particularly when compared to error statistics like for instance the Mean Squared Error. In this paper we present Regression Error Characteristic (REC) surfaces that introduce a further degree of detail by plotting the cumulative distribution function of the errors across the distribution of the target variable, i.e. the joint cumulative distribution function of the errors and the target variable. This provides a more detailed analysis of the performance of models when compared to REC curves. This extra detail is particularly relevant in applications with non-uniform error costs, where it is important to study the performance of models for specific ranges of the target variable. In this paper we present the notion of REC surfaces, describe how to use them to compare the performance of models, and illustrate their use with an important practical class of applications: the prediction of rare extreme values.
Luís Torgo
KDD1
2003 Predicting Outliers
Luís Torgo, Rita P. Ribeiro
PKDD1
2003 Clustered Partial Linear Regression
Luís Torgo, Joaquim Pinto da Costa
Mach. Learn.1
2000 Clustered Partial Linear Regression
Luís Torgo, Joaquim Pinto da Costa
ECML1
2000 Partial Linear Trees
Luís Torgo
ICML1
2000 Efficient and Comprehensible Local Regression
Luís Torgo
PAKDD1
1998 Error Estimators for Pruning Regression Trees
Luís Torgo
ECML1
1997 Search-Based Class Discretization
Luís Torgo, João Gama 0001
ECML1
1997 Functional Models for Regression Tree Leaves
Luís Torgo
ICML1
1997 Regression Using Classification Algorithms
abstract
This article presents an alternative approach to the problem of regression. The methodology we describe allows the use of classification algorithms in regression tasks. From a practical point of view this enables the use of a wide range of existing machine learning (ML) systems in regression problems. In effect, most of the widely available systems deal with classification. Our method works as a pre-processing step in which the continuous goal variable values are discretised into a set of intervals. We use misclassification costs as a means to reflect the implicit ordering among these intervals. We describe a set of alternative discretisation methods and, based on our experimental results, justify the need for a search-based approach to choose the best method. The discretisation process is isolated from the classification algorithm, thus being applicable to virtually any existing system. The implemented system (RECLA) can thus be seen as a generic pre-processing tool. We have tested RECLA with three different classification systems and evaluated it in several regression data sets. Our experimental results confirm the validity of our search-based approach to class discretisation, and reveal the accuracy benefits of adding misclassification costs.
Luís Torgo, João Gama 0001
Intell. Data Anal.1
1993 Controlled Redundancy in Incremental Rule Learning
Luís Torgo
ECML1
1993 Rule Combination in Inductive Learning
Luís Torgo
ECML1