Peter A. Flach

dblp:f/PeterAFlach · DBLP profile ↗
← Back
115ranked-venue papers
19as first author
12since 2021 · last 2024
0000-0001-6857-5810ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 92 · 15 first-author · 7 since 2021Databases, data management, data science and information retrieval · 27 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 1 since 2021Theory of computation · 11 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2024 Explaining a Probabilistic Prediction on the Simplex with Shapley Compositions
abstract
Originating in game theory, Shapley values are widely used for explaining a machine learning model’s prediction by quantifying the contribution of each feature’s value to the prediction. This requires a scalar prediction as in binary classification, whereas a multiclass probabilistic prediction is a discrete probability distribution, living on a multidimensional simplex. In such a multiclass setting the Shapley values are typically computed separately on each class in a one-vs-rest manner, ignoring the compositional nature of the output distribution. In this paper, we introduce Shapley compositions as a well-founded way to properly explain a multiclass probabilistic prediction, using the Aitchison geometry from compositional data analysis. We prove that the Shapley composition is the unique quantity satisfying linearity, symmetry and efficiency on the Aitchison simplex, extending the corresponding axiomatic properties of the standard Shapley value. We demonstrate this proper multiclass treatment in a range of scenarios.
Paul-Gauthier Noé, Miquel Perelló-Nieto, Jean-François Bonastre, Peter A. Flach
ECAI4
2024 Interpretable representations in explainable AI: from theory to practice
abstract
Abstract Interpretable representations are the backbone of many explainers that target black-box predictive systems based on artificial intelligence and machine learning algorithms. They translate the low-level data representation necessary for good predictive performance into high-level human-intelligible concepts used to convey the explanatory insights. Notably, the explanation type and its cognitive complexity are directly controlled by the interpretable representation, tweaking which allows to target a particular audience and use case. However, many explainers built upon interpretable representations overlook their merit and fall back on default solutions that often carry implicit assumptions, thereby degrading the explanatory power and reliability of such techniques. To address this problem, we study properties of interpretable representations that encode presence and absence of human-comprehensible concepts. We demonstrate how they are operationalised for tabular, image and text data; discuss their assumptions, strengths and weaknesses; identify their core building blocks; and scrutinise their configuration and parameterisation. In particular, this in-depth analysis allows us to pinpoint their explanatory properties, desiderata and scope for (malicious) manipulation in the context of tabular data where a linear model is used to quantify the influence of interpretable concepts on a black-box prediction. Our findings lead to a range of recommendations for designing trustworthy interpretable representations; specifically, the benefits of class-aware (supervised) discretisation of tabular data, e.g., with decision trees, and sensitivity of image interpretable representations to segmentation granularity and occlusion colour.
Kacper Sokol, Peter A. Flach
Data Min. Knowl. Discov.2
2023 Reconciling Training and Evaluation Objectives in Location Agnostic Surrogate Explainers
abstract
Transparency in AI models is crucial to designing, auditing, and deploying AI systems. However, 'black box' models are still used in practice for their predictive power despite their lack of transparency. This has led to a demand for post-hoc, model-agnostic surrogate explainers which provide explanations for decisions of any model by approximating its behaviour close to a query point with a surrogate model. However, it is often overlooked how the location of the query point in the decision surface of the black box model affects the faithfulness of the surrogate explainer. Here, we show that when using standard techniques, there is a decrease in agreement between the black box and the surrogate model for query points towards the edge of the test dataset and when moving away from the decision boundary. This originates from a mismatch between the data distributions used to train and evaluate surrogate explainers. We address this by leveraging knowledge about the test data distribution captured in the class labels of the black box model. By addressing this and encouraging users to take care in understanding the alignment of training and evaluation objectives, we empower them to construct more faithful surrogate explainers.
Matthew Clifford, Jonathan Erskine, Alexander Hepburn, Peter A. Flach, Raúl Santos-Rodríguez
CIKM4
2023 Co-designing opportunities for Human-Centred Machine Learning in supporting Type 1 diabetes decision-making
abstract
Type 1 Diabetes (T1D) self-management requires hundreds of daily decisions. Diabetes technologies that use machine learning have significant potential to simplify this process and provide better decision support, but often rely on cumbersome data logging and cognitively demanding reflection on collected data. We set out to use co-design to identify opportunities for machine learning to support diabetes self-management in everyday settings. However, over nine months of interviews and design workshops with 15 people with T1D, we had to re-assess our assumptions about user needs. Our participants reported confidence in their personal knowledge and rejected machine learning based decision support when coping with routine situations, but highlighted the need for technological support in the context of unfamiliar or unexpected situations (holidays, illness, etc.). However, these are the situations where prior data are often lacking and drawing data-driven conclusions is challenging. Reflecting this challenge, we provide suggestions on how machine learning and other artificial intelligence approaches, e.g., expert systems, could enable decision-making support in both routine and unexpected situations.
Katarzyna Stawarz, Dmitri S. Katz, Amid Ayobi, Paul Marshall, Taku Yamagata, Raúl Santos-Rodríguez, Peter A. Flach, Aisling Ann O'Kane
Int. J. Hum. Comput. Stud.7
2023 Classifier calibration: a survey on how to assess and improve predicted class probabilities
abstract
Abstract This paper provides both an introduction to and a detailed overview of the principles and practice of classifier calibration. A well-calibrated classifier correctly quantifies the level of uncertainty or confidence associated with its instance-wise predictions. This is essential for critical applications, optimal decision making, cost-sensitive classification, and for some types of context change. Calibration research has a rich history which predates the birth of machine learning as an academic field by decades. However, a recent increase in the interest on calibration has led to new methods and the extension from binary to the multiclass setting. The space of options and issues to consider is large, and navigating it requires the right set of concepts and tools. We provide both introductory material and up-to-date technical details of the main concepts and methods, including proper scoring rules and other evaluation metrics, visualisation approaches, a comprehensive account of post-hoc calibration methods for binary and multiclass classification, and several advanced topics.
Telmo de Menezes e Silva Filho, Hao Song 0007, Miquel Perelló-Nieto, Raúl Santos-Rodríguez, Meelis Kull, Peter A. Flach
Mach. Learn.6
2022 LIMESegment: Meaningful, Realistic Time Series Explanations
abstract
LIME (Locally Interpretable Model-Agnostic Explanations) has become a popular way of generating explanations for tabular, image and natural language models, providing insight into why an instance was given a particular classification. In this paper we adapt LIME to time series classification, an under-explored area with existing approaches failing to account for the structure of this kind of data. We frame the non-trivial challenge of adapting LIME to time series classification as the following open questions: “What is a meaningful interpretable representation of a time series?”, “How does one realistically perturb a time series?” and “What is a local neighbourhood around a time series?”. We propose solutions to all three questions and combine them into a novel time series explanation framework called LIMESegment, which outperforms existing adaptations of LIME to time series on a variety of classification tasks.
Torty Sivill, Peter A. Flach
AISTATS2
2022 Self-Enhancer: A Self-supervised Framework for Low-Supervision, Drifted Data with Significant Missing Values
Yu Chen 0092, Peter A. Flach
ICANN (4)2
2022 Understanding Reinforcement Learning Based Localisation as a Probabilistic Inference Algorithm
Taku Yamagata, Raúl Santos-Rodríguez, Robert J. Piechocki, Peter A. Flach
ICANN (2)4
2021 Multi-label thresholding for cost-sensitive classification
Reem Alotaibi, Peter A. Flach
Neurocomputing2
2021 Co-Designing Personal Health? Multidisciplinary Benefits and Challenges in Informing Diabetes Self-Care Technologies
abstract
Co-design is a widely applied design process with well-documented values, including mutual learning and collective creativity. However, the real-world challenges of conducting multidisciplinary co-design research to inform the design of self-care technologies are not well established. We provide a qualitative account of a multidisciplinary project that aimed to co-design machine learning applications for Type 1 Diabetes (T1D) self-management. Through interviews, we identify not only perceived social, technological and strategic benefits of co-design but also organisational, translational and pragmatic design challenges: participants with T1D experienced difficulties in co-designing systems that met their individual self-care needs as part of group activities; HCI and AI researchers described challenges resulting from applying co-design outcomes to data-driven ML work; and industry collaborators highlighted academic data sharing regulations as cross-organisational challenges that can impede co-design efforts. Based on this understanding, we discuss opportunities for supporting multidisciplinary collaborations and aligning individual health needs with collaborative co-design activities.
Amid Ayobi, Katarzyna Stawarz, Dmitri S. Katz, Paul Marshall, Taku Yamagata, Raúl Santos-Rodríguez, Peter A. Flach, Aisling Ann O'Kane
Proc. ACM Hum. Comput. Interact.7
2021 Human Activity Recognition Based on Dynamic Active Learning
abstract
Activity of daily living is an important indicator of the health status and functional capabilities of an individual. Activity recognition, which aims at understanding the behavioral patterns of people, has increasingly received attention in recent years. However, there are still a number of challenges confronting the task. First, labelling training data is expensive and time-consuming, leading to limited availability of annotations. Secondly, activities performed by individuals have considerable variability, which renders the generally used supervised learning with a fixed label set unsuitable. To address these issues, we propose a dynamic active learning-based activity recognition method in this work. Different from traditional active learning methods which select samples based on a fixed label set, the proposed method not only selects informative samples from known classes, but also dynamically identifies new activities which are not included in the predefined label set. Starting with a classifier that has access to a limited number of labelled samples, we iteratively extend the training set with informative labels by fully considering the uncertainty, diversity and representativeness of samples, based on which better-informed classifiers can be trained, further reducing the annotation cost. We evaluate the proposed method on two synthetic datasets and two existing benchmark datasets. Experimental results demonstrate that our method not only boosts the activity recognition performance with considerably reduced annotation cost, but also enables adaptive daily activity analysis allowing the presence and detection of novel activities and patterns.
Haixia Bi, Miquel Perelló-Nieto, Raúl Santos-Rodríguez, Peter A. Flach
IEEE J. Biomed. Health Informatics4
2021 CRISP-DM Twenty Years Later: From Data Mining Processes to Data Science Trajectories
abstract
CRISP-DM(CRoss-Industry Standard Process for Data Mining) has its origins in the second half of the nineties and is thus about two decades old. According to many surveys and user polls it is still the de facto standard for developing data mining and knowledge discovery projects. However, undoubtedly the field has moved on considerably in twenty years, with data science now the leading term being favoured over data mining. In this paper we investigate whether, and in what contexts, CRISP-DM is still fit for purpose for data science projects. We argue that if the project is goal-directed and process-driven the process model view still largely holds. On the other hand, when data science projects become more exploratory the paths that the project can take become more varied, and a more flexible model is called for. We suggest what the outlines of such a trajectory-based model might look like and how it can be used to categorise data science projects (goal-directed, exploratory or data management). We examine seven real-life exemplars where exploratory activities play an important role and compare them against 51 use cases extracted from the NIST Big Data Public Working Group. We anticipate this categorisation can help project planning in terms of time and cost characteristics.
Fernando Martínez-Plumed, Lidia Contreras Ochando, Cèsar Ferri, José Hernández-Orallo, Meelis Kull, Nicolas Lachiche, María José Ramírez-Quintana, Peter A. Flach
IEEE Trans. Knowl. Data Eng.8
2020 FACE: Feasible and Actionable Counterfactual Explanations
abstract
Work in Counterfactual Explanations tends to focus on the principle of "the closest possible world" that identifies small changes leading to the desired outcome. In this paper we argue that while this approach might initially seem intuitively appealing it exhibits shortcomings not addressed in the current literature. First, a counterfactual example generated by the state-of-the-art systems is not necessarily representative of the underlying data distribution, and may therefore prescribe unachievable goals (e.g., an unsuccessful life insurance applicant with severe disability may be advised to do more sports). Secondly, the counterfactuals may not be based on a "feasible path" between the current state of the subject and the suggested one, making actionable recourse infeasible (e.g., low-skilled unsuccessful mortgage applicants may be told to double their salary, which may be hard without first increasing their skill level). These two shortcomings may render counterfactual explanations impractical and sometimes outright offensive. To address these two major flaws, first of all, we propose a new line of Counterfactual Explanations research aimed at providing actionable and feasible paths to transform a selected instance into one that meets a certain goal. Secondly, we propose FACE: an algorithmically sound way of uncovering these "feasible paths" based on the shortest path distances defined via density-weighted metrics. Our approach generates counterfactuals that are coherent with the underlying data distribution and supported by the "feasible paths" of change, which are achievable and can be tailored to the problem at hand.
Rafael Poyiadzi, Kacper Sokol, Raúl Santos-Rodríguez, Tijl De Bie, Peter A. Flach
AIES5
2020 Polsar Image Classification via Robust Low-Rank Feature Extraction and Markov Random Field
abstract
Polarimetric synthetic aperture radar (PolSAR) image classification has been investigated rigorously in various remote sensing applications. However, it is still a challenging task nowadays. One significant barrier lies in the speckle effect embedded in the PolSAR imaging process, which significantly degrades the quality of the images and further complicates the classification. To address this issue, we present a novel Pol-SAR image classification method which removes speckle noise via robust low-rank feature extraction and enforces smoothness priors through Markov Random Field (MRF). Specifically, we employ the mixture of Gaussian (MoG) based low-rank matrix factorization (LRMF) to simultaneously extract robust features and remove noise. Then, a classification map is obtained by applying Random Forest (RF) classifier on the extracted LRMF features. Finally, we refine the classification map by Markov random field (MRF) to enforce contextual smoothness. We conduct experiments on two real benchmark PolSAR data sets. Experimental results indicate that the proposed method achieves promising classification performance and preferable spatial consistency.
Haixia Bi, Raúl Santos-Rodríguez, Peter A. Flach
IGARSS3
2020 Uni- and multivariate probability density models for numeric subgroup discovery
abstract
Subgroup Discovery is a supervised, exploratory data mining paradigm that aims to identify subsets of a dataset that show interesting behaviour with respect to some designated target attribute. The way in which such distributional differences are quantified varies with the target attribute type. This work concerns continuous targets, which are important in many practical applications. For such targets, differences are often quantified using z-score and similar measures that compare simple statistics such as the mean and variance of the subset and the data. However, most distributions are not fully determined by their mean and variance alone. As a result, measures of distributional difference solely based on such simple statistics will miss potentially interesting subgroups. This work proposes methods to recognise distributional differences in a much broader sense. To this end, density estimation is performed using histogram and kernel density estimation techniques. In the spirit of Exceptional Model Mining, the proposed methods are extended to deal with multiple continuous target attributes, such that comparisons are not restricted to univariate distributions, but are available for joint distributions of any dimensionality. The methods can be incorporated easily into existing Subgroup Discovery frameworks, so no new frameworks are developed.
Marvin Meeng, Harm de Vries, Peter A. Flach, Siegfried Nijssen, Arno J. Knobbe
Intell. Data Anal.3
2020 Reflections on reciprocity in research
Peter A. Flach
Mach. Learn.1
2019 Performance Evaluation in Machine Learning: The Good, the Bad, the Ugly, and the Way Forward
abstract
This paper gives an overview of some ways in which our understanding of performance evaluation measures for machine-learned classifiers has improved over the last twenty years. I also highlight a range of areas where this understanding is still lacking, leading to ill-advised practices in classifier evaluation. This suggests that in order to make further progress we need to develop a proper measurement theory of machine learning. I then demonstrate by example what such a measurement theory might look like and what kinds of new results it would entail. Finally, I argue that key properties such as classification ability and data set difficulty are unlikely to be directly observable, suggesting the need for latent-variable models and causal inference.
Peter A. Flach
AAAI1
2019 Desiderata for Interpretability: Explaining Decision Tree Predictions with Counterfactuals
abstract
Explanations in machine learning come in many forms, but a consensus regarding their desired properties is still emerging. In our work we collect and organise these explainability desiderata and discuss how they can be used to systematically evaluate properties and quality of an explainable system using the case of class-contrastive counterfactual statements. This leads us to propose a novel method for explaining predictions of a decision tree with counterfactuals. We show that our model-specific approach exploits all the theoretical advantages of counterfactual explanations, hence improves decision tree interpretability by decoupling the quality of the interpretation from the depth and width of the tree.
Kacper Sokol, Peter A. Flach
AAAI2
2019 $β^3$-IRT: A New Item Response Model and its Applications
abstract
Item Response Theory (IRT) aims to assess latent abilities of respondents based on the correctness of their answers in aptitude test items with different difficulty levels. In this paper, we propose the $\beta^3$-IRT model, which models continuous responses and can generate a much enriched family of Item Characteristic Curves. In experiments we applied the proposed model to data from an online exam platform, and show our model outperforms a more standard 2PL-ND model on all datasets. Furthermore, we show how to apply $\beta^3$-IRT to assess the ability of machine learning classifiers.This novel application results in a new metric for evaluating the quality of the classifier’s probability estimates, based on the inferred difficulty and discrimination of data instances.
Telmo de Menezes e Silva Filho, Ricardo B. C. Prudêncio, Tom Diethe, Peter A. Flach
AISTATS5
2019 Distribution calibration for regression
abstract
We are concerned with obtaining well-calibrated output distributions from regression models. Such distributions allow us to quantify the uncertainty that the model has regarding the predicted target value. We introduce the novel concept of distribution calibration, and demonstrate its advantages over the existing definition of quantile calibration. We further propose a post-hoc approach to improving the predictions from previously trained regression models, using multi-output Gaussian Processes with a novel Beta link function. The proposed method is experimentally verified on a set of common regression models and shows improvements for both distribution-level and quantile-level calibration.
Hao Song 0007, Tom Diethe, Meelis Kull, Peter A. Flach
ICML4
2019 Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with Dirichlet calibration
abstract
Class probabilities predicted by most multiclass classifiers are uncalibrated, often tending towards over-confidence. With neural networks, calibration can be improved by temperature scaling, a method to learn a single corrective multiplicative factor for inputs to the last softmax layer. On non-neural models the existing methods apply binary calibration in a pairwise or one-vs-rest fashion. We propose a natively multiclass calibration method applicable to classifiers from any model class, derived from Dirichlet distributions and generalising the beta calibration method from binary classification. It is easily implemented with neural nets since it is equivalent to log-transforming the uncalibrated probabilities, followed by one linear layer and softmax. Experiments demonstrate improved probabilistic predictions according to multiple measures (confidence-ECE, classwise-ECE, log-loss, Brier score) across a wide range of datasets and classifiers. Parameters of the learned Dirichlet calibration map provide insights to the biases in the uncalibrated model.
Meelis Kull, Miquel Perelló-Nieto, Markus Kängsepp, Telmo de Menezes e Silva Filho, Hao Song 0007, Peter A. Flach
NeurIPS6
2019 Setting decision thresholds when operating conditions are uncertain
abstract
The quality of the decisions made by a machine learning model depends on the data and the operating conditions during deployment. Often, operating conditions such as class distribution and misclassification costs have changed during the time since the model was trained and evaluated. When deploying a binary classifier that outputs scores, once we know the new class distribution and the new cost ratio between false positives and false negatives, there are several methods in the literature to help us choose an appropriate threshold for the classifier’s scores. However, on many occasions, the information that we have about this operating condition is uncertain . Previous work has considered ranges or distributions of operating conditions during deployment, with expected costs being calculated for ranges or intervals, but still the decision for each point is made as if the operating condition were certain. The implications of this assumption have received limited attention: a threshold choice that is best suited without uncertainty may be suboptimal under uncertainty. In this paper we analyse the effect of operating condition uncertainty on the expected loss for different threshold choice methods, both theoretically and experimentally. We model uncertainty as a second conditional distribution over the actual operation condition and study it theoretically in such a way that minimum and maximum uncertainty are both seen as special cases of this general formulation. This is complemented by a thorough experimental analysis investigating how different learning algorithms behave for a range of datasets according to the threshold choice method and the uncertainty level.
Cèsar Ferri, José Hernández-Orallo, Peter A. Flach
Data Min. Knowl. Discov.3
2019 An application of hierarchical Gaussian processes to the detection of anomalies in star light curves
Niall Twomey, Haoyan Chen, Tom Diethe, Peter A. Flach
Neurocomputing4
2018 Anomaly detection in star light curves using hierarchical Gaussian processes
Haoyan Chen, Tom Diethe, Niall Twomey, Peter A. Flach
ESANN4
2018 The Facets of Artificial Intelligence: A Framework to Track the Evolution of AI
abstract
We present nine facets for the analysis of the past and future evolution of AI. Each facet has also a set of edges that can summarise different trends and contours in AI. With them, we first conduct a quantitative analysis using the information from two decades of AAAI/IJCAI conferences and around 50 years of documents from AI topics, an official database from the AAAI, illustrated by several plots. We then perform a qualitative analysis using the facets and edges, locating AI systems in the intelligence landscape and the discipline as a whole. This analytical framework provides a more structured and systematic way of looking at the shape and boundaries of AI.
Fernando Martínez-Plumed, Bao Sheng Loe, Peter A. Flach, Seán Ó hÉigeartaigh, Karina Vold, José Hernández-Orallo
IJCAI3
2018 Conversational Explanations of Machine Learning Predictions Through Class-contrastive Counterfactual Statements
abstract
Machine learning models have become pervasive in our everyday life; they decide on important matters influencing our education, employment and judicial system. Many of these predictive systems are commercial products protected by trade secrets, hence their decision-making is opaque. Therefore, in our research we address interpretability and explainability of predictions made by machine learning models. Our work draws heavily on human explanation research in social sciences: contrastive and exemplar explanations provided through a dialogue. This user-centric design, focusing on a lay audience rather than domain experts, applied to machine learning allows explainees to drive the explanation to suit their needs instead of being served a precooked template.
Kacper Sokol, Peter A. Flach
IJCAI2
2018 Glass-Box: Explaining AI Decisions With Counterfactual Statements Through Conversation With a Voice-enabled Virtual Assistant
abstract
The prevalence of automated decision making, influencing important aspects of our lives -- e.g., school admission, job market, insurance and banking -- has resulted in increasing pressure from society and regulators to make this process more transparent and ensure its explainability, accountability and fairness. We demonstrate a prototype voice-enabled device, called Glass-Box, which users can question to understand automated decisions and identify the underlying model's biases and errors. Our system explains algorithmic predictions with class-contrastive counterfactual statements (e.g., ``Had a number of conditions been different:...the prediction would change...''), which show a difference in a particular scenario that causes an algorithm to ``change its mind''. Such explanations do not require any prior technical knowledge to understand, hence are suitable for a lay audience, who interact with the system in a natural way -- through an interactive dialogue. We demonstrate the capabilities of the device by allowing users to impersonate a loan applicant who can question the system to understand the automated decision that he received.
Kacper Sokol, Peter A. Flach
IJCAI2
2018 Releasing eHealth Analytics into the Wild: Lessons Learnt from the SPHERE Project
abstract
The SPHERE project is devoted to advancing eHealth in a smart-home context, and supports full-scale sensing and data analysis to enable a generic healthcare service. We describe, from a data-science perspective, our experience of taking the system out of the laboratory into more than thirty homes in Bristol, UK. We describe the infrastructure and processes that had to be developed along the way, describe how we train and deploy Machine Learning systems in this context, and give a realistic appraisal of the state of the deployed systems.
Tom Diethe, Mike Holmes, Meelis Kull, Miquel Perelló-Nieto, Kacper Sokol, Hao Song 0007, Emma Tonkin, Niall Twomey, Peter A. Flach
KDD9
2017 Beta calibration: a well-founded and easily implemented improvement on logistic calibration for binary classifiers
abstract
For optimal decision making under variable class distributions and misclassification costs a classifier needs to produce well-calibrated estimates of the posterior probability. Isotonic calibration is a powerful non-parametric method that is however prone to overfitting on smaller datasets; hence a parametric method based on the logistic curve is commonly used. While logistic calibration is designed for normally distributed per-class scores, we demonstrate experimentally that many classifiers including Naive Bayes and Adaboost suffer from a particular distortion where these score distributions are heavily skewed. In such cases logistic calibration can easily yield probability estimates that are worse than the original scores. Moreover, the logistic curve family does not include the identity function, and hence logistic calibration can easily uncalibrate a perfectly calibrated classifier. In this paper we solve all these problems with a richer class of calibration maps based on the beta distribution. We derive the method from first principles and show that fitting it is as easy as fitting a logistic curve. Extensive experiments show that beta calibration is superior to logistic calibration for Naive Bayes and Adaboost.
Meelis Kull, Telmo de Menezes e Silva Filho, Peter A. Flach
AISTATS3
2017 The Role of Textualisation and Argumentation in Understanding the Machine Learning Process
abstract
Understanding data, models and predictions is important for machine learning applications. Due to the limitations of our spatial perception and intuition, analysing high-dimensional data is inherently difficult. Furthermore, black-box models achieving high predictive accuracy are widely used, yet the logic behind their predictions is often opaque. Use of textualisation -- a natural language narrative of selected phenomena -- can tackle these shortcomings. When extended with argumentation theory we could envisage machine learning models and predictions arguing persuasively for their choices.
Kacper Sokol, Peter A. Flach
IJCAI2
2017 Unsupervised learning of sensor topologies for improving activity recognition in smart environments
Niall Twomey, Tom Diethe, Ian Craddock, Peter A. Flach
Neurocomputing4
2016 Declaratively Capturing Local Label Correlations with Multi-Label Trees
abstract
The goal of multi-label classification is to predict multiple labels per data point simultaneously. Real-world applications tend to have high-dimensional label spaces, employing hundreds or even thousands of labels. While these labels could be predicted separately, by capturing label correlation we might achieve better predictive performance. In contrast with previous attempts in the literature that have modelled label correlations globally, this paper proposes a novel algorithm to model correlations and cluster labels locally. LaCovaC is a multi-label decision tree classifier that clusters labels into several dependent subsets at various points during training. The clusters are obtained locally by identifying the conditionally-dependent labels in localised regions of the feature space using the label correlation matrix. LaCovaC interleaves between two main decisions on the label matrix with training instances in rows and labels in columns: splitting this matrix vertically by partitioning the labels into subsets, or splitting it horizontally using features in the conventional way. Experiments on 13 benchmark datasets demonstrate that our proposal achieves competitive performance over a wide range of evaluation metrics when compared with the state-of-the-art multi-label classifiers.
Reem Alotaibi, Meelis Kull, Peter A. Flach
ECAI3
2016 Active transfer learning for activity recognition
Tom Diethe, Niall Twomey, Peter A. Flach
ESANN3
2016 Background Check: A General Technique to Build More Reliable and Versatile Classifiers
abstract
We introduce a powerful technique to make classifiers more reliable and versatile. Background Check equips classifiers with the ability to assess the difference of unlabelled test data from the training data. In particular, Background Check gives classifiers the capability to (i) perform cautious classification with a reject option, (ii) identify outliers, and (iii) better assess the confidence in their predictions. We derive the method from first principles and consider four particular relationships between background and foreground distributions. One of these assumes an affine relationship with two parameters, and we show how this bivariate parameter space naturally interpolates between the above capabilities. We demonstrate the versatility of the approach by comparing it experimentally with published special-purpose solutions for outlier detection and confident classification on 41 benchmark datasets. Results show that Background Check can match and in many cases surpass the performances of specialised approaches.
Miquel Perelló-Nieto, Telmo de Menezes e Silva Filho, Meelis Kull, Peter A. Flach
ICDM4
2016 ADL™: A Topic Model for Discovery of Activities of Daily Living in a Smart Home
Tom Diethe, Peter A. Flach
IJCAI3
2016 Fast Unsupervised Online Drift Detection Using Incremental Kolmogorov-Smirnov Test
abstract
Data stream research has grown rapidly over the last decade. Two major features distinguish data stream from batch learning: stream data are generated on the fly, possibly in a fast and variable rate; and the underlying data distribution can be non-stationary, leading to a phenomenon known as concept drift. Therefore, most of the research on data stream classification focuses on proposing efficient models that can adapt to concept drifts and maintain a stable performance over time. However, specifically for the classification task, the majority of such methods rely on the instantaneous availability of true labels for all already classified instances. This is a strong assumption that is rarely fulfilled in practical applications. Hence there is a clear need for efficient methods that can detect concept drifts in an unsupervised way. One possibility is the well-known Kolmogorov-Smirnov test, a statistical hypothesis test that checks whether two samples differ. This work has two main contributions. The first one is the Incremental Kolmogorov-Smirnov algorithm that allows performing the Kolmogorov-Smirnov hypothesis test instantly using two samples that change over time, where the change is an insertion and/or removal of an observation. Our algorithm employs a randomized tree and is able to perform the insertion and removal operations in O(log N) with high probability and calculate the Kolmogorov-Smirnov test in O(1), where N is the number of sample observations. This is a significant speed-up compared to the O(N log N) cost of the non-incremental implementation. The second contribution is the use of the Incremental Kolmogorov-Smirnov test to detect concept drifts without true labels. Classification algorithms adapted to use the test rely on a limited portion of those labels just to update the classification model after a concept drift is detected.
Denis Moreira dos Reis, Peter A. Flach, Stan Matwin, Gustavo Batista
KDD2
2016 Subgroup Discovery with Proper Scoring Rules
Hao Song 0007, Meelis Kull, Peter A. Flach, Georgios Kalogridis
ECML/PKDD (2)3
2016 Cost-sensitive boosting algorithms: Do we really need them?
abstract
We provide a unifying perspective for two decades of work on cost-sensitive Boosting algorithms. When analyzing the literature 1997–2016, we find 15 distinct cost-sensitive variants of the original algorithm; each of these has its own motivation and claims to superiority—so who should we believe? In this work we critique the Boosting literature using four theoretical frameworks: Bayesian decision theory, the functional gradient descent view, margin theory, and probabilistic modelling. Our finding is that only three algorithms are fully supported—and the probabilistic model view suggests that all require their outputs to be calibrated for best performance. Experiments on 18 datasets across 21 degrees of imbalance support the hypothesis—showing that once calibrated, they perform equivalently, and outperform all others. Our final recommendation—based on simplicity, flexibility and performance—is to use the original Adaboost algorithm with a shifted decision threshold and calibrated probability estimates.
Nikolaos Nikolaou, Narayanan Unny Edakunni, Meelis Kull, Peter A. Flach, Gavin Brown 0001
Mach. Learn.4
2016 On the need for structure modelling in sequence prediction
abstract
There is no uniform approach in the literature for modelling sequential correlations in sequence classification problems. It is easy to find examples of unstructured models ( e.g. logistic regression) where correlations are not taken into account at all, but there are also many examples where the correlations are explicitly incorporated into a—potentially computationally expensive—structured classification model ( e.g. conditional random fields). In this paper we lay theoretical and empirical foundations for clarifying the types of problem which necessitate direct modelling of correlations in sequences, and the types of problem where unstructured models that capture sequential aspects solely through features are sufficient. The theoretical work in this paper shows that the rate of decay of auto-correlations within a sequence is related to the excess classification risk that is incurred by ignoring the structural aspect of the data. This is an intuitively appealing result, demonstrating the intimate link between the auto-correlations and excess classification risk. Drawing directly on this theory, we develop well-founded visual analytics tools that can be applied a priori on data sequences and we demonstrate how these tools can guide practitioners in specifying feature representations based on auto-correlation profiles. Empirical analysis is performed on three sequential datasets. With baseline feature templates, structured and unstructured models achieve similar performance, indicating no initial preference for either model. We then apply the visual analytics tools to the datasets, and show that classification performance in all cases is improved over baseline results when our tools are involved in defining feature representations.
Niall Twomey, Tom Diethe, Peter A. Flach
Mach. Learn.3
2016 Feature Construction and Calibration for Clustering Daily Load Curves from Smart-Meter Data
abstract
This paper proposes and compares feature construction and calibration methods for clustering daily electricity load curves. Such load curves describe electricity demand over a period of time. A rich body of the literature has studied clustering of load curves, usually using temporal features. This limits the potential to discover new knowledge, which may not be best represented as models consisting of all time points on load curves. This paper presents three new methods to construct features: 1) conditional filters on time-resolution-based features; 2) calibration and normalization; and 3) using profile errors. These new features extend the potential of clustering load curves. Moreover, smart metering is now generating high-resolution time series, and so the dimensionality reduction offered by these features is welcome. The clustering results using the proposed new features are compared with clusterings obtained from temporal features, as well as clusterings with Fourier features, using household electricity consumption time series as test data. The experimental results suggest that the proposed feature construction methods offer new means for gaining insight in energy-consumption patterns.
Reem Alotaibi, Nanlin Jin, Tom Wilcox, Peter A. Flach
IEEE Trans. Ind. Informatics4
2015 Reframing in Frequent Pattern Mining
abstract
Mining frequent patterns is a crucial task in data mining. Most of the existing frequent pattern mining methods find the complete set of frequent patterns from a given dataset. However, in real-life scenarios we often need to predict the future frequent patterns for different tasks such as business policy making, web page recommendation, stock-market behavior and road traffic analysis. Predicting future frequent patterns from the currently available set of frequent patterns is challenging due to dataset shift where data distributions may change from one dataset to another. In this paper, we propose a new approach called reframing in frequent pattern mining to solve this task. Moreover, we experimentally show the existence of dataset shift in two real-life transactional datasets and the capability of our approach to handle these unknown shifts.
Chowdhury Farhan Ahmed, Mohammad Samiullah 0001, Nicolas Lachiche, Meelis Kull, Peter A. Flach
ICTAI5
2015 Precision-Recall-Gain Curves: PR Analysis Done Right
abstract
Precision-Recall analysis abounds in applications of binary classification where true negatives do not add value and hence should not affect assessment of the classifier's performance. Perhaps inspired by the many advantages of receiver operating characteristic (ROC) curves and the area under such curves for accuracy-based performance assessment, many researchers have taken to report Precision-Recall (PR) curves and associated areas as performance metric. We demonstrate in this paper that this practice is fraught with difficulties, mainly because of incoherent scale assumptions -- e.g., the area under a PR curve takes the arithmetic mean of precision values whereas the $F_{\beta}$ score applies the harmonic mean. We show how to fix this by plotting PR curves in a different coordinate system, and demonstrate that the new Precision-Recall-Gain curves inherit all key advantages of ROC curves. In particular, the area under Precision-Recall-Gain curves conveys an expected $F_1$ score on a harmonic scale, and the convex hull of a Precision-Recall-Gain curve allows us to calibrate the classifier's scores so as to determine, for each operating point on the convex hull, the interval of $\beta$ values for which the point optimises $F_{\beta}$. We demonstrate experimentally that the area under traditional PR curves can easily favour models with lower expected $F_1$ score than others, and so the use of Precision-Recall-Gain curves will result in better model selection.
Peter A. Flach, Meelis Kull
NIPS1
2015 Versatile Decision Trees for Learning Over Multiple Contexts
Reem Alotaibi, Ricardo B. C. Prudêncio, Meelis Kull, Peter A. Flach
ECML/PKDD (1)4
2015 Bayesian Modelling of the Temporal Aspects of Smart Home Activity with Circular Statistics
Tom Diethe, Niall Twomey, Peter A. Flach
ECML/PKDD (2)3
2015 Novel Decompositions of Proper Scoring Rules for Classification: Score Adjustment as Precursor to Calibration
Meelis Kull, Peter A. Flach
ECML/PKDD (1)2
2014 A Machine Learning Approach to Objective Cardiac Event Detection
abstract
This paper presents an automated framework for the detection of the QRS complex from Electrocardiogram (ECG) signals. We introduce an artefact-tolerant pre-processing algorithm which emphasises a number of characteristics of the ECG that are representative of the QRS complex. With this processed ECG signal we train Logistic Regression and Support Vector Machine classification models. With our approach we obtain over 99.7% detection sensitivity and precision on the MIT-BIH database without using supplementary de-noising or pre-emphasis filters.
Niall Twomey, Peter A. Flach
CISIS2
2014 LaCova: A Tree-Based Multi-label Classifier Using Label Covariance as Splitting Criterion
abstract
Dealing with multiple labels is a supervised learning problem of increasing importance. Multi-label classifiers face the challenge of exploiting correlations between labels. While in existing work these correlations are often modelled globally, in this paper we use the divide-and-conquer approach of decision trees which enables taking local decisions about how best to model label dependency. The resulting algorithm establishes a tree-based multi-label classifier called LaCova which dynamically interpolates between two well-known baseline methods: Binary Relevance, which assumes all labels independent, and Label Power set, which learns the joint label distribution. The key idea is a splitting criterion based on the label covariance matrix at that node, which allows us to choose between a horizontal split (branching on a feature) and a vertical split (separating the labels). Empirical results on 12 data sets show strong performance of the proposed method, particularly on data sets with hundreds of labels.
Reem Alotaibi, Meelis Kull, Peter A. Flach
ICMLA3
2014 Reliability Maps: A Tool to Enhance Probability Estimates and Improve Classification Accuracy
Meelis Kull, Peter A. Flach
ECML/PKDD (2)2
2014 Rate-Constrained Ranking and the Rate-Weighted AUC
Louise A. C. Millard, Peter A. Flach, Julian P. T. Higgins
ECML/PKDD (2)2
2014 Rate-Oriented Point-Wise Confidence Bounds for ROC Curves
Louise A. C. Millard, Meelis Kull, Peter A. Flach
ECML/PKDD (2)3
2014 Subgroup Discovery in Smart Electricity Meter Data
abstract
This work presents data mining methods for discovering unusual consumption patterns and their associated descriptive models from smart electricity meter data. At present, data mining and knowledge discovery in electricity meter data suffer from three notable weaknesses: 1) insufficient focus on intelligent data analysis of subgroups (subsets) whose patterns vary significantly from aggregate patterns embodied in an entire dataset; 2) a lack of effort towards generating intuitively understandable and practically applicable knowledge for industrial practitioners to identify such subgroups; and 3) limited knowledge regarding the link between unusual consumption patterns and household consumers' socio-demographic characteristics. This paper addresses these practically important but technically challenging issues by applying subgroup discovery algorithms to a real smart electricity meter dataset. Subgroups whose patterns are unusual and whose sizes are large enough are discovered, and their descriptive and predictive models are generated. Furthermore, to enrich subgroup discovery algorithms, three new-quality measures for real-valued targets are proposed. The comparative studies empirically evaluate the effectiveness and usefulness of subgroup discovery on classification accuracy, predictive power, and computational resources. The methodologies and algorithms presented are generic, and therefore applicable to a wider range of data mining problems.
Nanlin Jin, Peter A. Flach, Tom Wilcox, Royston Sellman, Joshua Thumim, Arno J. Knobbe
IEEE Trans. Ind. Informatics2
2013 A Higher-order data flow model for heterogeneous Big Data
abstract
We introduce a data flow model that supports highly parallelisable design patterns and also has useful properties for analysing data serially over extended time periods without requiring traditional Big Data computing facilities. The model ranges over a class of higher-order relations which are sufficiently expressive to represent a wide variety of unstructured, semi-structured and structured data. Using JSONMatch, our web service implementation of the model, we show that the combination of this model and higher-order representation provides a powerful and extensible framework that is particularly well suited to analysing Big Variety data in a web application context.
Simon Price, Peter A. Flach
IEEE BigData2
2013 Guest editors' introduction: special section of selected papers from ECML-PKDD 2012
Tijl De Bie, Peter A. Flach
Data Min. Knowl. Discov.2
2013 SubSift web services and workflows for profiling and comparing scientists and their published works
Simon Price, Peter A. Flach, Sebastian Spiegler, Christopher Bailey 0001, Nikki Rogers
Future Gener. Comput. Syst.2
2013 Guest editors' introduction: special issue of selected papers from ECML-PKDD 2012
Tijl De Bie, Peter A. Flach
Mach. Learn.2
2013 ROC curves in cost space
José Hernández-Orallo, Peter A. Flach, Cèsar Ferri
Mach. Learn.2
2012 Caveats and pitfalls of ROC analysis in clinical microarray research (and how to avoid them)
abstract
The receiver operating characteristic (ROC) has emerged as the gold standard for assessing and comparing the performance of classifiers in a wide range of disciplines including the life sciences. ROC curves are frequently summarized in a single scalar, the area under the curve (AUC). This article discusses the caveats and pitfalls of ROC analysis in clinical microarray research, particularly in relation to (i) the interpretation of AUC (especially a value close to 0.5); (ii) model comparisons based on AUC; (iii) the differences between ranking and classification; (iv) effects due to multiple hypotheses testing; (v) the importance of confidence intervals for AUC; and (vi) the choice of the appropriate performance metric. With a discussion of illustrative examples and concrete real-world studies, this article highlights critical misconceptions that can profoundly impact the conclusions about the observed performance.
Daniel P. Berrar, Peter A. Flach
Briefings Bioinform.2
2012 A unified view of performance metrics: translating threshold choice into expected classification loss
José Hernández-Orallo, Peter A. Flach, Cèsar Ferri
J. Mach. Learn. Res.2
2012 ILP turns 20 - Biography and future challenges
abstract
Inductive Logic Programming (ILP) is an area of Machine Learning which has now reached its twentieth year. Using the analogy of a human biography this paper recalls the development of the subject from its infancy through childhood and teenage years. We show how in each phase ILP has been characterised by an attempt to extend theory and implementations in tandem with the development of novel and challenging real-world applications. Lastly, by projection we suggest directions for research which will help the subject coming of age.
Stephen H. Muggleton, Luc De Raedt, David Poole 0001, Ivan Bratko, Peter A. Flach, Katsumi Inoue, Ashwin Srinivasan 0001
Mach. Learn.5
2011 A Coherent Interpretation of AUC as a Measure of Aggregated Classification Performance
Peter A. Flach, José Hernández-Orallo, Cèsar Ferri
ICML1
2011 Brier Curves: a New Cost-Based Visualisation of Classifier Performance
José Hernández-Orallo, Peter A. Flach, Cèsar Ferri
ICML2
2011 Smooth Receiver Operating Characteristics (smROC) Curves
William Klement, Peter A. Flach, Nathalie Japkowicz, Stan Matwin
ECML/PKDD (2)2
2011 The Machine Learning journal: 25 years young
Peter A. Flach
Mach. Learn.1
2010 Enhanced Word Decomposition by Calibrating the Decision Threshold of Probabilistic Models and Using a Model Ensemble
Sebastian Spiegler, Peter A. Flach
ACL2
2010 Ukwabelana - An open-source morphological Zulu corpus
Sebastian Spiegler, Andrew van der Spuy, Peter A. Flach
COLING3
2010 SubSift Web Services and Workflows for Profiling and Comparing Scientists and Their Published Works
abstract
Scientific researchers, laboratories and organisations can be profiled and compared by analysing their published works, including documents ranging from academic papers to web sites, blog posts and Twitter feeds. This paper describes how the vector space model from information retrieval, more normally associated with full text search, has been employed in the open source Sub Sift software to support workflows to profile and compare such collections of documents. Sub Sift was originally designed to match submitted conference or journal papers to potential peer reviewers based on the similarity between the paper's abstract and the reviewer's publications as found in online bibliographic databases. The software is implemented as a family of Restful web services that, composed into a re-usable workflow, have already been used to support several major data mining conferences. Alternative workflows and service compositions are now enabling other interesting applications.
Simon Price, Peter A. Flach, Sebastian Spiegler, Christopher Bailey 0001, Nikki Rogers
eScience2
2010 The Advantages of Seed Examples in First-Order Multi-class Subgroup Discovery
abstract
Subgroup discovery is halfway between predictive and descriptive rule learning: while there is a target concept, the goal of subgroup discovery is not necessarily to achieve high accuracy in predicting the target, but rather to identify subsets of the population whose class distribution is significantly different from the overall distribution. The target concept helps us to achieve a trade-off between accuracy and interestingness. In previous work [1] we illustrated the usefulness of several multiclass subgroup evaluation measures in producing highly predictive subgroups in a propositional logic framework. In this paper we upgrade multi-class subgroup discovery to first-order logic employing a new weighted covering algorithm that takes advantage of learning from seed examples. In a rigourous experimental evaluation we show that the use of seed examples leads to considerable and statistically significant improvement of predictive power, both in terms of accuracy (3 percent points on average over the data sets) and AUC (7 percent points). The paper is structured as follows. Section 2 describes the ingredients of our Aleph-MSD++ algorithm. An empirical evaluation of multi-class subgroup discovery for feature construction over 11 data sets is presented in Section 3. Finally, we conclude the paper in Section 4.
Tarek Abudawood, Peter A. Flach
ECAI2
2010 Learning Multi-class Theories in ILP
Tarek Abudawood, Peter A. Flach
ILP2
2010 The Machine Learning journal: 250 issues and counting
Peter A. Flach
Mach. Learn.1
2009 Evaluation Measures for Multi-class Subgroup Discovery
Tarek Abudawood, Peter A. Flach
ECML/PKDD (1)2
2009 Towards Learning Morphology for Under-Resourced Fusional and Agglutinating Languages
abstract
In this paper, we describe a novel and effective approach for automatically decomposing a word into stem and suffixes. Russian and Turkish are used as exemplars of fusional and agglutinating languages. Rather than relying on corpus counts, we use a small number of word-pairs as training data, that can be particularly suited for under-resourced languages. For fusional languages, we initially learn a tree of aligned suffix rules (TASR) from word-pairs. The tree is built top-down, from general to specific rules, using suffix rule frequency and rule subsumption, and is executed bottom-up, i.e., the most specific rule that fires is chosen. TASR is used to segment a word form into a stem and suffix sequence. For fusional languages learning through generation (using TASR) is essential for proper stem extraction. Subsequently, an unsupervised segmentation algorithm graph-based unsupervised suffix segmentation (GBUSS) is used to segment the suffix sequence. GBUSS employs a suffix graph where node merging, guided by an information-theoretic measure, generates suffix sequences. The approach, experimentally validated on Russian, is shown to be highly effective. For agglutinating languages only the GBUSS is needed for word decomposition. Promising experimental results for Turkish are obtained.
Kseniya B. Shalonova, Bruno Golénia, Peter A. Flach
IEEE Trans. Speech Audio Process.3
2008 A Fast Method for Property Prediction in Graph-Structured Data from Positive and Unlabelled Examples
abstract
The analysis of large and complex networks, or graphs, is becoming increasingly important in many scientific areas including machine learning, social network analysis and bioinformatics. One natural type of question that can be asked in network analysis is “Given two sets R and T of individuals in a graph with complete and missing knowledge, respectively, about a property of interest, which individuals in T are closest to R with respect to this property?”. To answer this question, we can rank the individuals in T such that the individuals ranked highest are most likely to exhibit the property of interest. Several methods based on weighted paths in the graph and Markov chain models have been proposed to solve this task. In this paper, we show that we can improve previously published approaches by rephrasing this problem as the task of property prediction in graph-structured data from positive examples, the individuals in R, and unlabelled data, the individuals in T, and applying an inexpensive iterative neighbourhood's majority vote based prediction algorithm (“iNMV”) to this task. We evaluate our iNMV prediction algorithm and two previously proposed methods using Markov chains on three real world graphs in terms of ROC AUC statistic. iNMV obtains rankings that are either significantly better or not significantly worse than the rankings obtained from the more complex Markov chain based algorithms, while achieving a reduction in run time of one order of magnitude on large graphs.
Susanne Hoche, Peter A. Flach, David Hardcastle
ECAI2
2008 Querying and Merging Heterogeneous Data by Approximate Joins on Higher-Order Terms
Simon Price, Peter A. Flach
ILP2
2008 Learning the morphology of Zulu with different degrees of supervision
abstract
In this paper we compare different levels of supervision for learning the morphology of the indigenous South African language Zulu. After a preliminary analysis of the Zulu data used for our experiments, we concentrate on supervised, semi-supervised and unsupervised approaches comparing strengths and weaknesses of each method. The challenges we face are limited data availability and data sparsity in connection with morphological analysis of indigenous languages. At the end of the paper we draw conclusions for our future work towards a morphological analyzer for Zulu.
Sebastian Spiegler, Bruno Golénia, Kseniya B. Shalonova, Peter A. Flach, Roger C. F. Tucker
SLT4
2007 A Simple Lexicographic Ranker and Probability Estimator
Peter A. Flach, Edson Takashi Matsubara
ECML1
2007 An Improved Model Selection Heuristic for AUC
Shaomin Wu, Peter A. Flach, Cèsar Ferri
ECML2
2007 Putting Things in Order: On the Fundamental Role of Ranking in Classification and Probability Estimation
Peter A. Flach
ECML/PKDD1
2006 Towards Automating Simulation-Based Design Verification Using ILP
Kerstin Eder, Peter A. Flach, Hsiou-Wen Hsueh
ILP2
2005 Combining Bayesian Networks with Higher-Order Data Representations
Elias Gyftodimos, Peter A. Flach
IDA2
2005 Repairing Concavities in ROC Curves
Peter A. Flach, Shaomin Wu
IJCAI1
2005 ROCCER: An Algorithm for Rule Learning Based on ROC Analysis
Ronaldo C. Prati, Peter A. Flach
IJCAI2
2005 A Response to Webb and Ting's On the Application of ROC Analysis to Predict Classification Performance Under Varying Class Distributions
Tom Fawcett, Peter A. Flach
Mach. Learn.2
2005 ROC 'n' Rule Learning - Towards a Better Understanding of Covering Algorithms
Johannes Fürnkranz, Peter A. Flach
Mach. Learn.2
2004 An Analysis of Stopping and Filtering Criteria for Rule Learning
Johannes Fürnkranz, Peter A. Flach
ECML2
2004 Redundant feature elimination for multi-class problems
abstract
We consider the problem of eliminating redundant Boolean features for a given data set, where a feature is redundant if it separates the classes less well than another feature or set of features. Lavrač et al. proposed the algorithm REDUCE that works by pairwise comparison of features, i.e., it eliminates a feature if it is redundant with respect to another feature. Their algorithm operates in an ILP setting and is restricted to two-class problems. In this paper we improve their method and extend it to multiple classes. Central to our approach is the notion of a neighbourhood of examples: a set of examples of the same class where the number of different features between examples is relatively small. Redundant features are eliminated by applying a revised version of the REDUCE method to each pair of neighbourhoods of different class. We analyse the performance of our method on a range of data sets.
Annalisa Appice, Michelangelo Ceci, Simon Alan Rawles, Peter A. Flach
ICML4
2004 Delegating classifiers
abstract
A sensible use of classifiers must be based on the estimated reliability of their predictions. A cautious classifier would delegate the difficult or uncertain predictions to other, possibly more specialised, classifiers. In this paper we analyse and develop this idea of delegating classifiers in a systematic way. First, we design a two-step scenario where a first classifier chooses which examples to classify and delegates the difficult examples to train a second classifier. Secondly, we present an iterated scenario involving an arbitrary number of chained classifiers. We compare these scenarios to classical ensemble methods, such as bagging and boosting. We show experimentally that our approach is not far behind these methods in terms of accuracy, but with several advantages: (i) improved efficiency, since each classifier learns from fewer examples than the previous one; (ii) improved comprehensibility, since each classification derives from a single classifier; and (iii) the possibility to simplify the overall multi-classifier by removing the parts that lead to delegation.
Cèsar Ferri, Peter A. Flach, José Hernández-Orallo
ICML2
2004 Subgroup Discovery with CN2-SD
Nada Lavrac, Branko Kavsek, Peter A. Flach, Ljupco Todorovski
J. Mach. Learn. Res.3
2004 Naive Bayesian Classification of Structured Data
Peter A. Flach, Nicolas Lachiche
Mach. Learn.1
2004 Kernels and Distances for Structured Data
Thomas Gärtner 0001, John W. Lloyd, Peter A. Flach
Mach. Learn.3
2004 Decision Support Through Subgroup Discovery: Three Case Studies and the Lessons Learned
Nada Lavrac, Bojan Cestnik, Dragan Gamberger, Peter A. Flach
Mach. Learn.4
2004 Book review: Logic for Learning: Learning Comprehensible Theories from Structured Data by John W. Lloyd, Springer-Verlag, 2003, ISBN 3-540-42027-4
Peter A. Flach
Theory Pract. Log. Program.1
2003 Improving the AUC of Probabilistic Estimation Trees
Cèsar Ferri, Peter A. Flach, José Hernández-Orallo
ECML2
2003 The Geometry of ROC Space: Understanding Machine Learning Metrics through ROC Isometrics
Peter A. Flach
ICML1
2003 An Analysis of Rule Evaluation Metrics
Johannes Fürnkranz, Peter A. Flach
ICML2
2003 Improving Accuracy and Cost of Two-class and Multi-class Probabilistic Classifiers Using ROC Curves
Nicolas Lachiche, Peter A. Flach
ICML2
2003 Comparative Evaluation of Approaches to Propositionalization
Mark-A. Krogel, Simon Alan Rawles, Filip Zelezný, Peter A. Flach, Nada Lavrac, Stefan Wrobel
ILP4
2003 Improved Distances for Structured Data
Dimitrios Mavroeidis, Peter A. Flach
ILP2
2002 Improved Dataset Characterisation for Meta-learning
Yonghong Peng, Peter A. Flach, Carlos Soares, Pavel Brazdil
Discovery Science2
2002 Adapting classification rule induction to subgroup discovery
abstract
Rule learning is typically used for solving classification and prediction tasks. However learning of classification rules can be adapted also to subgroup discovery. This paper shows how this can be achieved by modifying the covering algorithm and the search heuristic, performing probabilistic classification of instances, and using an appropriate measure for evaluating the results of subgroup discovery. Experimental evaluation of the CN2-SD subgroup discovery algorithm on 17 UCI data sets demonstrates substantial reduction of the number of induced rules, increased rule coverage and rule significance, as well as slight improvements in terms of the area under the ROC curve.
Nada Lavrac, Peter A. Flach, Branko Kavsek, Ljupco Todorovski
ICDM2
2002 Learning Decision Trees Using the Area Under the ROC Curve
Cèsar Ferri, Peter A. Flach, José Hernández-Orallo
ICML2
2002 Multi-Instance Kernels
Thomas Gärtner 0001, Peter A. Flach, Adam Kowalczyk, Alexander J. Smola
ICML2
2002 Kernels for Structured Data
Thomas Gärtner 0001, John W. Lloyd, Peter A. Flach
ILP3
2002 1BC2: A True First-Order Bayesian Classifier
Nicolas Lachiche, Peter A. Flach
ILP2
2002 RSD: Relational Subgroup Discovery through First-Order Feature Construction
Nada Lavrac, Filip Zelezný, Peter A. Flach
ILP3
2001 WBCsvm: Weighted Bayesian Classification based on Support Vector Machines
Thomas Gärtner 0001, Peter A. Flach
ICML2
2001 On the state of the art in machine learning: A personal review
Peter A. Flach
Artif. Intell.1
2001 Editorial: Inductive Logic Programming is Coming of Age
Peter A. Flach, Saso Dzeroski
Mach. Learn.1
2001 Confirmation-Guided Discovery of First-Order Rules with Tertius
Peter A. Flach, Nicolas Lachiche
Mach. Learn.1
2001 An extended transformation approach to inductive logic programming
abstract
Inductive logic programming (ILP) is concerned with learning relational descriptions that typically have the form of logic programs. In a transformation approach, an ILP task is transformed into an equivalent learning task in a different representation formalism. Propositionalization is a particular transformation method, in which the ILP task is compiled to an attribute-value learning task. The main restriction of propositionalization methods such as LINUS is that they are unable to deal with nondeterminate local variables in the body of hypothesis clauses. In this paper we show how this limitation can be overcome., by systematic first-order feature construction using a particular individual-centered feature bias. The approach can be applied in any domain where there is a clear notion of individual. We also show how to improve upon exhaustive first-order feature construction by using a relevancy filter. The proposed approach is illustrated on the “trains” and “mutagenesis” ILP domains.
Nada Lavrac, Peter A. Flach
ACM Trans. Comput. Log.2
2000 Predictive Performance of Weghted Relative Accuracy
Ljupco Todorovski, Peter A. Flach, Nada Lavrac
PKDD2
2000 Discovery of multivalued dependencies from relations
Iztok Savnik, Peter A. Flach
Intell. Data Anal.2
1998 Comparing Consequence Relations
Peter A. Flach
KR1
1996 Rationality Postulates for Induction
Peter A. Flach
TARK1
1993 Predicate Invention in Inductive Data Engineering
Peter A. Flach
ECML1
1991 Towards a Theory of Inductive Logic Programming
Peter A. Flach
ISMIS1