EDBT 2026 Demo / reviewers in the wild / expert
Peter A. Flach
dblp:f/PeterAFlach
· DBLP profile ↗
115ranked-venue papers
19as first author
12since 2021 · last 2024
0000-0001-6857-5810ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 92 · 15 first-author · 7 since 2021Databases, data management, data science and information retrieval · 27 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 1 since 2021Theory of computation · 11 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Explaining a Probabilistic Prediction on the Simplex with Shapley CompositionsabstractOriginating in game theory, Shapley values are widely used for explaining a machine learning model’s prediction by quantifying the contribution of each feature’s value to the prediction. This requires a scalar prediction as in binary classification, whereas a multiclass probabilistic prediction is a discrete probability distribution, living on a multidimensional simplex. In such a multiclass setting the Shapley values are typically computed separately on each class in a one-vs-rest manner, ignoring the compositional nature of the output distribution. In this paper, we introduce Shapley compositions as a well-founded way to properly explain a multiclass probabilistic prediction, using the Aitchison geometry from compositional data analysis. We prove that the Shapley composition is the unique quantity satisfying linearity, symmetry and efficiency on the Aitchison simplex, extending the corresponding axiomatic properties of the standard Shapley value. We demonstrate this proper multiclass treatment in a range of scenarios. Paul-Gauthier Noé, Miquel Perelló-Nieto, Jean-François Bonastre, Peter A. Flach |
ECAI | 4 |
| 2024 | Interpretable representations in explainable AI: from theory to practiceabstractAbstract Interpretable representations are the backbone of many explainers that target black-box predictive systems based on artificial intelligence and machine learning algorithms. They translate the low-level data representation necessary for good predictive performance into high-level human-intelligible concepts used to convey the explanatory insights. Notably, the explanation type and its cognitive complexity are directly controlled by the interpretable representation, tweaking which allows to target a particular audience and use case. However, many explainers built upon interpretable representations overlook their merit and fall back on default solutions that often carry implicit assumptions, thereby degrading the explanatory power and reliability of such techniques. To address this problem, we study properties of interpretable representations that encode presence and absence of human-comprehensible concepts. We demonstrate how they are operationalised for tabular, image and text data; discuss their assumptions, strengths and weaknesses; identify their core building blocks; and scrutinise their configuration and parameterisation. In particular, this in-depth analysis allows us to pinpoint their explanatory properties, desiderata and scope for (malicious) manipulation in the context of tabular data where a linear model is used to quantify the influence of interpretable concepts on a black-box prediction. Our findings lead to a range of recommendations for designing trustworthy interpretable representations; specifically, the benefits of class-aware (supervised) discretisation of tabular data, e.g., with decision trees, and sensitivity of image interpretable representations to segmentation granularity and occlusion colour. Kacper Sokol, Peter A. Flach |
Data Min. Knowl. Discov. | 2 |
| 2023 | Reconciling Training and Evaluation Objectives in Location Agnostic Surrogate ExplainersabstractTransparency in AI models is crucial to designing, auditing, and deploying AI systems. However, 'black box' models are still used in practice for their predictive power despite their lack of transparency. This has led to a demand for post-hoc, model-agnostic surrogate explainers which provide explanations for decisions of any model by approximating its behaviour close to a query point with a surrogate model. However, it is often overlooked how the location of the query point in the decision surface of the black box model affects the faithfulness of the surrogate explainer. Here, we show that when using standard techniques, there is a decrease in agreement between the black box and the surrogate model for query points towards the edge of the test dataset and when moving away from the decision boundary. This originates from a mismatch between the data distributions used to train and evaluate surrogate explainers. We address this by leveraging knowledge about the test data distribution captured in the class labels of the black box model. By addressing this and encouraging users to take care in understanding the alignment of training and evaluation objectives, we empower them to construct more faithful surrogate explainers. Matthew Clifford, Jonathan Erskine, Alexander Hepburn, Peter A. Flach, Raúl Santos-Rodríguez |
CIKM | 4 |
| 2023 | Co-designing opportunities for Human-Centred Machine Learning in supporting Type 1 diabetes decision-makingabstractType 1 Diabetes (T1D) self-management requires hundreds of daily decisions. Diabetes technologies that use machine learning have significant potential to simplify this process and provide better decision support, but often rely on cumbersome data logging and cognitively demanding reflection on collected data. We set out to use co-design to identify opportunities for machine learning to support diabetes self-management in everyday settings. However, over nine months of interviews and design workshops with 15 people with T1D, we had to re-assess our assumptions about user needs. Our participants reported confidence in their personal knowledge and rejected machine learning based decision support when coping with routine situations, but highlighted the need for technological support in the context of unfamiliar or unexpected situations (holidays, illness, etc.). However, these are the situations where prior data are often lacking and drawing data-driven conclusions is challenging. Reflecting this challenge, we provide suggestions on how machine learning and other artificial intelligence approaches, e.g., expert systems, could enable decision-making support in both routine and unexpected situations. Katarzyna Stawarz, Dmitri S. Katz, Amid Ayobi, Paul Marshall, Taku Yamagata, Raúl Santos-Rodríguez, Peter A. Flach, Aisling Ann O'Kane |
Int. J. Hum. Comput. Stud. | 7 |
| 2023 | Classifier calibration: a survey on how to assess and improve predicted class probabilitiesabstractAbstract This paper provides both an introduction to and a detailed overview of the principles and practice of classifier calibration. A well-calibrated classifier correctly quantifies the level of uncertainty or confidence associated with its instance-wise predictions. This is essential for critical applications, optimal decision making, cost-sensitive classification, and for some types of context change. Calibration research has a rich history which predates the birth of machine learning as an academic field by decades. However, a recent increase in the interest on calibration has led to new methods and the extension from binary to the multiclass setting. The space of options and issues to consider is large, and navigating it requires the right set of concepts and tools. We provide both introductory material and up-to-date technical details of the main concepts and methods, including proper scoring rules and other evaluation metrics, visualisation approaches, a comprehensive account of post-hoc calibration methods for binary and multiclass classification, and several advanced topics. Telmo de Menezes e Silva Filho, Hao Song 0007, Miquel Perelló-Nieto, Raúl Santos-Rodríguez, Meelis Kull, Peter A. Flach |
Mach. Learn. | 6 |
| 2022 | LIMESegment: Meaningful, Realistic Time Series ExplanationsabstractLIME (Locally Interpretable Model-Agnostic Explanations) has become a popular way of generating explanations for tabular, image and natural language models, providing insight into why an instance was given a particular classification. In this paper we adapt LIME to time series classification, an under-explored area with existing approaches failing to account for the structure of this kind of data. We frame the non-trivial challenge of adapting LIME to time series classification as the following open questions: “What is a meaningful interpretable representation of a time series?”, “How does one realistically perturb a time series?” and “What is a local neighbourhood around a time series?”. We propose solutions to all three questions and combine them into a novel time series explanation framework called LIMESegment, which outperforms existing adaptations of LIME to time series on a variety of classification tasks. Torty Sivill, Peter A. Flach |
AISTATS | 2 |
| 2022 | Self-Enhancer: A Self-supervised Framework for Low-Supervision, Drifted Data with Significant Missing Values
Yu Chen 0092, Peter A. Flach |
ICANN (4) | 2 |
| 2022 | Understanding Reinforcement Learning Based Localisation as a Probabilistic Inference Algorithm
Taku Yamagata, Raúl Santos-Rodríguez, Robert J. Piechocki, Peter A. Flach |
ICANN (2) | 4 |
| 2021 | Multi-label thresholding for cost-sensitive classification
Reem Alotaibi, Peter A. Flach |
Neurocomputing | 2 |
| 2021 | Co-Designing Personal Health? Multidisciplinary Benefits and Challenges in Informing Diabetes Self-Care TechnologiesabstractCo-design is a widely applied design process with well-documented values, including mutual learning and collective creativity. However, the real-world challenges of conducting multidisciplinary co-design research to inform the design of self-care technologies are not well established. We provide a qualitative account of a multidisciplinary project that aimed to co-design machine learning applications for Type 1 Diabetes (T1D) self-management. Through interviews, we identify not only perceived social, technological and strategic benefits of co-design but also organisational, translational and pragmatic design challenges: participants with T1D experienced difficulties in co-designing systems that met their individual self-care needs as part of group activities; HCI and AI researchers described challenges resulting from applying co-design outcomes to data-driven ML work; and industry collaborators highlighted academic data sharing regulations as cross-organisational challenges that can impede co-design efforts. Based on this understanding, we discuss opportunities for supporting multidisciplinary collaborations and aligning individual health needs with collaborative co-design activities. Amid Ayobi, Katarzyna Stawarz, Dmitri S. Katz, Paul Marshall, Taku Yamagata, Raúl Santos-Rodríguez, Peter A. Flach, Aisling Ann O'Kane |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2021 | Human Activity Recognition Based on Dynamic Active LearningabstractActivity of daily living is an important indicator of the health status and functional capabilities of an individual. Activity recognition, which aims at understanding the behavioral patterns of people, has increasingly received attention in recent years. However, there are still a number of challenges confronting the task. First, labelling training data is expensive and time-consuming, leading to limited availability of annotations. Secondly, activities performed by individuals have considerable variability, which renders the generally used supervised learning with a fixed label set unsuitable. To address these issues, we propose a dynamic active learning-based activity recognition method in this work. Different from traditional active learning methods which select samples based on a fixed label set, the proposed method not only selects informative samples from known classes, but also dynamically identifies new activities which are not included in the predefined label set. Starting with a classifier that has access to a limited number of labelled samples, we iteratively extend the training set with informative labels by fully considering the uncertainty, diversity and representativeness of samples, based on which better-informed classifiers can be trained, further reducing the annotation cost. We evaluate the proposed method on two synthetic datasets and two existing benchmark datasets. Experimental results demonstrate that our method not only boosts the activity recognition performance with considerably reduced annotation cost, but also enables adaptive daily activity analysis allowing the presence and detection of novel activities and patterns. Haixia Bi, Miquel Perelló-Nieto, Raúl Santos-Rodríguez, Peter A. Flach |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | CRISP-DM Twenty Years Later: From Data Mining Processes to Data Science TrajectoriesabstractCRISP-DM(CRoss-Industry Standard Process for Data Mining) has its origins in the second half of the nineties and is thus about two decades old. According to many surveys and user polls it is still the de facto standard for developing data mining and knowledge discovery projects. However, undoubtedly the field has moved on considerably in twenty years, with data science now the leading term being favoured over data mining. In this paper we investigate whether, and in what contexts, CRISP-DM is still fit for purpose for data science projects. We argue that if the project is goal-directed and process-driven the process model view still largely holds. On the other hand, when data science projects become more exploratory the paths that the project can take become more varied, and a more flexible model is called for. We suggest what the outlines of such a trajectory-based model might look like and how it can be used to categorise data science projects (goal-directed, exploratory or data management). We examine seven real-life exemplars where exploratory activities play an important role and compare them against 51 use cases extracted from the NIST Big Data Public Working Group. We anticipate this categorisation can help project planning in terms of time and cost characteristics. Fernando Martínez-Plumed, Lidia Contreras Ochando, Cèsar Ferri, José Hernández-Orallo, Meelis Kull, Nicolas Lachiche, María José Ramírez-Quintana, Peter A. Flach |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2020 | FACE: Feasible and Actionable Counterfactual ExplanationsabstractWork in Counterfactual Explanations tends to focus on the principle of "the closest possible world" that identifies small changes leading to the desired outcome. In this paper we argue that while this approach might initially seem intuitively appealing it exhibits shortcomings not addressed in the current literature. First, a counterfactual example generated by the state-of-the-art systems is not necessarily representative of the underlying data distribution, and may therefore prescribe unachievable goals (e.g., an unsuccessful life insurance applicant with severe disability may be advised to do more sports). Secondly, the counterfactuals may not be based on a "feasible path" between the current state of the subject and the suggested one, making actionable recourse infeasible (e.g., low-skilled unsuccessful mortgage applicants may be told to double their salary, which may be hard without first increasing their skill level). These two shortcomings may render counterfactual explanations impractical and sometimes outright offensive. To address these two major flaws, first of all, we propose a new line of Counterfactual Explanations research aimed at providing actionable and feasible paths to transform a selected instance into one that meets a certain goal. Secondly, we propose FACE: an algorithmically sound way of uncovering these "feasible paths" based on the shortest path distances defined via density-weighted metrics. Our approach generates counterfactuals that are coherent with the underlying data distribution and supported by the "feasible paths" of change, which are achievable and can be tailored to the problem at hand. Rafael Poyiadzi, Kacper Sokol, Raúl Santos-Rodríguez, Tijl De Bie, Peter A. Flach |
AIES | 5 |
| 2020 | Polsar Image Classification via Robust Low-Rank Feature Extraction and Markov Random FieldabstractPolarimetric synthetic aperture radar (PolSAR) image classification has been investigated rigorously in various remote sensing applications. However, it is still a challenging task nowadays. One significant barrier lies in the speckle effect embedded in the PolSAR imaging process, which significantly degrades the quality of the images and further complicates the classification. To address this issue, we present a novel Pol-SAR image classification method which removes speckle noise via robust low-rank feature extraction and enforces smoothness priors through Markov Random Field (MRF). Specifically, we employ the mixture of Gaussian (MoG) based low-rank matrix factorization (LRMF) to simultaneously extract robust features and remove noise. Then, a classification map is obtained by applying Random Forest (RF) classifier on the extracted LRMF features. Finally, we refine the classification map by Markov random field (MRF) to enforce contextual smoothness. We conduct experiments on two real benchmark PolSAR data sets. Experimental results indicate that the proposed method achieves promising classification performance and preferable spatial consistency. Haixia Bi, Raúl Santos-Rodríguez, Peter A. Flach |
IGARSS | 3 |
| 2020 | Uni- and multivariate probability density models for numeric subgroup discoveryabstractSubgroup Discovery is a supervised, exploratory data mining paradigm that aims to identify subsets of a dataset that show interesting behaviour with respect to some designated target attribute. The way in which such distributional differences are quantified varies with the target attribute type. This work concerns continuous targets, which are important in many practical applications. For such targets, differences are often quantified using z-score and similar measures that compare simple statistics such as the mean and variance of the subset and the data. However, most distributions are not fully determined by their mean and variance alone. As a result, measures of distributional difference solely based on such simple statistics will miss potentially interesting subgroups. This work proposes methods to recognise distributional differences in a much broader sense. To this end, density estimation is performed using histogram and kernel density estimation techniques. In the spirit of Exceptional Model Mining, the proposed methods are extended to deal with multiple continuous target attributes, such that comparisons are not restricted to univariate distributions, but are available for joint distributions of any dimensionality. The methods can be incorporated easily into existing Subgroup Discovery frameworks, so no new frameworks are developed. Marvin Meeng, Harm de Vries, Peter A. Flach, Siegfried Nijssen, Arno J. Knobbe |
Intell. Data Anal. | 3 |
| 2020 | Reflections on reciprocity in research
Peter A. Flach |
Mach. Learn. | 1 |
| 2019 | Performance Evaluation in Machine Learning: The Good, the Bad, the Ugly, and the Way ForwardabstractThis paper gives an overview of some ways in which our understanding of performance evaluation measures for machine-learned classifiers has improved over the last twenty years. I also highlight a range of areas where this understanding is still lacking, leading to ill-advised practices in classifier evaluation. This suggests that in order to make further progress we need to develop a proper measurement theory of machine learning. I then demonstrate by example what such a measurement theory might look like and what kinds of new results it would entail. Finally, I argue that key properties such as classification ability and data set difficulty are unlikely to be directly observable, suggesting the need for latent-variable models and causal inference. Peter A. Flach |
AAAI | 1 |
| 2019 | Desiderata for Interpretability: Explaining Decision Tree Predictions with CounterfactualsabstractExplanations in machine learning come in many forms, but a consensus regarding their desired properties is still emerging. In our work we collect and organise these explainability desiderata and discuss how they can be used to systematically evaluate properties and quality of an explainable system using the case of class-contrastive counterfactual statements. This leads us to propose a novel method for explaining predictions of a decision tree with counterfactuals. We show that our model-specific approach exploits all the theoretical advantages of counterfactual explanations, hence improves decision tree interpretability by decoupling the quality of the interpretation from the depth and width of the tree. Kacper Sokol, Peter A. Flach |
AAAI | 2 |
| 2019 | $β^3$-IRT: A New Item Response Model and its ApplicationsabstractItem Response Theory (IRT) aims to assess latent abilities of respondents based on the correctness of their answers in aptitude test items with different difficulty levels. In this paper, we propose the $\beta^3$-IRT model, which models continuous responses and can generate a much enriched family of Item Characteristic Curves. In experiments we applied the proposed model to data from an online exam platform, and show our model outperforms a more standard 2PL-ND model on all datasets. Furthermore, we show how to apply $\beta^3$-IRT to assess the ability of machine learning classifiers.This novel application results in a new metric for evaluating the quality of the classifier’s probability estimates, based on the inferred difficulty and discrimination of data instances. Telmo de Menezes e Silva Filho, Ricardo B. C. Prudêncio, Tom Diethe, Peter A. Flach |
AISTATS | 5 |
| 2019 | Distribution calibration for regressionabstractWe are concerned with obtaining well-calibrated output distributions from regression models. Such distributions allow us to quantify the uncertainty that the model has regarding the predicted target value. We introduce the novel concept of distribution calibration, and demonstrate its advantages over the existing definition of quantile calibration. We further propose a post-hoc approach to improving the predictions from previously trained regression models, using multi-output Gaussian Processes with a novel Beta link function. The proposed method is experimentally verified on a set of common regression models and shows improvements for both distribution-level and quantile-level calibration. Hao Song 0007, Tom Diethe, Meelis Kull, Peter A. Flach |
ICML | 4 |
| 2019 | Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with Dirichlet calibrationabstractClass probabilities predicted by most multiclass classifiers are uncalibrated, often tending towards over-confidence. With neural networks, calibration can be improved by temperature scaling, a method to learn a single corrective multiplicative factor for inputs to the last softmax layer. On non-neural models the existing methods apply binary calibration in a pairwise or one-vs-rest fashion. We propose a natively multiclass calibration method applicable to classifiers from any model class, derived from Dirichlet distributions and generalising the beta calibration method from binary classification. It is easily implemented with neural nets since it is equivalent to log-transforming the uncalibrated probabilities, followed by one linear layer and softmax. Experiments demonstrate improved probabilistic predictions according to multiple measures (confidence-ECE, classwise-ECE, log-loss, Brier score) across a wide range of datasets and classifiers. Parameters of the learned Dirichlet calibration map provide insights to the biases in the uncalibrated model. Meelis Kull, Miquel Perelló-Nieto, Markus Kängsepp, Telmo de Menezes e Silva Filho, Hao Song 0007, Peter A. Flach |
NeurIPS | 6 |
| 2019 | Setting decision thresholds when operating conditions are uncertainabstractThe quality of the decisions made by a machine learning model depends on the data and the operating conditions during deployment. Often, operating conditions such as class distribution and misclassification costs have changed during the time since the model was trained and evaluated. When deploying a binary classifier that outputs scores, once we know the new class distribution and the new cost ratio between false positives and false negatives, there are several methods in the literature to help us choose an appropriate threshold for the classifier’s scores. However, on many occasions, the information that we have about this operating condition is uncertain . Previous work has considered ranges or distributions of operating conditions during deployment, with expected costs being calculated for ranges or intervals, but still the decision for each point is made as if the operating condition were certain. The implications of this assumption have received limited attention: a threshold choice that is best suited without uncertainty may be suboptimal under uncertainty. In this paper we analyse the effect of operating condition uncertainty on the expected loss for different threshold choice methods, both theoretically and experimentally. We model uncertainty as a second conditional distribution over the actual operation condition and study it theoretically in such a way that minimum and maximum uncertainty are both seen as special cases of this general formulation. This is complemented by a thorough experimental analysis investigating how different learning algorithms behave for a range of datasets according to the threshold choice method and the uncertainty level. Cèsar Ferri, José Hernández-Orallo, Peter A. Flach |
Data Min. Knowl. Discov. | 3 |
| 2019 | An application of hierarchical Gaussian processes to the detection of anomalies in star light curves
Niall Twomey, Haoyan Chen, Tom Diethe, Peter A. Flach |
Neurocomputing | 4 |
| 2018 | Anomaly detection in star light curves using hierarchical Gaussian processes
Haoyan Chen, Tom Diethe, Niall Twomey, Peter A. Flach |
ESANN | 4 |
| 2018 | The Facets of Artificial Intelligence: A Framework to Track the Evolution of AIabstractWe present nine facets for the analysis of the past and future evolution of AI. Each facet has also a set of edges that can summarise different trends and contours in AI. With them, we first conduct a quantitative analysis using the information from two decades of AAAI/IJCAI conferences and around 50 years of documents from AI topics, an official database from the AAAI, illustrated by several plots. We then perform a qualitative analysis using the facets and edges, locating AI systems in the intelligence landscape and the discipline as a whole. This analytical framework provides a more structured and systematic way of looking at the shape and boundaries of AI. Fernando Martínez-Plumed, Bao Sheng Loe, Peter A. Flach, Seán Ó hÉigeartaigh, Karina Vold, José Hernández-Orallo |
IJCAI | 3 |
| 2018 | Conversational Explanations of Machine Learning Predictions Through Class-contrastive Counterfactual StatementsabstractMachine learning models have become pervasive in our everyday life; they decide on important matters influencing our education, employment and judicial system. Many of these predictive systems are commercial products protected by trade secrets, hence their decision-making is opaque. Therefore, in our research we address interpretability and explainability of predictions made by machine learning models. Our work draws heavily on human explanation research in social sciences: contrastive and exemplar explanations provided through a dialogue. This user-centric design, focusing on a lay audience rather than domain experts, applied to machine learning allows explainees to drive the explanation to suit their needs instead of being served a precooked template. Kacper Sokol, Peter A. Flach |
IJCAI | 2 |
| 2018 | Glass-Box: Explaining AI Decisions With Counterfactual Statements Through Conversation With a Voice-enabled Virtual AssistantabstractThe prevalence of automated decision making, influencing important aspects of our lives -- e.g., school admission, job market, insurance and banking -- has resulted in increasing pressure from society and regulators to make this process more transparent and ensure its explainability, accountability and fairness. We demonstrate a prototype voice-enabled device, called Glass-Box, which users can question to understand automated decisions and identify the underlying model's biases and errors. Our system explains algorithmic predictions with class-contrastive counterfactual statements (e.g., ``Had a number of conditions been different:...the prediction would change...''), which show a difference in a particular scenario that causes an algorithm to ``change its mind''. Such explanations do not require any prior technical knowledge to understand, hence are suitable for a lay audience, who interact with the system in a natural way -- through an interactive dialogue. We demonstrate the capabilities of the device by allowing users to impersonate a loan applicant who can question the system to understand the automated decision that he received. Kacper Sokol, Peter A. Flach |
IJCAI | 2 |
| 2018 | Releasing eHealth Analytics into the Wild: Lessons Learnt from the SPHERE ProjectabstractThe SPHERE project is devoted to advancing eHealth in a smart-home context, and supports full-scale sensing and data analysis to enable a generic healthcare service. We describe, from a data-science perspective, our experience of taking the system out of the laboratory into more than thirty homes in Bristol, UK. We describe the infrastructure and processes that had to be developed along the way, describe how we train and deploy Machine Learning systems in this context, and give a realistic appraisal of the state of the deployed systems. Tom Diethe, Mike Holmes, Meelis Kull, Miquel Perelló-Nieto, Kacper Sokol, Hao Song 0007, Emma Tonkin, Niall Twomey, Peter A. Flach |
KDD | 9 |
| 2017 | Beta calibration: a well-founded and easily implemented improvement on logistic calibration for binary classifiersabstractFor optimal decision making under variable class distributions and misclassification costs a classifier needs to produce well-calibrated estimates of the posterior probability. Isotonic calibration is a powerful non-parametric method that is however prone to overfitting on smaller datasets; hence a parametric method based on the logistic curve is commonly used. While logistic calibration is designed for normally distributed per-class scores, we demonstrate experimentally that many classifiers including Naive Bayes and Adaboost suffer from a particular distortion where these score distributions are heavily skewed. In such cases logistic calibration can easily yield probability estimates that are worse than the original scores. Moreover, the logistic curve family does not include the identity function, and hence logistic calibration can easily uncalibrate a perfectly calibrated classifier. In this paper we solve all these problems with a richer class of calibration maps based on the beta distribution. We derive the method from first principles and show that fitting it is as easy as fitting a logistic curve. Extensive experiments show that beta calibration is superior to logistic calibration for Naive Bayes and Adaboost. Meelis Kull, Telmo de Menezes e Silva Filho, Peter A. Flach |
AISTATS | 3 |
| 2017 | The Role of Textualisation and Argumentation in Understanding the Machine Learning ProcessabstractUnderstanding data, models and predictions is important for machine learning applications. Due to the limitations of our spatial perception and intuition, analysing high-dimensional data is inherently difficult. Furthermore, black-box models achieving high predictive accuracy are widely used, yet the logic behind their predictions is often opaque. Use of textualisation -- a natural language narrative of selected phenomena -- can tackle these shortcomings. When extended with argumentation theory we could envisage machine learning models and predictions arguing persuasively for their choices. Kacper Sokol, Peter A. Flach |
IJCAI | 2 |
| 2017 | Unsupervised learning of sensor topologies for improving activity recognition in smart environments
Niall Twomey, Tom Diethe, Ian Craddock, Peter A. Flach |
Neurocomputing | 4 |
| 2016 | Declaratively Capturing Local Label Correlations with Multi-Label TreesabstractThe goal of multi-label classification is to predict multiple labels per data point simultaneously. Real-world applications tend to have high-dimensional label spaces, employing hundreds or even thousands of labels. While these labels could be predicted separately, by capturing label correlation we might achieve better predictive performance. In contrast with previous attempts in the literature that have modelled label correlations globally, this paper proposes a novel algorithm to model correlations and cluster labels locally. LaCovaC is a multi-label decision tree classifier that clusters labels into several dependent subsets at various points during training. The clusters are obtained locally by identifying the conditionally-dependent labels in localised regions of the feature space using the label correlation matrix. LaCovaC interleaves between two main decisions on the label matrix with training instances in rows and labels in columns: splitting this matrix vertically by partitioning the labels into subsets, or splitting it horizontally using features in the conventional way. Experiments on 13 benchmark datasets demonstrate that our proposal achieves competitive performance over a wide range of evaluation metrics when compared with the state-of-the-art multi-label classifiers. Reem Alotaibi, Meelis Kull, Peter A. Flach |
ECAI | 3 |
| 2016 | Active transfer learning for activity recognition
Tom Diethe, Niall Twomey, Peter A. Flach |
ESANN | 3 |
| 2016 | Background Check: A General Technique to Build More Reliable and Versatile ClassifiersabstractWe introduce a powerful technique to make classifiers more reliable and versatile. Background Check equips classifiers with the ability to assess the difference of unlabelled test data from the training data. In particular, Background Check gives classifiers the capability to (i) perform cautious classification with a reject option, (ii) identify outliers, and (iii) better assess the confidence in their predictions. We derive the method from first principles and consider four particular relationships between background and foreground distributions. One of these assumes an affine relationship with two parameters, and we show how this bivariate parameter space naturally interpolates between the above capabilities. We demonstrate the versatility of the approach by comparing it experimentally with published special-purpose solutions for outlier detection and confident classification on 41 benchmark datasets. Results show that Background Check can match and in many cases surpass the performances of specialised approaches. Miquel Perelló-Nieto, Telmo de Menezes e Silva Filho, Meelis Kull, Peter A. Flach |
ICDM | 4 |
| 2016 | ADL™: A Topic Model for Discovery of Activities of Daily Living in a Smart Home
Tom Diethe, Peter A. Flach |
IJCAI | 3 |
| 2016 | Fast Unsupervised Online Drift Detection Using Incremental Kolmogorov-Smirnov TestabstractData stream research has grown rapidly over the last decade. Two major features distinguish data stream from batch learning: stream data are generated on the fly, possibly in a fast and variable rate; and the underlying data distribution can be non-stationary, leading to a phenomenon known as concept drift. Therefore, most of the research on data stream classification focuses on proposing efficient models that can adapt to concept drifts and maintain a stable performance over time. However, specifically for the classification task, the majority of such methods rely on the instantaneous availability of true labels for all already classified instances. This is a strong assumption that is rarely fulfilled in practical applications. Hence there is a clear need for efficient methods that can detect concept drifts in an unsupervised way. One possibility is the well-known Kolmogorov-Smirnov test, a statistical hypothesis test that checks whether two samples differ. This work has two main contributions. The first one is the Incremental Kolmogorov-Smirnov algorithm that allows performing the Kolmogorov-Smirnov hypothesis test instantly using two samples that change over time, where the change is an insertion and/or removal of an observation. Our algorithm employs a randomized tree and is able to perform the insertion and removal operations in O(log N) with high probability and calculate the Kolmogorov-Smirnov test in O(1), where N is the number of sample observations. This is a significant speed-up compared to the O(N log N) cost of the non-incremental implementation. The second contribution is the use of the Incremental Kolmogorov-Smirnov test to detect concept drifts without true labels. Classification algorithms adapted to use the test rely on a limited portion of those labels just to update the classification model after a concept drift is detected. Denis Moreira dos Reis, Peter A. Flach, Stan Matwin, Gustavo Batista |
KDD | 2 |
| 2016 | Subgroup Discovery with Proper Scoring Rules
Hao Song 0007, Meelis Kull, Peter A. Flach, Georgios Kalogridis |
ECML/PKDD (2) | 3 |
| 2016 | Cost-sensitive boosting algorithms: Do we really need them?abstractWe provide a unifying perspective for two decades of work on cost-sensitive Boosting algorithms. When analyzing the literature 1997–2016, we find 15 distinct cost-sensitive variants of the original algorithm; each of these has its own motivation and claims to superiority—so who should we believe? In this work we critique the Boosting literature using four theoretical frameworks: Bayesian decision theory, the functional gradient descent view, margin theory, and probabilistic modelling. Our finding is that only three algorithms are fully supported—and the probabilistic model view suggests that all require their outputs to be calibrated for best performance. Experiments on 18 datasets across 21 degrees of imbalance support the hypothesis—showing that once calibrated, they perform equivalently, and outperform all others. Our final recommendation—based on simplicity, flexibility and performance—is to use the original Adaboost algorithm with a shifted decision threshold and calibrated probability estimates. Nikolaos Nikolaou, Narayanan Unny Edakunni, Meelis Kull, Peter A. Flach, Gavin Brown 0001 |
Mach. Learn. | 4 |
| 2016 | On the need for structure modelling in sequence predictionabstractThere is no uniform approach in the literature for modelling sequential correlations in sequence classification problems. It is easy to find examples of unstructured models ( e.g. logistic regression) where correlations are not taken into account at all, but there are also many examples where the correlations are explicitly incorporated into a—potentially computationally expensive—structured classification model ( e.g. conditional random fields). In this paper we lay theoretical and empirical foundations for clarifying the types of problem which necessitate direct modelling of correlations in sequences, and the types of problem where unstructured models that capture sequential aspects solely through features are sufficient. The theoretical work in this paper shows that the rate of decay of auto-correlations within a sequence is related to the excess classification risk that is incurred by ignoring the structural aspect of the data. This is an intuitively appealing result, demonstrating the intimate link between the auto-correlations and excess classification risk. Drawing directly on this theory, we develop well-founded visual analytics tools that can be applied a priori on data sequences and we demonstrate how these tools can guide practitioners in specifying feature representations based on auto-correlation profiles. Empirical analysis is performed on three sequential datasets. With baseline feature templates, structured and unstructured models achieve similar performance, indicating no initial preference for either model. We then apply the visual analytics tools to the datasets, and show that classification performance in all cases is improved over baseline results when our tools are involved in defining feature representations. Niall Twomey, Tom Diethe, Peter A. Flach |
Mach. Learn. | 3 |
| 2016 | Feature Construction and Calibration for Clustering Daily Load Curves from Smart-Meter DataabstractThis paper proposes and compares feature construction and calibration methods for clustering daily electricity load curves. Such load curves describe electricity demand over a period of time. A rich body of the literature has studied clustering of load curves, usually using temporal features. This limits the potential to discover new knowledge, which may not be best represented as models consisting of all time points on load curves. This paper presents three new methods to construct features: 1) conditional filters on time-resolution-based features; 2) calibration and normalization; and 3) using profile errors. These new features extend the potential of clustering load curves. Moreover, smart metering is now generating high-resolution time series, and so the dimensionality reduction offered by these features is welcome. The clustering results using the proposed new features are compared with clusterings obtained from temporal features, as well as clusterings with Fourier features, using household electricity consumption time series as test data. The experimental results suggest that the proposed feature construction methods offer new means for gaining insight in energy-consumption patterns. Reem Alotaibi, Nanlin Jin, Tom Wilcox, Peter A. Flach |
IEEE Trans. Ind. Informatics | 4 |
| 2015 | Reframing in Frequent Pattern MiningabstractMining frequent patterns is a crucial task in data mining. Most of the existing frequent pattern mining methods find the complete set of frequent patterns from a given dataset. However, in real-life scenarios we often need to predict the future frequent patterns for different tasks such as business policy making, web page recommendation, stock-market behavior and road traffic analysis. Predicting future frequent patterns from the currently available set of frequent patterns is challenging due to dataset shift where data distributions may change from one dataset to another. In this paper, we propose a new approach called reframing in frequent pattern mining to solve this task. Moreover, we experimentally show the existence of dataset shift in two real-life transactional datasets and the capability of our approach to handle these unknown shifts. Chowdhury Farhan Ahmed, Mohammad Samiullah 0001, Nicolas Lachiche, Meelis Kull, Peter A. Flach |
ICTAI | 5 |
| 2015 | Precision-Recall-Gain Curves: PR Analysis Done RightabstractPrecision-Recall analysis abounds in applications of binary classification where true negatives do not add value and hence should not affect assessment of the classifier's performance. Perhaps inspired by the many advantages of receiver operating characteristic (ROC) curves and the area under such curves for accuracy-based performance assessment, many researchers have taken to report Precision-Recall (PR) curves and associated areas as performance metric. We demonstrate in this paper that this practice is fraught with difficulties, mainly because of incoherent scale assumptions -- e.g., the area under a PR curve takes the arithmetic mean of precision values whereas the $F_{\beta}$ score applies the harmonic mean. We show how to fix this by plotting PR curves in a different coordinate system, and demonstrate that the new Precision-Recall-Gain curves inherit all key advantages of ROC curves. In particular, the area under Precision-Recall-Gain curves conveys an expected $F_1$ score on a harmonic scale, and the convex hull of a Precision-Recall-Gain curve allows us to calibrate the classifier's scores so as to determine, for each operating point on the convex hull, the interval of $\beta$ values for which the point optimises $F_{\beta}$. We demonstrate experimentally that the area under traditional PR curves can easily favour models with lower expected $F_1$ score than others, and so the use of Precision-Recall-Gain curves will result in better model selection. Peter A. Flach, Meelis Kull |
NIPS | 1 |
| 2015 | Versatile Decision Trees for Learning Over Multiple Contexts
Reem Alotaibi, Ricardo B. C. Prudêncio, Meelis Kull, Peter A. Flach |
ECML/PKDD (1) | 4 |
| 2015 | Bayesian Modelling of the Temporal Aspects of Smart Home Activity with Circular Statistics
Tom Diethe, Niall Twomey, Peter A. Flach |
ECML/PKDD (2) | 3 |
| 2015 | Novel Decompositions of Proper Scoring Rules for Classification: Score Adjustment as Precursor to Calibration
Meelis Kull, Peter A. Flach |
ECML/PKDD (1) | 2 |
| 2014 | A Machine Learning Approach to Objective Cardiac Event DetectionabstractThis paper presents an automated framework for the detection of the QRS complex from Electrocardiogram (ECG) signals. We introduce an artefact-tolerant pre-processing algorithm which emphasises a number of characteristics of the ECG that are representative of the QRS complex. With this processed ECG signal we train Logistic Regression and Support Vector Machine classification models. With our approach we obtain over 99.7% detection sensitivity and precision on the MIT-BIH database without using supplementary de-noising or pre-emphasis filters. Niall Twomey, Peter A. Flach |
CISIS | 2 |
| 2014 | LaCova: A Tree-Based Multi-label Classifier Using Label Covariance as Splitting CriterionabstractDealing with multiple labels is a supervised learning problem of increasing importance. Multi-label classifiers face the challenge of exploiting correlations between labels. While in existing work these correlations are often modelled globally, in this paper we use the divide-and-conquer approach of decision trees which enables taking local decisions about how best to model label dependency. The resulting algorithm establishes a tree-based multi-label classifier called LaCova which dynamically interpolates between two well-known baseline methods: Binary Relevance, which assumes all labels independent, and Label Power set, which learns the joint label distribution. The key idea is a splitting criterion based on the label covariance matrix at that node, which allows us to choose between a horizontal split (branching on a feature) and a vertical split (separating the labels). Empirical results on 12 data sets show strong performance of the proposed method, particularly on data sets with hundreds of labels. Reem Alotaibi, Meelis Kull, Peter A. Flach |
ICMLA | 3 |
| 2014 | Reliability Maps: A Tool to Enhance Probability Estimates and Improve Classification Accuracy
Meelis Kull, Peter A. Flach |
ECML/PKDD (2) | 2 |
| 2014 | Rate-Constrained Ranking and the Rate-Weighted AUC
Louise A. C. Millard, Peter A. Flach, Julian P. T. Higgins |
ECML/PKDD (2) | 2 |
| 2014 | Rate-Oriented Point-Wise Confidence Bounds for ROC Curves
Louise A. C. Millard, Meelis Kull, Peter A. Flach |
ECML/PKDD (2) | 3 |
| 2014 | Subgroup Discovery in Smart Electricity Meter DataabstractThis work presents data mining methods for discovering unusual consumption patterns and their associated descriptive models from smart electricity meter data. At present, data mining and knowledge discovery in electricity meter data suffer from three notable weaknesses: 1) insufficient focus on intelligent data analysis of subgroups (subsets) whose patterns vary significantly from aggregate patterns embodied in an entire dataset; 2) a lack of effort towards generating intuitively understandable and practically applicable knowledge for industrial practitioners to identify such subgroups; and 3) limited knowledge regarding the link between unusual consumption patterns and household consumers' socio-demographic characteristics. This paper addresses these practically important but technically challenging issues by applying subgroup discovery algorithms to a real smart electricity meter dataset. Subgroups whose patterns are unusual and whose sizes are large enough are discovered, and their descriptive and predictive models are generated. Furthermore, to enrich subgroup discovery algorithms, three new-quality measures for real-valued targets are proposed. The comparative studies empirically evaluate the effectiveness and usefulness of subgroup discovery on classification accuracy, predictive power, and computational resources. The methodologies and algorithms presented are generic, and therefore applicable to a wider range of data mining problems. Nanlin Jin, Peter A. Flach, Tom Wilcox, Royston Sellman, Joshua Thumim, Arno J. Knobbe |
IEEE Trans. Ind. Informatics | 2 |
| 2013 | A Higher-order data flow model for heterogeneous Big DataabstractWe introduce a data flow model that supports highly parallelisable design patterns and also has useful properties for analysing data serially over extended time periods without requiring traditional Big Data computing facilities. The model ranges over a class of higher-order relations which are sufficiently expressive to represent a wide variety of unstructured, semi-structured and structured data. Using JSONMatch, our web service implementation of the model, we show that the combination of this model and higher-order representation provides a powerful and extensible framework that is particularly well suited to analysing Big Variety data in a web application context. Simon Price, Peter A. Flach |
IEEE BigData | 2 |
| 2013 | Guest editors' introduction: special section of selected papers from ECML-PKDD 2012
Tijl De Bie, Peter A. Flach |
Data Min. Knowl. Discov. | 2 |
| 2013 | SubSift web services and workflows for profiling and comparing scientists and their published works
Simon Price, Peter A. Flach, Sebastian Spiegler, Christopher Bailey 0001, Nikki Rogers |
Future Gener. Comput. Syst. | 2 |
| 2013 | Guest editors' introduction: special issue of selected papers from ECML-PKDD 2012
Tijl De Bie, Peter A. Flach |
Mach. Learn. | 2 |
| 2013 | ROC curves in cost space
José Hernández-Orallo, Peter A. Flach, Cèsar Ferri |
Mach. Learn. | 2 |
| 2012 | Caveats and pitfalls of ROC analysis in clinical microarray research (and how to avoid them)abstractThe receiver operating characteristic (ROC) has emerged as the gold standard for assessing and comparing the performance of classifiers in a wide range of disciplines including the life sciences. ROC curves are frequently summarized in a single scalar, the area under the curve (AUC). This article discusses the caveats and pitfalls of ROC analysis in clinical microarray research, particularly in relation to (i) the interpretation of AUC (especially a value close to 0.5); (ii) model comparisons based on AUC; (iii) the differences between ranking and classification; (iv) effects due to multiple hypotheses testing; (v) the importance of confidence intervals for AUC; and (vi) the choice of the appropriate performance metric. With a discussion of illustrative examples and concrete real-world studies, this article highlights critical misconceptions that can profoundly impact the conclusions about the observed performance. Daniel P. Berrar, Peter A. Flach |
Briefings Bioinform. | 2 |
| 2012 | A unified view of performance metrics: translating threshold choice into expected classification loss
José Hernández-Orallo, Peter A. Flach, Cèsar Ferri |
J. Mach. Learn. Res. | 2 |
| 2012 | ILP turns 20 - Biography and future challengesabstractInductive Logic Programming (ILP) is an area of Machine Learning which has now reached its twentieth year. Using the analogy of a human biography this paper recalls the development of the subject from its infancy through childhood and teenage years. We show how in each phase ILP has been characterised by an attempt to extend theory and implementations in tandem with the development of novel and challenging real-world applications. Lastly, by projection we suggest directions for research which will help the subject coming of age. Stephen H. Muggleton, Luc De Raedt, David Poole 0001, Ivan Bratko, Peter A. Flach, Katsumi Inoue, Ashwin Srinivasan 0001 |
Mach. Learn. | 5 |
| 2011 | A Coherent Interpretation of AUC as a Measure of Aggregated Classification Performance
Peter A. Flach, José Hernández-Orallo, Cèsar Ferri |
ICML | 1 |
| 2011 | Brier Curves: a New Cost-Based Visualisation of Classifier Performance
José Hernández-Orallo, Peter A. Flach, Cèsar Ferri |
ICML | 2 |
| 2011 | Smooth Receiver Operating Characteristics (smROC) Curves
William Klement, Peter A. Flach, Nathalie Japkowicz, Stan Matwin |
ECML/PKDD (2) | 2 |
| 2011 | The Machine Learning journal: 25 years young
Peter A. Flach |
Mach. Learn. | 1 |
| 2010 | Enhanced Word Decomposition by Calibrating the Decision Threshold of Probabilistic Models and Using a Model Ensemble
Sebastian Spiegler, Peter A. Flach |
ACL | 2 |
| 2010 | Ukwabelana - An open-source morphological Zulu corpus
Sebastian Spiegler, Andrew van der Spuy, Peter A. Flach |
COLING | 3 |
| 2010 | SubSift Web Services and Workflows for Profiling and Comparing Scientists and Their Published WorksabstractScientific researchers, laboratories and organisations can be profiled and compared by analysing their published works, including documents ranging from academic papers to web sites, blog posts and Twitter feeds. This paper describes how the vector space model from information retrieval, more normally associated with full text search, has been employed in the open source Sub Sift software to support workflows to profile and compare such collections of documents. Sub Sift was originally designed to match submitted conference or journal papers to potential peer reviewers based on the similarity between the paper's abstract and the reviewer's publications as found in online bibliographic databases. The software is implemented as a family of Restful web services that, composed into a re-usable workflow, have already been used to support several major data mining conferences. Alternative workflows and service compositions are now enabling other interesting applications. Simon Price, Peter A. Flach, Sebastian Spiegler, Christopher Bailey 0001, Nikki Rogers |
eScience | 2 |
| 2010 | The Advantages of Seed Examples in First-Order Multi-class Subgroup DiscoveryabstractSubgroup discovery is halfway between predictive and descriptive rule learning: while there is a target concept, the goal of subgroup discovery is not necessarily to achieve high accuracy in predicting the target, but rather to identify subsets of the population whose class distribution is significantly different from the overall distribution. The target concept helps us to achieve a trade-off between accuracy and interestingness. In previous work [1] we illustrated the usefulness of several multiclass subgroup evaluation measures in producing highly predictive subgroups in a propositional logic framework. In this paper we upgrade multi-class subgroup discovery to first-order logic employing a new weighted covering algorithm that takes advantage of learning from seed examples. In a rigourous experimental evaluation we show that the use of seed examples leads to considerable and statistically significant improvement of predictive power, both in terms of accuracy (3 percent points on average over the data sets) and AUC (7 percent points). The paper is structured as follows. Section 2 describes the ingredients of our Aleph-MSD++ algorithm. An empirical evaluation of multi-class subgroup discovery for feature construction over 11 data sets is presented in Section 3. Finally, we conclude the paper in Section 4. Tarek Abudawood, Peter A. Flach |
ECAI | 2 |
| 2010 | Learning Multi-class Theories in ILP
Tarek Abudawood, Peter A. Flach |
ILP | 2 |
| 2010 | The Machine Learning journal: 250 issues and counting
Peter A. Flach |
Mach. Learn. | 1 |
| 2009 | Evaluation Measures for Multi-class Subgroup Discovery
Tarek Abudawood, Peter A. Flach |
ECML/PKDD (1) | 2 |
| 2009 | Towards Learning Morphology for Under-Resourced Fusional and Agglutinating LanguagesabstractIn this paper, we describe a novel and effective approach for automatically decomposing a word into stem and suffixes. Russian and Turkish are used as exemplars of fusional and agglutinating languages. Rather than relying on corpus counts, we use a small number of word-pairs as training data, that can be particularly suited for under-resourced languages. For fusional languages, we initially learn a tree of aligned suffix rules (TASR) from word-pairs. The tree is built top-down, from general to specific rules, using suffix rule frequency and rule subsumption, and is executed bottom-up, i.e., the most specific rule that fires is chosen. TASR is used to segment a word form into a stem and suffix sequence. For fusional languages learning through generation (using TASR) is essential for proper stem extraction. Subsequently, an unsupervised segmentation algorithm graph-based unsupervised suffix segmentation (GBUSS) is used to segment the suffix sequence. GBUSS employs a suffix graph where node merging, guided by an information-theoretic measure, generates suffix sequences. The approach, experimentally validated on Russian, is shown to be highly effective. For agglutinating languages only the GBUSS is needed for word decomposition. Promising experimental results for Turkish are obtained. Kseniya B. Shalonova, Bruno Golénia, Peter A. Flach |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | A Fast Method for Property Prediction in Graph-Structured Data from Positive and Unlabelled ExamplesabstractThe analysis of large and complex networks, or graphs, is becoming increasingly important in many scientific areas including machine learning, social network analysis and bioinformatics. One natural type of question that can be asked in network analysis is “Given two sets R and T of individuals in a graph with complete and missing knowledge, respectively, about a property of interest, which individuals in T are closest to R with respect to this property?”. To answer this question, we can rank the individuals in T such that the individuals ranked highest are most likely to exhibit the property of interest. Several methods based on weighted paths in the graph and Markov chain models have been proposed to solve this task. In this paper, we show that we can improve previously published approaches by rephrasing this problem as the task of property prediction in graph-structured data from positive examples, the individuals in R, and unlabelled data, the individuals in T, and applying an inexpensive iterative neighbourhood's majority vote based prediction algorithm (“iNMV”) to this task. We evaluate our iNMV prediction algorithm and two previously proposed methods using Markov chains on three real world graphs in terms of ROC AUC statistic. iNMV obtains rankings that are either significantly better or not significantly worse than the rankings obtained from the more complex Markov chain based algorithms, while achieving a reduction in run time of one order of magnitude on large graphs. Susanne Hoche, Peter A. Flach, David Hardcastle |
ECAI | 2 |
| 2008 | Querying and Merging Heterogeneous Data by Approximate Joins on Higher-Order Terms
Simon Price, Peter A. Flach |
ILP | 2 |
| 2008 | Learning the morphology of Zulu with different degrees of supervisionabstractIn this paper we compare different levels of supervision for learning the morphology of the indigenous South African language Zulu. After a preliminary analysis of the Zulu data used for our experiments, we concentrate on supervised, semi-supervised and unsupervised approaches comparing strengths and weaknesses of each method. The challenges we face are limited data availability and data sparsity in connection with morphological analysis of indigenous languages. At the end of the paper we draw conclusions for our future work towards a morphological analyzer for Zulu. Sebastian Spiegler, Bruno Golénia, Kseniya B. Shalonova, Peter A. Flach, Roger C. F. Tucker |
SLT | 4 |
| 2007 | A Simple Lexicographic Ranker and Probability Estimator
Peter A. Flach, Edson Takashi Matsubara |
ECML | 1 |
| 2007 | An Improved Model Selection Heuristic for AUC
Shaomin Wu, Peter A. Flach, Cèsar Ferri |
ECML | 2 |
| 2007 | Putting Things in Order: On the Fundamental Role of Ranking in Classification and Probability Estimation
Peter A. Flach |
ECML/PKDD | 1 |
| 2006 | Towards Automating Simulation-Based Design Verification Using ILP
Kerstin Eder, Peter A. Flach, Hsiou-Wen Hsueh |
ILP | 2 |
| 2005 | Combining Bayesian Networks with Higher-Order Data Representations
Elias Gyftodimos, Peter A. Flach |
IDA | 2 |
| 2005 | Repairing Concavities in ROC Curves
Peter A. Flach, Shaomin Wu |
IJCAI | 1 |
| 2005 | ROCCER: An Algorithm for Rule Learning Based on ROC Analysis
Ronaldo C. Prati, Peter A. Flach |
IJCAI | 2 |
| 2005 | A Response to Webb and Ting's On the Application of ROC Analysis to Predict Classification Performance Under Varying Class Distributions
Tom Fawcett, Peter A. Flach |
Mach. Learn. | 2 |
| 2005 | ROC 'n' Rule Learning - Towards a Better Understanding of Covering Algorithms
Johannes Fürnkranz, Peter A. Flach |
Mach. Learn. | 2 |
| 2004 | An Analysis of Stopping and Filtering Criteria for Rule Learning
Johannes Fürnkranz, Peter A. Flach |
ECML | 2 |
| 2004 | Redundant feature elimination for multi-class problemsabstractWe consider the problem of eliminating redundant Boolean features for a given data set, where a feature is redundant if it separates the classes less well than another feature or set of features. Lavrač et al. proposed the algorithm REDUCE that works by pairwise comparison of features, i.e., it eliminates a feature if it is redundant with respect to another feature. Their algorithm operates in an ILP setting and is restricted to two-class problems. In this paper we improve their method and extend it to multiple classes. Central to our approach is the notion of a neighbourhood of examples: a set of examples of the same class where the number of different features between examples is relatively small. Redundant features are eliminated by applying a revised version of the REDUCE method to each pair of neighbourhoods of different class. We analyse the performance of our method on a range of data sets. Annalisa Appice, Michelangelo Ceci, Simon Alan Rawles, Peter A. Flach |
ICML | 4 |
| 2004 | Delegating classifiersabstractA sensible use of classifiers must be based on the estimated reliability of their predictions. A cautious classifier would delegate the difficult or uncertain predictions to other, possibly more specialised, classifiers. In this paper we analyse and develop this idea of delegating classifiers in a systematic way. First, we design a two-step scenario where a first classifier chooses which examples to classify and delegates the difficult examples to train a second classifier. Secondly, we present an iterated scenario involving an arbitrary number of chained classifiers. We compare these scenarios to classical ensemble methods, such as bagging and boosting. We show experimentally that our approach is not far behind these methods in terms of accuracy, but with several advantages: (i) improved efficiency, since each classifier learns from fewer examples than the previous one; (ii) improved comprehensibility, since each classification derives from a single classifier; and (iii) the possibility to simplify the overall multi-classifier by removing the parts that lead to delegation. Cèsar Ferri, Peter A. Flach, José Hernández-Orallo |
ICML | 2 |
| 2004 | Subgroup Discovery with CN2-SD
Nada Lavrac, Branko Kavsek, Peter A. Flach, Ljupco Todorovski |
J. Mach. Learn. Res. | 3 |
| 2004 | Naive Bayesian Classification of Structured Data
Peter A. Flach, Nicolas Lachiche |
Mach. Learn. | 1 |
| 2004 | Kernels and Distances for Structured Data
Thomas Gärtner 0001, John W. Lloyd, Peter A. Flach |
Mach. Learn. | 3 |
| 2004 | Decision Support Through Subgroup Discovery: Three Case Studies and the Lessons Learned
Nada Lavrac, Bojan Cestnik, Dragan Gamberger, Peter A. Flach |
Mach. Learn. | 4 |
| 2004 | Book review: Logic for Learning: Learning Comprehensible Theories from Structured Data by John W. Lloyd, Springer-Verlag, 2003, ISBN 3-540-42027-4
Peter A. Flach |
Theory Pract. Log. Program. | 1 |
| 2003 | Improving the AUC of Probabilistic Estimation Trees
Cèsar Ferri, Peter A. Flach, José Hernández-Orallo |
ECML | 2 |
| 2003 | The Geometry of ROC Space: Understanding Machine Learning Metrics through ROC Isometrics
Peter A. Flach |
ICML | 1 |
| 2003 | An Analysis of Rule Evaluation Metrics
Johannes Fürnkranz, Peter A. Flach |
ICML | 2 |
| 2003 | Improving Accuracy and Cost of Two-class and Multi-class Probabilistic Classifiers Using ROC Curves
Nicolas Lachiche, Peter A. Flach |
ICML | 2 |
| 2003 | Comparative Evaluation of Approaches to Propositionalization
Mark-A. Krogel, Simon Alan Rawles, Filip Zelezný, Peter A. Flach, Nada Lavrac, Stefan Wrobel |
ILP | 4 |
| 2003 | Improved Distances for Structured Data
Dimitrios Mavroeidis, Peter A. Flach |
ILP | 2 |
| 2002 | Improved Dataset Characterisation for Meta-learning
Yonghong Peng, Peter A. Flach, Carlos Soares, Pavel Brazdil |
Discovery Science | 2 |
| 2002 | Adapting classification rule induction to subgroup discoveryabstractRule learning is typically used for solving classification and prediction tasks. However learning of classification rules can be adapted also to subgroup discovery. This paper shows how this can be achieved by modifying the covering algorithm and the search heuristic, performing probabilistic classification of instances, and using an appropriate measure for evaluating the results of subgroup discovery. Experimental evaluation of the CN2-SD subgroup discovery algorithm on 17 UCI data sets demonstrates substantial reduction of the number of induced rules, increased rule coverage and rule significance, as well as slight improvements in terms of the area under the ROC curve. Nada Lavrac, Peter A. Flach, Branko Kavsek, Ljupco Todorovski |
ICDM | 2 |
| 2002 | Learning Decision Trees Using the Area Under the ROC Curve
Cèsar Ferri, Peter A. Flach, José Hernández-Orallo |
ICML | 2 |
| 2002 | Multi-Instance Kernels
Thomas Gärtner 0001, Peter A. Flach, Adam Kowalczyk, Alexander J. Smola |
ICML | 2 |
| 2002 | Kernels for Structured Data
Thomas Gärtner 0001, John W. Lloyd, Peter A. Flach |
ILP | 3 |
| 2002 | 1BC2: A True First-Order Bayesian Classifier
Nicolas Lachiche, Peter A. Flach |
ILP | 2 |
| 2002 | RSD: Relational Subgroup Discovery through First-Order Feature Construction
Nada Lavrac, Filip Zelezný, Peter A. Flach |
ILP | 3 |
| 2001 | WBCsvm: Weighted Bayesian Classification based on Support Vector Machines
Thomas Gärtner 0001, Peter A. Flach |
ICML | 2 |
| 2001 | On the state of the art in machine learning: A personal review
Peter A. Flach |
Artif. Intell. | 1 |
| 2001 | Editorial: Inductive Logic Programming is Coming of Age
Peter A. Flach, Saso Dzeroski |
Mach. Learn. | 1 |
| 2001 | Confirmation-Guided Discovery of First-Order Rules with Tertius
Peter A. Flach, Nicolas Lachiche |
Mach. Learn. | 1 |
| 2001 | An extended transformation approach to inductive logic programmingabstractInductive logic programming (ILP) is concerned with learning relational descriptions that typically have the form of logic programs. In a transformation approach, an ILP task is transformed into an equivalent learning task in a different representation formalism. Propositionalization is a particular transformation method, in which the ILP task is compiled to an attribute-value learning task. The main restriction of propositionalization methods such as LINUS is that they are unable to deal with nondeterminate local variables in the body of hypothesis clauses. In this paper we show how this limitation can be overcome., by systematic first-order feature construction using a particular individual-centered feature bias. The approach can be applied in any domain where there is a clear notion of individual. We also show how to improve upon exhaustive first-order feature construction by using a relevancy filter. The proposed approach is illustrated on the “trains” and “mutagenesis” ILP domains. Nada Lavrac, Peter A. Flach |
ACM Trans. Comput. Log. | 2 |
| 2000 | Predictive Performance of Weghted Relative Accuracy
Ljupco Todorovski, Peter A. Flach, Nada Lavrac |
PKDD | 2 |
| 2000 | Discovery of multivalued dependencies from relations
Iztok Savnik, Peter A. Flach |
Intell. Data Anal. | 2 |
| 1998 | Comparing Consequence Relations
Peter A. Flach |
KR | 1 |
| 1996 | Rationality Postulates for Induction
Peter A. Flach |
TARK | 1 |
| 1993 | Predicate Invention in Inductive Data Engineering
Peter A. Flach |
ECML | 1 |
| 1991 | Towards a Theory of Inductive Logic Programming
Peter A. Flach |
ISMIS | 1 |