EDBT 2026 Demo / reviewers in the wild / expert
Eibe Frank
dblp:f/EibeFrank
· DBLP profile ↗
78ranked-venue papers
17as first author
12since 2021 · last 2025
0000-0001-6152-7111ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 12 first-author · 9 since 2021Databases, data management, data science and information retrieval · 38 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multiple Instance VerificationabstractWe explore multiple instance verification, a problem setting in which a query instance is verified against a bag of target instances with heterogeneous, unknown relevancy. We show that naive adaptations of attention-based multiple instance learning (MIL) methods and standard verification methods like Siamese neural networks are unsuitable for this setting: directly combining state-of-the-art (SOTA) MIL methods and Siamese networks is shown to be no better, and sometimes significantly worse, than a simple baseline model. Postulating that this may be caused by the failure of the representation of the target bag to incorporate the query instance, we introduce a new pooling approach named “cross-attention pooling” (CAP). Under the CAP framework, we propose two novel attention functions to address the challenge of distinguishing between highly similar instances in a target bag. Through empirical studies on three different verification tasks, we demonstrate that CAP outperforms adaptations of SOTA MIL methods and the baseline by substantial margins, in terms of both classification accuracy and the ability to detect key instances. The superior ability to identify key instances is attributed to the new attention functions by ablation studies. Eibe Frank, Geoff Holmes 0001 |
J. Mach. Learn. Res. | 2 |
| 2024 | Hitting the target: stopping active learning at the cost-based optimumabstractAbstract Active learning allows machine learning models to be trained using fewer labels while retaining similar performance to traditional supervised learning. An active learner selects the most informative data points, requests their labels, and retrains itself. While this approach is promising, it raises the question of how to determine when the model is ‘good enough’ without the additional labels required for traditional evaluation. Previously, different stopping criteria have been proposed aiming to identify the optimal stopping point. Yet, optimality can only be expressed as a domain-dependent trade-off between accuracy and the number of labels, and no criterion is superior in all applications. As a further complication, a comparison of criteria for a particular real-world application would require practitioners to collect additional labelled data they are aiming to avoid by using active learning in the first place. This work enables practitioners to employ active learning by providing actionable recommendations for which stopping criteria are best for a given real-world scenario. We contribute the first large-scale comparison of stopping criteria for pool-based active learning, using a cost measure to quantify the accuracy/label trade-off, public implementations of all stopping criteria we evaluate, and an open-source framework for evaluating stopping criteria. Our research enables practitioners to substantially reduce labelling costs by utilizing the stopping criterion which best suits their domain. Zac Pullar-Strecker, Katharina Dost, Eibe Frank, Jörg Wicker |
Mach. Learn. | 3 |
| 2024 | Feature extractor stacking for cross-domain few-shot learning
Hongyu Wang 0008, Eibe Frank, Bernhard Pfahringer, Michael Mayo, Geoff Holmes 0001 |
Mach. Learn. | 2 |
| 2023 | Large scale K-means clustering using GPUsabstractAbstract The k-means algorithm is widely used for clustering, compressing, and summarizing vector data. We present a fast and memory-efficient GPU-based algorithm for exact k-means, Asynchronous Selective Batched K-means (ASB K-means). Unlike most GPU-based k-means algorithms that require loading the whole dataset onto the GPU for clustering, the amount of GPU memory required to run our algorithm can be chosen to be much smaller than the size of the whole dataset. Thus, our algorithm can cluster datasets whose size exceeds the available GPU memory. The algorithm works in a batched fashion and applies the triangle inequality in each k-means iteration to omit a data point if its membership assignment, i.e., the cluster it belongs to, remains unchanged, thus significantly reducing the number of data points that need to be transferred between the CPU’s RAM and the GPU’s global memory and enabling the algorithm to very efficiently process large datasets. Our algorithm can be substantially faster than a GPU-based implementation of standard k-means even in situations when application of the standard algorithm is feasible because the whole dataset fits into GPU memory. Experiments show that ASB K-means can run up to 15x times faster than a standard GPU-based implementation of k-means, and it also outperforms the GPU-based k-means implementation in NVIDIA’s open-source RAPIDS machine learning library on all the datasets used in our experiments. Eibe Frank, Bernhard Pfahringer |
Data Min. Knowl. Discov. | 2 |
| 2023 | teex: A toolbox for the evaluation of explanationsabstractWe present teex, a Python toolbox for the evaluation of explanations. teex focuses on the evaluation of local explanations of the predictions of machine learning models by comparing them to ground-truth explanations. It supports several types of explanations: feature importance vectors, saliency maps, decision rules, and word importance maps. A collection of evaluation metrics is provided for each type. Real-world datasets and generators of synthetic data with ground-truth explanations are also contained within the library. teex contributes to research on explainable AI by providing tested, streamlined, user-friendly tools to compute quality metrics for the evaluation of explanation methods. Source code and a basic overview can be found at github.com/chus-chus/teex, and tutorials and full API documentation are at teex.readthedocs.io. Jesus Antonanzas, Yunzhe Jia, Eibe Frank, Albert Bifet, Bernhard Pfahringer |
Neurocomputing | 3 |
| 2022 | Efficiently correcting machine learning: considering the role of example ordering in human-in-the-loop training of image classification modelsabstractArguably the most popular application task in artificial intelligence is image classification using transfer learning. Transfer learning enables models pre-trained on general classes of images, available in large numbers, to be refined for a specific application. This enables domain experts with their own—generally, substantially smaller—collections of images to build deep learning models. The good performance of such models poses the question of whether it is possible to further reduce the effort required to label training data by adopting a human-in-the-loop interface that presents the expert with the current predictions of the model on a new batch of data and only requires correction of these predictions—rather than de novo labelling by the expert—before retraining the model on the extended data. This paper looks at how to order the data in this iterative training scheme to achieve the highest model performance while minimising the effort needed to correct misclassified examples. Experiments are conducted involving five methods of ordering, using four image classification datasets, and three popular pre-trained models. Two of the methods we consider order the examples a priori whereas the other three employ an active learning approach where the ordering is updated iteratively after each new batch of data and retraining of the model. The main finding is that it is important to consider accuracy of the model in relation to the number of corrections that are required: using accuracy in relation to the number of labelled training examples—as is common practice in the literature—can be misleading. More specifically, active methods require more cumulative corrections than a priori methods for a given level of accuracy. Within their groups, active and a priori methods perform similarly. Preliminary evidence is provided that suggests that for “simple” problems, i.e., those involving fewer examples and classes, no method improves upon random selection of examples. For more complex problems, an a priori strategy based on a greedy sample selection method known as “kernel herding” performs best. Geoff Holmes 0001, Eibe Frank, Dale Fletcher, Corey Sterling |
IUI | 2 |
| 2022 | A simple but strong baseline for online continual learning: Repeated Augmented RehearsalabstractOnline continual learning (OCL) aims to train neural networks incrementally from a non-stationary data stream with a single pass through data. Rehearsal-based methods attempt to approximate the observed input distributions over time with a small memory and revisit them later to avoid forgetting. Despite their strong empirical performance, rehearsal methods still suffer from a poor approximation of past data’s loss landscape with memory samples. This paper revisits the rehearsal dynamics in online settings. We provide theoretical insights on the inherent memory overfitting risk from the viewpoint of biased and dynamic empirical risk minimization, and examine the merits and limits of repeated rehearsal.Inspired by our analysis, a simple and intuitive baseline, repeated augmented rehearsal (RAR), is designed to address the underfitting-overfitting dilemma of online rehearsal. Surprisingly, across four rather different OCL benchmarks,this simple baseline outperforms vanilla rehearsal by 9\%-17\% and also significantly improves the state-of-the-art rehearsal-based methods MIR, ASER, and SCR. We also demonstrate that RAR successfully achieves an accurate approximation of the loss landscape of past data and high-loss ridge aversion in its learning trajectory. Extensive ablation studies are conducted to study the interplay between repeated and augmented rehearsal, and reinforcement learning (RL) is applied to dynamically adjust the hyperparameters of RAR to balance the stability-plasticity trade-off online. Yaqian Zhang 0004, Bernhard Pfahringer, Eibe Frank, Albert Bifet, Nick Jin Sean Lim, Yunzhe Jia |
NeurIPS | 3 |
| 2022 | Sampling Permutations for Shapley Value EstimationabstractGame-theoretic attribution techniques based on Shapley values are used to interpret black-box machine learning models, but their exact calculation is generally NP-hard, requiring approximation methods for non-trivial models. As the computation of Shapley values can be expressed as a summation over a set of permutations, a common approach is to sample a subset of these permutations for approximation. Unfortunately, standard Monte Carlo sampling methods can exhibit slow convergence, and more sophisticated quasi-Monte Carlo methods have not yet been applied to the space of permutations. To address this, we investigate new approaches based on two classes of approximation methods and compare them empirically. First, we demonstrate quadrature techniques in a RKHS containing functions of permutations, using the Mallows kernel in combination with kernel herding and sequential Bayesian quadrature. The RKHS perspective also leads to quasi-Monte Carlo type error bounds, with a tractable discrepancy measure defined on permutations. Second, we exploit connections between the hypersphere $\mathbb{S}^{d-2}$ and permutations to create practical algorithms for generating permutation samples with good properties. Experiments show the above techniques provide significant improvements for Shapley value estimates over existing methods, converging to a smaller RMSE in the same number of model evaluations. Rory Mitchell, Joshua N. Cooper, Eibe Frank, Geoff Holmes 0001 |
J. Mach. Learn. Res. | 3 |
| 2021 | Studying and Exploiting the Relationship Between Model Accuracy and Explanation Quality
Yunzhe Jia, Eibe Frank, Bernhard Pfahringer, Albert Bifet, Nick Jin Sean Lim |
ECML/PKDD (2) | 2 |
| 2021 | Classifier Chains: A Review and PerspectivesabstractThe family of methods collectively known as classifier chains has become a popular approach to multi-label learning problems. This approach involves chaining together off-the-shelf binary classifiers in a directed structure, such that individual label predictions become features for other classifiers. Such methods have proved flexible and effective and have obtained state-of-the-art empirical performance across many datasets and multi-label evaluation metrics. This performance led to further studies of the underlying mechanism and efficacy, and investigation into how it could be improved. In the recent decade, numerous studies have explored the theoretical underpinnings of classifier chains, and many improvements have been made to the training and inference procedures, such that this method remains among the best options for multi-label learning. Given this past and ongoing interest, which covers a broad range of applications and research themes, the goal of this work is to provide a review of classifier chains, a survey of the techniques and extensions provided in the literature, as well as perspectives for this approach in the domain of multi-label classification in the future. We conclude positively, with a number of recommendations for researchers and practitioners, as well as outlining key issues for future research. Jesse Read, Bernhard Pfahringer, Geoff Holmes 0001, Eibe Frank |
J. Artif. Intell. Res. | 4 |
| 2021 | Regularisation of neural networks by enforcing Lipschitz continuityabstractAbstract We investigate the effect of explicitly enforcing the Lipschitz continuity of neural networks with respect to their inputs. To this end, we provide a simple technique for computing an upper bound to the Lipschitz constant—for multiple p-norms—of a feed forward neural network composed of commonly used layer types. Our technique is then used to formulate training a neural network with a bounded Lipschitz constant as a constrained optimisation problem that can be solved using projected stochastic gradient methods. Our evaluation study shows that the performance of the resulting models exceeds that of models trained with other common regularisers. We also provide evidence that the hyperparameters are intuitive to tune, demonstrate how the choice of norm for computing the Lipschitz constant impacts the resulting model, and show that the performance gains provided by our method are particularly noticeable when only a small amount of training data is available. Henry Gouk, Eibe Frank, Bernhard Pfahringer, Michael J. Cree |
Mach. Learn. | 2 |
| 2021 | An Empirical Study of Moment Estimators for Quantile ApproximationabstractWe empirically evaluate lightweight moment estimators for the single-pass quantile approximation problem, including maximum entropy methods and orthogonal series with Fourier, Cosine, Legendre, Chebyshev and Hermite basis functions. We show how to apply stable summation formulas to offset numerical precision issues for higher-order moments, leading to reliable single-pass moment estimators up to order 15. Additionally, we provide an algorithm for GPU-accelerated quantile approximation based on parallel tree reduction. Experiments evaluate the accuracy and runtime of moment estimators against the state-of-the-art KLL quantile estimator on 14,072 real-world datasets drawn from the OpenML database. Our analysis highlights the effectiveness of variants of moment-based quantile approximation for highly space efficient summaries: their average performance using as few as five sample moments can approach the performance of a KLL sketch containing 500 elements. Experiments also illustrate the difficulty of applying the method reliably and showcases which moment-based approximations can be expected to fail or perform poorly. Rory Mitchell, Eibe Frank, Geoff Holmes 0001 |
ACM Trans. Database Syst. | 2 |
| 2020 | Adaptive XGBoost for Evolving Data StreamsabstractBoosting is an ensemble method that combines base models in a sequential manner to achieve high predictive accuracy. A popular learning algorithm based on this ensemble method is eXtreme Gradient Boosting (XGB). We present an adaptation of XGB for classification of evolving data streams. In this setting, new data arrives over time and the relationship between the class and the features may change in the process, thus exhibiting concept drift. The proposed method creates new members of the ensemble from mini-batches of data as new data becomes available. The maximum ensemble size is fixed, but learning does not stop when this size is reached because the ensemble is updated on new data to ensure consistency with the current concept. We also explore the use of concept drift detection to trigger a mechanism to update the ensemble. We test our method on real and synthetic data with concept drift and compare it against batch-incremental and instance-incremental classification methods for data streams. Jacob Montiel, Rory Mitchell, Eibe Frank, Bernhard Pfahringer, Talel Abdessalem, Albert Bifet |
IJCNN | 3 |
| 2020 | Embedding Java Classes with code2vec: Improvements from Variable ObfuscationabstractAutomatic source code analysis in key areas of software engineering, such as code security, can benefit from Machine Learning (ML). However, many standard ML approaches require a numeric representation of data and cannot be applied directly to source code. Thus, to enable ML, we need to embed source code into numeric feature vectors while maintaining the semantics of the code as much as possible. code2vec is a recently released embedding approach that uses the proxy task of method name prediction to map Java methods to feature vectors. However, experimentation with code2vec shows that it learns to rely on variable names for prediction, causing it to be easily fooled by typos or adversarial attacks. Moreover, it is only able to embed individual Java methods and cannot embed an entire collection of methods such as those present in a typical Java class, making it difficult to perform predictions at the class level (e.g., for the identification of malicious Java classes). Both shortcomings are addressed in the research presented in this paper. We investigate the effect of obfuscating variable names during training of a code2vec model to force it to rely on the structure of the code rather than specific names and consider a simple approach to creating class-level embeddings by aggregating sets of method embeddings. Our results, obtained on a challenging new collection of source-code classification problems, indicate that obfuscating variable names produces an embedding model that is both impervious to variable naming and more accurately reflects code semantics. The datasets, models, and code are shared1 for further ML research on source code. Rhys Compton, Eibe Frank, Panos Patros, Abigail M. Y. Koay |
MSR | 2 |
| 2019 | Stochastic Gradient TreesabstractWe present an algorithm for learning decision trees using stochastic gradient information as the source of supervision. In contrast to previous approaches to gradient-based tree learning, our method operates in the incremental learning setting rather than the batch learning setting, and does not make use of soft splits or require the construction of a new tree for every update. We demonstrate how one can apply these decision trees to different problems by changing only the loss function, using classification, regression, and multi-instance learning as example applications. In the experimental evaluation, our method performs similarly to standard incremental classification trees, outperforms state of the art incremental regression trees, and achieves comparable performance with batch multi-instance learning methods. Henry Gouk, Bernhard Pfahringer, Eibe Frank |
ACML | 3 |
| 2019 | On Calibration of Nested Dichotomies
Tim Leathart, Eibe Frank, Bernhard Pfahringer, Geoff Holmes 0001 |
PAKDD (1) | 2 |
| 2019 | Ensembles of Nested Dichotomies with Multiple Subset Evaluation
Tim Leathart, Eibe Frank, Bernhard Pfahringer, Geoff Holmes 0001 |
PAKDD (1) | 2 |
| 2019 | AffectiveTweets: a Weka Package for Analyzing Affect in TweetsabstractAffectiveTweets is a set of programs for analyzing emotion and sentiment of social media messages such as tweets. It is implemented as a package for the Weka machine learning workbench and provides methods for calculating state-of-the-art affect analysis features from tweets that can be fed into machine learning algorithms implemented in Weka. It also implements methods for building affective lexicons and distant supervision methods for training affective models from unlabeled tweets. The package was used by several teams in the shared tasks: EmoInt 2017 and Affect in Tweets SemEval 2018 Task 1. Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer, Saif M. Mohammad |
J. Mach. Learn. Res. | 2 |
| 2019 | WekaDeeplearning4j: A deep learning package for Weka based on Deeplearning4j
Steven Lang, Felipe Bravo-Marquez, Christopher Beckham, Mark A. Hall, Eibe Frank |
Knowl. Based Syst. | 5 |
| 2018 | MaxGain: Regularisation of Neural Networks by Constraining Activation Magnitudes
Henry Gouk, Bernhard Pfahringer, Eibe Frank, Michael J. Cree |
ECML/PKDD (1) | 3 |
| 2018 | Online estimation of discrete, continuous, and conditional joint densities using classifier chains
Michael Geilke, Andreas Karwath, Eibe Frank, Stefan Kramer 0001 |
Data Min. Knowl. Discov. | 3 |
| 2018 | Transferring sentiment knowledge between words and tweetsabstractMessage-level and word-level polarity classification are two popular tasks in Twitter sentiment analysis. They have been commonly addressed by training supervised models from labelled data. The main limitation of these models is the high cost of data annotation. Transferring existing labels from a related problem domain is one possible solution for this problem. In this paper, we study how to transfer sentiment labels from the word domain to the tweet domain and vice versa by making their corresponding instances compatible. We model instances of these two domains as the aggregation of instances from the other (i.e., tweets are treated as collections of the words they contain and words are treated as collections of the tweets in which they occur) and perform aggregation by averaging the corresponding constituents. We study two different setups for averaging tweet and word vectors: 1) representing tweets by standard NLP features such as unigrams and part-of-speech tags and words by averaging the vectors of the tweets in which they occur, and 2) representing words using skip-gram embeddings and tweets as the average embedding vector of their words. A consequence of our approach is that instances of both domains reside in the same feature space. Thus, a sentiment classifier trained on labelled data from one domain can be used to classify instances from the other one. We evaluate this approach in two transfer learning tasks: 1) sentiment classification of tweets by applying a word-level sentiment classifier, and 2) induction of a polarity lexicon by applying a tweet-level polarity classifier. Our results show that the proposed model can successfully classify words and tweets after transfer. Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer |
Web Intell. | 2 |
| 2017 | Probability Calibration TreesabstractObtaining accurate and well calibrated probability estimates from classifiers is useful in many applications, for example, when minimising the expected cost of classifications. Existing methods of calibrating probability estimates are applied globally, ignoring the potential for improvements by applying a more fine-grained model. We propose probability calibration trees, a modification of logistic model trees that identifies regions of the input space in which different probability calibration models are learned to improve performance. We compare probability calibration trees to two widely used calibration methods—isotonic regression and Platt scaling—and show that our method results in lower root mean squared error on average than both methods, for estimates produced by a variety of base learners. Tim Leathart, Eibe Frank, Geoff Holmes 0001, Bernhard Pfahringer |
ACML | 2 |
| 2017 | Learning Through Utility Optimization in Regression TasksabstractAccounting for misclassification costs is important in many practical applications of machine learning, and cost-sensitive techniques for classification have been studied extensively. Utility-based learning provides a generalization of purely cost-based approaches that considers both costs and benefits, enabling application to domains with complex cost-benefit settings. However, there is little work on utility- or cost-based learning for regression. In this paper, we formally define the problem of utility-based regression and propose a strategy for maximizing the utility of regression models. We verify our findings in a large set of experiments that show the advantage of our proposal in a diverse set of domains, learning algorithms and cost/benefit settings. Paula Branco, Luís Torgo, Rita P. Ribeiro, Eibe Frank, Bernhard Pfahringer, Markus Michael Rau |
DSAA | 4 |
| 2016 | Annotate-Sample-Average (ASA): A New Distant Supervision Approach for Twitter Sentiment AnalysisabstractThe classification of tweets into polarity classes is a popular task in sentiment analysis. State-of-the-art solutions to this problem are based on supervised machine learning models trained from manually annotated examples. A drawback of these approaches is the high cost involved in data annotation. Two freely available resources that can be exploited to solve the problem are: 1) large amounts of unlabelled tweets obtained from the Twitter API and 2) prior lexical knowledge in the form of opinion lexicons. In this paper, we propose Annotate-Sample-Average (ASA), a distant supervision method that uses these two resources to generate synthetic training data for Twitter polarity classification. Positive and negative training instances are generated by sampling and averaging unlabelled tweets containing words with the corresponding polarity. Polarity of words is determined from a given polarity lexicon. Our experimental results show that the training data generated by ASA (after tuning its parameters) produces a classifier that performs significantly better than a classifier trained from tweets annotated with emoticons and a classifier trained, without any sampling and averaging, from tweets annotated according to the polarity of their words. Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer |
ECAI | 2 |
| 2016 | Building Ensembles of Adaptive Nested Dichotomies with Random-Pair Selection
Tim Leathart, Bernhard Pfahringer, Eibe Frank |
ECML/PKDD (2) | 3 |
| 2016 | Determining Word-Emotion Associations from Tweets by Multi-label ClassificationabstractThe automatic detection of emotions in Twitter posts is a challenging task due to the informal nature of the language used in this platform. In this paper, we propose a methodology for expanding the NRC word-emotion association lexicon for the language used in Twitter. We perform this expansion using multi-label classification of words and compare different word-level features extracted from unlabelled tweets such as unigrams, Brown clusters, POS tags, and word2vec embeddings. The results show that the expanded lexicon achieves major improvements over the original lexicon when classifying tweets into emotional categories. In contrast to previous work, our methodology does not depend on tweets annotated with emotional hashtags, thus enabling the identification of emotional words from any domain-specific collection using unlabelled tweets. Felipe Bravo-Marquez, Eibe Frank, Saif M. Mohammad, Bernhard Pfahringer |
WI | 2 |
| 2016 | From Opinion Lexicons to Sentiment Classification of Tweets and Vice Versa: A Transfer Learning ApproachabstractMessage-level and word-level polarity classification are two popular tasks in Twitter sentiment analysis. They have been commonly addressed by training supervised models from labelled data. The main limitation of these models is the high cost of data annotation. Transferring existing labels from a related problem domain is one possible solution for this problem. In this paper, we propose a simple model for transferring sentiment labels from words to tweets and vice versa by representing both tweets and words using feature vectors residing in the same feature space. Tweets are represented by standard NLP features such as unigrams and part-of-speech tags. Words are represented by averaging the vectors of the tweets in which they occur. We evaluate our approach in two transfer learning problems: 1) training a tweet-level polarity classifier from a polarity lexicon, and 2) inducing a polarity lexicon from a collection of polarity-annotated tweets. Our results show that the proposed approach can successfully classify words and tweets after transfer. Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer |
WI | 2 |
| 2016 | Building a Twitter opinion lexicon from automatically-annotated tweets
Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer |
Knowl. Based Syst. | 2 |
| 2015 | Positive, Negative, or Neutral: Learning an Expanded Opinion Lexicon from Emoticon-Annotated Tweets
Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer |
IJCAI | 2 |
| 2015 | From Unlabelled Tweets to Twitter-specific Opinion WordsabstractIn this article, we propose a word-level classification model for automatically generating a Twitter-specific opinion lexicon from a corpus of unlabelled tweets. The tweets from the corpus are represented by two vectors: a bag-of-words vector and a semantic vector based on word-clusters. We propose a distributional representation for words by treating them as the centroids of the tweet vectors in which they appear. The lexicon generation is conducted by training a word-level classifier using these centroids to form the instance space and a seed lexicon to label the training instances. Experimental results show that the two types of tweet vectors complement each other in a statistically significant manner and that our generated lexicon produces significant improvements for tweet-level polarity classification. Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer |
SIGIR | 2 |
| 2013 | Online Estimation of Discrete DensitiesabstractWe address the problem of estimating a discrete joint density online, that is, the algorithm is only provided the current example and its current estimate. The proposed online estimator of discrete densities, EDDO (Estimation of Discrete Densities Online), uses classifier chains to model dependencies among features. Each classifier in the chain estimates the probability of one particular feature. Because a single chain may not provide a reliable estimate, we also consider ensembles of classifier chains and ensembles of weighted classifier chains. For all density estimators, we provide consistency proofs and propose algorithms to perform certain inference tasks. The empirical evaluation of the estimators is conducted in several experiments and on data sets of up to several million instances: We compare them to density estimates computed from Bayesian structure learners, evaluate them under the influence of noise, measure their ability to deal with concept drift, and measure the run-time performance. Our experiments demonstrate that, even though designed to work online, EDDO delivers estimators of competitive accuracy compared to batch Bayesian structure learners and batch variants of EDDO. Michael Geilke, Eibe Frank, Andreas Karwath, Stefan Kramer 0001 |
ICDM | 2 |
| 2012 | Learning a concept-based document similarity measureabstractDocument similarity measures are crucial components of many text‐analysis tasks, including information retrieval, document classification, and document clustering. Conventional measures are brittle: They estimate the surface overlap between documents based on the words they mention and ignore deeper semantic connections. We propose a new measure that assesses similarity at both the lexical and semantic levels, and learns from human judgments how to combine them by using machine‐learning techniques. Experiments show that the new measure produces values for documents that are more consistent with people's judgments than people are with each other. We also use it to classify and cluster large document sets covering different genres and topics, and find that it improves both classification and clustering performance. Anna-Lan Huang, David N. Milne, Eibe Frank, Ian H. Witten |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2012 | Ensembles of Restricted Hoeffding TreesabstractThe success of simple methods for classification shows that is is often not necessary to model complex attribute interactions to obtain good classification accuracy on practical problems. In this article, we propose to exploit this phenomenon in the data stream context by building an ensemble of Hoeffding trees that are each limited to a small subset of attributes. In this way, each tree is restricted to model interactions between attributes in its corresponding subset. Because it is not known a priori which attribute subsets are relevant for prediction, we build exhaustive ensembles that consider all possible attribute subsets of a given size. As the resulting Hoeffding trees are not all equally important, we weigh them in a suitable manner to obtain accurate classifications. This is done by combining the log-odds of their probability estimates using sigmoid perceptrons, with one perceptron per class. We propose a mechanism for setting the perceptrons’ learning rate using the change detection method for data streams, and also use to reset ensemble members (i.e., Hoeffding trees) when they no longer perform well. Our experiments show that the resulting ensemble classifier outperforms bagging for data streams in terms of accuracy when both are used in conjunction with adaptive naive Bayes Hoeffding trees, at the expense of runtime and memory consumption. We also show that our stacking method can improve the performance of a bagged ensemble. Albert Bifet, Eibe Frank, Geoff Holmes 0001, Bernhard Pfahringer |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2011 | Classifier chains for multi-label classificationabstractThe widely known binary relevance method for multi-label classification, which considers each label as an independent binary problem, has often been overlooked in the literature due to the perceived inadequacy of not directly modelling label correlations. Most current methods invest considerable complexity to model interdependencies between labels. This paper shows that binary relevance-based methods have much to offer, and that high predictive performance can be obtained without impeding scalability to large datasets. We exemplify this with a novel classifier chains method that can model label correlations while maintaining acceptable computational complexity. We extend this approach further in an ensemble framework. An extensive empirical evaluation covers a broad range of multi-label datasets with a variety of evaluation metrics. The results illustrate the competitiveness of the chaining method against related and state-of-the-art methods, both in terms of predictive performance and time complexity. Jesse Read, Bernhard Pfahringer, Geoff Holmes 0001, Eibe Frank |
Mach. Learn. | 4 |
| 2010 | Fast Conditional Density Estimation for Quantitative Structure-Activity RelationshipsabstractMany methods for quantitative structure-activity relationships (QSARs) deliver point estimates only, without quantifying the uncertainty inherent in the prediction. One way to quantify the uncertainy of a QSAR prediction is to predict the conditional density of the activity given the structure instead of a point estimate. If a conditional density estimate is available, it is easy to derive prediction intervals of activities. In this paper, we experimentally evaluate and compare three methods for conditional density estimation for their suitability in QSAR modeling. In contrast to traditional methods for conditional density estimation, they are based on generic machine learning schemes, more specifically, class probability estimators. Our experiments show that a kernel estimator based on class probability estimates from a random forest classifier is highly competitive with Gaussian process regression, while taking only a fraction of the time for training. Therefore, generic machine-learning based methods for conditional density estimation may be a good and fast option for quantifying uncertainty in QSAR modeling. Fabian Buchwald, Tobias Girschick, Eibe Frank, Stefan Kramer 0001 |
AAAI | 3 |
| 2010 | Sentiment Knowledge Discovery in Twitter Streaming Data
Albert Bifet, Eibe Frank |
Discovery Science | 2 |
| 2010 | Speeding Up and Boosting Diverse Density Learning
James R. Foulds, Eibe Frank |
Discovery Science | 2 |
| 2010 | Fast Perceptron Decision Tree Learning from Evolving Data Streams
Albert Bifet, Geoff Holmes 0001, Bernhard Pfahringer, Eibe Frank |
PAKDD (2) | 4 |
| 2010 | WEKA - Experiences with a Java Open-Source Project
Remco R. Bouckaert, Eibe Frank, Mark A. Hall, Geoff Holmes 0001, Bernhard Pfahringer, Peter Reutemann, Ian H. Witten |
J. Mach. Learn. Res. | 2 |
| 2010 | A Study of Hierarchical and Flat Classification of ProteinsabstractAutomatic classification of proteins using machine learning is an important problem that has received significant attention in the literature. One feature of this problem is that expert-defined hierarchies of protein classes exist and can potentially be exploited to improve classification performance. In this article, we investigate empirically whether this is the case for two such hierarchies. We compare multiclass classification techniques that exploit the information in those class hierarchies and those that do not, using logistic regression, decision trees, bagged decision trees, and support vector machines as the underlying base learners. In particular, we compare hierarchical and flat variants of ensembles of nested dichotomies. The latter have been shown to deliver strong classification performance in multiclass settings. We present experimental results for synthetic, fold recognition, enzyme classification, and remote homology detection data. Our results show that exploiting the class hierarchy improves performance on the synthetic data but not in the case of the protein classification problems. Based on this, we recommend that strong flat multiclass methods be used as a baseline to establish the benefit of exploiting class hierarchies in this area. Arthur Zimek, Fabian Buchwald, Eibe Frank, Stefan Kramer 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2009 | Conditional Density Estimation with Class Probability Estimators
Eibe Frank, Remco R. Bouckaert |
ACML | 1 |
| 2009 | Large-scale attribute selection using wrappersabstractScheme-specific attribute selection with the wrapper and variants of forward selection is a popular attribute selection technique for classification that yields good results. However, it can run the risk of overfitting because of the extent of the search and the extensive use of internal cross-validation. Moreover, although wrapper evaluators tend to achieve superior accuracy compared to filters, they face a high computational cost. The problems of overfitting and high runtime occur in particular on high-dimensional datasets, like microarray data. We investigate Linear Forward Selection, a technique to reduce the number of attributes expansions in each forward selection step. Our experiments demonstrate that this approach is faster, finds smaller subsets and can even increase the accuracy compared to standard forward selection. We also investigate a variant that applies explicit subset size determination in forward selection to combat overfitting, where the search is forced to stop at a precomputed “optimal” subset size. We show that this technique reduces subset size while maintaining comparable accuracy. Martin Gütlein, Eibe Frank, Mark A. Hall, Andreas Karwath |
CIDM | 2 |
| 2009 | Human-competitive tagging using automatic keyphrase extraction
Olena Medelyan, Eibe Frank, Ian H. Witten |
EMNLP | 2 |
| 2009 | Clustering Documents Using a Wikipedia-Based Concept Representation
Anna-Lan Huang, David N. Milne, Eibe Frank, Ian H. Witten |
PAKDD | 3 |
| 2009 | Classifier Chains for Multi-label Classification
Jesse Read, Bernhard Pfahringer, Geoff Holmes 0001, Eibe Frank |
ECML/PKDD (2) | 4 |
| 2009 | Accuracy of machine learning models versus "hand crafted" expert systems - A credit scoring case study
Arie Ben-David, Eibe Frank |
Expert Syst. Appl. | 2 |
| 2008 | Clustering Documents with Active Learning Using WikipediaabstractWikipedia has been applied as a background knowledge base to various text mining problems, but very few attempts have been made to utilize it for document clustering. In this paper we propose to exploit the semantic knowledge in Wikipedia for clustering, enabling the automatic grouping of documents with similar themes. Although clustering is intrinsically unsupervised, recent research has shown that incorporating supervision improves clustering performance, even when limited supervision is provided. The approach presented in this paper applies supervision using active learning. We first utilize Wikipedia to create a concept-based representation of a text document, with each concept associated to a Wikipedia article. We then exploit the semantic relatedness between Wikipedia concepts to find pair-wise instance-level constraints for supervised clustering, guiding clustering towards the direction indicated by the constraints. We test our approach on three standard text document datasets. Empirical results show that our basic document representation strategy yields comparable performance to previous attempts; and adding constraints improves clustering performance further by up to 20%. Anna-Lan Huang, David N. Milne, Eibe Frank, Ian H. Witten |
ICDM | 3 |
| 2008 | One-Class Classification by Combining Density and Class Probability Estimation
Kathryn Hempstalk, Eibe Frank, Ian H. Witten |
ECML/PKDD (1) | 2 |
| 2007 | An Empirical Comparison of Exact Nearest Neighbour Algorithms
Ashraf M. Kibriya, Eibe Frank |
PKDD | 2 |
| 2006 | Improving on Bagging with Input Smearing
Eibe Frank, Bernhard Pfahringer |
PAKDD | 1 |
| 2006 | Naive Bayes for Text Classification with Unbalanced Classes
Eibe Frank, Remco R. Bouckaert |
PKDD | 1 |
| 2005 | Ensembles of Balanced Nested Dichotomies for Multi-class Problems
Eibe Frank, Stefan Kramer 0001 |
PKDD | 2 |
| 2005 | Unsupervised Discretization Using Tree-Based Density Estimation
Gabi Schmidberger, Eibe Frank |
PKDD | 2 |
| 2005 | Speeding Up Logistic Model Tree Induction
Marc Sumner, Eibe Frank, Mark A. Hall |
PKDD | 2 |
| 2005 | Logistic Model Trees
Niels Landwehr, Mark A. Hall, Eibe Frank |
Mach. Learn. | 3 |
| 2004 | Ensembles of nested dichotomies for multi-class problemsabstractNested dichotomies are a standard statistical technique for tackling certain polytomous classification problems with logistic regression. They can be represented as binary trees that recursively split a multi-class clas-sification task into a system of dichotomies and provide a statistically sound way of applying two-class learning algorithms to multi-class prob-lems (assuming these algorithms generate class probability estimates). However, there are usually many candidate trees for a given problem and in the standard approach the choice of a particular tree is based on do-main knowledge that may not be available in practice. An alternative is to treat every system of nested dichotomies as equally likely and to form an ensemble classifier based on this assumption. We show that this approach produces more accurate classifications than applying C4.5 and logistic regression directly to multi-class problems. Our results also show that ensembles of nested dichotomies produce more accurate classifiers than pairwise classification if both techniques are used with C4.5, and compa-rable results for logistic regression. Compared to error-correcting output codes, they are preferable if logistic regression is used, and comparable in the case of C4.5. An additional benefit is that they generate class proba-bility estimates. Consequently they appear to be a good general-purpose method for applying binary classifiers to multi-class problems. 1 Eibe Frank, Stefan Kramer 0001 |
ICML | 1 |
| 2004 | Evaluating the Replicability of Significance Tests for Comparing Learning Algorithms
Remco R. Bouckaert, Eibe Frank |
PAKDD | 2 |
| 2004 | Logistic Regression and Boosting for Labeled Bags of Instances
Eibe Frank |
PAKDD | 2 |
| 2004 | Data mining in bioinformatics using WekaabstractUNLABELLED: The Weka machine learning workbench provides a general-purpose environment for automatic classification, regression, clustering and feature selection-common data mining problems in bioinformatics research. It contains an extensive collection of machine learning algorithms and data pre-processing methods complemented by graphical user interfaces for data exploration and the experimental comparison of different machine learning techniques on the same problem. Weka can process data given in the form of a single relational table. Its main objectives are to (a) assist users in extracting useful information from data and (b) enable them to easily identify a suitable algorithm for generating an accurate predictive model from it. AVAILABILITY: http://www.cs.waikato.ac.nz/ml/weka. Eibe Frank, Mark A. Hall, Leonard E. Trigg, Geoff Holmes 0001, Ian H. Witten |
Bioinform. | 1 |
| 2004 | Predicting Library of Congress classifications from Library of Congress subject headingsabstractAbstract This paper addresses the problem of automatically assigning a Library of Congress Classification (LCC) to a work given its set of Library of Congress Subject Headings (LCSH). LCCs are organized in a tree: The root node of this hierarchy comprises all possible topics, and leaf nodes correspond to the most specialized topic areas defined. We describe a procedure that, given a resource identified by its LCSH, automatically places that resource in the LCC hierarchy. The procedure uses machine learning techniques and training data from a large library catalog to learn a model that maps from sets of LCSH to classifications from the LCC tree. We present empirical results for our technique showing its accuracy on an independent collection of 50,000 LCSH/LCC pairs. Eibe Frank, Gordon W. Paynter |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | Logistic Model Trees
Niels Landwehr, Mark A. Hall, Eibe Frank |
ECML | 3 |
| 2003 | A Two-Level Learning Method for Generalized Multi-instance Problems
Nils B. Weidmann, Eibe Frank, Bernhard Pfahringer |
ECML | 2 |
| 2003 | Visualizing Class Probability Estimators
Eibe Frank, Mark A. Hall |
PKDD | 1 |
| 2003 | Locally Weighted Naive Bayes
Eibe Frank, Mark A. Hall, Bernhard Pfahringer |
UAI | 1 |
| 2002 | Racing Committees for Large Datasets
Eibe Frank, Geoff Holmes 0001, Richard Kirkby, Mark A. Hall |
Discovery Science | 1 |
| 2002 | Multiclass Alternating Decision Trees
Geoff Holmes 0001, Bernhard Pfahringer, Richard Kirkby, Eibe Frank, Mark A. Hall |
ECML | 4 |
| 2001 | A Simple Approach to Ordinal Classification
Eibe Frank, Mark A. Hall |
ECML | 1 |
| 2001 | Determining Progression in Glaucoma Using Visual Fields
Andrew Turpin, Eibe Frank, Mark A. Hall, Ian H. Witten, Chris A. Johnson 0002 |
PAKDD | 2 |
| 2001 | Interactive machine learning: letting users build classifiers
Malcolm Ware, Eibe Frank, Geoff Holmes 0001, Mark A. Hall, Ian H. Witten |
Int. J. Hum. Comput. Stud. | 2 |
| 2000 | Text Categorization Using Compression ModelsabstractSummary form only given. Test categorization is the assignment of natural language texts to predefined categories based on their concept. The use of predefined categories implies a "supervised learning" approach to categorization, where already-classified articles which effectively define the categories are used as "training data" to build a model that can be used for classifying new articles that comprise "the data". Typical approaches extract features from articles and use the feature vectors as input to a machine learning scheme that learns how to classify articles. The features are generally words. It has often been observed that compression seems to provide a very promising alternative approach to categorization. The overall compression of an article with respect to different models can be compared to see which one it fits most closely. Such a scheme has several potential advantages: it yields an overall judgement on the document as a whole, rather than discarding information by pre-selecting features it avoids the messy and rather artificial problem of defining word boundaries; it deals uniformly with morphological variants of words; depending on the model (and its order), it can take account of phrasal effects that span word boundaries; it offers a uniform way of dealing with different types of documents for example, arbitrary files in a computer system; it generally minimizes arbitrary decisions that inevitably need to be taken to render any learning scheme practical. Eibe Frank, Chang Chui, Ian H. Witten |
Data Compression Conference | 1 |
| 2000 | Naive Bayes for Regression (Technical Note)abstractAbstract. Despite its simplicity, the naive Bayes learning scheme performs well on most classification tasks, and is often significantly more accurate than more sophisticated methods. Although the probability estimates that it produces can be inaccurate, it often assigns maximum probability to the correct class. This suggests that its good performance might be restricted to situations where the output is categorical. It is therefore interesting to see how it performs in domains where the predicted value is numeric, because in this case, predictions are more sensitive to inaccurate probability estimates. This paper shows how to apply the naive Bayes methodology to numeric prediction (i.e., regression) tasks by modeling the probability distribution of the target value with kernel density estimators, and compares it to linear regression, locally weighted linear regression, and a method that produces “model trees”—decision trees with linear regression functions at the leaves. Although we exhibit an artificial dataset for which naive Bayes is the method of choice, on real-world datasets it is almost uniformly worse than locally weighted linear regression and model trees. The comparison with linear regression depends on the error measure: for one measure naive Bayes performs similarly, while for another it is worse. We also show that standard naive Bayes applied to regression problems by discretizing the target value performs similarly badly. We then present empirical evidence that isolates naive Bayes ’ independence assumption as the culprit for its poor performance in the regression setting. These results indicate that the simplistic statistical assumption that naive Bayes makes is indeed more restrictive for regression than for classification. Eibe Frank, Leonard E. Trigg, Geoff Holmes 0001, Ian H. Witten |
Mach. Learn. | 1 |
| 1999 | Making Better Use of Global Discretization
Eibe Frank, Ian H. Witten |
ICML | 1 |
| 1999 | Domain-Specific Keyphrase Extraction
Eibe Frank, Gordon W. Paynter, Ian H. Witten, Carl Gutwin, Craig G. Nevill-Manning |
IJCAI | 1 |
| 1999 | Improving browsing in digital libraries with keyphrase indexes
Carl Gutwin, Gordon W. Paynter, Ian H. Witten, Craig G. Nevill-Manning, Eibe Frank |
Decis. Support Syst. | 5 |
| 1998 | Generating Accurate Rule Sets Without Global Optimization
Eibe Frank, Ian H. Witten |
ICML | 1 |
| 1998 | Using a Permutation Test for Attribute Selection in Decision Trees
Eibe Frank, Ian H. Witten |
ICML | 1 |
| 1998 | Using Model Trees for Classification
Eibe Frank, Stuart Inglis, Geoff Holmes 0001, Ian H. Witten |
Mach. Learn. | 1 |