Bianca Zadrozny

dblp:77/5253 · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
2since 2021 · last 2022
0000-0002-7260-2057ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 6 first-author · 1 since 2021Databases, data management, data science and information retrieval · 17 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Trustworthy machine learning · 21% Learning theory · 17% Representation and self-supervised learning · 14%
Databases, data mining, and information retrieval
7 papers
Data mining · 90% Information retrieval · 10%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Smart cities and intelligent transportation · 97% Computational social science and digital humanities · 3%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › text embedding
character embedding
0.212014
Learning Character-level Representations for Part-of-Speech Tagging · ICML 2014
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging
0.212014
Learning Character-level Representations for Part-of-Speech Tagging · ICML 2014
Smart cities and intelligent transportation › public transit
bus travel time prediction
0.212014
Bus Travel Time Predictions Using Additive Models · ICDM 2014
Smart cities and intelligent transportation
public transit
0.212014
Bus Travel Time Predictions Using Additive Models · ICDM 2014
Data mining › predictive modeling
regression
0.212014
Bus Travel Time Predictions Using Additive Models · ICDM 2014
Machine learning › Trustworthy machine learning › uncertainty estimation
probability calibration
0.142002
Transforming classifier scores into accurate multiclass probability estimates · KDD 2002
Reducing multiclass to binary by coupling probability estimates · NIPS 2001
Learning and making decisions when costs and probabilities are both unknown · KDD 2001
Machine learning › Trustworthy machine learning › dataset bias
selection bias
0.122005
An Improved Categorization of Classifier's Sensitivity on Sample Selection Bias · ICDM 2005
Learning and evaluating classifiers under sample selection bias · ICML 2004
Machine learning › Learning paradigms
cost-sensitive learning
0.122004
An iterative method for multi-class cost-sensitive learning · KDD 2004
Cost-Sensitive Learning by Cost-Proportionate Example Weighting · ICDM 2003
Data mining › predictive modeling › classification
cost-sensitive learning
0.122002
Sequential cost-sensitive decision making with reinforcement learning · KDD 2002
Learning and making decisions when costs and probabilities are both unknown · KDD 2001
Machine learning › Efficient and distributed learning
active learning
0.112006
Outlier detection by active learning · KDD 2006
Machine learning › Time series and sequential data
anomaly detection
0.112006
Outlier detection by active learning · KDD 2006
Machine learning › Learning theory › query learning
selective sampling
0.112006
Outlier detection by active learning · KDD 2006
Machine learning › Representation and self-supervised learning
word representation
0.112014
Learning Character-level Representations for Part-of-Speech Tagging · ICML 2014
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.122004
Cost-Sensitive Learning by Cost-Proportionate Example Weighting · ICDM 2003
An iterative method for multi-class cost-sensitive learning · KDD 2004
Machine learning › Learning paradigms › cost-sensitive learning
cost-sensitive classification
0.112005
Error limiting reductions between classification tasks · ICML 2005
Machine learning › Reinforcement learning
policy evaluation
0.112005
Relating reinforcement learning performance to classification performance · ICML 2005
Machine learning › Trustworthy machine learning
robustness
0.112005
An Improved Categorization of Classifier's Sensitivity on Sample Selection Bias · ICDM 2005
Machine learning › Learning paradigms
supervised learning
0.112005
Error limiting reductions between classification tasks · ICML 2005
Data mining › predictive modeling
classification
0.112005
An Improved Categorization of Classifier's Sensitivity on Sample Selection Bias · ICDM 2005
Information retrieval
evaluation
0.112005
Ranking-Based Evaluation of Regression Models · ICDM 2005
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.012004
An iterative method for multi-class cost-sensitive learning · KDD 2004
Machine learning › Trustworthy machine learning
debiasing
0.012004
Learning and evaluating classifiers under sample selection bias · ICML 2004
Machine learning › Kernel, tree and ensemble methods
gradient boosting
0.012004
An iterative method for multi-class cost-sensitive learning · KDD 2004
Machine learning › Reinforcement learning
offline reinforcement learning
0.012002
Empirical Comparison of Various Reinforcement Learning Strategies for Sequential Targeted Marketing · ICDM 2002
Machine learning › Learning theory
classification
0.012001
Reducing multiclass to binary by coupling probability estimates · NIPS 2001
Machine learning › Kernel, tree and ensemble methods
decision tree
0.012001
Obtaining calibrated probability estimates from decision trees and naive Bayesian classifiers · ICML 2001
Machine learning › Learning theory › classification
multiclass classification
0.012001
Reducing multiclass to binary by coupling probability estimates · NIPS 2001
Machine learning › Trustworthy machine learning
uncertainty estimation
0.012001
Obtaining calibrated probability estimates from decision trees and naive Bayesian classifiers · ICML 2001
Data mining
anomaly detection
0.012007
High-quantile modeling for customer wallet estimation and other applications · KDD 2007
Data mining › anomaly detection
fraud detection
0.012007
High-quantile modeling for customer wallet estimation and other applications · KDD 2007

Methods — techniques the papers use, named apart from their topics

additive model · 0.4GPS data analysis · 0.4distributed word representation · 0.2deep neural network · 0.2naive bayes · 0.2reduction · 0.1spearman correlation · 0.1kendall correlation · 0.1regression tree · 0.1nearest neighbor · 0.1boosting · 0.1selective sampling · 0.1classification reduction · 0.1active learning · 0.1binary classification oracle · 0.1binary classification · 0.1ROC curves · 0.1ROC curve · 0.1
YearPublicationVenuePosition
2022 Controlling Weather Field Synthesis Using Variational Autoencoders
abstract
One of the consequences of climate change is an observed increase in the frequency of extreme climate events. That poses a challenge for weather forecast and generation algorithms, which learn from historical data but should embed an often uncertain bias to create correct scenarios. This paper investigates how mapping climate data to a known distribution using variational autoencoders might help explore such biases and control the synthesis of weather fields towards scenarios with more frequent extreme weather events. We experimented using a monsoon-affected precipitation dataset from southwest India, which should give a roughly stable pattern of rainy days and ease investigating the suitability of our solution. We report compelling results showing that mapping complex weather data to a known distribution implements an efficient control for weather field synthesis towards more (or less) extreme scenarios.
Dário A. B. Oliveira, Jorge Guevara Diaz, Bianca Zadrozny, Campbell D. Watson, Xiao Xiang Zhu 0001
IGARSS3
2021 A lazy feature selection method for multi-label classification
abstract
In many important application domains, such as text categorization, biomolecular analysis, scene or video classification and medical diagnosis, instances are naturally associated with more than one class label, giving rise to multi-label classification problems. This has led, in recent years, to a substantial amount of research in multi-label classification. More specifically, feature selection methods have been developed to allow the identification of relevant and informative features for multi-label classification. This work presents a new feature selection method based on the lazy feature selection paradigm and specific for the multi-label context. Experimental results show that the proposed technique is competitive when compared to multi-label feature selection techniques currently used in the literature, and is clearly more scalable, in a scenario where there is an increasing amount of data.
Rafael B. Pereira, Alexandre Plastino 0001, Bianca Zadrozny, Luiz H. C. Merschmann
Intell. Data Anal.3
2018 Efficient Classification of Seismic Textures
abstract
One of the most critical activities for the oil and gas industry is the discovery of new possibles reserves. Geoscientists must rely on indirect measures of the subsurface to scrutinize huge areas looking for leads of hydrocarbon reservoirs. Usually, to study the Earth's crust, geoscientists examine seismic images. Although deep learning has become popular in the last decade, only a few published results have demonstrated the application of such techniques to seismic images. In this paper, we present deep neural models specifically for the task of seismic facies analysis, using state-of-the-art concepts and tools to train and classify seismic facies efficiently. Our results show that we can train a neural network in 4 minutes using less than 5% of the dataset, and yet obtain 88% of accuracy. Moreover, we can reach up to 97% of accuracy in 30 minutes using 60% of the dataset.
Daniel Salles Chevitarese, Daniela Szwarcman, Emilio Vital Brazil, Bianca Zadrozny
IJCNN4
2018 Correlation analysis of performance measures for multi-label classification
Rafael B. Pereira, Alexandre Plastino 0001, Bianca Zadrozny, Luiz H. C. Merschmann
Inf. Process. Manag.3
2015 Detecting Semantically Equivalent Questions in Online User Forums
abstract
Two questions asking the same thing could be too different in terms of vocabulary and syntactic structure, which makes identifying their semantic equivalence challenging. This study aims to detect semantically equivalent questions in online user forums. We perform an extensive number of experiments using data from two different Stack Exchange forums. We compare standard machine learning methods such as Support Vector Machines (SVM) with a convolutional neural network (CNN). The proposed CNN generates distributed vector representations for pairs of questions and scores them using a similarity metric. We evaluate in-domain word embeddings versus the ones trained with Wikipedia, estimate the impact of the training set size, and evaluate some aspects of domain adaptation. Our experimental results show that the convolutional neural network with in-domain word embeddings achieves high performance even with limited training data.
Dasha Bogdanova, Cícero Nogueira dos Santos, Luciano Barbosa, Bianca Zadrozny
CoNLL4
2015 USapiens: A System for Urban Trajectory Data Analytics
abstract
In the past few years a growing number of cities have started monitoring the position of public transportation vehicles using GPS devices. In this paper, we focus on a particularly important urban dataset: GPS bus data. Buses are valuable sensors and information associated with buses can provide unprecedented insight into many different aspects of city's life, from human behavior to mobility patterns. But analyzing these large urban datasets presents many challenges. Urban datasets are complex, containing location and temporal components in addition that they are commonly released in their raw format. Furthermore, urban datasets may have noisy and missing data, locations gathered in a low sampling rate and not mapped to the underlying road network, among other issues which makes it difficult for citizens, administrators and developers to get insights. In this paper, we present a system, called USapiens, for analyzing large urban trajectory data. We first describe the architecture of the proposed system for pre-processing and analyzing urban trajectory data. We then detail five use cases we build using very large GPS dataset obtained from buses operating in the city of Rio de Janeiro to get insights into various aspects of public transportation in the city.
Marcos R. Vieira, Luciano Barbosa, Matthias Kormaksson, Bianca Zadrozny
MDM (1)4
2014 Bus Travel Time Predictions Using Additive Models
abstract
Many factors can affect the predictability of public bus services such as traffic, weather, day of week, and hour of day. However, the exact nature of such relationships between travel times and predictor variables is, in most situations, not known. In this paper we develop a framework that allows for flexible modeling of bus travel times through the use of Additive Models. The proposed class of models provides a principled statistical framework that is highly flexible in terms of model building. The experimental results demonstrate uniformly superior performance of our best model as compared to previous prediction methods when applied to a very large GPS data set obtained from buses operating in the city of Rio de Janeiro.
Matthias Kormaksson, Luciano Barbosa, Marcos R. Vieira, Bianca Zadrozny
ICDM4
2014 Learning Character-level Representations for Part-of-Speech Tagging
abstract
Distributed word representations have recently been proven to be an invaluable resource for NLP. These representations are normally learned using neural networks and capture syntactic and semantic information about words. Information about word morphology and shape is normally ignored when learning word representations. However, for tasks like part-of-speech tagging, intra-word information is extremely useful, specially when dealing with morphologically rich languages. In this paper, we propose a deep neural network that learns character-level representation of words and associate them with usual word representations to perform POS tagging. Using the proposed approach, while avoiding the use of any handcrafted feature, we produce state-of-the-art POS taggers for two languages: English, with 97.32% accuracy on the Penn Treebank WSJ corpus; and Portuguese, with 97.47% accuracy on the Mac-Morpho corpus, where the latter represents an error reduction of 12.2% on the best previous known result.
Cícero Nogueira dos Santos, Bianca Zadrozny
ICML2
2011 Lazy attribute selection: Choosing attributes at classification time
abstract
Attribute selection is a data preprocessing step which aims at identifying relevant attributes for the target machine learning task – namely classification in this paper. In this paper, we propose a new attribute selection strategy – based on a lazy
Rafael B. Pereira, Alexandre Plastino 0001, Bianca Zadrozny, Luiz H. C. Merschmann, Alex Alves Freitas
Intell. Data Anal.3
2011 Sensor data analysis for equipment monitoring
Ana Cristina Bicharra Garcia, Cristiana Bentes, Rafael Heitor C. de Melo, Bianca Zadrozny, Thadeu J. P. Penna
Knowl. Inf. Syst.4
2009 Tutorial summary: Reductions in machine learning
abstract
No abstract available.
Alina Beygelzimer, John Langford 0001, Bianca Zadrozny
ICML3
2008 Guest editorial: special issue on utility-based data mining
Gary Weiss 0001, Bianca Zadrozny, Maytal Saar-Tsechansky
Data Min. Knowl. Discov.2
2007 High-quantile modeling for customer wallet estimation and other applications
abstract
In this paper we discuss the important practical problem of customer wallet estimation, i.e., estimation of potential spending by customers(rather than their expected spending). For this purpose we utilize quantile modeling, whose goal is to estimate a quantile of the discriminative conditional distribution of the response, rather than the mean, which is the implicit goal of most standard regression approaches. We argue that a notion of wallet can be captured through high quantile modeling (e.g, estimating the 90th percentile), and describe a wallet estimation implementation within IBM's Market Alignment Program (MAP). We also discuss the wide range of domains where high-quantile modeling can be practically important: estimating opportunities in sales and marketing domains, defining 'surprising' patterns for outlier and fraud detection and more. We survey some existing approaches for quantile modeling, and propose adaptations of nearest-neighbor and regression-tree approaches to quantile modeling. We demonstrate the various models' performance in high quantile estimation in several domains, including our motivating problem of estimating the 'realistic' IT wallets of IBM customers.
Claudia Perlich, Saharon Rosset, Richard D. Lawrence, Bianca Zadrozny
KDD4
2007 Ranking-based evaluation of regression models
Saharon Rosset, Claudia Perlich, Bianca Zadrozny
Knowl. Inf. Syst.3
2006 Using secure coprocessors for privacy preserving collaborative data mining and analysis
abstract
Secure coprocessors have traditionally been used as a keystone of a security subsystem, eliminating the need to protect the rest of the subsystem with physical security measures. With technological advances and hardware miniaturization they have become increasingly powerful. This opens up the possibility of using them for non traditional use. This paper describes a solution for privacy preserving data sharing and mining using cryptographically secure but resource limited coprocessors. It uses memory light data mining methodologies along with a light weight database engine with federation capability, running on a coprocessor. The data to be shared resides with the enterprises that want to collaborate. This system will allow multiple enterprises, which are generally not allowed to share data, to do so solely for the purpose of detecting particular types of anomalies and for generating alerts. We also present results from experiments which demonstrate the value of such collaborations.
Bishwaranjan Bhattacharjee, Naoki Abe, Kenneth A. Goldman, Bianca Zadrozny, Vamsavardhana R. Chillakuru, Marysabel del Carpio, Chidanand Apté
DaMoN4
2006 Outlier detection by active learning
abstract
Most existing approaches to outlier detection are based on density estimation methods. There are two notable issues with these methods: one is the lack of explanation for outlier flagging decisions, and the other is the relatively high computational requirement. In this paper, we present a novel approach to outlier detection based on classification, in an attempt to address both of these issues. Our approach isbased on two key ideas. First, we present a simple reduction of outlier detection to classification, via a procedure that involves applying classification to a labeled data set containing artificially generated examples that play the role of potential outliers. Once the task has been reduced to classification, we then invoke a selective sampling mechanism based on active learning to the reduced classification problem. We empirically evaluate the proposed approach using a number of data sets, and find that our method is superior to other methods based on the same reduction to classification, but using standard classification methods. We also show that it is competitive to the state-of-the-art outlier detection methods in the literature based on density estimation, while significantly improving the computational complexity and explanatory power.
Naoki Abe, Bianca Zadrozny, John Langford 0001
KDD2
2006 Predicting Conditional Quantiles via Reduction to Classification
John Langford 0001, Roberto Oliveira 0001, Bianca Zadrozny
UAI3
2005 Weighted One-Against-All
Alina Beygelzimer, John Langford 0001, Bianca Zadrozny
AAAI3
2005 An Improved Categorization of Classifier's Sensitivity on Sample Selection Bias
abstract
A recent paper categorizes classifier learning algorithms according to their sensitivity to a common type of sample selection bias where the chance of an example being selected into the training sample depends on its feature vector x but not (directly) on its class label y. A classifier learner is categorized as "local" if it is insensitive to this type of sample selection bias, otherwise, it is considered "global". In that paper, the true model is not clearly distinguished from the model that the algorithm outputs. In their discussion of Bayesian classifiers, logistic regression and hard-margin SVMs, the true model (or the model that generates the true class label for every example) is implicitly assumed to be contained in the model space of the learner, and the true class probabilities and model estimated class probabilities are assumed to asymptotically converge as the training data set size increases. However, in the discussion of naive Bayes, decision trees and soft-margin SVMs, the model space is assumed not to contain the true model, and these three algorithms are instead argued to be "global learners". We argue that most classifier learners may or may not be affected by sample selection bias; this depends on the dataset as well as the heuristics or inductive bias implied by the learning algorithm and their appropriateness to the particular dataset.
Wei Fan 0001, Ian Davidson, Bianca Zadrozny, Philip S. Yu
ICDM3
2005 Ranking-Based Evaluation of Regression Models
abstract
We suggest the use of ranking-based evaluation measures for regression models, as a complement to the commonly used residual-based evaluation. We argue that in some cases, such as the case study we present, ranking can be the main underlying goal in building a regression model, and ranking performance is the correct evaluation metric. However, even when ranking is not the contextually correct performance metric, the measures we explore still have significant advantages: They are robust against extreme outliers in the evaluation set; and they are interpretable. The two measures we consider correspond closely to non-parametric correlation coefficients commonly used in data analysis (Spearman's p and Kendall's r); and they both have interesting graphical representations, which, similarly to ROC curves, offer useful "partial" model performance views, in addition to a one-number summary in the area under the curve. We illustrate our methods on a case study of evaluating IT wallet size estimation models for IBM's customers.
Saharon Rosset, Claudia Perlich, Bianca Zadrozny
ICDM3
2005 Error limiting reductions between classification tasks
abstract
We introduce a reduction-based model for analyzing supervised learning tasks. We use this model to devise a new reduction from multi-class cost-sensitive classification to binary classification with the following guarantee: If the learned binary classifier has error rate at most ε then the cost-sensitive classifier has cost at most 2ε times the expected sum of costs of all possible lables. Since cost-sensitive classification can embed any bounded loss finite choice supervised learning task, this result shows that any such task can be solved using a binary classification oracle. Finally, we present experimental results showing that our new reduction outperforms existing algorithms for multi-class cost-sensitive learning.
Alina Beygelzimer, Varsha Dani, Thomas P. Hayes, John Langford 0001, Bianca Zadrozny
ICML5
2005 Relating reinforcement learning performance to classification performance
abstract
We prove a quantitative connection between the expected sum of rewards of a policy and binary classification performance on created subproblems. This connection holds without any unobservable assumptions (no assumption of independence, small mixing time, fully observable states, or even hidden states) and the resulting statement is independent of the number of states or actions. The statement is critically dependent on the size of the rewards and prediction performance of the created classifiers.We also provide some general guidelines for obtaining good classification performance on the created subproblems. In particular, we discuss possible methods for generating training examples for a classifier learning algorithm.
John Langford 0001, Bianca Zadrozny
ICML2
2004 Learning and evaluating classifiers under sample selection bias
abstract
Classifier learning methods commonly assume that the training data consist of randomly drawn examples from the same distribution as the test examples about which the learned model is expected to make predictions. In many practical situations, however, this assumption is violated, in a problem known in econometrics as sample selection bias. In this paper, we formalize the sample selection bias problem in machine learning terms and study analytically and experimentally how a number of well-known classifier learning methods are affected by it. We also present a bias correction method that is particularly useful for classifier evaluation under sample selection bias.
Bianca Zadrozny
ICML1
2004 An iterative method for multi-class cost-sensitive learning
abstract
Cost-sensitive learning addresses the issue of classification in the presence of varying costs associated with different types of misclassification. In this paper, we present a method for solving multi-class cost-sensitive learning problems using any binary classification algorithm. This algorithm is derived using hree key ideas: 1) iterative weighting; 2) expanding data space; and 3) gradient boosting with stochastic ensembles. We establish some theoretical guarantees concerning the performance of this method. In particular, we show that a certain variant possesses the boosting property, given a form of weak learning assumption on the component binary classifier. We also empirically evaluate the performance of the proposed method using benchmark data sets and verify that our method generally achieves better results than representative methods for cost-sensitive learning, in terms of predictive performance (cost minimization) and, in many cases, computational efficiency.
Naoki Abe, Bianca Zadrozny, John Langford 0001
KDD2
2003 Cost-Sensitive Learning by Cost-Proportionate Example Weighting
abstract
We propose and evaluate a family of methods for converting classifier learning algorithms and classification theory into cost-sensitive algorithms and theory. The proposed conversion is based on cost-proportionate weighting of the training examples, which can be realized either by feeding the weights to the classification algorithm (as often done in boosting), or by careful subsampling. We give some theoretical performance guarantees on the proposed methods, as well as empirical evidence that they are practical alternatives to existing approaches. In particular, we propose costing, a method based on cost-proportionate rejection sampling and ensemble aggregation, which achieves excellent predictive performance on two publicly available datasets, while drastically reducing the computation required by other methods.
Bianca Zadrozny, John Langford 0001, Naoki Abe
ICDM1
2002 Empirical Comparison of Various Reinforcement Learning Strategies for Sequential Targeted Marketing
abstract
We empirically evaluate the performance of various reinforcement learning methods in applications to sequential targeted marketing. In particular we propose and evaluate a progression of reinforcement learning methods, ranging from the "direct" or "batch" methods to "indirect" or "simulation based" methods, and those that we call "semidirect" methods that fall between them. We conduct a number of controlled experiments to evaluate the performance of these competing methods. Our results indicate that while the indirect methods can perform better in a situation in which nearly perfect modeling is possible, under the more realistic situations in which the system's modeling parameters have restricted attention, the indirect methods' performance tend to degrade. We also show that semi-direct methods are effective in reducing the amount of computation necessary to attain a given level of performance, and often result in more profitable policies.
Naoki Abe, Edwin P. D. Pednault, Haixun Wang, Bianca Zadrozny, Wei Fan 0001, Chidanand Apté
ICDM4
2002 Sequential cost-sensitive decision making with reinforcement learning
abstract
Recently, there has been increasing interest in the issues of cost-sensitive learning and decision making in a variety of applications of data mining. A number of approaches have been developed that are effective at optimizing cost-sensitive decisions when each decision is considered in isolation. However, the issue of sequential decision making, with the goal of maximizing total benefits accrued over a period of time instead of immediate benefits, has rarely been addressed. In the present paper, we propose a novel approach to sequential decision making based on the reinforcement learning framework. Our approach attempts to learn decision rules that optimize a sequence of cost-sensitive decisions so as to maximize the total benefits accrued over time. We use the domain of targeted' marketing as a testbed for empirical evaluation of the proposed method. We conducted experiments using approximately two years of monthly promotion data derived from the well-known KDD Cup 1998 donation data set. The experimental results show that the proposed method for optimizing total accrued benefits out performs the usual targeted-marketing methodology of optimizing each promotion in isolation. We also analyze the behavior of the targeting rules that were obtained and discuss their appropriateness to the application domain.
Edwin P. D. Pednault, Naoki Abe, Bianca Zadrozny
KDD3
2002 Transforming classifier scores into accurate multiclass probability estimates
abstract
Class membership probability estimates are important for many applications of data mining in which classification outputs are combined with other sources of information for decision-making, such as example-dependent misclassification costs, the outputs of other classifiers, or domain knowledge. Previous calibration methods apply only to two-class problems. Here, we show how to obtain accurate probability estimates for multiclass problems by combining calibrated binary probability estimates. We also propose a new method for obtaining calibrated two-class probability estimates that can be applied to any classifier that produces a ranking of examples. Using naive Bayes and support vector machine classifiers, we give experimental results from a variety of two-class and multiclass domains, including direct marketing, text categorization and digit recognition.
Bianca Zadrozny, Charles Elkan
KDD1
2001 Obtaining calibrated probability estimates from decision trees and naive Bayesian classifiers
Bianca Zadrozny, Charles Elkan
ICML1
2001 Learning and making decisions when costs and probabilities are both unknown
abstract
In many data mining domains, misclassification costs are different for different examples, in the same way that class membership probabilities are example-dependent. In these domains, both costs and probabilities are unknown for test examples, so both cost estimators and probability estimators must be learned. After discussing how to make optimal decisions given cost and probability estimates, we present decision tree and naive Bayesian learning methods for obtaining well-calibrated probability estimates. We then explain how to obtain unbiased estimators for example-dependent costs, taking into account the difficulty that in general, probabilities and costs are not independent random variables, and the training examples for which costs are known are not representative of all examples. The latter problem is called sample selection bias in econometrics. Our solution to it is based on Nobel prize-winning work due to the economist James Heckman. We show that the methods we propose perform better than MetaCost and all other known methods, in a comprehensive experimental comparison that uses the well-known, large, and challenging dataset from the KDD'98 data mining contest.
Bianca Zadrozny, Charles Elkan
KDD1
2001 Reducing multiclass to binary by coupling probability estimates
abstract
This paper presents a method for obtaining class membership probability esti- mates for multiclass classification problems by coupling the probability estimates produced by binary classifiers. This is an extension for arbitrary code matrices of a method due to Hastie and Tibshirani for pairwise coupling of probability estimates. Experimental results with Boosted Naive Bayes show that our method produces calibrated class membership probability estimates, while having similar classification accuracy as loss-based decoding, a method for obtaining the most likely class that does not generate probability estimates.
Bianca Zadrozny
NIPS1