EDBT 2026 Demo / reviewers in the wild / expert
Mahesh V. Joshi
dblp:66/1618
· DBLP profile ↗
10ranked-venue papers
7as first author
0since 2021 · last 2004
0009-0005-6324-6635ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 6 first-authorArtificial intelligence and machine learning · 5 · 4 first-authorSystems, architecture and hardware · 1 · 1 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Data mining · 92% Information retrieval · 8% | |
| Artificial intelligence
1 paper |
Kernel, tree and ensemble methods · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › predictive modeling
classification |
0.1 | 4 | 2002 | Predicting rare classes: can boosting make any weak learner strong? · KDD 2002 On Evaluating Performance of Classifiers for Rare Classes · ICDM 2002 Mining Needle in a Haystack: Classifying Rare Classes via Two-phase Rule Induction · SIGMOD Conference 2001 |
Data mining › predictive modeling › classification › imbalanced classification
rare class classification |
0.1 | 4 | 2002 | Predicting rare classes: can boosting make any weak learner strong? · KDD 2002 On Evaluating Performance of Classifiers for Rare Classes · ICDM 2002 Mining Needle in a Haystack: Classifying Rare Classes via Two-phase Rule Induction · SIGMOD Conference 2001 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting |
0.0 | 1 | 2002 | Predicting rare classes: can boosting make any weak learner strong? · KDD 2002 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.0 | 1 | 2002 | Predicting rare classes: can boosting make any weak learner strong? · KDD 2002 |
Data mining › predictive modeling › classification
classifier evaluation |
0.0 | 1 | 2002 | On Evaluating Performance of Classifiers for Rare Classes · ICDM 2002 |
Information retrieval › retrieval evaluation
precision/recall metrics |
0.0 | 1 | 2002 | On Evaluating Performance of Classifiers for Rare Classes · ICDM 2002 |
Data mining › predictive modeling › classification › ensemble learning
boosting |
0.0 | 1 | 2001 | Evaluating Boosting Algorithms to Classify Rare Classes: Comparison and Improvements · ICDM 2001 |
Data mining › predictive modeling › classification
ensemble learning |
0.0 | 1 | 2001 | Evaluating Boosting Algorithms to Classify Rare Classes: Comparison and Improvements · ICDM 2001 |
Data mining › predictive modeling › classification
rule induction |
0.0 | 1 | 2001 | Mining Needle in a Haystack: Classifying Rare Classes via Two-phase Rule Induction · SIGMOD Conference 2001 |
Network security › intrusion detection and prevention
intrusion detection |
0.0 | 1 | 2001 | Mining Needle in a Haystack: Classifying Rare Classes via Two-phase Rule Induction · SIGMOD Conference 2001 |
Methods — techniques the papers use, named apart from their topics
boosting · 0.1adacost · 0.1RIPPER · 0.1two-phase rule induction · 0.1sequential covering · 0.1point-metric analysis · 0.0weight updating · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2004 | CREDOS: Classification Using Ripple Down Structure (A Case for Rare Classes)abstractRipple down rules (RDRs) are commonly used by the expert systems community because they make knowledge bases easy to use and efficient to maintain. We observe that RDRs offer a unique tree-based representation that generalizes the decision tree and disjunctive normal form (DNF) rule-based models, and specializes a generic form of the PNrule model. In this paper, we explore their use for learning predictive classifier models. Such models require to have a generalization capability, most commonly achieved with the help of pruning methods. Existing RDR induction algorithms are developed to build an initial knowledge base that will be used and modified by humans to explain every case correctly. They do not look at RDR as a predictive model, and hence offer very little measures against over-fitting. Existing pruning strategies developed by the data mining community cannot be directly used for pruning a RDR structure because of the uniqueness of the structure and the prediction process. In this paper, we propose a novel induction algorithm CREDOS. The key characteristic of CREDOS is its generic pruning framework. We provide a specific instantiation of it based on the minimum description length (MDL) principle. Using real-world datasets requiring prediction of rare classes, we compare CREDOS to other state-of-the-art algorithms. It exhibits significantly better or comparable performance, especially in predicting a wide variety of rarely occurring events. Mahesh V. Joshi, Vipin Kumar 0001 |
SDM | 1 |
| 2003 | Topic Learning from Few Examples
Huaiyu Zhu 0001, Shivakumar Vaithyanathan, Mahesh V. Joshi |
PKDD | 3 |
| 2002 | On Evaluating Performance of Classifiers for Rare ClassesabstractPredicting rare classes effectively is an important problem. The definition of effective classifier, embodied in the classifier evaluation metric, is however very subjective, dependent on the application domain. In this paper a wide variety of point-metrics are put into a common analytical context defined by the recall and precision of the target rare class. This enables us to compare various metrics in an objective, domain-independent manner. We judge their suitability for the rare class problems along the dimensions of learning difficulty and levels of rarity. This yields many valuable insights. In order to address the goal of achieving better recall and precision, we also propose a way of comparing classifiers directly based on the relationships between recall and precision values. It resorts to a composite point-metric only when recall-precision based comparisons yield conflicting results. Mahesh V. Joshi |
ICDM | 1 |
| 2002 | Predicting rare classes: can boosting make any weak learner strong?abstractBoosting is a strong ensemble-based learning algorithm with the promise of iteratively improving the classification accuracy using any base learner, as long as it satisfies the condition of yielding weighted accuracy > 0.5. In this paper, we analyze boosting with respect to this basic condition on the base learner, to see if boosting ensures prediction of rarely occurring events with high recall and precision. First we show that a base learner can satisfy the required condition even for poor recall or precision levels, especially for very rare classes. Furthermore, we show that the intelligent weight updating mechanism in boosting, even in its strong cost-sensitive form, does not prevent cases where the base learner always achieves high precision but poor recall or high recall but poor precision, when mapped to the original distribution. In either of these cases, we show that the voting mechanism of boosting falls to achieve good overall recall and precision for the ensemble. In effect, our analysis indicates that one cannot be blind to the base learner performance, and just rely on the boosting mechanism to take care of its weakness. We validate our arguments empirically on variety of real and synthetic rare class problems. In particular, using AdaCost as the boosting algorithm, and variations of PNrule and RIPPER as the base learners, we show that if algorithm A achieves better recall-precision balance than algorithm B, then using A as the base learner in AdaCost yields significantly better performance than using B as the base learner. Mahesh V. Joshi, Ramesh C. Agarwal, Vipin Kumar 0001 |
KDD | 1 |
| 2002 | Predicting Rare Classes: Comparing Two-Phase Rule Induction to Cost-Sensitive Boosting
Mahesh V. Joshi, Ramesh C. Agarwal, Vipin Kumar 0001 |
PKDD | 1 |
| 2001 | Evaluating Boosting Algorithms to Classify Rare Classes: Comparison and ImprovementsabstractClassification of rare events has many important data mining applications. Boosting is a promising meta-technique that improves the classification performance of any weak classifier. So far, no systematic study has been conducted to evaluate how boosting performs for the task of mining rare classes. The authors evaluate three existing categories of boosting algorithms from the single viewpoint of how they update the example weights in each iteration, and discuss their possible effect on recall and precision of the rare class. We propose enhanced algorithms in two of the categories, and justify their choice of weight updating parameters theoretically. Using some specially designed synthetic datasets, we compare the capability of all the algorithms from the rare class perspective. The results support our qualitative analysis, and also indicate that our enhancements bring an extra capability for achieving better balance between recall and precision in mining rare classes. Mahesh V. Joshi, Vipin Kumar 0001, Ramesh C. Agarwal |
ICDM | 1 |
| 2001 | PNrule: A New Framework for Learning Classifier Models in Data Mining (A Case-Study in Network Intrusion Detection)abstract1 Introduction and Motivation Learning classifier models is an important problem in data mining. Observations from the real world are often recorded as a set of records, each characterized by multiple attributes. Associated with each record is a categorical attribute called class. Given a training set of records with known class labels, the problem is to learn a model for the class in terms of other attributes. The goal is to use this model to predict the class of any given set of records, such that certain objective function based on the predicted and actual classes is optimized. Traditionally, the goal has been to minimize the number of misclassified records; i.e. to maximize accuracy. Various techniques exist today to build classifier models[11]. Although no single technique is proven to be the best in all situations, techniques that learn rule-based models are especially popular in the domain of data mining. This can be contributed to the easy interpretability of the rules by humans, and competitive performance exhibited by rule-based models in many application domains. Ramesh C. Agarwal, Mahesh V. Joshi |
SDM | 2 |
| 2001 | Mining Needle in a Haystack: Classifying Rare Classes via Two-phase Rule InductionabstractLearning models to classify rarely occurring target classes is an important problem with applications in network intrusion detection, fraud detection, or deviation detection in general. In this paper, we analyze our previously proposed two-phase rule induction method in the context of learning complete and precise signatures of rare classes. The key feature of our method is that it separately conquers the objectives of achieving high recall and high precision for the given target class. The first phase of the method aims for high recall by inducing rules with high support and a reasonable level of accuracy. The second phase then tries to improve the precision by learning rules to remove false positives in the collection of the records covered by the first phase rules. Existing sequential covering techniques try to achieve high precision for each individual disjunct learned. In this paper, we claim that such approach is inadequate for rare classes, because of two problems: splintered false positives and error-prone small disjuncts. Motivated by the strengths of our two-phase design, we design various synthetic data models to identify and analyze the situations in which two state-of-the-art methods, RIPPER and C4.5 rules, either fail to learn a model or learn a very poor model. In all these situations, our two-phase approach learns a model with significantly better recall and precision levels. We also present a comparison of the three methods on a challenging real-life network intrusion detection dataset. Our method is significantly better or comparable to the best competitor in terms of achieving better balance between recall and precision. Mahesh V. Joshi, Ramesh C. Agarwal, Vipin Kumar 0001 |
SIGMOD Conference | 1 |
| 1998 | The Design, Implementation, and Evaluation of a Symmetric Banded Linear Solver for Distributed-Memory Parallel ComputersabstractThis article describes the design, implementation, and evaluation of a parallel algorithm for the Cholesky factorization of symmetric banded matrices. The algorithm is part of IBM's parallel engineering and scientific subroutine library version 1.2 and is compatible with ScaLAPACK's banded solver. Analysis, as well as experiments on an IBM SP2 distributed-memory parallel computer, shows that the algorithm efficiently factors banded matrices with wide bandwidth. For example, a 31-mode SP2 factors a large matrix more than 16 times faster than a single node would factor it using the best sequential algorithm, and more than 20 times faster than a single node would using LAPACK's DPBTRF. The algorithm uses novel ideas in the area of distributed dense-matrix computations that include the use of a dynamic schedule for a blocked systolic-like algorithm and the separation of the input and output layouts from the layout the algorithm uses internally. The algorithm alson uses known techniques such as blocking to improve its communication-to-computation ratio and its data-cache behavior. Fred G. Gustavson, Mahesh V. Joshi, Sivan Toledo |
ACM Trans. Math. Softw. | 3 |
| 1997 | A high performance two dimensional scalable parallel algorithm for solving sparse triangular systemsabstractSolving a system of equations of the form Tx=y, where T is a sparse triangular matrix, is required after the factorization phase in the direct methods of solving systems of linear equations. A few parallel formulations have been proposed recently. The common belief in parallelizing this problem is that the parallel formulation utilizing a two dimensional distribution of T is unscalable. We propose the first known efficient scalable parallel algorithm which uses a two dimensional block cyclic distribution of T. The algorithm is shown to be applicable to dense as well as sparse triangular solvers. Since most of the known highly scalable algorithms employed in the factorization phase yield a two dimensional distribution of T, our algorithm avoids the redistribution cost incurred by the one dimensional algorithms. We present the parallel runtime and scalability analyses of the proposed two dimensional algorithm. The dense triangular solver is shown to be scalable. The sparse triangular solver is shown to be at least as scalable as the dense solver. We also show that it is optimal for one class of sparse systems. The experimental results of the sparse triangular solver show that it has good speedup characteristics and yields high performance for a variety of sparse systems. Mahesh V. Joshi, George Karypis, Vipin Kumar 0001 |
HiPC | 1 |