Kushankur Ghosh

dblp:256/8386 · DBLP profile ↗
← Back
2ranked-venue papers in the field
2as first author
2since 2021 · last 2024
0000-0002-4761-120XORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2 (2 first)
YearPublicationVenuePosition
2024 Unsupervised Parameter-free Outlier Detection using HDBSCAN* Outlier Profiles
abstract
In machine learning and data mining, outliers are data points that significantly differ from the dataset and often introduce irrelevant information that can induce bias in its statistics and models. Therefore, unsupervised methods are crucial to detect outliers if there is limited or no information about them. Global-Local Outlier Scores based on Hierarchies (GLOSH) is an unsupervised outlier detection method within HDBSCAN*, a state-of-the-art hierarchical clustering method. GLOSH estimates outlier scores for each data point by comparing its density to the highest density of the region they reside in the HDBSCAN* hierarchy. GLOSH may be sensitive to HDBSCAN*’s minptsparameter that influences density estimation. With limited knowledge about the data, choosing an appropriate minptsvalue beforehand is challenging as one or some minptsvalues may better represent the underlying cluster structure than others. Additionally, in the process of searching for "potential outliers", one has to define the number of outliers n a dataset has, which may be impractical and is often unknown. In this paper, we propose an unsupervised strategy to find the "best" minptsvalue, leveraging the range of GLOSH scores across minptsvalues to identify the value for which GLOSH scores can best identify outliers from the rest of the dataset. Moreover, we propose an unsupervised strategy to estimate a threshold for classifying points into inliers and (potential) outliers without the need to pre-define any value. Our experiments show that our strategies can automatically find the minptsvalue and threshold that yield the best or near best outlier detection results using GLOSH.
Kushankur Ghosh, Murilo Coelho Naldi, Jörg Sander 0001, Euijin Choo
IEEE Big Data1
2021 On the combined effect of class imbalance and concept complexity in deep learning
abstract
Structural concept complexity, class overlap, and data scarcity are some of the most important factors influencing the performance of classifiers under class imbalance conditions. When these effects were uncovered in the early 2000s, understandably, the classifiers on which they were demonstrated belonged to the classical rather than Deep Learning categories of approaches. As Deep Learning is gaining ground over classical machine learning and is beginning to be used in critical applied settings, it is important to assess systematically how well they respond to the kind of challenges their classical counterparts have struggled with in the past two decades. The purpose of this paper is to study the behavior of deep learning systems in settings that have previously been deemed challenging to classical machine learning systems to find out whether the depth of the systems is an asset in such settings. The results in both artificial and real-world image datasets show that these settings remain mostly challenging for Deep Learning systems. Deeper architectures help with structural concept complexity but not with data scarcity and class overlap.
Kushankur Ghosh, Colin Bellinger, Roberto Corizzo, Bartosz Krawczyk, Nathalie Japkowicz
IEEE BigData1