Sahand Hariri

dblp:220/9162 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2021
0000-0003-0860-3300ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
anomaly detection
0.512021
Extended Isolation Forest · IEEE Trans. Knowl. Data Eng. 2021
Data mining › anomaly detection › unsupervised anomaly detection
isolation forest
0.512021
Extended Isolation Forest · IEEE Trans. Knowl. Data Eng. 2021
Data mining › anomaly detection › statistical anomaly detection
non-parametric anomaly detection
0.112021
Extended Isolation Forest · IEEE Trans. Knowl. Data Eng. 2021

Methods — techniques the papers use, named apart from their topics

random hyperplane slicing · 0.5ensemble · 0.5
YearPublicationVenuePosition
2021 Extended Isolation Forest
abstract
We present an extension to the model-free anomaly detection algorithm, Isolation Forest. This extension, named Extended Isolation Forest (EIF), resolves issues with assignment of anomaly score to given data points. We motivate the problem using heat maps for anomaly scores. These maps suffer from artifacts generated by the criteria for branching operation of the binary tree. We explain this problem in detail and demonstrate the mechanism by which it occurs visually. We then propose two different approaches for improving the situation. First we propose transforming the data randomly before creation of each tree, which results in averaging out the bias. Second, which is the preferred way, is to allow the slicing of the data to use hyperplanes with random slopes. This approach results in remedying the artifact seen in the anomaly score heat maps. We show that the robustness of the algorithm is much improved using this method by looking at the variance of scores of data points distributed along constant level sets. We report AUROC and AUPRC for our synthetic datasets, along with real-world benchmark datasets. We find no appreciable difference in the rate of convergence nor in computation time between the standard Isolation Forest and EIF.
Sahand Hariri, Matias Carrasco Kind, Robert J. Brunner
IEEE Trans. Knowl. Data Eng.1