Aniza Mohamed Din

dblp:55/541 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0002-5859-023XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 71% Database theory · 29%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Database theory › query complexity
data complexity
1.422024
An Investigation of SMOTE Based Methods for Imbalanced Datasets with Data Complexity Analysis (Extended Abstract) · ICDE 2024
An Investigation of SMOTE Based Methods for Imbalanced Datasets With Data Complexity Analysis · IEEE Trans. Knowl. Data Eng. 2023
Data mining › predictive modeling › classification › imbalanced classification
oversampling
1.022024
An Investigation of SMOTE Based Methods for Imbalanced Datasets with Data Complexity Analysis (Extended Abstract) · ICDE 2024
An Investigation of SMOTE Based Methods for Imbalanced Datasets With Data Complexity Analysis · IEEE Trans. Knowl. Data Eng. 2023
Data mining › predictive modeling
classification
0.922024
An Investigation of SMOTE Based Methods for Imbalanced Datasets With Data Complexity Analysis · IEEE Trans. Knowl. Data Eng. 2023
An Investigation of SMOTE Based Methods for Imbalanced Datasets with Data Complexity Analysis (Extended Abstract) · ICDE 2024
Data mining › predictive modeling › classification
imbalanced classification
0.922024
An Investigation of SMOTE Based Methods for Imbalanced Datasets With Data Complexity Analysis · IEEE Trans. Knowl. Data Eng. 2023
An Investigation of SMOTE Based Methods for Imbalanced Datasets with Data Complexity Analysis (Extended Abstract) · ICDE 2024
Data mining › predictive modeling › classification
class imbalance
0.812024
An Investigation of SMOTE Based Methods for Imbalanced Datasets with Data Complexity Analysis (Extended Abstract) · ICDE 2024

Methods — techniques the papers use, named apart from their topics

SMOTE · 1.4f1-score · 0.8synthetic minority oversampling · 0.7
YearPublicationVenuePosition
2024 An Investigation of SMOTE Based Methods for Imbalanced Datasets with Data Complexity Analysis (Extended Abstract)
abstract
This extended abstract highlights challenges with imbalanced datasets in real-world applications, where issues like noise, class overlap, and small subsets of data impact classification accuracy. While the Synthetic Minority Oversampling Technique (SMOTE) addresses imbalanced datasets by increasing minority class examples, it struggles with handling these data complexities and might worsen the situation. As a result, several SMOTE variants have emerged, aiming to improve its effectiveness by integrating it with other methods or altering its approach. This paper offers a comparative analysis of these variants, examining how each tackles specific data complexities. Through experiments on 24 imbalanced datasets, changes in complexity measures resulting from these SMOTE variants, in terms of F1-Score and data complexity metrics are observed and demonstrated.
Nur Athirah Azhar, Muhammad Syafiq Mohd Pozi, Aniza Mohamed Din, Adam Jatowt
ICDE3
2023 An Investigation of SMOTE Based Methods for Imbalanced Datasets With Data Complexity Analysis
abstract
Many binary class datasets in real-life applications are affected by class imbalance problem. Data complexities like noise examples, class overlap and small disjuncts problems are observed to play a key role in producing poor classification performance. These complexities tend to exist in tandem with class imbalance problem. Synthetic Minority Oversampling Technique (SMOTE) is a well-known method to re-balance the number of examples in imbalanced datasets. However, this technique cannot effectively tackle data complexities and it also has the capability of magnifying the degree of complexities. Also, the performance of the SMOTE is still not satisfactory. Therefore, various SMOTE variants have been proposed to overcome the downsides of SMOTE either by combining SMOTE with other algorithms or modifying the existing SMOTE algorithm. This paper aims to comparatively review the algorithms applied in SMOTE variants and investigate which data complexities are being addressed in what variants. Series of experiments are conducted on 24 binary class imbalanced datasets to observe the changes in the data complexity measures after SMOTE variants were applied in these datasets. The evaluation metrics like G-Mean and F1-Score are also analyzed to investigate the difference in classification performance between SMOTE variants.
Nur Athirah Azhar, Muhammad Syafiq Mohd Pozi, Aniza Mohamed Din, Adam Jatowt
IEEE Trans. Knowl. Data Eng.3