Tanzira Najnin

dblp:384/0000 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0005-3647-4253ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › predictive modeling
classification
0.812024
A Novel Feature Space Augmentation Method to Improve Classification Performance and Evaluation Reliability · KDD 2024
Data mining › predictive modeling › classification
class imbalance
0.812024
A Novel Feature Space Augmentation Method to Improve Classification Performance and Evaluation Reliability · KDD 2024
Data mining › feature engineering
feature-space augmentation
0.812024
A Novel Feature Space Augmentation Method to Improve Classification Performance and Evaluation Reliability · KDD 2024

Methods — techniques the papers use, named apart from their topics

uniform random sampling · 0.8synthetic instance generation · 0.8
YearPublicationVenuePosition
2026 An Integrated Approach to Knowledge and Prediction Modeling of Breast Cancer Metastasis Using Gene Regulatory Networks
abstract
Understanding why only some breast cancers become metastatic, and predicting metastatic risks at earlier cancer stages, are two major goals of breast cancer research. These goals are clearly connected and synergistic, but are rarely integrated within a single study. Knowledge discovery-oriented research has identified critical biological pathways related to metastasis, but the interplay between these pathways remains elusive, resulting in models that lack prediction accuracy. Conversely, complex machine learning models achieve high prediction accuracy without detailed biological knowledge, making the explanation challenging. Here, we propose a novel computational framework that addresses both goals simultaneously. To support knowledge discovery, we construct gene regulatory networks (GRNs) to model the cellular states of metastatic and non-metastatic patients. To enable explainable metastasis prediction, we introduce a dysregulation score based on the GRN models. Experimental results demonstrate that our method not only identified significant metastasis-associated GRN changes, but also revealed the loss of co-regulation among key biological processes in metastatic patients. Leveraging the dysregulation score, our model-free classifier outperformed complex machine learning models under rigorous evaluation. This work bridges a significant gap between knowledge discovery and accurate, explainable prediction by employing carefully designed knowledge models and knowledge-based prediction, with potential applicability in other disease contexts.
Tanzira Najnin, Sakhawat Hossain Saimon, Maryam Zand, Nahim Adnan, Tim Hui-Ming Huang, Jianhua Ruan
IEEE Trans. Comput. Biol. Bioinform.1
2024 A Novel Feature Space Augmentation Method to Improve Classification Performance and Evaluation Reliability
abstract
Classification tasks in many real-world domains are exacerbated by class imbalance, relatively small sample sizes compared to high dimensionality, and measurement uncertainty. The problem of class imbalance has been extensively studied, and data augmentation methods based on interpolation of minority class instances have been proposed as a viable solution to mitigate imbalance. It remains to be seen whether augmentation can be applied to improve the overall performance while maintaining stability, especially with a limited number of samples. In this paper, we present a novel feature-space augmentation technique that can be applied to high-dimensional data for classification tasks and address these issues. Our method utilizes uniform random sampling and introduces synthetic instances by taking advantage of the local distributions of individual features in the observed instances. The core augmentation algorithm is class-invariant, which opens up an unexplored avenue of simultaneously improving and stabilizing performance by augmenting unlabeled instances. The proposed method is evaluated using a comprehensive performance analysis involving multiple classifiers and metrics. Comparative analysis with existing feature space augmentation methods strongly suggests that the proposed algorithm can result in improved classification performance while also increasing the overall reliability of the performance evaluation.
Sakhawat Hossain Saimon, Tanzira Najnin, Jianhua Ruan
KDD2