VLDB 2026 Research / reviewers in the wild / expert
Haydar Demirhan
dblp:133/6157
· DBLP profile ↗
8ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-8565-4710ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A clustering framework for skewed features with low true cluster separationabstractAbstract Features with considerably larger or smaller observations than the rest of the dataset, causing noticeable skewness in the feature distributions, are prevalent in practical applications. Traditional clustering methods often assume symmetric data, leading to poor performance with skewed features. This challenge becomes further complicated in datasets with low separation between true clusters in the feature space. These problems are encountered in a wide range of important practical areas, such as cell grouping, forest fires, maritime search and rescue, urbanization studies, and neuroimaging. Bayesian model-based clustering methods can accurately capture the skewness in the data and centers of poorly separated true clusters. However, they are computationally inefficient due to their Bayesian nature. We propose a Bayesian model-based clustering framework to address these issues by utilizing the generalized multivariate log-gamma distribution with a Dirichlet process mixture. Comparative numerical experiments on 30 benchmark datasets with traditional and Bayesian model-based clustering algorithms demonstrate the superior performance of the proposed method, particularly for skewed datasets with low true cluster separability. The proposed approach, implemented in R, also shows better computational efficiency than its Bayesian alternatives. The computer codes to implement our approach are provided to facilitate practical applications. Muntazir Mehdi, Haydar Demirhan, Sona Taheri |
Data Min. Knowl. Discov. | 2 |
| 2025 | Mixed fuzzy C-means clusteringabstractClustering analysis becomes challenging when the dataset has mixed data types comprising categorical (nominal or ordinal scale) and numerical (interval scale) features. Mainstream distance metrics cannot handle the information in categorical data about the similarity between the observations and cluster centers, leading to performance loss. Various methods are introduced in the literature to handle the mixed data types in clustering. However, each method has disadvantages in capturing categorical information about similarity, adjusting the contribution of categorical information to clustering, and computational or implementation inefficiency. This study proposes a mixed fuzzy C-means clustering method for mixed data types. Two new distance metrics are developed to handle binary and multi-class nominal features. The scaled entropy of each data type is used to adjust the weight of each data type in the overall similarity metric, providing a lower bias since no user-specified weight is required. A comparative numerical study is conducted with twenty real datasets and seven benchmark methods using five cluster validation statistics. The mixed fuzzy C-means clustering performs better than the benchmark methods and is computationally efficient in practice. Since all the computer codes for implementing mixed fuzzy C-means clustering are given, the proposed method is readily applicable to practical problems. Haydar Demirhan |
Inf. Sci. | 1 |
| 2024 | An online fuzzy fraud detection framework for credit card transactionsabstractCredit card transaction fraud is one of the challenging security concerns for financial firms globally. The continuously changing environment of fraud characteristics and the class imbalance and complete separation issues in fraud data create difficulties in accurately and efficiently predicting fraudulent transactions and implementing fraud detection systems in real-time. The study aims to develop a new real-time fraud detection framework that can efficiently be implemented online and address the issues caused by non-stationary changes in transaction and fraud characteristics, class imbalance, and complete separation. We propose a new approach to handle the impact of non-stationary changes in fraud transaction patterns. It enables efficiency in model training, given the sheer size of datasets. By implementing a robust fuzzy logistic regression model against class imbalance and separation problems, we address the challenge of having a very low rate of fraudulent transactions in the dataset and having separation issues due to specific characteristics of transactions. The analysis of the performance versus efficiency nexus of the proposed methodology reveals that the proposed framework shows strong performance results with specificity and sensitivity greater than 0.90 and Matthew’s correlation coefficient greater than 0.80 even on small sample sizes and produces highly accurate results in identifying fraudulent and non-fraudulent transactions with an accuracy greater than 0.99. Benchmarking with machine learning and other fraud detection approaches reveals that the proposed framework provides better detection performance while maintaining a higher rate of identifying non-fraudulent transactions than the alternative approaches. Improved classification performance leads to detecting fraudulent transactions with higher precision while avoiding misclassifying legitimate transfers, resulting in lower financial losses and improved customer satisfaction. Georgios Charizanos, Haydar Demirhan, Duygu Içen |
Expert Syst. Appl. | 2 |
| 2024 | A Monte Carlo fuzzy logistic regression framework against imbalance and separationabstractThis article proposes a new fuzzy logistic regression framework with high classification performance against imbalance and separation while keeping the interpretability of classical logistic regression. Separation and imbalance are two core problems in logistic regression, which can result in biased coefficient estimates and inaccurate predictions. Existing research on fuzzy logistic regression primarily focuses on developing possibilistic models instead of using a logit link function that converts log-odds ratios to probabilities. At the same time, little consideration is given to issues of separation and imbalance. Our study aims to address these challenges by proposing new methods of fuzzifying binary variables and classifying subjects based on a comparison against a fuzzy threshold. We use combinations of fuzzy and crisp predictors, output, and coefficients to understand which combinations perform better under imbalance and separation. Numerical experiments with synthetic and real datasets are conducted to demonstrate the usefulness and superiority of the proposed framework. Seven crisp machine learning models are implemented for benchmarking in the numerical experiments. The proposed framework shows consistently strong performance results across datasets with imbalance or separation and performs equally well when such issues are absent. Meanwhile, the considered machine learning methods are significantly impacted by the imbalanced datasets. Georgios Charizanos, Haydar Demirhan, Duygu Içen |
Inf. Sci. | 2 |
| 2024 | Hierarchical fuzzy regression functions for mixed predictors and an application to real estate price predictionabstractAbstract Categorical features appear in datasets from almost every practice area, including real estate datasets. One of the most critical handicaps of machine learning algorithms is that they are not designed to capture the qualitative nature of the categorical features, leading to sub-optimal predictions for the datasets with categorical observations. This study focuses on a new fuzzy regression functions framework, namely hierarchical fuzzy regression functions, that can handle categorical features properly for the regression task. The proposed framework is benchmarked with linear regression, support vector machines, deep neural networks, and adaptive neuro-fuzzy inference systems with real estate data having categorical features from six markets. It is observed that the proposed method produces better prediction performance for real estate price prediction than the benchmark methods in a wide variety of real estate markets. Since we provide all the required software codes to implement the proposed hierarchical fuzzy regression functions framework, our approach offers practitioners a readily applicable, high-performing tool for real estate price prediction and other regression problems involving categorical independent features. Haydar Demirhan, Furkan Baser |
Neural Comput. Appl. | 1 |
| 2022 | Modified fuzzy regression functions with a noise cluster against outlier contamination
Srinivas Chakravarty, Haydar Demirhan, Furkan Baser |
Expert Syst. Appl. | 2 |
| 2020 | A clinical coding recommender system
Mani Suleiman, Haydar Demirhan, Leanne Boyd, Federico Girosi, Vural Aksakalli |
Knowl. Based Syst. | 2 |
| 2019 | A bagging algorithm for the imputation of missing values in time series
Agung Andiojaya, Haydar Demirhan |
Expert Syst. Appl. | 2 |