VLDB 2026 Research / reviewers in the wild / expert
Sona Taheri
dblp:139/8435
· DBLP profile ↗
10ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-1779-4567ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A clustering framework for skewed features with low true cluster separationabstractAbstract Features with considerably larger or smaller observations than the rest of the dataset, causing noticeable skewness in the feature distributions, are prevalent in practical applications. Traditional clustering methods often assume symmetric data, leading to poor performance with skewed features. This challenge becomes further complicated in datasets with low separation between true clusters in the feature space. These problems are encountered in a wide range of important practical areas, such as cell grouping, forest fires, maritime search and rescue, urbanization studies, and neuroimaging. Bayesian model-based clustering methods can accurately capture the skewness in the data and centers of poorly separated true clusters. However, they are computationally inefficient due to their Bayesian nature. We propose a Bayesian model-based clustering framework to address these issues by utilizing the generalized multivariate log-gamma distribution with a Dirichlet process mixture. Comparative numerical experiments on 30 benchmark datasets with traditional and Bayesian model-based clustering algorithms demonstrate the superior performance of the proposed method, particularly for skewed datasets with low true cluster separability. The proposed approach, implemented in R, also shows better computational efficiency than its Bayesian alternatives. The computer codes to implement our approach are provided to facilitate practical applications. Muntazir Mehdi, Haydar Demirhan, Sona Taheri |
Data Min. Knowl. Discov. | 3 |
| 2026 | Absolute indices for determining compactness, separability and number of clustersabstractIdentifying “true” clusters in a dataset is inherently challenging. Clustering models and algorithms may fail to produce compact, well-separated groups or to correctly determine the optimal number of clusters. Cluster validity indices are often applied to find such clusters. However, most existing indices are relative measures, designed primarily for comparing clustering algorithms or tuning their parameters. In this paper, we introduce new absolute cluster validity indices that assess both cluster compactness and separability. For each cluster, we define a compactness function, and for each pair of clusters, we define an associated set of neighboring points. The compactness function measures the cohesion of individual clusters as well as the entire clustering distribution, while the neighboring-point sets are used to define the margin between cluster pairs and the overall distribution margin. These compactness and separability indices are then employed to estimate the true number of clusters. We evaluate the proposed indices on a variety of synthetic and real-world datasets and compare their performance with several widely used cluster validity measures. Adil M. Bagirov, Ramiz M. Aliguliyev, Nargiz Sultanova, Sona Taheri |
Pattern Recognit. | 4 |
| 2025 | ConASD: Contrastive Few Shot Learning for Detecting Autism Spectrum Disorder via Eye Tracking ScanpathabstractAbstract Detecting Autism Spectrum Disorder (ASD) using Eye Tracking (ET) datasets is a challenging task and has been a long-standing problem. Recently, there has been a trend of developing ASD diagnosis models based on machine learning (ML), especially deep learning techniques. In this paper, we show that these existing methods still struggle to make accurate diagnoses in few-shot learning (FSL) settings, where the data available for training is limited in amount and imbalanced in nature. To address this challenge, we propose a model, named ConASD, for effective diagnosis of ASD under the FSL setting. The proposed model is a two-stage framework: it first trains an encoder for ET images using supervised contrastive learning, followed by fine-tuning a classifier for final diagnosis. With the contrastive learning strategy, the pre-trained encoder can better capture the discriminative features of the eye-tracking images, even with limited training data, and ultimately leads to better diagnosis accuracy and better generalization to unseen data. We evaluate the proposed ConASD model using two real-world ET datasets. The results demonstrate that ConASD outperforms existing approaches, particularly in few-shot scenarios, by up-to 7% improvement in terms of F1 scores. The results in this paper highlight the potential of using contrastive learning as a powerful tool, particularly in real-world medical scenarios where class imbalance is frequent and the data is limited. Sharifah Mousli, Sona Taheri, Estrid He |
Multim. Syst. | 2 |
| 2024 | Robust clustering algorithm: The use of soft trimming approachabstractThe presence of noise or outliers in data sets may heavily affect the performance of clustering algorithms and lead to unsatisfactory results. The majority of conventional clustering algorithms are sensitive to noise and outliers. Robust clustering algorithms often overcome difficulties associated with noise and outliers and find true cluster structures. We introduce a soft trimming approach for the hard clustering problem where its objective is modeled as a sum of the cluster function and a function represented as a composition of the algebraic and distance functions. We utilize the composite function to estimate the degree of the significance of each data point in clustering. A robust clustering algorithm based on the new model and a procedure for generating starting cluster centers is developed. We demonstrate the performance of the proposed algorithm using some synthetic and real-world data sets containing noise and outliers. We also compare its performance with that of some well-known clustering techniques. Results show that the new algorithm is robust to noise and outliers and finds true cluster structures. Sona Taheri, Adil M. Bagirov, Nargiz Sultanova, Burak Ordin |
Pattern Recognit. Lett. | 1 |
| 2023 | Methods and Applications of Clusterwise Linear Regression: A Survey and ComparisonabstractClusterwise linear regression (CLR) is a well-known technique for approximating a data using more than one linear function. It is based on the combination of clustering and multiple linear regression methods. This article provides a comprehensive survey and comparative assessments of CLR including model formulations, description of algorithms, and their performance on small to large-scale synthetic and real-world datasets. Some applications of the CLR algorithms and possible future research directions are also discussed. Qiang Long, Adil M. Bagirov, Sona Taheri, Nargiz Sultanova |
ACM Trans. Knowl. Discov. Data | 3 |
| 2023 | Nonsmooth Optimization-Based Model and Algorithm for Semisupervised ClusteringabstractUsing a nonconvex nonsmooth optimization approach, we introduce a model for semisupervised clustering (SSC) with pairwise constraints. In this model, the objective function is represented as a sum of three terms: the first term reflects the clustering error for unlabeled data points, the second term expresses the error for data points with must-link (ML) constraints, and the third term represents the error for data points with cannot-link (CL) constraints. This function is nonconvex and nonsmooth. To find its optimal solutions, we introduce an adaptive SSC (A-SSC) algorithm. This algorithm is based on the combination of the nonsmooth optimization method and an incremental approach, which involves the auxiliary SSC problem. The algorithm constructs clusters incrementally starting from one cluster and gradually adding one cluster center at each iteration. The solutions to the auxiliary SSC problem are utilized as starting points for solving the nonconvex SSC problem. The discrete gradient method (DGM) of nonsmooth optimization is applied to solve the underlying nonsmooth optimization problems. This method does not require subgradient evaluations and uses only function values. The performance of the A-SSC algorithm is evaluated and compared with four benchmarking SSC algorithms on one synthetic and 12 real-world datasets. Results demonstrate that the proposed algorithm outperforms the other four algorithms in identifying compact and well-separated clusters while satisfying most constraints. Adil M. Bagirov, Sona Taheri, Fusheng Bai, Fangying Zheng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Missing Value Imputation via Clusterwise Linear RegressionabstractIn this paper a new method of preprocessing incomplete data is introduced. The method is based on clusterwise linear regression and it combines two well-known approaches for missing value imputation: linear regression and clustering. The idea is to approximate missing values using only those data points that are somewhat similar to the incomplete data point. A similar idea is used also in clustering based imputation methods. Nevertheless, here the linear regression approach is used within each cluster to accurately predict the missing values, and this is done simultaneously to clustering. The proposed method is tested using some synthetic and real-world data sets and compared with other algorithms for missing value imputations. Numerical results demonstrate that this method produces the most accurate imputations in MCAR and MAR data sets with a clear structure and the percentages of missing data no more than 25 percent. Napsu Karmitsa, Sona Taheri, Adil M. Bagirov, Pauliina Mäkinen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2018 | Clustering in large data sets with the limited memory bundle method
Napsu Karmitsa, Adil M. Bagirov, Sona Taheri |
Pattern Recognit. | 3 |
| 2016 | Nonsmooth DC programming approach to the minimum sum-of-squares clustering problems
Adil M. Bagirov, Sona Taheri, Julien Ugon |
Pattern Recognit. | 2 |
| 2014 | Attribute weighted Naive Bayes classifier using a local optimization
Sona Taheri, John Yearwood, Musa A. Mammadov, Sattar Seifollahi |
Neural Comput. Appl. | 1 |