VLDB 2026 Research / reviewers in the wild / expert
Farhad Pourkamali-Anaraki
dblp:136/5128
· DBLP profile ↗
21ranked-venue papers
16as first author
11since 2021 · last 2026
0000-0003-4078-1676ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 10 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Probabilistic Neural Networks (PNNs) with t-distributed outputs: adaptive prediction intervals beyond Gaussian assumptions
Farhad Pourkamali-Anaraki |
Neural Comput. Appl. | 1 |
| 2025 | Out-of-the-Box Uncertainty: Reducing Confident Errors with Dirichlet Classifiers
Courtney Franzen, Farhad Pourkamali-Anaraki |
ICMLA | 2 |
| 2025 | Kolmogorov-Arnold Networks in Low-Data Regimes: A Comparative Study with Multilayer PerceptronsabstractMultilayer Perceptrons (MLPs) have long been a cornerstone in deep learning, known for their capacity to model complex relationships. Recently, Kolmogorov-Arnold Networks (KANs) have emerged as a compelling alternative, utilizing highly flexible learnable activation functions directly on network edges, a departure from the neuron-centric approach of MLPs. However, KANs significantly increase the number of learnable parameters, raising concerns about their effectiveness in data-scarce environments. Hence, this paper presents a comprehensive comparative study of MLPs and KANs from both algorithmic and experimental perspectives, with a focus on low-data regimes. We introduce an effective technique for designing MLPs with unique, parameterized activation functions for each neuron, enabling a more balanced comparison with KANs. Using empirical evaluations on two real-world data sets from medicine and engineering, we explore the trade-offs between model complexity and accuracy. Our findings show that MLPs with individualized activation functions achieve higher predictive accuracy with only a modest increase in parameters, especially when the sample size is limited to around one hundred. For example, in a classification problem within additive manufacturing, MLPs achieve a median accuracy of 0.91, significantly outperforming KANs, which only reach a median accuracy of 0.53. In addition, we analyze key factors affecting KAN performance, including the polynomial order and grid size of spline-based activation functions. Farhad Pourkamali-Anaraki |
ICMLA | 1 |
| 2024 | Two-stage surrogate modeling for data-driven design optimization with application to composite microstructure generation
Farhad Pourkamali-Anaraki, Jamal F. Husseini, Evan J. Pineda, Brett A. Bednarcyk, Scott E. Stapleton |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Adaptive activation functions for predictive modeling with sparse experimental data
Farhad Pourkamali-Anaraki, Tahamina Nasrin, Robert E. Jensen, Amy M. Peterson |
Neural Comput. Appl. | 1 |
| 2023 | Evaluating Regression Models with Partial Data: A Sampling ApproachabstractMachine learning methods rely on data to uncover relationships between inputs and outputs of complex systems, making it crucial to have sufficient amounts of representative data. Therefore, recent research has focused on choosing informative input-output pairs, i.e., labeled data, to facilitate the adoption of machine learning in science and engineering applications. Despite these efforts, estimating the test error with a limited amount of labeled data still needs to be explored. Hence, this paper investigates a novel framework for selecting informative labeled samples from a set of unlabeled testing instances to evaluate regression models with the quadratic loss function. Key contributions of this work include the design of nonuniform sampling distributions over candidate testing points and the deployment of an unbiased estimator to achieve desirable tradeoffs between estimation accuracy and testing data size. Comprehensive experimental results corroborate the impressive performance and flexibility of the proposed approach in real-world applications, such as reducing the standard deviation of the resulting estimator by almost a factor of two compared to uniform sampling. The paper concludes with practical advice for researchers and practitioners who encounter difficulties related to limited labeled data. Farhad Pourkamali-Anaraki, Mohammad Amin Hariri-Ardebili |
CoDIT | 1 |
| 2023 | Advancing Precision Medicine: An Evaluative Study of Feature Selection MethodsabstractFeature selection methods are instrumental in deciphering intricate patterns within complex datasets. They enable a nuanced balance between the dimensionality of the input and predictive performance, thereby significantly boosting computational efficiency, reliability, and the potential identification of a small set of features as key predictors of disease outcomes. In this study, we extensively explore feature selection methods, focusing on two distinct biomedical datasets, and unveil insights into their performance and utility. Our first dataset is a high-dimensional gene array encompassing 7129 genes, associated with cancer classification tasks. Implementing a novel combinatorial approach, we distilled an optimal subset of 4 genes (99.9% dimensionality decrease). This subset improves the F1 score from 92% to 100%. In addition, our second dataset comprises a more manageable set of 33 features. These features, ranging from clinical and demographic data to symptomatic indicators, are collated from a cohort of 282 patients for the purpose of classifying the presence or absence of heart disease. Our approach improved the F1 score from 95% to 100% on this dataset. By identifying crucial predictive features, we can significantly enhance diagnostic accuracy and help pave the path for individualized treatment approaches, and improve patient outcomes. Our study underlines the importance of feature selection methods and processes as potent instruments in data-driven decision-making and strategy formulation within the healthcare domain. Carol M. Kiekhaefer, Farhad Pourkamali-Anaraki |
ICMLA | 2 |
| 2023 | Evaluation of classification models in limited data scenarios with application to additive manufacturing
Farhad Pourkamali-Anaraki, Tahamina Nasrin, Robert E. Jensen, Amy M. Peterson |
Eng. Appl. Artif. Intell. | 1 |
| 2022 | D-CBRS: Accounting for Intra-Class Diversity in Continual LearningabstractContinual learning – accumulating knowledge from a sequence of learning experiences – is an important yet challenging problem. In this paradigm, the model’s performance for previously encountered instances may substantially drop as additional data are seen. When dealing with class-imbalanced data, forgetting is further exacerbated. Prior work has proposed replay-based approaches which aim at reducing forgetting by intelligently storing instances for future replay. Although Class-Balancing Reservoir Sampling (CBRS) has been successful in dealing with imbalanced data, the intraclass diversity has not been accounted for, implicitly assuming that each instance of a class is equally informative. We present Diverse-CBRS (D-CBRS), an algorithm that allows us to consider within class diversity when storing instances in the memory. Our results show that D-CBRS outperforms state-of-the-art memory management continual learning algorithms on data sets with considerable intra-class diversity. Yasin Findik, Farhad Pourkamali-Anaraki |
ICIP | 2 |
| 2022 | Structural uncertainty quantification with partial information
Mohammad Amin Hariri-Ardebili, Farhad Pourkamali-Anaraki |
Expert Syst. Appl. | 2 |
| 2021 | An Empirical Evaluation of the t-SNE Algorithm for Data Visualization in Structural EngineeringabstractA fundamental task in machine learning involves visualizing high-dimensional data sets that arise in high-impact application domains. When considering the context of large imbalanced data, this problem becomes much more challenging. In this paper, the t-Distributed Stochastic Neighbor Embedding (t-SNE) algorithm is used to reduce the dimensions of an earthquake engineering related data set for visualization purposes. Since imbalanced data sets greatly affect the accuracy of classifiers, we employ Synthetic Minority Oversampling Technique (SMOTE) to tackle the imbalanced nature of such data set. We present the result obtained from t-SNE and SMOTE and compare it to the basic approaches with various aspects. Considering four options and six classification algorithms, we show that using t-SNE on the imbalanced data and SMOTE on the training data set, neural network classifiers have promising results without sacrificing accuracy. Hence, we can transform the studied scientific data into a two-dimensional (2D) space, enabling the visualization of the classifier and the resulting decision surface using a 2D plot. Parisa Hajibabaee, Farhad Pourkamali-Anaraki, Mohammad Amin Hariri-Ardebili |
ICMLA | 2 |
| 2020 | Kernel Ridge Regression Using Importance Sampling with Application to Seismic Response PredictionabstractScalable kernel methods, including kernel ridge regression, often rely on low-rank matrix approximations using the Nyström method, which involves selecting landmark points from large data sets. The existing approaches to selecting landmarks are typically computationally demanding as they require manipulating and performing computations with large matrices in the input or feature space. In this paper, our contribution is twofold. The first contribution is to propose a novel landmark selection method that promotes diversity using an efficient twostep approach. Our landmark selection technique follows a coarse to fine strategy, where the first step computes importance scores with a single pass over the whole data. The second step performs K-means clustering on the constructed coreset to use the obtained centroids as landmarks. Hence, the introduced method provides tunable trade-offs between accuracy and efficiency. Our second contribution is to investigate the performance of several landmark selection techniques using a novel application of kernel methods for predicting structural responses due to earthquake load and material uncertainties. Our experiments exhibit the merits of our proposed landmark selection scheme against baselines. Farhad Pourkamali-Anaraki, Mohammad Amin Hariri-Ardebili, Lydia Morawiec |
ICMLA | 1 |
| 2020 | Efficient Solvers for Sparse Subspace Clustering
Farhad Pourkamali-Anaraki, James Folberth, Stephen Becker |
Signal Process. | 1 |
| 2019 | Improved fixed-rank Nyström approximation via QR decomposition: Practical and theoretical aspects
Farhad Pourkamali-Anaraki, Stephen Becker |
Neurocomputing | 1 |
| 2018 | Randomized Clustered Nystrom for Large-Scale Kernel MachinesabstractThe Nystrom method is a popular technique for generating low-rank approximations of kernel matrices that arise in many machine learning problems. The approximation quality of the Nystrom method depends crucially on the number of selected landmark points and the selection procedure. In this paper, we introduce a randomized algorithm for generating landmark points that is scalable to large high-dimensional data sets. The proposed method performs K-means clustering on low-dimensional random projections of a data set and thus leads to significant savings for high-dimensional data sets. Our theoretical results characterize the tradeoffs between accuracy and efficiency of the proposed method. Moreover, numerical experiments on classification and regression tasks demonstrate the superior performance and efficiency of our proposed method compared with existing approaches. Farhad Pourkamali-Anaraki, Stephen Becker, Michael B. Wakin |
AAAI | 1 |
| 2017 | Preconditioned Data Sparsification for Big Data With Applications to PCA and K-MeansabstractWe analyze a compression scheme for large data sets that randomly keeps a small percentage of the components of each data sample. The benefit is that the output is a sparse matrix, and therefore, subsequent processing, such as principal component analysis (PCA) or K-means, is significantly faster, especially in a distributed-data setting. Furthermore, the sampling is single-pass and applicable to streaming data. The sampling mechanism is a variant of previous methods proposed in the literature combined with a randomized preconditioning to smooth the data. We provide guarantees for PCA in terms of the covariance matrix, and guarantees for K-means in terms of the error in the center estimators at a given step. We present numerical evidence to show both that our bounds are nearly tight and that our algorithms provide a real benefit when applied to standard test data sets, as well as providing certain benefits over related sampling approaches. Farhad Pourkamali-Anaraki, Stephen Becker |
IEEE Trans. Inf. Theory | 1 |
| 2016 | Estimation of the sample covariance matrix from compressive measurementsabstractThis study focuses on the estimation of the sample covariance matrix from low‐dimensional random projections of data known as compressive measurements. In particular, the authors present an unbiased estimator to extract the covariance structure from compressive measurements obtained by a general class of random projection matrices consisting of independent and identically distributed zero‐mean entries and finite first four moments. In contrast to previous works, they make no structural assumptions about the underlying covariance matrix such as being low‐rank. In fact, their analysis is based on a non‐Bayesian data setting which requires no distributional assumptions on the set of data samples. Furthermore, inspired by the generality of the projection matrices, they propose an approach to covariance estimation that utilises sparse Rademacher matrices. Therefore, their algorithm can be used to estimate the covariance matrix in applications with limited memory and computation power at the acquisition devices. Experimental results demonstrate that their approach allows for accurate estimation of the sample covariance matrix on several real‐world data sets including video data. Farhad Pourkamali-Anaraki |
IET Signal Process. | 1 |
| 2014 | Efficient recovery of principal components from compressive measurements with application to Gaussian mixture model estimationabstractThere has been growing interest in performing signal processing tasks directly on compressive measurements, e.g. low-dimensional linear measurements of signals taken with Gaussian random vectors. In this paper, we present a highly efficient algorithm to recover the covariance matrix of high-dimensional data from compressive measurements. We show that, as the number of data samples increases, the eigenvectors (principal components) of the empirical covariance matrix of a simple matrix-vector multiplication of the compressive measurements converge to the true principal components of the original data. Also, we investigate the perturbation of eigenvalues of the covariance matrix under random projection of the data to find conditions for approximate recovery of them. Furthermore, we introduce an important application of our proposed method for efficient estimation of the parameters of Gaussian Mixture Models from compressive measurements. We present experimental results demonstrating the performance and efficiency of our proposed algorithms. Farhad Pourkamali-Anaraki, Shannon M. Hughes |
ICASSP | 1 |
| 2014 | Memory and Computation Efficient PCA via Very Sparse Random ProjectionsabstractAlgorithms that can efficiently recover principal components in very high-dimensional, streaming, and/or distributed data settings have become an important topic in the literature. In this paper, we propose an approach to principal component estimation that utilizes projections onto very sparse random vectors with Bernoulli-generated nonzero entries. Indeed, our approach is simultaneously efficient in memory/storage space, efficient in computation, and produces accurate PC estimates, while also allowing for rigorous theoretical performance analysis. Moreover, one can tune the sparsity of the random vectors deliberately to achieve a desired point on the tradeoffs between memory, computation, and accuracy. We rigorously characterize these tradeoffs and provide statistical performance guarantees. In addition to these very sparse random vectors, our analysis also applies to more general random projections. We present experimental results demonstrating that this approach allows for simultaneously achieving a substantial reduction of the computational complexity and memory/storage space, with little loss in accuracy, particularly for very high-dimensional data. Farhad Pourkamali-Anaraki, Shannon M. Hughes |
ICML | 1 |
| 2013 | Compressive K-SVDabstractDictionary learning algorithms design a dictionary that is specifically tailored to enable sparse representation of a given set of training signals. In turn, the increased sparsity of the signals with respect to this dictionary enables significantly improved performance in a variety of state-of-the-art signal processing tasks, e.g. compressive sensing. However, while these algorithms typically assume that all training data is fully available, this may not be the case in practice. In fact, the high cost of acquiring each signal or the sheer amount of data to be acquired may motivate us to take a compressive sensing (CS) approach, taking only a few CS measurements of each signal. In this paper, we present a novel algorithm for learning a dictionary on a set of training signals using only compressive sensing measurements of them. Our proposed algorithm is a generalization of the well-known K-SVD algorithm and preserves its convergence properties. Experimental results on synthetically generated data verify that our proposed algorithm can recover the generating dictionary atoms from CS measurements alone (so long as enough measurements of enough training signals are available), even for the case of noisy measurements. Finally, we show that compressive K-SVD (CK-SVD) can also be used to aid in signal reconstruction and compressive classification on the CS measurements. Farhad Pourkamali-Anaraki, Shannon M. Hughes |
ICASSP | 1 |
| 2013 | Kernel compressive sensingabstractCompressive sensing allows us to recover signals that are linearly sparse in some basis from a smaller number of measurements than traditionally required. However, it has been shown that many classes of images or video can be more efficiently modeled as lying on a nonlinear manifold, and hence described as a non-linear function of a few underlying parameters. Recently, there has been growing interest in using these manifold models to reduce the required number of compressive sensing measurements. However, the complexity of manifold models has been an obstacle to their use in efficient data acquisition. In this paper, we introduce a new algorithm for applying manifold models in compressive sensing using kernel methods. Our proposed algorithm, kernel compressive sensing (KCS), is the kernel version of classical compressive sensing. It uses dictionary learning in the feature space to build an efficient model for one or more signal manifolds. It then is able to formulate the problem of recovering the signal's coordinates in the manifold representation as an underdetermined linear inverse problem as in traditional compressive sensing. Standard compressive sensing recovery methods can thus be used to recover these coordinates, avoiding additional computational complexity. We present experimental results demonstrating the efficiency and efficacy of this algorithm in manifold-based compressive sensing. Farhad Pourkamali-Anaraki, Shannon M. Hughes |
ICIP | 1 |