Bala Rajaratnam

dblp:24/666 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0003-2570-8059ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 since 2021Theory of computation · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Large Scale Partial Correlation Screening With Uncertainty Quantification
abstract
Identifying multivariate dependencies in high-dimensional data is an important problem in large-scale inference. This problem has motivated recent advances in mining (partial) correlations, which focus on the challenging ultra-high dimensional setting where the sample size,n, is fixed, while the number of features,p, grows without bound. The state-of-the-art method for partial correlation screening can lead to undesirable results. This paper introduces a novel principled framework for partial correlation screening with error control (PARSEC), which leverages the connection between partial correlations and regression coefficients. We establish the inferential properties of PARSEC whennis fixed andpgrows super-exponentially. First, we provide “fixed-n-large-p” asymptotic expressions for the familywise error rate (FWER) andk-FWER. Equally importantly, our analysis leads to a novel discovery which permits the calculation of exact marginal p-values for controlling the false discovery rate (FDR), and also the positive FDR (pFDR). To our knowledge, no other competing approach in the “fixed-n-large-p” setting allows for error control across the spectrum of multiple hypothesis testing metrics. We establish the computational complexity of PARSEC and rigorously demonstrate its scalability to the largepsetting. The theory and methods are successfully validated on simulated and real data, and PARSEC is shown to outperform the current state-of-the-art.
Emily Neo, Peter Radchenko, Bala Rajaratnam
IEEE Trans. Inf. Theory3
2023 Hierarchical Relational Learning for Few-Shot Knowledge Graph Completion
Han Wu 0009, Jie Yin 0001, Bala Rajaratnam, Jianyuan Guo
ICLR3
2023 A Unified Framework for Correlation Mining in Ultra-High Dimension
abstract
Many applications benefit from theory relevant to the identification of variables having large correlations or partial correlations in high dimension. Recently there has been progress in the ultra-high dimensional setting when the sample size$n$is fixed and the dimension$p$tends to infinity. Despite these advances, the correlation screening framework suffers from practical, methodological and theoretical deficiencies. For instance, previous correlation screening theory requires that the population covariance matrix be sparse and block diagonal. This block sparsity assumption is however restrictive in practical applications. As a second example, correlation and partial correlation screening requires the estimation of dependence measures, which can be computationally prohibitive. In this paper, we propose a unifying approach to correlation and partial correlation mining that is not restricted to block diagonal correlation structure, thus yielding a methodology that is suitable for modern applications. By making connections to random geometric graphs, the number of highly correlated or partial correlated variables are shown to have compound Poisson finite-sample characterizations, which hold for both the finite$p$case and when$p$tends to infinity. The unifying framework also demonstrates a duality between correlation and partial correlation screening with theoretical and practical consequences.
Bala Rajaratnam, Alfred O. Hero III
IEEE Trans. Inf. Theory2
2019 A scalable sparse Cholesky based approach for learning high-dimensional covariance matrices in ordered data
Kshitij Khare, Sang-Yun Oh, Syed Rahman, Bala Rajaratnam
Mach. Learn.4
2017 Generalized Pseudolikelihood Methods for Inverse Covariance Estimation
abstract
We introduce PseudoNet, a new pseudolikelihood-based estimator of the inverse covariance matrix, that has a number of useful statistical and computational properties. We show, through detailed experiments with synthetic and also real-world finance as well as wind power data, that PseudoNet outperforms related methods in terms of estimation error and support recovery, making it well-suited for use in a downstream application, where obtaining low estimation error can be important. We also show, under regularity conditions, that PseudoNet is consistent. Our proof assumes the existence of accurate estimates of the diagonal entries of the underlying inverse covariance matrix; we additionally provide a two-step method to obtain these estimates, even in a high-dimensional setting, going beyond the proofs for related methods. Unlike other pseudolikelihood-based methods, we also show that PseudoNet does not saturate, i.e., in high dimensions, there is no hard limit on the number of nonzero entries in the PseudoNet estimate. We present a fast algorithm as well as screening rules that make computing the PseudoNet estimate over a range of tuning parameters tractable.
Alnur Ali, Kshitij Khare, Sang-Yun Oh, Bala Rajaratnam
AISTATS4
2017 Two-Stage Sampling, Prediction and Adaptive Regression via Correlation Screening
abstract
This paper proposes a general adaptive procedure for budget-limited predictor design in high dimensions called two-stage Sampling, Prediction and Adaptive Regression via Correlation Screening (SPARCS). The SPARCS can be applied to high-dimensional prediction problems in experimental science, medicine, finance, and engineering, as illustrated by the following. Suppose that one wishes to run a sequence of experiments to learn a sparse multivariate predictor of a dependent variable Y (disease prognosis for instance) based on a p dimensional set of independent variables X = [X1,..., Xp]T(assayed biomarkers). Assume that the cost of acquiring the full set of variables X increases linearly in its dimension. The SPARCS breaks the data collection into two stages in order to achieve an optimal tradeoff between sampling cost and predictor performance. In the first stage, we collect a few (n) expensive samples {yi, si}i=1n, at the full dimension p ≫ n of X, winnowing the number of variables down to a smaller dimension l <; p using a type of cross correlation or regression coefficient screening. In the second stage, we collect a larger number (t - n) of cheaper samples of the l variables that passed the screening of the first stage. At the second stage, a low-dimensional predictor is constructed by solving the standard regression problem using all t samples of the selected variables. The SPARCS is an adaptive online algorithm that implements false positive control on the selected variables, is well suited to small sample sizes, and is scalable to high dimensions. We establish asymptotic bounds for the familywise error rate, specify high dimensional convergence rates for support recovery, and establish optimal sample allocation rules to the first and second stages.
Hamed Firouzi, Alfred O. Hero III, Bala Rajaratnam
IEEE Trans. Inf. Theory3
2016 Foundational Principles for Large-Scale Inference: Illustrations Through Correlation Mining
abstract
When can reliable inference be drawn in the “Big Data” context? This paper presents a framework for answering this fundamental question in the context of correlation mining, with implications for general large-scale inference. In large-scale data applications like genomics, connectomics, and eco-informatics, the data set is often variable rich but sample starved: a regime where the number n of acquired samples (statistical replicates) is far fewer than the number p of observed variables (genes, neurons, voxels, or chemical constituents). Much of recent work has focused on understanding the computational complexity of proposed methods for “Big Data.” Sample complexity, however, has received relatively less attention, especially in the setting when the sample size n is fixed, and the dimension p grows without bound. To address this gap, we develop a unified statistical framework that explicitly quantifies the sample complexity of various inferential tasks. Sampling regimes can be divided into several categories: 1) the classical asymptotic regime where the variable dimension is fixed and the sample size goes to infinity; 2) the mixed asymptotic regime where both variable dimension and sample size go to infinity at comparable rates; and 3) the purely high-dimensional asymptotic regime where the variable dimension goes to infinity and the sample size is fixed. Each regime has its niche but only the latter regime applies to exa-scale data dimension. We illustrate this high-dimensional framework for the problem of correlation mining, where it is the matrix of pairwise and partial correlations among the variables that are of interest. Correlation mining arises in numerous applications and subsumes the regression context as a special case. We demonstrate various regimes of correlation mining based on the unifying perspective of high-dimensional learning rates and sample complexity for different structured covariance models and different inference tasks.
Alfred O. Hero III, Bala Rajaratnam
Proc. IEEE2
2014 Optimization Methods for Sparse Pseudo-Likelihood Graphical Model Selection
Sang-Yun Oh, Onkar Dalal, Kshitij Khare, Bala Rajaratnam
NIPS4
2013 Predictive Correlation Screening: Application to Two-stage Predictor Design in High Dimension
abstract
We introduce a new approach to variable selection, called Predictive Correlation Screening, for predictor design. Predictive Correlation Screening (PCS) implements false positive control on the selected variables, is well suited to small sample sizes, and is scalable to high dimensions. We establish asymptotic bounds for Familywise Error Rate (FWER), and resultant mean square error of a linear predictor on the selected variables. We apply Predictive Correlation Screening to the following two-stage predictor design problem. An experimenter wants to learn a multivariate predictor of gene expressions based on successive biological samples assayed on mRNA arrays. She assays the whole genome on a few samples and from these assays she selects a small number of variables using Predictive Correlation Screening. To reduce assay cost, she subsequently assays only the selected variables on the remaining samples, to learn the predictor coefficients. We show superiority of Predictive Correlation Screening relative to LASSO and correlation learning (sometimes popularly referred to in the literature as marginal regression or simple thresholding) in terms of performance and computational complexity.
Hamed Firouzi, Bala Rajaratnam, Alfred O. Hero III
AISTATS2
2012 Iterative Thresholding Algorithm for Sparse Inverse Covariance Estimation
abstract
Sparse graphical modelling/inverse covariance selection is an important problem in machine learning and has seen significant advances in recent years. A major focus has been on methods which perform model selection in high dimensions. To this end, numerous convex $\ell_1$ regularization approaches have been proposed in the literature. It is not however clear which of these methods are optimal in any well-defined sense. A major gap in this regard pertains to the rate of convergence of proposed optimization methods. To address this, an iterative thresholding algorithm for numerically solving the $\ell_1$-penalized maximum likelihood problem for sparse inverse covariance estimation is presented. The proximal gradient method considered in this paper is shown to converge at a linear rate, a result which is the first of its kind for numerically solving the sparse inverse covariance estimation problem. The convergence rate is provided in closed form, and is related to the condition number of the optimal point. Numerical results demonstrating the proven rate of convergence are presented.
Benjamin T. Rolfs, Bala Rajaratnam, Dominique Guillot, Ian Wong, Arian Maleki
NIPS2
2012 Hub Discovery in Partial Correlation Graphs
abstract
One of the most important problems in large-scale inference problems is the identification of variables that are highly dependent on several other variables. When dependence is measured by partial correlations, these variables identify those rows of the partial correlation matrix that have several entries with large magnitudes, i.e., hubs in the associated partial correlation graph. This paper develops theory and algorithms for discovering such hubs from a few observations of these variables. We introduce a hub screening framework in which the user specifies both a minimum (partial) correlation$\rho $and a minimum degree$\delta $to screen the vertices. The choice of$\rho $and$\delta $can be guided by our mathematical expressions for the phase transition correlation threshold$\rho _{c}$governing the average number of discoveries. They can also be guided by our asymptotic expressions for familywise discovery rates under the assumption of large number$p$of variables, fixed number$n$of multivariate samples, and weak dependence. Under the null hypothesis that the dispersion (covariance) matrix is sparse, these limiting expressions can be used to enforce familywise error constraints and to rank the discoveries in order of increasing statistical significance. For$n\ll p$, the computational complexity of the proposed partial correlation screening method is low and is therefore highly scalable. Thus, it can be applied to significantly larger problems than previous approaches. The theory is applied to discovering hubs in a high-dimensional gene microarray dataset.
Alfred O. Hero III, Bala Rajaratnam
IEEE Trans. Inf. Theory2
2011 A local dependence measure and its application to screening for high correlations in large data sets
Kumar Sricharan, Alfred O. Hero III, Bala Rajaratnam
FUSION3
2008 Component-wise parameter smoothing for learning mixture models
abstract
In this paper, we propose a novel component-wise smoothing algorithm that constructs a hierarchy (or family) of smoothened log-likelihood surfaces. Our approach first smoothens the likelihood function and then applies the EM algorithm to obtain a promising solution on this smoothened surface. Using the most promising solutions as initial guesses, the EM algorithm is applied again on the original likelihood. This effective optimization procedure eliminates extensive search in the non-promising regions of the parameter space. Empirical results on some standard datasets show the reduction of the number of local maxima and improvements in the log-likelihood values.
Chandan K. Reddy, Bala Rajaratnam
ICPR2
2008 TRUST-TECH-Based Expectation Maximization for Learning Finite Mixture Models
abstract
In spite of the initialization problem, the Expectation-Maximization (EM) algorithm is widely used for estimating the parameters of finite mixture models. Most popular model-based clustering techniques might yield poor clusters if the parameters are not initialized properly. To reduce the sensitivity of initial points, a novel algorithm for learning mixture models from multivariate data is introduced in this paper. The proposed algorithm takes advantage of TRUST-TECH (TRansformation Under STability-reTaining Equilibra CHaracterization) to compute neighborhood local maxima on likelihood surface using stability regions. Basically, our method coalesces the advantages of the traditional EM with that of the dynamic and geometric characteristics of the stability regions of the corresponding nonlinear dynamical system of the log-likelihood function. Two phases namely, the EM phase and the stability region phase, are repeated alternatively in the parameter space to achieve improvements in the maximum likelihood. The EM phase obtains the local maximum of the likelihood function and the stability region phase helps to escape out of the local maximum by moving towards the neighboring stability regions. The algorithm has been tested on both synthetic and real datasets and the improvements in the performance compared to other approaches are demonstrated. The robustness with respect to initialization is also illustrated experimentally.
Chandan K. Reddy, Hsiao-Dong Chiang, Bala Rajaratnam
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Stability Region Based Expectation Maximization for Model-based Clustering
abstract
In spite of the initialization problem, the expectation-maximization (EM) algorithm is widely used for estimating the parameters in several data mining related tasks. Most popular model-based clustering techniques might yield poor clusters if the parameters are not initialized properly. To reduce the sensitivity of initial points, a novel algorithm for learning mixture models from multivariate data is introduced in this paper. The proposed algorithm takes advantage of TRUST-TECH (TRansformation Under STability- reTaining Equilibra CHaracterization) to compute neighborhood local maxima on likelihood surface using stability regions. Basically, our method coalesces the advantages of the traditional EM with that of the dynamic and geometric characteristics of the stability regions of the corresponding nonlinear dynamical system of the log-likelihood function. Two phases namely, the EM phase and the stability region phase, are repeated alternatively in the parameter space to achieve improvements in the maximum likelihood. Though applied to Gaussian mixtures in this paper, our technique can be easily generalized to any other parametric finite mixture model. The algorithm has been tested on both synthetic and real datasets and the improvements in the performance compared to other approaches are demonstrated. The robustness with respect to initialization is also illustrated experimentally.
Chandan K. Reddy, Hsiao-Dong Chiang, Bala Rajaratnam
ICDM3