Harsha Honnappa

dblp:19/6913 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0002-0834-054XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Mathematical optimization · 100%
Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 78% Learning theory · 22%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization
distributed optimization
1.322023
Distributed (ATC) Gradient Descent for High Dimension Sparse Regression · IEEE Trans. Inf. Theory 2023
Distributed Sparse Regression via Penalization · J. Mach. Learn. Res. 2023
Mathematical optimization › statistical estimation › regression › sparse regression
lasso
1.322023
Distributed (ATC) Gradient Descent for High Dimension Sparse Regression · IEEE Trans. Inf. Theory 2023
Distributed Sparse Regression via Penalization · J. Mach. Learn. Res. 2023
Mathematical optimization › statistical estimation › regression
sparse regression
1.322023
Distributed (ATC) Gradient Descent for High Dimension Sparse Regression · IEEE Trans. Inf. Theory 2023
Distributed Sparse Regression via Penalization · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning
bayesian decision theory
0.712023
On the Statistical Consistency of Risk-Sensitive Bayesian Decision-Making · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
variational bayesian inference
0.712023
On the Statistical Consistency of Risk-Sensitive Bayesian Decision-Making · NeurIPS 2023
Mathematical optimization › distributed optimization
consensus optimization
0.712023
Distributed Sparse Regression via Penalization · J. Mach. Learn. Res. 2023
Mathematical optimization › statistical estimation
high-dimensional estimation
0.712023
Distributed Sparse Regression via Penalization · J. Mach. Learn. Res. 2023
Mathematical optimization
statistical learning theory
0.712023
Distributed Sparse Regression via Penalization · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference
0.412020
Asymptotic Consistency of α-Rényi-Approximate Posteriors · J. Mach. Learn. Res. 2020
Machine learning › Learning theory › statistical estimation
asymptotic consistency
0.412020
Asymptotic Consistency of α-Rényi-Approximate Posteriors · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian asymptotics
posterior consistency
0.412020
Asymptotic Consistency of α-Rényi-Approximate Posteriors · J. Mach. Learn. Res. 2020
Machine learning › Learning theory › statistical estimation
statistical consistency
0.212023
On the Statistical Consistency of Risk-Sensitive Bayesian Decision-Making · NeurIPS 2023
Distributed systems
consensus
0.212023
Distributed (ATC) Gradient Descent for High Dimension Sparse Regression · IEEE Trans. Inf. Theory 2023
Distributed systems › distributed coordination › multi-agent systems
multi-agent networks
0.212023
Distributed (ATC) Gradient Descent for High Dimension Sparse Regression · IEEE Trans. Inf. Theory 2023
Mathematical optimization › continuous optimization
convex optimization
0.212023
Distributed Sparse Regression via Penalization · J. Mach. Learn. Res. 2023
Mathematical optimization › continuous optimization › convex optimization › proximal methods
proximal gradient method
0.212023
Distributed Sparse Regression via Penalization · J. Mach. Learn. Res. 2023

Methods — techniques the papers use, named apart from their topics

restricted strong convexity · 1.3convergence analysis · 1.3adapt-then-combine · 1.3proximal gradient · 0.7penalty method · 0.7entropic risk measure · 0.7dual representation · 0.7rényi divergence · 0.4asymptotic analysis · 0.4
YearPublicationVenuePosition
2023 On the Statistical Consistency of Risk-Sensitive Bayesian Decision-Making
abstract
We study data-driven decision-making problems in the Bayesian framework, where the expectation in the Bayes risk is replaced by a risk-sensitive entropic risk measure with respect to the posterior distribution. We focus on problems where calculating the posterior distribution is intractable, a typical situation in modern applications with large datasets and complex data generating models. We leverage a dual representation of the entropic risk measure to introduce a novel risk-sensitive variational Bayesian (RSVB) framework for jointly computing a risk-sensitive posterior approximation and the corresponding decision rule. Our general framework includes \textit{loss-calibrated} VB (Lacoste-Julien et al. [2011] ) as a special case. We also study the impact of these computational approximations on the predictive performance of the inferred decision rules. We compute the convergence rates of the RSVB approximate posterior and the corresponding optimal value. We illustrate our theoretical findings in parametric and nonparametric settings with the help of three examples.
Prateek Jaiswal, Harsha Honnappa, Vinayak A. Rao
NeurIPS2
2023 Distributed Sparse Regression via Penalization
abstract
We study sparse linear regression over a network of agents, modeled as an undirected graph (with no centralized node). The estimation problem is formulated as the minimization of the sum of the local LASSO loss functions plus a quadratic penalty of the consensus constraint—the latter being instrumental to obtain distributed solution methods. While penalty-based consensus methods have been extensively studied in the optimization literature, their statistical and computational guarantees in the high dimensional setting remain unclear. This work provides an answer to this open problem. Our contribution is two-fold. First, we establish statistical consistency of the estimator: under a suitable choice of the penalty parameter, the optimal solution of the penalized problem achieves near optimal minimax rate $O(s \log d/N)$ in $\ell_2$-loss, where $s$ is the sparsity value, $d$ is the ambient dimension, and $N$ is the total sample size in the network—this matches centralized sample rates. Second, we show that the proximal-gradient algorithm applied to the penalized problem, which naturally leads to distributed implementations, converges linearly up to a tolerance of the order of the centralized statistical error---the rate scales as $O(d)$, revealing an unavoidable speed-accuracy dilemma. Numerical results demonstrate the tightness of the derived sample rate and convergence rate scalings.
Gesualdo Scutari, Ying Sun 0003, Harsha Honnappa
J. Mach. Learn. Res.4
2023 Distributed (ATC) Gradient Descent for High Dimension Sparse Regression
abstract
We study linear regression from data distributed over a network of agents (with no server node) by means of LASSO estimation, in high-dimension, which allows the ambient dimension to grow faster than the sample size. While there is a vast literature of distributed algorithms applicable to the problem, statistical and computational guarantees of most of them remain unclear in high dimension. This paper provides a first statistical study of the Distributed Gradient Descent (DGD) in the Adapt-Then-Combine (ATC) form. Our theory shows that, under standard notions of restricted strong convexity and smoothness of the loss functions–which hold with high probability for standard data generation models–suitable conditions on the network connectivity and algorithm tuning, DGD-ATC converges globally at a linear rate to an estimate that is within thecentralizedstatistical precision of the model. In the worst-case scenario, the total number of communications to statistical optimality grows logarithmically with the ambient dimension, which improves on the communication complexity of DGD in the Combine-Then-Adapt (CTA) form, scaling linearly with the dimension. This reveals that mixing gradient information among agents, as DGD-ATC does, is critical in high-dimensions to obtain favorable rate scalings.
Gesualdo Scutari, Ying Sun 0003, Harsha Honnappa
IEEE Trans. Inf. Theory4
2020 Asymptotic Consistency of α-Rényi-Approximate Posteriors
Prateek Jaiswal, Vinayak A. Rao, Harsha Honnappa
J. Mach. Learn. Res.3