Reese Pathak

dblp:251/8459 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 2 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Learning theory · 34% Probabilistic and Bayesian machine learning · 22% Deep learning architectures and training · 22%
Theoretical computer science
2 papers
Mathematical optimization · 80% Information theory · 20%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
mixture of linear regressions
0.812024
Transformers can optimally learn regression mixture models · ICLR 2024
Machine learning › Deep learning architectures and training › neural network expressivity
transformer expressivity
0.812024
Transformers can optimally learn regression mixture models · ICLR 2024
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift
0.612022
A new similarity measure for covariate shift with applications to nonparametric regression · ICML 2022
Machine learning › Learning theory › statistical estimation › minimax estimation
minimax rates
0.612022
A new similarity measure for covariate shift with applications to nonparametric regression · ICML 2022
Machine learning › Learning theory
nonparametric regression
0.612022
A new similarity measure for covariate shift with applications to nonparametric regression · ICML 2022
Mathematical optimization › continuous optimization › matrix optimization › matrix recovery
matrix completion
0.512021
Weighted Matrix Completion From Non-Random, Non-Uniform Sampling Patterns · IEEE Trans. Inf. Theory 2021
Information theory › signal processing › sampling theory
nonuniform sampling
0.512021
Weighted Matrix Completion From Non-Random, Non-Uniform Sampling Patterns · IEEE Trans. Inf. Theory 2021
Mathematical optimization › continuous optimization
convex optimization
0.412020
FedSplit: an algorithmic framework for fast federated optimization · NeurIPS 2020
Mathematical optimization
distributed optimization
0.412020
FedSplit: an algorithmic framework for fast federated optimization · NeurIPS 2020
Mathematical optimization › distributed optimization
federated optimization
0.412020
FedSplit: an algorithmic framework for fast federated optimization · NeurIPS 2020
Mathematical optimization
fixed point analysis
0.412020
FedSplit: an algorithmic framework for fast federated optimization · NeurIPS 2020
Mathematical optimization › continuous optimization › convex optimization
operator splitting
0.412020
FedSplit: an algorithmic framework for fast federated optimization · NeurIPS 2020
Information theory › signal processing
compressed sensing
0.112021
Weighted Matrix Completion From Non-Random, Non-Uniform Sampling Patterns · IEEE Trans. Inf. Theory 2021
Machine learning › Efficient and distributed learning
federated learning
0.112020
FedSplit: an algorithmic framework for fast federated optimization · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

operator splitting · 0.9convergence rate analysis · 0.9mean squared error analysis · 0.8exponential weights · 0.8integral probability metric · 0.6h{ö}lder class analysis · 0.6debiased projection · 0.5
YearPublicationVenuePosition
2024 Transformers can optimally learn regression mixture models
abstract
Mixture models arise in many regression problems, but most methods have seen limited adoption partly due to these algorithms' highly-tailored and model-specific nature. On the other hand, transformers are flexible, neural sequence models that present the intriguing possibility of providing general-purpose prediction methods, even in this mixture setting. In this work, we investigate the hypothesis that transformers can learn an optimal predictor for mixtures of regressions. We construct a generative process for a mixture of linear regressions for which the decision-theoretic optimal procedure is given by data-driven exponential weights on a finite set of parameters. We observe that transformers achieve low mean-squared error on data generated via this process. By probing the transformer's output at inference time, we also show that transformers typically make predictions that are close to the optimal predictor. Our experiments also demonstrate that transformers can learn mixtures of regressions in a sample-efficient fashion and are somewhat robust to distribution shifts. We complement our experimental observations by proving constructively that the decision-theoretic optimal procedure is indeed implementable by a transformer.
Reese Pathak, Rajat Sen, Weihao Kong, Abhimanyu Das
ICLR1
2022 A new similarity measure for covariate shift with applications to nonparametric regression
abstract
We study covariate shift in the context of nonparametric regression. We introduce a new measure of distribution mismatch between the source and target distributions using the integrated ratio of probabilities of balls at a given radius. We use the scaling of this measure with respect to the radius to characterize the minimax rate of estimation over a family of H{ö}lder continuous functions under covariate shift. In comparison to the recently proposed notion of transfer exponent, this measure leads to a sharper rate of convergence and is more fine-grained. We accompany our theory with concrete instances of covariate shift that illustrate this sharp difference.
Reese Pathak, Cong Ma 0001, Martin J. Wainwright
ICML1
2021 Weighted Matrix Completion From Non-Random, Non-Uniform Sampling Patterns
abstract
We study the matrix completion problem when the observation pattern is deterministic and possibly non-uniform. We propose a simple and efficient debiased projection scheme for recovery from noisy observations and analyze the error under a suitable weighted metric. We introduce a simple function of the weight matrix and the sampling pattern that governs the accuracy of the recovered matrix. We derive theoretical guarantees that upper bound the recovery error and nearly matching lower bounds that showcase optimality in several regimes. Our numerical experiments demonstrate the computational efficiency and accuracy of our approach, and show that debiasing is essential when using non-uniform sampling patterns.
Simon Foucart, Deanna Needell, Reese Pathak, Yaniv Plan, Mary Wootters
IEEE Trans. Inf. Theory3
2020 FedSplit: an algorithmic framework for fast federated optimization
abstract
Motivated by federated learning, we consider the hub-and-spoke model of distributed optimization in which a central authority coordinates the computation of a solution among many agents while limiting communication. We first study some past procedures for federated optimization, and show that their fixed points need not correspond to stationary points of the original optimization problem, even in simple convex settings with deterministic updates. In order to remedy these issues, we introduce FedSplit, a class of algorithms based on operator splitting procedures for solving distributed convex minimization with additive structure. We prove that these procedures have the correct fixed points, corresponding to optima of the original optimization problem, and we characterize their convergence rates under different settings. Our theory shows that these methods are provably robust to inexact computation of intermediate local quantities. We complement our theory with some experiments that demonstrate the benefits of our methods in practice.
Reese Pathak, Martin J. Wainwright
NeurIPS1