Shreyas Chaudhari

dblp:209/9835 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-8826-2253ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Peer-to-Peer Learning Dynamics of Wide Neural Networks
abstract
Peer-to-peer learning is an increasingly popular framework that enables beyond-5G distributed edge devices to collaboratively train deep neural networks in a privacy-preserving manner without the aid of a central server. Neural network training algorithms for emerging environments, e.g., smart cities, have many design considerations that are difficult to tune in deployment settings – such as neural network architectures and hyperparameters. This presents a critical need for characterizing the training dynamics of distributed optimization algorithms used to train highly nonconvex neural networks in peer-to-peer learning environments. In this work, we provide an explicit characterization of the learning dynamics of wide neural networks trained using popular distributed gradient descent (DGD) algorithms. Our results leverage both recent advancements in neural tangent kernel (NTK) theory and extensive previous work on distributed learning and consensus. We validate our analytical results by accurately predicting the parameter and error dynamics of wide neural networks trained for classification tasks.
Shreyas Chaudhari, Srinivasa Pranav, Emile Anand, José M. F. Moura
ICASSP1
2024 Distributional Off-Policy Evaluation for Slate Recommendations
abstract
Recommendation strategies are typically evaluated by using previously logged data, employing off-policy evaluation methods to estimate their expected performance. However, for strategies that present users with slates of multiple items, the resulting combinatorial action space renders many of these methods impractical. Prior work has developed estimators that leverage the structure in slates to estimate the expected off-policy performance, but the estimation of the entire performance distribution remains elusive. Estimating the complete distribution allows for a more comprehensive evaluation of recommendation strategies, particularly along the axes of risk and fairness that employ metrics computable from the distribution. In this paper, we propose an estimator for the complete off-policy performance distribution for slates and establish conditions under which the estimator is unbiased and consistent. This builds upon prior work on off-policy evaluation for slates and off-policy distribution estimation in reinforcement learning. We validate the efficacy of our method empirically on synthetic data as well as on a slate recommendation simulator constructed from real-world data (MovieLens-20M). Our results show a significant reduction in estimation variance and improved sample efficiency over prior work across a range of slate structures.
Shreyas Chaudhari, David T. Arbour, Georgios Theocharous, Nikos Vlassis
AAAI1
2024 From Past to Future: Rethinking Eligibility Traces
abstract
In this paper, we introduce a fresh perspective on the challenges of credit assignment and policy evaluation. First, we delve into the nuances of eligibility traces and explore instances where their updates may result in unexpected credit assignment to preceding states. From this investigation emerges the concept of a novel value function, which we refer to as the ????????????? ????? ????????. Unlike traditional state value functions, bidirectional value functions account for both future expected returns (rewards anticipated from the current state onward) and past expected returns (cumulative rewards from the episode's start to the present). We derive principled update equations to learn this value function and, through experimentation, demonstrate its efficacy in enhancing the process of policy evaluation. In particular, our results indicate that the proposed learning approach can, in certain challenging contexts, perform policy evaluation more rapidly than TD(λ)–a method that learns forward value functions, v^π, ????????. Overall, our findings present a new perspective on eligibility traces and potential advantages associated with the novel value function it inspires, especially for policy evaluation.
Dhawal Gupta, Scott M. Jordan, Shreyas Chaudhari, Bo Liu 0006, Philip S. Thomas, Bruno C. da Silva 0001
AAAI3
2024 Graph Convolutional Neural Networks In The Companion Model
abstract
Graph Convolutional Neural Networks (graph CNNs) adapt the traditional CNN architecture for use on graphs, replacing convolution layers with graph convolution layers. Although similar in architecture, graph CNNs are used for geometric deep learning whereas conventional CNNs are used for deep learning on grid-based data, such as audio or images, with seemingly no direct relationship between the two classes of neural networks.This paper shows that under certain conditions traditional CNNs can be used with graph data as a good approximation to graph CNNs, avoiding the need for graph CNNs. We show this by using an alternative graph signal representation – the graph companion model that we recently proposed in [1]. Instead of using the given graph and signal in the nodal domain, the graph companion model uses the equivalent companion graph and signal representation in the companion domain. By this way, the graph CNN architecture in the nodal domain is equivalent to our deep learning architecture: a traditional CNN in the companion domain with appropriate boundary conditions (b.c.). The paper shows that we obtain similar results on graph classification experiments using a traditional CNN in the companion domain vs. the usual graph CNNs in the nodal domain.
John Shi, Shreyas Chaudhari, José M. F. Moura
ICASSP2
2024 Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
abstract
Evaluating policies using off-policy data is crucial for applying reinforcement learning to real-world problems such as healthcare and autonomous driving. Previous methods for *off-policy evaluation* (OPE) generally suffer from high variance or irreducible bias, leading to unacceptably high prediction errors. In this work, we introduce STAR, a framework for OPE that encompasses a broad range of estimators -- which include existing OPE methods as special cases -- that achieve lower mean squared prediction errors. STAR leverages state abstraction to distill complex, potentially continuous problems into compact, discrete models which we call *abstract reward processes* (ARPs). Predictions from ARPs estimated from off-policy data are provably consistent (asymptotically correct). Rather than proposing a specific estimator, we present a new framework for OPE and empirically demonstrate that estimators within STAR outperform existing methods. The best STAR estimator outperforms baselines in all twelve cases studied, and even the median STAR estimator surpasses the baselines in seven out of the twelve cases.
Shreyas Chaudhari, Ameet Deshpande, Bruno C. da Silva 0001, Philip S. Thomas
NeurIPS1
2023 Learning Gradients of Convex Functions with Monotone Gradient Networks
abstract
While much effort has been devoted to deriving and analyzing effective convex formulations of signal processing problems, the gradients of convex functions also have critical applications ranging from gradient-based optimization to optimal transport. Recent works have explored data-driven methods for learning convex objective functions, but learning their monotone gradients is seldom studied. In this work, we propose C-MGN and M-MGN, two monotone gradient neural network architectures for directly learning the gradients of convex functions. We show that, compared to state of the art methods, our networks are easier to train, learn monotone gradient fields more accurately, and use significantly fewer parameters. We further demonstrate their ability to learn optimal transport mappings to augment driving image data.
Shreyas Chaudhari, Srinivasa Pranav, José M. F. Moura
ICASSP1
2021 Unsupervised Clustering of Time Series Signals Using Neuromorphic Energy-Efficient Temporal Neural Networks
abstract
Unsupervised time series clustering is a challenging problem with diverse industrial applications such as anomaly detection, bio-wearables, etc. These applications typically involve small, low-power devices on the edge that collect and process real-time sensory signals. State-of-the-art time-series clustering methods perform some form of loss minimization that is extremely computationally intensive from the perspective of edge devices. In this work, we propose a neuromorphic approach to unsupervised time series clustering based on Temporal Neural Networks that is capable of ultra low-power, continuous online learning. We demonstrate its clustering performance on a subset of UCR Time Series Archive datasets. Our results show that the proposed approach either outperforms or performs similarly to most of the existing algorithms while being far more amenable for efficient hardware implementation. Our hardware assessment analysis shows that in 7 nm CMOS the proposed architecture, on average, consumes only about 0.005 mm2die area and 22 μW power and can process each signal with about 5 ns latency.
Shreyas Chaudhari, Harideep Nair, José M. F. Moura, John Paul Shen
ICASSP1
2021 A Unified Approach to Translate Classical Bandit Algorithms to Structured Bandits
abstract
We consider a finite-armed structured bandit problem in which mean rewards of different arms are known functions of a common hidden parameter θ*. This problem setting subsumes several previously studied frameworks that assume linear or invertible reward functions. We propose a novel approach to gradually estimate the hidden θ*and use the estimate together with the mean reward functions to substantially reduce exploration of sub-optimal arms. This approach enables us to fundamentally generalize any classic bandit algorithm including UCB and Thompson Sampling to the structured bandit setting. We prove via regret analysis that our proposed UCB-C and TS-C algorithms (structured bandit versions of UCB and Thompson Sampling, respectively) pull only a subset of the sub-optimal arms O(log T ) times while the other sub-optimal arms (referred to as non-competitive arms) are pulled O(1) times. As a result, in cases where all sub-optimal arms are noncompetitive, which can happen in many practical scenarios, the proposed algorithms achieve bounded regret.
Samarth Gupta, Shreyas Chaudhari, Subhojyoti Mukherjee, Gauri Joshi, Osman Yagan
ICASSP2
2021 Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers
Benjamin Eysenbach, Shreyas Chaudhari, Swapnil Asawa, Sergey Levine, Ruslan Salakhutdinov
ICLR2
2021 Multi-Armed Bandits With Correlated Arms
abstract
We consider a multi-armed bandit framework where the rewards obtained by pulling different arms are correlated. We develop a unified approach to leverage these reward correlations and present fundamental generalizations of classic bandit algorithms to the correlated setting. We present a unified proof technique to analyze the proposed algorithms. Rigorous analysis of C-UCB (the correlated bandit versions of Upper-confidence-bound) reveals that the algorithm end up pulling certain sub-optimal arms, termed as non-competitive, only O(1) times, as opposed to the O(logT) pulls required by classic bandit algorithms such as UCB, TS etc. We present regret-lower bound and show that when arms are correlated through a latent random source, our algorithms obtain order-optimal regret. We validate the proposed algorithms via experiments on the MovieLens and Goodreads datasets, and show significant improvement over classical bandit algorithms.
Samarth Gupta, Shreyas Chaudhari, Gauri Joshi, Osman Yagan
IEEE Trans. Inf. Theory2