Shantanu Thakoor

dblp:218/7437 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
8since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Reinforcement learning · 51% Graph learning · 24% Representation and self-supervised learning · 10%
Theoretical computer science
1 paper
Automated reasoning and model checking · 100%

Topics — the 28 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
1.322023
Understanding Self-Predictive Learning for Reinforcement Learning · ICML 2023
Representations and Exploration for Deep Reinforcement Learning using Singular Value Decomposition · ICML 2023
Machine learning › Reinforcement learning
exploration
1.222023
Representations and Exploration for Deep Reinforcement Learning using Singular Value Decomposition · ICML 2023
BYOL-Explore: Exploration by Bootstrapped Prediction · NeurIPS 2022
Machine learning › Graph learning
graph neural network
1.222023
Half-Hop: A graph upsampling approach for slowing down message passing · ICML 2023
Large-Scale Representation Learning on Graphs via Bootstrapping · ICLR 2022
Computer vision › Video understanding and tracking
action anticipation
0.712023
Relax, it doesn't matter how you get there: A new self-supervised approach for multi-timescale behavior analysis · NeurIPS 2023
Computer vision › Video understanding and tracking › video analytics › behavior analysis
animal behavior analysis
0.712023
Relax, it doesn't matter how you get there: A new self-supervised approach for multi-timescale behavior analysis · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning
behavior representation learning
0.712023
Relax, it doesn't matter how you get there: A new self-supervised approach for multi-timescale behavior analysis · NeurIPS 2023
Machine learning › Reinforcement learning › exploration › novelty-based exploration
count-based exploration
0.712023
Representations and Exploration for Deep Reinforcement Learning using Singular Value Decomposition · ICML 2023
Machine learning › Graph learning
graph augmentation
0.712023
Half-Hop: A graph upsampling approach for slowing down message passing · ICML 2023
Machine learning › Graph learning › graph neural network
message passing
0.712023
Half-Hop: A graph upsampling approach for slowing down message passing · ICML 2023
Machine learning › Reinforcement learning
value-based reinforcement learning
0.712023
Understanding Self-Predictive Learning for Reinforcement Learning · ICML 2023
Natural language and speech › Information extraction and text analysis
bootstrapping
0.612022
Large-Scale Representation Learning on Graphs via Bootstrapping · ICLR 2022
Machine learning › Reinforcement learning › exploration › intrinsic motivation
curiosity-driven exploration
0.612022
BYOL-Explore: Exploration by Bootstrapped Prediction · NeurIPS 2022
Machine learning › Reinforcement learning › value-based reinforcement learning
generalized policy improvement
0.612022
Generalised Policy Improvement with Geometric Policy Composition · ICML 2022
Machine learning › Graph learning
graph representation learning
0.612022
Large-Scale Representation Learning on Graphs via Bootstrapping · ICLR 2022
Machine learning › Reinforcement learning
model-based reinforcement learning
0.612022
Generalised Policy Improvement with Geometric Policy Composition · ICML 2022
Machine learning › Reinforcement learning › policy optimization
policy improvement
0.612022
Generalised Policy Improvement with Geometric Policy Composition · ICML 2022
Machine learning › Graph learning › graph representation learning
scalable graph representation learning
0.612022
Large-Scale Representation Learning on Graphs via Bootstrapping · ICLR 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment
0.512021
Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning › value function estimation
future-dependent value function
0.512021
Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.512021
Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021
Automated reasoning and model checking
neural network verification
0.412019
The Marabou Framework for Verification and Analysis of Deep Neural Networks · CAV (1) 2019
Automated reasoning and model checking › satisfiability modulo theories
SMT-based verification
0.412019
The Marabou Framework for Verification and Analysis of Deep Neural Networks · CAV (1) 2019
Program synthesis and code generation
programming by example
0.312018
Synthesis of Programs from Multimodal Datasets · AAAI 2018
Machine learning › Representation and self-supervised learning › representation learning
latent representation learning
0.212023
Understanding Self-Predictive Learning for Reinforcement Learning · ICML 2023
Machine learning › Generative modeling › generative adversarial network
mode collapse mitigation
0.212023
Understanding Self-Predictive Learning for Reinforcement Learning · ICML 2023
Machine learning › Representation and self-supervised learning › matrix factorization
singular value decomposition
0.212023
Representations and Exploration for Deep Reinforcement Learning using Singular Value Decomposition · ICML 2023
Machine learning › Reinforcement learning
model-free reinforcement learning
0.112021
Counterfactual Credit Assignment in Model-Free Reinforcement Learning · ICML 2021
Machine learning › Trustworthy machine learning › verification
robustness verification
0.112019
The Marabou Framework for Verification and Analysis of Deep Neural Networks · CAV (1) 2019

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 1.3spectral decomposition · 0.7singular value decomposition · 0.7semi-gradient updates · 0.7predictive state representation · 0.7multi-scale architecture · 0.7mini-batch training · 0.7message passing · 0.7graph upsampling · 0.7bootstrapping · 0.6parallel execution · 0.4constraint satisfaction · 0.4SMT solving · 0.4least general generalization · 0.3learning-to-rank · 0.3
YearPublicationVenuePosition
2023 Half-Hop: A graph upsampling approach for slowing down message passing
abstract
Message passing neural networks have shown a lot of success on graph-structured data. However, there are many instances where message passing can lead to over-smoothing or fail when neighboring nodes belong to different classes. In this work, we introduce a simple yet general framework for improving learning in message passing neural networks. Our approach essentially upsamples edges in the original graph by adding "slow nodes" at each edge that can mediate communication between a source and a target node. Our method only modifies the input graph, making it plug-and-play and easy to use with existing models. To understand the benefits of slowing down message passing, we provide theoretical and empirical analyses. We report results on several supervised and self-supervised benchmarks, and show improvements across the board, notably in heterophilic conditions where adjacent nodes are more likely to have different labels. Finally, we show how our approach can be used to generate augmentations for self-supervised learning, where slow nodes are randomly introduced into different edges in the graph to generate multi-scale views with variable path lengths.
Mehdi Azabou, Venkataramana Ganesh, Shantanu Thakoor, Chi-Heng Lin, Lakshmi Sathidevi, Michal Valko, Petar Velickovic, Eva L. Dyer
ICML3
2023 Representations and Exploration for Deep Reinforcement Learning using Singular Value Decomposition
abstract
Representation learning and exploration are among the key challenges for any deep reinforcement learning agent. In this work, we provide a singular value decomposition based method that can be used to obtain representations that preserve the underlying transition structure in the domain. Perhaps interestingly, we show that these representations also capture the relative frequency of state visitations, thereby providing an estimate for pseudo-counts for free. To scale this decomposition method to large-scale domains, we provide an algorithm that never requires building the transition matrix, can make use of deep networks, and also permits mini-batch training. Further, we draw inspiration from predictive state representations and extend our decomposition method to partially observable environments. With experiments on multi-task settings with partially observable domains, we show that the proposed method can not only learn useful representation on DM-Lab-30 environments (that have inputs involving language instructions, pixel images, rewards, among others) but it can also be effective at hard exploration tasks in DM-Hard-8 environments.
Yash Chandak, Shantanu Thakoor, Zhaohan Guo, Yunhao Tang, Rémi Munos, Will Dabney, Diana Borsa
ICML2
2023 Understanding Self-Predictive Learning for Reinforcement Learning
abstract
We study the learning dynamics of self-predictive learning for reinforcement learning, a family of algorithms that learn representations by minimizing the prediction error of their own future latent representations. Despite its recent empirical success, such algorithms have an apparent defect: trivial representations (such as constants) minimize the prediction error, yet it is obviously undesirable to converge to such solutions. Our central insight is that careful designs of the optimization dynamics are critical to learning meaningful representations. We identify that a faster paced optimization of the predictor and semi-gradient updates on the representation, are crucial to preventing the representation collapse. Then in an idealized setup, we show self-predictive learning dynamics carries out spectral decomposition on the state transition matrix, effectively capturing information of the transition dynamics. Building on the theoretical insights, we propose bidirectional self-predictive learning, a novel self-predictive algorithm that learns two representations simultaneously. We examine the robustness of our theoretical insights with a number of small-scale experiments and showcase the promise of the novel representation learning algorithm with large-scale experiments.
Yunhao Tang, Zhaohan Guo, Pierre H. Richemond, Bernardo Ávila Pires, Yash Chandak, Rémi Munos, Mark Rowland 0001, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, András György 0001, Shantanu Thakoor, Will Dabney, Bilal Piot, Daniele Calandriello, Michal Valko
ICML12
2023 Relax, it doesn't matter how you get there: A new self-supervised approach for multi-timescale behavior analysis
abstract
Unconstrained and natural behavior consists of dynamics that are complex and unpredictable, especially when trying to predict what will happen multiple steps into the future. While some success has been found in building representations of animal behavior under constrained or simplified task-based conditions, many of these models cannot be applied to free and naturalistic settings where behavior becomes increasingly hard to model. In this work, we develop a multi-task representation learning model for animal behavior that combines two novel components: (i) an action-prediction objective that aims to predict the distribution of actions over future timesteps, and (ii) a multi-scale architecture that builds separate latent spaces to accommodate short- and long-term dynamics. After demonstrating the ability of the method to build representations of both local and global dynamics in robots in varying environments and terrains, we apply our method to the MABe 2022 Multi-Agent Behavior challenge, where our model ranks first overall on both mice and fly benchmarks. In all of these cases, we show that our model can build representations that capture the many different factors that drive behavior and solve a wide range of downstream tasks.
Mehdi Azabou, Michael Mendelson, Nauman Ahad, Maks Sorokin, Shantanu Thakoor, Carolina Urzay, Eva L. Dyer
NeurIPS5
2022 Large-Scale Representation Learning on Graphs via Bootstrapping
Shantanu Thakoor, Corentin Tallec, Mohammad Gheshlaghi Azar, Mehdi Azabou, Eva L. Dyer, Rémi Munos, Petar Velickovic, Michal Valko
ICLR1
2022 Generalised Policy Improvement with Geometric Policy Composition
abstract
We introduce a method for policy improvement that interpolates between the greedy approach of value-based reinforcement learning (RL) and the full planning approach typical of model-based RL. The new method builds on the concept of a geometric horizon model (GHM, also known as a \gamma-model), which models the discounted state-visitation distribution of a given policy. We show that we can evaluate any non-Markov policy that switches between a set of base Markov policies with fixed probability by a careful composition of the base policy GHMs, without any additional learning. We can then apply generalised policy improvement (GPI) to collections of such non-Markov policies to obtain a new Markov policy that will in general outperform its precursors. We provide a thorough theoretical analysis of this approach, develop applications to transfer and standard RL, and empirically demonstrate its effectiveness over standard GPI on a challenging deep RL continuous control task. We also provide an analysis of GHM training methods, proving a novel convergence result regarding previously proposed methods and showing how to train these models stably in deep RL settings.
Shantanu Thakoor, Mark Rowland 0001, Diana Borsa, Will Dabney, Rémi Munos, André Barreto 0001
ICML1
2022 BYOL-Explore: Exploration by Bootstrapped Prediction
abstract
We present BYOL-Explore, a conceptually simple yet general approach for curiosity-driven exploration in visually complex environments. BYOL-Explore learns the world representation, the world dynamics and the exploration policy all-together by optimizing a single prediction loss in the latent space with no additional auxiliary objective. We show that BYOL-Explore is effective in DM-HARD-8, a challenging partially-observable continuous-action hard-exploration benchmark with visually rich 3-D environment. On this benchmark, we solve the majority of the tasks purely through augmenting the extrinsic reward with BYOL-Explore intrinsic reward, whereas prior work could only get off the ground with human demonstrations. As further evidence of the generality of BYOL-Explore, we show that it achieves superhuman performance on the ten hardest exploration games in Atari while having a much simpler design than other competitive agents.
Zhaohan Guo, Shantanu Thakoor, Miruna Pislar, Bernardo Ávila Pires, Florent Altché, Corentin Tallec, Alaa Saade, Daniele Calandriello, Jean-Bastien Grill, Yunhao Tang, Michal Valko, Rémi Munos, Mohammad Gheshlaghi Azar, Bilal Piot
NeurIPS2
2021 Counterfactual Credit Assignment in Model-Free Reinforcement Learning
abstract
Credit assignment in reinforcement learning is the problem of measuring an action’s influence on future rewards. In particular, this requires separating skill from luck, i.e. disentangling the effect of an action on rewards from that of external factors and subsequent actions. To achieve this, we adapt the notion of counterfactuals from causality theory to a model-free RL setup. The key idea is to condition value functions on future events, by learning to extract relevant information from a trajectory. We formulate a family of policy gradient algorithms that use these future-conditional value functions as baselines or critics, and show that they are provably low variance. To avoid the potential bias from conditioning on future information, we constrain the hindsight information to not contain information about the agent’s actions. We demonstrate the efficacy and validity of our algorithm on a number of illustrative and challenging problems.
Thomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, Will Dabney, Thomas S. Stepleton, Nicolas Heess, Arthur Guez, Eric Moulines, Marcus Hutter, Lars Buesing, Rémi Munos
ICML4
2019 The Marabou Framework for Verification and Analysis of Deep Neural Networks
abstract
Deep neural networks are revolutionizing the way complex systems are designed. Consequently, there is a pressing need for tools and techniques for network analysis and certification. To help in addressing that need, we present Marabou, a framework for verifying deep neural networks. Marabou is an SMT-based tool that can answer queries about a network’s properties by transforming these queries into constraint satisfaction problems. It can accommodate networks with different activation functions and topologies, and it performs high-level reasoning on the network that can curtail the search space and improve performance. It also supports parallel execution to further enhance scalability. Marabou accepts multiple input formats, including protocol buffer files generated by the popular TensorFlow framework for neural networks. We describe the system architecture and main components, evaluate the technique and discuss ongoing work.
Guy Katz, Derek A. Huang, Duligur Ibeling, Kyle Julian, Christopher Lazarus, Rachel Lim, Parth Shah 0003, Shantanu Thakoor, Haoze Wu 0001, Aleksandar Zeljic, David L. Dill, Mykel J. Kochenderfer, Clark W. Barrett
CAV (1)8
2018 Synthesis of Programs from Multimodal Datasets
abstract
We describe MultiSynth, a framework for synthesizing domain-specific programs from a multimodal dataset of examples. Given a domain-specific language (DSL), a dataset is multimodal if there is no single program in the DSL that generalizes over all the examples. Further, even if the examples in the dataset were generalized in terms of a set of programs, the domains of these programs may not be disjoint, thereby leading to ambiguity in synthesis. MultiSynth is a framework that incorporates concepts of synthesizing programs with minimum generality, while addressing the need of accurate prediction. We show how these can be achieved through (i) transformation driven partitioning of the dataset, (ii) least general generalization, for a generalized specification of the input and the output, and (iii) learning to rank, for estimating feature weights in order to map an input to the most appropriate mode in case of ambiguity. We show the effectiveness of our framework in two domains: in the first case, we extend an existing approach for synthesizing programs for XML tree transformations to ambiguous multimodal datasets. In the second case, MultiSynth is used to preorder words for machine translation, by learning permutations of productions in the parse trees of the source side sentences. Our evaluations reflect the effectiveness of our approach.
Shantanu Thakoor, Simoni Shah, Ganesh Ramakrishnan, Amitabha Sanyal
AAAI1