Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ankesh Anand

dblp:192/1710 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 43% Representation and self-supervised learning · 26% Language models and text generation · 23%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
model-based reinforcement learning
1.222023
Investigating the Role of Model-Based Learning in Exploration and Transfer · ICML 2023
Procedural generalization by planning with self-supervised world models · ICLR 2022
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
1.022021
Pretraining Representations for Data-Efficient Reinforcement Learning · NeurIPS 2021
Data-Efficient Reinforcement Learning with Self-Predictive Representations · ICLR 2021
Natural language and speech › Language models and text generation
in-context learning
0.812024
Many-Shot In-Context Learning · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model reasoning
0.812024
Many-Shot In-Context Learning · NeurIPS 2024
Natural language and speech › Language models and text generation › in-context learning
many-shot in-context learning
0.812024
Many-Shot In-Context Learning · NeurIPS 2024
Machine learning › Reinforcement learning
exploration
0.712023
Investigating the Role of Model-Based Learning in Exploration and Transfer · ICML 2023
Machine learning › Learning theory
generalization
0.612022
Procedural generalization by planning with self-supervised world models · ICLR 2022
Machine learning › Reinforcement learning › generalization in reinforcement learning
procedural generalization
0.612022
Procedural generalization by planning with self-supervised world models · ICLR 2022
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.612022
Procedural generalization by planning with self-supervised world models · ICLR 2022
Machine learning › Representation and self-supervised learning › pre-training
unsupervised pre-training
0.512021
Pretraining Representations for Data-Efficient Reinforcement Learning · NeurIPS 2021
Machine learning › Representation and self-supervised learning
contrastive learning
0.412019
Unsupervised State Representation Learning in Atari · NeurIPS 2019
Machine learning › Representation and self-supervised learning
mutual information maximization
0.412019
Unsupervised State Representation Learning in Atari · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning
0.412019
Unsupervised State Representation Learning in Atari · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.412019
Unsupervised State Representation Learning in Atari · NeurIPS 2019
Machine learning › Transfer learning and domain adaptation
knowledge transfer
0.212023
Investigating the Role of Model-Based Learning in Exploration and Transfer · ICML 2023

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 1.7supervised fine-tuning · 0.8chain-of-thought prompting · 0.8model-based learning · 0.7planning · 0.6latent dynamics modeling · 0.5goal-conditioned reinforcement learning · 0.5neural encoder · 0.4mutual information maximization · 0.4contrastive learning · 0.4
YearPublicationVenuePosition
2024 Many-Shot In-Context Learning
abstract
Large language models (LLMs) excel at few-shot in-context learning (ICL) -- learning from a few examples provided in context at inference, without any weight updates. Newly expanded context windows allow us to investigate ICL with hundreds or thousands of examples – the many-shot regime. Going from few-shot to many-shot, we observe significant performance gains across a wide variety of generative and discriminative tasks. While promising, many-shot ICL can be bottlenecked by the available amount of human-generated outputs. To mitigate this limitation, we explore two new settings: (1) "Reinforced ICL" that uses model-generated chain-of-thought rationales in place of human rationales, and (2) "Unsupervised ICL" where we remove rationales from the prompt altogether, and prompts the model only with domain-specific inputs. We find that both Reinforced and Unsupervised ICL can be quite effective in the many-shot regime, particularly on complex reasoning tasks. We demonstrate that, unlike few-shot learning, many-shot learning is effective at overriding pretraining biases, can learn high-dimensional functions with numerical inputs, and performs comparably to supervised fine-tuning. Finally, we reveal the limitations of next-token prediction loss as an indicator of downstream ICL performance.
Rishabh Agarwal, Avi Singh, Bernd Bohnet, Luis Rosias, Stephanie C. Y. Chan, Ankesh Anand, Zaheer Abbas, Azade Nova, John D. Co-Reyes, Eric Chu, Feryal M. P. Behbahani, Aleksandra Faust, Hugo Larochelle
NeurIPS8
2023 Investigating the Role of Model-Based Learning in Exploration and Transfer
abstract
State of the art reinforcement learning has enabled training agents on tasks of ever increasing complexity. However, the current paradigm tends to favor training agents from scratch on every new task or on collections of tasks with a view towards generalizing to novel task configurations. The former suffers from poor data efficiency while the latter is difficult when test tasks are out-of-distribution. Agents that can effectively transfer their knowledge about the world pose a potential solution to these issues. In this paper, we investigate transfer learning in the context of model-based agents. Specifically, we aim to understand where exactly environment models have an advantage and why. We find that a model-based approach outperforms controlled model-free baselines for transfer learning. Through ablations, we show that both the policy and dynamics model learnt through exploration matter for successful transfer. We demonstrate our results across three domains which vary in their requirements for transfer: in-distribution procedural (Crafter), in-distribution identical (RoboDesk), and out-of-distribution (Meta-World). Our results show that intrinsic exploration combined with environment models present a viable direction towards agents that are self-supervised and able to generalize to novel reward functions.
Jacob C. Walker, Eszter Vértes, Yazhe Li, Gabriel Dulac-Arnold, Ankesh Anand, Theophane Weber, Jessica B. Hamrick
ICML5
2022 Procedural generalization by planning with self-supervised world models
Ankesh Anand, Jacob C. Walker, Yazhe Li, Eszter Vértes, Julian Schrittwieser, Sherjil Ozair, Theophane Weber, Jessica B. Hamrick
ICLR1
2021 Data-Efficient Reinforcement Learning with Self-Predictive Representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm, Aaron C. Courville, Philip Bachman
ICLR2
2021 Pretraining Representations for Data-Efficient Reinforcement Learning
abstract
Data efficiency is a key challenge for deep reinforcement learning. We address this problem by using unlabeled data to pretrain an encoder which is then finetuned on a small amount of task-specific data. To encourage learning representations which capture diverse aspects of the underlying MDP, we employ a combination of latent dynamics modelling and unsupervised goal-conditioned RL. When limited to 100k steps of interaction on Atari games (equivalent to two hours of human experience), our approach significantly surpasses prior work combining offline representation pretraining with task-specific finetuning, and compares favourably with other pretraining methods that require orders of magnitude more data. Our approach shows particular promise when combined with larger models as well as more diverse, task-aligned observational data -- approaching human-level performance and data-efficiency on Atari in our best setting.
Max Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand, Laurent Charlin, R. Devon Hjelm, Philip Bachman, Aaron C. Courville
NeurIPS4
2019 Unsupervised State Representation Learning in Atari
abstract
State representation learning, or the ability to capture latent generative factors of an environment is crucial for building intelligent agents that can perform a wide variety of tasks. Learning such representations in an unsupervised manner without supervision from rewards is an open problem. We introduce a method that tries to learn better state representations by maximizing mutual information across spatially and temporally distinct features of a neural encoder of the observations. We also introduce a new benchmark based on Atari 2600 games where we evaluate representations based on how well they capture the ground truth state. We believe this new framework for evaluating representation learning models will be crucial for future representation learning research. Finally, we compare our technique with other state-of-the-art generative and contrastive representation learning methods.
Ankesh Anand, Evan Racah, Sherjil Ozair, Yoshua Bengio, Marc-Alexandre Côté, R. Devon Hjelm
NeurIPS1
2018 Phishing URL Detection with Oversampling based on Text Generative Adversarial Networks
abstract
The problem of imbalanced classes arises frequently in binary classification tasks. If one class outnumbers another, trained classifiers become heavily biased towards the majority class. For phishing URL detection, it is very natural that the number of collected benign URLs (i.e., the majority class) is much larger than the number of collected phishy URLs (i.e., the minority class). Oversampling the minority class can be a powerful tool to overcome this situation. However, existing methods perform the oversampling task in the feature space where the original data format is removed and URLs are succinctly represented by vectors. These methods are successful only if feature definitions are correct and the dataset is diverse and not too sparse. In this paper, we propose an oversampling technique in the data space. We train text generative adversarial networks (text-GANs) with URLs in the minority class and generate synthetic URLs that can be made part of the training set. We crawl a crowd-sourced URL repository to collect recently discovered phishy and benign URLs. Our experiments demonstrate significant performance improvements after using the proposed oversampling technique. Interestingly, some of the original test URLs are exactly regenerated by the proposed text generative model.
Ankesh Anand, Kshitij Gorde, Joel Ruben Antony Moniz, Noseong Park, Tanmoy Chakraborty 0002, Bei-tseng Chu
IEEE BigData1
2018 MMGAN: Manifold-Matching Generative Adversarial Networks
abstract
It is well-known that GANs are difficult to train, and several different techniques have been proposed in order to stabilize their training. In this paper, we propose a novel training method called manifold-matching, and a new GAN model called manifold-matching GAN (MMGAN). MMGAN finds two manifolds representing the vector representations of real and fake images. If these two manifolds match, it means that real and fake images are statistically identical. To assist the manifold-matching task, we also use i) kernel tricks to find better manifold structures, ii) moving-averaged manifolds across mini-batches, and iii) a regularizer based on correlation matrix to suppress mode collapse. We conduct in-depth experiments with three image datasets and compare with several state-of-the-art GAN models. 32.4% of images generated by the proposed MMGAN are recognized as fake images during our user study (16% enhancement compared to other state-of-the-art model). MMGAN achieved an unsupervised inception score of 7.8 for CIFAR-10.
Noseong Park, Ankesh Anand, Joel Ruben Antony Moniz, Kookjin Lee, Jaegul Choo, David Keetae Park, Tanmoy Chakraborty 0002, Hongkyu Park
ICPR2
2017 FairScholar: Balancing Relevance and Diversity for Scientific Paper Recommendation
Ankesh Anand, Tanmoy Chakraborty 0002, Amitava Das 0001
ECIR1
2017 We Used Neural Networks to Detect Clickbaits: You Won't Believe What Happened Next!
Ankesh Anand, Tanmoy Chakraborty 0002, Noseong Park
ECIR1