EDBT 2026 Demo / reviewers in the wild / expert
Vivek Veeriah
dblp:162/0205
· DBLP profile ↗
12ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-0728-4953ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Reinforcement learning · 70% Optimization for machine learning · 9% Generative modeling · 7% | |
| Human-computer interaction and pervasive computing
1 paper |
Games and playful interaction · 100% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
auxiliary tasks |
0.9 | 2 | 2021 | Learning State Representations from Random Deep Action-conditional Predictions · NeurIPS 2021 Discovery of Useful Questions as Auxiliary Tasks · NeurIPS 2019 |
Machine learning › Generative modeling › computational creativity
creative generation |
0.9 | 1 | 2025 | Generating Creative Chess Puzzles · NeurIPS 2025 |
Machine learning › Reinforcement learning › meta-reinforcement learning
meta-gradient reinforcement learning |
0.9 | 2 | 2020 | A Self-Tuning Actor-Critic Algorithm · NeurIPS 2020 How Should an Agent Practice? · AAAI 2020 |
Machine learning › Reinforcement learning
reward design |
0.9 | 1 | 2025 | Generating Creative Chess Puzzles · NeurIPS 2025 |
Machine learning › Reinforcement learning › value function estimation
general value functions |
0.8 | 2 | 2020 | Learning Retrospective Knowledge with Reverse Reinforcement Learning · NeurIPS 2020 Discovery of Useful Questions as Auxiliary Tasks · NeurIPS 2019 |
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process |
0.7 | 1 | 2023 | ReLOAD: Reinforcement Learning with Optimistic Ascent-Descent for Last-Iterate Convergence in Constrained MDPs · ICML 2023 |
Machine learning › Reinforcement learning
constrained reinforcement learning |
0.7 | 1 | 2023 | ReLOAD: Reinforcement Learning with Optimistic Ascent-Descent for Last-Iterate Convergence in Constrained MDPs · ICML 2023 |
Machine learning › Optimization for machine learning › convergence guarantees
last-iterate convergence |
0.7 | 1 | 2023 | ReLOAD: Reinforcement Learning with Optimistic Ascent-Descent for Last-Iterate Convergence in Constrained MDPs · ICML 2023 |
Machine learning › Reinforcement learning
actor-critic methods |
0.6 | 2 | 2021 | A Self-Tuning Actor-Critic Algorithm · NeurIPS 2020 Learning State Representations from Random Deep Action-conditional Predictions · NeurIPS 2021 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.5 | 1 | 2021 | Discovery of Options via Meta-Learned Subgoals · NeurIPS 2021 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery |
0.5 | 1 | 2021 | Discovery of Options via Meta-Learned Subgoals · NeurIPS 2021 |
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning |
0.5 | 1 | 2021 | Learning State Representations from Random Deep Action-conditional Predictions · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning |
0.5 | 1 | 2021 | Learning State Representations from Random Deep Action-conditional Predictions · NeurIPS 2021 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.4 | 1 | 2020 | A Self-Tuning Actor-Critic Algorithm · NeurIPS 2020 |
Machine learning › Reinforcement learning › reward learning
intrinsic reward learning |
0.4 | 1 | 2020 | How Should an Agent Practice? · AAAI 2020 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.4 | 2 | 2015 | Deep Learning Architecture with Dynamically Programmed Layers for Brain Connectome Prediction · KDD 2015 Differential Recurrent Neural Networks for Action Recognition · ICCV 2015 |
Machine learning › Reinforcement learning
value function |
0.4 | 1 | 2020 | Learning Retrospective Knowledge with Reverse Reinforcement Learning · NeurIPS 2020 |
Machine learning › Reinforcement learning
value function estimation |
0.4 | 1 | 2019 | Discovery of Useful Questions as Auxiliary Tasks · NeurIPS 2019 |
Games and playful interaction › board games
chess |
0.3 | 1 | 2025 | Generating Creative Chess Puzzles · NeurIPS 2025 |
Computer vision › Video understanding and tracking
action recognition |
0.2 | 1 | 2015 | Differential Recurrent Neural Networks for Action Recognition · ICCV 2015 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.2 | 1 | 2015 | Deep Learning Architecture with Dynamically Programmed Layers for Brain Connectome Prediction · KDD 2015 |
Machine learning › Time series and sequential data
anomaly detection |
0.1 | 1 | 2020 | Learning Retrospective Knowledge with Reverse Reinforcement Learning · NeurIPS 2020 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.7chess engine search statistics · 1.7meta-gradient · 1.3optimistic ascent-descent · 0.7gradient descent ascent · 0.7manager-worker decomposition · 0.5general value functions · 0.5action-conditional prediction · 0.5leaky v-trace · 0.4intrinsic reward · 0.4dynamically programmed layer · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AugInsert: Learning Robust Visual-Force Policies via Data Augmentation for Object Assembly TasksabstractOperating in unstructured environments like households requires robotic policies that are robust to out-of-distribution conditions. Although much work has been done in evaluating robustness for visuomotor policies, the robustness evaluation of a multisensory approach that includes force-torque sensing remains largely unexplored. This work introduces a novel, factor-based evaluation framework with the goal of assessing the robustness of multisensory policies in a peg-in-hole assembly task. To this end, we develop a multisensory policy framework utilizing the Perceiver IO architecture to learn the task. We investigate which factors pose the greatest generalization challenges in object assembly and explore a simple multisensory data augmentation technique to enhance out-of-distribution performance. We provide a simulation environment enabling controlled evaluation of these factors. Our results reveal that multisensory variations such as Grasp Pose present the most significant challenges for robustness, and naive unisensory data augmentation applied independently to each sensory modality proves insufficient to overcome them. Additionally, we find force-torque sensing to be the most informative modality for our contact-rich assembly task, with vision being the least informative. Finally, we briefly discuss supporting real-world experimental results. For additional experiments and qualitative results, we refer to the project webpage https://rpm-lab-umn.github.io/auginsert/. Ryan Diaz, Adam Imdieke, Vivek Veeriah, Karthik Desingh |
IROS | 3 |
| 2025 | Generating Creative Chess PuzzlesabstractWhile Generative AI rapidly advances in various domains, generating truly creative, aesthetic, and counter-intuitive outputs remains a challenge. This paper presents an approach to tackle these difficulties in the domain of chess puzzles. We start by benchmarking Generative AI architectures, and then introduce an RL framework with novel rewards based on chess engine search statistics to overcome some of those shortcomings. The rewards are designed to enhance a puzzle's uniqueness, counter-intuitiveness, diversity, and realism. Our RL approach dramatically increases counter-intuitive puzzle generation by 10x, from 0.22\% (supervised) to 2.5\%, surpassing existing dataset rates (2.1\%) and the best Lichess-trained model (0.4\%). Our puzzles meet novelty and diversity benchmarks, retain aesthetic themes, and are rated by human experts as more creative, enjoyable, and counter-intuitive than composed book puzzles, even approaching classic compositions. Our final outcome is a curated booklet of these novel AI-generated puzzles, which is acknowledged for creativity by three world-renowned experts. Xidong Feng, Vivek Veeriah, Marcus Chiam, Michael Dennis 0001, Federico Barbero, Johan S. Obando-Ceron, Jiaxin Shi, Satinder Singh 0001, Shaobo Hou, Nenad Tomasev, Tom Zahavy |
NeurIPS | 2 |
| 2023 | ReLOAD: Reinforcement Learning with Optimistic Ascent-Descent for Last-Iterate Convergence in Constrained MDPsabstractIn recent years, reinforcement learning (RL) has been applied to real-world problems with increasing success. Such applications often require to put constraints on the agent’s behavior. Existing algorithms for constrained RL (CRL) rely on gradient descent-ascent, but this approach comes with a caveat. While these algorithms are guaranteed to converge on average, they do not guarantee last-iterate convergence, i.e., the current policy of the agent may never converge to the optimal solution. In practice, it is often observed that the policy alternates between satisfying the constraints and maximizing the reward, rarely accomplishing both objectives simultaneously. Here, we address this problem by introducing Reinforcement Learning with Optimistic Ascent-Descent (ReLOAD), a principled CRL method with guaranteed last-iterate convergence. We demonstrate its empirical effectiveness on a wide variety of CRL problems including discrete MDPs and continuous control. In the process we establish a benchmark of challenging CRL problems. Ted Moskovitz, Brendan O'Donoghue, Vivek Veeriah, Sebastian Flennerhag, Satinder Singh 0001, Tom Zahavy |
ICML | 3 |
| 2021 | Discovery of Options via Meta-Learned SubgoalsabstractTemporal abstractions in the form of options have been shown to help reinforcement learning (RL) agents learn faster. However, despite prior work on this topic, the problem of discovering options through interaction with an environment remains a challenge. In this paper, we introduce a novel meta-gradient approach for discovering useful options in multi-task RL environments. Our approach is based on a manager-worker decomposition of the RL agent, in which a manager maximises rewards from the environment by learning a task-dependent policy over both a set of task-independent discovered-options and primitive actions. The option-reward and termination functions that define a subgoal for each option are parameterised as neural networks and trained via meta-gradients to maximise their usefulness. Empirical analysis on gridworld and DeepMind Lab tasks show that: (1) our approach can discover meaningful and diverse temporally-extended options in multi-task RL domains, (2) the discovered options are frequently used by the agent while learning to solve the training tasks, and (3) that the discovered options help a randomly initialised manager learn faster in completely new tasks. Vivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu, Junhyuk Oh, Iurii Kemaev, Hado van Hasselt, David Silver 0001, Satinder Singh 0001 |
NeurIPS | 1 |
| 2021 | Learning State Representations from Random Deep Action-conditional PredictionsabstractOur main contribution in this work is an empirical finding that random General Value Functions (GVFs), i.e., deep action-conditional predictions---random both in what feature of observations they predict as well as in the sequence of actions the predictions are conditioned upon---form good auxiliary tasks for reinforcement learning (RL) problems. In particular, we show that random deep action-conditional predictions when used as auxiliary tasks yield state representations that produce control performance competitive with state-of-the-art hand-crafted auxiliary tasks like value prediction, pixel control, and CURL in both Atari and DeepMind Lab tasks. In another set of experiments we stop the gradients from the RL part of the network to the state representation learning part of the network and show, perhaps surprisingly, that the auxiliary tasks alone are sufficient to learn state representations good enough to outperform an end-to-end trained actor-critic baseline. We opensourced our code at https://github.com/Hwhitetooth/random_gvfs. Vivek Veeriah, Risto Vuorio, Richard L. Lewis, Satinder Singh 0001 |
NeurIPS | 2 |
| 2020 | How Should an Agent Practice?abstractWe present a method for learning intrinsic reward functions to drive the learning of an agent during periods of practice in which extrinsic task rewards are not available. During practice, the environment may differ from the one available for training and evaluation with extrinsic rewards. We refer to this setup of alternating periods of practice and objective evaluation as practice-match, drawing an analogy to regimes of skill acquisition common for humans in sports and games. The agent must effectively use periods in the practice environment so that performance improves during matches. In the proposed method the intrinsic practice reward is learned through a meta-gradient approach that adapts the practice reward parameters to reduce the extrinsic match reward loss computed from matches. We illustrate the method on a simple grid world, and evaluate it in two games in which the practice environment differs from match: Pong with practice against a wall without an opponent, and PacMan with practice in a maze without ghosts. The results show gains from learning in practice in addition to match periods over learning in matches only. Janarthanan Rajendran, Richard L. Lewis, Vivek Veeriah, Honglak Lee, Satinder Singh 0001 |
AAAI | 3 |
| 2020 | A Self-Tuning Actor-Critic AlgorithmabstractReinforcement learning algorithms are highly sensitive to the choice of hyperparameters, typically requiring significant manual effort to identify hyperparameters that perform well on a new domain. In this paper, we take a step towards addressing this issue by using metagradients to automatically adapt hyperparameters online by meta-gradient descent (Xu et al., 2018). We apply our algorithm, Self-Tuning Actor-Critic (STAC), to self-tune all the differentiable hyperparameters of an actor-critic loss function, to discover auxiliary tasks, and to improve off-policy learning using a novel leaky V-trace operator. STAC is simple to use, sample efficient and does not require a significant increase in compute. Ablative studies show that the overall performance of STAC improved as we adapt more hyperparameters. When applied to the Arcade Learning Environment (Bellemare et al. 2012), STAC improved the median human normalized score in 200M steps from 243% to 364%. When applied to the DM Control suite (Tassa et al., 2018), STAC improved the mean score in 30M steps from 217 to 389 when learning with features, from 108 to 202 when learning from pixels, and from 195 to 295 in the Real-World Reinforcement Learning Challenge (Dulac-Arnold et al., 2020). Tom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel, Junhyuk Oh, Hado van Hasselt, David Silver 0001, Satinder Singh 0001 |
NeurIPS | 3 |
| 2020 | Learning Retrospective Knowledge with Reverse Reinforcement LearningabstractWe present a Reverse Reinforcement Learning (Reverse RL) approach for representing retrospective knowledge. General Value Functions (GVFs) have enjoyed great success in representing predictive knowledge, i.e., answering questions about possible future outcomes such as “how much fuel will be consumed in expectation if we drive from A to B?”. GVFs, however, cannot answer questions like “how much fuel do we expect a car to have given it is at B at time t?”. To answer this question, we need to know when that car had a full tank and how that car came to B. Since such questions emphasize the influence of possible past events on the present, we refer to their answers as retrospective knowledge. In this paper, we show how to represent retrospective knowledge with Reverse GVFs, which are trained via Reverse RL. We demonstrate empirically the utility of Reverse GVFs in both representation learning and anomaly detection. Shangtong Zhang, Vivek Veeriah, Shimon Whiteson |
NeurIPS | 2 |
| 2019 | Discovery of Useful Questions as Auxiliary TasksabstractArguably, intelligent agents ought to be able to discover their own questions so that in learning answers for them they learn unanticipated useful knowledge and skills; this departs from the focus in much of machine learning on agents learning answers to externally defined questions. We present a novel method for a reinforcement learning (RL) agent to discover questions formulated as general value functions or GVFs, a fairly rich form of knowledge representation. Specifically, our method uses non-myopic meta-gradients to learn GVF-questions such that learning answers to them, as an auxiliary task, induces useful representations for the main task faced by the RL agent. We demonstrate that auxiliary tasks based on the discovered GVFs are sufficient, on their own, to build representations that support main task learning, and that they do so better than popular hand-designed auxiliary tasks from the literature. Furthermore, we show, in the context of Atari2600 videogames, how such auxiliary tasks, meta-learned alongside the main task, can improve the data efficiency of an actor-critic agent. Vivek Veeriah, Matteo Hessel, Zhongwen Xu, Janarthanan Rajendran, Richard L. Lewis, Junhyuk Oh, Hado van Hasselt, David Silver 0001, Satinder Singh 0001 |
NeurIPS | 1 |
| 2017 | Crossprop: Learning Representations by Stochastic Meta-Gradient Descent in Neural Networks
Vivek Veeriah, Shangtong Zhang, Richard S. Sutton |
ECML/PKDD (1) | 1 |
| 2015 | Differential Recurrent Neural Networks for Action RecognitionabstractThe long short-term memory (LSTM) neural network is capable of processing complex sequential information since it utilizes special gating schemes for learning representations from long input sequences. It has the potential to model any time-series or sequential data, where the current hidden state has to be considered in the context of the past hidden states. This property makes LSTM an ideal choice to learn the complex dynamics of various actions. Unfortunately, the conventional LSTMs do not consider the impact of spatio-temporal dynamics corresponding to the given salient motion patterns, when they gate the information that ought to be memorized through time. To address this problem, we propose a differential gating scheme for the LSTM neural network, which emphasizes on the change in information gain caused by the salient motions between the successive frames. This change in information gain is quantified by Derivative of States (DoS), and thus the proposed LSTM model is termed as differential Recurrent Neural Network (dRNN). We demonstrate the effectiveness of the proposed model by automatically recognizing actions from the real-world 2D and 3D human action datasets. Our study is one of the first works towards demonstrating the potential of learning complex time-series representations via high-order derivatives of states. Vivek Veeriah, Naifan Zhuang, Guo-Jun Qi |
ICCV | 1 |
| 2015 | Deep Learning Architecture with Dynamically Programmed Layers for Brain Connectome PredictionabstractThis paper explores the idea of using deep neural network architecture with dynamically programmed layers for brain connectome prediction problem. Understanding the brain connectome structure is a very interesting and a challenging problem. It is critical in the research for epilepsy and other neuropathological diseases. We introduce a new deep learning architecture that exploits the spatial and temporal nature of the neuronal activation data. The architecture consists of a combination of Convolutional layer and a Recurrent layer for predicting the connectome of neurons based on their time-series of activation data. The key contribution of this paper is a dynamically programmed layer that is critical in determining the alignment between the neuronal activations of pair-wise combinations of neurons. Vivek Veeriah, Rohit Durvasula, Guo-Jun Qi |
KDD | 1 |