EDBT 2026 Demo / reviewers in the wild / expert
Timothy A. Mann
dblp:217/3322 · also Timothy Arthur Mann
· DBLP profile ↗
26ranked-venue papers
6as first author
5since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
21 papers |
Reinforcement learning · 43% Trustworthy machine learning · 39% Generative modeling · 8% | |
| Databases, data mining, and information retrieval
3 papers |
Recommender systems · 100% |
Topics — the 30 heaviest of 59, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
robustness |
2.2 | 5 | 2021 | Data Augmentation Can Improve Robustness · NeurIPS 2021 Self-supervised Adversarial Robustness for the Low-label, High-data Regime · ICLR 2021 Achieving Robustness in the Wild via Adversarial Mixing With Disentangled Representations · CVPR 2020 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
2.1 | 4 | 2022 | Defending Against Image Corruptions Through Adversarial Augmentations · ICLR 2022 Data Augmentation Can Improve Robustness · NeurIPS 2021 Improving Robustness using Generated Data · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
1.3 | 3 | 2021 | Data Augmentation Can Improve Robustness · NeurIPS 2021 Achieving Robustness in the Wild via Adversarial Mixing With Disentangled Representations · CVPR 2020 Scalable Verified Training for Provably Robust Image Classification · ICCV 2019 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
1.0 | 4 | 2018 | Learning Robust Options · AAAI 2018 Adaptive Skills Adaptive Partitions (ASAP) · NIPS 2016 Time-Regularized Interrupting Options (TRIO) · ICML 2014 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery |
0.8 | 3 | 2018 | Learning Robust Options · AAAI 2018 Adaptive Skills Adaptive Partitions (ASAP) · NIPS 2016 Time-Regularized Interrupting Options (TRIO) · ICML 2014 |
Machine learning › Reinforcement learning
robust reinforcement learning |
0.8 | 2 | 2020 | Robust Reinforcement Learning for Continuous Control with Model Misspecification · ICLR 2020 Learning Robust Options · AAAI 2018 |
Machine learning › Trustworthy machine learning › adversarial machine learning
adversarial data augmentation |
0.6 | 1 | 2022 | Defending Against Image Corruptions Through Adversarial Augmentations · ICLR 2022 |
Machine learning › Trustworthy machine learning › robustness › corruption robustness
image corruption robustness |
0.6 | 1 | 2022 | Defending Against Image Corruptions Through Adversarial Augmentations · ICLR 2022 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
options framework |
0.5 | 2 | 2017 | Approximate Value Iteration with Temporally Extended Actions (Extended Abstract) · IJCAI 2017 Adaptive Skills Adaptive Partitions (ASAP) · NIPS 2016 |
Machine learning › Reinforcement learning
constrained reinforcement learning |
0.5 | 1 | 2021 | Balancing Constraints and Rewards with Meta-Gradient D4PG · ICLR 2021 |
Machine learning › Reinforcement learning › large-scale reinforcement learning
distributed reinforcement learning |
0.5 | 1 | 2021 | Balancing Constraints and Rewards with Meta-Gradient D4PG · ICLR 2021 |
Machine learning › Deep learning architectures and training › data augmentation
generative data augmentation |
0.5 | 1 | 2021 | Improving Robustness using Generated Data · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › robustness
robust learning |
0.5 | 1 | 2021 | Improving Robustness using Generated Data · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
robust overfitting |
0.5 | 1 | 2021 | Data Augmentation Can Improve Robustness · NeurIPS 2021 |
Machine learning › Generative modeling
synthetic training data |
0.5 | 1 | 2021 | Improving Robustness using Generated Data · NeurIPS 2021 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
temporally extended actions |
0.5 | 2 | 2017 | Approximate Value Iteration with Temporally Extended Actions (Extended Abstract) · IJCAI 2017 Scaling Up Approximate Value Iteration with Options: Better Policies with Fewer Iterations · ICML 2014 |
Machine learning › Reinforcement learning
bandit |
0.4 | 1 | 2020 | Non-Stationary Delayed Bandits with Intermediate Observations · ICML 2020 |
Machine learning › Reinforcement learning
continuous control |
0.4 | 1 | 2020 | Robust Reinforcement Learning for Continuous Control with Model Misspecification · ICLR 2020 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.4 | 1 | 2020 | Achieving Robustness in the Wild via Adversarial Mixing With Disentangled Representations · CVPR 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2020 | Achieving Robustness in the Wild via Adversarial Mixing With Disentangled Representations · CVPR 2020 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
model misspecification |
0.4 | 1 | 2020 | Robust Reinforcement Learning for Continuous Control with Model Misspecification · ICLR 2020 |
Machine learning › Reinforcement learning › multi-armed bandit
non-stationary bandits |
0.4 | 1 | 2020 | Non-Stationary Delayed Bandits with Intermediate Observations · ICML 2020 |
Machine learning › Reinforcement learning
regret minimization |
0.4 | 1 | 2020 | Non-Stationary Delayed Bandits with Intermediate Observations · ICML 2020 |
Machine learning › Generative modeling › generative adversarial network
StyleGAN |
0.4 | 1 | 2020 | Achieving Robustness in the Wild via Adversarial Mixing With Disentangled Representations · CVPR 2020 |
Machine learning › Reinforcement learning › regret minimization
sublinear regret |
0.4 | 1 | 2020 | Non-Stationary Delayed Bandits with Intermediate Observations · ICML 2020 |
Mathematical optimization
linear programming |
0.4 | 1 | 2020 | The NodeHopper: Enabling Low Latency Ranking with Constraints via a Fast Dual Solver · KDD 2020 |
Machine learning › Trustworthy machine learning › robustness › certified robustness
certified adversarial robustness |
0.4 | 1 | 2019 | A Dual Approach to Verify and Train Deep Networks · IJCAI 2019 |
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder |
0.4 | 1 | 2019 | Beyond Greedy Ranking: Slate Optimization via List-CVAE · ICLR (Poster) 2019 |
Machine learning › Trustworthy machine learning › verification
formal verification of neural networks |
0.4 | 1 | 2019 | A Dual Approach to Verify and Train Deep Networks · IJCAI 2019 |
Machine learning › Trustworthy machine learning › robustness › certified robustness
interval bound propagation |
0.4 | 1 | 2019 | Scalable Verified Training for Provably Robust Image Classification · ICCV 2019 |
Methods — techniques the papers use, named apart from their topics
adversarial training · 1.8data augmentation · 1.1linear programming · 0.9dual optimization · 0.9self-supervised learning · 0.5model weight averaging · 0.5meta-gradient · 0.5generative model · 0.5gaussian sampling · 0.5distributed distributional deterministic policy gradients · 0.5disentangled representation · 0.4adversarial mixing · 0.4neural network · 0.4list-CVAE · 0.4factorization · 0.4regret analysis · 0.2distribution-norm · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Defending Against Image Corruptions Through Adversarial Augmentations
Dan Andrei Calian, Florian Stimberg, Olivia Wiles, Sylvestre-Alvise Rebuffi, András György 0001, Timothy A. Mann, Sven Gowal |
ICLR | 6 |
| 2021 | Balancing Constraints and Rewards with Meta-Gradient D4PG
Dan Andrei Calian, Daniel J. Mankowitz, Tom Zahavy, Zhongwen Xu, Junhyuk Oh, Nir Levine, Timothy A. Mann |
ICLR | 7 |
| 2021 | Self-supervised Adversarial Robustness for the Low-label, High-data Regime
Sven Gowal, Po-Sen Huang, Aäron van den Oord, Timothy A. Mann, Pushmeet Kohli |
ICLR | 4 |
| 2021 | Improving Robustness using Generated DataabstractRecent work argues that robust training requires substantially larger datasets than those required for standard classification. On CIFAR-10 and CIFAR-100, this translates into a sizable robust-accuracy gap between models trained solely on data from the original training set and those trained with additional data extracted from the "80 Million Tiny Images" dataset (TI-80M). In this paper, we explore how generative models trained solely on the original training set can be leveraged to artificially increase the size of the original training set and improve adversarial robustness to $\ell_p$ norm-bounded perturbations. We identify the sufficient conditions under which incorporating additional generated data can improve robustness, and demonstrate that it is possible to significantly reduce the robust-accuracy gap to models trained with additional real data. Surprisingly, we even show that even the addition of non-realistic random data (generated by Gaussian sampling) can improve robustness. We evaluate our approach on CIFAR-10, CIFAR-100, SVHN and TinyImageNet against $\ell_\infty$ and $\ell_2$ norm-bounded perturbations of size $\epsilon = 8/255$ and $\epsilon = 128/255$, respectively. We show large absolute improvements in robust accuracy compared to previous state-of-the-art methods. Against $\ell_\infty$ norm-bounded perturbations of size $\epsilon = 8/255$, our models achieve 66.10% and 33.49% robust accuracy on CIFAR-10 and CIFAR-100, respectively (improving upon the state-of-the-art by +8.96% and +3.29%). Against $\ell_2$ norm-bounded perturbations of size $\epsilon = 128/255$, our model achieves 78.31% on CIFAR-10 (+3.81%). These results beat most prior works that use external data. Sven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg, Dan Andrei Calian, Timothy A. Mann |
NeurIPS | 6 |
| 2021 | Data Augmentation Can Improve RobustnessabstractAdversarial training suffers from robust overfitting, a phenomenon where the robust test accuracy starts to decrease during training. In this paper, we focus on reducing robust overfitting by using common data augmentation schemes. We demonstrate that, contrary to previous findings, when combined with model weight averaging, data augmentation can significantly boost robust accuracy. Furthermore, we compare various augmentations techniques and observe that spatial composition techniques work the best for adversarial training. Finally, we evaluate our approach on CIFAR-10 against $\ell_\infty$ and $\ell_2$ norm-bounded perturbations of size $\epsilon = 8/255$ and $\epsilon = 128/255$, respectively. We show large absolute improvements of +2.93% and +2.16% in robust accuracy compared to previous state-of-the-art methods. In particular, against $\ell_\infty$ norm-bounded perturbations of size $\epsilon = 8/255$, our model reaches 60.07% robust accuracy without using any external data. We also achieve a significant performance boost with this approach while using other architectures and datasets such as CIFAR-100, SVHN and TinyImageNet. Sylvestre-Alvise Rebuffi, Sven Gowal, Dan Andrei Calian, Florian Stimberg, Olivia Wiles, Timothy A. Mann |
NeurIPS | 6 |
| 2020 | Achieving Robustness in the Wild via Adversarial Mixing With Disentangled RepresentationsabstractRecent research has made the surprising finding that state-of-the-art deep learning models sometimes fail to generalize to small variations of the input. Adversarial training has been shown to be an effective approach to overcome this problem. However, its application has been limited to enforcing invariance to analytically defined transformations like lp-norm bounded perturbations. Such perturbations do not necessarily cover plausible real-world variations that preserve the semantics of the input (such as a change in lighting conditions). In this paper, we propose a novel approach to express and formalize robustness to these kinds of real-world transformations of the input. The two key ideas underlying our formulation are (1) leveraging disentangled representations of the input to define different factors of variations, and (2) generating new input images by adversarially composing the representations of different images. We use a StyleGAN model to demonstrate the efficacy of this framework. Specifically, we leverage the disentangled latent representations computed by a StyleGAN model to generate perturbations of an image that are similar to real-world variations (like adding make-up, or changing the skin-tone of a person) and train models to be invariant to these perturbations. Extensive experiments show that our method improves generalization and reduces the effect of spurious correlations (reducing the error rate of a "smile" detector by 21% for example). Sven Gowal, Chongli Qin, Po-Sen Huang, A. Taylan Cemgil, Krishnamurthy Dvijotham, Timothy A. Mann, Pushmeet Kohli |
CVPR | 6 |
| 2020 | Robust Reinforcement Learning for Continuous Control with Model Misspecification
Daniel J. Mankowitz, Nir Levine, Rae Jeong, Abbas Abdolmaleki, Jost Tobias Springenberg, Jackie Kay, Todd Hester, Timothy A. Mann, Martin A. Riedmiller |
ICLR | 9 |
| 2020 | Non-Stationary Delayed Bandits with Intermediate ObservationsabstractOnline recommender systems often face long delays in receiving feedback, especially when optimizing for some long-term metrics. While mitigating the effects of delays in learning is well-understood in stationary environments, the problem becomes much more challenging when the environment changes. In fact, if the timescale of the change is comparable to the delay, it is impossible to learn about the environment, since the available observations are already obsolete. However, the arising issues can be addressed if intermediate signals are available without delay, such that given those signals, the long-term behavior of the system is stationary. To model this situation, we introduce the problem of stochastic, non-stationary, delayed bandits with intermediate observations. We develop a computationally efficient algorithm based on UCRL, and prove sublinear regret guarantees for its performance. Experimental results demonstrate that our method is able to learn in non-stationary delayed environments where existing methods fail. Claire Vernade, András György 0001, Timothy A. Mann |
ICML | 3 |
| 2020 | The NodeHopper: Enabling Low Latency Ranking with Constraints via a Fast Dual SolverabstractModern recommender systems need to deal with multiple objectives like balancing user engagement with recommending diverse and fresh content. An appealing way to optimally trade these off is by imposing constraints on the ranking according to which items are presented to a user. This results in a constrained ranking optimization problem that can be solved as a linear program (LP). However, off-the-shelf LP solvers are unable to meet the severe latency constraints in systems that serve live traffic. To address this challenge, we exploit the structure of the dual optimization problem to develop a fast solver. We analyze theoretical properties of our solver and show experimentally that it is able to solve constrained ranking problems on synthetic and real-world recommendation datasets an order of magnitude faster than off-the-shelf solvers, thereby enabling their deployment under severe latency constraints. Anton Zhernov, Krishnamurthy Dvijotham, Ivan Lobov, Dan Andrei Calian, Michelle X. Gong, Natarajan Chandrashekar, Timothy A. Mann |
KDD | 7 |
| 2019 | Scalable Verified Training for Provably Robust Image ClassificationabstractRecent work has shown that it is possible to train deep neural networks that are provably robust to norm-bounded adversarial perturbations. Most of these methods are based on minimizing an upper bound on the worst-case loss over all possible adversarial perturbations. While these techniques show promise, they often result in difficult optimization procedures that remain hard to scale to larger networks. Through a comprehensive analysis, we show how a simple bounding technique, interval bound propagation (IBP), can be exploited to train large provably robust neural networks that beat the state-of-the-art in verified accuracy. While the upper bound computed by IBP can be quite weak for general networks, we demonstrate that an appropriate loss and clever hyper-parameter schedule allow the network to adapt such that the IBP bound is tight. This results in a fast and stable learning algorithm that outperforms more sophisticated methods and achieves state-of-the-art results on MNIST, CIFAR-10 and SVHN. It also allows us to train the largest model to be verified beyond vacuous bounds on a downscaled version of IMAGENET. Sven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel, Chongli Qin, Jonathan Uesato, Relja Arandjelovic, Timothy A. Mann, Pushmeet Kohli |
ICCV | 8 |
| 2019 | Beyond Greedy Ranking: Slate Optimization via List-CVAE
Ray Jiang, Sven Gowal, Yuqiu Qian, Timothy A. Mann, Danilo Jimenez Rezende |
ICLR (Poster) | 4 |
| 2019 | Learning from Delayed Outcomes via Proxies with Applications to Recommender SystemsabstractPredicting delayed outcomes is an important problem in recommender systems (e.g., if customers will finish reading an ebook). We formalize the problem as an adversarial, delayed online learning problem and consider how a proxy for the delayed outcome (e.g., if customers read a third of the book in 24 hours) can help minimize regret, even though the proxy is not available when making a prediction. Motivated by our regret analysis, we propose two neural network architectures: Factored Forecaster (FF) which is ideal if the proxy is informative of the outcome in hindsight, and Residual Factored Forecaster (RFF) that is robust to a non-informative proxy. Experiments on two real-world datasets for predicting human behavior show that RFF outperforms both FF and a direct forecaster that does not make use of the proxy. Our results suggest that exploiting proxies by factorization is a promising way to mitigate the impact of long delays in human-behavior prediction tasks. Timothy A. Mann, Sven Gowal, András György 0001, Huiyi Hu, Ray Jiang, Balaji Lakshminarayanan, Prav Srinivasan |
ICML | 1 |
| 2019 | A Dual Approach to Verify and Train Deep NetworksabstractThis paper addressed the problem of formally verifying desirable properties of neural networks, i.e., obtaining provable guarantees that neural networks satisfy specifications relating their inputs and outputs (e.g., robustness to bounded norm adversarial perturbations). Most previous work on this topic was limited in its applicability by the size of the network, network architecture and the complexity of properties to be verified. In contrast, our framework applies to a general class of activation functions and specifications. We formulate verification as an optimization problem (seeking to find the largest violation of the specification) and solve a Lagrangian relaxation of the optimization problem to obtain an upper bound on the worst case violation of the specification being verified. Our approach is anytime, i.e., it can be stopped at any time and a valid bound on the maximum violation can be obtained. Finally, we highlight how this approach can be used to train models that are amenable to verification. Sven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Timothy A. Mann, Pushmeet Kohli |
IJCAI | 4 |
| 2019 | Adaptive Temporal-Difference Learning for Policy Evaluation with Per-State Uncertainty EstimatesabstractWe consider the core reinforcement-learning problem of on-policy value function approximation from a batch of trajectory data, and focus on various issues of Temporal Difference (TD) learning and Monte Carlo (MC) policy evaluation. The two methods are known to achieve complementary bias-variance trade-off properties, with TD tending to achieve lower variance but potentially higher bias. In this paper, we argue that the larger bias of TD can be a result of the amplification of local approximation errors. We address this by proposing an algorithm that adaptively switches between TD and MC in each state, thus mitigating the propagation of errors. Our method is based on learned confidence intervals that detect biases of TD estimates. We demonstrate in a variety of policy evaluation tasks that this simple adaptive algorithm performs competitively with the best approach in hindsight, suggesting that learned confidence intervals are a powerful technique for adapting policy evaluation to use TD or MC returns in a data-driven way. Carlos Riquelme, Hugo Penedones, Damien Vincent, Hartmut Maennel, Sylvain Gelly, Timothy A. Mann, André Barreto 0001, Gergely Neu |
NeurIPS | 6 |
| 2019 | A Bayesian Approach to Robust Reinforcement Learning
Esther Derman, Daniel J. Mankowitz, Timothy A. Mann, Shie Mannor |
UAI | 3 |
| 2018 | Learning Robust OptionsabstractRobust reinforcement learning aims to produce policies that have strong guarantees even in the face of environments/transition models whose parameters have strong uncertainty. Existing work uses value-based methods and the usual primitive action setting. In this paper, we propose robust methods for learning temporally abstract actions, in the framework of options. We present a Robust Options Policy Iteration (ROPI) algorithm with convergence guarantees, which learns options that are robust to model uncertainty. We utilize ROPI to learn robust options with the Robust Options Deep Q Network (RO-DQN) that solves multiple tasks and mitigates model misspecification due to model uncertainty. We present experimental results which suggest that policy iteration with linear features may have an inherent form of robustness when using coarse feature representations. In addition, we present experimental results which demonstrate that robustness helps policy iteration implemented on top of deep neural networks to generalize over a much broader range of dynamics than non-robust policy iteration. Daniel J. Mankowitz, Timothy A. Mann, Pierre-Luc Bacon, Doina Precup, Shie Mannor |
AAAI | 2 |
| 2018 | Soft-Robust Actor-Critic Policy-Gradient
Esther Derman, Daniel J. Mankowitz, Timothy A. Mann, Shie Mannor |
UAI | 3 |
| 2018 | A Dual Approach to Scalable Verification of Deep Networks
Krishnamurthy Dvijotham, Robert Stanforth, Sven Gowal, Timothy A. Mann, Pushmeet Kohli |
UAI | 4 |
| 2017 | Approximate Value Iteration with Temporally Extended Actions (Extended Abstract)abstractThe options framework provides a concrete way to implement and reason about temporally extended actions. Existing literature has demonstrated the value of planning with options empirically, but there is a lack of theoretical analysis formalizing when planning with options is more efficient than planning with primitive actions. We provide a general analysis of the convergence rate of a popular Approximate Value Iteration (AVI) algorithm called Fitted Value Iteration (FVI) with options. Our analysis reveals that longer duration options and a pessimistic estimate of the value function both lead to faster convergence. Furthermore, options can improve convergence even when they are suboptimal and sparsely distributed throughout the state space. Next we consider generating useful options for planning based on a subset of landmark states. This suggests a new algorithm, Landmark-based AVI (LAVI), that represents the value function only at landmark states. We analyze OFVI and LAVI using the proposed landmark-based options and compare the two algorithms. Our theoretical and experimental results demonstrate that options can play an important role in AVI by decreasing approximation error and inducing fast convergence. Timothy A. Mann, Shie Mannor, Doina Precup |
IJCAI | 1 |
| 2016 | Adaptive Skills Adaptive Partitions (ASAP)abstractWe introduce the Adaptive Skills, Adaptive Partitions (ASAP) framework that (1) learns skills (i.e., temporally extended actions or options) as well as (2) where to apply them. We believe that both (1) and (2) are necessary for a truly general skill learning framework, which is a key building block needed to scale up to lifelong learning agents. The ASAP framework is also able to solve related new tasks simply by adapting where it applies its existing learned skills. We prove that ASAP converges to a local optimum under natural conditions. Finally, our experimental results, which include a RoboCup domain, demonstrate the ability of ASAP to learn where to reuse skills as well as solve multiple tasks with considerably less experience than solving each task from scratch. Daniel J. Mankowitz, Timothy A. Mann, Shie Mannor |
NIPS | 2 |
| 2015 | Off-policy Model-based Learning under Unknown Factored DynamicsabstractOff-policy learning in dynamic decision problems is essential for providing strong evidence that a new policy is better than the one in use. But how can we prove superiority without testing the new policy? To answer this question, we introduce the G-SCOPE algorithm that evaluates a new policy based on data generated by the existing policy. Our algorithm is both computationally and sample efficient because it greedily learns to exploit factored structure in the dynamics of the environment. We present a finite sample analysis of our approach and show through experiments that the algorithm scales well on high-dimensional problems with few samples. Assaf Hallak, François Schnitzler, Timothy A. Mann, Shie Mannor |
ICML | 3 |
| 2015 | Approximate Value Iteration with Temporally Extended ActionsabstractTemporally extended actions have proven useful for reinforcement learning, but their duration also makes them valuable for efficient planning. The options framework provides a concrete way to implement and reason about temporally extended actions. Existing literature has demonstrated the value of planning with options empirically, but there is a lack of theoretical analysis formalizing when planning with options is more efficient than planning with primitive actions. We provide a general analysis of the convergence rate of a popular Approximate Value Iteration (AVI) algorithm called Fitted Value Iteration (FVI) with options. Our analysis reveals that longer duration options and a pessimistic estimate of the value function both lead to faster convergence. Furthermore, options can improve convergence even when they are suboptimal and sparsely distributed throughout the state-space. Next we consider the problem of generating useful options for planning based on a subset of landmark states. This suggests a new algorithm, Landmark-based AVI (LAVI), that represents the value function only at the landmark states. We analyze both FVI and LAVI using the proposed landmark-based options and compare the two algorithms. Our experimental results in three different domains demonstrate the key properties from the analysis. Our theoretical and experimental results demonstrate that options can play an important role in AVI by decreasing approximation error and inducing fast convergence. Timothy A. Mann, Shie Mannor, Doina Precup |
J. Artif. Intell. Res. | 1 |
| 2014 | Scaling Up Approximate Value Iteration with Options: Better Policies with Fewer IterationsabstractWe show how options, a class of control structures encompassing primitive and temporally extended actions, can play a valuable role in planning in MDPs with continuous state-spaces. Analyzing the convergence rate of Approximate Value Iteration with options reveals that for pessimistic initial value function estimates, options can speed up convergence compared to planning with only primitive actions even when the temporally extended actions are suboptimal and sparsely scattered throughout the state-space. Our experimental results in an optimal replacement task and a complex inventory management task demonstrate the potential for options to speed up convergence in practice. We show that options induce faster convergence to the optimal value function, which implies deriving better policies with fewer iterations. Timothy A. Mann, Shie Mannor |
ICML | 1 |
| 2014 | Time-Regularized Interrupting Options (TRIO)abstractHigh-level skills relieve planning algorithms from low-level details. But when the skills are poorly designed for the domain, the resulting plan may be severely suboptimal. Sutton et al. 1999 made an important step towards resolving this problem by introducing a rule that automatically improves a set of skills called options. This rule terminates an option early whenever switching to another option gives a higher value than continuing with the current option. However, they only analyzed the case where the improvement rule is applied once. We show conditions where this rule converges to the optimal set of options. A new Bellman-like operator that simultaneously improves the set of options is at the core of our analysis. One problem with the update rule is that it tends to favor lower-level skills. Therefore we introduce a regularization term that favors longer duration skills. Experimental results demonstrate that this approach can derive a good set of high-level skills even when the original set of skills cannot solve the problem. Timothy A. Mann, Daniel J. Mankowitz, Shie Mannor |
ICML | 1 |
| 2014 | How hard is my MDP?" The distribution-norm to the rescue"
Odalric-Ambrym Maillard, Timothy A. Mann, Shie Mannor |
NIPS | 2 |
| 2011 | Scaling Up Reinforcement Learning through Targeted ExplorationabstractRecent Reinforcement Learning (RL) algorithms, such as R-MAX, make (with high probability) only a small number of poor decisions. In practice, these algorithms do not scale well as the number of states grows because the algorithms spend too much effort exploring. We introduce an RL algorithm State TArgeted R-MAX (STAR-MAX) that explores a subset of the state space, called the exploration envelope ξ. When ξ equals the total state space, STAR-MAX behaves identically to R-MAX. When ξ is a subset of the state space, to keep exploration within ξ, a recovery rule β is needed. We compared existing algorithms with our algorithm employing various exploration envelopes. With an appropriate choice of ξ, STAR-MAX scales far better than existing RL algorithms as the number of states increases. A possible drawback of our algorithm is its dependence on a good choice of ξ and β. However, we show that an effective recovery rule β can be learned on-line and ξ can be learned from demonstrations. We also find that even randomly sampled exploration envelopes can improve cumulative rewards compared to R-MAX. We expect these results to lead to more efficient methods for RL in large-scale problems. Timothy A. Mann, Yoonsuck Choe |
AAAI | 1 |