VLDB 2026 Research / reviewers in the wild / expert
Sam Toyer
dblp:203/9103
· DBLP profile ↗
7ranked-venue papers
4as first author
2since 2021 · last 2024
0000-0002-6665-6593ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
2 papers |
Security and privacy of machine learning · 80% Systems and software security · 20% | |
| Artificial intelligence
3 papers |
Reinforcement learning · 47% Planning, search and constraint satisfaction · 19% Trustworthy machine learning · 13% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
imitation learning |
0.8 | 2 | 2020 | The MAGICAL Benchmark for Robust Imitation · NeurIPS 2020 Variational Discriminator Bottleneck: Improving Imitation Learning, Inverse RL, and GANs by Constraining Information Flow · ICLR (Poster) 2019 |
Security and privacy of machine learning › large language model safety
fine-tuning safety |
0.8 | 1 | 2024 | A StrongREJECT for Empty Jailbreaks · NeurIPS 2024 |
Security and privacy of machine learning › adversarial attack
jailbreak attack |
0.8 | 1 | 2024 | A StrongREJECT for Empty Jailbreaks · NeurIPS 2024 |
Security and privacy of machine learning
large language model safety |
0.8 | 1 | 2024 | A StrongREJECT for Empty Jailbreaks · NeurIPS 2024 |
Security and privacy of machine learning
prompt injection |
0.8 | 1 | 2024 | Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game · ICLR 2024 |
Systems and software security
vulnerability discovery |
0.8 | 1 | 2024 | Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game · ICLR 2024 |
Machine learning › Reinforcement learning › imitation learning
robust imitation learning |
0.4 | 1 | 2020 | The MAGICAL Benchmark for Robust Imitation · NeurIPS 2020 |
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift |
0.4 | 1 | 2020 | The MAGICAL Benchmark for Robust Imitation · NeurIPS 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2019 | Variational Discriminator Bottleneck: Improving Imitation Learning, Inverse RL, and GANs by Constraining Information Flow · ICLR (Poster) 2019 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.4 | 1 | 2019 | Variational Discriminator Bottleneck: Improving Imitation Learning, Inverse RL, and GANs by Constraining Information Flow · ICLR (Poster) 2019 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › generalized planning
general policies |
0.3 | 1 | 2018 | Action Schema Networks: Generalised Policies With Deep Learning · AAAI 2018 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
probabilistic planning |
0.3 | 1 | 2018 | Action Schema Networks: Generalised Policies With Deep Learning · AAAI 2018 |
Machine learning › Graph learning › graph neural network
relational network |
0.3 | 1 | 2018 | Action Schema Networks: Generalised Policies With Deep Learning · AAAI 2018 |
Methods — techniques the papers use, named apart from their topics
large language model · 0.8benchmark construction · 0.8automated evaluation · 0.8benchmark suite · 0.4variational information bottleneck · 0.4weight sharing · 0.3supervised training · 0.3deep learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Tensor Trust: Interpretable Prompt Injection Attacks from an Online GameabstractWhile Large Language Models (LLMs) are increasingly being used in real-world applications, they remain vulnerable to *prompt injection attacks*: malicious third party prompts that subvert the intent of the system designer. To help researchers study this problem, we present a dataset of over 563,000 prompt injection attacks and 118,000 prompt-based "defenses" against prompt injection, all created by players of an online game called Tensor Trust. To the best of our knowledge, this is the first dataset that includes both human-generated attacks and defenses for instruction-following LLMs. The attacks in our dataset have easily interpretable structure, and shed light on the weaknesses of LLMs. We also use the dataset to create a benchmark for resistance to two types of prompt injection, which we refer to as *prompt extraction* and *prompt hijacking*. Our benchmark results show that many models are vulnerable to the attack strategies in the Tensor Trust dataset. Furthermore, we show that some attack strategies from the dataset generalize to deployed LLM-based applications, even though they have a very different set of constraints to the game. We release data and code at [tensortrust.ai/paper](https://tensortrust.ai/paper) Sam Toyer, Olivia Watkins, Ethan Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, Stuart Russell 0001 |
ICLR | 1 |
| 2024 | A StrongREJECT for Empty JailbreaksabstractMost jailbreak papers claim the jailbreaks they propose are highly effective, often boasting near-100% attack success rates. However, it is perhaps more common than not for jailbreak developers to substantially exaggerate the effectiveness of their jailbreaks. We suggest this problem arises because jailbreak researchers lack a standard, high-quality benchmark for evaluating jailbreak performance, leaving researchers to create their own. To create a benchmark, researchers must choose a dataset of forbidden prompts to which a victim model will respond, along with an evaluation method that scores the harmfulness of the victim model’s responses. We show that existing benchmarks suffer from significant shortcomings and introduce the StrongREJECT benchmark to address these issues. StrongREJECT's dataset contains prompts that victim models must answer with specific, harmful information, while its automated evaluator measures the extent to which a response gives useful information to forbidden prompts. In doing so, the StrongREJECT evaluator achieves state-of-the-art agreement with human judgments of jailbreak effectiveness. Notably, we find that existing evaluation methods significantly overstate jailbreak effectiveness compared to human judgments and the StrongREJECT evaluator. We describe a surprising and novel phenomenon that explains this discrepancy: jailbreaks bypassing a victim model’s safety fine-tuning tend to reduce its capabilities. Together, our findings underscore the need for researchers to use a high-quality benchmark, such as StrongREJECT, when developing new jailbreak attacks. We release the StrongREJECT code and data at https://strong-reject.readthedocs.io/. Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, Sam Toyer |
NeurIPS | 11 |
| 2020 | The MAGICAL Benchmark for Robust ImitationabstractImitation Learning (IL) algorithms are typically evaluated in the same environment that was used to create demonstrations. This rewards precise reproduction of demonstrations in one particular environment, but provides little information about how robustly an algorithm can generalise the demonstrator's intent to substantially different deployment settings. This paper presents the MAGICAL benchmark suite, which permits systematic evaluation of generalisation by quantifying robustness to different kinds of distribution shift that an IL algorithm is likely to encounter in practice. Using the MAGICAL suite, we confirm that existing IL algorithms overfit significantly to the context in which demonstrations are provided. We also show that standard methods for reducing overfitting are effective at creating narrow perceptual invariances, but are not sufficient to enable transfer to contexts that require substantially different behaviour, which suggests that new approaches will be needed in order to robustly generalise demonstrator intent. Code and data for the MAGICAL suite is available at https://github.com/qxcv/magical/ Sam Toyer, Rohin Shah, Andrew Critch, Stuart Russell 0001 |
NeurIPS | 1 |
| 2020 | ASNets: Deep Learning for Generalised PlanningabstractIn this paper, we discuss the learning of generalised policies for probabilistic and classical planning problems using Action Schema Networks (ASNets). The ASNet is a neural network architecture that exploits the relational structure of (P)PDDL planning problems to learn a common set of weights that can be applied to any problem in a domain. By mimicking the actions chosen by a traditional, non-learning planner on a handful of small problems in a domain, ASNets are able to learn a generalised reactive policy that can quickly solve much larger instances from the domain. This work extends the ASNet architecture to make it more expressive, while still remaining invariant to a range of symmetries that exist in PPDDL problems. We also present a thorough experimental evaluation of ASNets, including a comparison with heuristic search planners on seven probabilistic and deterministic domains, an extended evaluation on over 18,000 Blocksworld instances, and an ablation study. Finally, we show that sparsity-inducing regularisation can produce ASNets that are compact enough for humans to understand, yielding insights into how the structure of ASNets allows them to generalise across a domain. Sam Toyer, Sylvie Thiébaux, Felipe W. Trevizan, Lexing Xie |
J. Artif. Intell. Res. | 1 |
| 2019 | Variational Discriminator Bottleneck: Improving Imitation Learning, Inverse RL, and GANs by Constraining Information Flow
Xue Bin Peng, Angjoo Kanazawa, Sam Toyer, Pieter Abbeel, Sergey Levine |
ICLR (Poster) | 3 |
| 2019 | Guiding Search with Generalized Policies for Probabilistic PlanningabstractWe examine techniques for combining generalized policies with search algorithms to exploit the strengths and overcome the weaknesses of each when solving probabilistic planning problems. The Action Schema Network (ASNet) is a recent contribution to planning that uses deep learning and neural networks to learn generalized policies for probabilistic planning problems. ASNets are well suited to problems where local knowledge of the environment can be exploited to improve performance, but may fail to generalize to problems they were not trained on. Monte-Carlo Tree Search (MCTS) is a forward-chaining state space search algorithm for optimal decision making which performs simulations to incrementally build a search tree and estimate the values of each state. Although MCTS can achieve state-of-the-art results when paired with domain-specific knowledge, without this knowledge, MCTS requires a large number of simulations in order to obtain reliable state-value estimates. By combining ASNets with MCTS, we are able to improve the capability of an ASNet to generalize beyond the distribution of problems it was trained on, as well as enhance the navigation of the search space by MCTS. William Shen, Felipe W. Trevizan, Sam Toyer, Sylvie Thiébaux, Lexing Xie |
SOCS | 3 |
| 2018 | Action Schema Networks: Generalised Policies With Deep LearningabstractIn this paper, we introduce the Action Schema Network (ASNet): a neural network architecture for learning generalised policies for probabilistic planning problems. By mimicking the relational structure of planning problems, ASNets are able to adopt a weight sharing scheme which allows the network to be applied to any problem from a given planning domain. This allows the cost of training the network to be amortised over all problems in that domain. Further, we propose a training method which balances exploration and supervised training on small problems to produce a policy which remains robust when evaluated on larger problems. In experiments, we show that ASNet's learning capability allows it to significantly outperform traditional non-learning planners in several challenging domains. Sam Toyer, Felipe W. Trevizan, Sylvie Thiébaux, Lexing Xie |
AAAI | 1 |