VLDB 2026 Research / reviewers in the wild / expert
Peter R. Wurman
dblp:79/5768
· DBLP profile ↗
20ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-9349-0624ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Theory of computation · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 40% Motion planning and robot control · 20% Trustworthy machine learning · 12% | |
| Theoretical computer science
2 papers |
Algorithmic game theory and mechanism design · 100% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
1.2 | 2 | 2026 | Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026 Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy Exploration · NeurIPS 2024 |
Machine learning › Reinforcement learning › multi-task reinforcement learning
contextual reinforcement learning |
1.0 | 1 | 2026 | Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026 |
Robotics › Motion planning and robot control
robot control |
1.0 | 1 | 2026 | Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026 |
Robotics › Motion planning and robot control › robot control
robust control |
1.0 | 1 | 2026 | Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.9 | 1 | 2025 | SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning · ICLR 2025 |
Machine learning › Learning theory › inductive bias
simplicity bias |
0.9 | 1 | 2025 | SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning · ICLR 2025 |
Machine learning › Reinforcement learning › policy optimization
diverse policy learning |
0.8 | 1 | 2024 | Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy Exploration · NeurIPS 2024 |
Machine learning › Reinforcement learning
actor-critic methods |
0.6 | 1 | 2022 | Value Function Decomposition for Iterative Design of Reinforcement Learning Agents · NeurIPS 2022 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition |
0.6 | 1 | 2022 | Value Function Decomposition for Iterative Design of Reinforcement Learning Agents · NeurIPS 2022 |
Machine learning › Efficient and distributed learning
model compression |
0.5 | 1 | 2021 | Efficient Real-Time Inference in Temporal Convolution Networks · ICRA 2021 |
Machine learning › Efficient and distributed learning › inference efficiency
real-time inference |
0.5 | 1 | 2021 | Efficient Real-Time Inference in Temporal Convolution Networks · ICRA 2021 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.1 | 1 | 2021 | Efficient Real-Time Inference in Temporal Convolution Networks · ICRA 2021 |
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
temporal convolutional network |
0.1 | 1 | 2021 | Efficient Real-Time Inference in Temporal Convolution Networks · ICRA 2021 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search › best-first search
a* search |
0.1 | 1 | 2008 | PBA*: Using Proactive Search to Make A* Robust to Unplanned Deviations · AAAI 2008 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search |
0.1 | 1 | 2008 | PBA*: Using Proactive Search to Make A* Robust to Unplanned Deviations · AAAI 2008 |
Knowledge, reasoning and agents › Multi-agent systems
multi-robot coordination |
0.1 | 1 | 2007 | Coordinating Hundreds of Cooperative, Autonomous Vehicles in Warehouses · AAAI 2007 |
Robotics › Robot manipulation
warehouse automation |
0.1 | 1 | 2007 | Coordinating Hundreds of Cooperative, Autonomous Vehicles in Warehouses · AAAI 2007 |
Algorithmic game theory and mechanism design › auction theory
combinatorial auction |
0.1 | 2 | 2003 | Computing the outcome of proxy bidding in combinatorial auctions extended abstract · EC 2003 AkBA: a progressive, anonymous-price combinatorial auction · EC 2000 |
Algorithmic game theory and mechanism design › auction theory › bidding strategy
proxy bidding |
0.0 | 1 | 2003 | Computing the outcome of proxy bidding in combinatorial auctions extended abstract · EC 2003 |
Algorithmic game theory and mechanism design › market equilibrium
competitive equilibrium |
0.0 | 1 | 2000 | AkBA: a progressive, anonymous-price combinatorial auction · EC 2000 |
Algorithmic game theory and mechanism design › mechanism design › auction design
iterative auction |
0.0 | 1 | 2000 | AkBA: a progressive, anonymous-price combinatorial auction · EC 2000 |
Algorithmic game theory and mechanism design › mechanism design › auction design › multi-unit auction
uniform-price auction |
0.0 | 1 | 2000 | AkBA: a progressive, anonymous-price combinatorial auction · EC 2000 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
multi-agent path finding |
0.0 | 1 | 2007 | Coordinating Hundreds of Cooperative, Autonomous Vehicles in Warehouses · AAAI 2007 |
Methods — techniques the papers use, named apart from their topics
single-phase adaptation · 1.0context encoder · 1.0residual connections · 0.9observation normalization · 0.9layer normalization · 0.9successor features · 0.8diversity objective with constraints · 0.8value decomposition · 0.6soft actor-critic · 0.6RT-TCN algorithm · 0.5outcome computation · 0.0equilibrium price construction · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single PolicyabstractGeneralization to unseen environments is a significant challenge in the field of robotics and control. In this work, we focus on contextual reinforcement learning, where agents act within environments with varying contexts, such as self-driving cars or quadrupedal robots that need to operate in different terrains or weather conditions than they were trained for. We tackle the critical task of generalizing to out-of-distribution (OOD) settings, without access to explicit context information at test time. Recent work has addressed this problem by training a context encoder and a history adaptation module in separate stages. While promising, this two-phase approach is cumbersome to implement and train. We simplify the methodology and introduce SPARC: single-phase adaptation for robust control. We test SPARC on varying contexts within the high-fidelity racing simulator Gran Turismo 7 and wind-perturbed MuJoCo environments, and find that it achieves reliable and robust OOD generalization. Bram Grooten, Patrick MacAlpine, Kaushik Subramanian, Peter Stone 0001, Peter R. Wurman |
AAAI | 5 |
| 2025 | SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement LearningabstractRecent advances in CV and NLP have been largely driven by scaling up the number of network parameters, despite traditional theories suggesting that larger networks are prone to overfitting.
These large networks avoid overfitting by integrating components that induce a simplicity bias, guiding models toward simple and generalizable solutions.
However, in deep RL, designing and scaling up networks have been less explored.
Motivated by this opportunity, we present SimBa, an architecture designed to scale up parameters in deep RL by injecting a simplicity bias. SimBa consists of three components: (i) an observation normalization layer that standardizes inputs with running statistics, (ii) a residual feedforward block to provide a linear pathway from the input to output, and (iii) a layer normalization to control feature magnitudes.
By scaling up parameters with SimBa, the sample efficiency of various deep RL algorithms—including off-policy, on-policy, and unsupervised methods—is consistently improved.
Moreover, solely by integrating SimBa architecture into SAC, it matches or surpasses state-of-the-art deep RL methods with high computational efficiency across DMC, MyoSuite, and HumanoidBench.
These results demonstrate SimBa's broad applicability and effectiveness across diverse RL algorithms and environments. Dongyoon Hwang, Donghu Kim, Hyunseung Kim, Jun Jet Tai, Kaushik Subramanian, Peter R. Wurman, Jaegul Choo, Peter Stone 0001, Takuma Seno |
ICLR | 7 |
| 2024 | Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy ExplorationabstractThe ability to approach the same problem from different angles is a cornerstone of human intelligence that leads to robust solutions and effective adaptation to problem variations. In contrast, current RL methodologies tend to lead to policies that settle on a single solution to a given problem, making them brittle to problem variations. Replicating human flexibility in reinforcement learning agents is the challenge that we explore in this work. We tackle this challenge by extending state-of-the-art approaches to introduce DUPLEX, a method that explicitly defines a diversity objective with constraints and makes robust estimates of policies’ expected behavior through successor features. The trained agents can (i) learn a diverse set of near-optimal policies in complex highly-dynamic environments and (ii) exhibit competitive and diverse skills in out-of-distribution (OOD) contexts. Empirical results indicate that DUPLEX improves over previous methods and successfully learns competitive driving styles in a hyper-realistic simulator (i.e., GranTurismo ™ 7) as well as diverse and effective policies in several multi-context robotics MuJoCo simulations with OOD gravity forces and height limits. To the best of our knowledge, our method is the first to achieve diverse solutions in complex driving simulators and OOD robotic contexts. DUPLEX agents demonstrating diverse behaviors can be found at https://ai.sony/publications/Discovering-Creative-Behaviors-through-DUPLEX-Diverse-Universal-Features-for-Policy-Exploration/. Borja G. León, Francesco Riccio, Kaushik Subramanian, Peter R. Wurman, Peter Stone 0001 |
NeurIPS | 4 |
| 2023 | Composing Efficient, Robust Tests for Policy SelectionabstractModern reinforcement learning systems produce many high-quality policies throughout the learning process. However, to choose which policy to actually deploy in the real world, they must be tested under an intractable number of environmental conditions. We introduce RPOSST, an algorithm to select a small set of test cases from a larger pool based on a relatively small number of sample evaluations. RPOSST treats the test case selection problem as a two-player game and optimizes a solution with provable $k$-of-$N$ robustness, bounding the error relative to a test that used all the test cases in the pool. Empirical results demonstrate that RPOSST finds a small set of test cases that identify high quality policies in a toy one-shot game, poker datasets, and a high-fidelity racing simulator. Dustin Morrill, Thomas J. Walsh 0001, Daniel Hernandez, Peter R. Wurman, Peter Stone 0001 |
UAI | 4 |
| 2022 | Value Function Decomposition for Iterative Design of Reinforcement Learning AgentsabstractDesigning reinforcement learning (RL) agents is typically a difficult process that requires numerous design iterations. Learning can fail for a multitude of reasons and standard RL methods provide too few tools to provide insight into the exact cause. In this paper, we show how to integrate \textit{value decomposition} into a broad class of actor-critic algorithms and use it to assist in the iterative agent-design process. Value decomposition separates a reward function into distinct components and learns value estimates for each. These value estimates provide insight into an agent's learning and decision-making process and enable new training methods to mitigate common problems. As a demonstration, we introduce SAC-D, a variant of soft actor-critic (SAC) adapted for value decomposition. SAC-D maintains similar performance to SAC, while learning a larger set of value predictions. We also introduce decomposition-based tools that exploit this information, including a new reward \textit{influence} metric, which measures each reward component's effect on agent decision-making. Using these tools, we provide several demonstrations of decomposition's use in identifying and addressing problems in the design of both environments and agents. Value decomposition is broadly applicable and easy to incorporate into existing algorithms and workflows, making it a powerful tool in an RL practitioner's toolbox. James MacGlashan, Evan Archer, Alisa Devlic, Takuma Seno, Craig Sherstan, Peter R. Wurman, Peter Stone 0001 |
NeurIPS | 6 |
| 2021 | Efficient Real-Time Inference in Temporal Convolution NetworksabstractIt has been recently demonstrated that Temporal Convolution Networks (TCNs) provide state-of-the-art results in many problem domains where the input data is a time-series. TCNs typically incorporate information from a long history of inputs (the receptive field) into a single output using many convolution layers. Real-time inference using a trained TCN can be challenging on devices with limited compute and memory, especially if the receptive field is large. This paper introduces the RT-TCN algorithm that reuses the output of prior convolution operations to minimize the computational requirements and persistent memory footprint of a TCN during real-time inference. We also show that when a TCN is trained using time slices of the input time-series, it can be executed in realtime continually using RT-TCN. In addition, we provide TCN architecture guidelines that ensure that real-time inference can be performed within memory and computational constraints. Piyush Khandelwal, James MacGlashan, Peter R. Wurman, Peter Stone 0001 |
ICRA | 3 |
| 2021 | Agent-Based Markov Modeling for Improved COVID-19 Mitigation PoliciesabstractThe year 2020 saw the covid-19 virus lead to one of the worst global pandemics in history. As a result, governments around the world have been faced with the challenge of protecting public health while keeping the economy running to the greatest extent possible. Epidemiological models provide insight into the spread of these types of diseases and predict the effects of possible intervention policies. However, to date, even the most data-driven intervention policies rely on heuristics. In this paper, we study how reinforcement learning (RL) and Bayesian inference can be used to optimize mitigation policies that minimize economic impact without overwhelming hospital capacity. Our main contributions are (1) a novel agent-based pandemic simulator which, unlike traditional models, is able to model fine-grained interactions among people at specific locations in a community; (2) an RLbased methodology for optimizing fine-grained mitigation policies within this simulator; and (3) a Hidden Markov Model for predicting infected individuals based on partial observations regarding test results, presence of symptoms, and past physical contacts. This article is part of the special track on AI and COVID-19. Roberto Capobianco, Varun Raj Kompella, James Ault, Guni Sharon, Stacy Jong, Spencer J. Fox, Lauren Ancel Meyers, Peter R. Wurman, Peter Stone 0001 |
J. Artif. Intell. Res. | 8 |
| 2018 | Analysis and Observations From the First Amazon Picking ChallengeabstractThis paper presents an overview of the inaugural Amazon Picking Challenge along with a summary of a survey conducted among the 26 participating teams. The challenge goal was to design an autonomous robot to pick items from a warehouse shelf. This task is currently performed by human workers, and there is hope that robots can someday help increase efficiency and throughput while lowering cost. We report on a 28-question survey posed to the teams to learn about each team's background, mechanism design, perception apparatus, planning, and control approach. We identify trends in this data, correlate it with each team's success in the competition, and discuss observations and lessons learned based on survey results and the authors' personal experiences during the challenge. Nikolaus Correll, Kostas E. Bekris, Dmitry Berenson, Oliver Brock, Albert J. Causo, Kris Hauser, Kei Okada, Alberto Rodriguez 0003, Joseph M. Romano, Peter R. Wurman |
IEEE Trans Autom. Sci. Eng. | 10 |
| 2008 | PBA*: Using Proactive Search to Make A* Robust to Unplanned Deviations
Paul Breimyer, Peter R. Wurman |
AAAI | 2 |
| 2007 | The game of scale: decision making with economies of scaleabstractWhile diffusion of innovation topics in economics and majority games in game theory have been widely studied, the impact of economy-of-scale effects in aggregated decision making has received little attention. In this paper, we present a basic model, the Game of Scale, to study the effects of economy-of-scale in decision making among a large pool of self-interested agents. We solve the model's static equilibria and present two dynamic decision models, one myopic and one trend-following. Most of the parameter space converges quickly; however, the behaviors exhibited near critical input values show drastic changes. We demonstrate how trend-following can improve global outcomes over myopic decision making. Finally, we describe how the game can be externally controlled. Christopher J. Hazard, Peter R. Wurman |
ICEC | 2 |
| 2007 | Coordinating Hundreds of Cooperative, Autonomous Vehicles in Warehouses
Peter R. Wurman, Raffaello D'Andrea, Mick Mountz |
AAAI | 1 |
| 2005 | Applying metaheuristic techniques to search the space of bidding strategies in combinatorial auctionsabstractMany non-cooperative settings that could potentially be studied using game theory are characterized by having very large strategy spaces and payoffs that are costly to compute. Best response dynamics is a method of searching for pure-strategy equilibria in games that is attractive for its simplicity and scalability (relative to more analytical approaches). However, when the cost of determining the outcome of a particular set of joint strategies is high, it is impractical to compute the payoffs of all possible responses to the other players actions. Thus, we study metaheuristic approaches--genetic algorithms and tabu search in particular--to explore the strategy space. We configure the parameters of metaheuristics to adapt to the problem of finding the best response strategy and present how it can be helpful in finding Nash equilibria of combinatorial auctions which is an important solution concept in game theory. Ashish Sureka, Peter R. Wurman |
GECCO | 2 |
| 2005 | Monte Carlo approximation in incomplete information, sequential auction games
Gangshu (George) Cai, Peter R. Wurman |
Decis. Support Syst. | 2 |
| 2003 | A comparison of two algorithms for multi-unit k-double auctionsabstractWe develop two algorithms to manage bid data in flexible, multi-unit double auctions. The first algorithm is a multi-unit extension of the 4-HEAP algorithm, and the second is a novel algorithm based on the Internal Path Reduction tree. To facilitate the generation of price quotes, we enhance the IPR tree algorithm to maintain information about the current Mth and (M + 1)st units. Our experiments show that when bids are relatively unordered, AUC-IPR outperforms 4-HEAP by a factor of about 2.6. However, under some scenarios in which bids are received in order, AUC-IPR is not always better than 4-HEAP. Shengli Bao, Peter R. Wurman |
ICEC | 2 |
| 2003 | An algorithm for computing the outcome of combinatorial auctions with proxy biddingabstractProxy bidding has proved useful in a variety of real auction formats, such as eBay, and has been proposed for some combinatorial auctions. Previous work on proxy bidding in combinatorial auctions requires the auctioneer essentially run the auction with myopic bidders to determine the outcome. In addition to being computationally costly, this process is only as accurate as the bid increment, and decreasing the bid increment to improve accuracy greatly increases the running time. In this paper, we present an algorithm that computes the outcome of the proxy auction by examining only the events that cause the proxy bidders to change their behaviors. This algorithm is much faster than the alternative, and computes exact solutions. Peter R. Wurman, Gangshu (George) Cai, Ashish Sureka |
ICEC | 1 |
| 2003 | Computing the outcome of proxy bidding in combinatorial auctions extended abstractabstractNo abstract available. Peter R. Wurman, Gangshu (George) Cai, Ashish Sureka |
EC | 1 |
| 2000 | AkBA: a progressive, anonymous-price combinatorial auctionabstractThe allocation of discrete, complementary resources is a fundamental problem in economics and of direct interest to e-commerce applications.Combinatorial auctions account for complementarities by optimizing over oers expressed in terms of bundles.Progressive v ersions of combinatorial auctions alleviate the burden on bidders of expressing offers for all bundles of interest by p r o viding interim feedback based on partial sets of bids.Feedback i n t e r m s o f h ypothetical prices is particularly useful, as it directs bidders toward those bundles potentially yielding the greatest surplus.For a general class of discrete resource allocation problems with free disposal, we establish by construction the existence of competitive equilibrium prices on bundles that support the eÆcient allocation.We i n troduce AkBA, a family of progressive auctions that use these equilibrium bundle prices.We examine a particular instance of the family, called A1BA, and present some empirical data on its performance. Peter R. Wurman, Michael P. Wellman |
EC | 1 |
| 1998 | Some Economics of Market-Based Distributed SchedulingabstractMarket mechanisms solve distributed scheduling problems by allocating the scheduled resources according to market prices. We model distributed scheduling as a discrete resource allocation problem, and demonstrate the applicability of economic analysis to this framework. Drawing on results from the literature, we discuss the existence of equilibrium prices for some general classes of scheduling problems, and the quality of equilibrium solutions. We then present two auction protocols for implementing solutions, and analyze their computational and economic properties. William E. Walsh, Michael P. Wellman, Peter R. Wurman, Jeffrey K. MacKie-Mason |
ICDCS | 3 |
| 1998 | Flexible double auctions for electronic commerce: theory and implementation
Peter R. Wurman, William E. Walsh, Michael P. Wellman |
Decis. Support Syst. | 1 |
| 1996 | Optimal Factory Scheduling using Stochastic Dominance A*
Peter R. Wurman, Michael P. Wellman |
UAI | 1 |