Peter R. Wurman

dblp:79/5768 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-9349-0624ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Theory of computation · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Reinforcement learning · 40% Motion planning and robot control · 20% Trustworthy machine learning · 12%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 100%

Topics — the 23 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
out-of-distribution generalization
1.222026
Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026
Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy Exploration · NeurIPS 2024
Machine learning › Reinforcement learning › multi-task reinforcement learning
contextual reinforcement learning
1.012026
Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026
Robotics › Motion planning and robot control
robot control
1.012026
Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026
Robotics › Motion planning and robot control › robot control
robust control
1.012026
Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy · AAAI 2026
Machine learning › Reinforcement learning
deep reinforcement learning
0.912025
SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning · ICLR 2025
Machine learning › Learning theory › inductive bias
simplicity bias
0.912025
SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning › policy optimization
diverse policy learning
0.812024
Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy Exploration · NeurIPS 2024
Machine learning › Reinforcement learning
actor-critic methods
0.612022
Value Function Decomposition for Iterative Design of Reinforcement Learning Agents · NeurIPS 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition
0.612022
Value Function Decomposition for Iterative Design of Reinforcement Learning Agents · NeurIPS 2022
Machine learning › Efficient and distributed learning
model compression
0.512021
Efficient Real-Time Inference in Temporal Convolution Networks · ICRA 2021
Machine learning › Efficient and distributed learning › inference efficiency
real-time inference
0.512021
Efficient Real-Time Inference in Temporal Convolution Networks · ICRA 2021
Machine learning › Efficient and distributed learning
inference efficiency
0.112021
Efficient Real-Time Inference in Temporal Convolution Networks · ICRA 2021
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
temporal convolutional network
0.112021
Efficient Real-Time Inference in Temporal Convolution Networks · ICRA 2021
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search › best-first search
a* search
0.112008
PBA*: Using Proactive Search to Make A* Robust to Unplanned Deviations · AAAI 2008
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search
0.112008
PBA*: Using Proactive Search to Make A* Robust to Unplanned Deviations · AAAI 2008
Knowledge, reasoning and agents › Multi-agent systems
multi-robot coordination
0.112007
Coordinating Hundreds of Cooperative, Autonomous Vehicles in Warehouses · AAAI 2007
Robotics › Robot manipulation
warehouse automation
0.112007
Coordinating Hundreds of Cooperative, Autonomous Vehicles in Warehouses · AAAI 2007
Algorithmic game theory and mechanism design › auction theory
combinatorial auction
0.122003
Computing the outcome of proxy bidding in combinatorial auctions extended abstract · EC 2003
AkBA: a progressive, anonymous-price combinatorial auction · EC 2000
Algorithmic game theory and mechanism design › auction theory › bidding strategy
proxy bidding
0.012003
Computing the outcome of proxy bidding in combinatorial auctions extended abstract · EC 2003
Algorithmic game theory and mechanism design › market equilibrium
competitive equilibrium
0.012000
AkBA: a progressive, anonymous-price combinatorial auction · EC 2000
Algorithmic game theory and mechanism design › mechanism design › auction design
iterative auction
0.012000
AkBA: a progressive, anonymous-price combinatorial auction · EC 2000
Algorithmic game theory and mechanism design › mechanism design › auction design › multi-unit auction
uniform-price auction
0.012000
AkBA: a progressive, anonymous-price combinatorial auction · EC 2000
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
multi-agent path finding
0.012007
Coordinating Hundreds of Cooperative, Autonomous Vehicles in Warehouses · AAAI 2007

Methods — techniques the papers use, named apart from their topics

single-phase adaptation · 1.0context encoder · 1.0residual connections · 0.9observation normalization · 0.9layer normalization · 0.9successor features · 0.8diversity objective with constraints · 0.8value decomposition · 0.6soft actor-critic · 0.6RT-TCN algorithm · 0.5outcome computation · 0.0equilibrium price construction · 0.0
YearPublicationVenuePosition
2026 Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy
abstract
Generalization to unseen environments is a significant challenge in the field of robotics and control. In this work, we focus on contextual reinforcement learning, where agents act within environments with varying contexts, such as self-driving cars or quadrupedal robots that need to operate in different terrains or weather conditions than they were trained for. We tackle the critical task of generalizing to out-of-distribution (OOD) settings, without access to explicit context information at test time. Recent work has addressed this problem by training a context encoder and a history adaptation module in separate stages. While promising, this two-phase approach is cumbersome to implement and train. We simplify the methodology and introduce SPARC: single-phase adaptation for robust control. We test SPARC on varying contexts within the high-fidelity racing simulator Gran Turismo 7 and wind-perturbed MuJoCo environments, and find that it achieves reliable and robust OOD generalization.
Bram Grooten, Patrick MacAlpine, Kaushik Subramanian, Peter Stone 0001, Peter R. Wurman
AAAI5
2025 SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning
abstract
Recent advances in CV and NLP have been largely driven by scaling up the number of network parameters, despite traditional theories suggesting that larger networks are prone to overfitting. These large networks avoid overfitting by integrating components that induce a simplicity bias, guiding models toward simple and generalizable solutions. However, in deep RL, designing and scaling up networks have been less explored. Motivated by this opportunity, we present SimBa, an architecture designed to scale up parameters in deep RL by injecting a simplicity bias. SimBa consists of three components: (i) an observation normalization layer that standardizes inputs with running statistics, (ii) a residual feedforward block to provide a linear pathway from the input to output, and (iii) a layer normalization to control feature magnitudes. By scaling up parameters with SimBa, the sample efficiency of various deep RL algorithms—including off-policy, on-policy, and unsupervised methods—is consistently improved. Moreover, solely by integrating SimBa architecture into SAC, it matches or surpasses state-of-the-art deep RL methods with high computational efficiency across DMC, MyoSuite, and HumanoidBench. These results demonstrate SimBa's broad applicability and effectiveness across diverse RL algorithms and environments.
Dongyoon Hwang, Donghu Kim, Hyunseung Kim, Jun Jet Tai, Kaushik Subramanian, Peter R. Wurman, Jaegul Choo, Peter Stone 0001, Takuma Seno
ICLR7
2024 Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy Exploration
abstract
The ability to approach the same problem from different angles is a cornerstone of human intelligence that leads to robust solutions and effective adaptation to problem variations. In contrast, current RL methodologies tend to lead to policies that settle on a single solution to a given problem, making them brittle to problem variations. Replicating human flexibility in reinforcement learning agents is the challenge that we explore in this work. We tackle this challenge by extending state-of-the-art approaches to introduce DUPLEX, a method that explicitly defines a diversity objective with constraints and makes robust estimates of policies’ expected behavior through successor features. The trained agents can (i) learn a diverse set of near-optimal policies in complex highly-dynamic environments and (ii) exhibit competitive and diverse skills in out-of-distribution (OOD) contexts. Empirical results indicate that DUPLEX improves over previous methods and successfully learns competitive driving styles in a hyper-realistic simulator (i.e., GranTurismo ™ 7) as well as diverse and effective policies in several multi-context robotics MuJoCo simulations with OOD gravity forces and height limits. To the best of our knowledge, our method is the first to achieve diverse solutions in complex driving simulators and OOD robotic contexts. DUPLEX agents demonstrating diverse behaviors can be found at https://ai.sony/publications/Discovering-Creative-Behaviors-through-DUPLEX-Diverse-Universal-Features-for-Policy-Exploration/.
Borja G. León, Francesco Riccio, Kaushik Subramanian, Peter R. Wurman, Peter Stone 0001
NeurIPS4
2023 Composing Efficient, Robust Tests for Policy Selection
abstract
Modern reinforcement learning systems produce many high-quality policies throughout the learning process. However, to choose which policy to actually deploy in the real world, they must be tested under an intractable number of environmental conditions. We introduce RPOSST, an algorithm to select a small set of test cases from a larger pool based on a relatively small number of sample evaluations. RPOSST treats the test case selection problem as a two-player game and optimizes a solution with provable $k$-of-$N$ robustness, bounding the error relative to a test that used all the test cases in the pool. Empirical results demonstrate that RPOSST finds a small set of test cases that identify high quality policies in a toy one-shot game, poker datasets, and a high-fidelity racing simulator.
Dustin Morrill, Thomas J. Walsh 0001, Daniel Hernandez, Peter R. Wurman, Peter Stone 0001
UAI4
2022 Value Function Decomposition for Iterative Design of Reinforcement Learning Agents
abstract
Designing reinforcement learning (RL) agents is typically a difficult process that requires numerous design iterations. Learning can fail for a multitude of reasons and standard RL methods provide too few tools to provide insight into the exact cause. In this paper, we show how to integrate \textit{value decomposition} into a broad class of actor-critic algorithms and use it to assist in the iterative agent-design process. Value decomposition separates a reward function into distinct components and learns value estimates for each. These value estimates provide insight into an agent's learning and decision-making process and enable new training methods to mitigate common problems. As a demonstration, we introduce SAC-D, a variant of soft actor-critic (SAC) adapted for value decomposition. SAC-D maintains similar performance to SAC, while learning a larger set of value predictions. We also introduce decomposition-based tools that exploit this information, including a new reward \textit{influence} metric, which measures each reward component's effect on agent decision-making. Using these tools, we provide several demonstrations of decomposition's use in identifying and addressing problems in the design of both environments and agents. Value decomposition is broadly applicable and easy to incorporate into existing algorithms and workflows, making it a powerful tool in an RL practitioner's toolbox.
James MacGlashan, Evan Archer, Alisa Devlic, Takuma Seno, Craig Sherstan, Peter R. Wurman, Peter Stone 0001
NeurIPS6
2021 Efficient Real-Time Inference in Temporal Convolution Networks
abstract
It has been recently demonstrated that Temporal Convolution Networks (TCNs) provide state-of-the-art results in many problem domains where the input data is a time-series. TCNs typically incorporate information from a long history of inputs (the receptive field) into a single output using many convolution layers. Real-time inference using a trained TCN can be challenging on devices with limited compute and memory, especially if the receptive field is large. This paper introduces the RT-TCN algorithm that reuses the output of prior convolution operations to minimize the computational requirements and persistent memory footprint of a TCN during real-time inference. We also show that when a TCN is trained using time slices of the input time-series, it can be executed in realtime continually using RT-TCN. In addition, we provide TCN architecture guidelines that ensure that real-time inference can be performed within memory and computational constraints.
Piyush Khandelwal, James MacGlashan, Peter R. Wurman, Peter Stone 0001
ICRA3
2021 Agent-Based Markov Modeling for Improved COVID-19 Mitigation Policies
abstract
The year 2020 saw the covid-19 virus lead to one of the worst global pandemics in history. As a result, governments around the world have been faced with the challenge of protecting public health while keeping the economy running to the greatest extent possible. Epidemiological models provide insight into the spread of these types of diseases and predict the effects of possible intervention policies. However, to date, even the most data-driven intervention policies rely on heuristics. In this paper, we study how reinforcement learning (RL) and Bayesian inference can be used to optimize mitigation policies that minimize economic impact without overwhelming hospital capacity. Our main contributions are (1) a novel agent-based pandemic simulator which, unlike traditional models, is able to model fine-grained interactions among people at specific locations in a community; (2) an RLbased methodology for optimizing fine-grained mitigation policies within this simulator; and (3) a Hidden Markov Model for predicting infected individuals based on partial observations regarding test results, presence of symptoms, and past physical contacts. This article is part of the special track on AI and COVID-19.
Roberto Capobianco, Varun Raj Kompella, James Ault, Guni Sharon, Stacy Jong, Spencer J. Fox, Lauren Ancel Meyers, Peter R. Wurman, Peter Stone 0001
J. Artif. Intell. Res.8
2018 Analysis and Observations From the First Amazon Picking Challenge
abstract
This paper presents an overview of the inaugural Amazon Picking Challenge along with a summary of a survey conducted among the 26 participating teams. The challenge goal was to design an autonomous robot to pick items from a warehouse shelf. This task is currently performed by human workers, and there is hope that robots can someday help increase efficiency and throughput while lowering cost. We report on a 28-question survey posed to the teams to learn about each team's background, mechanism design, perception apparatus, planning, and control approach. We identify trends in this data, correlate it with each team's success in the competition, and discuss observations and lessons learned based on survey results and the authors' personal experiences during the challenge.
Nikolaus Correll, Kostas E. Bekris, Dmitry Berenson, Oliver Brock, Albert J. Causo, Kris Hauser, Kei Okada, Alberto Rodriguez 0003, Joseph M. Romano, Peter R. Wurman
IEEE Trans Autom. Sci. Eng.10
2008 PBA*: Using Proactive Search to Make A* Robust to Unplanned Deviations
Paul Breimyer, Peter R. Wurman
AAAI2
2007 The game of scale: decision making with economies of scale
abstract
While diffusion of innovation topics in economics and majority games in game theory have been widely studied, the impact of economy-of-scale effects in aggregated decision making has received little attention. In this paper, we present a basic model, the Game of Scale, to study the effects of economy-of-scale in decision making among a large pool of self-interested agents. We solve the model's static equilibria and present two dynamic decision models, one myopic and one trend-following. Most of the parameter space converges quickly; however, the behaviors exhibited near critical input values show drastic changes. We demonstrate how trend-following can improve global outcomes over myopic decision making. Finally, we describe how the game can be externally controlled.
Christopher J. Hazard, Peter R. Wurman
ICEC2
2007 Coordinating Hundreds of Cooperative, Autonomous Vehicles in Warehouses
Peter R. Wurman, Raffaello D'Andrea, Mick Mountz
AAAI1
2005 Applying metaheuristic techniques to search the space of bidding strategies in combinatorial auctions
abstract
Many non-cooperative settings that could potentially be studied using game theory are characterized by having very large strategy spaces and payoffs that are costly to compute. Best response dynamics is a method of searching for pure-strategy equilibria in games that is attractive for its simplicity and scalability (relative to more analytical approaches). However, when the cost of determining the outcome of a particular set of joint strategies is high, it is impractical to compute the payoffs of all possible responses to the other players actions. Thus, we study metaheuristic approaches--genetic algorithms and tabu search in particular--to explore the strategy space. We configure the parameters of metaheuristics to adapt to the problem of finding the best response strategy and present how it can be helpful in finding Nash equilibria of combinatorial auctions which is an important solution concept in game theory.
Ashish Sureka, Peter R. Wurman
GECCO2
2005 Monte Carlo approximation in incomplete information, sequential auction games
Gangshu (George) Cai, Peter R. Wurman
Decis. Support Syst.2
2003 A comparison of two algorithms for multi-unit k-double auctions
abstract
We develop two algorithms to manage bid data in flexible, multi-unit double auctions. The first algorithm is a multi-unit extension of the 4-HEAP algorithm, and the second is a novel algorithm based on the Internal Path Reduction tree. To facilitate the generation of price quotes, we enhance the IPR tree algorithm to maintain information about the current Mth and (M + 1)st units. Our experiments show that when bids are relatively unordered, AUC-IPR outperforms 4-HEAP by a factor of about 2.6. However, under some scenarios in which bids are received in order, AUC-IPR is not always better than 4-HEAP.
Shengli Bao, Peter R. Wurman
ICEC2
2003 An algorithm for computing the outcome of combinatorial auctions with proxy bidding
abstract
Proxy bidding has proved useful in a variety of real auction formats, such as eBay, and has been proposed for some combinatorial auctions. Previous work on proxy bidding in combinatorial auctions requires the auctioneer essentially run the auction with myopic bidders to determine the outcome. In addition to being computationally costly, this process is only as accurate as the bid increment, and decreasing the bid increment to improve accuracy greatly increases the running time. In this paper, we present an algorithm that computes the outcome of the proxy auction by examining only the events that cause the proxy bidders to change their behaviors. This algorithm is much faster than the alternative, and computes exact solutions.
Peter R. Wurman, Gangshu (George) Cai, Ashish Sureka
ICEC1
2003 Computing the outcome of proxy bidding in combinatorial auctions extended abstract
abstract
No abstract available.
Peter R. Wurman, Gangshu (George) Cai, Ashish Sureka
EC1
2000 AkBA: a progressive, anonymous-price combinatorial auction
abstract
The allocation of discrete, complementary resources is a fundamental problem in economics and of direct interest to e-commerce applications.Combinatorial auctions account for complementarities by optimizing over oers expressed in terms of bundles.Progressive v ersions of combinatorial auctions alleviate the burden on bidders of expressing offers for all bundles of interest by p r o viding interim feedback based on partial sets of bids.Feedback i n t e r m s o f h ypothetical prices is particularly useful, as it directs bidders toward those bundles potentially yielding the greatest surplus.For a general class of discrete resource allocation problems with free disposal, we establish by construction the existence of competitive equilibrium prices on bundles that support the eÆcient allocation.We i n troduce AkBA, a family of progressive auctions that use these equilibrium bundle prices.We examine a particular instance of the family, called A1BA, and present some empirical data on its performance.
Peter R. Wurman, Michael P. Wellman
EC1
1998 Some Economics of Market-Based Distributed Scheduling
abstract
Market mechanisms solve distributed scheduling problems by allocating the scheduled resources according to market prices. We model distributed scheduling as a discrete resource allocation problem, and demonstrate the applicability of economic analysis to this framework. Drawing on results from the literature, we discuss the existence of equilibrium prices for some general classes of scheduling problems, and the quality of equilibrium solutions. We then present two auction protocols for implementing solutions, and analyze their computational and economic properties.
William E. Walsh, Michael P. Wellman, Peter R. Wurman, Jeffrey K. MacKie-Mason
ICDCS3
1998 Flexible double auctions for electronic commerce: theory and implementation
Peter R. Wurman, William E. Walsh, Michael P. Wellman
Decis. Support Syst.1
1996 Optimal Factory Scheduling using Stochastic Dominance A*
Peter R. Wurman, Michael P. Wellman
UAI1