Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Gonzalo Gonzalez-Pumariega

dblp:271/8258 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0004-5425-7319ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Planning, search and constraint satisfaction · 53% Reinforcement learning · 23% Robot manipulation · 17%
Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 38% Performance modeling and evaluation · 38% Processor architecture and microarchitecture · 23%

Topics — the 8 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
agent planning
0.912025
Robotouille: An Asynchronous Planning Benchmark for LLM Agents · ICLR 2025
Machine learning › Reinforcement learning
multi-turn reinforcement learning
0.912025
Multi-Turn Code Generation Through Single-Step Rewards · ICML 2025
Cloud and datacenter computing
cluster resource management and scheduling
0.412020
CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores · MICRO 2020
Performance modeling and evaluation
workload characterization
0.412020
CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores · MICRO 2020
Natural language and speech › Language models and text generation
code generation
0.312025
Multi-Turn Code Generation Through Single-Step Rewards · ICML 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › multi-agent planning
collaborative planning
0.312025
Robotouille: An Asynchronous Planning Benchmark for LLM Agents · ICLR 2025
Processor architecture and microarchitecture
multicore design
0.112020
CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores · MICRO 2020
Processor architecture and microarchitecture › chip multiprocessor
reconfigurable multicore
0.112020
CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores · MICRO 2020

Methods — techniques the papers use, named apart from their topics

large language model · 2.2reinforcement learning · 1.7execution feedback · 1.7recursive summarization · 1.3chain-of-thought prompting · 1.3verifier models · 0.9verifier model · 0.9react · 0.9dynamically dimensioned search · 0.4data mining · 0.4collaborative filtering · 0.4
YearPublicationVenuePosition
2025 Robotouille: An Asynchronous Planning Benchmark for LLM Agents
abstract
Effective asynchronous planning, or the ability to efficiently reason and plan over states and actions that must happen in parallel or sequentially, is essential for agents that must account for time delays, reason over diverse long-horizon tasks, and collaborate with other agents. While large language model (LLM) agents show promise in high-level task planning, current benchmarks focus primarily on short-horizon tasks and do not evaluate such asynchronous planning capabilities. We introduce Robotouille, a challenging benchmark environment designed to test LLM agents' ability to handle long-horizon asynchronous scenarios. Our synchronous and asynchronous datasets capture increasingly complex planning challenges that go beyond existing benchmarks, requiring agents to manage over- lapping tasks and interruptions Our results show that ReAct (gpt-4o) achieves 47% on synchronous tasks but only 11% on asynchronous tasks, highlighting significant room for improvement. We further analyze failure modes, demonstrating the need for LLM agents to better incorporate long-horizon feedback and self-audit their reasoning during task execution.
Gonzalo Gonzalez-Pumariega, Leong Su Yean, Neha Sunkara, Sanjiban Choudhury
ICLR1
2025 Multi-Turn Code Generation Through Single-Step Rewards
abstract
We address the problem of code generation from multi-turn execution feedback. Existing methods either generate code without feedback or use complex, hierarchical reinforcement learning to optimize multi-turn rewards. We propose a simple yet scalable approach, $\mu$CODE, that solves multi-turn code generation using only single-step rewards. Our key insight is that code generation is a one-step recoverable MDP, where the correct code can be recovered from any intermediate code state in a single turn. $\mu$CODE iteratively trains both a generator to provide code solutions conditioned on multi-turn execution feedback and a verifier to score the newly generated code. Experimental evaluations show that our approach achieves significant improvements over state-of-the-art baselines. We provide analysis of the design choices of the reward models and policy, and show the efficacy of $\mu$CODE at utilizing the execution feedback.
Arnav Kumar Jain, Gonzalo Gonzalez-Pumariega, Wayne Chen, Alexander M. Rush, Sanjiban Choudhury
ICML2
2024 Affinity Diagramming with a Robot
abstract
We investigate what it might look like for a robot to work with a human on a need-finding design task using an affinity diagram. While some recent projects have examined how human–robot teams might explore solutions to design problems, human–robot collaboration in the sensemaking aspects of the design process has not been studied. Designers use affinity diagrams to make sense of unstructured information by clustering paper notes on a work surface. To explore human–robot collaboration on a sensemaking design activity, we developed HIRO, an autonomous robot that constructs affinity diagrams with humans. In a within-user study, 56 participants affinity-diagrammed themes to characterize needs in quotes taken from real-world user data, once alone and once with HIRO. Users spent more time on the task with HIRO than alone, without strong evidence for corresponding effects on cognitive load. In addition, a majority of participants said they preferred to work with HIRO. From post-interaction interviews, we identified eight themes leading to four guidelines for robots that collaborate with humans on sensemaking design tasks: (1) account for the robot’s speed, (2) pursue mutual understanding rather than just correctness, (3) identify opportunities for constructive disagreements, and (4) use other modes of communication in addition to physical materials.
Matthew V. Law, Nnamdi Nwagwu, Amritansh Kwatra, Daniel M. Diangelis, Naifang Yu, Gonzalo Gonzalez-Pumariega, Amit Rajesh, Guy Hoffman
ACM Trans. Hum. Robot Interact.7
2023 Demo2Code: From Summarizing Demonstrations to Synthesizing Code via Extended Chain-of-Thought
abstract
Language instructions and demonstrations are two natural ways for users to teach robots personalized tasks. Recent progress in Large Language Models (LLMs) has shown impressive performance in translating language instructions into code for robotic tasks. However, translating demonstrations into task code continues to be a challenge due to the length and complexity of both demonstrations and code, making learning a direct mapping intractable. This paper presents Demo2Code, a novel framework that generates robot task code from demonstrations via an extended chain-of-thought and defines a common latent specification to connect the two. Our framework employs a robust two-stage process: (1) a recursive summarization technique that condenses demonstrations into concise specifications, and (2) a code synthesis approach that expands each function recursively from the generated specifications. We conduct extensive evaluation on various robot task benchmarks, including a novel game benchmark Robotouille, designed to simulate diverse cooking tasks in a kitchen environment.
Yuki Wang, Gonzalo Gonzalez-Pumariega, Sanjiban Choudhury
NeurIPS2
2020 CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores
abstract
Multi-tenancy for latency-critical applications leads to resource interference and unpredictable performance. Core reconfiguration opens up more opportunities for application colocation, as it allows the hardware to adjust to the dynamic performance and power needs of a specific mix of co-scheduled services. However, reconfigurability also introduces challenges, as even for a small number of reconfigurable cores, exploring the design space becomes more time- and resource-demanding.We present CuttleSys, a runtime for reconfigurable multicores that leverages scalable and lightweight data mining to quickly identify suitable core and cache configurations for a set of co-scheduled applications. The runtime combines collaborative filtering to infer the behavior of each job on every core and cache configuration, with Dynamically Dimensioned Search to efficiently explore the configuration space. We evaluate CuttleSys on multicores with tens of reconfigurable cores and show up to 2.46× and 1.55× performance improvements compared to core-level gating and oracle-like asymmetric multicores respectively, under stringent power constraints.
Neeraj Kulkarni, Gonzalo Gonzalez-Pumariega, Amulya Khurana, Christine A. Shoemaker, Christina Delimitrou, David H. Albonesi
MICRO2