Minae Kwon

dblp:177/8692 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
4since 2021 · last 2024
0000-0002-3116-330XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-authorSystems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 33% Motion planning and robot control · 23% Knowledge representation and reasoning · 10%
Human-computer interaction and pervasive computing
3 papers
Human-robot interaction · 100%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
active perception
0.812024
Toward Grounded Commonsense Reasoning · ICRA 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.812024
Toward Grounded Commonsense Reasoning · ICRA 2024
Robotics › Motion planning and robot control › robot learning › manipulation learning
language-conditioned manipulation
0.812024
Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections · ICRA 2024
Robotics › Robot manipulation › learning from demonstration
learning from corrections
0.812024
Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections · ICRA 2024
Robotics › Motion planning and robot control
robot learning
0.812024
Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections · ICRA 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
task planning
0.812024
Toward Grounded Commonsense Reasoning · ICRA 2024
Machine learning › Reinforcement learning › reward learning
LLM-based reward generation
0.712023
Reward Design with Language Models · ICLR 2023
Machine learning › Reinforcement learning
reward design
0.712023
Reward Design with Language Models · ICLR 2023
Knowledge, reasoning and agents › Multi-agent systems
automated negotiation
0.512021
Targeted Data Acquisition for Evolving Negotiation Agents · ICML 2021
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.512021
Targeted Data Acquisition for Evolving Negotiation Agents · ICML 2021
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.512021
Targeted Data Acquisition for Evolving Negotiation Agents · ICML 2021
Human-robot interaction
human modeling
0.412020
When Humans Aren't Optimal: Robots that Collaborate with Risk-Aware Humans · HRI 2020
Human-robot interaction
human-robot collaboration
0.412020
When Humans Aren't Optimal: Robots that Collaborate with Risk-Aware Humans · HRI 2020
Robotics › Motion planning and robot control
trajectory optimization
0.312018
Expressing Robot Incapability · HRI 2018
Human-robot interaction › cognitive human-robot interaction
mental models of robots
0.212016
Human Expectations of Social Robots · HRI 2016
Machine learning › Reinforcement learning
imitation learning
0.112021
Targeted Data Acquisition for Evolving Negotiation Agents · ICML 2021
Machine learning › Reinforcement learning
human behavior modeling
0.112020
When Humans Aren't Optimal: Robots that Collaborate with Risk-Aware Humans · HRI 2020

Methods — techniques the papers use, named apart from their topics

user study · 1.8large language model · 1.5planning · 0.9cumulative prospect theory · 0.9vision-language model · 0.8retrieval-augmented generation · 0.8trajectory optimization · 0.7language model · 0.7supervised learning · 0.5expert oracle · 0.5
YearPublicationVenuePosition
2024 Toward Grounded Commonsense Reasoning
abstract
Consider a robot tasked with tidying a desk with a meticulously constructed Lego sports car. A human may recognize that it is not appropriate to disassemble the sports car and put it away as part of the "tidying." How can a robot reach that conclusion? Although large language models (LLMs) have recently been used to enable commonsense reasoning, grounding this reasoning in the real world has been challenging. To reason in the real world, robots must go beyond passively querying LLMs and actively gather information from the environment that is required to make the right decision. For instance, after detecting that there is an occluded car, the robot may need to actively perceive the car to know whether it is an advanced model car made out of Legos or a toy car built by a toddler. We propose an approach that leverages an LLM and vision language model (VLM) to help a robot actively perceive its environment to perform grounded commonsense reasoning. To evaluate our framework at scale, we release the MessySurfaces dataset which contains images of 70 real-world surfaces that need to be cleaned. We additionally illustrate our approach with a robot on 2 carefully designed surfaces. We find an average 12.9% improvement on the MessySurfaces benchmark and an average 15% improvement on the robot experiments over baselines that do not use active perception. The dataset, code, and videos of our approach can be found at https://minaek.github.io/grounded_commonsense_reasoning/.
Minae Kwon, Hengyuan Hu, Vivek Myers, Siddharth Karamcheti, Anca D. Dragan, Dorsa Sadigh
ICRA1
2024 Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections
abstract
Today’s robot policies exhibit subpar performance when faced with the challenge of generalizing to novel environments. Human corrective feedback is a crucial form of guidance to enable such generalization. However, adapting to and learning from online human corrections is a non-trivial endeavor: not only do robots need to remember human feedback over time to retrieve the right information in new settings and reduce the intervention rate, but also they would need to be able to respond to feedback that can be arbitrary corrections about high-level human preferences to low-level adjustments to skill parameters. In this work, we present Distillation and Retrieval of Online Corrections (DROC), a large language model (LLM)-based system that can respond to arbitrary forms of language feedback, distill generalizable knowledge from corrections, and retrieve relevant past experiences based on textual and visual similarity for improving performance in novel settings. DROC is able to respond to a sequence of online language corrections that address failures in both high-level task plans and low-level skill primitives. We demonstrate that DROC effectively distills the relevant information from the sequence of online corrections in a knowledge base and retrieves that knowledge in settings with new task or object instances. DROC outperforms other techniques that directly generate robot code via LLMs [1] by using only half of the total number of corrections needed in the first round and requires little to no corrections after two iterations. We show further results and videos on our project website: https://sites.google.com/stanford.edu/droc.
Lihan Zha, Yuchen Cui, Li-Heng Lin, Minae Kwon, Montse Gonzalez Arenas, Andy Zeng 0001, Fei Xia 0002, Dorsa Sadigh
ICRA4
2023 Reward Design with Language Models
Minae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa Sadigh
ICLR1
2021 Targeted Data Acquisition for Evolving Negotiation Agents
abstract
Successful negotiators must learn how to balance optimizing for self-interest and cooperation. Yet current artificial negotiation agents often heavily depend on the quality of the static datasets they were trained on, limiting their capacity to fashion an adaptive response balancing self-interest and cooperation. For this reason, we find that these agents can achieve either high utility or cooperation, but not both. To address this, we introduce a targeted data acquisition framework where we guide the exploration of a reinforcement learning agent using annotations from an expert oracle. The guided exploration incentivizes the learning agent to go beyond its static dataset and develop new negotiation strategies. We show that this enables our agents to obtain higher-reward and more Pareto-optimal solutions when negotiating with both simulated and human partners compared to standard supervised learning and reinforcement learning methods. This trend additionally holds when comparing agents using our targeted data acquisition framework to variants of agents trained with a mix of supervised learning and reinforcement learning, or to agents using tailored reward functions that explicitly optimize for utility and Pareto-optimality.
Minae Kwon, Siddharth Karamcheti, Mariano-Florentino Cuellar, Dorsa Sadigh
ICML1
2020 Continual Adaptation for Efficient Machine Communication
abstract
To communicate with new partners in new contexts, humans rapidly form new linguistic conventions.Recent neural language models are able to comprehend and produce the existing conventions present in their training data, but are not able to flexibly and interactively adapt those conventions on the fly as humans do.We introduce an interactive repeated reference task as a benchmark for models of adaptation in communication and propose a regularized continual learning framework that allows an artificial agent initialized with a generic language model to more accurately and efficiently communicate with a partner over time.We evaluate this framework through simulations on COCO and in real-time reference game experiments with human partners.
Robert D. Hawkins, Minae Kwon, Dorsa Sadigh, Noah D. Goodman
CoNLL2
2020 When Humans Aren't Optimal: Robots that Collaborate with Risk-Aware Humans
abstract
In order to collaborate safely and efficiently, robots need to anticipate how their human partners will behave. Some of today's robots model humans as if they were also robots, and assume users are always optimal. Other robots account for human limitations, and relax this assumption so that the human is noisily rational. Both of these models make sense when the human receives deterministic rewards: i.e., gaining either $100 or $130 with certainty. But in real-world scenarios, rewards are rarely deterministic. Instead, we must make choices subject to risk and uncertainty-and in these settings, humans exhibit a cognitive bias towards suboptimal behavior. For example, when deciding between gaining $100 with certainty or $130 only 80% of the time, people tend to make the risk-averse choice-even though it leads to a lower expected gain! In this paper, we adopt a well-known Risk-Aware human model from behavioral economics called Cumulative Prospect Theory and enable robots to leverage this model during human-robot interaction (HRI). In our user studies, we offer supporting evidence that the Risk-Aware model more accurately predicts suboptimal human behavior. We find that this increased modeling accuracy results in safer and more efficient human-robot collaboration. Overall, we extend existing rational human models so that collaborative robots can anticipate and plan around suboptimal human behavior during HRI.
Minae Kwon, Erdem Biyik, Aditi Talati, Karan Bhasin, Dylan P. Losey, Dorsa Sadigh
HRI1
2018 Expressing Robot Incapability
abstract
Our goal is to enable robots to express their incapability, and to do so in a way that communicates both what they are trying to accomplish and why they are unable to accomplish it. We frame this as a trajectory optimization problem: maximize the similarity between the motion expressing incapability and what would amount to successful task execution, while obeying the physical limits of the robot. We introduce and evaluate candidate similarity measures, and show that one in particular generalizes to a range of tasks, while producing expressive motions that are tailored to each task. Our user study supports that our approach automatically generates motions expressing incapability that communicate both what and why to end-users, and improve their overall perception of the robot and willingness to collaborate with it in the future.
Minae Kwon, Sandy H. Huang, Anca D. Dragan
HRI1
2018 Planning with Verbal Communication for Human-Robot Collaboration
abstract
Human collaborators coordinate effectively their actions through both verbal and non-verbal communication. We believe that the the same should hold for human-robot teams. We propose a formalism that enables a robot to decide optimally between taking a physical action toward task completion and issuing an utterance to the human teammate. We focus on two types of utterances: verbal commands, where the robot asks the human to take a physical action, and state-conveying actions, where the robot informs the human about its internal state, which captures the information that the robot uses in its decision making. Human subject experiments show that enabling the robot to issue verbal commands is the most effective form of communicating objectives, while retaining user trust in the robot. Communicating information about the robot’s state should be done judiciously, since many participants questioned the truthfulness of the robot statements when the robot did not provide sufficient explanation about its actions.
Stefanos Nikolaidis, Minae Kwon, Jodi Forlizzi, Siddhartha S. Srinivasa
ACM Trans. Hum. Robot Interact.2
2016 Human Expectations of Social Robots
abstract
A key assumption that drives much of HRI research is that human robot collaboration can be improved by advancing a robot's capabilities. We argue that this assumption has potentially negative implications, as increasing social capabilities in robots can produce an expectations gap where humans develop unrealistically high expectations of social robots due to generalization from human mental models. By conducting two studies with 674 participants, we examine how people develop and adjust mental models of robots. We find that both a robot's physical appearance and its behavior influence how we form these models. This suggests it is possible for a robot to unintentionally manipulate a human into building an inaccurate mental model of its overall abilities simply by displaying a few capabilities that humans possess, such as speaking and turn-taking. We conclude that this expectations gap, if not corrected for, could ironically result in less effective collaborations as robot capabilities improve.
Minae Kwon, Malte F. Jung, Ross A. Knepper
HRI1