So Yeon Min

dblp:78/84 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
8since 2021 · last 2024
0000-0001-7466-3948ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 44% Robot navigation and mapping · 17% Planning, search and constraint satisfaction · 16%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 15 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
instruction following
1.932024
Situated Instruction Following · ECCV (61) 2024
FILM: Following Instructions in Language with Modular Methods · ICLR 2022
Don't Copy the Teacher: Data and Model Challenges in Embodied Dialogue · EMNLP 2022
Robotics › Robot navigation and mapping
embodied instruction following
0.812024
Situated Instruction Following · ECCV (61) 2024
Robotics › Robot navigation and mapping
social navigation
0.812024
Habitat 3.0: A Co-Habitat for Humans, Avatars, and Robots · ICLR 2024
Natural language and speech › Language models and text generation › LLM agents
tool use
0.812024
Tools Fail: Detecting Silent Errors in Faulty Tools · EMNLP 2024
Natural language and speech › Language models and text generation › prompting
chain-of-thought prompting
0.712023
SPRING: Studying Papers and Reasoning to play Games · NeurIPS 2023
Machine learning › Reinforcement learning › exploration
embodied exploration
0.712023
EXCALIBUR: Encouraging and Evaluating Embodied Exploration · CVPR 2023
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering
0.712023
EXCALIBUR: Encouraging and Evaluating Embodied Exploration · CVPR 2023
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game playing
0.712023
SPRING: Studying Papers and Reasoning to play Games · NeurIPS 2023
Natural language and speech › Language models and text generation › in-context learning
in-context reasoning
0.712023
SPRING: Studying Papers and Reasoning to play Games · NeurIPS 2023
Natural language and speech › Language models and text generation
large language model reasoning
0.712023
SPRING: Studying Papers and Reasoning to play Games · NeurIPS 2023
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
embodied conversational agents
0.612022
Don't Copy the Teacher: Data and Model Challenges in Embodied Dialogue · EMNLP 2022
Robotics › Robot navigation and mapping
embodied AI simulation
0.212024
Habitat 3.0: A Co-Habitat for Humans, Avatars, and Robots · ICLR 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › agent planning
embodied planning
0.212024
Tools Fail: Detecting Silent Errors in Faulty Tools · EMNLP 2024
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
interactive agents
0.212023
EXCALIBUR: Encouraging and Evaluating Embodied Exploration · CVPR 2023
Machine learning › Reinforcement learning
imitation learning
0.212022
Don't Copy the Teacher: Data and Model Challenges in Embodied Dialogue · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

humanoid simulation · 1.5VR interface · 1.5large language model · 1.4multimodal grounding · 0.8learned policy · 0.8learned policies · 0.8virtual reality interface · 0.7directed acyclic graph · 0.7chain-of-thought prompting · 0.7imitation learning · 0.6empirical comparison · 0.6
YearPublicationVenuePosition
2024 Situated Instruction Following
So Yeon Min, Xavier Puig, Devendra Singh Chaplot, Tsung-Yen Yang, Akshara Rai, Priyam Parashar, Ruslan Salakhutdinov, Yonatan Bisk, Roozbeh Mottaghi
ECCV (61)1
2024 Tools Fail: Detecting Silent Errors in Faulty Tools
abstract
Tools have become a mainstay of LLMs, allowing them to retrieve knowledge not in their weights, to perform tasks on the web, and even to control robots.However, most ontologies and surveys of tool-use have assumed the core challenge for LLMs is choosing the tool.Instead, we introduce a framework for tools more broadly which guides us to explore a model's ability to detect "silent" tool errors, and reflect on how to plan.This more directly aligns with the increasingly popular use of models as tools.We provide an initial approach to failure recovery with promising results both on a controlled calculator setting and embodied agent planning.1
Jimin Sun, So Yeon Min, Yingshan Chang, Yonatan Bisk
EMNLP2
2024 Habitat 3.0: A Co-Habitat for Humans, Avatars, and Robots
abstract
We present Habitat 3.0: a simulation platform for studying collaborative human-robot tasks in home environments. Habitat 3.0 offers contributions across three dimensions: (1) Accurate humanoid simulation: addressing challenges in modeling complex deformable bodies and diversity in appearance and motion, all while ensuring high simulation speed. (2) Human-in-the-loop infrastructure: enabling real human interaction with simulated robots via mouse/keyboard or a VR interface, facilitating evaluation of robot policies with human input. (3) Collaborative tasks: studying two collaborative tasks, Social Navigation and Social Rearrangement. Social Navigation investigates a robot's ability to locate and follow humanoid avatars in unseen environments, whereas Social Rearrangement addresses collaboration between a humanoid and robot while rearranging a scene. These contributions allow us to study end-to-end learned and heuristic baselines for human-robot collaboration in-depth, as well as evaluate them with humans in the loop. Our experiments demonstrate that learned robot policies lead to efficient task completion when collaborating with unseen humanoid agents and human partners that might exhibit behaviors that the robot has not seen before. Additionally, we observe emergent behaviors during collaborative task execution, such as the robot yielding space when obstructing a humanoid agent, thereby allowing the effective completion of the task by the humanoid agent. Furthermore, our experiments using the human-in-the-loop tool demonstrate that our automated evaluation with humanoids can provide an indication of the relative ordering of different policies when evaluated with real human collaborators. Habitat 3.0 unlocks interesting new features in simulators for Embodied AI, and we hope it paves the way for a new frontier of embodied human-AI interaction capabilities. For more details and visualizations, visit: https://aihabitat.org/habitat3.
Xavier Puig, Eric Undersander, Andrew Szot, Mikael Dallaire Cote, Tsung-Yen Yang, Ruslan Partsey, Ruta Desai, Alexander Clegg, Michal Hlavac, So Yeon Min, Vladimir Vondrus, Théophile Gervet, Vincent-Pierre Berges, John M. Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakrishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara Rai, Roozbeh Mottaghi
ICLR10
2023 EXCALIBUR: Encouraging and Evaluating Embodied Exploration
abstract
Experience precedes understanding. Humans constantly explore and learn about their environment out of curiosity, gather information, and update their models of the world. On the other hand, machines are either trained to learn passively from static and fixed datasets, or taught to complete specific goal-conditioned tasks. To encourage the development of exploratory interactive agents, we present the EXCALIBUR benchmark. EXCALIBUR allows agents to explore their environment for long durations and then query their understanding of the physical world via inquiries like: “is the small heavy red bowl made from glass?” or “is there a silver spoon heavier than the egg?”. This design encourages agents to perform free-form home exploration without myopia induced by goal conditioning. Once the agents have answered a series of questions, they can renter the scene to refine their knowledge, update their beliefs, and improve their performance on the questions. Our experiments demonstrate the challenges posed by this dataset for the present-day state-of-the-art embodied systems and the headroom afforded to develop new innovative methods. Finally, we present a virtual reality interface that enables humans to seamlessly interact within the simulated world and use it to gather human performance measures. EXCALIBUR affords unique challenges in comparison to presentday benchmarks and represents the next frontier for embodied AI research.
Hao Zhu 0011, Raghav Kapoor, So Yeon Min, Winson Han, Jiatai Li, Kaiwen Geng, Graham Neubig, Yonatan Bisk, Aniruddha Kembhavi, Luca Weihs
CVPR3
2023 Self-Supervised Object Goal Navigation with In-Situ Finetuning
abstract
A household robot should be able to navigate to target objects without requiring users to first annotate everything in their home. Most current approaches to object navigation do not test on real robots and rely solely on reconstructed scans of houses and their expensively labeled semantic 3D meshes. In this work, our goal is to build an agent that builds self-supervised models of the world via exploration, the same as a child might - thus we (1) eschew the expense of labeled 3D mesh and (2) enable self-supervised in-situ finetuning in the real world. We identify a strong source of self-supervision (Location Consistency - LocCon) that can train all components of an ObjectNav agent, using unannotated simulated houses. Our key insight is that embodied agents can leverage location consistency as a self-supervision signal - collecting images from different views/angles and applying contrastive learning. We show that our agent can perform competitively in the real world and simulation. Our results also indicate that supervised training with 3D mesh annotations causes models to learn simulation artifacts, which are not transferrable to the real world. In contrast, our LocCon shows the most robust transfer in the real world among the set of models we compare to, and that the real-world performance of all models can be further improved with self-supervised LocCon in-situ training.
So Yeon Min, Yao-Hung Tsai, Ali Farhadi, Ruslan Salakhutdinov, Yonatan Bisk, Jian Zhang 0050
IROS1
2023 SPRING: Studying Papers and Reasoning to play Games
abstract
Open-world survival games pose significant challenges for AI algorithms due to their multi-tasking, deep exploration, and goal prioritization requirements. Despite reinforcement learning (RL) being popular for solving games, its high sample complexity limits its effectiveness in complex open-world games like Crafter or Minecraft. We propose a novel approach, SPRING, to read Crafter's original academic paper and use the knowledge learned to reason and play the game through a large language model (LLM). Prompted with the LaTeX source as game context and a description of the agent's current observation, our SPRING framework employs a directed acyclic graph (DAG) with game-related questions as nodes and dependencies as edges. We identify the optimal action to take in the environment by traversing the DAG and calculating LLM responses for each node in topological order, with the LLM's answer to final node directly translating to environment actions. In our experiments, we study the quality of in-context "reasoning" induced by different forms of prompts under the setting of the Crafter environment. Our experiments suggest that LLMs, when prompted with consistent chain-of-thought, have great potential in completing sophisticated high-level trajectories. Quantitatively, SPRING with GPT-4 outperforms all state-of-the-art RL baselines, trained for 1M steps, without any training. Finally, we show the potential of Crafter as a test bed for LLMs. Code at github.com/holmeswww/SPRING
Yue Wu 0001, So Yeon Min, Shrimai Prabhumoye, Yonatan Bisk, Ruslan Salakhutdinov, Amos Azaria, Tom M. Mitchell, Yuanzhi Li
NeurIPS2
2022 Don't Copy the Teacher: Data and Model Challenges in Embodied Dialogue
abstract
Embodied dialogue instruction following requires an agent to complete a complex sequence of tasks from a natural language exchange.The recent introduction of benchmarks (Padmakumar et al., 2022) raises the question of how best to train and evaluate models for this multi-turn, multi-agent, long-horizon task.This paper contributes to that conversation, by arguing that imitation learning (IL) and related low-level metrics are actually misleading and do not align with the goals of embodied dialogue research and may hinder progress.We provide empirical comparisons of metrics, analysis of three models, and make suggestions for how the field might best progress.First, we observe that models trained with IL take spurious actions during evaluation.Second, we find that existing models fail to ground query utterances, which are essential for task completion.Third, we argue evaluation should focus on higher-level semantic goals. 1
So Yeon Min, Hao Zhu 0011, Ruslan Salakhutdinov, Yonatan Bisk
EMNLP1
2022 FILM: Following Instructions in Language with Modular Methods
So Yeon Min, Devendra Singh Chaplot, Pradeep Ravikumar, Yonatan Bisk, Ruslan Salakhutdinov
ICLR1
2007 An Improvement of the Processing Delay for the G.723.1 Vocoder
Kwang-Hyoung Lee, So Yeon Min, Keun-Wang Lee, Jeong Gyu Jee
MMM (2)2
2006 A Design Technique of CBD Meta-model Based on Graph Theory
Eun Sook Cho, So Yeon Min, Chul Jin Kim
ICCSA (4)2
2006 A Study on the Pitch Extraction Detection by Linear Approximation of Sub-band
Keun-Wang Lee, Kwang-Hyoung Lee, So Yeon Min
ICCSA (2)3
2006 High Speed Codebook Searching Algorithm for the CELP Vocoder in the Internet-Based Environment
So Yeon Min, Eun Sook Cho, Chul Jin Kim
ICCSA (2)1