VLDB 2026 Research / reviewers in the wild / expert
Byeonghwi Kim
dblp:280/2943
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-3775-2778ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Robot navigation and mapping · 39% Language models and text generation · 15% Planning, search and constraint satisfaction · 15% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
instruction following |
1.6 | 3 | 2025 | Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents · ICCV 2023 Multi-Level Compositional Reasoning for Interactive Instruction Following · AAAI 2023 Multi-Modal Grounded Planning and Efficient Replanning for Learning Embodied Agents with a Few Examples · AAAI 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › agent planning
embodied planning |
1.5 | 2 | 2025 | Multi-Modal Grounded Planning and Efficient Replanning for Learning Embodied Agents with a Few Examples · AAAI 2025 Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents · ICCV 2023 |
Robotics › Robot navigation and mapping
embodied instruction following |
1.5 | 2 | 2024 | Online Continual Learning for Interactive Instruction Following Agents · ICLR 2024 ReALFRED: An Embodied Instruction Following Benchmark in Photo-Realistic Environments · ECCV (10) 2024 |
Robotics › Robot navigation and mapping
embodied navigation |
1.2 | 2 | 2023 | Multi-Level Compositional Reasoning for Interactive Instruction Following · AAAI 2023 Factorizing Perception and Policy for Interactive Instruction Following · ICCV 2021 |
Robotics › Robot navigation and mapping
embodied perception |
0.9 | 1 | 2025 | Multi-Modal Grounded Planning and Efficient Replanning for Learning Embodied Agents with a Few Examples · AAAI 2025 |
Machine learning › Reinforcement learning › non-stationary reinforcement learning
continual reinforcement learning |
0.8 | 1 | 2024 | Online Continual Learning for Interactive Instruction Following Agents · ICLR 2024 |
Computer vision › Vision and language › multimodal reasoning
compositional reasoning |
0.7 | 1 | 2023 | Multi-Level Compositional Reasoning for Interactive Instruction Following · AAAI 2023 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
hierarchical policy learning |
0.7 | 1 | 2023 | Multi-Level Compositional Reasoning for Interactive Instruction Following · AAAI 2023 |
Robotics › Motion planning and robot control
robot control |
0.7 | 1 | 2023 | Multi-Level Compositional Reasoning for Interactive Instruction Following · AAAI 2023 |
Robotics › Robot navigation and mapping › visual navigation
language-guided navigation |
0.5 | 1 | 2021 | Factorizing Perception and Policy for Interactive Instruction Following · ICCV 2021 |
Computer vision › Vision and language
vision-and-language navigation |
0.2 | 1 | 2024 | ReALFRED: An Embodied Instruction Following Benchmark in Photo-Realistic Environments · ECCV (10) 2024 |
Robotics › Robot manipulation › physical interaction
object interaction |
0.2 | 1 | 2023 | Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
visual grounding · 0.9large language model · 0.9few-shot learning · 0.9task-free continual learning · 0.8confidence-aware moving average · 0.8reinforcement learning · 0.7policy composition · 0.7imitation learning · 0.7algorithmic planning · 0.7modular architecture · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Modal Grounded Planning and Efficient Replanning for Learning Embodied Agents with a Few ExamplesabstractLearning a perception and reasoning module for robotic assistants to plan steps to perform complex tasks based on natural language instructions often requires large free-form language annotations, especially for short high-level instructions. To reduce the cost of annotation, large language models (LLMs) are used as a planner with few data. However, when elaborating the steps, even the state-of-the-art planner that uses LLMs mostly relies on linguistic common sense, often neglecting the status of the environment at command reception, resulting in inappropriate plans. To generate plans grounded in the environment, we propose FLARE (Few-shot Language with environmental Adaptive Replanning Embodied agent), which improves task planning using both language command and environmental perception. As language instructions often contain ambiguities or incorrect expressions, we additionally propose to correct the mistakes using visual cues from the agent. The proposed scheme allows us to use a few language pairs thanks to the visual cues and outperforms state-of-the-art approaches. Our code and the dataset are publicly available to facilitate further research. Taewoong Kim, Byeonghwi Kim |
AAAI | 2 |
| 2024 | ReALFRED: An Embodied Instruction Following Benchmark in Photo-Realistic Environments
Taewoong Kim, Cheolhong Min, Byeonghwi Kim, Jinyeon Kim, Wonje Jeung |
ECCV (10) | 3 |
| 2024 | Online Continual Learning for Interactive Instruction Following AgentsabstractIn learning an embodied agent executing daily tasks via language directives, the literature largely assumes that the agent learns all training data at the beginning. We argue that such a learning scenario is less realistic, since a robotic agent is supposed to learn the world continuously as it explores and perceives it. To take a step towards a more realistic embodied agent learning scenario, we propose two continual learning setups for embodied agents; learning new behaviors (Behavior Incremental Learning, Behavior-IL) and new environments (Environment Incremental Learning, Environment-IL) For the tasks, previous ‘data prior’ based continual learning methods maintain logits for the past tasks. However, the stored information is often insufficiently learned information and requires task boundary information, which might not always be available. Here, we propose to update them based on confidence scores without task boundary information (i.e., task-free) in a moving average fashion, named Confidence-Aware Moving Average (CAMA). In the proposed challenging Behavior-IL and Environment-IL setups, our simple CAMA outperforms prior arts in our empirical validations by noticeable margins. Byeonghwi Kim, Minhyuk Seo |
ICLR | 1 |
| 2023 | Multi-Level Compositional Reasoning for Interactive Instruction FollowingabstractRobotic agents performing domestic chores by natural language directives are required to master the complex job of navigating environment and interacting with objects in the environments. The tasks given to the agents are often composite thus are challenging as completing them require to reason about multiple subtasks, e.g., bring a cup of coffee. To address the challenge, we propose to divide and conquer it by breaking the task into multiple subgoals and attend to them individually for better navigation and interaction. We call it Multi-level Compositional Reasoning Agent (MCR-Agent). Specifically, we learn a three-level action policy. At the highest level, we infer a sequence of human-interpretable subgoals to be executed based on language instructions by a high-level policy composition controller. At the middle level, we discriminatively control the agent’s navigation by a master policy by alternating between a navigation policy and various independent interaction policies. Finally, at the lowest level, we infer manipulation actions with the corresponding object masks using the appropriate interaction policy. Our approach not only generates human interpretable subgoals but also achieves 2.03% absolute gain to comparable state of the arts in the efficiency metric (PLWSR in unseen set) without using rule-based planning or a semantic spatial memory. The code is available at https://github.com/yonseivnl/mcr-agent. Suvaansh Bhambri, Byeonghwi Kim |
AAAI | 2 |
| 2023 | Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied AgentsabstractAccomplishing household tasks requires to plan step-by-step actions considering the consequences of previous actions. However, the state-of-the-art embodied agents often make mistakes in navigating the environment and interacting with proper objects due to imperfect learning by imitating experts or algorithmic planners without such knowledge. To improve both visual navigation and object interaction, we propose to consider the consequence of taken actions by CAPEAM (Context-Aware Planning and Environment-Aware Memory) that incorporates semantic context (e.g., appropriate objects to interact with) in a sequence of actions, and the changed spatial arrangement and states of interacted objects (e.g., location that the object has been moved to) in inferring the subsequent actions. We empirically show that the agent with the proposed CAPEAM achieves state-of-the-art performance in various metrics using a challenging interactive instruction following benchmark in both seen and unseen environments by large margins (up to +10.70% in unseen env.). Byeonghwi Kim, Jinyeon Kim, Yuyeong Kim, Cheolhong Min |
ICCV | 1 |
| 2021 | Factorizing Perception and Policy for Interactive Instruction FollowingabstractPerforming simple household tasks based on language directives is very natural to humans, yet it remains an open challenge for AI agents. The ‘interactive instruction following’ task attempts to make progress towards building agents that jointly navigate, interact, and reason in the environment at every step. To address the multifaceted problem, we propose a model that factorizes the task into interactive perception and action policy streams with enhanced components and name it as MOCA, a Modular Object-Centric Approach. We empirically validate that MOCA outperforms prior arts by significant margins on the ALFRED benchmark with improved generalization. Kunal Pratap Singh, Suvaansh Bhambri, Byeonghwi Kim, Roozbeh Mottaghi |
ICCV | 3 |