Yen-Ling Kuo

dblp:120/3172 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
17since 2021 · last 2025
0000-0002-6433-6713ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 10 since 2021Systems, architecture and hardware · 6 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MuMA-ToM: Multi-modal Multi-Agent Theory of Mind
abstract
Understanding people's social interactions in complex real-world scenarios often relies on intricate mental reasoning. To truly understand how and why people interact with one another, we must infer the underlying mental states that give rise to the social interactions, i.e., Theory of Mind reasoning in multi-agent interactions. Additionally, social interactions are often multi-modal -- we can watch people's actions, hear their conversations, and/or read about their past behaviors. For AI systems to successfully and safely interact with people in real-world environments, they also need to understand people's mental states as well as their inferences about each other's mental states based on multi-modal information about their interactions. For this, we introduce MuMA-ToM, a Multi-modal Multi-Agent Theory of Mind benchmark. MuMA-ToM is the first multi-modal Theory of Mind benchmark that evaluates mental reasoning in embodied multi-agent interactions. In MuMA-ToM, we provide video and text descriptions of people's multi-modal behavior in realistic household environments. Based on the context, we then ask questions about people's goals, beliefs, and beliefs about others' goals. We validated MuMA-ToM in a human experiment and provided a human baseline. We also proposed a novel multi-modal, multi-agent ToM model, LIMP (Language model-based Inverse Multi-agent Planning). Our experimental results show that LIMP significantly outperforms state-of-the-art methods, including large multi-modal models (e.g., GPT-4o, Gemini-1.5 Pro) and a recent multi-modal ToM model, BIP-ALM.
Haojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin, Leyla Isik, Yen-Ling Kuo, Tianmin Shu
AAAI6
2025 Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual Foraging
abstract
Imagine searching a collection of coins for quarters (0.25), dimes (0.10), nickels (0.05), and pennies (0.01)—a hybrid foraging task where observers look for multiple instances of multiple target types. In such tasks, how do target values and their prevalence influence foraging and eye movement behaviors (e.g., should you prioritize rare quarters or common nickels)? To explore this, we conducted human psychophysics experiments, revealing that humans are proficient reward foragers. Their eye fixations are drawn to regions with higher average rewards, fixation durations are longer on more valuable targets, and their cumulative rewards exceed chance, approaching the upper bound of optimal foragers. To probe these decision-making processes of humans, we developed a transformer-based Visual Forager (VF) model trained via reinforcement learning. Our VF model takes a series of targets, their corresponding values, and the search image as inputs, processes the images using foveated vision, and produces a sequence of eye movements along with decisions on whether to collect each fixated item. Our model outperforms all baselines, achieves cumulative rewards comparable to those of humans, and approximates human foraging behavior in eye movements and foraging biases within time-limited environments. Furthermore, stress tests on out-of-distribution tasks with novel targets, unseen values, and varying set sizes demonstrate the VF model’s effective generalization. Our work offers valuable insights into the relationship between eye movements and decision-making, with our model serving as a powerful tool for further exploration of this connection. All data, code, and models are available at https://github.com/ZhangLab-DeepNeuroCogLab/visual-forager.
Dingwei Tan, Yen-Ling Kuo, Zhaowei Sun, Jeremy M. Wolfe, Tat-Jen Cham, Mengmi Zhang
CVPR3
2025 Diff-Dagger: Uncertainty Estimation With Diffusion Policy for Robotic Manipulation
abstract
Recently, diffusion policy has shown impressive results in handling multi-modal tasks in robotic manipulation. However, it has fundamental limitations in out-of-distribution failures that persist due to compounding errors and its limited capability to extrapolate. One way to address these limitations is robot-gated DAgger, an interactive imitation learning with a robot query system to actively seek expert help during policy rollout. While robot-gated DAgger has high potential for learning at scale, existing methods like Ensemble-DAgger struggle with highly expressive policies: They often misinterpret policy disagreements as uncertainty at multi-modal decision points. To address this problem, we introduce Diff-DAgger, an efficient robot-gated DAgger algorithm that leverages the training objective of diffusion policy. We evaluate Diff-DAgger across different robot tasks including stacking, pushing, and plugging, and show that Diff-DAgger improves the task failure prediction by 39.0 %, the task completion rate by 20.6 %, and reduces the wall-clock time by a factor of 7.8. We hope that this work opens up a path for efficiently incorporating expressive yet data-hungry policies into interactive robot learning settings. The project website is available at: https://diffdagger.github.io.
Sung-Wook Lee, Xuhui Kang, Yen-Ling Kuo
ICRA3
2025 Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits
abstract
Understanding the progress of a task allows humans to not only track what has been done but also to better plan for future goals. We demonstrate TaKSIE, a novel framework that incorporates task progress knowledge into visual sub-goal generation for robotic manipulation tasks. We jointly train a recurrent network with a latent diffusion model to generate the next visual subgoal based on the robot's cur-rent observation and the input language command. At exe-cution time, the robot leverages a visual progress represen-tation to monitor the task progress and adaptively samples the next visual subgoal from the model to guide the manip-ulation policy. We train and validate our model in simu-lated and real-world robotic tasks, achieving state-of-the-art performance on the CALVIN manipulation benchmark. We find that the inclusion of task progress knowledge can improve the robustness of trained policy for different initial robot poses or various movement speeds during demonstrations. The project page is available at https://live-robotics-uva.github.io/TaKSIE/.
Xuhui Kang, Yen-Ling Kuo
WACV2
2024 Neural Amortized Inference for Nested Multi-Agent Reasoning
abstract
Multi-agent interactions, such as communication, teaching, and bluffing, often rely on higher-order social inference, i.e., understanding how others infer oneself. Such intricate reasoning can be effectively modeled through nested multi-agent reasoning. Nonetheless, the computational complexity escalates exponentially with each level of reasoning, posing a significant challenge. However, humans effortlessly perform complex social inferences as part of their daily lives. To bridge the gap between human-like inference capabilities and computational limitations, we propose a novel approach: leveraging neural networks to amortize high-order social inference, thereby expediting nested multi-agent reasoning. We evaluate our method in two challenging multi-agent interaction domains. The experimental results demonstrate that our method is computationally efficient while exhibiting minimal degradation in accuracy.
Kunal Jha, Tuan Anh Le 0001, Chuanyang Jin, Yen-Ling Kuo, Josh Tenenbaum, Tianmin Shu
AAAI4
2024 Learning Representations for Robust Human-Robot Interaction
abstract
For robots to robustly and flexibly interact with humans, they need to acquire skills to use across scenarios. One way to enable the generalization of skills is to learn representations that are useful for downstream tasks. Learning a representation for interactions requires an understanding of what (e.g., objects) as well as how (e.g., actions, controls, and manners) to interact with. However, most existing language or visual representations mainly focus on objects. To enable robust human-robot interactions, we need a representation that is not just grounded at the object level but to reason at the action level. The ability to reason about an agent’s own actions and other’s actions will be crucial for long-tail interactions. My research focuses on leveraging the compositional nature of language and reward functions to learn representations that generalize to novel scenarios. Together with the information from multiple modalities, the learned representation can reason about task progress, future behaviors, and the goals/beliefs of an agent. The above ideas have been demonstrated in my research on building robots to understand language and engage in social interactions.
Yen-Ling Kuo
AAAI1
2024 MMToM-QA: Multimodal Theory of Mind Question Answering
abstract
Chuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer Ullman, Antonio Torralba, Joshua Tenenbaum, Tianmin Shu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Chuanyang Jin, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer D. Ullman, Antonio Torralba 0001, Josh Tenenbaum, Tianmin Shu
ACL (1)5
2024 Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation
abstract
We study object interaction anticipation in egocentric videos. This task requires an understanding of the spatio-temporal context formed by past actions on objects, coined action context. We propose TransFusion, a multimodal transformer-based architecture for short-term object interaction anticipation. Our method exploits the representational power of language by summarizing the action con-text textually, after leveraging pre-trained vision-language foundation models to extract the action context from past video frames. The summarized action context and the last observed video frame are processed by the multimodal fusion module to forecast the next object interaction. Experiments on the Ego4D next active object interaction dataset show the effectiveness of our multimodal fusion model and highlight the benefits of using the power of foundation models and language-based context summaries in a task where vision may appear to suffice. Our novel approach outperforms all state-of-the-art methods on both versions of the Ego4D dataset. A project video and code are available at https://eth-ait.github.io/transfusion-proj/.
Razvan-George Pasca, Alexey Gavryushin, Yen-Ling Kuo, Kaichun Mo, Luc Van Gool, Otmar Hilliges, Xi Wang 0021
CVPR4
2024 A Transformer-Based Model for the Prediction of Human Gaze Behavior on Videos
abstract
Eye-tracking applications that utilize the human gaze in video understanding tasks have become increasingly important. To effectively automate the process of video analysis based on eye-tracking data, it is important to accurately replicate human gaze behavior. However, this task presents significant challenges due to the inherent complexity and ambiguity of human gaze patterns. In this work, we introduce a novel method for simulating human gaze behavior. Our approach uses a transformer-based reinforcement learning algorithm to train an agent that acts as a human observer, with the primary role of watching videos and simulating human gaze behavior. We employed an eye-tracking dataset gathered from videos generated by the VirtualHome simulator, with a primary focus on activity recognition. Our experimental results demonstrate the effectiveness of our gaze prediction method by highlighting its capability to replicate human gaze behavior and its applicability for downstream tasks where real human-gaze is used as input.
Süleyman Özdel, Yao Rong 0001, Mert Albaba, Yen-Ling Kuo, Xi Wang 0021, Enkelejda Kasneci
ETRA4
2024 Gaze-Guided Graph Neural Network for Action Anticipation Conditioned on Intention
abstract
Humans utilize their gaze to concentrate on essential information while perceiving and interpreting intentions in videos. Incorporating human gaze into computational algorithms can significantly enhance model performance in video understanding tasks. In this work, we address a challenging and innovative task in video understanding: predicting the actions of an agent in a video based on a partial video. We introduce the Gaze-guided Action Anticipation algorithm, which establishes a visual-semantic graph from the video input. Our method utilizes a Graph Neural Network to recognize the agent’s intention and predict the action sequence to fulfill this intention. To assess the efficiency of our approach, we collect a dataset containing household activities generated in the VirtualHome environment, accompanied by human gaze data of viewing videos. Our method outperforms state-of-the-art techniques, achieving a 7% improvement in accuracy for 18-class intention recognition. This highlights the efficiency of our method in learning important features from human gaze data.
Süleyman Özdel, Yao Rong 0001, Mert Albaba, Yen-Ling Kuo, Xi Wang 0021, Enkelejda Kasneci
ETRA4
2023 Zero-Shot Linear Combinations of Grounded Social Interactions with Linear Social MDPs
abstract
Humans and animals engage in rich social interactions. It is often theorized that a relatively small number of basic social interactions give rise to the full range of behavior observed. But no computational theory explaining how social interactions combine together has been proposed before. We do so here. We take a model, the Social MDP, which is able to express a range of social interactions, and extend it to represent linear combinations of social interactions. Practically for robotics applications, such models are now able to not just express that an agent should help another agent, but to express goal-centric social interactions. Perhaps an agent is helping someone get dressed, but preventing them from falling, and is happy to exchange stories in the meantime. How an agent responds socially, should depend on what it thinks the other agent is doing at that point in time. To encode this notion, we take linear combinations of social interactions as defined in Social MDPs, and compute the weights on those combinations on the fly depending on the estimated goals of other agents. This new model, the Linear Social MDP, enables zero-shot reasoning about complex social interactions, provides a mathematical basis for the long-standing intuition that social interactions should compose, and leads to interesting new behaviors that we validate using human observers. Complex social interactions are part of the future of intelligent agents, and having principled mathematical models built on a foundation like MDPs will make it possible to bring social interactions to every robotic application.
Ravi Tejwani, Yen-Ling Kuo, Tianmin Shu, Bennett Stankovits, Dan Gutfreund, Josh Tenenbaum, Boris Katz, Andrei Barbu
AAAI2
2022 Reconstructing Action-Conditioned Human-Object Interactions Using Commonsense Knowledge Priors
abstract
We present a method for inferring diverse 3D models of human-object interactions from images. Reasoning about how humans interact with objects in complex scenes from a single 2D image is a challenging task given ambiguities arising from the loss of information through projection. In addition, modeling 3D interactions requires the generalization ability towards diverse object categories and interaction types. We propose an action-conditioned modeling of interactions that allows us to infer diverse 3D arrangements of humans and objects without supervision on contact regions or 3D scene geometry. Our method extracts high-level commonsense knowledge from large language models (such as GPT-3), and applies them to perform 3D reasoning of human-object interactions. Our key insight is priors extracted from large language models can help in reasoning about human-object contacts from textural prompts only. We quantitatively evaluate the inferred 3D models on a large human-object interaction dataset and show how our method leads to better 3D reconstructions. We further qualitatively evaluate the effectiveness of our method on real images and demonstrate its generalizability towards interaction types and object categories.
Xi Wang 0021, Gen Li 0010, Yen-Ling Kuo, Muhammed Kocabas, Emre Aksan, Otmar Hilliges
3DV3
2022 Motion-centric Tools to Reflect on Digital Creative Experiences and Created Outputs
abstract
Craft and motion are strongly connected, and reflection about this connection enhances our understanding of digital creative experiences. We summarize seven key steps for reflection in craft and create machine-assisted motion-centric tools to support reflection for each key step. These tools reveal different properties of the recorded movement sequence and the relationship between movement and created models. We tested the tools with a sample of novice users who were asked to create a wire-based jewelry model with a set of creation intentions. The resultant qualitative data and visual analysis charts show that introducing the motion information helps the user to meaningfully reflect on their creative experiences and ultimately enhance their decision-making in terms of selecting the models and creating subsequent digital craft-making movements. We therefore hope that our study will stimulate future discussion to bring motion-centric reflection to digital craft processes.
Yen-Ting Cho, Yen-Ling Kuo, Yen-Ting Yeh 0001, Huai-Hsuan Liang, Yu-Ting Li
Creativity & Cognition2
2022 Quantifying the Emergence of Symbolic Communication
Emily Cheng, Yen-Ling Kuo, Josefina Correa, Boris Katz, Ignacio Cases, Andrei Barbu
CogSci2
2022 Trajectory Prediction with Linguistic Representations
abstract
Language allows humans to build mental models that interpret what is happening around them resulting in more accurate long-term predictions. We present a novel trajectory prediction model that uses linguistic intermediate representations to forecast trajectories, and is trained using trajectory samples with partially-annotated captions. The model learns the meaning of each of the words without direct per-word supervision. At inference time, it generates a linguistic description of trajectories which captures maneuvers and interactions over an extended time interval. This generated description is used to refine predictions of the trajectories of multiple agents. We train and validate our model on the Argoverse dataset, and demonstrate improved accuracy results in trajectory prediction. In addition, our model is more interpretable: it presents part of its reasoning in plain language as captions, which can aid model development and can aid in building confidence in the model before deploying it.
Yen-Ling Kuo, Xin Huang 0018, Andrei Barbu, Stephen G. McGill, Boris Katz, John J. Leonard, Guy Rosman
ICRA1
2022 Incorporating Rich Social Interactions Into MDPs
abstract
Much of what we do as humans is engage socially with other agents, a skill that robots must also eventually possess. We demonstrate that a rich theory of social interactions originating from microsociology can be formalized by extending a nested MDP where agents reason about arbitrary functions of each other's rewards. This extended Social MDP allows us to encode the five basic interactions that underlie microsociology: cooperation, conflict, coercion, competition, and exchange. The result is a robotic agent capable of executing social interactions in new environments with no interaction-specific training; like humans it can engage socially in novel ways even without a single example of that social interaction. Moreover, the estimations of these Social MDPs align closely with the judge-ments of humans when considering which social interaction is taking place in an environment. This method both sheds light on the nature of social interactions, by providing concrete mathematical definitions, and brings rich social interactions into a mathematical framework that has proven to be natural for robotics.
Ravi Tejwani, Yen-Ling Kuo, Tianmin Shu, Bennett Stankovits, Dan Gutfreund, Josh Tenenbaum, Boris Katz, Andrei Barbu
ICRA2
2021 IntuModels: Enabling Interactive Modeling for the Novice through Idea Generation and Selection
abstract
We present IntuModels, a machine-assisted interactive modeling workflow to enable the novice to create 3D models. The workflow uses a phase-driven approach, including idea generation and selection, to stimulate creativity and assist users who are not familiar with generating design ideas and 3D modeling techniques. By transforming parametric models using continual input data controlled by users, IntuModels motivates users to intuitively generate a huge amount of 3D model options with a good diversity. For selecting from the created models, we design a balanced overview showing the models using clustering and staged tools to help the user view and understand the models correctly. We tested IntuModels with a sample of novices who were asked to create a wire-based jewelry model. Presented in thematic networks and quantitative charts, the results showed that the novice considered IntuModels to be intuitive to use and useful for creating models that exceeded their expectations for post-production.
Yen-Ting Cho, Yen-Ling Kuo, Yen-Ting Yeh 0001, Yen-Yi Huang, Po-Lun Huang
Creativity & Cognition2
2020 Deep compositional robotic planners that follow natural language commands
abstract
We demonstrate how a sampling-based robotic planner can be augmented to learn to understand a sequence of natural language commands in a continuous configuration space to move and manipulate objects. Our approach combines a deep network structured according to the parse of a complex command that includes objects, verbs, spatial relations, and attributes, with a sampling-based planner, RRT. A recurrent hierarchical deep network controls how the planner explores the environment, determines when a planned path is likely to achieve a goal, and estimates the confidence of each move to trade off exploitation and exploration between the network and the planner. Planners are designed to have near-optimal behavior when information about the task is missing, while networks learn to exploit observations which are available from the environment, making the two naturally complementary. Combining the two enables generalization to new maps, new kinds of obstacles, and more complex sentences that do not occur in the training set. Little data is required to train the model despite it jointly acquiring a CNN that extracts features from the environment as it learns the meanings of words. The model provides a level of interpretability through the use of attention maps allowing users to see its reasoning steps despite being an end-to-end model. This end-to-end model allows robots to learn to follow natural language commands in challenging continuous environments.
Yen-Ling Kuo, Boris Katz, Andrei Barbu
ICRA1
2020 Encoding formulas as deep networks: Reinforcement learning for zero-shot execution of LTL formulas
abstract
We demonstrate a reinforcement learning agent which uses a compositional recurrent neural network that takes as input an LTL formula and determines satisfying actions. The input LTL formulas have never been seen before, yet the network performs zero-shot generalization to satisfy them. This is a novel form of multi-task learning for RL agents where agents learn from one diverse set of tasks and generalize to a new set of diverse tasks. The formulation of the network enables this capacity to generalize. We demonstrate this ability in two domains. In a symbolic domain, the agent finds a sequence of letters that is accepted. In a Minecraft-like environment, the agent finds a sequence of actions that conform to the formula. While prior work could learn to execute one formula reliably given examples of that formula, we demonstrate how to encode all formulas reliably. This could form the basis of new multitask agents that discover sub-tasks and execute them without any additional training, as well as the agents which follow more complex linguistic commands. The structures required for this generalization are specific to LTL formulas, which opens up an interesting theoretical question: what structures are required in neural networks for zero-shot generalization to different logics?
Yen-Ling Kuo, Boris Katz, Andrei Barbu
IROS1
2019 MovIPrint: Move, Explore and Fabricate
abstract
MovIPrint is a user-friendly, interactive installation that uses software and a depth-sensing camera to capture human body movement. After inputting digital data such as images or video into the software, MovIPrint offers people innovative and user-friendly ways to explore that data by manipulating it with their body movement. We use media content and/or wireframe design to enable people to then fabricate their own moving images and 3D digital models.
Yen-Ting Cho, Yen-Ling Kuo, Yen-Ting Yeh 0001, Yi-Chin Lee
ACM Multimedia2
2018 Deep Sequential Models for Sampling-Based Planning
abstract
We demonstrate how a sequence model and a sampling-based planner can influence each other to produce efficient plans and how such a model can automatically learn to take advantage of observations of the environment. Sampling-based planners such as RRT generally know nothing of their environments even if they have traversed similar spaces many times. A sequence model, such as an HMM or LSTM, guides the search for good paths. The resulting model, called DeRRT*, observes the state of the planner and the local environment to bias the next move and next planner state. The neural-network-based models avoid manual feature engineering by co-training a convolutional network which processes map features and observations from sensors. We incorporate this sequence model in a manner that combines its likelihood with the existing bias for searching large unexplored Voronoi regions. This leads to more efficient trajectories with fewer rejected samples even in difficult domains such as when escaping bug traps. This model can also be used for dimensionality reduction in multi-agent environments with dynamic obstacles. Instead of planning in a high-dimensional space that includes the configurations of the other agents, we plan in a low-dimensional subspace relying on the sequence model to bias samples using the observed behavior of the other agents. The techniques presented here are general, include both graphical models and deep learning approaches, and can be adapted to a range of planners.
Yen-Ling Kuo, Andrei Barbu, Boris Katz
IROS1
2012 Planning for Reasoning with Multiple Common Sense Knowledge Bases
abstract
Intelligent user interfaces require common sense knowledge to bridge the gap between the functionality of applications and the user’s goals. While current reasoning methods have been used to provide contextual information for interface agents, the quality of their reasoning results is limited by the coverage of their underlying knowledge bases. This article presents reasoning composition , a planning-based approach to integrating reasoning methods from multiple common sense knowledge bases to answer queries. The reasoning results of one reasoning method are passed to other reasoning methods to form a reasoning chain to the target context of a query. By leveraging different weak reasoning methods, we are able to find answers to queries that cannot be directly answered by querying a single common sense knowledge base. By conducting experiments on ConceptNet and WordNet, we compare the reasoning results of reasoning composition, directly querying merged knowledge bases, and spreading activation. The results show an 11.03% improvement in coverage over directly querying merged knowledge bases and a 49.7% improvement in accuracy over spreading activation. Two case studies are presented, showing how reasoning composition can improve performance of retrieval in a video editing system and a dialogue assistant.
Yen-Ling Kuo, Yung-Jen Hsu 0001
ACM Trans. Interact. Intell. Syst.1
2011 Resource-Bounded Crowd-Sourcing of Commonsense Knowledge
abstract
Knowledge acquisition is the essential process of extracting and encoding knowledge, both domain specific and commonsense, to be used in intelligent systems. While many large knowledge bases have been constructed, none is close to complete. This paper presents an approach to improving a knowledge base efficiently under resource constraints. Using a guiding knowledge base, questions are generated from a weak form of similarity-based inference given the glossary mapping between two knowledge bases. The candidate questions are prioritized in terms of the concept coverage of the target knowledge. Experiments were conducted to find questions to grow the Chinese ConceptNet using the English ConceptNet as a guide. The results were evaluated by online users to verify that 94.17% of the questions and 85.77% of the answers are good. In addition, the answers collected in a six-week period showed consistent improvement to a 36.33% increase in concept coverage of the Chinese commonsense knowledge base against the English ConceptNet.
Yen-Ling Kuo, Yung-Jen Hsu 0001
IJCAI1
2011 ACTraversal: Ranking Crowdsourced Commonsense Assertions and Certifications
Tao-Hsuan Chang, Yen-Ling Kuo, Yung-Jen Hsu 0001
PRIMA2
2011 Capability Modeling of Knowledge-Based Agents for Commonsense Knowledge Integration
Yen-Ling Kuo, Yung-Jen Hsu 0001
PRIMA1