Yijie Guo

dblp:166/3547 · DBLP profile ↗
← Back
34ranked-venue papers
8as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 5 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 13 · 3 first-author · 13 since 2021Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Designing Character-Bound In-Public Companions: Interactional Insights from ACG Practices
Yijie Guo, Ruhan Wang, Yuanling Feng, Jini Tao, Zhiling Xu, Yaowen Shen, Qifei Zhou, Zhihao Yao 0004, Zhenhan Huang, Haipeng Mi
DIS1
2026 Designing for Long-Term Emotion Regulation: A Breathing Biofeedback Game for Women in Compulsory Isolation Drug Rehabilitation Centers
Qi Chen 0020, Jiachen Du, Zhihao Yao 0004, Michael Detsiang Li Jr., Yu Cheng 0024, Yijie Guo, Yanzhi Yang, Xijing Chen, Haipeng Mi
CHI7
2026 GestuProp: 3D Virtual Reality Prop Generation with Co-Speech Gestures
abstract
Virtual Reality (VR) has been widely adopted in domains such as gaming, education, and healthcare, where 3D props play a central role in enabling immersive interaction. With the advancement of generative AI, 3D props can now be created rapidly; however, little research has explored how gestures and speech can be integrated to support prop generation. To address this gap, we introduce GestuProp, a VR prop generation system driven by co-speech gestures. Building on a formative study with 30 participants, we proposed a gesture design space and developed the VR system GestuProp. We then conducted a user study with 14 participants, which showed that GestuProp demonstrates good usability and favorable user experiences, while also revealing how object categories influence gesture use and interaction. These findings highlight the potential of gesture–speech synergy to advance prop generation in VR.
Zhihao Yao 0004, Xiwen Yao, Haowei Xiong, Yuanling Feng, Qirui Sun, Yijie Guo, Haipeng Mi
CHI6
2026 From Text to Movement: LLM-driven Swarm User Interfaces for Embodied and Interactive Storytelling
abstract
This paper introduces PuppetLine, an interactive storytelling system that translates natural language narratives into coordinated performances using tabletop robots. The system combines large language models with a constrained set of action and emotion primitives to generate physically executable and interpretable multi-robot motions. We describe the system design and report findings from a designer-centered user study examining how people interpret and reflect on robot enactments. The results provide formative insights into mapping narrative intent to embodied interaction and inform the design of narrative Swarm User Interfaces.
Ruhan Wang, Shuowen Li, Danqi Huang, Yijie Guo, Haipeng Mi
IUI5
2025 A Card-based Co-Design Toolkit for Exploring Smart Material Applications with Multiple Stakeholders: A Case Study on Automotive Interior Design
abstract
Smart materials have garnered significant attention in both academia and industry, yet identifying pragmatically impactful applications still requires contributions from multiple stakeholders, including researchers, designers, and industry professionals.Although previous research has explored novel technical approaches or user-centered applications of smart materials, this study focuses on how to stimulate effective dialogue among stakeholders to explore impactful smart material applications.
Tianyu Yu 0001, Yao Lu 0038, Kejin Yu, Xiwen Yao, Wenjing Deng, Xueqing Li 0005, Yue Yang 0005, Yijie Guo, Guanhong Liu, Haipeng Mi
Conference on Designing Interactive Systems11
2025 Exploring the Design of LLM-based Agent in Enhancing Self-disclosure Among the Older Adults
Yijie Guo, Ruhan Wang, Zhenhan Huang, Tongtong Jin, Xiwen Yao, Yuanling Feng, Haipeng Mi
CHI1
2025 "Would You Please Help Me?" A Study on People's Reaction to a Tactile Paving Detection Robot
abstract
This paper presents AccessiBot, a tactile paving detection robot developed to address urban accessibility challenges. Beyond collecting real-time data for evidence-based governance, this study explores AccessiBot's potential to encourage citizen engagement. Using the Wizard-of-Oz methodology, we conducted an in-the-wild study to simulate real-world interactions and observe how people respond to the robot. Our observation findings show that people do help robots, and their engagement with the robot varied significantly, ranging from passive acknowledgement to active assistance. Besides, post-hoc interviews indicated that participants recognized the robot's social value.
Wenjing Deng, Zhuoyi Cui, Xintong Wu, Yijie Guo, Haipeng Mi
HRI5
2025 Whisk: Inducing Altruistic Behavior to Relieve Child Dental Anxiety
abstract
This study presents Whisk, an interactive robot designed to alleviate Child Dental Anxiety (CDA) and promote health and well-being in line with the United Nations Sustainable Development Goals. Modeled as a hedgehog, Whisk simulates frightened behaviors such as curling up and raising its quills, interacting with children and guiding them to soothe the robot, thereby reducing their own anxiety. The robot's tactile feedback mechanism helps children shift their attention from fears of dental treatment to caring for the vulnerable robot, improving treatment compliance and alleviating anxiety. User experiments indicate that Whisk is effective in reducing CDA, with the robot's weakness eliciting altruistic behaviors from children. Future work will focus on optimizing emotion recognition and feedback mechanisms, as well as exploring its applications in other medical contexts. Whisk offers a novel solution for children's dental treatment, aiming to enhance the treatment experience and promote health and well-being.
Shuzi Yin, Bingjie Gao, Jiahe Lin, Yeonsu Shin, Yijie Guo, Haipeng Mi
HRI6
2025 Chewiebot: A Non-Electric Chewing Practice Robot for Toddlers
abstract
Chewing is an essential developmental activity for toddlers, especially those aged 2 to 3 years. However, some children have underdeveloped chewing abilities and lack adequate practice. To address this, we propose Chewiebot, a standardized chewable training robot to enhance toddlers' chewing skills. Chewiebot's interaction design utilizes sweet taste rewards as feedback, incorporating playing house elements to make the chewing practice more engaging and encourage children to practice. Chewiebot's design challenges traditional electronic robots by replacing conventional circuits with fluid circuits, ensuring a safe chewing experience for children. Paving the way for safe and eco-friendly designs in child-robot interactions.
Hsiangning Li, Yujie Du, Yijie Guo, Haipeng Mi
HRI5
2025 AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
abstract
Robotic manipulation in open-world settings requires not only task execution but also the ability to detect and learn from failures. While recent advances in vision-language models (VLMs) and large language models (LLMs) have improved robots' spatial reasoning and problem-solving abilities, they still struggle with failure recognition, limiting their real-world applicability. We introduce AHA, an open-source VLM designed to detect and reason about failures in robotic manipulation using natural language. By framing failure detection as a free-form reasoning task, AHA identifies failures and provides detailed, adaptable explanations across different robots, tasks, and environments. We fine-tuned AHA using FailGen, a scalable framework that generates the first large-scale dataset of robotic failure trajectories, the AHA dataset. FailGen achieves this by procedurally perturbing successful demonstrations from simulation. Despite being trained solely on the AHA dataset, AHA generalizes effectively to real-world failure datasets, robotic systems, and unseen tasks. It surpasses the second-best model (GPT-4o in-context learning) by 10.3% and exceeds the average performance of six compared models including five state-of-the-art VLMs by 35.3% across multiple metrics and datasets. We integrate AHA into three manipulation frameworks that utilize LLMs/VLMs for reinforcement learning, task and motion planning, and zero-shot trajectory generation. AHA’s failure feedback enhances these policies' performances by refining dense reward functions, optimizing task planning, and improving sub-task verification, boosting task success rates by an average of 21.4% across all three tasks compared to GPT-4 models. Project page: https://aha-vlm.github.io
Jiafei Duan, Wilbert Pumacay, Nishanth Kumar, Yi Ru Wang, Shulin Tian, Ranjay Krishna, Dieter Fox, Ajay Mandlekar, Yijie Guo
ICLR10
2025 SRSA: Skill Retrieval and Adaptation for Robotic Assembly Tasks
abstract
Enabling robots to learn novel tasks in a data-efficient manner is a long-standing challenge. Common strategies involve carefully leveraging prior experiences, especially transition data collected on related tasks. Although much progress has been made for general pick-and-place manipulation, far fewer studies have investigated contact-rich assembly tasks, where precise control is essential. We introduce SRSA} (Skill Retrieval and Skill Adaptation), a novel framework designed to address this problem by utilizing a pre-existing skill library containing policies for diverse assembly tasks. The challenge lies in identifying which skill from the library is most relevant for fine-tuning on a new task. Our key hypothesis is that skills showing higher zero-shot success rates on a new task are better suited for rapid and effective fine-tuning on that task. To this end, we propose to predict the transfer success for all skills in the skill library on a novel task, and then use this prediction to guide the skill retrieval process. We establish a framework that jointly captures features of object geometry, physical dynamics, and expert actions to represent the tasks, allowing us to efficiently learn the transfer success predictor. Extensive experiments demonstrate that SRSA significantly outperforms the leading baseline. When retrieving and fine-tuning skills on unseen tasks, SRSA achieves a 19% relative improvement in success rate, exhibits 2.6x lower standard deviation across random seeds, and requires 2.4x fewer transition samples to reach a satisfactory success rate, compared to the baseline. In a continual learning setup, SRSA efficiently learns policies for new tasks and incorporates them into the skill library, enhancing future policy learning. Furthermore, policies trained with SRSA in simulation achieve a 90% mean success rate when deployed in the real world. Please visit our project webpage https://srsa2024.github.io/.
Yijie Guo, Bingjie Tang, Iretiayo Akinola, Dieter Fox, Abhishek Gupta 0004, Yashraj Narang
ICLR1
2025 ES-Parkour: Advanced Robot Parkour with Bio-inspired Event Camera and Spiking Neural Network
abstract
In recent years, quadruped robotics has advanced significantly, particularly in perception and motion control via reinforcement learning, enabling complex motions in challenging environments. Visual sensors like depth cameras enhance stability and robustness but face limitations, such as low operating frequencies relative to joint control and sensitivity to lighting, which hinder outdoor deployment. Additionally, deep neural networks in sensor and control systems increase computational demands. To address these issues, we introduce spiking neural networks (SNNs) and event cameras to perform a challenging quadruped parkour task. Event cameras capture dynamic visual data, while SNNs efficiently process spike sequences, mimicking biological perception. Experimental results demonstrate that this approach significantly outperforms traditional models, achieving excellent parkour performance with just 11.7% of the energy consumption of an artificial neural network (ANN)-based model, yielding an 88.3% energy reduction. By integrating event cameras with SNNs, our work advances robotic reinforcement learning and opens new possibilities for applications in demanding environments.
Qiang Zhang 0029, Jiahang Cao, Jingkai Sun, Yecheng Shao, Yijie Guo, Renjing Xu
ICME7
2025 Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models
abstract
Diffusion models have been widely employed in the field of 3D manipulation due to their efficient capability to learn distributions, allowing for precise prediction of action trajectories. However, diffusion models typically rely on large parameter UNet backbones as policy networks, which can be challenging to deploy on resource-constrained devices. Recently, the Mamba model has emerged as a promising solution for efficient modeling, offering low computational complexity and strong performance in sequence modeling. In this work, we propose the Mamba Policy, a lighter but stronger policy that reduces the parameter count by over 80% compared to the original policy network while achieving superior performance. Specifically, we introduce the XMamba Block, which effectively integrates input information with conditional features and leverages a combination of Mamba and Attention mechanisms for deep feature extraction. Extensive experiments demonstrate that the Mamba Policy excels on the Adroit, Dexart, and MetaWorld datasets, requiring significantly fewer computational resources. Additionally, we highlight the Mamba Policy’s enhanced robustness in long-horizon scenarios compared to baseline methods and explore the performance of various Mamba variants within the Mamba Policy framework. Real-world experiments are also conducted to further validate its effectiveness. Our open-source project page can be found at https://andycao1125.github.io/mamba_policy/.
Jiahang Cao, Qiang Zhang 0029, Jingkai Sun, Hao Cheng 0015, Yulin Li 0001, Jun Ma 0008, Kun Wu 0001, Yecheng Shao, Yijie Guo, Renjing Xu
IROS13
2025 Distillation-PPO: A Novel Two-Stage Reinforcement Learning Framework for Humanoid Robot Perceptive Locomotion
abstract
In recent years, humanoid robots have garnered significant attention from both academia and industry due to their high adaptability to environments and human-like characteristics. With the rapid advancement of reinforcement learning, substantial progress has been made in the walking control of humanoid robots. However, existing methods still face challenges when dealing with complex environments and irregular terrains. In the field of perceptive locomotion, existing approaches are generally divided into two-stage methods and end-to-end methods. Two-stage methods first train a teacher policy in a simulated environment and then use distillation techniques, such as DAgger, to transfer the privileged information learned as latent features or actions to the student policy. End-to-end methods, on the other hand, forgo the learning of privileged information and directly learn policies from a partially observable Markov decision process (POMDP) through reinforcement learning. However, due to the lack of supervision from a teacher policy, end-to-end methods often face difficulties in training and exhibit unstable performance in real-world applications. This paper proposes an innovative two-stage perceptive locomotion framework that combines the advantages of teacher policies learned in a fully observable Markov decision process (MDP) to regularize and supervise the student policy. At the same time, it leverages the characteristics of reinforcement learning to ensure that the student policy can continue to learn in a POMDP, thereby enhancing the model’s upper bound. Our experimental results demonstrate that our two-stage training framework achieves higher training efficiency and stability in simulated environments, while also exhibiting better robustness and generalization capabilities in real-world applications.
Qiang Zhang 0029, Jingkai Sun, Jiahang Cao, Yijie Guo, Renjing Xu
IROS8
2025 AI-Gadget Kit: Integrating Swarm User Interfaces with LLM-driven Agents for Tabletop Game Applications
abstract
While Swarm User Interfaces (SUIs) have succeeded in enriching tangible interaction experiences, their limitations in autonomous action planning have hindered the potential for personalized and dynamic interaction generation in tabletop games. Based on the AI-Gadget Kit we developed, this paper explores how to integrate LLM-driven agents to enable SUIs to execute interaction tasks within tabletop games. After defining the design space of this kit, we elucidate the method for designing agents that can extend the meta-actions of SUIs to motion planning. Furthermore, we introduce an add-on prompt method that simplifies the design process for four interaction relationships in tabletop games. Lastly, we present an example that illustrates the potential of AI-Gadget Kit to construct personalized complex interactions in SUI tabletop games.
Yijie Guo, Ruhan Wang, Zhenhan Huang, Zhihao Yao 0004, Tianyu Yu 0001, Zhiling Xu, Xueqing Li 0005, Haipeng Mi
RO-MAN1
2025 An efficient quantized GEMV implementation for large language models inference with matrix core
Lu Lu 0011, Yijie Guo, Zhanyu Yang
J. Supercomput.4
2024 FinePOSE: Fine-Grained Prompt-Driven 3D Human Pose Estimation via Diffusion Models
abstract
The 3D Human Pose Estimation (3D HPE) task uses 2D images or videos to predict human joint coordinates in 3D space. Despite recent advancements in deep learning-based methods, they mostly ignore the capability of coupling accessible texts and naturally feasible knowledge of humans, missing out on valuable implicit supervision to guide the 3D HPE task. Moreover, previous efforts often study this task from the perspective of the whole human body, neglecting fine-grained guidance hidden in different body parts. To this end, we present a new Fine-Grained Prompt-Driven Denoiser based on a diffusion model for 3D HPE, named FinePOSE. It consists of three core blocks enhancing the reverse process of the diffusion model: (1) Fine-grained Part-aware Prompt learning (FPP) block constructs fine-grained part-aware prompts via coupling accessible texts and naturally feasible knowledge of body parts with learnable prompts to model implicit guidance. (2) Fine-grained Prompt-pose Communication (FPC) block establishes fine-grained communications between learned part-aware prompts and poses to improve the denoising quality. (3) Prompt-driven Timestamp Stylization (PTS) block integrates learned prompt embedding and temporal information related to the noise level to enable adaptive adjustment at each denoising step. Extensive experiments on public single-human pose estimation datasets show that FinePOSE outperforms state-of-the-art methods. We further extend FinePOSE to multi-human pose estimation. Achieving 34.3mm average MPJPE on the EgoHumans dataset demonstrates the potential of FinePOSE to deal with complex multi-human scenarios. Code is available at https://github.com/PKU-ICST-MIPL/FinePOSE_CVPR2024.
Jinglin Xu, Yijie Guo, Yuxin Peng 0001
CVPR2
2024 Geometric Fabrics: a Safe Guiding Medium for Policy Learning
abstract
Robotics policies are always subjected to complex, second order dynamics that entangle their actions with resulting states. In reinforcement learning (RL) contexts, policies have the burden of deciphering these complicated interactions over massive amounts of experience and complex reward functions to learn how to accomplish tasks. Moreover, policies typically issue actions directly to controllers like Operational Space Control (OSC) or joint PD control, which induces straightline motion towards these action targets in task or joint space. However, straightline motion in these spaces for the most part do not capture the rich, nonlinear behavior our robots need to exhibit, shifting the burden of discovering these behaviors more completely to the agent. Unlike these simpler controllers, geometric fabrics capture a much richer and desirable set of behaviors via artificial, second order dynamics grounded in nonlinear geometry. These artificial dynamics shift the uncontrolled dynamics of a robot via an appropriate control law to form behavioral dynamics. Behavioral dynamics unlock a new action space and safe, guiding behavior over which RL policies are trained. Behavioral dynamics enable bang-bang-like RL policy actions that are still safe for real robots, simplify reward engineering, and help sequence real-world, high-performance policies. We describe the framework more generally and create a specific instantiation for the problem of dexterous, in-hand reorientation of a cube by a highly actuated robot hand.
Karl Van Wyk, Ankur Handa, Viktor Makoviychuk, Yijie Guo, Arthur Allshire, Nathan D. Ratliff
ICRA4
2024 Reinforcement Learning with Generalizable Gaussian Splatting
abstract
An excellent representation is crucial for reinforcement learning (RL) performance, especially in vision-based reinforcement learning tasks. The quality of the environment representation directly influences the achievement of the learning task. Previous vision-based RL typically uses explicit or implicit ways to represent environments, such as images, points, voxels, and neural radiance fields. However, these representations contain several drawbacks. They cannot either describe complex local geometries or generalize well to unseen scenes, or require precise foreground masks. Moreover, these implicit neural representations are akin to a "black box", significantly hindering interpretability. 3D Gaussian Splatting (3DGS), with its explicit scene representation and differentiable rendering nature, is considered a revolutionary change for reconstruction and representation methods. In this paper, we propose a novel Generalizable Gaussian Splatting framework to be the representation of RL tasks, called GSRL. Through validation in the RoboMimic environment, our method achieves better results than other baselines in multiple tasks, improving the performance by 10%, 44%, and 15% compared with baselines on the hardest task. This work is the first attempt to leverage generalizable 3DGS as a representation for RL.
Qiang Zhang 0029, Jingkai Sun, Jiahang Cao, Weining Zhang, Yecheng Shao, Yijie Guo, Renjing Xu
IROS9
2024 Whole-body Humanoid Robot Locomotion with Human Reference
abstract
Recently, humanoid robots have made significant advances in their ability to perform challenging tasks due to the deployment of Reinforcement Learning (RL), however, the inherent complexity of humanoid robots, including the difficulty of designing complicated reward functions and training entire sophisticated systems, still poses a notable challenge. To conquer these challenges, after many iterations and in-depth investigations, we have meticulously developed a full-size humanoid robot, "Adam", whose innovative structural design greatly improves the efficiency and effectiveness of the imitation learning process. In addition, we have developed a novel imitation learning framework based on an adversarial motion prior, which applies not only to Adam but also to humanoid robots in general. Using the framework, Adam can exhibit unprecedented human-like characteristics in locomotion tasks. Our experimental results demonstrate that the proposed framework enables Adam to achieve human-comparable performance in complex locomotion tasks, marking the first time that human locomotion data has been used for imitation learning in a full-size humanoid robot. For more video demonstrations, please visit our YouTube channel: https://www.youtube.com/watch?v=7hK2ySYBa1I
Qiang Zhang 0029, Peter Cui, David Yan, Jingkai Sun, Yiqun Duan, Weining Zhang, Yijie Guo, Arthur Zhang, Renjing Xu
IROS9
2023 Fluidic Computation Kit: Towards Electronic-free Shape-changing Interfaces
abstract
Although fluidic computation has been utilized to develop interactive devices in the field of Human-Computer Interaction (HCI), the limited computation complexity of previous work hinders the exploration of richer interaction modalities. Based on the Fluidic Computation Kit we developed, this paper explores how unconventional mechanical computing can be leveraged to design shape-changing interfaces that integrate input sensing, output, and complex computation. After introducing the design space enabled by the Kit, we explain how to design four types of elementary computational components and six categories of operators. We end by providing several application scenarios which illustrate the Fluidic Computation Kit’s potential to build sophisticated circuits (e.g., a parallel processor) for use in the field of HCI.
Qiuyu Lu, Haiqing Xu 0001, Yijie Guo, Joey Yu Wang, Lining Yao
CHI3
2023 A Plug-In Weight-Shifting Module That Adds Emotional Expressiveness to Inanimate Objects in Handheld Interaction
abstract
A plug-in weight-shifting module that can be inserted into a variety of objects is presented. The module is equipped with a movable weight inside its body. Three-dimensional weight shifts are presented by controlling one-dimensional translational and two-dimensional rotational movements. To explore the use case of this weight-shifting module, eight weight shift patterns expressing certain emotions were created through a workshop and a qualitative analysis. User tests, to which three different embodiments and scenarios were applied, examined the following three cases: the weight shift patterns were presented to the user by a) a stuffed toy-style robot that mediated human messaging, b) a cushion that made the user relax, and c) a container that enhanced the user's movie-watching experience. User interviews revealed the feasibility of the module and its weight shift patterns for the user's perception of emotions.
Yohei Noguchi 0001, Yijie Guo, Fumihide Tanaka
ICRA2
2023 Yousu: A mythical character robot design for public scene interaction
abstract
With the advancement of interactive technology in the information age, the problem of “visual blindness” in the field of display design has become increasingly prevalent in public scene interaction design. Therefore, it has become crucial to address how new forms of interaction and interaction scenes can be adopted to attract the public. In the context of robotics’ continuous development, robots are playing an increasingly prominent role as interactive subjects in public scene interaction experiences. In new fields such as digital entertainment, spatial experience, and new media art, various typical scenes of robot interaction have emerged. In this study, a window robot named “Yousu” was developed based on an ancient Chinese mythological character and deployed in the window of a bookstore in Beijing. A user experiment was conducted to investigate how to design a reasonable and effective character robot interaction in public scenes to enhance the interaction scenes’ attractiveness.
Qirui Sun, Yijie Guo, Zhihao Yao 0004, Haipeng Mi
RO-MAN2
2023 Exploring the Design of Robot Mediation with Bodily Contact for Remote Conflict
abstract
Interpersonal conflicts are often more difficult to mediate when communicating remotely. The lack of social cues and external mediation makes it difficult for positive conflict behaviors to occur. To this end, robots have been shown to have the potential as mediators. In this paper, we attempt to discuss how to design appropriate bodily contact interactions for the different roles of a robot mediator so as to facilitate the effectiveness of its mediation. We first conduct a pilot interview to probe the potential roles and design elements of robot contact in this study. Then, we explore the relationship between these roles and design elements through a 16-participant design workshop. Finally, we analyze these findings and propose design suggestions for future robot mediator design.
Ruhan Wang, Chih-Heng Li, Yijie Guo, Fumihide Tanaka, Haipeng Mi
RO-MAN3
2023 Novel accelerated methods for convolution neural network with matrix core
Yijie Guo, Songxiang Zhu
J. Supercomput.1
2022 Learning Action Translator for Meta Reinforcement Learning on Sparse-Reward Tasks
abstract
Meta reinforcement learning (meta-RL) aims to learn a policy solving a set of training tasks simultaneously and quickly adapting to new tasks. It requires massive amounts of data drawn from training tasks to infer the common structure shared among tasks. Without heavy reward engineering, the sparse rewards in long-horizon tasks exacerbate the problem of sample efficiency in meta-RL. Another challenge in meta-RL is the discrepancy of difficulty level among tasks, which might cause one easy task dominating learning of the shared policy and thus preclude policy adaptation to new tasks. This work introduces a novel objective function to learn an action translator among training tasks. We theoretically verify that the value of the transferred policy with the action translator can be close to the value of the source policy and our objective function (approximately) upper bounds the value difference. We propose to combine the action translator with context-based meta-RL algorithms for better data collection and moreefficient exploration during meta-training. Our approach em-pirically improves the sample efficiency and performance ofmeta-RL algorithms on sparse-reward tasks.
Yijie Guo, Qiucheng Wu, Honglak Lee
AAAI1
2021 Batch Reinforcement Learning Through Continuation Method
Yijie Guo, Shengyu Feng, Nicolas Le Roux, Ed H. Chi, Honglak Lee, Minmin Chen
ICLR1
2020 Memory Based Trajectory-conditioned Policies for Learning from Sparse Rewards
abstract
Reinforcement learning with sparse rewards is challenging because an agent can rarely obtain non-zero rewards and hence, gradient-based optimization of parameterized policies can be incremental and slow. Recent work demonstrated that using a memory buffer of previous successful trajectories can result in more effective policies. However, existing methods may overly exploit past successful experiences, which can encourage the agent to adopt sub-optimal and myopic behaviors. In this work, instead of focusing on good experiences with limited diversity, we propose to learn a trajectory-conditioned policy to follow and expand diverse past trajectories from a memory buffer. Our method allows the agent to reach diverse regions in the state space and improve upon the past trajectories to reach new states. We empirically show that our approach significantly outperforms count-based exploration methods (parametric approach) and self-imitation learning (parametric approach with non-parametric memory) on various complex tasks with local optima. In particular, without using expert demonstrations or resetting to arbitrary states, we achieve the state-of-the-art scores under five billion number of frames, on challenging Atari games such as Montezuma’s Revenge and Pitfall.
Yijie Guo, Marcin Moczulski, Shengyu Feng, Samy Bengio, Mohammad Norouzi 0002, Honglak Lee
NeurIPS1
2020 Predictive Information Accelerates Learning in RL
abstract
The Predictive Information is the mutual information between the past and the future, I(Xpast; Xfuture). We hypothesize that capturing the predictive information is useful in RL, since the ability to model what will happen next is necessary for success on many tasks. To test our hypothesis, we train Soft Actor-Critic (SAC) agents from pixels with an auxiliary task that learns a compressed representation of the predictive information of the RL environment dynamics using a contrastive version of the Conditional Entropy Bottleneck (CEB) objective. We refer to these as Predictive Information SAC (PI-SAC) agents. We show that PI-SAC agents can substantially improve sample efficiency over challenging baselines on tasks from the DM Control suite of continuous control environments. We evaluate PI-SAC agents by comparing against uncompressed PI-SAC agents, other compressed and uncompressed agents, and SAC agents directly trained from pixels. Our implementation is given on GitHub.
Kuang-Huei Lee, Ian Fischer, Anthony Z. Liu, Yijie Guo, Honglak Lee, John F. Canny, Sergio Guadarrama
NeurIPS4
2019 Contingency-Aware Exploration in Reinforcement Learning
Yijie Guo, Marcin Moczulski, Junhyuk Oh, Neal Wu, Mohammad Norouzi 0002, Honglak Lee
ICLR (Poster)2
2018 Unsupervised Discovery of Object Landmarks as Structural Representations
abstract
Deep neural networks can model images with rich latent representations, but they cannot naturally conceptualize structures of object categories in a human-perceptible way. This paper addresses the problem of learning object structures in an image modeling process without supervision. We propose an autoencoding formulation to discover landmarks as explicit structural representations. The encoding module outputs landmark coordinates, whose validity is ensured by constraints that reflect the necessary properties for landmarks. The decoding module takes the landmarks as a part of the learnable input representations in an end-to-end differentiable framework. Our discovered landmarks are semantically meaningful and more predictive of manually annotated landmarks than those discovered by previous methods. The coordinates of our landmarks are also complementary features to pretrained deep-neural-network representations in recognizing visual attributes. In addition, the proposed method naturally creates an unsupervised, perceptible interface to manipulate object shapes and decode images with controllable structures.
Yuting Zhang 0001, Yijie Guo, Yixin Jin, Yijun Luo, Zhiyuan He 0003, Honglak Lee
CVPR2
2018 Self-Imitation Learning
abstract
This paper proposes Self-Imitation Learning (SIL), a simple off-policy actor-critic algorithm that learns to reproduce the agent’s past good decisions. This algorithm is designed to verify our hypothesis that exploiting past good experiences can indirectly drive deep exploration. Our empirical results show that SIL significantly improves advantage actor-critic (A2C) on several hard exploration Atari games and is competitive to the state-of-the-art count-based exploration methods. We also show that SIL improves proximal policy optimization (PPO) on MuJoCo tasks.
Junhyuk Oh, Yijie Guo, Satinder Singh 0001, Honglak Lee
ICML2
2017 Discriminative Bimodal Networks for Visual Localization and Detection with Natural Language Queries
abstract
Associating image regions with text queries has been recently explored as a new way to bridge visual and linguistic representations. A few pioneering approaches have been proposed based on recurrent neural language models trained generatively (e.g., generating captions), but achieving somewhat limited localization accuracy. To better address natural-language-based visual entity localization, we propose a discriminative approach. We formulate a discriminative bimodal neural network (DBNet), which can be trained by a classifier with extensive use of negative samples. Our training objective encourages better localization on single images, incorporates text phrases in a broad range, and properly pairs image regions with text phrases into positive and negative examples. Experiments on the Visual Genome dataset demonstrate the proposed DBNet significantly outperforms previous state-of-the-art methods both for localization on single images and for detection on multiple images. We we also establish an evaluation protocol for natural-language visual detection. Code is available at: http://ytzhang.net/projects/dbnet.
Yuting Zhang 0001, Luyao Yuan, Yijie Guo, Zhiyuan He 0003, I-An Huang, Honglak Lee
CVPR3
2016 Perspective Transformer Nets: Learning Single-View 3D Object Reconstruction without 3D Supervision
abstract
Understanding the 3D world is a fundamental problem in computer vision. However, learning a good representation of 3D objects is still an open problem due to the high dimensionality of the data and many factors of variation involved. In this work, we investigate the task of single-view 3D object reconstruction from a learning agent's perspective. We formulate the learning process as an interaction between 3D and 2D representations and propose an encoder-decoder network with a novel projection loss defined by the projective transformation. More importantly, the projection loss enables the unsupervised learning using 2D observation without explicit 3D supervision. We demonstrate the ability of the model in generating 3D volume from a single 2D image with three sets of experiments: (1) learning from single-class objects; (2) learning from multi-class objects and (3) testing on novel object classes. Results show superior performance and better generalization ability for 3D object reconstruction when the projection loss is involved.
Xinchen Yan, Jimei Yang, Ersin Yumer, Yijie Guo, Honglak Lee
NIPS4