Yanjiang Guo

dblp:330/4785 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-5841-6170ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 31% Motion planning and robot control · 25% Robot manipulation · 20%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Motion planning and robot control
robot learning
1.722025
Improving Vision-Language-Action Model with Online Reinforcement Learning · ICRA 2025
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
1.722025
Improving Vision-Language-Action Model with Online Reinforcement Learning · ICRA 2025
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent · ICML 2025
Machine learning › Generative modeling
diffusion model
1.022025
Prediction with Action: Visual Policy Learning via Joint Denoising Process · NeurIPS 2024
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025
Machine learning › Reinforcement learning
embodied control
0.912025
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent · ICML 2025
Robotics › Motion planning and robot control › robot learning › robot policy learning
generalist robot policy
0.912025
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025
Robotics › Motion planning and robot control › robot dynamics
inverse dynamics
0.912025
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025
Computer vision › 3D vision
spatial understanding
0.912025
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent · ICML 2025
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning
0.912025
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025
Robotics › Robot manipulation
diffusion policy
0.812024
Prediction with Action: Visual Policy Learning via Joint Denoising Process · NeurIPS 2024
Machine learning › Reinforcement learning › deep reinforcement learning
visual policy learning
0.812024
Prediction with Action: Visual Policy Learning via Joint Denoising Process · NeurIPS 2024
Machine learning › Reinforcement learning
meta-reinforcement learning
0.712023
Zero-Shot Policy Transfer with Disentangled Task Representation of Meta-Reinforcement Learning · ICRA 2023
Machine learning › Reinforcement learning › generalization in reinforcement learning
policy generalization
0.712023
Zero-Shot Policy Transfer with Disentangled Task Representation of Meta-Reinforcement Learning · ICRA 2023
Machine learning › Reinforcement learning › transfer learning in reinforcement learning › policy transfer
zero-shot policy transfer
0.712023
Zero-Shot Policy Transfer with Disentangled Task Representation of Meta-Reinforcement Learning · ICRA 2023
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining
0.312025
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent · ICML 2025
Machine learning › Reinforcement learning › online decision making
online reinforcement learning
0.312025
Improving Vision-Language-Action Model with Online Reinforcement Learning · ICRA 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.312025
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations · ICML 2025
Machine learning › Reinforcement learning
imitation learning
0.212024
Prediction with Action: Visual Policy Learning via Joint Denoising Process · NeurIPS 2024
Robotics › Robot manipulation › assembly
insertion task
0.212023
Zero-Shot Policy Transfer with Disentangled Task Representation of Meta-Reinforcement Learning · ICRA 2023
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
policy transfer
0.212023
Zero-Shot Policy Transfer with Disentangled Task Representation of Meta-Reinforcement Learning · ICRA 2023

Methods — techniques the papers use, named apart from their topics

vision-language model pretraining · 0.9vision-language model · 0.9video diffusion model · 0.9supervised fine-tuning · 0.9reinforcement learning · 0.9inverse dynamics learning · 0.9future prediction objective · 0.9fine-tuning on robot data · 0.9joint denoising · 0.8diffusion transformer · 0.8
YearPublicationVenuePosition
2025 Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
abstract
Visual representations play a crucial role in developing generalist robotic policies. Previous vision encoders, typically pre-trained with single-image reconstruction or two-image contrastive learning, tend to capture static information, often neglecting the dynamic aspects vital for embodied tasks. Recently, video diffusion models (VDMs) demonstrate the ability to predict future frames and showcase a strong understanding of physical world. We hypothesize that VDMs inherently produce visual representations that encompass both current static information and predicted future dynamics, thereby providing valuable guidance for robot action learning. Based on this hypothesis, we propose the Video Prediction Policy (VPP), which learns implicit inverse dynamics model conditioned on predicted future representations inside VDMs. To predict more precise future, we fine-tune pre-trained video foundation model on robot datasets along with internet human manipulation data. In experiments, VPP achieves a 18.6% relative improvement on the Calvin ABC-D generalization benchmark compared to the previous state-of-the-art, and demonstrates a 31.6% increase in success rates for complex real-world dexterous manipulation tasks. For your convenience, videos can be found at https://video-prediction-policy.github.io/
Yanjiang Guo, Pengchao Wang, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, Jianyu Chen 0002
ICML2
2025 UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
abstract
Recent advancements in Vision-Language-Action (VLA) models have leveraged pre-trained Vision-Language Models (VLMs) to improve the generalization capabilities. VLMs, typically pre-trained on vision-language understanding tasks, provide rich semantic knowledge and reasoning abilities. However, prior research has shown that VLMs often focus on high-level semantic content and neglect low-level features, limiting their ability to capture detailed spatial information and understand physical dynamics. These aspects, which are crucial for embodied control tasks, remain underexplored in existing pre-training paradigms. In this paper, we investigate the training paradigm for VLAs, and introduce UP-VLA, a Unified VLA model training with both multi-modal Understanding and future Prediction objectives, enhancing both high-level semantic comprehension and low-level spatial understanding. Experimental results show that UP-VLA achieves a 33% improvement on the Calvin ABC-D benchmark compared to the previous state-of-the-art method. Additionally, UP-VLA demonstrates improved success rates in real-world manipulation tasks, particularly those requiring precise spatial information.
Jianke Zhang, Yanjiang Guo, Jianyu Chen 0002
ICML2
2025 Improving Vision-Language-Action Model with Online Reinforcement Learning
abstract
Recent studies have successfully integrated large vision-language models (VLMs) into low-level robotic control by supervised fine-tuning (SFT) with expert robotic datasets, resulting in what we term vision-language-action (VLA) models. Although the VLA models are powerful, how to improve these large models during interaction with environments remains an open question. In this paper, we explore how to further improve these VLA models via Reinforcement Learning (RL), a commonly used fine-tuning technique for large models. However, we find that directly applying online RL to large VLA models presents significant challenges, including training instability that severely impacts the performance of large models, and computing burdens that exceed the capabilities of most local machines. To address these challenges, we propose iRe-VLA framework, which iterates between Reinforcement Learning and Supervised Learning to effectively improve VLA models, leveraging the exploratory benefits of RL while maintaining the stability of supervised learning. Experiments in two simulated benchmarks and a real-world manipulation suite validate the effectiveness of our method.
Yanjiang Guo, Jianke Zhang, Yen-Jen Wang, Jianyu Chen 0002
ICRA1
2024 DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
abstract
Large language models (LLMs) encode a vast amount of semantic knowledge and possess remarkable understanding and reasoning capabilities. Previous work has explored how to ground LLMs in robotic tasks to generate feasible and executable textual plans. However, low-level execution in the physical world may deviate from the high-level textual plan due to environmental perturbations or imperfect controller design. In this paper, we propose DoReMi, a novel language model grounding framework that enables immediate Detection and Recovery from Misalignments between plan and execution. Specifically, we leverage LLMs to play a dual role, aiding not only in high-level planning but also generating constraints that can indicate misalignment during execution. Then vision language models (VLMs) are utilized to detect constraint violations continuously. Our pipeline can monitor the low-level execution and enable timely recovery if certain plan-execution misalignment occurs. Experiments on various complex tasks including robot arms and humanoid robots demonstrate that our method can lead to higher task success rates and shorter task completion times.
Yanjiang Guo, Yen-Jen Wang, Lihan Zha, Jianyu Chen 0002
IROS1
2024 Prediction with Action: Visual Policy Learning via Joint Denoising Process
abstract
Diffusion models have demonstrated remarkable capabilities in image generation tasks, including image editing and video creation, representing a good understanding of the physical world. On the other line, diffusion models have also shown promise in robotic control tasks by denoising actions, known as diffusion policy. Although the diffusion generative model and diffusion policy exhibit distinct capabilities—image prediction and robotic action, respectively—they technically follow similar denoising process. In robotic tasks, the ability to predict future images and generate actions is highly correlated since they share the same underlying dynamics of the physical world. Building on this insight, we introduce \textbf{PAD}, a novel visual policy learning framework that unifies image \textbf{P}rediction and robot \textbf{A}ction within a joint \textbf{D}enoising process. Specifically, PAD utilizes Diffusion Transformers (DiT) to seamlessly integrate images and robot states, enabling the simultaneous prediction of future images and robot actions. Additionally, PAD supports co-training on both robotic demonstrations and large-scale video datasets and can be easily extended to other robotic modalities, such as depth images. PAD outperforms previous methods, achieving a significant 38.9\% relative improvement on the full Metaworld benchmark, by utilizing a single text-conditioned visual policy within a data-efficient imitation learning setting. Furthermore, PAD demonstrates superior generalization to unseen tasks in real-world robot manipulation settings with 28.0\% success rate increase compared to the strongest baseline. Videos of PAD can be found at https://sites.google.com/view/pad-paper
Yanjiang Guo, Jianke Zhang, Yen-Jen Wang, Chaochao Lu, Jianyu Chen 0002
NeurIPS1
2023 Zero-Shot Policy Transfer with Disentangled Task Representation of Meta-Reinforcement Learning
abstract
Humans are capable of abstracting various tasks as different combinations of multiple attributes. This perspective of compositionality is vital for human rapid learning and adaption since previous experiences from related tasks can be combined to generalize across novel compositional settings. In this work, we aim to achieve zero-shot policy generalization of Reinforcement Learning (RL) agents by leveraging the task compositionality. Our proposed method is a meta-RL algorithm with disentangled task representation, explicitly encoding different aspects of the tasks. Policy generalization is then performed by inferring unseen compositional task representations via the obtained disentanglement without extra exploration. The evaluation is conducted on three simulated tasks and a challenging real-world robotic insertion task. Experimental results demonstrate that our proposed method achieves policy generalization to unseen compositional tasks in a zero-shot manner.
Zheng Wu 0002, Yichen Xie 0002, Wenzhao Lian, Yanjiang Guo, Jianyu Chen 0002, Stefan Schaal, Masayoshi Tomizuka
ICRA5