Yihang Yao

dblp:305/7045 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0005-6093-853XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 68% Generative modeling · 10% Trustworthy machine learning · 8%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
safe reinforcement learning
3.552024
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning · NeurIPS 2024
Feasibility Consistent Representation Learning for Safe Reinforcement Learning · ICML 2024
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
offline reinforcement learning
2.232024
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning · NeurIPS 2024
Learning from Sparse Offline Datasets via Conservative Density Estimation · ICLR 2024
Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement learning for language models
0.912025
Behavior Injection: Preparing Language Models for Reinforcement Learning · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model training › post-training
reinforcement learning post-training
0.912025
Behavior Injection: Preparing Language Models for Reinforcement Learning · NeurIPS 2025
Machine learning › Deep learning architectures and training
data augmentation
0.812024
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.812024
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning · NeurIPS 2024
Machine learning › Generative modeling
synthetic data generation
0.812024
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.712023
Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data · ICML 2023
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy learning
0.712023
Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization
0.712023
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning › offline reinforcement learning
decision transformer
0.712023
Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning › safe reinforcement learning
offline safe reinforcement learning
0.712023
Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data · ICML 2023
Robotics › Motion planning and robot control
robot learning
0.312026
Tailored Primitive Initialization is the Secret Key to Reinforcement Learning · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.312025
Behavior Injection: Preparing Language Models for Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning
off-policy reinforcement learning
0.212023
Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data · ICML 2023
Mathematical optimization
multi-objective optimization
0.212023
Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.9variational inference · 1.3primitive initialization · 1.0supervised fine-tuning · 0.9data augmentation · 0.9self-supervised learning · 0.8representation learning · 0.8importance sampling · 0.8density estimation · 0.8conservative constraints · 0.8zero-shot adaptation · 0.7multi-objective optimization · 0.7decision transformer · 0.7
YearPublicationVenuePosition
2026 Tailored Primitive Initialization is the Secret Key to Reinforcement Learning
abstract
Yihang Yao, Guangtao Zeng, Raina Wu, Yang Zhang, Ding Zhao, Zhang-Wei Hong, Chuang Gan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yihang Yao, Guangtao Zeng, Raina Wu, Yang Zhang 0001, Ding Zhao, Zhang-Wei Hong, Chuang Gan 0001
ACL (1)1
2025 QuietPaw: Learning Quadrupedal Locomotion with Versatile Noise Preference Alignment
abstract
When operating at their full capacity, quadrupedal robots can produce loud footstep noise, which can be disruptive in human-centered environments like homes, offices, and hospitals. As a result, balancing locomotion performance with noise constraints is crucial for the successful real-world deployment of quadrupedal robots. However, achieving adaptive noise control is challenging due to (a) the trade-off between agility and noise minimization, (b) the need for generalization across diverse deployment conditions, and (c) the difficulty of effectively adjusting policies based on noise requirements. We propose QuietPaw, a framework incorporating our Conditional Noise-Constrained Policy (CNCP), a constrained learning-based algorithm that enables flexible, noise-aware locomotion by conditioning policy behavior on noise-reduction levels. We leverage value representation decomposition in the critics, disentangling state representations from condition-dependent representations and this allows a single versatile policy to generalize across noise levels without retraining while improving the Pareto trade-off between agility and noise reduction. We validate our approach in simulation and the real world, demonstrating that CNCP can effectively balance locomotion performance and noise constraints, achieving continuously adjustable noise reduction.
Yuyou Zhang, Yihang Yao, Shiqi Liu 0005, Yaru Niu, Changyi Lin, Yuxiang Yang 0007, Wenhao Yu 0003, Tingnan Zhang, Jie Tan 0001, Ding Zhao
IROS2
2025 Behavior Injection: Preparing Language Models for Reinforcement Learning
abstract
Reinforcement learning (RL) has emerged as a powerful post-training technique to incentivize the reasoning ability of large language models (LLMs). However, LLMs can respond very inconsistently to RL finetuning: some show substantial performance gains, while others plateau or even degrade. To understand this divergence, we analyze the per-step influence of the RL objective and identify two key conditions for effective post-training: (1) RL-informative rollout accuracy, and (2) strong data co-influence, which quantifies how much the training data affects performance on other samples. Guided by these insights, we propose behavior injection, a task-agnostic data augmentation scheme applied prior to RL. Behavior injection enriches the supervised finetuning (SFT) data by seeding exploratory and exploitative behaviors, effectively making the model more RL-ready. We evaluate our method across two reasoning benchmarks with multiple base models. The results demonstrate that our theoretically motivated augmentation can significantly increase the performance gain from RL over the pre-RL model.
Zhepeng Cen, Yihang Yao, William Jongwon Han, Zuxin Liu, Ding Zhao
NeurIPS2
2024 Learning from Sparse Offline Datasets via Conservative Density Estimation
abstract
Offline reinforcement learning (RL) offers a promising direction for learning policies from pre-collected datasets without requiring further interactions with the environment. However, existing methods struggle to handle out-of-distribution (OOD) extrapolation errors, especially in sparse reward or scarce data settings. In this paper, we propose a novel training algorithm called Conservative Density Estimation (CDE), which addresses this challenge by explicitly imposing constraints on the state-action occupancy stationary distribution. CDE overcomes the limitations of existing approaches, such as the stationary distribution correction method, by addressing the support mismatch issue in marginal importance sampling. Our method achieves state-of-the-art performance on the D4RL benchmark. Notably, CDE consistently outperforms baselines in challenging tasks with sparse rewards or insufficient data, demonstrating the advantages of our approach in addressing the extrapolation error problem in offline RL.
Zhepeng Cen, Zuxin Liu, Zitong Wang 0005, Yihang Yao, Henry Lam, Ding Zhao
ICLR4
2024 Feasibility Consistent Representation Learning for Safe Reinforcement Learning
abstract
In the field of safe reinforcement learning (RL), finding a balance between satisfying safety constraints and optimizing reward performance presents a significant challenge. A key obstacle in this endeavor is the estimation of safety constraints, which is typically more difficult than estimating a reward metric due to the sparse nature of the constraint signals. To address this issue, we introduce a novel framework named Feasibility Consistent Safe Reinforcement Learning (FCSRL). This framework combines representation learning with feasibility-oriented objectives to identify and extract safety-related information from the raw state for safe RL. Leveraging self-supervised learning techniques and a more learnable safety metric, our approach enhances the policy learning and constraint estimation. Empirical evaluations across a range of vector-state and image-based tasks demonstrate that our method is capable of learning a better safety-aware embedding and achieving superior performance than previous representation learning baselines.
Zhepeng Cen, Yihang Yao, Zuxin Liu, Ding Zhao
ICML2
2024 OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning
abstract
Offline safe reinforcement learning (RL) aims to train a policy that satisfies con- straints using a pre-collected dataset. Most current methods struggle with the mismatch between imperfect demonstrations and the desired safe and rewarding performance. In this paper, we mitigate this issue from a data-centric perspective and introduce OASIS (cOnditionAl diStributIon Shaping), a new paradigm in offline safe RL designed to overcome these critical limitations. OASIS utilizes a conditional diffusion model to synthesize offline datasets, thus shaping the data dis- tribution toward a beneficial target domain. Our approach makes compliance with safety constraints through effective data utilization and regularization techniques to benefit offline safe RL training. Comprehensive evaluations on public benchmarks and varying datasets showcase OASIS’s superiority in benefiting offline safe RL agents to achieve high-reward behavior while satisfying the safety constraints, out- performing established baselines. Furthermore, OASIS exhibits high data efficiency and robustness, making it suitable for real-world applications, particularly in tasks where safety is imperative and high-quality demonstrations are scarce. More details are available at the website https://sites.google.com/view/saferl-oasis/home.
Yihang Yao, Zhepeng Cen, Wenhao Ding, Haohong Lin, Shiqi Liu 0005, Tingnan Zhang, Wenhao Yu 0003, Ding Zhao
NeurIPS1
2023 Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data
abstract
Previous work demonstrates that the optimal safe reinforcement learning policy in a noise-free environment is vulnerable and could be unsafe under observational attacks. While adversarial training effectively improves robustness and safety, collecting samples by attacking the behavior agent online could be expensive or prohibitively dangerous in many applications. We propose the robuSt vAriational ofF-policy lEaRning (SAFER) approach, which only requires benign training data without attacking the agent. SAFER obtains an optimal non-parametric variational policy distribution via convex optimization and then uses it to improve the parameterized policy robustly via supervised learning. The two-stage policy optimization facilitates robust training, and extensive experiments on multiple robot platforms show the efficiency of SAFER in learning a robust and safe policy: achieving the same reward with much fewer constraint violations during training than on-policy baselines.
Zuxin Liu, Zijian Guo 0002, Zhepeng Cen, Huan Zhang 0001, Yihang Yao, Hanjiang Hu, Ding Zhao
ICML5
2023 Constrained Decision Transformer for Offline Safe Reinforcement Learning
abstract
Safe reinforcement learning (RL) trains a constraint satisfaction policy by interacting with the environment. We aim to tackle a more challenging problem: learning a safe policy from an offline dataset. We study the offline safe RL problem from a novel multi-objective optimization perspective and propose the $\epsilon$-reducible concept to characterize problem difficulties. The inherent trade-offs between safety and task performance inspire us to propose the constrained decision transformer (CDT) approach, which can dynamically adjust the trade-offs during deployment. Extensive experiments show the advantages of the proposed method in learning an adaptive, safe, robust, and high-reward policy. CDT outperforms its variants and strong offline safe RL baselines by a large margin with the same hyperparameters across all tasks, while keeping the zero-shot adaptation capability to different constraint thresholds, making our approach more suitable for real-world RL under constraints.
Zuxin Liu, Zijian Guo 0002, Yihang Yao, Zhepeng Cen, Wenhao Yu 0003, Tingnan Zhang, Ding Zhao
ICML3
2023 Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
abstract
Safe reinforcement learning (RL) focuses on training reward-maximizing agents subject to pre-defined safety constraints. Yet, learning versatile safe policies that can adapt to varying safety constraint requirements during deployment without retraining remains a largely unexplored and challenging area. In this work, we formulate the versatile safe RL problem and consider two primary requirements: training efficiency and zero-shot adaptation capability. To address them, we introduce the Conditioned Constrained Policy Optimization (CCPO) framework, consisting of two key modules: (1) Versatile Value Estimation (VVE) for approximating value functions under unseen threshold conditions, and (2) Conditioned Variational Inference (CVI) for encoding arbitrary constraint thresholds during policy optimization. Our extensive experiments demonstrate that CCPO outperforms the baselines in terms of safety and task performance while preserving zero-shot adaptation capabilities to different constraint thresholds data-efficiently. This makes our approach suitable for real-world dynamic applications.
Yihang Yao, Zuxin Liu, Zhepeng Cen, Wenhao Yu 0003, Tingnan Zhang, Ding Zhao
NeurIPS1
2023 An Integrated in Situ Image Acquisition and Annotation Scheme for Instance Segmentation Models in Open Scenes With a Human-Robot Interaction Approach
abstract
A large amount of data acquisition and annotation work is required to train a supervised machine learning model for open scenes. However, traditional manual approaches are inefficient. Here, a method is proposed for on-site image acquisition and semiautomatic annotation based on eye-tracking. This method uses the recognition capabilities and computational advantages of humans and machines to improve annotation efficiency, overcoming the bottleneck of AI-based approaches to the natural scenery understanding of field robots. The proposed method contains three advancements. First, we designed a head-mounted display with a built-in pose measurement module to achieve first-person teleoperation data acquisition, where a pseudoframe interpolation algorithm is designed to overcome the latency problem in immersive remote data transmission and to achieve efficient field data acquisition. Second, we propose an adaptive superpixel segmentation algorithm to reduce human–machine interactions based on eye-tracking. Third, since traditionally, the annotation process cannot provide feedback to the acquisition process and results in a low conversion rate. We proposed a new conversion rate index denoting the rate of transforming collected data into valid data to quantify the acquisition quality in real time. While achieving an annotation quality of 0.964 per the Dice index, which is approximately equal to that of the manual method, the proposed method improves the annotation efficiency by more than 3 times. Finally, the agricultural field experiments containing a real-life scene of robotic tomato-picking verified that the proposed method based on human–computer interaction can make full use of human perception and recognition intelligence.
Yihang Yao, Binhao Chen, Xiaofeng Du, Yidong He, Chengliang Liu 0001
IEEE Trans. Hum. Mach. Syst.3