Junyang Liang

dblp:395/3035 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 70% Motion planning and robot control · 30%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
reward design
1.722025
Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill Learning · ICRA 2025
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution · AAAI 2025
Robotics › Motion planning and robot control › robot learning
robot skill learning
1.722025
Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill Learning · ICRA 2025
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution · AAAI 2025
Machine learning › Reinforcement learning › reward learning
LLM-based reward generation
0.912025
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution · AAAI 2025
Machine learning › Reinforcement learning › reward design
reward shaping
0.912025
Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill Learning · ICRA 2025
Machine learning › Reinforcement learning
policy optimization
0.312025
Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill Learning · ICRA 2025
Machine learning › Reinforcement learning
value function
0.312025
Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill Learning · ICRA 2025

Methods — techniques the papers use, named apart from their topics

large language model · 1.7policy hybridization · 0.9multi-branch value network · 0.9bayesian optimization · 0.9
YearPublicationVenuePosition
2025 Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution
abstract
The ability to autonomously explore and resolve tasks with minimal human guidance is crucial for the self-development of embodied intelligence. Although reinforcement learning methods can largely ease human effort, it's challenging to design reward functions for real-world tasks, especially for high-dimensional robotic control, due to complex relationships among joints and tasks. Recent advancements large language models (LLMs) enable automatic reward function design. However, approaches evaluate reward functions by re-training policies from scratch placing an undue burden on the reward function, expecting it to be effective throughout the whole policy improvement process. We argue for a more practical strategy in robotic autonomy, focusing on refining existing policies with policy-dependent reward functions rather than a universal one. To this end, we propose a novel reward-policy co-evolution framework where the reward function and the learned policy benefit from each other's progressive on-the-fly improvements, resulting in more efficient and higher-performing skill acquisition. Specifically, the reward evolution process translates the robot's previous best reward function, descriptions of tasks and environment into text inputs. These inputs are used to query LLMs to generate a dynamic amount of reward function candidates, ensuring continuous improvement at each round of evolution. For policy evolution, our method generates new policy populations by hybridizing historically optimal and random policies. Through an improved Bayesian optimization, our approach efficiently and robustly identifies the most capable and plastic reward-policy combination, which then proceeds to the next round of co-evolution. Despite using less data, our approach demonstrates an average normalized improvement of 95.3\% across various high-dimensional robotic skill learning tasks.
Changxin Huang, Yanbin Chang, Junfan Lin, Junyang Liang, Runhao Zeng, Jianqiang Li 0001
AAAI4
2025 Automated Hybrid Reward Scheduling Via Large Language Models for Robotic Skill Learning
abstract
Enabling a high-degree-of-freedom robot to learn specific skills is a challenging task due to the complexity of robotic dynamics. Reinforcement learning (RL) has emerged as a promising solution; however, addressing such problems requires the design of multiple reward functions to account for various constraints in robotic motion. Existing approaches typically sum all reward components indiscriminately to optimize the RL value function and policy. We argue that this uniform inclusion of all reward components in policy optimization is inefficient and limits the robot's learning performance. To address this, we propose an Automated Hybrid Reward Scheduling (AHRS) framework based on Large Language Models (LLMs). This paradigm dynamically adjusts the learning intensity of each reward component throughout the policy optimization process, enabling robots to acquire skills in a gradual and structured manner. Specifically, we design a multi-branch value network, where each branch corresponds to a distinct reward component. During policy optimization, each branch is assigned a weight that reflects its importance, and these weights are automatically computed based on rules designed by LLMs. The LLM generates a rule set in advance, derived from the task description, and during training, it selects a weight calculation rule from the library based on language prompts that evaluate the performance of each branch. Experimental results demonstrate that the AHRS method achieves an average$\mathbf{6. 4 8 \%}$performance improvement across multiple high-degree-of-freedom robotic tasks.
Changxin Huang, Junyang Liang, Yanbin Chang, Jingzhao Xu, Jianqiang Li 0001
ICRA2