Xiyue Peng

dblp:347/6307 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0009-2470-8082ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 83% Trustworthy machine learning · 10% Language models and text generation · 7%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
1.012026
BARouter: A Budget-adaptive Online Large Language Model Router Framework · WWW 2026
Cloud and datacenter computing › inference serving
LLM serving
1.012026
BARouter: A Budget-adaptive Online Large Language Model Router Framework · WWW 2026
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization
0.912025
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization · NeurIPS 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization · NeurIPS 2025
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization · NeurIPS 2025
Security and privacy of machine learning › large language model alignment
safety alignment
0.912025
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization · NeurIPS 2025
Machine learning › Reinforcement learning
actor-critic methods
0.812024
Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process
0.812024
Safe and Efficient: A Primal-Dual Method for Offline Convex CMDPs under Partial Data Coverage · NeurIPS 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning › safe reinforcement learning
offline safe reinforcement learning
0.812024
Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning › offline reinforcement learning
pessimism
0.812024
Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning › safe reinforcement learning
safe policy improvement
0.812024
Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning
safe reinforcement learning
0.812024
Safe and Efficient: A Primal-Dual Method for Offline Convex CMDPs under Partial Data Coverage · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model
0.312026
BARouter: A Budget-adaptive Online Large Language Model Router Framework · WWW 2026
Natural language and speech › Language models and text generation
alignment
0.312025
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization · NeurIPS 2025
Mathematical optimization
linear programming
0.212024
Safe and Efficient: A Primal-Dual Method for Offline Convex CMDPs under Partial Data Coverage · NeurIPS 2024
Mathematical optimization
primal-dual method
0.212024
Safe and Efficient: A Primal-Dual Method for Offline Convex CMDPs under Partial Data Coverage · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

regret minimization · 2.0online learning · 2.0rectified policy gradient · 1.7constrained markov decision process · 1.7uncertainty parameters · 1.5primal-dual algorithm · 1.5linear programming · 1.5stackelberg game · 0.8no-regret optimization · 0.8importance weighting · 0.8
YearPublicationVenuePosition
2026 BARouter: A Budget-adaptive Online Large Language Model Router Framework
abstract
With the rapid advancement of large language models (LLMs), a diverse ecosystem of models with different scales and domain specializations has emerged, including LLM-based web agents and online multi-LLM server systems. LLM routing, which opportunistically leverages this diversity to balance response quality and computational cost, has become a central problem in optimizing the performance of LLM serving systems. We propose the Budget-Adaptive Router (BARouter), a budget-adaptive online routing framework for LLM serving systems. BARouter dynamically adjusts its routing policy for incoming queries based on the estimated response quality, query costs, and real-time budget consumption, enabling seamless adaptation to varying initial budgets and evolving input distributions without manual hyperparameter tuning. Theoretically, we prove that BARouter achieves sublinear regret over the time horizon T. The extensive experiments show that BARouter effectively allocates the right budget to the queries at the right time, consistently outperforming baseline algorithms, and remains robust to varying budget levels and shifting query distributions.
Lingkai Zu, Xiyue Peng, Xin Liu 0049
WWW2
2025 Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
abstract
Balancing helpfulness and safety (harmlessness) is a critical challenge in aligning large language models (LLMs). Current approaches often decouple these two objectives, training separate preference models for helpfulness and safety, while framing safety as a constraint within a constrained Markov Decision Process (CMDP) framework. This paper identifies a potential issue when using the widely adopted expected safety constraints for LLM safety alignment, termed "safety compensation'', where the constraints are satisfied on expectation, but individual prompts may trade off safety, resulting in some responses being overly restrictive while others remain unsafe. To address this issue, we propose **Rectified Policy Optimization (RePO)**, which replaces the expected safety constraint with critical safety constraints imposed on every prompt. At the core of RePO is a policy update mechanism driven by rectified policy gradients, which penalizes the strict safety violation of every prompt, thereby enhancing safety across nearly all prompts. Our experiments demonstrate that RePO outperforms strong baseline methods and significantly enhances LLM safety alignment.
Xiyue Peng, Hengquan Guo, Dongqing Zou, Ziyu Shao, Honghao Wei, Xin Liu 0049
NeurIPS1
2024 Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning
abstract
We propose WSAC (Weighted Safe Actor-Critic), a novel algorithm for Safe Offline Reinforcement Learning (RL) under functional approximation, which can robustly optimize policies to improve upon an arbitrary reference policy with limited data coverage. WSAC is designed as a two-player Stackelberg game to optimize a refined objective function. The actor optimizes the policy against two adversarially trained value critics with small importance-weighted Bellman errors, which focus on scenarios where the actor's performance is inferior to the reference policy. In theory, we demonstrate that when the actor employs a no-regret optimization oracle, WSAC achieves a number of guarantees: $(i)$ For the first time in the safe offline RL setting, we establish that WSAC can produce a policy that outperforms {\bf any} reference policy while maintaining the same level of safety, which is critical to designing a safe algorithm for offline RL. $(ii)$ WSAC achieves the optimal statistical convergence rate of $1/\sqrt{N}$ to the reference policy, where $N$ is the size of the offline dataset. $(iii)$ We theoretically show that WSAC guarantees a safe policy improvement across a broad range of hyperparameters that control the degree of pessimism, indicating its practical robustness. Additionally, we offer a practical version of WSAC and compare it with existing state-of-the-art safe offline RL algorithms in several continuous control environments. WSAC outperforms all baselines across a range of tasks, supporting the theoretical results.
Honghao Wei, Xiyue Peng, Arnob Ghosh, Xin Liu 0049
NeurIPS2
2024 Safe and Efficient: A Primal-Dual Method for Offline Convex CMDPs under Partial Data Coverage
abstract
Offline safe reinforcement learning (RL) aims to find an optimal policy using a pre-collected dataset when data collection is impractical or risky. We propose a novel linear programming (LP) based primal-dual algorithm for convex MDPs that incorporates ``uncertainty'' parameters to improve data efficiency while requiring only partial data coverage assumption. Our theoretical results achieve a sample complexity of $\mathcal{O}(1/(1-\gamma)\sqrt{n})$ under general function approximation, improving the current state-of-the-art by a factor of $1/(1-\gamma)$, where $n$ is the number of data samples in an offline dataset, and $\gamma$ is the discount factor. The numerical experiments validate our theoretical findings, demonstrating the practical efficacy of our approach in achieving improved safety and learning efficiency in safe offline settings.
Xiyue Peng, Honghao Wei, Xin Liu 0049
NeurIPS2