EDBT 2026 Demo / reviewers in the wild / expert
Xiyue Peng
dblp:347/6307
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0009-2470-8082ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 83% Trustworthy machine learning · 10% Language models and text generation · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 17 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
cluster resource management and scheduling |
1.0 | 1 | 2026 | BARouter: A Budget-adaptive Online Large Language Model Router Framework · WWW 2026 |
Cloud and datacenter computing › inference serving
LLM serving |
1.0 | 1 | 2026 | BARouter: A Budget-adaptive Online Large Language Model Router Framework · WWW 2026 |
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization |
0.9 | 1 | 2025 | Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization · NeurIPS 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.9 | 1 | 2025 | Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization · NeurIPS 2025 |
Security and privacy of machine learning › large language model alignment
safety alignment |
0.9 | 1 | 2025 | Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization · NeurIPS 2025 |
Machine learning › Reinforcement learning
actor-critic methods |
0.8 | 1 | 2024 | Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process |
0.8 | 1 | 2024 | Safe and Efficient: A Primal-Dual Method for Offline Convex CMDPs under Partial Data Coverage · NeurIPS 2024 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.8 | 1 | 2024 | Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning › safe reinforcement learning
offline safe reinforcement learning |
0.8 | 1 | 2024 | Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning › offline reinforcement learning
pessimism |
0.8 | 1 | 2024 | Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning › safe reinforcement learning
safe policy improvement |
0.8 | 1 | 2024 | Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning · NeurIPS 2024 |
Machine learning › Reinforcement learning
safe reinforcement learning |
0.8 | 1 | 2024 | Safe and Efficient: A Primal-Dual Method for Offline Convex CMDPs under Partial Data Coverage · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2026 | BARouter: A Budget-adaptive Online Large Language Model Router Framework · WWW 2026 |
Natural language and speech › Language models and text generation
alignment |
0.3 | 1 | 2025 | Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization · NeurIPS 2025 |
Mathematical optimization
linear programming |
0.2 | 1 | 2024 | Safe and Efficient: A Primal-Dual Method for Offline Convex CMDPs under Partial Data Coverage · NeurIPS 2024 |
Mathematical optimization
primal-dual method |
0.2 | 1 | 2024 | Safe and Efficient: A Primal-Dual Method for Offline Convex CMDPs under Partial Data Coverage · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
regret minimization · 2.0online learning · 2.0rectified policy gradient · 1.7constrained markov decision process · 1.7uncertainty parameters · 1.5primal-dual algorithm · 1.5linear programming · 1.5stackelberg game · 0.8no-regret optimization · 0.8importance weighting · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BARouter: A Budget-adaptive Online Large Language Model Router FrameworkabstractWith the rapid advancement of large language models (LLMs), a diverse ecosystem of models with different scales and domain specializations has emerged, including LLM-based web agents and online multi-LLM server systems. LLM routing, which opportunistically leverages this diversity to balance response quality and computational cost, has become a central problem in optimizing the performance of LLM serving systems. We propose the Budget-Adaptive Router (BARouter), a budget-adaptive online routing framework for LLM serving systems. BARouter dynamically adjusts its routing policy for incoming queries based on the estimated response quality, query costs, and real-time budget consumption, enabling seamless adaptation to varying initial budgets and evolving input distributions without manual hyperparameter tuning. Theoretically, we prove that BARouter achieves sublinear regret over the time horizon T. The extensive experiments show that BARouter effectively allocates the right budget to the queries at the right time, consistently outperforming baseline algorithms, and remains robust to varying budget levels and shifting query distributions. Lingkai Zu, Xiyue Peng, Xin Liu 0049 |
WWW | 2 |
| 2025 | Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy OptimizationabstractBalancing helpfulness and safety (harmlessness) is a critical challenge in aligning large language models (LLMs). Current approaches often decouple these two objectives, training separate preference models for helpfulness and safety, while framing safety as a constraint within a constrained Markov Decision Process (CMDP) framework. This paper identifies a potential issue when using the widely adopted expected safety constraints for LLM safety alignment, termed "safety compensation'', where the constraints are satisfied on expectation, but individual prompts may trade off safety, resulting in some responses being overly restrictive while others remain unsafe. To address this issue, we propose **Rectified Policy Optimization (RePO)**, which replaces the expected safety constraint with critical safety constraints imposed on every prompt. At the core of RePO is a policy update mechanism driven by rectified policy gradients, which penalizes the strict safety violation of every prompt, thereby enhancing safety across nearly all prompts. Our experiments demonstrate that RePO outperforms strong baseline methods and significantly enhances LLM safety alignment. Xiyue Peng, Hengquan Guo, Dongqing Zou, Ziyu Shao, Honghao Wei, Xin Liu 0049 |
NeurIPS | 1 |
| 2024 | Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement LearningabstractWe propose WSAC (Weighted Safe Actor-Critic), a novel algorithm for Safe Offline Reinforcement Learning (RL) under functional approximation, which can robustly optimize policies to improve upon an arbitrary reference policy with limited data coverage. WSAC is designed as a two-player Stackelberg game to optimize a refined objective function. The actor optimizes the policy against two adversarially trained value critics with small importance-weighted Bellman errors, which focus on scenarios where the actor's performance is inferior to the reference policy. In theory, we demonstrate that when the actor employs a no-regret optimization oracle, WSAC achieves a number of guarantees: $(i)$ For the first time in the safe offline RL setting, we establish that WSAC can produce a policy that outperforms {\bf any} reference policy while maintaining the same level of safety, which is critical to designing a safe algorithm for offline RL. $(ii)$ WSAC achieves the optimal statistical convergence rate of $1/\sqrt{N}$ to the reference policy, where $N$ is the size of the offline dataset. $(iii)$ We theoretically show that WSAC guarantees a safe policy improvement across a broad range of hyperparameters that control the degree of pessimism, indicating its practical robustness. Additionally, we offer a practical version of WSAC and compare it with existing state-of-the-art safe offline RL algorithms in several continuous control environments. WSAC outperforms all baselines across a range of tasks, supporting the theoretical results. Honghao Wei, Xiyue Peng, Arnob Ghosh, Xin Liu 0049 |
NeurIPS | 2 |
| 2024 | Safe and Efficient: A Primal-Dual Method for Offline Convex CMDPs under Partial Data CoverageabstractOffline safe reinforcement learning (RL) aims to find an optimal policy using a pre-collected dataset when data collection is impractical or risky. We propose a novel linear programming (LP) based primal-dual algorithm for convex MDPs that incorporates ``uncertainty'' parameters to improve data efficiency while requiring only partial data coverage assumption. Our theoretical results achieve a sample complexity of $\mathcal{O}(1/(1-\gamma)\sqrt{n})$ under general function approximation, improving the current state-of-the-art by a factor of $1/(1-\gamma)$, where $n$ is the number of data samples in an offline dataset, and $\gamma$ is the discount factor. The numerical experiments validate our theoretical findings, demonstrating the practical efficacy of our approach in achieving improved safety and learning efficiency in safe offline settings. Xiyue Peng, Honghao Wei, Xin Liu 0049 |
NeurIPS | 2 |