Chenbei Lu

dblp:264/1791 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-7715-4927ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 75% Learning theory · 12% Transfer learning and domain adaptation · 12%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
model-based reinforcement learning
1.722025
Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach · NeurIPS 2025
Overcoming the Curse of Dimensionality in Reinforcement Learning Through Approximate Factorization · ICML 2025
Machine learning › Reinforcement learning › markov decision process
factored markov decision process
0.912025
Overcoming the Curse of Dimensionality in Reinforcement Learning Through Approximate Factorization · ICML 2025
Machine learning › Reinforcement learning
model-free reinforcement learning
0.912025
Overcoming the Curse of Dimensionality in Reinforcement Learning Through Approximate Factorization · ICML 2025
Machine learning › Reinforcement learning
offline learning
0.912025
Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation › model adaptation
online adaptation
0.912025
Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach · NeurIPS 2025
Machine learning › Learning theory
sample complexity
0.912025
Overcoming the Curse of Dimensionality in Reinforcement Learning Through Approximate Factorization · ICML 2025
Machine learning › Reinforcement learning
value function
0.912025
Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach · NeurIPS 2025
Mathematical optimization
stochastic optimization
0.312025
Overcoming the Curse of Dimensionality in Reinforcement Learning Through Approximate Factorization · ICML 2025

Methods — techniques the papers use, named apart from their topics

graph-coloring-based synchronous sampling · 1.7bellman-jensen gap analysis · 0.9bayesian value learning · 0.9
YearPublicationVenuePosition
2025 Overcoming the Curse of Dimensionality in Reinforcement Learning Through Approximate Factorization
abstract
Factored Markov Decision Processes (FMDPs) offer a promising framework for overcoming the curse of dimensionality in reinforcement learning (RL) by decomposing high-dimensional MDPs into smaller and independently evolving components. Despite their potential, existing studies on FMDPs face three key limitations: reliance on perfectly factorizable models, suboptimal sample complexity guarantees for model-based algorithms, and the absence of model-free algorithms. To address these challenges, we introduce approximate factorization, which extends FMDPs to handle imperfectly factored models. Moreover, we develop a model-based algorithm and a model-free algorithm (in the form of variance-reduced Q-learning), both achieving the first near-minimax sample complexity guarantees for FMDPs. A key novelty in the design of these two algorithms is the development of a graph-coloring-based optimal synchronous sampling strategy. Numerical simulations based on the wind farm storage control problem corroborate our theoretical findings.
Chenbei Lu, Laixi Shi, Zaiwei Chen, Chenye Wu, Adam Wierman
ICML1
2025 Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach
abstract
Traditional reinforcement learning (RL) assumes the agents make decisions based on Markov decision processes (MDPs) with one-step transition models. In many real-world applications, such as energy management and stock investment, agents can access multi-step predictions of future states, which provide additional advantages for decision making. However, multi-step predictions are inherently high-dimensional: naively embedding these predictions into an MDP leads to an exponential blow-up in state space and the curse of dimensionality. Moreover, existing RL theory provides few tools to analyze prediction-augmented MDPs, as it typically works on one-step transition kernels and cannot accommodate multi-step predictions with errors or partial action-coverage. We address these challenges with three key innovations: First, we propose the \emph{Bayesian value function} to characterize the optimal prediction-aware policy tractably. Second, we develop a novel \emph{Bellman–Jensen Gap} analysis on the Bayesian value function, which enables characterizing the value of imperfect predictions. Third, we introduce BOLA (Bayesian Offline Learning with Online Adaptation), a two-stage model-based RL algorithm that separates offline Bayesian value learning from lightweight online adaptation to real-time predictions. We prove that BOLA remains sample-efficient even under imperfect predictions. We validate our theory and algorithm on synthetic MDPs and a real-world wind energy storage control problem.
Chenbei Lu, Zaiwei Chen, Tongxin Li 0001, Chenye Wu, Adam Wierman
NeurIPS1
2021 Efficiency or Fairness?: Carpooling Design for Online Ride-hailing Platform in Transport Hubs at Midnight
abstract
The online ride-hailing platform has revolutionized urban transport. However, there is much room for improvement. We consider meeting the demand for an online ride-hailing at transport hubs late at night, when the public transport system stops its operations. Passengers arriving late at night face a long wait before service. We launch the ride-hailing model in the theoretic framework of queueing and introduce the arrival and the service processes. To improve the efficiency of the ride-hailing platform, as well as to maintain fairness between different types of passengers, we study three variations of carpool service policies. We then provide practical guidelines on the trade-off between efficiency and fairness to assist the online platform designers. Specifically, we derive the analytical trade-off bounds with the passenger parameters. Furthermore, we suggest that these bounds can be good performance estimators for the empirical trade-off when only limited passenger information is available. This analysis motivates us to design the optimal service rate for the entire platform. Finally, we conduct numerical studies based on field data retrieved from Didi Chuxing, highlighting the remarkable performance of our proposed method in terms of improving the quality of online ride-hailing service.
Chenbei Lu, Jiaman Wu, Chenye Wu, Yongli Qin, Qun (Tracy) Li
SIGSPATIAL/GIS1