VLDB 2026 Research / reviewers in the wild / expert
Lin Li 0090
dblp:73/2252-90
· DBLP profile ↗
15ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0002-5446-6100ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 12 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pseudo-distribution elite critics: Enhancing accuracy in reinforcement learning value estimation
Yujia Zhang 0013, Lin Li 0090, Wei Wei 0018, Jiye Liang |
Neural Networks | 2 |
| 2025 | Improving Generalization in Offline Reinforcement Learning via Latent Distribution Representation LearningabstractDealing with the distribution shift is a significant challenge when building offline reinforcement learning (RL) models that can generalize from a static dataset to out-of-distribution (OOD) scenarios. Previous approaches have employed pessimism or conservatism strategies. More recently, data-driven work has taken a distributional perspective, treating offline data as a domain adaptation problem. However, these methods use heuristic techniques to simulate distribution shifts, resulting in a limited diversity of artificially created distribution gaps. In this paper, we propose a novel perspective: offline datasets inherently contain multiple latent distributions, with behavior data from diverse policies potentially following different distributions and data from the same policy across various time phases also exhibiting distribution variance. We introduce the Latent Distribution Representation Learning (LAD) framework, which aims to characterize the multiple latent distributions within offline data and reduce the distribution gaps between any pair of them. LAD consists of a min-max adversarial process: it first identifies the "worst-case" distributions to enlarge the diversity of distribution gaps and then reduces these gaps to learn invariant representations for generalization. We derive a generalization error bound to support LAD theoretically and verify its effectiveness through extensive experiments. Lin Li 0090, Wei Wei 0018, Qixian Yu, Jianye Hao, Jiye Liang |
AAAI | 2 |
| 2025 | Dynamic Uncertainty Estimation for Offline Reinforcement LearningabstractOffline reinforcement learning confronts the distributional shift challenge, a consequence of learning policy from static datasets. Current methods primarily handle this issue by aligning the learned policy with the behavior policy or conservatively estimating Q-values for out-of-distribution (OOD) actions. However, these approaches can lead to overly pessimistic estimation of Q-values of the OOD actions in unfamiliar situations, resulting in a suboptimal policy. To address this, we propose a new method, Dynamic Uncertainty estimation for Offline Reinforcement Learning. This method introduces a base density-truncated OOD data sampling approach to reduce the impact of extrapolation errors on uncertainty estimation. It enables conservative estimation of Q-values for OOD actions while avoiding negative impacts on in-distribution data. We also develop a dynamic uncertainty estimation mechanism to prevent excessive pessimism and enhance the generalization of the Q-function. This mechanism dynamically adjusts the degree of pessimism in the Q-function by minimizing the error between target and estimated values. Our method outperforms existing algorithms, as demonstrated by experimental results based on the D4RL benchmark, and proves its superiority in addressing the distributional shift challenge. Jiesheng Wang, Lin Li 0090, Wei Wei 0018, Yujia Zhang 0013, Xin Yang 0012 |
AAAI | 2 |
| 2025 | AGOD: Enhancing Multi-Agent Generalization via Attribution-Guided Observation DropoutabstractGeneralization remains a fundamental challenge in the field of Multi-Agent Reinforcement Learning (MARL). In real-world settings, agents frequently encounter changing environments and novel configurations of entities, resulting in unstable performance and significant degradation in Out-of-Distribution (OOD) scenarios. From the perspective of improving generalization, existing methods address this challenge through various approaches, such as dynamic policy adaptation, role specialization, and regularization techniques. Among them, dropout is a representative approach that enhances generalization by randomly omitting parts of the input during training. However, most existing dropout strategies overlook the varying importance of different entities, which can result in the removal of critical information and the retention of irrelevant observations, thereby limiting the model’s performance and stability in complex, dynamic environments. This oversight often results in unstable policies and limited generalization capabilities. In this work, we propose a new generalization method called Attribution-Guided Observation Dropout (AGOD). This method introduces an attribution coefficient to measure the contribution of each observed entity. It selectively drops those with higher attribution values during training, thereby encouraging agents to avoid over-reliance on key information and enhancing their generalization ability in complex environments. The proposed AGOD method considers entity importance, avoiding indiscriminate dropout, and shows sustained superior performance on the MPEV2 benchmark, proving its effectiveness in complex dynamic environments. Wei Wei 0018, Binchao Ma, Lin Li 0090, Huizhong Song, Fengjiao Li |
DAI | 3 |
| 2025 | Risk-aware Direct Preference Optimization under Nested Risk MeasureabstractWhen fine-tuning pre-trained Large Language Models (LLMs) to align with human values and intentions, maximizing the estimated reward can lead to superior performance, but it also introduces potential risks due to deviations from the reference model's intended behavior. Most existing methods typically introduce KL divergence to constrain deviations between the trained model and the reference model; however, this may not be sufficient in certain applications that require tight risk control. In this paper, we introduce Risk-aware Direct Preference Optimization (Ra-DPO), a novel approach that incorporates risk-awareness by employing a class of nested risk measures. This approach formulates a constrained risk-aware advantage function maximization problem and then converts the Bradley-Terry model into a token-level representation. The objective function maximizes the likelihood of the policy while suppressing the deviation between a trained model and the reference model using a sequential risk ratio, thereby enhancing the model's risk-awareness. Experimental results across three open-source datasets: IMDb Dataset, Anthropic HH Dataset, and AlpacaEval, demonstrate the proposed method's superior performance in balancing alignment performance and model drift. Lin Li 0090, Yajie Qi, Huizhong Song, Yaodong Yang 0001, Jun Wang 0012, Wei Wei 0018 |
NeurIPS | 2 |
| 2025 | Rethinking exploration-exploitation trade-off in reinforcement learning via cognitive consistency
Wei Wei 0018, Lin Li 0090, Jiye Liang |
Neural Networks | 3 |
| 2024 | Improving Generalization in Offline Reinforcement Learning via Adversarial Data SplittingabstractOffline Reinforcement Learning (RL) commonly suffers from the out-of-distribution (OOD) overestimation issue due to the distribution shift. Prior work gradually shifts their focus from suppressing OOD overestimation to avoiding overly conservative learning from suboptimal behavior policies to improve generalization. However, most approaches explicitly delimit boundaries for OOD actions based on the support in the dataset, which can potentially impede the data near these boundaries from acquiring realistic estimates. This paper investigates how to loosen the rigid demarcation of OOD boundaries, adaptively extracting knowledge from empirical data to implicitly improve the model's generalization to nearby unseen data. We introduce an adversarial data splitting (ADS) framework that enforces the model to generalize the distribution shifts simulated from the train/validation subsets splitting of the dataset. Specifically, ADS is modeled as a min-max optimization problem inspired by meta-learning and solved by iterating over the following two steps. First, we train the model on the train-subset to minimize its loss on the validation-subset. Then, we adversarially generate the "hardest" train/validation subsets with the maximum distribution shift, making the model incapable of generalization at that splitting. We derive a generalization error bound for theoretically understanding ADS and verify the effectiveness with extensive experiments. Code is available at https://github.com/DkING-lv6/ADS. Lin Li 0090, Wei Wei 0018, Qixian Yu, Jianye Hao, Jiye Liang |
ICML | 2 |
| 2024 | Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningabstractA challenging problem in seeking to bring multi-agent reinforcement learning (MARL) techniques into real-world applications, such as autonomous driving and drone swarms, is how to control multiple agents safely and cooperatively to accomplish tasks. Most existing safe MARL methods learn the centralized value function by introducing a global state to guide safety cooperation. However, the global coupling arising from agents’ safety constraints and the exponential growth of the state-action space size limit their applicability in instant communication or computing resource-constrained systems and larger multi-agent systems. In this paper, we develop a novel scalable and theoretically-justified multi-agent constrained policy optimization method. This method utilizes the rigorous bounds of the trust region method and the bounds of the truncated advantage function to provide a new local policy optimization objective for each agent. Also, we prove that the safety constraints and the joint policy improvement can be met when each agent adopts a sequential update scheme to optimize a $\kappa$-hop policy. Then, we propose a practical algorithm called Scalable MAPPO-Lagrangian (Scal-MAPPO-L). The proposed method’s effectiveness is verified on a collection of benchmark tasks, and the results support our theory that decentralized training with local interactions can still improve reward performance and satisfy safe constraints. Lin Li 0090, Wei Wei 0018, Huizhong Song, Yaodong Yang 0001, Jiye Liang |
NeurIPS | 2 |
| 2024 | Controlling estimation error in reinforcement learning via Reinforced Operation
Yujia Zhang 0013, Lin Li 0090, Wei Wei 0018, Xiu You, Jiye Liang |
Inf. Sci. | 2 |
| 2024 | Re-attentive experience replay in off-policy reinforcement learning
Wei Wei 0018, Lin Li 0090, Jiye Liang |
Mach. Learn. | 3 |
| 2024 | A unified framework to control estimation error in reinforcement learning
Yujia Zhang 0013, Lin Li 0090, Wei Wei 0018, Yunpeng Lv, Jiye Liang |
Neural Networks | 2 |
| 2023 | Set-membership Belief State-based Reinforcement Learning for POMDPsabstractReinforcement learning (RL) has made significant progress in areas such as Atari games and robotic control, where the agents have perfect sensing capabilities. However, in many real-world sequential decision-making tasks, the observation data could be noisy or incomplete due to the intrinsic low quality of the sensors or unexpected malfunctions; that is, the agent’s perceptions are rarely perfect. The current POMDP RL methods, such as particle-based and Gaussian-based, can only provide a probability estimate of hidden states rather than certain belief regions, which may lead to inefficient and even wrong decision-making. This paper proposes a novel algorithm called Set-membership Belief state-based Reinforcement Learning (SBRL), which consists of two parts: a Set-membership Belief state learning Model (SBM) for learning bounded belief state sets and an RL controller for making decisions based on SBM. We prove that our belief estimation method can provide a series of belief state sets that always contain the true states under the unknown-but-bounded (UBB) noise. The effectiveness of the proposed method is verified on a collection of benchmark tasks, and the results show that our method outperforms the state-of-the-art methods. Wei Wei 0018, Lin Li 0090, Huizhong Song, Jiye Liang |
ICML | 3 |
| 2023 | Multiple metric learning via local metric fusion
Xinyao Guo, Lin Li 0090, Chuangyin Dang, Jiye Liang, Wei Wei 0018 |
Inf. Sci. | 2 |
| 2023 | Multi-actor mechanism for actor-critic reinforcement learning
Lin Li 0090, Wei Wei 0018, Yujia Zhang 0013, Jiye Liang |
Inf. Sci. | 1 |
| 2022 | Controlling Underestimation Bias in Reinforcement Learning via Quasi-median OperationabstractHow to get a good value estimation is one of the key problems in reinforcement learning (RL). Current off-policy methods, such as Maxmin Q-learning, TD3 and TADD, suffer from the underestimation problem when solving the overestimation problem. In this paper, we propose the Quasi-Median Operation, a novel way to mitigate the underestimation bias by selecting the quasi-median from multiple state-action values. Based on the quasi-median operation, we propose Quasi-Median Q-learning (QMQ) for the discrete action tasks and Quasi-Median Delayed Deep Deterministic Policy Gradient (QMD3) for the continuous action tasks. Theoretically, the underestimation bias of our method is improved while the estimation variance is significantly reduced compared to Maxmin Q-learning, TD3 and TADD. We conduct extensive experiments on the discrete and continuous action tasks, and results show that our method outperforms the state-of-the-art methods. Wei Wei 0018, Yujia Zhang 0013, Jiye Liang, Lin Li 0090 |
AAAI | 4 |