VLDB 2026 Research / reviewers in the wild / expert
Shuo Shen 0002
dblp:65/602-2
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0001-5401-4078ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 77% Knowledge representation and reasoning · 15% Optimization for machine learning · 8% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
language-conditioned reinforcement learning |
1.0 | 1 | 2026 | Complex Instruction Following with Diverse Style Policies in Football Games · AAAI 2026 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
1.0 | 1 | 2026 | Complex Instruction Following with Diverse Style Policies in Football Games · AAAI 2026 |
Machine learning › Reinforcement learning
policy learning |
0.9 | 1 | 2025 | Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement Learning · AAAI 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
symbolic regression |
0.9 | 1 | 2025 | Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement Learning · AAAI 2025 |
Machine learning › Reinforcement learning
constrained reinforcement learning |
0.8 | 1 | 2024 | Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C Estimation · AAAI 2024 |
Machine learning › Reinforcement learning
value function estimation |
0.8 | 1 | 2024 | Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C Estimation · AAAI 2024 |
Machine learning › Optimization for machine learning
constrained optimization |
0.2 | 1 | 2024 | Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C Estimation · AAAI 2024 |
Machine learning › Optimization for machine learning › constrained optimization
lagrangian methods |
0.2 | 1 | 2024 | Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C Estimation · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
style parameters · 1.0style interpreter · 1.0reinforcement learning · 0.9mixed path entropy · 0.9dynamic gating · 0.9ensemble learning · 0.8adaptive ensemble c-learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Complex Instruction Following with Diverse Style Policies in Football GamesabstractDespite advancements in language-controlled reinforcement learning (LC-RL) for basic domains and straightforward commands (e.g., object manipulation and navigation), effectively extending LC-RL to comprehend and execute high-level or abstract instructions in complex, multi-agent environments, such as football games, remains a significant challenge. To address this gap, we introduce Language-Controlled Diverse Style Policies (LCDSP), a novel LC-RL paradigm specifically designed for complex scenarios. LCDSP comprises two key components: a Diverse Style Training (DST) method and a Style Interpreter (SI). The DST method efficiently trains a single policy capable of exhibiting a wide range of diverse behaviors by modulating agent actions through style parameters (SP). The SI is designed to accurately and rapidly translate high-level language instructions into these corresponding SP. Through extensive experiments in a complex 5v5 football environment, we demonstrate that LCDSP effectively comprehends abstract tactical instructions and accurately executes the desired diverse behavioral styles, showcasing its potential for complex, real-world applications. Chenglu Sun, Shuo Shen 0002, Haonan Hu, Wei Zhou 0063, Chen Chen 0039 |
AAAI | 2 |
| 2025 | Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement LearningabstractSymbolic regression (SR) has emerged as a pivotal technique for uncovering the intrinsic information within data and enhancing the interpretability of AI models. However, current state-of-the-art (sota) SR methods struggle to perform correct recovery of symbolic expressions from high-noise data. To address this issue, we introduce a novel noise-resilient SR (NRSR) method capable of recovering expressions from high-noise data. Our method leverages a novel reinforcement learning (RL) approach in conjunction with a designed noise-resilient gating module (NGM) to learn symbolic selection policies. The gating module can dynamically filter the meaningless information from high-noise data, thereby demonstrating a high noise-resilient capability for the SR process. And we also design a mixed path entropy (MPE) bonus term in the RL process to increase the exploration capabilities of the policy. Experimental results demonstrate that our method significantly outperforms several popular baselines on benchmarks with high-noise data. Furthermore, our method also can achieve sota performance on benchmarks with clean data, showcasing its robustness and efficacy in SR tasks. Chenglu Sun, Shuo Shen 0002, Wenzhi Tao, Deyi Xue, Zixia Zhou |
AAAI | 2 |
| 2025 | Enhancing Offline Safe Reinforcement Learning with Trajectory-Constrained Diffusion Planning
Youfang Lin, Shuo Shen 0002, Hanfeng Lin, Peng Cheng 0013, Sheng Han 0001, Kai Lv 0002 |
AAMAS | 3 |
| 2025 | Enhancing AI-Bot Strength and Strategy Diversity in Adversarial Games: A Novel Deep Reinforcement Learning FrameworkabstractDeep reinforcement learning (DRL) has emerged as a leading technique for designing AI-bots in the gaming industry. However, practical implementation of DRL-trained bots often encounter two significant challenges: improving strength and diversifying strategies to satisfy player expectations. We observe that the strength of AI-bots are intrinsically tied to the diversity of emerged strategies. Considering this relationship, we introduce diversity is strength (DIS), a novel DRL training framework capable of concurrently training multiple types of AI-bots for adversarial games. These bots are interconnected through an elaborated history model pool (HMP) structure, thereby improving their strength and strategy diversity to tackle the aforementioned challenges. We further devise a model evaluation and sampling scheme to form the HMP, identify superior models, and enrich the model strategies. The DIS can generate diverse and reliable strategies without the need for human data. This method is validated by achieving first-place finishes in two AI competitions based on complex adversarial games, including Google Research Football and Olympic Games. Experiments demonstrate that bots trained using DIS attain an excellent performance and plentiful strategies. Specifically, diversity analysis demonstrates that the trained bots possess a wealth of strategies, and ablation studies confirm the beneficial impact of the designed modules on the training process. Chenglu Sun, Shuo Shen 0002, Deyi Xue, Wenzhi Tao, Zixia Zhou |
IEEE Trans. Games | 2 |
| 2024 | Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C EstimationabstractIn the domain of real-world agents, the application of Reinforcement Learning (RL) remains challenging due to the necessity for safety constraints. Previously, Constrained Reinforcement Learning (CRL) has predominantly focused on on-policy algorithms. Although these algorithms exhibit a degree of efficacy, their interactivity efficiency in real-world settings is sub-optimal, highlighting the demand for more efficient off-policy methods. However, off-policy CRL algorithms grapple with challenges in precise estimation of the C-function, particularly due to the fluctuations in the constrained Lagrange multiplier. Addressing this gap, our study focuses on the nuances of C-value estimation in off-policy CRL and introduces the Adaptive Ensemble C-learning (AEC) approach to reduce these inaccuracies. Building on state-of-the-art off-policy algorithms, we propose AEC-based CRL algorithms designed for enhanced task optimization. Extensive experiments on nine constrained robotics tasks reveal the superior interaction efficiency and performance of our algorithms in comparison to preceding methods. Youfang Lin, Shuo Shen 0002, Sheng Han 0001, Kai Lv 0002 |
AAAI | 3 |