Shuo Shen 0002

dblp:65/602-2 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0001-5401-4078ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 77% Knowledge representation and reasoning · 15% Optimization for machine learning · 8%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
language-conditioned reinforcement learning
1.012026
Complex Instruction Following with Diverse Style Policies in Football Games · AAAI 2026
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.012026
Complex Instruction Following with Diverse Style Policies in Football Games · AAAI 2026
Machine learning › Reinforcement learning
policy learning
0.912025
Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement Learning · AAAI 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
symbolic regression
0.912025
Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement Learning · AAAI 2025
Machine learning › Reinforcement learning
constrained reinforcement learning
0.812024
Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C Estimation · AAAI 2024
Machine learning › Reinforcement learning
value function estimation
0.812024
Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C Estimation · AAAI 2024
Machine learning › Optimization for machine learning
constrained optimization
0.212024
Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C Estimation · AAAI 2024
Machine learning › Optimization for machine learning › constrained optimization
lagrangian methods
0.212024
Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C Estimation · AAAI 2024

Methods — techniques the papers use, named apart from their topics

style parameters · 1.0style interpreter · 1.0reinforcement learning · 0.9mixed path entropy · 0.9dynamic gating · 0.9ensemble learning · 0.8adaptive ensemble c-learning · 0.8
YearPublicationVenuePosition
2026 Complex Instruction Following with Diverse Style Policies in Football Games
abstract
Despite advancements in language-controlled reinforcement learning (LC-RL) for basic domains and straightforward commands (e.g., object manipulation and navigation), effectively extending LC-RL to comprehend and execute high-level or abstract instructions in complex, multi-agent environments, such as football games, remains a significant challenge. To address this gap, we introduce Language-Controlled Diverse Style Policies (LCDSP), a novel LC-RL paradigm specifically designed for complex scenarios. LCDSP comprises two key components: a Diverse Style Training (DST) method and a Style Interpreter (SI). The DST method efficiently trains a single policy capable of exhibiting a wide range of diverse behaviors by modulating agent actions through style parameters (SP). The SI is designed to accurately and rapidly translate high-level language instructions into these corresponding SP. Through extensive experiments in a complex 5v5 football environment, we demonstrate that LCDSP effectively comprehends abstract tactical instructions and accurately executes the desired diverse behavioral styles, showcasing its potential for complex, real-world applications.
Chenglu Sun, Shuo Shen 0002, Haonan Hu, Wei Zhou 0063, Chen Chen 0039
AAAI2
2025 Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement Learning
abstract
Symbolic regression (SR) has emerged as a pivotal technique for uncovering the intrinsic information within data and enhancing the interpretability of AI models. However, current state-of-the-art (sota) SR methods struggle to perform correct recovery of symbolic expressions from high-noise data. To address this issue, we introduce a novel noise-resilient SR (NRSR) method capable of recovering expressions from high-noise data. Our method leverages a novel reinforcement learning (RL) approach in conjunction with a designed noise-resilient gating module (NGM) to learn symbolic selection policies. The gating module can dynamically filter the meaningless information from high-noise data, thereby demonstrating a high noise-resilient capability for the SR process. And we also design a mixed path entropy (MPE) bonus term in the RL process to increase the exploration capabilities of the policy. Experimental results demonstrate that our method significantly outperforms several popular baselines on benchmarks with high-noise data. Furthermore, our method also can achieve sota performance on benchmarks with clean data, showcasing its robustness and efficacy in SR tasks.
Chenglu Sun, Shuo Shen 0002, Wenzhi Tao, Deyi Xue, Zixia Zhou
AAAI2
2025 Enhancing Offline Safe Reinforcement Learning with Trajectory-Constrained Diffusion Planning
Youfang Lin, Shuo Shen 0002, Hanfeng Lin, Peng Cheng 0013, Sheng Han 0001, Kai Lv 0002
AAMAS3
2025 Enhancing AI-Bot Strength and Strategy Diversity in Adversarial Games: A Novel Deep Reinforcement Learning Framework
abstract
Deep reinforcement learning (DRL) has emerged as a leading technique for designing AI-bots in the gaming industry. However, practical implementation of DRL-trained bots often encounter two significant challenges: improving strength and diversifying strategies to satisfy player expectations. We observe that the strength of AI-bots are intrinsically tied to the diversity of emerged strategies. Considering this relationship, we introduce diversity is strength (DIS), a novel DRL training framework capable of concurrently training multiple types of AI-bots for adversarial games. These bots are interconnected through an elaborated history model pool (HMP) structure, thereby improving their strength and strategy diversity to tackle the aforementioned challenges. We further devise a model evaluation and sampling scheme to form the HMP, identify superior models, and enrich the model strategies. The DIS can generate diverse and reliable strategies without the need for human data. This method is validated by achieving first-place finishes in two AI competitions based on complex adversarial games, including Google Research Football and Olympic Games. Experiments demonstrate that bots trained using DIS attain an excellent performance and plentiful strategies. Specifically, diversity analysis demonstrates that the trained bots possess a wealth of strategies, and ablation studies confirm the beneficial impact of the designed modules on the training process.
Chenglu Sun, Shuo Shen 0002, Deyi Xue, Wenzhi Tao, Zixia Zhou
IEEE Trans. Games2
2024 Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C Estimation
abstract
In the domain of real-world agents, the application of Reinforcement Learning (RL) remains challenging due to the necessity for safety constraints. Previously, Constrained Reinforcement Learning (CRL) has predominantly focused on on-policy algorithms. Although these algorithms exhibit a degree of efficacy, their interactivity efficiency in real-world settings is sub-optimal, highlighting the demand for more efficient off-policy methods. However, off-policy CRL algorithms grapple with challenges in precise estimation of the C-function, particularly due to the fluctuations in the constrained Lagrange multiplier. Addressing this gap, our study focuses on the nuances of C-value estimation in off-policy CRL and introduces the Adaptive Ensemble C-learning (AEC) approach to reduce these inaccuracies. Building on state-of-the-art off-policy algorithms, we propose AEC-based CRL algorithms designed for enhanced task optimization. Extensive experiments on nine constrained robotics tasks reveal the superior interaction efficiency and performance of our algorithms in comparison to preceding methods.
Youfang Lin, Shuo Shen 0002, Sheng Han 0001, Kai Lv 0002
AAAI3