VLDB 2026 Research / reviewers in the wild / expert
Shengyi Huang
dblp:251/8731
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-4986-1365ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 74% Language models and text generation · 23% Efficient and distributed learning · 3% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Performance modeling and evaluation · 72% Distributed systems · 28% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
deep reinforcement learning |
1.3 | 2 | 2024 | Cleanba: A Reproducible and Efficient Distributed Reinforcement Learning Platform · ICLR 2024 CleanRL: High-quality Single-file Implementations of Deep Reinforcement Learning Algorithms · J. Mach. Learn. Res. 2022 |
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation
instruction following |
0.9 | 1 | 2025 | Generalizing Verifiable Instruction Following · NeurIPS 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models · ICLR 2025 |
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards |
0.9 | 1 | 2025 | Generalizing Verifiable Instruction Following · NeurIPS 2025 |
Machine learning › Reinforcement learning › large-scale reinforcement learning
distributed reinforcement learning |
0.8 | 1 | 2024 | Cleanba: A Reproducible and Efficient Distributed Reinforcement Learning Platform · ICLR 2024 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.7 | 1 | 2023 | Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 Tricks · NeurIPS 2023 |
Machine learning › Reinforcement learning
policy optimization |
0.7 | 1 | 2023 | Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 Tricks · NeurIPS 2023 |
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization |
0.7 | 1 | 2023 | Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 Tricks · NeurIPS 2023 |
Performance modeling and evaluation › simulation › parallel and distributed simulation
parallel simulation |
0.6 | 1 | 2022 | EnvPool: A Highly Parallel Reinforcement Learning Environment Execution Engine · NeurIPS 2022 |
Machine learning › Efficient and distributed learning › efficient training
compute-efficient training |
0.3 | 1 | 2025 | Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models · ICLR 2025 |
Distributed systems › distributed machine learning
distributed training |
0.2 | 1 | 2024 | Cleanba: A Reproducible and Efficient Distributed Reinforcement Learning Platform · ICLR 2024 |
Machine learning › Reinforcement learning
reinforcement learning training |
0.2 | 1 | 2022 | EnvPool: A Highly Parallel Reinforcement Learning Environment Execution Engine · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
actor-learner framework · 1.5PPO · 1.5IMPALA · 1.5parallel environment execution · 1.1asynchronous simulation · 1.1reinforcement learning with verifiable rewards · 0.9online DPO · 0.9off-policy reinforcement learning · 0.9constraint verification · 0.9ablation study · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language ModelsabstractThe dominant paradigm for RLHF is *online* and *on-policy* RL: synchronously generating from the large language model (LLM) policy, labelling with a reward model, and learning using feedback on the LLM's own outputs. While performant, this paradigm is computationally inefficient. Inspired by classical deep RL literature, we propose separating generation and learning in RLHF. This enables asynchronous generation of new samples while simultaneously training on old samples, leading to faster training and more compute-optimal scaling. However, asynchronous training relies on an underexplored regime, online but *off-policy* RLHF: learning on samples from previous iterations of our model which give a worse training signal. We tackle the fundamental challenge in this regime: how much off-policyness can we tolerate for asynchronous training to speed up learning but maintain performance? Among several RLHF algorithms we test, online DPO is found to be most robust to off-policy data, and robustness increases with the scale of the policy model. We study further compute optimizations for asynchronous RLHF but find that they come at a performance cost, giving rise to a trade-off. We verify the scalability of asynchronous RLHF by training a general-purpose chatbot from LLaMA 3.1 8B on an instruction-following task $\sim$40\% faster than a synchronous run while matching final performance. Finally, we extend our results to math and reasoning to demonstrate asynchronous RL can finetune Rho 1B on GSM8k $\sim$70\% faster while matching synchronous accuracy. Michael Noukhovitch, Shengyi Huang, Sophie Xhonneux, Seyed Arian Hosseini, Rishabh Agarwal, Aaron C. Courville |
ICLR | 2 |
| 2025 | Generalizing Verifiable Instruction FollowingabstractA crucial factor for successful human and AI interaction is the ability of language models or chatbots to follow human instructions precisely. A common feature of instructions are output constraints like only answer with yes or no" ormention the word `abracadabra' at least 3 times" that the user adds to craft a more useful answer.Even today's strongest models struggle with fulfilling such constraints. We find that most models strongly overfit on a small set of verifiable constraints from the benchmarks that test these abilities, a skill called precise instruction following, and are not able to generalize well to unseen output constraints. We introduce a new benchmark, IFBench, to evaluate precise instruction following generalization on 58 new, diverse, and challenging verifiable out-of-domain constraints. In addition, we perform an extensive analysis of how and on what data models can be trained to improve precise instruction following generalization. Specifically, we carefully design constraint verification modules and show that reinforcement learning with verifiable rewards (RLVR) significantly improves instruction following. In addition to IFBench, we release 29 additional new hand-annotated training constraints and verification functions, RLVR training prompts, and code. Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison, Shengyi Huang, Pradeep Dasigi, Nathan Lambert 0001, Hannaneh Hajishirzi |
NeurIPS | 5 |
| 2024 | Cleanba: A Reproducible and Efficient Distributed Reinforcement Learning PlatformabstractDistributed Deep Reinforcement Learning (DRL) aims to leverage more computational resources to train autonomous agents with less training time. Despite recent progress in the field, reproducibility issues have not been sufficiently explored. This paper first shows that the typical actor-learner framework can have reproducibility issues even if hyperparameters are controlled. We then introduce Cleanba, a new open-source platform for distributed DRL that proposes a highly reproducible architecture. Cleanba implements highly optimized distributed variants of PPO and IMPALA. Our Atari experiments show that these variants can obtain equivalent or higher scores than strong IMPALA baselines in moolib and torchbeast and PPO baseline in CleanRL. However, Cleanba variants present 1) shorter training time and 2) more reproducible learning curves in different hardware settings. Shengyi Huang, Jiayi Weng, Rujikorn Charakorn, Zhongwen Xu, Santiago Ontañón |
ICLR | 1 |
| 2023 | Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 TricksabstractMost reinforcement learning methods rely heavily on dense, well-normalized environment rewards. DreamerV3 recently introduced a model-based method with a number of tricks that mitigate these limitations, achieving state-of-the-art on a wide range of benchmarks with a single set of hyperparameters. This result sparked discussion about the generality of the tricks, since they appear to be applicable to other reinforcement learning algorithms. Our work applies DreamerV3's tricks to PPO and is the first such empirical study outside of the original work. Surprisingly, we find that the tricks presented do not transfer as general improvements to PPO. We use a high quality PPO reference implementation and present extensive ablation studies totaling over 10,000 A100 hours on the Arcade Learning Environment and the DeepMind Control Suite. Though our experiments demonstrate that these tricks do not generally outperform PPO, we identify cases where they succeed and offer insight into the relationship between the implementation tricks. In particular, PPO with these tricks performs comparably to PPO on Atari games with reward clipping and significantly outperforms PPO without reward clipping. Ryan Sullivan, Akarsh Kumar, Shengyi Huang, John Dickerson 0001, Joseph Suarez |
NeurIPS | 3 |
| 2022 | EnvPool: A Highly Parallel Reinforcement Learning Environment Execution EngineabstractThere has been significant progress in developing reinforcement learning (RL) training systems. Past works such as IMPALA, Apex, Seed RL, Sample Factory, and others, aim to improve the system's overall throughput. In this paper, we aim to address a common bottleneck in the RL training system, i.e., parallel environment execution, which is often the slowest part of the whole system but receives little attention. With a curated design for paralleling RL environments, we have improved the RL environment simulation speed across different hardware setups, ranging from a laptop and a modest workstation, to a high-end machine such as NVIDIA DGX-A100. On a high-end machine, EnvPool achieves one million frames per second for the environment execution on Atari environments and three million frames per second on MuJoCo environments. When running EnvPool on a laptop, the speed is 2.8x that of the Python subprocess. Moreover, great compatibility with existing RL training libraries has been demonstrated in the open-sourced community, including CleanRL, rl_games, DeepMind Acme, etc. Finally, EnvPool allows researchers to iterate their ideas at a much faster pace and has great potential to become the de facto RL environment execution engine. Example runs show that it only takes five minutes to train agents to play Atari Pong and MuJoCo Ant on a laptop. EnvPool is open-sourced at https://github.com/sail-sg/envpool. Jiayi Weng, Shengyi Huang, Bo Liu 0039, Denys Makoviichuk, Viktor Makoviychuk, Ting Luo 0009, Zhongwen Xu, Shuicheng Yan |
NeurIPS | 3 |
| 2022 | CleanRL: High-quality Single-file Implementations of Deep Reinforcement Learning AlgorithmsabstractCleanRL is an open-source library that provides high-quality single-file implementations of Deep Reinforcement Learning (DRL) algorithms. These single-file implementations are self-contained algorithm variant files such as dqn.py, ppo.py, and ppo_atari.py that individually include all algorithm variant's implementation details. Such a paradigm significantly reduces the complexity and the lines of code (LOC) in each implemented variant, which makes them quicker and easier to understand. This paradigm gives the researchers the most fine-grained control over all aspects of the algorithm in a single file, allowing them to prototype novel features quickly. Despite having succinct implementations, CleanRL's codebase is thoroughly documented and benchmarked to ensure performance is on par with reputable sources. As a result, CleanRL produces a repository tailor-fit for two purposes: 1) understanding all implementation details of DRL algorithms and 2) quickly prototyping novel features. CleanRL's source code can be found at https://github.com/vwxyzjn/cleanrl. Shengyi Huang, Rousslan Fernand Julien Dossa, Chang Ye, Jeff Braga, Dipam Chakraborty, Kinal Mehta, João G. M. Araújo |
J. Mach. Learn. Res. | 1 |
| 2021 | Gym-µRTS: Toward Affordable Full Game Real-time Strategy Games Research with Deep Reinforcement LearningabstractIn recent years, researchers have achieved great success in applying Deep Reinforcement Learning (DRL) algorithms to Real-time Strategy (RTS) games, creating strong autonomous agents that could defeat professional players in StarCraft II. However, existing approaches to tackle full games have high computational costs, usually requiring the use of thousands of GPUs and CPUs for weeks. This paper has two main contributions to address this issue: 1) We introduce Gym-JLRTS (pronounced “gym-micro-RTS”) as a fast-to-run RL environment for full-game RTS research and 2) we present a collection of techniques to scale DRL to play full-game µRTS as well as ablation studies to demonstrate their empirical importance. Our best-trained bot can defeat every µRTS bot we tested from the past µRTS competitions when working in a single-map setting, resulting in a state-of-the-art DRL agent while only taking about 60 hours of training using a single machine (one GPU, three vCPU. 16GB RAM). Shengyi Huang, Santiago Ontañón, Chris Bamford 0001, Lukasz Grela |
CoG | 1 |