Zhiyu Mei

dblp:299/5277 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 33% Language models and text generation · 29% Motion planning and robot control · 14%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 50% Parallel and multicore computing · 25% High-performance computing · 25%
Databases, data mining, and information retrieval
1 paper
Data stream processing · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
0.812024
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study · ICML 2024
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.812024
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study · ICML 2024
Machine learning › Reinforcement learning › large-scale reinforcement learning
distributed reinforcement learning
0.812024
SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores · ICLR 2024
Robotics › Legged, aerial and field robots › legged robots
quadruped robot
0.812024
LAGOON: Language-Guided Motion Control · ICRA 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study · ICML 2024
Robotics › Motion planning and robot control
robot learning
0.812024
LAGOON: Language-Guided Motion Control · ICRA 2024
Distributed systems › distributed machine learning
distributed reinforcement learning
0.812024
SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores · ICLR 2024
Distributed systems › distributed machine learning
distributed training
0.812024
SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores · ICLR 2024
High-performance computing › large-scale training
large-scale distributed training
0.812024
SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores · ICLR 2024
Parallel and multicore computing › parallel computing › parallel machine learning
parallel training
0.812024
SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores · ICLR 2024
Data stream processing
sketch-based stream processing
0.512021
DHS: Adaptive Memory Layout Organization of Sketch Slots for Fast and Accurate Data Stream Processing · KDD 2021
Machine learning › Transfer learning and domain adaptation › sim-to-real transfer
domain randomization
0.212024
LAGOON: Language-Guided Motion Control · ICRA 2024
Machine learning › Reinforcement learning
policy optimization
0.212024
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study · ICML 2024
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.212024
LAGOON: Language-Guided Motion Control · ICRA 2024

Methods — techniques the papers use, named apart from their topics

dataflow abstraction · 1.5reward modeling · 0.8reinforcement learning · 0.8proximal policy optimization · 0.8pretrained motion generation · 0.8domain randomization · 0.8direct preference optimization · 0.8sketch · 0.5hashing · 0.5
YearPublicationVenuePosition
2025 AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
abstract
Reinforcement learning (RL) has become a trending paradigm for training large language models (LLMs), particularly for reasoning tasks. Effective RL for LLMs requires massive parallelization and poses an urgent need for efficient training systems. Most existing large-scale RL systems for LLMs are synchronous by alternating generation and training in a batch setting, where the rollouts in each training batch are generated by the same (or latest) model. This stabilizes RL training but suffers from severe system-level inefficiency. Generation must wait until the longest output in the batch is completed before model update, resulting in GPU underutilization. We present AReaL, a fully asynchronous RL system that completely decouples generation from training. Rollout workers in AReaL continuously generate new outputs without waiting, while training workers update the model whenever a batch of data is collected. AReaL also incorporates a collection of system-level optimizations, leading to substantially higher GPU utilization. To stabilize RL training, AReaL balances the workload of rollout and training workers to control data staleness, and adopts a staleness-enhanced PPO variant to better handle outdated training samples. Extensive experiments on math and code reasoning benchmarks show that AReaL achieves up to 2.77x training speedup compared to synchronous systems with the same number of GPUs and matched or even improved final performance. The code of AReaL is available at https://github.com/inclusionAI/AReaL/.
Jiaxuan Gao, Xujie Shen, Zhiyu Mei, Chuyi He, Shusheng Xu, Jun Mei, Tongkai Yang, Binhang Yuan
NeurIPS5
2025 How Far Are We from Optimal Reasoning Efficiency?
abstract
Large Reasoning Models (LRMs) demonstrate remarkable problem-solving capabilities through extended Chain-of-Thought (CoT) reasoning but often produce excessively verbose and redundant reasoning traces. This inefficiency incurs high inference costs and limits practical deployment. While existing fine-tuning methods aim to improve reasoning efficiency, assessing their efficiency gains remains challenging due to inconsistent evaluations. In this work, we introduce the ***reasoning efficiency frontiers***, empirical upper bounds derived from fine-tuning a base LRM (DeepSeek-R1-Distill-Qwen-1.5B/7B) across diverse approaches and training configurations. Based on these frontiers, we propose the ***Reasoning Efficiency Gap (REG)***, a unified metric quantifying deviations of any fine-tuned LRMs from these frontiers. Systematic evaluation on challenging mathematical benchmarks, AMC23, AIME24, and AIME25, reveals significant gaps in current methods: they either sacrifice accuracy for short length or use excessive tokens to achieve sub-optimal accuracies despite high overall accuracy. To reduce the efficiency gap, we propose ***REO-RL***, a Reinforcement Learning algorithm that optimizes reasoning efficiency by targeting a sparse set of token budgets. Leveraging numerical integration over strategically selected budgets, REO-RL approximates the full efficiency objective with low error using a small set of token budgets. Experiments show that, compared to vanilla RL with outcome reward, REO-RL reduces the reasoning efficiency gap by 74.5\% and 64.2\% in the 1.5B and 7B settings. The 7B LRM fine-tuned with REO-RL achieves reasoning conciseness surpassing frontier LRMs like Qwen3 and Claude Sonnet 3.7. Ablation studies confirm the efficacy of our token budget strategy and highlight REO-RL’s flexibility across design choices. This work establishes a systematic framework for evaluating and optimizing reasoning efficiency in LRMs. We will release the related code, data, and models to support future research on efficient reasoning in LRMs.
Jiaxuan Gao, Shu Yan, Qixin Tan, Shusheng Xu, Zhiyu Mei, Kaifeng Lyu, Yi Wu 0013
NeurIPS7
2024 SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores
abstract
The ever-growing complexity of reinforcement learning (RL) tasks demands a distributed system to efficiently generate and process a massive amount of data. However, existing open-source libraries suffer from various limitations, which impede their practical use in challenging scenarios where large-scale training is necessary. In this paper, we present a novel abstraction on the dataflows of RL training, which unifies diverse RL training applications into a general framework. Following this abstraction, we develop a scalable, efficient, and extensible distributed RL system called ReaLly Scalable RL (SRL), which allows efficient and massively parallelized training and easy development of customized algorithms. Our evaluation shows that SRL outperforms existing academic libraries, reaching at most 21x higher training throughput in a distributed setting. On learning performance, beyond performing and scaling well on common RL benchmarks with different RL algorithms, SRL can reproduce the same solution in the challenging hide-and-seek environment as reported by OpenAI with up to 5x speedup in wallclock time. Notably, SRL is the first in the academic community to perform RL experiments at a large scale with over 15k CPU cores. SRL anonymous repository is available at: https://anonymous.4open.science/r/srl-1E45/.
Zhiyu Mei, Jiaxuan Gao, Guangju Wang, Huanchen Zhang, Yi Wu 0013
ICLR1
2024 Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
abstract
Reinforcement Learning from Human Feedback (RLHF) is currently the most widely used method to align large language models (LLMs) with human preferences. Existing RLHF methods can be roughly categorized as either reward-based or reward-free. Novel applications such as ChatGPT and Claude leverage reward-based methods that first learn a reward model and apply actor-critic algorithms, such as Proximal Policy Optimization (PPO). However, in academic benchmarks, state-of-the-art results are often achieved via reward-free methods, such as Direct Preference Optimization (DPO). Is DPO truly superior to PPO? Why does PPO perform poorly on these benchmarks? In this paper, we first conduct both theoretical and empirical studies on the algorithmic properties of DPO and show that DPO may have fundamental limitations. Moreover, we also comprehensively examine PPO and reveal the key factors for the best performances of PPO in fine-tuning LLMs. Finally, we benchmark DPO and PPO across a collection of RLHF testbeds, ranging from dialogue to code generation. Experiment results demonstrate that PPO is able to surpass other alignment methods in all cases and achieve state-of-the-art results in challenging code competitions.
Shusheng Xu, Jiaxuan Gao, Zhiyu Mei, Guangju Wang, Chao Yu 0005, Yi Wu 0013
ICML6
2024 LAGOON: Language-Guided Motion Control
abstract
We aim to control a robot to physically behave in the real world following any high-level language command like "cartwheel" or "kick". Although human motion datasets exist, this task remains particularly challenging since generative models can produce physically unrealistic motions, which will be more severe for robots due to different body structures and physical properties. Deploying such a motion to a physical robot can cause even greater difficulties due to the sim2real gap. We develop LAnguage-Guided mOtion cONtrol (LAGOON), a multi-phase reinforcement learning (RL) method to generate physically realistic robot motions under language commands. LAGOON first leverages a pretrained model to generate a human motion from a language command. Then an RL phase trains a control policy in simulation to mimic the generated human motion. Finally, with domain randomization, our learned policy can be deployed to a quadrupedal robot, leading to a quadrupedal robot that can take diverse behaviors in the real world under natural language commands.
Shusheng Xu, Huaijie Wang, Yutao Ouyang, Jiaxuan Gao, Zhiyu Mei, Chao Yu 0005, Yi Wu 0013
ICRA5
2021 DHS: Adaptive Memory Layout Organization of Sketch Slots for Fast and Accurate Data Stream Processing
abstract
Data stream processing is a crucial computation task in data mining applications. The rigid and fixed data structures in existing solutions limit their accuracy, throughput, and generality in measurement tasks. We propose Dynamic Hierarchical Sketch (DHS), a sketch-based hybrid solution targeting these properties. During the online stream processing, DHS hashes items to buckets and organizes cells in each bucket dynamically; the size of all cells in a bucket is adjusted adaptively to the actual size and distribution of flows. Thus, memory is efficiently used to precisely record elephant flows and cover more mice flows. Implementation and evaluation show that DHS achieves high accuracy, high throughput, and high generality on five measurement tasks: flow size estimation, flow size distribution estimation, heavy hitter detection, heavy changer detection, and entropy estimation.
Bohan Zhao, Xiang Li 0156, Boyu Tian, Zhiyu Mei, Wenfei Wu
KDD4