Shuozhe Li

dblp:296/3004 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0004-1289-5830ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 70% Efficient and distributed learning · 15% Language models and text generation · 15%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › offline reinforcement learning
dual reinforcement learning
0.912025
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning
imitation learning
0.912025
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning · ICLR 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning · NeurIPS 2025
Cloud and datacenter computing › inference serving
LLM serving
0.912025
StitchLLM: Serving LLMs, One Block at a Time · ACL (1) 2025
Cloud and datacenter computing
serverless computing
0.912025
StitchLLM: Serving LLMs, One Block at a Time · ACL (1) 2025
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning
0.312025
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning
policy optimization
0.312025
ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning · NeurIPS 2025
Cloud and datacenter computing › datacenter storage
block-level scheduling
0.312025
StitchLLM: Serving LLMs, One Block at a Time · ACL (1) 2025
Cloud and datacenter computing
cluster resource management and scheduling
0.312025
StitchLLM: Serving LLMs, One Block at a Time · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

memory management · 1.7block-level scheduling · 1.7visitation distribution ratio estimation · 0.9self-explanation · 0.9discriminator-weighted behavior cloning · 0.9chain-of-thought · 0.9GRPO · 0.9
YearPublicationVenuePosition
2025 StitchLLM: Serving LLMs, One Block at a Time
abstract
Bodun Hu, Shuozhe Li, Saurabh Agarwal, Myungjin Lee, Akshay Jajoo, Jiamin Li, Le Xu, Geon-Woo Kim, Donghyun Kim, Hong Xu, Amy Zhang, Aditya Akella. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Bodun Hu, Shuozhe Li, Myungjin Lee, Akshay Jajoo, Jiamin Li 0002, Geon-Woo Kim, Donghyun Kim 0002, Hong Xu 0001, Amy Zhang 0001, Aditya Akella
ACL (1)2
2025 An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
abstract
We introduce Iterative Dual Reinforcement Learning (IDRL), a new method that takes an optimal discriminator-weighted imitation view of solving RL. Our method is motivated by a simple experiment in which we find training a discriminator using the offline dataset plus an additional expert dataset and then performing discriminator-weighted behavior cloning gives strong results on various types of datasets. That optimal discriminator weight is quite similar to the learned visitation distribution ratio in Dual-RL, however, we find that current Dual-RL methods do not correctly estimate that ratio. In IDRL, we propose a correction method to iteratively approach the optimal visitation distribution ratio in the offline dataset given no addtional expert dataset. During each iteration, IDRL removes zero-weight suboptimal transitions using the learned ratio from the previous iteration and runs Dual-RL on the remaining subdataset. This can be seen as replacing the behavior visitation distribution with the optimized visitation distribution from the previous iteration, which theoretically gives a curriculum of improved visitation distribution ratios that are closer to the optimal discriminator weight. We verify the effectiveness of IDRL on various kinds of offline datasets, including D4RL datasets and more realistic corrupted demonstrations. IDRL beats strong Primal-RL and Dual-RL baselines in terms of both performance and stability, on all datasets.
Haoran Xu 0003, Shuozhe Li, Harshit Sikchi, Scott Niekum, Amy Zhang 0001
ICLR2
2025 ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
abstract
Self-improvement via RL often fails on complex reasoning tasks because GRPO-style post-training methods rely on the model’s initial ability to generate positive samples. Without guided exploration, these approaches merely reinforce what the model already knows (distribution-sharpening) rather than enabling the model to solve problems where it initially generates no correct solutions. To unlock reasoning ability in such settings, the model must explore new reasoning trajectories beyond its current output distribution. Such exploration requires access to sufficiently good positive samples to guide the learning. While expert demonstrations seem like a natural solution, we find that they are often ineffective in RL post-training. Instead, we identify two key properties of effective positive samples: they should (1) be likely under the current policy, and (2) increase the model’s likelihood of predicting the correct answer. Based on these insights, we propose \textbf{Self-Explanation Policy Optimization (ExPO)}—a simple and modular framework that generates such samples by conditioning on the ground-truth answer. ExPO enables efficient exploration and guides the model to produce reasoning trajectories more aligned with its policy than expert-written CoTs, while ensuring higher quality than its own (incorrect) samples. Experiments show that ExPO improves both learning efficiency and final performance on reasoning benchmarks, surpassing expert-demonstration-based methods in challenging settings such as MATH level-5, where the model initially struggles the most.
Ruiyang Zhou, Shuozhe Li, Amy Zhang 0001, Liu Leqi
NeurIPS2
2021 Rotation Sensing Using Passive RFID Tags
abstract
Rotational movement is important in many applications, yet has been under-explored. In this paper, we explore the feasibility of using a single RFID reader antenna to simultaneously sense rotation and translation movement (i.e., rotation axis, rotation speed, and translation speed). We exploit the polarization in RFID to enable motion sensing. We develop an analytical model to capture the impact of polarization on the received signal and an optimization framework to incorporate the model to estimate the movement. We implement our system, Tag-based Inertial Measurement Unit (TIMU), and demonstrate its effectiveness through an extensive evaluation. To our knowledge, this is the first system that tracks general motion using a single RFID reader antenna.
Swadhin Pradhan, Shuozhe Li, Lili Qiu
MobiHoc2