Zaijun Wang

dblp:216/1545 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 65% Efficient and distributed learning · 35%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 61% GPUs and heterogeneous computing · 21% Performance modeling and evaluation · 18%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model inference
1.012026
PiLLM: Resource-Efficient LLM Inference Using Workload Prediction · EuroSys 2026
Machine learning › Efficient and distributed learning
resource allocation
1.012026
PiLLM: Resource-Efficient LLM Inference Using Workload Prediction · EuroSys 2026
Natural language and speech › Language models and text generation › text generation
structured generation
0.912025
Pre³: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation · ACL (1) 2025
Machine learning and data management
inference serving
0.912025
Past-Future Scheduler for LLM Serving under SLA Guarantees · ASPLOS (2) 2025
Compilers and program optimization
parsing
0.912025
Pre³: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation · ACL (1) 2025
Cloud and datacenter computing
inference serving
0.912025
Past-Future Scheduler for LLM Serving under SLA Guarantees · ASPLOS (2) 2025
GPUs and heterogeneous computing
GPU resource management
0.312026
PiLLM: Resource-Efficient LLM Inference Using Workload Prediction · EuroSys 2026

Methods — techniques the papers use, named apart from their topics

workload prediction · 2.0dynamic resource allocation · 2.0past-future scheduling · 1.7memory requirement estimation · 1.7constrained decoding · 1.7
YearPublicationVenuePosition
2026 PiLLM: Resource-Efficient LLM Inference Using Workload Prediction
abstract
LLM inference demands substantial GPU resources, but its highly variable workload characteristics create efficiency challenges at both inter-GPU and intra-GPU levels. To meet Service Level Objectives (SLOs), existing systems typically overprovision resources in two ways: allocating excess GPUs to handle peak loads and reserving excessive memory per request to prevent out-of-memory during token generation. We introduce PiLLM (Predictable inference for LLMs), a system that addresses these inefficiencies through accurate workload prediction and dynamic resource allocation.
Yunqian Fan, Shihao Bai, Ruihao Gong, Zaijun Wang
EuroSys4
2026 PMARL: Multi-Agent Reinforcement Learning in Large-Scale Systems
abstract
Large-scale multi-agent systems face two core challenges: inefficient policy learning and the explosion of state dimensions. Existing methods often rely on manually designed task sequences to guide agents’ learning in stages, but these designs lack adaptability to agents’ learning abilities, making it difficult to ensure the rationality of task difficulty. Moreover, the representation capability of current network structures is limited, making it challenging to efficiently handle high-dimensional state information and complex interaction relationships. To address these issues, we propose a Progressive Multi-Agent Reinforcement Learning (PMARL) framework. PMARL introduces a task adapter that adaptively selects task difficulty based on agents’ learning abilities, eliminating reliance on manual experience. Additionally, a Dynamic Dimension Adaptive Network (DDAN) is designed, incorporating hypernetwork and self-attention mechanisms to achieve adaptive feature extraction of high-dimensional states and efficient representation of agent interaction relationships. Experimental results demonstrate that PMARL exhibits higher efficiency and better adaptability compared to existing methods when addressing large-scale multi-agent tasks.
Baofu Fang, Hao Wang 0008, Kui Yu, Zaijun Wang
ACM Trans. Intell. Syst. Technol.5
2025 Pre³: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation
abstract
Junyi Chen, Shihao Bai, Zaijun Wang, Siyu Wu, Chuheng Du, Hailong Yang, Ruihao Gong, Shengzhong Liu, Fan Wu, Guihai Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shihao Bai, Zaijun Wang, Siyu Wu 0001, Chuheng Du, Hailong Yang 0002, Ruihao Gong, Shengzhong Liu, Fan Wu 0006, Guihai Chen
ACL (1)3
2025 Past-Future Scheduler for LLM Serving under SLA Guarantees
abstract
The exploration and application of Large Language Models (LLMs) is thriving. To reduce deployment costs, continuous batching has become an essential feature in current service frameworks. The effectiveness of continuous batching relies on an accurate estimate of the memory requirements of requests. However, due to the diversity in request output lengths, existing frameworks tend to adopt aggressive or conservative schedulers, which often result in significant overestimation or underestimation of memory consumption. Consequently, they suffer from harmful request evictions or prolonged queuing times, failing to achieve satisfactory throughput under strict Service Level Agreement (SLA) guarantees (a.k.a. goodput), across various LLM application scenarios with differing input-output length distributions. To address this issue, we propose a novel Past-Future scheduler that precisely estimates the peak memory resources required by the running batch via considering the historical distribution of request output lengths and calculating memory occupancy at each future time point. It adapts to applications with all types of input-output length distributions, balancing the trade-off between request queuing and harmful evictions, thereby consistently achieving better goodput. Furthermore, to validate the effectiveness of the proposed scheduler, we developed a high-performance LLM serving framework, LightLLM, that implements the Past-Future scheduler. Compared to existing aggressive or conservative schedulers, LightLLM demonstrates superior goodput, achieving up to 2-3× higher goodput than other schedulers under heavy loads. LightLLM is open source to boost the research in such direction (https://github.com/ModelTC/lightllm).
Ruihao Gong, Shihao Bai, Siyu Wu 0001, Yunqian Fan, Zaijun Wang, Hailong Yang 0002, Xianglong Liu 0001
ASPLOS (2)5
2021 Visual SLAM for robot navigation in healthcare facility
Baofu Fang, Gaofei Mei, Xiaohui Yuan 0001, Zaijun Wang, Junyang Wang 0004
Pattern Recognit.5
2020 Distributed task allocation method based on self-awareness of autonomous robots
Zaijun Wang, Jinlin Zhu, YunTing Ma, Zifan Li
J. Supercomput.1
2019 Collaborative task assignment of interconnected, affective robots towards autonomous healthcare assistant
Baofu Fang, Zaijun Wang, Mohamed Elhoseny, Xiaohui Yuan 0001
Future Gener. Comput. Syst.3