VLDB 2026 Research / reviewers in the wild / expert
Yiyuan He
dblp:387/1240
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0003-2128-2852ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI InfrastructureabstractABSTRACT Objective Large Language Models (LLMs) are increasingly deployed in modern AI infrastructure, creating a strong demand for high‐throughput and resource‐efficient serving systems. Disaggregated LLM serving, which decouples prompt prefill from auto‐regressive decode to accommodate their heterogeneous compute and memory characteristics, has emerged as a promising architecture. However, existing disaggregated serving systems suffer from three fundamental limitations: static resource allocation that fails to adapt to highly dynamic workloads, severe load imbalance between compute‐bound prefill and memory‐bound decode stages, and prefix‐cache‐aware routing that skews load distribution and creates performance hotspots. These issues collectively limit resource utilization, scalability, and the ability to meet service level objectives (SLOs) under real‐world workloads. Methods To address these challenges, we propose BanaServe, a dynamic orchestration framework for disaggregated LLM serving that continuously rebalances both computational and memory resources across prefill and decode instances. BanaServe introduces three key mechanisms: (i) layer‐level weight migration to enable coarse‐grained redistribution of computation, (ii) attention‐level Key–Value (KV) cache migration for fine‐grained memory load balancing, and (iii) a Global KV Cache Store with layer‐wise overlapped transmission to decouple routing decisions from cache placement. Together, these mechanisms eliminate cache‐induced hotspots and allow routers to perform purely load‐aware scheduling with minimal latency overhead. BanaServe is implemented on top of state‐of‐the‐art LLM serving frameworks, including vLLM and DistServe. Results We evaluate BanaServe under diverse and challenging workloads, including long‐context inference, bursty request arrivals, and mixed prompt–generation patterns. Experimental results show that, compared to vLLM, BanaServe improves throughput by 1.2–3.9× and reduces total processing time by 3.9%–78.4%. In comparison with DistServe, BanaServe achieves 1.1–2.8× higher throughput while reducing latency by 1.4%–70.1%. These gains are consistent across workload variations, demonstrating BanaServe's robustness under highly dynamic serving conditions. Conclusion BanaServe demonstrates that dynamic, multi‐granularity resource rebalancing and cache‐decoupled routing are essential for efficient disaggregated LLM serving. By jointly addressing resource elasticity, stage imbalance, and cache‐induced load skew, BanaServe substantially improves throughput, latency, and resource utilization in real‐world deployments. This work provides a practical and scalable foundation for next‐generation LLM serving systems operating under dynamic and heterogeneous workloads. Yiyuan He, Minxian Xu, Jingfeng Wu, Jianmin Hu, Chong Ma 0005, Cheng-Zhong Xu 0001, Lin Qu, Kejiang Ye |
Softw. Pract. Exp. | 1 |
| 2025 | Cloudnativesim: A Toolkit for Modeling and Simulation of Cloud-Native ApplicationsabstractABSTRACT Background Cloud‐native applications are increasingly becoming popular in modern software design. Employing a microservice‐based architecture into these applications is a prevalent strategy that enhances system availability and flexibility. However, cloud‐native applications introduce new challenges, including frequent inter‐service communication and the management of heterogeneous codebases and hardware, resulting in unpredictable complexity and dynamism. Furthermore, as applications scale, only limited research teams or enterprises possess the resources for large‐scale deployment and testing, which impedes progress in the cloud‐native domain. Aims To address these challenges, we propose CloudNativeSim, a simulator for cloud‐native applications with a microservice‐based architecture. Results CloudNativeSim offers several key benefits: (i) comprehensive and dynamic modeling for cloud‐native applications, (ii) an extended simulation framework with new policy interfaces for scheduling cloud‐native applications, and (iii) support for customized application scenarios and user feedback based on Quality of Service (QoS) metrics. Conclusion CloudNativeSim can be easily deployed on standard computers to manage a high volume of requests and services. Its performance was validated through a case study, demonstrating higher than 94.5% accuracy in terms of response time simulation. The study further highlights the feasibility of CloudNativeSim by illustrating the effects of various scaling policies. Jingfeng Wu, Minxian Xu, Yiyuan He, Kejiang Ye, Cheng-Zhong Xu 0001 |
Softw. Pract. Exp. | 3 |
| 2024 | Resource Management for GPT-Based Model Deployed on Clouds: Challenges, Solutions, and Future Directions
Yongkang Dang, Yiyuan He, Minxian Xu, Kejiang Ye |
ICA3PP (2) | 2 |
| 2024 | UELLM: A Unified and Efficient Approach for Large Language Model Inference Serving
Yiyuan He, Minxian Xu, Jingfeng Wu, Wanyi Zheng, Kejiang Ye, Cheng-Zhong Xu 0001 |
ICSOC (1) | 1 |