Zhaolin Duan

dblp:346/7080 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0002-0427-4482ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 iRoute: Local Routing Table-based Workflow Management in Serverless Computing
Laiping Zhao, Zhiyuan Su, Wenhao Huang 0005, Kang Chen 0001, Zhaolin Duan, Jingjie Zong, Wenxin Li 0001, Deze Zeng, Wenyu Qu
EuroSys7
2026 µShare: Non-Intrusive Kernel Co-Locating on NVIDIA GPUs
abstract
The hardware scheduler on NVIDIA GPUs is highly inefficient in utilizing micro-architectural hardware resources. It places blocks from the same kernel within the same GPU Streaming Multiprocessor (SM) core, resulting in a stacking colocating problem, where identical blocks are placed within the same SM core, saturating only a subset of intra-SM hardware resources while leaving others underutilized. The primary challenge in addressing this issue is that the NVIDIA hardware is closed-source, preventing us from directly modifying the hardware scheduler. To bridge the semantic gap between the resource demands of kernels and the scheduler, we introduce µ Share, which enables intra-SM scattered colocating of kernels through a non-intrusive half-plus blocksize shaping method. It shapes the blocksize of kernels to a halfplus blocksize (i.e., slightly more than half of the SM's thread capacity), scattering identical blocks of the same kernel across different SMs. It further adopts a time-shifted launching method to reduce intra-SM resource contention. Compared to state-of-the-art systems, µ Share does not require intrusive modifications to hardware or kernel code, yet it can still improve inference throughput by 26.90%-54.09% and increases low-level hardware utilization by 38.53%-61.15%.
Wenhao Huang 0005, Zhaolin Duan, Laiping Zhao, Yuhao Zhang 0006, Yichi Chen 0001, Zhihang Tang, Kang Chen 0001, Deze Zeng, Wenxin Li 0001, Keqiu Li
HPCA2
2026 Symbiotic MLLM Serving: Dynamically Balancing Parallelism Across GPUs and Resources Within GPUs
Yangyu Zhang, Zhaolin Duan, Shuoming Zhang, Shuaijiang Li, Donglin Yu, Yuan Wen, Chunwei Xia, Xiyu Shi, Huimin Cui
ISCA4
2024 FUYAO: DPU-enabled Direct Data Transfer for Serverless Computing
abstract
Serverless computing typically relies on the third-party forwarding method to transmit data between functions. This method couples control flow and data flow together, resulting in significantly slow data transmission speeds. This challenge makes it difficult for the serverless computing paradigm to meet the low-latency requirements of web services.
Laiping Zhao, Zhaolin Duan, Sheng Chen 0015, Yitao Hu, Zhiyuan Su, Wenyu Qu
ASPLOS (3)4
2024 Photia: Cross-Layer Cache Optimization for Function Startup in Serverless Computing
abstract
Serverless computing enhances application development and deployment through elastic resource provisioning. However, the cold start problem, caused by the stateless nature of serverless functions, remains a significant challenge, particularly in high-traffic scenarios. Existing solutions focus on optimizing function instance caches and reducing sandbox initialization times, yet resource inefficiencies in multi-level cache architectures persist. This paper proposes Photia, a cross-layer collaborative cache management system that integrates decision-making between function instance and sandbox caches. Experimental results demonstrate that Photia reduces function startup latency by up to 2x and improves sandbox cache resource utilization by 20%.
Zhaolin Duan, Laiping Zhao
HPCC1
2024 Strong Session Serializability for Serverless Computing
abstract
Existing works explore various consistency levels to address transactional consistency gaps in serverless platforms but have not explored achieving strong session serializability (3SER) for serverless. For applications with 3SER requirements, stronger transactional consistency is unnecessary and can compromise performance. However, the stateless nature of serverless architectures and their fault-tolerance mechanisms pose challenges in designing a 3SER transaction protocol. We propose TFaaS, the first transactional serverless platform that provides a 3SER guarantee for serverless workflows. TFaaS employs caching to manage session states, transaction execution states, and workflow logs. It selectively persists the necessary states based on time fence techniques to ensure consistency and fault-tolerance requirements are met. Our evaluation demonstrates that TFaaS can effectively reduce latency by up to 6.5× and improve throughput by up to 8×.
Zhaolin Duan, Laiping Zhao
HPCC2