EDBT 2026 Demo / reviewers in the wild / expert
Qiangyu Pei
dblp:314/7066
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0001-8870-4309ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QRoute: A QoE-aware route planning system for enhancing in-vehicle video streaming
Jiahai Hu, Qiangyu Pei, Fangming Liu |
Comput. Networks | 4 |
| 2026 | Sonnet: A Workflow-Aware Serverless Platform for Time-Sensitive Edge Computing With WebAssemblyabstractThe serverless computing paradigm has emerged as a promising solution to address the resource underutilization and inflexible service scaling in edge environments by decoupling the monolithic application into a serverless workflow. However, existing serverless platforms are primarily designed for cloud centers, relying on heavyweight isolation mechanisms that are illsuited for resource-constrained edge computing. These limitations result in high latency, low deployment density, and restricted parallelism. In this paper, we proposeSonnet, a serverless platform tailored for edge computing, capable of rapidly responding to user requests and supporting efficient and elastic service scaling. Sonnet offers these features by (i) employing lightweight WebAssembly as the execution environment for functions, (ii) leveraging serverless workflow information to optimize function deployment on resource-constrained edge environments, and (iii) designing a function deployment algorithm that achieves dynamic load balancing within the cluster. An extensive evaluation ofSonnetwith real-world serverless workflows demonstrates its effectiveness and practical applicability. Compared with SOTA and commonly used edge computing serverless solutions, our experiments show that Sonnet can reduce end-to-end latency by 27% and improve throughput by 2.83×. Quanfeng Deng, Jing Wu 0024, Qiangyu Pei, Chuangxun Lin, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 3 |
| 2026 | Cooling as You Wish: Component-Level Cooling for Heterogeneous Edge DatacentersabstractAs computing shifts toward the edge, edge datacenters are becoming essential for supporting diverse real-time applications. Unlike traditional cloud datacenters, edge datacenters face unique cooling challenges due to their requirements forproximity to end users, high density, and hardware heterogeneity. While warm water cooling is a promising technique for this infrastructure, current one-size-fits-all cooling strategies significantly compromise efficiency due to severe inter- and intra-component hotspots. In this work, we present CoolEdge+, a cost-effective component–level water cooling system for enhancing the cooling efficiency of edge datacenters. Specifically, CoolEdge+dynamically adjusts the inlet water temperature for each component through a carefully designed water circulation architecture to mitigate inter-component hotspots. To address intra-component hotspots, it employs vapor chamber–based cold plates that rapidly dissipate heat without manual intervention or additional energy consumption. We further design a fine-grained cooling control framework that leverages a well-managed power capping approach to decide on customized inlet water temperatures and hardware power limits. Based on a hardware prototype and a real-world trace from Alibaba PAI, evaluation results show that CoolEdge+reduces cooling energy consumption by up to 27.19% compared to existing coarse-grained systems, while maintaining performance guarantees. Compared to the state-of-the-art CoolEdge, CoolEdge+saves 35.24% more cooling costs with comparable energy consumption and no latency violations. Fangming Liu, Qiangyu Pei, Yongjie Yuan, Qixia Zhang, Ziyang Jia, Fei Xu 0009, Bingheng Yan |
IEEE Trans. Computers | 2 |
| 2025 | A comparative measurement study of cross-layer 5G performance under different mobility scenarios
Jiahai Hu, Qiangyu Pei, Fangming Liu |
Comput. Networks | 4 |
| 2025 | Working Smarter Not Harder: Hybrid Cooling for Deep Learning in Edge DatacentersabstractThe proliferation of deep-learning-based mobile and IoT applications has driven the increasing deployment of edge datacenters equipped with domain-specific accelerators. The unprecedented computing power offered by these accelerators puts a heavy burden on the cooling system, motivating more potent cooling techniques like cold water cooling. However, we observe that cold water cooling results in significant energy waste in edge datacenters due to the fluctuating resource utilization both spatially and temporally. To tackle this issue, we propose the concept of “working smarter” by slowing down accelerators deliberately whenever possible and enabling warm water cooling during these times to achieve cooling efficiency. Based on this concept, we develop Hyco—a hybrid water cooling system tailored for edge datacenters running deep learning workloads. First, Hyco features a zone-based cooling architecture enabling dynamic switching between cold water and warm water cooling. Then, based on a lightweight latency estimation method, Hyco incorporates a learning-based scheduling scheme to determine “which” accelerator workers and “when” to slow down through an adaptive and intelligent power-latency trade-off for deep learning models. The simulation with real-world traces shows that Hyco reduces the cooling energy consumption by up to 34.74× while satisfying latency constraints more than 99% of the time for deep-learning-based applications. Qiangyu Pei, Yongjie Yuan, Haichuan Hu, Lin Wang 0015, Bingheng Yan, Chen Yu 0003, Fangming Liu |
IEEE Trans. Sustain. Comput. | 1 |
| 2024 | InferCool: Enhancing AI Inference Cooling through Transparent, Non-Intrusive Task ReassignmentabstractThe increasing power consumption of AI inference in modern datacenters has escalated cooling demands significantly, necessitating the adoption of potent cooling approaches like water cooling. Unlike traditional cloud workloads, AI inference has unique characteristics that create substantial gaps in achieving optimal cooling efficiency. In this work, we present the first comprehensive measurement study of AI inference cooling across various models within an industrial-ready scheduling framework, highlighting significant inefficiencies and their causes. To fill the gap while following the fundamental requirements of cooling systems, we explore a new opportunity presented by modern Multi-Instance GPU-enabled inference serving, where the scheduling dimension is naturally orthogonal to the cooling dimension. Building on this insight, we develop InferCool, a cooling middleware designed to enhance cooling efficiency for inference serving through transparent, non-intrusive task reassignment. It includes a streamlined power and temperature prediction approach and a thermal-aware, adaptive application deployment and request scheduling mechanism. Real-world experiments on a water-cooled testbed and a three-node cluster demonstrate that InferCool can reduce the maximum GPU temperature by 5°C across eight A100 GPUs, equivalent to cooling energy savings of about 20%. Importantly, InferCool requires no modifications to existing cooling infrastructures and is compatible with existing scheduling systems. Qiangyu Pei, Lin Wang 0015, Bingheng Yan, Chen Yu 0003, Fangming Liu |
SoCC | 1 |
| 2024 | λGrapher: A Resource-Efficient Serverless System for GNN Serving through Graph SharingabstractGraph Neural Networks (GNNs) have been increasingly adopted for graph analysis in web applications such as social networks. Yet, efficient GNN serving remains a critical challenge due to high workload fluctuations and intricate GNN operations. Serverless computing, thanks to its flexibility and agility, offers on-demand serving of GNN inference requests. Alas, the request-centric serverless model is still too coarse-grained to avoid resource waste. Haichuan Hu, Fangming Liu, Qiangyu Pei, Yongjie Yuan, Zichen Xu 0001, Lin Wang 0015 |
WWW | 3 |
| 2023 | AsyFunc: A High-Performance and Resource-Efficient Serverless Inference System via Asymmetric FunctionsabstractRecent advances in deep learning (DL) have spawned various intelligent cloud services with well-trained DL models. Nevertheless, it is nontrivial to maintain the desired end-to-end latency under bursty workloads, raising critical challenges on high-performance while resource-efficient inference services. To handle burstiness, some inference services have migrated to the serverless paradigm for its rapid elasticity. However, they neglect the impact of the time-consuming and resource-hungry model-loading process when scaling out function instances, leading to considerable resource inefficiency for maintaining high performance under burstiness. Qiangyu Pei, Yongjie Yuan, Haichuan Hu, Fangming Liu |
SoCC | 1 |
| 2022 | CoolEdge: hotspot-relievable warm water cooling for energy-efficient edge datacentersabstractAs the computing frontier drifts to the edge, edge datacenters play a crucial role in supporting various real-time applications. Different from cloud datacenters, the requirements of proximity to end-users, high density, and heterogeneity, present new challenges to cool the edge datacenters efficiently. Although warm water cooling has become a promising cooling technique for this infrastructure, the one-size-fits-all cooling control would lower the cooling efficiency considerably because of the severe thermal imbalance across servers, hardware, and even inside one hardware component in an edge datacenter. In this work, we propose CoolEdge, a hotspot-relievable warm water cooling system for improving the cooling efficiency and saving costs of edge datacenters. Specifically, through the elaborate design of water circulations, CoolEdge can dynamically adjust the water temperature and flow rate for each heterogeneous hardware component to eliminate the hardware-level hotspots. By redesigning cold plates, CoolEdge can quickly disperse the chip-level hotspots without manual intervention. We further quantify the power saving achieved by the warm water cooling theoretically, and propose a custom-designed cooling solution to decide an appropriate water temperature and flow rate periodically. Based on a hardware prototype and real-world traces from SURFsara, the evaluation results show that CoolEdge reduces the cooling energy by 81.81% and 71.92%, respectively, compared with conventional and state-of-the-art water cooling systems. Qiangyu Pei, Qixia Zhang, Fangming Liu, Ziyang Jia, Yishuo Wang, Yongjie Yuan |
ASPLOS | 1 |
| 2022 | HiTDL: High-Throughput Deep Learning Inference at the Hybrid Mobile EdgeabstractDeep neural networks (DNNs) have become a critical component for inference in modern mobile applications, but the efficient provisioning of DNNs is non-trivial. Existing mobile- and server-based approaches compromise either the inference accuracy or latency. Instead, a hybrid approach can reap the benefits of the two by splitting the DNN at an appropriate layer and running the two parts separately on the mobile and the server respectively. Nevertheless, the DNN throughput in the hybrid approach has not been carefully examined, which is particularly important for edge servers where limited compute resources are shared among multiple DNNs. This article presents HiTDL, a runtime framework for managing multiple DNNs provisioned following the hybrid approach at the edge. HiTDL's mission is to improve edge resource efficiency by optimizing the combined throughput of all co-located DNNs, while still guaranteeing their SLAs. To this end, HiTDL first builds comprehensive performance models for DNN inference latency and throughout with respect to multiple factors including resource availability, DNN partition plan, and cross-DNN interference. HiTDL then uses these models to generate a set of candidate partition plans with SLA guarantees for each DNN. Finally, HiTDL makes global throughput-optimal resource allocation decisions by selecting partition plans from the candidate set for each DNN via solving a fairness-aware multiple-choice knapsack problem. Experimental results based on a prototype implementation show that HiTDL improves the overall throughput of the edge by$4.3\times$compared with the state-of-the-art. Jing Wu 0024, Lin Wang 0015, Qiangyu Pei, Xingqi Cui, Fangming Liu, Tingting Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |