Qianlong Sang

dblp:356/4559 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0005-1563-9434ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 since 2021Computer networks · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Rethinking Serverless Keep-Alive by Decoupling Eviction Priority From Execution State
abstract
Keeping runtime alive is critical for mitigating cold start issues in Function-as-a-Service (FaaS) platforms. State-of-the-Art (SOTA) keep-alive policies often draw an analogy to data caching, adapting classic replacement algorithms to manage runtime pools. Our analysis reveals that this analogy is fundamentally flawed: 1) they violate their own eviction priorities due to conflicts with running containers, and 2) they ignore the exorbitant memory-time resource cost of runtime replacement. To address these challenges, we present FaaShadow, a lightweight keep-alive policy that recasts the problem from simple caching to cost-aware resource allocation. FaaShadow introduces the concept of a shadow pool, a per-function data structure that enables online estimation of the marginal utility of memory adjustments. By quantifying both the potential performance gain from allocating new containers and the performance loss from removing existing ones, FaaShadow makes data-driven reallocation decisions that maximize the global warm start rate. Experimental results show that FaaShadow achieves a 95% warm start rate using only 60% of the memory required by the best-in-class baseline. When paired with our dynamic scaling mechanism, FaaShadow reduces average memory consumption by a staggering 81.71% compared to the SOTA predictive scaler, while upholding performance targets.
Yili Gong, Xinquan Cai, Qianlong Sang, Tianheng Lu, Chuang Hu, Dazhao Cheng
IEEE Trans. Computers3
2026 Trident: Identifying, Constraining and Multi-Domain Governing for Resource Management on Mobile Devices
abstract
Mobile applications such as browsers, video, and other interactive software are tightly coupled with frame rendering, which is critical for user experience. Frame rendering requires the collaboration of CPU, GPU, and memory to ensure energy efficiency and maintain Quality of Experience (QoE). However, there are notable deficiencies in resource management for these components. Our observations reveal three critical issues: 1) the operating system fails to accurately identify rendering-related threads, 2) thread groups lack strict resource constraints, leading to insufficient resources for rendering-related threads, and 3) frequency scaling across components is not coordinated, resulting in performance degradation and power inefficiencies. To address these issues, we propose Trident, a holistic resource management framework. Trident includes a cross-layer thread tracer to identify rendering-related threads, a reinforcement learning-based governor to coordinate the frequency of multiple hardware components, and a gain scheduling-based share controller to constrain resources among thread groups dynamically. Our framework aims to minimize power consumption while maintaining QoE. We implement Trident as a system service on five distinct smartphones, from older models to recent flagships, and evaluate its effectiveness on popular applications under various workloads. The results demonstrate that Trident reduces power consumption by up to 16.8% compared to three state-of-the-art techniques while ensuring QoE on mobile platforms. Additionally, the overhead introduced by Trident is minimal, making it an efficient solution for real-world deployment.
Qianlong Sang, Chuang Hu, Yili Gong, Dazhao Cheng
IEEE Trans. Mob. Comput.1
2024 Incendio: Priority-Based Scheduling for Alleviating Cold Start in Serverless Computing
abstract
In serverless computing, cold start results in long response latency. Existing approaches strive to alleviate the issue by reducing the number of cold starts. However, our measurement based on real-world production traces shows that the minimum number of cold starts does not equate to the minimum response latency, and solely focusing on optimizing the number of cold starts will lead to sub-optimal performance. The root cause is that functions have different priorities in terms of latency benefits by transferring a cold start to a warm start. In this paper, we proposeIncendio, a serverless computing framework exploiting priority-based scheduling to minimize the overall response latency from the perspective of cloud providers. We reveal the priority of a function is correlated to multiple factors and design a priority model based on Spearman’s rank correlation coefficient. We integrate a hybrid Prophet-LightGBM prediction model to dynamically manage runtime pools, which enables the system to prewarm containers in advance and terminate containers at the appropriate time. Furthermore, to satisfy the low-cost and high-accuracy requirements in serverless computing, we propose a Clustered Reinforcement Learning-based function scheduling strategy. The evaluations show that Incendio speeds up the native system by 1.4×, and achieves 23% and 14.8% latency reductions compared to two state-of-the-art approaches.
Xinquan Cai, Qianlong Sang, Chuang Hu, Yili Gong, Kun Suo, Xiaobo Zhou 0002, Dazhao Cheng
IEEE Trans. Computers2
2024 Corrections to "DNN Surgery: Accelerating DNN Inference on the Edge through Layer Partitioning"
abstract
In this paper, we reference the previous conference version and complete the grant number mentioned in the acknowledgments of the conference version.
Huanghuang Liang, Qianlong Sang, Chuang Hu, Dazhao Cheng, Xiaobo Zhou 0002, Dan Wang 0002, Wei Bao 0001, Yu Wang 0003
IEEE Trans. Cloud Comput.2
2024 QoS-Aware Power Management via Scheduling and Governing Co-Optimization on Mobile Devices
abstract
Scheduling and governing are two key technologies to trade off the Quality of Service (QoS) against the power consumption on mobile devices with heterogeneous cores. However, there are still defects in the use of them, among which two of the decoupling issues are critical and need to be resolved. First, both the scheduling and governing decouple from QoS, one of the most important metrics of user experience on mobile platforms. Second, scheduling and governing also decouple from each other in mobile systems and they might weaken each other when being effective at the same time. To address the above issues, we propose Orthrus, a comprehensive QoS-aware power management approach that involves a governing approach based on deep reinforcement learning to adjust the frequency of heterogeneous cores, a scheduling algorithm based on finite state machine that assigns cores to QoS-related threads, and expert fuzzy control-based coordination mechanism between the two to manage the impact between scheduling and governing. Our proposed approach aims to minimize power consumption while guaranteeing the QoS. We implement Orthrus on Google Pixel 3 as the system service of Android and evaluate it using several widespread mobile applications. The performance evaluation demonstrates that Orthrus reduces the average power consumption by up to 35.7% compared to three state-of-the-art techniques while ensuring the QoS on mobile platforms.
Qianlong Sang, Jinqi Yan, Chuang Hu, Kun Suo, Dazhao Cheng
IEEE Trans. Mob. Comput.1
2024 Redundancy-Free and Load-Balanced TGNN Training With Hierarchical Pipeline Parallelism
abstract
Recently, Temporal Graph Neural Networks (TGNNs), as an extension of Graph Neural Networks, have demonstrated remarkable effectiveness in handling dynamic graph data. Distributed TGNN training requires efficiently tackling temporal dependency, which often leads to excessive cross-device communication that generates significant redundant data. However, existing systems are unable to remove the redundancy in data reuse and transfer, and suffer from severe communication overhead in a distributed setting. This work introduces Sven, a co-designed algorithm-system library aimed at accelerating TGNN training on a multi-GPU platform. Exploiting dependency patterns of TGNN models, we develop a redundancy-free graph organization to mitigate redundant data transfer. Additionally, we investigate communication imbalance issues among devices and formulate the graph partitioning problem as minimizing the maximum communication balance cost, which is proved to be an NP-hard problem. We propose an approximation algorithm called Re-FlexBiCut to tackle this problem. Furthermore, we incorporate prefetching, adaptive micro-batch pipelining, and asynchronous pipelining to present a hierarchical pipelining mechanism that mitigates the communication overhead. Sven represents the first comprehensive optimization solution for scaling memory-based TGNN training. Through extensive experiments conducted on a 64-GPU cluster, Sven demonstrates impressive speedup, ranging from 1.9x to 3.5x, compared to State-of-the-Art approaches. Additionally, Sven achieves up to 5.26x higher communication efficiency and reduces communication imbalance by up to 59.2%.
Yaqi Xia, Zheng Zhang 0036, Donglin Yang, Chuang Hu, Xiaobo Zhou 0002, Hongyang Chen 0001, Qianlong Sang, Dazhao Cheng
IEEE Trans. Parallel Distributed Syst.7
2023 TAPU: A Transmission-Analytics Processing Unit for Accelerating Multifunctions in IoT Gateways
abstract
Internet of Things (IoT) gateways integrate various sensors and compute initial decisions before transmitting data to the cloud for further processing. As the functions they need to support become increasingly complex, gateways must upgrade their hardware. Network functions (NF) and video analytics (VAs) are two typical examples of hardware requirements: NFs need specialized hardware accelerators, while VAs need parallel processing power. However, gateways are typically constrained by factors, such as power, size, and cost, leading to a need to multiplex functions and minimize hardware overprovisioning. This article proposes a novel accelerator, the transmission-analytic processing unit (TAPU), which uses multi-image FPGA to accelerate VAs and NFs for IoT gateways. We preconfigure one image for VAs and one image for NFs, then multiplex the FPGA resources in the time dimension. The TAPU system design requires both hardware and software revisions. In the hardware design, we discuss our considerations on hardware choice and present a new abstraction of hardware functions to overcome the challenge of application development on different multi-image FPGAs. For the software, we develop a fully functional TAPU system to adapt to dynamic network and VAs workloads. Our evaluation shows that TAPU utilization can reach 92%, considerably increasing VAs and network processing throughput over the current approach. We further evaluate TAPU through two case studies that support a campus traffic monitoring system and an office surveillance system, demonstrating excellent performance improvement and low overhead.
Huanghuang Liang, Qianlong Sang, Chuang Hu, Yili Gong, Dazhao Cheng, Xiaobo Zhou 0002, Yu Wang 0003
IEEE Internet Things J.2
2023 An Edge-Side Real-Time Video Analytics System With Dual Computing Resource Control
abstract
Video analytics systems conduct video preprocessing to filter out unnecessary frames and model inference using appropriately selected neural networks for high analytics speed. Video preprocessing is instruction-intensive computing (IIC) executed by CPU, and model inference is data-intensive computing (DIC) executed by GPU. In this paper, we show the analytics accuracy of existing systems can largely vary in fields, caused by thedynamicIIC and DIC workloads of differentcontentsin applications. Unfortunately, cameras havefixedCPU/GPU resources and cannot effectively adapt to workload dynamics. We develop Gemini, a new edge-side real-time video analytics system enhanced by a dual-image FPGA. We take the advantage of negligible image switching time of dual-image FPGAs, pre-configure one CPU image and one GPU image and elastically multiplex the dual CPU-GPU resources intimedimension. Gemini requires both hardware and software revisions. In hardware, we overcome challenges of hardware-dependent application development, low communication efficiency between the microprocessor and FPGA, and high programming complexity by hardware abstraction, asynchronous data transfer mechanism and stub-skeleton middleware. In software, we overcome the challenge of adapting to the dynamic workloads by a bandit learning approach. We implement Gemini and show that Gemini can improve the analytics accuracy to 90.35%.
Chuang Hu, Qianlong Sang, Huanghuang Liang, Dan Wang 0002, Dazhao Cheng, Jin Zhang 0001, Qing Li 0006, Junkun Peng
IEEE Trans. Computers3
2023 DNN Surgery: Accelerating DNN Inference on the Edge Through Layer Partitioning
abstract
Recent advances in deep neural networks have substantially improved the accuracy and speed of various intelligent applications. Nevertheless, one obstacle is that DNN inference imposes a heavy computation burden on end devices, but offloading inference tasks to the cloud causes a large volume of data transmission. Motivated by the fact that the data size of some intermediate DNN layers is significantly smaller than that of raw input data, we designed the DNN surgery, which allows partitioned DNN to be processed at both the edge and cloud while limiting the data transmission. The challenge is twofold: (1) Network dynamics substantially influence the performance of DNN partition, and (2) State-of-the-art DNNs are characterized by a directed acyclic graph rather than a chain, so that partition is incredibly complicated. To solve the issues, We design a Dynamic Adaptive DNN Surgery(DADS) scheme, which optimally partitions the DNN under different network conditions. We also study the partition problem under the cost-constrained system, where the resource of the cloud for inference is limited. Then, a real-world prototype based on the selif-driving car video dataset is implemented, showing that compared with current approaches, DNN surgery can improve latency up to 6.45 times and improve throughput up to 8.31 times. We further evaluate DNN surgery through two case studies where we use DNN surgery to support an indoor intrusion detection application and a campus traffic monitor application, and DNN surgery shows consistently high throughput and low latency.
Huanghuang Liang, Qianlong Sang, Chuang Hu, Dazhao Cheng, Xiaobo Zhou 0002, Dan Wang 0002, Wei Bao 0001, Yu Wang 0003
IEEE Trans. Cloud Comput.2