Jingrun Zhang

dblp:323/7870 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-1455-7817ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 TraceWizard: End-to-End Distributed Tracing Across Host and Network Devices in Cloud
abstract
The rise of microservice architecture in cloud computing has introduced additional complexities in diagnosing faults, as traditional end-to-end tracing systems often fail to address issues beyond the application layer, such as network devices. To overcome this limitation, we introduce TraceWizard, an enhanced end-to-end tracing system that integrates eBPF and SDN (Software Defined Network) technologies to track requests across host and network devices. By enabling full life-cycle tracing and maintaining consistent trace contexts, TraceWizard provides fine-grained insights into faults across applications, the OS kernel, and network devices. Our evaluations show that its data enables more effective fault detection, achieving an average accuracy of 91.6 % across different algorithms—significantly outperforming application-layer monitoring tools. Additionally, it helps operators identify root causes with minimal overhead, reducing QPS by 2.2%, increasing QCT by 2.2%, and adding 3.41 % CPU and 2.43% memory usage.
Kuangyuan Li, Jingrun Zhang, Pengfei Chen 0002, Hongyang Chen 0002, Ruipeng Hong, Wanqi Yang, Chen Sun 0005
CLOUD2
2023 DeepPower: Deep Reinforcement Learning based Power Management for Latency Critical Applications in Multi-core Systems
abstract
Latency-critical (LC) applications are widely deployed in modern datacenters. Effective power management for LC applications can yield significant cost savings. However, it poses a significant challenge in maintaining the desired Service Level Aggrement (SLA) levels. Prior researches have mainly emphasized predicting the service time of request and utilize heuristic algorithms for CPU frequency adjustment. Unfortunately, the control granularity is limited to the request level and manual feature selection is needed.
Jingrun Zhang, Guangba Yu, Liang Ai, Pengfei Chen 0002
ICPP1
2023 FaaSDeliver: Cost-Efficient and QoS-Aware Function Delivery in Computing Continuum
abstract
Serverless Function-as-a-Service (FaaS) is a rapidly growing computing paradigm in the cloud era. To provide rapid service response and save network bandwidth, traditional cloud-based FaaS platforms have been extended to the edge. However, launching functions in a heterogeneous computing continuum (HCC) that includes the cloud, fog, and the edge brings new challenges: determiningwhere functions should be delivered and how many resources should be allocated.To optimize the cost of running functions in the HCC, we propose an adaptive and efficient function delivery engine, namedFaaSDeliver, which automatically unearths a cost-efficient function delivery policy (FDP) for each function, including the FaaS platform selection and resource allocation. Real system implementation and evaluations in a practical HCC demonstrate thatFaaSDelivercan unearth the most cost-efficient FDPs from among 180,200 FDPs after a few trials.FaaSDeliverreduces the average cost of function execution from 38% to 78% compared to some state-of-the-art approaches.
Guangba Yu, Pengfei Chen 0002, Zibin Zheng, Jingrun Zhang
IEEE Trans. Serv. Comput.4
2022 A Transferable Time Series Forecasting Service Using Deep Transformer Model for Online Systems
abstract
Many real-world online systems expect to forecast the future trend of software quality to better automate operational processes, optimize software resource cost and ensure software reliability. To achieve that, all kinds of time series metrics collected from online software systems are adopted to characterize and monitor the quality of software services. To meet relevant software engineers’ requirements, we focus on time series forecasting and aim to provide an event-driven and self-adaptive forecasting service. In this paper, we present TTSF-transformer, a transferable time series forecasting service using deep transformer model. TTSF-transformer normalizes multiple metric frequencies to ensure the model sharing across multi-source systems, employs a deep transformer model with Bayesian estimation to generate the predictive marginal distribution, and introduces transfer learning and incremental learning into the training process to ensure the performance of long-term prediction. We conduct experiments on real-world time series metrics from two different types of game business in Tencent®. The results show that TTSF-transformer significantly outperforms other state-of-the-art methods and is suitable for wide deployment in large online systems.
Tao Huang 0021, Pengfei Chen 0002, Jingrun Zhang
ASE3