Hongyang Chen 0002

dblp:13/3715-2 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0002-9419-3768ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 5 since 2021Computer networks · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 TraceWizard: End-to-End Distributed Tracing Across Host and Network Devices in Cloud
abstract
The rise of microservice architecture in cloud computing has introduced additional complexities in diagnosing faults, as traditional end-to-end tracing systems often fail to address issues beyond the application layer, such as network devices. To overcome this limitation, we introduce TraceWizard, an enhanced end-to-end tracing system that integrates eBPF and SDN (Software Defined Network) technologies to track requests across host and network devices. By enabling full life-cycle tracing and maintaining consistent trace contexts, TraceWizard provides fine-grained insights into faults across applications, the OS kernel, and network devices. Our evaluations show that its data enables more effective fault detection, achieving an average accuracy of 91.6 % across different algorithms—significantly outperforming application-layer monitoring tools. Additionally, it helps operators identify root causes with minimal overhead, reducing QPS by 2.2%, increasing QCT by 2.2%, and adding 3.41 % CPU and 2.43% memory usage.
Kuangyuan Li, Jingrun Zhang, Pengfei Chen 0002, Hongyang Chen 0002, Ruipeng Hong, Wanqi Yang, Chen Sun 0005
CLOUD4
2025 MoTor: Resource-efficient cloud-native network acceleration with programmable switches
Hongyang Chen 0002, Pengfei Chen 0002, Zibin Zheng, Kaibin Fang
Comput. Networks1
2025 NetScope: Fault Localization in Programmable Networking Systems With Low-Cost In-Band Network Telemetry and In-Network Detection
abstract
Recently, Software Defined Networking (SDN) has gained widespread adoption as a network infrastructure. Although the openness and programmability of SDN facilitate large complex network construction, diagnosing faults in datacenter-scale network remains challenging. Previous network diagnosis tools pose significant overhead in fine-grained telemetry and typically lack automated fine-grained fault diagnosis capabilities. Although on-demand monitoring methods have been proposed to reduce telemetry overhead, they struggle with effectively setting fixed thresholds, which requires expert experience. This paper presents NetScope, a lightweight system for real-time anomaly detection with self-adaptive thresholds and automatic root cause localization in programmable networking systems. NetScope estimates latency medians for each Flow (i.e., a pair of source and sink switches) within the switch using the proposed per-Flow quantile sketch and calculates the threshold accordingly for anomaly detection. Upon detecting anomalies, NetScope collects aggregated packet-level telemetry on demand and generates a ranked list of fine-grained fault culprits at multiple levels, including port-level, Flow-level, and switch-level. Extensive experiments demonstrate the effectiveness and efficiency of NetScope in anomaly detection and fault localization. Specifically, NetScope achieves a 32%~116% relative improvement in anomaly detection and 6%~197% improvement in root cause analysis compared with other baselines without causing any network bandwidth in anomaly detection while consuming 64.2% less telemetry bandwidth for localization.
Hongyang Chen 0002, Benran Wang, Guangba Yu, Pengfei Chen 0002, Chen Sun 0005, Zibin Zheng
IEEE Trans. Netw.1
2025 Subgraphs as First-Class Citizens in Incident Management for Large-Scale Online Systems: An Evolution-Aware Framework
Pengfei Chen 0002, Yu Luo 0019, Qiuyu Yan, Hongyang Chen 0002, Guangba Yu, Zibin Zheng
IEEE Trans. Software Eng.5
2024 Real-Time Intrusion Detection and Prevention with Neural Network in Kernel Using eBPF
abstract
With the development of public cloud, real-time intrusion detection is becoming necessary. Current methods neither address the overhead of real-time network data capturing, nor effectively balance security level with performance. These issues can be addressed by offloading intrusion detection and prevention to the extended Berkeley Packet Filter (eBPF). However, current eBPF-based methods suffer from shortcomings in model performance or inference overhead. Moreover, they overlook the issues of eBPF in real-time scenarios, such as maximum eBPF instruction limitations. In this paper, we redesign the Neural Network inference mechanism to address the limitations of eBPF. Then, we propose a thread-safe parameter hot-updating mechanism without explicit spin lock. Evaluations indicate that our method achieves model performance comparable to the current best eBPF-based method while reducing memory overhead (5KB) and inference time (3000-5000ns per flow). Our method achieve F1-scores of 0.933 and 0.992 on the offline and online datasets, respectively.
Pengfei Chen 0002, Hongyang Chen 0002
DSN4
2024 Graph neural network based robust anomaly detection at service level in SDN driven microservice system
Hongyang Chen 0002, Pengfei Chen 0002, Benran Wang, Dandan Ma, Zibin Zheng
Comput. Networks1
2024 MicroFI: Non-Intrusive and Prioritized Request-Level Fault Injection for Microservice Applications
abstract
Microservice is a widely-adopted architecture for constructing cloud-native applications. To test application resiliency, chaos engineering is widely used to inject faults proactively in applications. However, the searching space formed by possible injection locations is huge due to the scale and complexity of the application. Although some methods are proposed to effectively explore injection space, they cannot prioritize high-impact injection solutions. Additionally, the blast radius of faults injected by existing methods is typically full of uncertainty, causing faults of multiple application functions. Although some tools are designed to conduct request-level injection, they require instrumentation on application code. To tackle these problems, this paper presents MicroFI, a non-intrusive fault injection framework, aiming to efficiently test different application functions with request-level injection. Request-level injection limits the blast radius to specified requests without any source code modification. Additionally, MicroFI leverages historical injection results and parallel technique to accelerate the searching. Moreover, An enhanced PageRank is used to measure the impact of faults and prioritize high-impact faults that fail more functions. Evaluations on three microservice applications show that MicroFI precisely injects faults and reduces up to 91% redundant faults on average. Additionally, by employing prioritization, MicroFI reduces an average of 47.3% injection budgets to cover all high-impact faults.
Hongyang Chen 0002, Pengfei Chen 0002, Guangba Yu
IEEE Trans. Dependable Secur. Comput.1
2023 MARS: Fault Localization in Programmable Networking Systems with Low-cost In-Band Network Telemetry
abstract
Recently, the adoption of Software Defined Networking (SDN) as a network infrastructure has gained significant popularity. Although the openness and programmability of SDN ease the construction of large complex networks, it is still challenging to diagnose faults in a complex datacenter-scale network, which is crucial to guarantee rigorous service level agreement (SLA) of upper-layer applications. Previous network diagnosis tools incur significant overhead in fine-grained telemetry, and usually lack the ability to automatically diagnose fine-grained faults. Although on-demand monitoring methods is proposed to reduce telemetry overhead, they struggle to effectively set static thresholds, which requires expert experience. In this paper, we present MARS, a lightweight system for anomaly detection with dynamic threshold and automatic root cause localization in programmable networking systems. MARS collects aggregated packet-level telemetry on demand and generates a ranked list of fine-grained fault culprits at multiple levels, including port-level, switch-level, and flow-level. Experimental evaluations show the cost-effectiveness of MARS, both in terms of network bandwidth and switch memory usage. Moreover, MARS achieves a 0.97 F1 score in anomaly detection, and 0.95 Recall at Top-2 and an overall 0.3 Exam Score in root cause localization.
Benran Wang, Hongyang Chen 0002, Pengfei Chen 0002, Guangba Yu
ICPP2
2023 MARS: Fault Localization in Programmable Networking Systems with Low-cost In-Band Network Telemetry
abstract
This paper presents MARS, a lightweight system for anomaly detection with dynamic threshold and automatic root cause localization in programmable networking systems. MARS collects aggregated packet-level telemetry on demand and generates a ranked list of fine-grained fault culprits at port-level, switch-level, and flow-level.
Benran Wang, Hongyang Chen 0002, Pengfei Chen 0002, Guangba Yu
IWQoS2
2023 Nezha: Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability Data
abstract
Root cause analysis (RCA) in large-scale microservice systems is a critical and challenging task. To understand and localize root causes of unexpected faults, modern observability tools collect and preserve multi-modal observability data, including metrics, traces, and logs. Since system faults may manifest as anomalies in different data sources, existing RCA approaches that rely on single-modal data are constrained in the granularity and interpretability of root causes. In this study, we present Nezha, an interpretable and fine-grained RCA approach that pinpoints root causes at the code region and resource type level by incorporative analysis of multi-modal data. Nezha transforms heterogeneous multi-modal data into a homogeneous event representation and extracts event patterns by constructing and mining event graphs. The core idea of Nezha is to compare event patterns in the fault-free phase with those in the fault-suffering phase to localize root causes in an interpretable way. Practical implementation and experimental evaluations on two microservice applications show that Nezha achieves a high top1 accuracy (89.77%) on average at the code region and resource type level and outperforms state-of-the-art approaches by a large margin. Two ablation studies further confirm the contributions of incorporating multi-modal data.
Guangba Yu, Pengfei Chen 0002, Hongyang Chen 0002, Zibin Zheng
ESEC/SIGSOFT FSE4
2022 Going through the Life Cycle of Faults in Clouds: Guidelines on Fault Handling
abstract
Faults are the primary culprits of breaking the high availability of cloud systems, even leading to costly outages. As the scale and complexity of clouds increase, it becomes extraordinarily difficult to understand, detect and diagnose faults. During outages, engineers record the detailed information of the whole life cycle of faults (i.e., fault occurrence, fault detection, fault identification, and fault mitigation) in the form of postmortems. In this paper, we conduct a quantitative and qualitative study on 354 public post-mortems collected in three popular large-scale clouds, 97.7% of which spans from 2015 to 2021. By reviewing and analyzing post-mortems, we go through the life cycle of faults in clouds and obtain 10 major findings. Based on these findings, we further reach a series of actionable guidelines for better fault handling.
Guangba Yu, Pengfei Chen 0002, Hongyang Chen 0002, Zhekang Chen
ISSRE4
2022 Graph based Incident Extraction and Diagnosis in Large-Scale Online Systems
abstract
With the ever increasing scale and complexity of online systems, incidents are gradually becoming commonplace. Without appropriate handling, they can seriously harm the system availability. However, in large-scale online systems, these incidents are usually drowning in a slew of issues (i.e., something abnormal, while not necessarily an incident), rendering them difficult to handle. Typically, these issues will result in a cascading effect across the system, and a proper management of the incidents depends heavily on a thorough analysis of this effect. Therefore, in this paper, we propose a method to automatically analyze the cascading effect of availability issues in online systems and extract the corresponding graph based issue representations incorporating both of the issue symptoms and affected service attributes. With the extracted representations, we train and utilize a graph neural networks based model to perform incident detection. Then, for the detected incident, we leverage the PageRank algorithm with a flexible transition matrix design to locate its root cause. We evaluate our approach using real-world data collected from the WeChat online service system, the largest instant message system in China. The results confirm the effectiveness of our approach. Moreover, our approach is successfully deployed in the company and eases the burden of operators in the face of a flood of issues and related alert signals.
Pengfei Chen 0002, Yu Luo 0019, Qiuyu Yan, Hongyang Chen 0002, Guangba Yu
ASE5
2021 Sieve: Attention-based Sampling of End-to-End Trace Data in Distributed Microservice Systems
abstract
End-to-end tracing plays an important role in understanding and monitoring distributed microservice systems. The trace data are valuable to help find out the anomalous or erroneous behavior of the system. However, the volume of trace data is huge leading to a heavy burden on analyzing and storing them. To reduce the volume of trace data, the sampling technique is widely adopted. However, existing uniform sampling approaches are unable to capture uncommon traces that are more interesting and informative. To tackle this problem, we design and implement Sieve, an online sampler that aims to bias sampling towards uncommon traces by taking advantage of the attention mechanism. The evaluation results on the trace datasets collected from real-world and experimental microservice systems show that Sieve is effective to increase sampling probabilities of the structurally and temporally uncommon traces and reduce the storage space to a large extent by taking a low sampling rate.
Pengfei Chen 0002, Guangba Yu, Hongyang Chen 0002, Zibin Zheng
ICWS4
2021 MicroRank: End-to-End Latency Issue Localization with Extended Spectrum Analysis in Microservice Environments
abstract
With the advantages of flexible scalability and fast delivery, microservice has become a popular software architecture in the modern IT industry. However, the explosion in the number of service instances and complex dependencies make the troubleshooting extremely challenging in microservice environments. To help understand and troubleshoot a microservice system, the end-to-end tracing technology has been widely applied to capture the execution path of each request. Nevertheless, the tracing data are not fully leveraged by cloud and application providers when conducting latency issue localization in the microservice environment.
Guangba Yu, Pengfei Chen 0002, Hongyang Chen 0002, Zijie Guan, Linxiao Jing, Tianjun Weng, Xinmeng Sun
WWW3