Huxing Zhang

dblp:71/10367 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0007-1761-9044ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 50% Performance modeling and evaluation · 25% Cloud and datacenter computing · 25%
Software engineering, system software, and programming languages
2 papers
Services computing and microservices · 56% Debugging and program repair · 28% Operating systems · 17%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › observability › distributed monitoring
distributed tracing
1.722025
ZeroTracer: In-Band eBPF-Based Trace Generator With Zero Instrumentation for Microservice Systems · IEEE Trans. Parallel Distributed Syst. 2025
Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis · ASPLOS (1) 2025
Debugging and program repair
fault localization
0.912025
Famos: Fault Diagnosis for Microservice Systems Through Effective Multi-Modal Data Fusion · ICSE 2025
Services computing and microservices › microservice architecture
microservice failure diagnosis
0.912025
Famos: Fault Diagnosis for Microservice Systems Through Effective Multi-Modal Data Fusion · ICSE 2025
Services computing and microservices › microservice architecture
microservice system
0.912025
ZeroTracer: In-Band eBPF-Based Trace Generator With Zero Instrumentation for Microservice Systems · IEEE Trans. Parallel Distributed Syst. 2025
Performance modeling and evaluation
trace sampling
0.912025
Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis · ASPLOS (1) 2025
Operating systems › extensible operating systems › kernel extensibility
eBPF
0.312025
ZeroTracer: In-Band eBPF-Based Trace Generator With Zero Instrumentation for Microservice Systems · IEEE Trans. Parallel Distributed Syst. 2025
Operating systems › kernel instrumentation
kernel tracing
0.312025
ZeroTracer: In-Band eBPF-Based Trace Generator With Zero Instrumentation for Microservice Systems · IEEE Trans. Parallel Distributed Syst. 2025

Methods — techniques the papers use, named apart from their topics

in-band trace generation · 1.7eBPF · 1.7multi-modal data fusion · 0.9gaussian attention · 0.9feature extraction · 0.9cross-attention · 0.9commonality and variability analysis · 0.9
YearPublicationVenuePosition
2025 Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis
abstract
Distributed traces contain valuable information but are often massive in volume, posing a core challenge in tracing framework design: balancing the tradeoff between preserving essential trace information and reducing trace volume. To address this tradeoff, previous approaches typically used a '1 or 0' sampling strategy: retaining sampled traces while completely discarding unsampled ones. However, based on an empirical study on real-world production traces, we discover that the '1 or 0' strategy actually fails to effectively balance this tradeoff.
Haiyu Huang 0002, Cheng Chen 0056, Kunyi Chen, Pengfei Chen 0002, Guangba Yu, Yilun Wang 0001, Huxing Zhang, Qi Zhou 0001
ASPLOS (1)8
2025 EagerLog: Active Learning Enhanced Retrieval Augmented Generation for Log-based Anomaly Detection
abstract
Logs record essential information about system operations and serve as a critical source for anomaly detection, which has generated growing research interest. Utilizing large language models (LLMs) within a retrieval-augmented generation (RAG) framework for log-based anomaly detection is an effective approach due to its strong generalization capabilities and efficient few-shot performance. However, the effectiveness of this method hinges on the quality of the knowledge source, which can be impacted by noise and changes within the software systems. Facing these problems, in this paper, we propose a novel log-based anomaly detection method named EagerLog, employing active learning to choose the logs for humans to label, thereby adding them to the knowledge source, thus enhancing the knowledge source and maintaining its quality. Our experiments on three open datasets (BGL, Thunderbird, Zookeeper) and one industrial dataset demonstrate that EagerLog can achieve 93.65% F1 score with approximately 10 labeled log sequences, surpassing existing methods by 15.32%.
Chiming Duan, Yong Yang 0011, Guiyang Liu, Jinbu Liu, Huxing Zhang, Qi Zhou 0001, Ying Li 0012, Gang Huang 0001
ICASSP6
2025 Famos: Fault Diagnosis for Microservice Systems Through Effective Multi-Modal Data Fusion
abstract
Accurately diagnosing the fault that causes the failure is crucial for maintaining the reliability of a microservice system after a failure occurs. Mainstream fault diagnosis approaches are data-driven and mainly rely on three modalities of runtime data: traces, logs, and metrics. Diagnosing faults with multiple modalities of data in microservice systems has been a clear trend in recent years because different types of faults and corresponding failures tend to manifest in data of various modalities. Accurately diagnosing faults by fully leveraging multiple modalities of data is confronted with two challenges: 1) how to minimize information loss when extracting features for data of each modality; 2) how to correctly capture and utilize the relationships among data of different modalities. To address these challenges, we propose FAMOS, a Fault diagnosis Approach for MicrOservice Systems through effective multi-modal data fusion. On the one hand, FAMOS employs independent feature extractors to preserve the intrinsic features for each modality. On the other hand, FAMOS introduces a new Gaussian-attention mechanism to accurately correlate data of different modalities and then captures the inter-modality relationship with a crossattention mechanism. We evaluated FAMOS on two datasets constructed by injecting comprehensive and abundant faults into an open-source microservice system and a real-world industrial microservice system. Experimental results demonstrate the FAMOS's effectiveness in fault diagnosis, achieving significant improvements in F1 scores compared to state-of-the-art (SOTA) methods, with an increase of 20.33 %.
Chiming Duan, Yong Yang 0011, Guiyang Liu, Jinbu Liu, Huxing Zhang, Qi Zhou 0001, Ying Li 0012, Gang Huang 0001
ICSE6
2025 ZeroTracer: In-Band eBPF-Based Trace Generator With Zero Instrumentation for Microservice Systems
abstract
Microservice enables agility in modern cloud-native applications but introduces challenges in fault troubleshooting due to its complex service coordination and cooperation. To tackle these challenges, distributed tracing has emerged for end-to-end request tracing and system understanding. However, existing tracing solutions often suffer from code instrumentation, trace loss and inaccuracy. To overcome these limitations, we introduce ZeroTracer, an in-kernel online distributed tracing system equipped with an eBPF-based (extended Berkeley Packet Filter) trace generator. ZeroTracer tailors for tracking HTTP requests due to its popularity in microservice systems. In our evaluations, ZeroTracer achieves remarkable trace accuracy (i.e., over 91%) and maintains stable performance under different workload concurrency. Moreover, ZeroTracer outperforms other non-invasive approaches which fail to reconcile accurate request causality. Notably, ZeroTracer effectively tracks end-to-end requests in multi-threaded microservice applications, which is absent in existing invasive distributed tracing systems with third-party library instrumentation. Moreover, ZeroTracer introduces a negligible overhead, with latency increasing by only 0.5%–1.2% and a modest 3%–5.8% increase in CPU and memory consumption.
Wanqi Yang, Pengfei Chen 0002, Huxing Zhang
IEEE Trans. Parallel Distributed Syst.4
2024 Network shortcut in data plane of service mesh with eBPF
Wanqi Yang, Pengfei Chen 0002, Guangba Yu, Huxing Zhang
J. Netw. Comput. Appl.5
2013 A decentralized approach for mining event correlations in distributed system monitoring
Gang Wu 0008, Huxing Zhang, Meikang Qiu, Zhong Ming 0001, Xiao Qin 0001
J. Parallel Distributed Comput.2
2012 CAR: Securing PCM Main Memory System with Cache Address Remapping
abstract
Phase Change Memory (PCM) has emerged as a promising alternative of DRAM to provide energy-efficient and high-capacity memory for high performance servers. A new DRAM + PCM hybrid memory architecture has been proposed to leverage PCM's high density and DRAM's robustness and performance. One of the big challenges of PCM is its limited write endurance (107~ 108times per cell). By knowing the association between DRAM and PCM, malicious software can easily force DRAM cache to be flushed continuously, which produces writes to certain PCM cells repeatedly (known as selective attack) and wears out PCM. Although existing wear-leveling approaches could evenly distribute writes under selective attack, the overall endurance of PCM is still severely impacted, and therefore it is suboptimal. In this paper, we propose Cache Address Remapping (CAR), that can adaptively remap DRAM cache address, to hide the association between DRAM and PCM. Moreover, CAR can minimize the write-back traffic to PCM under selective attack by uniformly distributing the writes to a single cache set into different cache sets. We propose a practical and low overhead implementation of CAR, called RanCAR. Experimental results show that CAR could reduce DRAM cache miss rate by ~4600x under selective attack, and prolong PCM lifetime from several minutes to 13.8 years on average.
Gang Wu 0008, Huxing Zhang, Yaozu Dong, Jingtong Hu
ICPADS2
2011 Improving PCM Endurance with Randomized Address Remapping in Hybrid Memory System
abstract
Phase-Change-Memory (PCM) has emerged as a promising alternative of DRAM main memory. A new hybrid memory architecture, where DRAM serves as cache of PCM main memory, has been proposed to leverage PCM's high scalability and DRAM's fast access time. One biggest issue of PCM is the limited number of writes to storage cells. We argue that good cache mechanism will decrease PCM writes dramatically in hybrid memory system. In this paper, we demonstrate that traditional set associative cache is susceptible to malicious attacks, which lead certain PCM cells to wear-out by constant cache flushes. A novel approach called Randomized Address Remapping (RAR) is proposed to hide the mapping details between DRAM and PCM. With this approach, the attacks based on set associative cache do not work, while the efficiency of caching still remains. We present Static Randomized Address Remapping (SRAR) and Dynamic Randomized Address Remapping (DRAR) in this paper. SRAR invalidates set associative cache based attacks by distributing their address accesses to different sets. DRAR uses a region-based approach to change the mapping dynamically, in case that the static mapping relationship is discovered by attacker compromising operating system. Experimental results show that RAR approaches can prevent malicious attacks and improve PCM endurance greatly.
Gang Wu 0008, Huxing Zhang, Yaozu Dong
CLUSTER3