EDBT 2026 Demo / reviewers in the wild / expert
Haifeng Sun 0004
dblp:00/11044-4
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0001-9358-8808ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 5 first-author · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic MeasurementabstractLarge language model (LLM) training is prone to anomalies due to its long duration and large scale, which can lead to significant performance degradation or even training crashes. Due to the synchronization nature of LLM training, anomalies exhibit the cascading effect, making their diagnosis challenging. Existing approaches rely on collecting communication operator information via code instrumentation, which yields only coarse-grained monitoring data and requires modifications to training code or communication libraries. We propose Pulse, a fine-grained, non-intrusive, and easy-to-deploy monitoring system. Our key idea is to enable fine-grained monitoring via traffic measurement. Pulse conducts microsecond-level RDMA traffic measurement on NICs, and transforms flow-level measurements into communication operator measurements, thereby enabling fine-grained and non-intrusive monitoring. We deploy Pulse on a testbed with 64 H200 GPUs and evaluate its anomaly localization capability under common failure scenarios. Pulse achieves machine-level localization in 10 out of 12 scenarios, while existing methods succeed in only 4 and even misdiagnose 2 of the remaining scenarios. Additionally, Pulse achieves over 90% precision and 100% recall, supports up to 2000 concurrent RDMA flow measurements per NIC, and imposes negligible overhead on training performance, making it a practical solution for real-world LLM training environments. Yibo Xiao, Haifeng Sun 0004, Qingkai Meng 0001, Jiong Duan, Xiaohe Hu, Rong Gu 0001, Guihai Chen, Chen Tian 0001 |
ASPLOS (2) | 3 |
| 2026 | NutCracker: A Compilation Framework for Hybrid DPU ArchitecturesabstractSoC-based SmartNICs, or data processing units (DPUs), are becoming a viable option for offloading infrastructure services. However, developers need deep hardware-level knowledge to fully utilize the available hardware accelerators on a target DPU. Hardware heterogeneity also makes porting across DPUs a formidable task. In this work, we propose a new compiler framework, NutCracker. Using NutCracker, programmers develop DPU applications using high-level target-independent languages. NutCracker applies a two-stage compilation process. It first performs progressive lowering to convert the source program to candidate intermediate representations (IRs) of the target DPU. Next, NutCracker applies cost-guided mapping optimization using equality saturation to select a final implementation on the target hardware with a configurable optimization goal. Evaluated on eight applications, NutCracker reduces developer effort by nearly 90% while delivering performance within 3% of handcrafted implementations for seven of the workloads. Moreover, its compilation time is comparable to standard toolchains such as GCC. Haifeng Sun 0004, Antoine Kaufmann, Jialin Li 0001 |
EuroSys | 2 |
| 2026 | SketchPlan: Full-Visibility Sketch-Based Telemetry with Limited Programmable Switch Coverage
Jinbo Sun, Haifeng Sun 0004, Jintao He, Qun Huang 0001, Sa Wang, Yungang Bao |
IWQoS | 2 |
| 2026 | Theseus: Runtime-Adaptive GPU Collective Communication with Hot-Swappable SchedulesabstractCurrent GPU Collective Communication Libraries (CCLs) employ predefined schedules optimized for stable environments. Their supported schedules and selection logic are fixed at communicator initialization, which fails to account for evolving runtime conditions, such as workload characteristics and hardware health status. Consequently, long-running GPU jobs experience suboptimal performance after hours or days of execution, which translates into longer job completion times and wasted GPU cluster resources. To address this problem, we present Theseus, a novel CCL backend that provides schedule-level runtime adaptivity. It admits user-defined schedules and selection policies. As runtime conditions change, Theseus selects suitable schedules using cluster-wide runtime attributes beyond CCL-internal metrics. Moreover, it hot-swaps from the previous schedule consistently across GPUs with low overhead. Theseus acts as a drop-in replacement to facilitate integration. We evaluate Theseus extensively on various GPU workloads with intuitive policies. Compared with NCCL, Theseus achieves up to 1.61X speedup of communication time in stable environments and 2.46X in dynamic environments. It improves end-to-end job completion time by up to 1.84X while incurring comparable or lower overhead. Rui Ding 0014, Xiandong Lu, Xunpeng Liu, Xuran Hao, Houyuan Zhu, Anyi Xu, Sinuo Cao, Haifeng Sun 0004, Qun Huang 0001, Jiamin Cao |
SIGCOMM | 10 |
| 2026 | Enabling General and Efficient Window Mechanism for In-Network TelemetryabstractRecent network telemetry solutions typically target programmable switches to achieve high performance and in-network visibility. They partition the packet stream into windows and then apply various stream processing techniques to summarize flow-level statistics. However, existing studies focus on the measurement within each window. Window management is still a missing piece due to the resource limitation of programmable switches. In this paper, we propose OmniWindow, a general and efficient window mechanism framework. OmniWindow splits the original window into fine-grained sub-windows such that the sub-windows can be merged into various window types. To deal with the resource restriction, OmniWindow carefully designs its data plane memory layout and proposes a window synchronization method. It also employs a collaborative architecture that can collect and reset stateful data in sub-windows within a limited time. We prototype OmniWindow on Tofino. We incorporate OmniWindow into a SOTA query-driven telemetry system and eight sketch-based telemetry algorithms. Our experiments demonstrate that OmniWindow enables these telemetry solutions to achieve higher accuracy than conventional window mechanism. Haifeng Sun 0004, Jintao He, Jie Gui, Qun Huang 0001 |
IEEE Trans. Netw. | 1 |
| 2025 | Poby: SmartNIC-accelerated Image Provisioning for Coldstart in Clouds
Zihao Chang, Haifeng Sun 0004, Yunlong Xie, Kan Shi, Ninghui Sun, Yungang Bao, Sa Wang |
USENIX ATC | 3 |
| 2024 | RB2: Narrow the Gap between RDMA Abstraction and Performance via a Middle LayerabstractAlthough the native RDMA interface allows for high throughput and low latency, its low-level abstraction raises significant programming challenges. Consequently, numerous systems encapsulate the RDMA interface into more user-friendly high-level abstractions such as Socket, MPI, and RPC. However, this ease of development often incurs considerable performance degradation. To address this trade-off, this paper introduces RB2, a high-performance RDMA-based Distributed Ring Buffer (DRB). RB2serves as a middle layer that effectively conceals the low-level details of the RDMA interface while also facilitating extension to other high-level abstractions.Nonetheless, it is non-trivial for DRBs to preserve the RDMA performance. We optimize the performance of RB2in three aspects. First, we perform micro-benchmarks to identify the pointer synchronization methods that are seemingly counter-intuitive but offer optimal performance improvements. Second, we propose an adaptive batching mechanism to alleviate the limitations of conventional fixed batching. Finally, we build an efficient memory subsystem using various optimization techniques. RB2outperforms SOTA designs by achieving 2.5 × to 7.5 × throughput while maintaining comparable tail latency for small messages. Haifeng Sun 0004, Yixuan Tan, Yongtong Wu, Qun Huang 0001, Xin Yao 0008, Gong Zhang 0001 |
INFOCOM | 1 |
| 2024 | AutoSketch: Automatic Sketch-Oriented Compiler for Query-driven Network Telemetry
Haifeng Sun 0004, Qun Huang 0001, Jinbo Sun, Wei Wang 0011, Fuliang Li, Yungang Bao, Xin Yao 0008, Gong Zhang 0001 |
NSDI | 1 |
| 2024 | Distributed Network Telemetry With Resource Efficiency and Full AccuracyabstractNetwork telemetry is essential for administrators to monitor massive data traffic in a network-wide manner. Existing telemetry solutions often face the dilemma between resource efficiency (i.e., low CPU, memory, and bandwidth overhead) and full accuracy (i.e., error-free and holistic measurement). We break this dilemma via a network-wide architectural design, which simultaneously achieves resource efficiency and full accuracy in flow-level telemetry for large-scale data centers. carefully coordinates the collaboration among different types of entities in the whole network to execute telemetry operations, such that the resource constraints of each entity are satisfied without compromising full accuracy. It further addresses consistency in network-wide epoch synchronization and accountability in error-free packet loss inference. We prototype in DPDK and P4. Testbed experiments on commodity servers and Tofino switches demonstrate the effectiveness of over state-of-the-art solutions. Haifeng Sun 0004, Qun Huang 0001, Patrick P. C. Lee, Wei Bai 0001, Yungang Bao |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | OmniWindow: A General and Efficient Window Mechanism Framework for Network TelemetryabstractRecent network telemetry solutions typically target programmable switches to achieve high performance and in-network visibility. They partition the packet stream into windows and then apply various stream processing techniques to summarize flow-level statistics. However, existing studies focus on the measurement within each window. Window management is still a missing piece due to the resource limitation of programmable switches. In this paper, we propose OmniWindow, a general and efficient window mechanism framework. OmniWindow splits the original window into fine-grained sub-windows such that the sub-windows can be merged into various window types. To deal with the resource restriction, OmniWindow carefully designs its data plane memory layout and proposes a window synchronization method. It also employs a collaborative architecture that can collect and reset stateful data in sub-windows within a limited time. We prototype OmniWindow on Tofino. We incorporate OmniWindow into a SOTA query-driven telemetry system and eight sketch-based telemetry algorithms. Our experiments demonstrate that OmniWindow enables these telemetry solutions to achieve higher accuracy than conventional window mechanism. Haifeng Sun 0004, Jintao He, Jie Gui, Qun Huang 0001 |
SIGCOMM | 1 |
| 2023 | Noah: Reinforcement-Learning-Based Rate Limiter for Microservices in Large-Scale E-Commerce ServicesabstractModern large-scale online service providers typically deploy microservices into containers to achieve flexible service management. One critical problem in such container-based microservice architectures is to control the arrival rate of requests in the containers to avoid containers from being overloaded. In this article, we present our experience of rate limit for the containers in Alibaba, one of the largest e-commerce services in the world. Given the highly diverse characteristics of containers in Alibaba, we point out that the existing rate limit mechanisms cannot meet our demand. Thus, we design Noah, a dynamic rate limiter that can automatically adapt to the specific characteristic of each container without human efforts. The key idea of Noah is to use deep reinforcement learning (DRL) that automatically infers the most suitable configuration for each container. To fully embrace the advantages of DRL in our context, Noah addresses two technical challenges. First, Noah uses a lightweight system monitoring mechanism to collect container status. In this way, it minimizes the monitoring overhead while ensuring a timely reaction to system load changes. Second, Noah injects synthetic extreme data when training its models. Thus, its model gains knowledge on unseen special events and hence remains highly available in extreme scenarios. To guarantee model convergence with the injected training data, Noah adopts task-specific curriculum learning to train the model from normal data to extreme data gradually. Noah has been deployed in the production of Alibaba for two years, serving more than 50000 containers and around 300 types of microservice applications. Experimental results show that Noah can well adapt to three common scenarios in the production environment. It effectively achieves better system availability and shorter request response time compared with four state-of-the-art rate limiters. Zhao Li 0007, Haifeng Sun 0004, Zheng Xiong, Qun Huang 0001, Zehong Hu, Shasha Ruan, Hai Hong, Jie Gui, Jintao He, Zebin Xu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | NfvInsight: A Framework for Automatically Deploying and Benchmarking VNF Chains
Tianni Xu, Haifeng Sun 0004, Xiao-Ming Zhou, Xiufeng Sui, Sa Wang, Qun Huang 0001, Yungang Bao |
J. Comput. Sci. Technol. | 2 |
| 2020 | OmniMon: Re-architecting Network Telemetry with Resource Efficiency and Full AccuracyabstractNetwork telemetry is essential for administrators to monitor massive data traffic in a network-wide manner. Existing telemetry solutions often face the dilemma between resource efficiency (i.e., low CPU, memory, and bandwidth overhead) and full accuracy (i.e., error-free and holistic measurement). We break this dilemma via a network-wide architectural design OmniMon, which simultaneously achieves resource efficiency and full accuracy in flow-level telemetry for large-scale data centers. OmniMon carefully coordinates the collaboration among different types of entities in the whole network to execute telemetry operations, such that the resource constraints of each entity are satisfied without compromising full accuracy. It further addresses consistency in network-wide epoch synchronization and accountability in error-free packet loss inference. We prototype OmniMon in DPDK and P4. Testbed experiments on commodity servers and Tofino switches demonstrate the effectiveness of OmniMon over state-of-the-art telemetry designs. Qun Huang 0001, Haifeng Sun 0004, Patrick P. C. Lee, Wei Bai 0001, Yungang Bao |
SIGCOMM | 2 |