VLDB 2026 Research / reviewers in the wild / expert
Xiaohe Hu
dblp:154/2753
· DBLP profile ↗
19ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0003-1487-2419ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 1 first-author · 11 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic MeasurementabstractLarge language model (LLM) training is prone to anomalies due to its long duration and large scale, which can lead to significant performance degradation or even training crashes. Due to the synchronization nature of LLM training, anomalies exhibit the cascading effect, making their diagnosis challenging. Existing approaches rely on collecting communication operator information via code instrumentation, which yields only coarse-grained monitoring data and requires modifications to training code or communication libraries. We propose Pulse, a fine-grained, non-intrusive, and easy-to-deploy monitoring system. Our key idea is to enable fine-grained monitoring via traffic measurement. Pulse conducts microsecond-level RDMA traffic measurement on NICs, and transforms flow-level measurements into communication operator measurements, thereby enabling fine-grained and non-intrusive monitoring. We deploy Pulse on a testbed with 64 H200 GPUs and evaluate its anomaly localization capability under common failure scenarios. Pulse achieves machine-level localization in 10 out of 12 scenarios, while existing methods succeed in only 4 and even misdiagnose 2 of the remaining scenarios. Additionally, Pulse achieves over 90% precision and 100% recall, supports up to 2000 concurrent RDMA flow measurements per NIC, and imposes negligible overhead on training performance, making it a practical solution for real-world LLM training environments. Yibo Xiao, Haifeng Sun 0004, Qingkai Meng 0001, Jiong Duan, Xiaohe Hu, Rong Gu 0001, Guihai Chen, Chen Tian 0001 |
ASPLOS (2) | 6 |
| 2026 | Cross-Data Center Training with Heterogeneous Accelerators: Protocols and Evaluation
Bohua Xu, Xiongyan Tang, Dongyue Zhang, Xiaohe Hu, Lexi Xu, Xiaoxiang Wang, Menghao Zhang 0001 |
ICC | 6 |
| 2026 | MEGATRACE: Troubleshooting Hang and Slowdown in Large-Scale LLM Training Clusters
Fangzheng Jiao, Menghao Zhang 0001, Jiaxun Huang, Yanmin Jia, Xiaohe Hu, Bohua Xu, Chunming Hu |
ICDCS | 6 |
| 2026 | CoPT: Collaborative Machine-Learning-Based Mechanism for Adaptive Congestion Control Parameter Tuning in Data Center Networks
Ziwen Yang, Ruya Gu, Duokun Xu, Shuyong Zhu, Yanmin Jia, Xiaohe Hu |
IWQoS | 7 |
| 2026 | EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
Yitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou, Siyuan Cao, Xujie Fan, Yuchen Xu 0003, Junkai Chen, Chenqi Zhao, Nengyuan Zhang, Shaoke Fang, Jiangyuan Chen, Yuanfeng Chen, Zhan Wang 0003, Yuchao Zhang 0004, Yang Liu 0038, Xiangrui Yang 0002, Xiaohe Hu, Limin Xiao 0001, Weifeng Zhang 0003, Yazhu Lan, Jianbo Dong, Binzhang Fu, Wenfei Wu |
SIGCOMM | 21 |
| 2025 | Mnemosyne: Lightweight and Fast Error Recovery for LLM Training in a Just-In-Time Manner
Jinyi Xia, Menghao Zhang 0001, Jiaxun Huang, Yuezheng Liu, Xiaohe Hu, Xudong Liu 0001, Chunming Hu |
APNet | 5 |
| 2025 | THEMIS: Addressing Congestion-Induced Unfairness in Long-Haul RDMA NetworksabstractRDMA is promising for enhancing the performance of cross-datacenter (DC) services. However, deploying RDMA over wide-area networks introduces severe congestion control unfairness, primarily due to asymmetric congestion feedback delays between inter-DC flows and intra-DC flows. As a result, intra-DC flows often bear the full burden of congestion response, leading to drastically increased flow completion times (FCT). In this work, we identify two key forms of unfairness — near-source and near-destination — depending on whether congestion occurs near the sender or receiver of inter-DC flows. Based on this, we propose THEMIS, a fairness maintenance patch for long-haul RDMA networks. To mitigate near-source unfairness, THEMIS devises a Proactive Notification Point to shorten the congestion feedback loop within a single DC. To alleviate near-destination unfairness, THEMIS introduces a Temporary Reaction Point to temporarily slow down the target inter-DC flow until the sender receives the corresponding congestion feedback. We implement an open-source prototype of THEMIS, and evaluate it on both real-world testbed and large-scale simulations. Compared to DCQCN, Annulus and BiCC, THEMIS reduces the intra-DC FCT by up to 79.2%, 63.6% and 55.6%, and decreases overall FCT by up to 61.2%, 31.9% and 59.5% respectively. Zihan Niu, Menghao Zhang 0001, Renjie Xie, Yuan Yang 0001, Xiaohe Hu |
ICNP | 6 |
| 2025 | Hawkeye: Diagnosing RDMA Network Performance Anomalies with PFC ProvenanceabstractRDMA is becoming increasingly prevalent from private data centers to public multi-tenant clouds, due to its remarkable performance improvement. However, its lossless traffic control, i.e., PFC, introduces new complexities in network performance anomalies (NPAs) due to its cascading congestion spreading property, which usually incurs complaints from customers/applications about certain flows' performance degradation. Existing studies fall short in fine-grained visibility of PFC impact and traceability of PFC causality, and are thus ineffective in diagnosing the root causes for RDMA NPAs. In this paper, we propose Hawkeye, an accurate and efficient RDMA NPA diagnosis system based on PFC provenance. Hawkeye comprises 1) a fine-grained PFC-aware telemetry mechanism to record the PFC impact on flows; 2) an in-network PFC causality analysis and tracing mechanism to quickly and efficiently collect causal telemetry for diagnosis; and 3) a provenance-based diagnosis algorithm to comprehensively present the anomaly breakdown, identifying the anomaly type and root causes accurately. Through extensive evaluations on both NS-3 simulations and a Tofino testbed, Hawkeye can quickly and accurately diagnose multiple RDMA NPAs with over 90% precision and 1–4 orders of magnitude lower overhead than baselines. Menghao Zhang 0001, Xiao Li 0044, Qiyang Peng, Mingwei Xu 0001, Xiaohe Hu, Jiahai Yang 0001, Xingang Shi |
SIGCOMM | 8 |
| 2024 | Megatuner: An Offline DCQCN Parameters Tuner For Large-scale ModelsabstractAs high-performance computing clusters (HPCC) scale up and single-node computing power improves, the bottle-neck of large-scale models training is shifting from computation to network communication. Datacenter Quantized Congestion Notification (DCQCN), as the most widely used congestion control algorithm in lossless Remote Direct Memory Access (RDMA) networks today, provides HPCC with low-latency, high-bandwidth, and high-quality network services. However, DCQCN has more than ten adjustable parameters on RDMA-enabled network inter-face cards (RNICs) and switches, and currently, relying on expert experience to adjust these parameters in engineering is inefficient. Additionally, even minor increases in single-iteration time during training can lead to significant resource wastage. Therefore, after conducting an in-depth analysis of different heuristic search algorithms and loss functions, we propose Megatuner—an offline DCQCN parameters tuner tailored for the traffic pattern of large-scale models. We conducted traffic simulations on the collective communications at the foundation of large-scale models, testing both on simulation platforms and in real environments. The results show that parameters tuned by Megatuner significantly outperform those set based on expert experience. We also addresses the issue of bandwidth degradation in real environments, ensuring long-term optimization of network performance. Xiaoxiang Wang, Fangzheng Jiao, Xiaohe Hu, Yang Jing |
GLOBECOM | 5 |
| 2024 | Paraleon: Automatic and Adaptive Tuning for DCQCN Parameters in RDMA NetworksabstractRDMA is a kernel-bypass and transport-offload technology that provides high throughput and low delay for datacenter networks, and DCQCN is the default and most widely used congestion control algorithm in large-scale RDMA networks. DCQCN involves over 10 parameters at RNICs and switches, and their settings significantly affect network performance, currently relying heavily on exhaustive manual tuning. Although some automatic methods are proposed to tune a subset of DCQCN parameters, none of them comprehensively address all parameters at both RNICs and switches, resulting in compromised network performance. In this paper, we propose Paraleon, an automatic and adaptive system to tune DCQCN parameters comprehensively. We design a millisecond-level sketch-based monitoring mechanism for accurate network-wide measurement, which collects runtime metrics as feedback to guide the tuning process. We also analyze the complicated parameter impacts on network performance, and leverage an improved heuristic searching algorithm for timely performance optimization with better efficiency and convergence. We implement Paraleon and conduct extensive experiments in both NS3 simulations and a real-world testbed. The results show that Paraleon achieves$3.8 \% \sim 61.4 \%$higher performance than existing tuning schemes. Ziteng Chen, Menghao Zhang 0001, Jiahao Cao 0001, Yang Jing, Mingwei Xu 0001, Renjie Xie, Fangzheng Jiao, Xiaohe Hu |
ICNP | 10 |
| 2023 | Kano: Efficient Cloud Native Network Policy VerificationabstractCloud-native computing has become a prevailing paradigm with lightweight runtime-level isolation and fast delivery for scalable applications. Cloud-native network policies (CNNPs) are used to realize network isolation with respect to security and availability. Due to the dynamic environment, CNNPs are label-based instead of IP-based and take the form of attribute-based access control (ABAC) to obtain good expressivity. To ensure the correctness of network isolation, CNNP verification is an essential but challenging problem given the large scale and frequent updates of cloud-native environments and the operation automation demand. Thus, we design Kano, an efficient, i.e., easy-to-use and fast-to-execute, system for verifying large scale CNNPs at runtime. Kano is operation-friendly, with a proposed intent-based verification language. A bit matrix model with a prefiltration algorithm and a partial-update method is proposed to support fast complete and incremental verification. Kano also generates fix plans for violations to assist operators. Kano is implemented as a CNNP verification system that is used in ABAC cloud-native platforms and is integrated into the popular Kubernetes orchestrator. An evaluation on a large scale network of 100k nodes and about 68k policies shows the efficiency of Kano, with 12.51 seconds for all reachable invariant verification and 0.299 milliseconds for policy addition verification. Xiaohe Hu, Chengjun Jia |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2022 | FPGA-Based Updatable Packet Classification Using TSS-Combined Bit-Selecting TreeabstractOpenFlow switches are being deployed in SDN to enable a wide spectrum of non-traditional applications. As a promising alternative to brutal force TCAMs, FPGA-based packet classification is being actively investigated. However, none of the existing FPGA designs can achieve high performance on both search and update for large-scale rule sets. To address this issue, we propose TcbTree, an FPGA-based algorithmic scheme for packet classification. Specifically, at the algorithmic side, i) a two-stage framework consisting of heterogeneous algorithms is proposed, where most rules can be mapped into several balanced trees without rule replications, ii) for the remaining few rules, a centralized TSS (Tuple Space Search) architecture together with a real-time feedback scheme is designed to enhance the efficiency of TSS search on FPGA, and iii) a tree dilution method is designed to equalize rule distribution in trees, so that the latency of tree search can be reduced. At the hardware side, i) an efficient data structure set is designed to convert tree traversal to addressing process, which breaks the constraints of limited tree depth and imbalanced node distribution, and ii) distinct from fully pipelined designs, multiple levels of parallelism are efficiently explored with multi-core, multi-search-engine and coarse-grained pipelines herein. Experimental results using ClassBench show that, with the implementation of TcbTree on FPGA, the average classification throughputs for 1k, 10k, 32k and 100k rule sets achieve 788.8 MPPS, 404.3 MPPS, 237 MPPS and 41.8 MPPS, respectively, and the update throughput for all benchmark rule sets is above 1 MUPS. Yao Xin, Wenjun Li 0004, Guoming Tang, Tong Yang 0003, Xiaohe Hu, Yi Wang 0004 |
IEEE/ACM Trans. Netw. | 5 |
| 2021 | Mahjong: A Generic Framework for Network Data Plane VerificationabstractExisting network data plane verification approaches check network correctness with various models and algorithms. With respect to a specific scenario, it is hard to judge which network model provides sufficient functionality and suitable performance, because existing verification approaches are implemented with different languages and evaluated against different datasets on different hardware platforms in their papers. A network operator usually has to try out a number of complex verification approaches to find the best one for her/his network and intents. Mahjong has a modular system architecture, a unified input format, and three classic verification tools built-in. Leveraging its well-defined partition interfaces and straight-forward configuration file, not only existing approaches can be refactored and merged into Mahjong, new approaches can also be introduced and evaluated with ease. Chengjun Jia, Xiaohe Hu, Jun Li 0003 |
ANCS | 3 |
| 2018 | Multi-core HTB for bandwidth sharingabstractRate limiting with bandwidth-sharing is important and widely used in various scenarios such as multi-tenant cloud. We propose a new rate limiting architecture which fully utilizes the parallel computing capabilities on the multi-core platforms. With Bandwidth Allocator allocating the bandwidth to Rate Limiters, we expand HTB into mHTB, which could provide scalable and flexible rate limiting on multi-core platforms. Chengjun Jia, Zhe Fu 0006, Xiaohe Hu, Shui Cao, Jun Li 0003 |
ANCS | 3 |
| 2018 | APF: fast network all-pair reachability calculationabstractNetwork verification, which is mainly about checking network states against network invariants or operator beliefs, has been a popular research thesis recently. Fast calculation of all-pair reachability of the network beforehand for further query or overall analysis can be a huge assistance to network verification. In this paper, we propose a new fast all-pair reachability calculation algorithm Atomic Predicates Flooding(APF). Experiments on real-life datasets show that the new algorithm based on network atomization is about three to four orders of magnitude faster than existing algorithms without network atomization. On most kinds of datasets, APF is even 2 to 5 times faster than the Warshall based all pair reachability algorithm with atomization. We believe that our method is essential for developing more practical and more efficient network and verification tools. Jinghan Zhou, Xiaohe Hu, Shui Cao, Jun Li 0003 |
ANCS | 3 |
| 2018 | Preserving Privacy at IXPsabstractAutonomous systems (ASes) on the Internet increasingly rely on Internet Exchange Points (IXPs) for peering. A single IXP may interconnect several 100s or 1000s of participants (ASes) all of which might peer with each other through BGP sessions. IXPs have addressed this scaling challenge through the use of route servers. However, route servers require participants to trust the IXP and reveal their policies, a drastic change from the accepted norm where all policies are kept private. In this paper we look at techniques to build route servers which provide the same functionality as existing route servers without requiring participants to reveal their policies thus preserving the status quo and enabling wider adoption of IXPs. Prior work has looked at secure multiparty computation (SMPC) as a means of implementing such route servers however this affects performance and reduces policy flexibility. In this paper we take a different tack and build on trusted execution environments (TEEs) such as Intel SGX to keep policies private and flexible. We present results from an initial route server implementation that runs under Intel SGX and show that our approach has 20x better performance than SMPC based approaches. Furthermore, we demonstrate that the additional privacy provided by our approach comes at minimal cost and our implementation is at worse 2.1x slower than a current route server implementation (and in some situations up to 2x faster). Xiaohe Hu, Arpit Gupta, Nick Feamster, Aurojit Panda, Scott Shenker |
APNet | 1 |
| 2017 | AHI: Efficient policy space set operationsabstractWith the fast industrial deployment of software-defined networking (SDN) and network function virtualization (NFV) technologies, network function policy enforcement in large scale virtualized networks becomes a key challenge for network management. Distributed policy enforcement heavily involves network policy space analysis, where set operations consume most of the computation. Based on spatial projection and bitmap indexing, a novel algorithm AHI (Atomic Hyper-Rectangle Indexing) is proposed for fast policy space set operations. Experiments with real datasets demonstrated that AHI improves set operation speed by two to three orders of magnitude and achieves the same least space cost, comparing to existing state-of-the-art algorithms r-BDD, wildcard expression, and PSA. Xiaohe Hu, Linli Wan, Zhi Liu 0001, Jun Li 0003 |
ICC | 2 |
| 2017 | SCL: Simplifying Distributed SDN Control Planes
Aurojit Panda, Wenting Zheng, Xiaohe Hu, Arvind Krishnamurthy, Scott Shenker |
NSDI | 3 |
| 2014 | Practical regular expression matching free of scalability and performance barriers
Kai Wang 0041, Zhe Fu 0006, Xiaohe Hu, Jun Li 0003 |
Comput. Commun. | 3 |