EDBT 2026 Demo / reviewers in the wild / expert
Xiaohu Xu
dblp:70/7881
· DBLP profile ↗
13ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | xCCLTuner: Treating xCCL as Black-Box and Automatically Tuning
Chenxu Wang 0007, Zhehao Lin, Peirui Cao, Xiaohu Xu, Wan-Chun Dou, Guihai Chen, Chen Tian 0001 |
INFOCOM | 6 |
| 2026 | Achieving Efficient and Robust Multi-Job Resource Scheduling in Deep Learning Clusters
Jianfeng Bao, Wentao Fan 0002, Gongming Zhao, Hongli Xu 0001, Peng Yang 0022, Xiaohu Xu |
IEEE Trans. Netw. | 6 |
| 2026 | Fine-Grained Scheduling of In-Network Aggregation Resources for Efficient Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Cam-Tu Nguyen, Shaoling Sun, Xiaohu Xu, Yongqiang Xiong, Wei Wang 0002, Xiaoliang Wang 0001, Guihai Chen |
IEEE Trans. Netw. | 12 |
| 2026 | Rail: ReArranging Inter-GPU Links for GPU-Centric ClustersabstractIn modern GPU-centric clusters, large-scale AI training relies on two distinct communication domains: a high-bandwidth intra-node domain using proprietary interconnects (e.g., NVLink), and a scale-out inter-node network domain (e.g., RDMA). We observe that the widely-used ring algorithm, often create a significant load imbalance across these domains. This leads to the counter-intuitive scenario where the expensive, high-bandwidth intra-node domain becomes a performance bottleneck, while the inter-node network remains underutilized. This inefficiency is further exacerbated by the disparity in bandwidth provisioning: inter-node network bandwidth is generally more cost-effective and accessible, whereas intra-node bandwidth is often proprietary and more costly to scale. To address this fundamental imbalance, we propose RAIL, aimed at resolving the intra-node bottleneck by strategically rearranging inter-GPU communication paths. This rebalancing ensures that traffic loads are appropriately matched with the distinct transmission capabilities of each domain, thereby maximizing overall communication performance. RAIL incorporates a Load Distributing Strategy (LDS) that can accurately partition physical nodes into logical nodes based on the a transmission capabilities of both domains, shifting excess traffic from the overloaded intra-node domain to the underutilized network domain. Additionally, the Intra-Rail Strategy (IRS) leverages topological characteristics to ensure optimal communication paths through the network domain between logical nodes. Our evaluation demonstrates that RAIL effectively mitigates congestion and achieves a 30.7% average increase in collective communication bus bandwidth compared to the widely-used NCCL solution. Haixin Nan, Jun Xu 0037, Peirui Cao, Zhaochen Zhang, Yizhi Wang 0004, Zhehao Lin, Yuhang Li 0002, Chengyuan Huang, Xiaohu Xu, Zhongming Ji, Shengju Zhang, Lingkun Meng, Rong Gu 0001, Guihai Chen, Chen Tian 0001 |
IEEE Trans. Netw. | 10 |
| 2025 | Mina: Fine-Grained In-network Aggregation Resource Scheduling for Machine Learning Service
Shichen Dong, Zhixiong Niu, Mingchao Zhang, Zhiying Xu, Chuntao Hu, Pengzhi Zhu, Qingchun Song, Peng Cheng 0005, Cam-Tu Nguyen, Shaoling Sun, Xiaohu Xu, Yongqiang Xiong, Wei Wang 0002, Xiaoliang Wang 0001 |
INFOCOM | 12 |
| 2025 | HiReC: High-Throughput and Reliable Cross-Cluster VPC Communication in CloudsabstractThe increasing demands of tenants are driving the growth of single virtual private cloud (VPC), leading to a trend towards cross-cluster VPC deployments, which fuels an increasing demand for cross-cluster VPC communication. However, with the rapid increase in cross-cluster traffic and its inherent dynamism, existing solutions fail to meet tenants' demands for throughput and reliability, thereby leading to network performance bottlenecks in cross-cluster communication. To address this issue, we present HiReC, a system designed to achieve high-throughput and reliable cross-cluster VPC communication. To improve throughput, HiReC leverages multiple gateways with a rounding-based mapping algorithm for load balancing to forward cross-cluster traffic. Moreover, we further enhance the forwarding capabilities of gateways with the eXpress Data Path (XDP) technology. To enhance reliability, HiReC employs a low-overhead, eBPF-based monitoring module and adaptive load adjustment mechanism to dynamically adjust traffic distribution among gateways, effectively handling gateway node or link failures. We implement our system and evaluate its performance through testbed experiments. The results show that HiReC can effectively improve the throughput of cross-cluster communication and deal with abnormal events. For example, HiReC improves the throughput by$3.88 \times$and reduces the failure recovery latency by$19 \times$compared with state-of-the-art solutions. Baoqing Wang, Gongming Zhao, Hongli Xu 0001, Wentao Fan 0002, Xiaohu Xu |
IWQoS | 7 |
| 2025 | CROP: Efficient and Robust Multi-Job Placement in Deep Learning ClustersabstractDeep learning (DL) has seen a growing dataset, an expanding model scale, and increasing applications in recent years. There is a notable trend of shifting DL training jobs from local computing units to powerful DL clusters built by cloud providers. These clusters allocate physical training nodes to DL jobs through a process referred to as multi-job placement. Existing multi-job placement strategies fail to achieve high efficiency in resource utilization, DL training, and robustness simultaneously, resulting in poor performance when resources are limited or when abnormalities occur in some devices. To tackle these challenges, we present CROP, an approach that performs efficient and robust multi-job placement in DL clusters. We formulate the efficient and robust multi-job placement problem as a non-linear program and prove its NP-hardness. To solve this problem, we present an effective submodular-based algorithm with a tight approximation factor of ($1-1/e$). We evaluate CROP on a small-scale testbed consisting of 8 physical GPUs and a large-scale simulation employing real-world job traces. Experimental results demonstrate that CROP achieves nearoptimal communication overhead while improving the training throughput of the DL cluster by up to$57.5\%$compared to state-of-the-art solutions. Peng Yang 0022, Gongming Zhao, Hongli Xu 0001, Haibo Wang 0004, Wentao Fan 0002, Xiaohu Xu |
IWQoS | 7 |
| 2025 | CARD: Cost-Efficient and Availability-Aware Application Deployment in Geo-Distributed CloudsabstractThe growing reliance on cloud services has made availability critical for global business continuity. To mitigate disruptions caused by cloud outages, many large-scale applications maintain core functionality through multi-instance deployments. To support this, cloud providers enable cross-AZ deployments to deliver high availability. However, pursuing high availability must be balanced against cost efficiency, which presents three key challenges: electricity price disparity, application affinity requirement, and disaster recovery demand. Existing research primarily focuses on single-region optimization, often overlooking the potential benefits of multi-region deployment in cost and availability. While some works explore multi-region deployment strategies, they fail to address application affinity or disaster recovery requirements, resulting in low application availability. To address this issue, we propose CARD, a cost-efficient and availability-aware application deployment scheme in geodistributed clouds. Specifically, we formulates this problem as a mixed-integer nonlinear program and designs an efficient approximation algorithm based on submodular function, achieving an approximation ratio of ($1-1 / e$). Large-scale simulations on realworld datasets demonstrate the algorithm's effectiveness, overall reducing costs by 20% – 50% and improving availability by 90% compared to existing solutions. Gongming Zhao, Hongli Xu 0001, Wentao Fan 0002, Xiaohu Xu |
IWQoS | 6 |
| 2024 | Advancing Malware Detection in Network Traffic With Self-Paced Class Incremental LearningabstractEnsuring network security, effective malware detection is of paramount importance. Traditional methods often struggle to accurately learn and process the characteristics of network traffic data, and must balance rapid processing with retaining memory for previously encountered malware categories as new ones emerge. To tackle these challenges, we propose a cutting-edge approach using self-paced class incremental learning (SPCIL). This method harnesses network traffic data for enhanced class incremental learning (CIL). A pivotal technique in deep learning, CIL facilitates the integration of new malware classes while preserving recognition of prior categories. The unique loss function in our SPCIL-driven malware detection combines sparse pairwise loss with sparse loss, striking an optimal balance between model simplicity and accuracy. Experimental results reveal that SPCIL proficiently identifies both existing and emerging malware classes, adeptly addressing catastrophic forgetting. In comparison to other incremental learning approaches, SPCIL stands out in performance and efficiency. It operates with a minimal model parameter count (8.35 million) and in increments of 2, 4, and 5, achieves impressive accuracy rates of 89.61%, 94.74%, and 97.21% respectively, underscoring its effectiveness and operational efficiency. Xiaohu Xu, Xixi Zhang 0001, Qianyun Zhang 0001, Yu Wang 0078, Bamidele Adebisi, Tomoaki Ohtsuki, Hikmet Sari, Guan Gui 0001 |
IEEE Internet Things J. | 1 |
| 2024 | Dynamic Compliant Force Control Strategy for Suppressing Vibrations and Over-Grinding of Robotic Belt Grinding SystemabstractThis work develops a dynamic compliant force control (DCFC) strategy for the robotic belt grinding system to suppress the vibrations and over-grinding phenomenon. First, the vibration mechanism is investigated, and the corresponding vibration models before and after contact are constructed, both of which can decompose the vibrations into three components: free, accompanying and forced vibrations. Next, the extra compliant hardware is equipped to the grinder to realize the dynamic adjustment of equivalent damping. The DCFC strategy considering mechanical compliance accompanied by the dynamic closed-loop control of the grinder damping is presented based on the empirical wavelet transform and multi-scale permutation entropy. Moreover, the cutting fluctuation ratio index is proposed to evaluate the severity of the over-grinding together with over-grinding time. Experiments demonstrate that compared with the general force control, the DCFC strategy can reduce the vibration amplitude from 3.01 mm/s$^2$to 0.97 mm/s$^2$, and the over-grinding time/cutting fluctuation ratio from 1.05 s/24.71% to 0.54 s/13.18%, consequently enhancing the grinding stability and quality.Note to Practitioners—This work is motivated by the need to maintain the grinding stability and suppress the over-grinding phenomenon for the robotic belt grinding system. The vibrations and over-grinding phenomenon caused by the weak rigidity and poor precision of the robot always result in incomplete grinding of the workpiece that needs partially manual regrinding. The proposed DCFC strategy enables a dynamic compliant contact force between the grinding tool and the robot, thus effectively suppressing the over-grinding phenomenon and enhancing the consistency of grinding quality, which is particularly suitable for considering material removal consistency in cut-in and cut-out areas. This method can be implemented by the user as described or be acquired as a standalone device, but some extra hardware needs to be equipped with the grinding tool to use this method. Zeyuan Yang 0003, Xiaohu Xu, Minxing Kuang, Dahu Zhu, Sijie Yan, Shuzhi Sam Ge, Han Ding 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2022 | An Automatic Pavement Crack Detection System with FocusCrack DatasetabstractRoad safety has always been one of the main concerns. With the development of deep learning, computer vision has begun to be used in road damage detection. It has the advantages of faster detection speed, lower cost, easier deployment, etc., which greatly reduces traffic accidents caused by the road. We design an automatic detection system for road damages and deploy it on NVIDIA Jetson Xavier NX. A dataset named FocusCrack is collected under diversiform roads and various lighting conditions. It contains six types of diseases, a total of 4181 images and 5812 labels. Compared with performance of several mainstream algorithms like Faster R-CNN and Single Shot MultiBox Detector (SSD), the model adopts the You Only Look Once v5s(YOLOv5s) algorithm. After experimental testing, the precision, recall, [email protected], and [email protected]:0.95 are 90.1%, 91.3%, 93.8%, and 51.9%. The system has achieved good results in practical application. Xinyun Yan, Xiaohu Xu, Zhengran He, Chishe Wang, Zhiyi Lu |
VTC Fall | 3 |
| 2016 | DNS with mapping service in identifier locator split architectureabstractThis paper mainly introduces the basic functions of DNS, and its usage for mapping service in the Identifier Locator Split (ILS) schemes, as one transitionally functional component in future network architecture. As well known, the overloaded semantics of IP address, being used for both endpoint identifier and routing locator in traditional internet, has hindered the smooth support for mobility of mobile users in network layer, thus lots of ILS schemes have been proposed to solve this problem, with considerations of enabling node mobility, multi-homing, universal connectivity and optimized routing scalability simultaneously. In these proposed ILS schemes, DNS is usually involved in different levels. In general, there exist four types of DNS's roles with Identifier Locator Mapping System (ILMS) in ILS architectures: Usage in a traditional way for merely translating hostnames to IP addresses; Support direct resolution from identifier to locator during node mobility, using newly defined Resource Records (RRs); Serve as a redirection agent for nodes to pinpoint the right mapping servers for further queries; And function in a hybrid manner for both direct resolution with new RRs and redirection service. In addition, some key performance metrics of DNS for ILMS are highlighted as well, and competing approaches for ILMS other than DNS are also discussed. As a result, through our detailed analyses in this survey, a better understanding for DNS's functions, especially with mapping service in ILMS, could be achieved. Bin Da, Xiaohu Xu, Kunyang Bi, Xiuli Zheng |
APCC | 2 |
| 2009 | Enhanced MILSA Architecture for Naming, Addressing, Routing and Security Issues in the Next Generation InternetabstractMILSA (Mobility and Multihoming supporting Identifier Locator Split Architecture) has been proposed to address the naming and addressing challenges for NGI (next generation Internet), we present several design enhancements for MILSA which include a hybrid architectural design that combines "core-edge separation approach" and "split approach", a security-enabled and logically oriented hierarchical identifier system, a three-level identifier resolution system, a new hierarchical code based design for locator structure, cooperative mechanisms among the three planes in MILSA model to assist mapping and routing, and an integrated MILSA service model. The underlying design rationale is also discussed along with the design descriptions. Further analysis addressing the IRTF (Internet Research Task Force) RRG (Routing Research Group) design goals shows that the enhanced MILSA provides comprehensive benefits in routing scalability, traffic engineering, mobility and multihoming, renumbering, security, and deployability. Jianli Pan, Raj Jain, Subharthi Paul, Mic Bowman, Xiaohu Xu, Shanzhi Chen |
ICC | 5 |