Xumiao Zhang

dblp:207/1740 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-3551-4074ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 2 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 AIDA: Accelerating Root Cause Analysis for Multi-Vendor Device Failures with LLM-Powered Reasoning
abstract
Root cause analysis (RCA) of network device failures is critical to cloud reliability. While monitoring can identify which device has failed, diagnosing why remains a slow, manual process, increasing the risk of recurring failures and cascading service disruptions. Existing automated methods are inadequate: traditional methods lack precision, while prior machine learning (ML) and large language model (LLM) approaches are often too coarse-grained, require heavy manual configuration, or fail to produce verifiable reasoning essential for operator trust. This paper presents AIDA, the first system to deliver automated, fine-grained RCA of network device failures, deployed at scale in Alibaba Cloud's production network. AIDA's contributions include: (1) fine-tuning an LLM with reinforcement learning to distill expert logic into interpretable reasoning chains; (2) synthesizing these chains into an evolving knowledge graph (KG) to support retrieval-augmented generation (RAG); and (3) employing RAG-driven multi-step inference wherein the LLM is sequentially guided by the KG to construct robust, verifiable reasoning. Deployed for over a year, AIDA has achieved 95.4% precision with interpretable output and reduced the median RCA time from 72.6 hours to 1.6 minutes. Notably, it curtails the 90th-percentile diagnosis latency from 329.9 hours to 19.4 hours.
Xuan Zeng 0002, Xumiao Zhang, Xiaoxi Zhang 0001, Deke Guo, Ennan Zhai
SIGCOMM3
2026 Evolution of AliYANG: Model-driven and LLM-assisted Network Configuration Management
abstract
Configuration management in large-scale cloud networks is increasingly challenging due to vendor heterogeneity, diverse configuration interfaces, and rapid configuration evolution across the network life cycle. Existing approaches rely heavily on vendor- and interface-specific templates and scripts, which are difficult to validate, costly to maintain, and scale poorly. We introduce AliYANG, a YANG-based configuration modeling framework that unifies configuration representation across vendors and management interfaces. It extends YANG to capture CLI semantics and derives a vendor-agnostic core model that separates configuration semantics from vendor-specific implementations. We further present NetCMDB, the production software infrastructure for AliYANG, which compiles models into typed configuration objects and supports end-to-end, model-driven configuration workflows. As networks evolve, manually constructing and maintaining models becomes a bottleneck. We incorporate LLM-assisted automation to facilitate vendor model augmentation, core model design, and bidirectional translation code generation. We report our three-year production deployment experience managing hundreds of thousands of devices, present evaluation results and case studies, and share lessons from operating a model-driven configuration system at cloud scale.
Mohan Yu, Xumiao Zhang, Zhe An, En Wang
SIGCOMM2
2025 Learning Production-Optimized Congestion Control Selection for Alibaba Cloud CDN
Xuan Zeng 0002, Xumiao Zhang, Xiaoxi Zhang 0001, Xu Chen 0004, Guihai Chen, Yubing Qiu, Chong Hao, Ennan Zhai
NSDI4
2025 Towards LLM-Based Failure Localization in Production-Scale Networks
abstract
Root causing and failure localization are critical to maintain reliability in cloud network operations. When an incident is reported, network operators must review massive volumes of monitoring data and identify the root cause (i.e., error device) as fast as possible, making it extremely challenging even for experienced operators. Large language models (LLMs) have shown great potential in text understanding and reasoning. In this paper, we present BiAn, an LLM-based framework designed to assist operators in efficient incident investigation. BiAn processes monitoring data and generates error device rankings with detailed explanations. To date, BiAn has been deployed in our network infrastructure for 10 months and it has successfully assisted operators in identifying error devices more quickly, reducing time to root causing by 20.5% (55.2% for high-risk incidents). Extensive performance evaluations based on 17 months of real cases further demonstrate that BiAn achieves accurate and fast failure localization. It improves accuracy by 9.2% compared to the baseline approach.
Chenxu Wang 0007, Xumiao Zhang, Runwei Lu, Xianshang Lin, Xuan Zeng 0002, Zhe An, Gongwei Wu, Chen Tian 0001, Guihai Chen, Guyue Liu, Yuhong Liao, Dennis Cai, Ennan Zhai
SIGCOMM2
2025 SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud Infrastructures
abstract
For providers operating large-scale global networks, the timeliness of network failure recovery significantly affects the reliability of network services. Ideally, a network monitoring system should have enough coverage to detect even minor issues, but high coverage means alert floods during severe network failures. In practice, there is a gap between the flooding raw alerts data collected by network monitoring tools and the readable information needed for failure diagnosis. Existing solutions using limited network monitoring data sources and heuristic diagnostic rules, lack comprehensive coverage and the capability to address severe failures, especially which network operators have never handled a similar one before. This paper presents SkyNet, a network analysis system to extract scope and severity information from alert floods. SkyNet ensures comprehensive coverage by integrating multiple monitoring data sources through a uniform input format, enhancing extensibility for new network monitoring tools. During alert floods, SkyNet groups alerts, assesses their severity, and filters out insignificant ones to aid network operators in mitigating network failures. To date, SkyNet has been running stably on our network for one and a half years without any false negatives and has successfully reduced the time-to-mitigation for over 80% of network failures since its deployment in production.
Huanwu Hu, Yunguang Li, Xiangyu Tang, Bingchuan Tian, Gongwei Wu, Xumiao Zhang, Ennan Zhai, Yuhong Liao, Dennis Cai
SIGCOMM9
2025 Roaming Free in the VR World with MP2
Xumiao Zhang, Yuning Chen, Xuan Zeng 0002, Zhilong Zheng, Xianshang Lin, Yanmei Liu, Songwu Lu, Z. Morley Mao, Wan Du, Dennis Cai, Ennan Zhai
USENIX ATC2
2024 OASIS: Collaborative Neural-Enhanced Mobile Video Streaming
abstract
Neural-enhanced video streaming (e.g., super-resolution) is an ongoing revolution which can provide extremely high-quality video streaming services breaking the restriction of bandwidth. However, such enhancements require intense computation power that is not affordable for a single mobile device, which hinders their real-world deployment. To address the limitation, we propose OASIS, the first system that facilitates multiple users in close proximity to execute intense neural-enhanced video streaming in realtime. To this end, OASIS intelligently distributes computation tasks among multiple mobile devices, selects appropriate video bitrates and super-resolution models, and optimizes video chunk delivery. As a result, the expensive neural-enhanced streaming is done through distributed collaboration, achieving optimal quality of experience (QoE). We implement and evaluate OASIS on commodity smartphones from different vendors, under various network and computation conditions. Extensive experiments demonstrate the high efficiency of OASIS: it improves the video streaming QoE by 40%-200% and reduces each participant's energy consumption by 60% when the system scales up from a single device to six devices.
Shuowei Jin, Ruiyang Zhu, Ahmad Hassan 0004, Xiao Zhu 0001, Xumiao Zhang, Z. Morley Mao, Feng Qian 0001, Zhi-Li Zhang
MMSys5
2024 Vulcan: Automatic Query Planning for Live ML Analytics
Yiwen Zhang 0008, Xumiao Zhang, Ganesh Ananthanarayanan, Anand Padmanabha Iyer, Yuanchao Shu, Paramvir Bahl, Z. Morley Mao, Mosharaf Chowdhury
NSDI2
2024 Boosting Collaborative Vehicular Perception on the Edge with Vehicle-to-Vehicle Communication
abstract
Collaborative Vehicular Perception (CVP) enables connected and autonomous vehicles (CAVs) to cooperatively extend their views through wirelessly sharing their sensor data. Existing CVP systems employ either a vehicle-to-vehicle (V2V) or vehicle-to-infrastructure (V2I) view exchange paradigm. In this paper, we advocate a hybrid CVP design: our developed system, Harbor, employs V2I as its fundamental underlying framework, and opportunistically employs V2V to boost the performance. In Harbor, vehicles (helpers) may serve as relays to assist other vehicles (helpees) in reaching an edge node, which performs sensor data merging to produce the extended view. We judiciously partition the workload between the edge and vehicles, develop a robust helper-helpee assignment model, and solve it efficiently at runtime. We conduct both real-world tests and large-scale emulation experiments using two prevailing CAV applications: drivable space detection and object detection. Our real-world evaluation conducted at one of the world's first purpose-built autonomous driving testbeds demonstrates that Harbor outperforms state-of-the-art V2V- or V2I-only CVP schemes by up to 36% in detection accuracy, resulting in significantly fewer collisions under dangerous driving scenarios.
Ruiyang Zhu, Xiao Zhu 0001, Anlan Zhang, Xumiao Zhang, Feng Qian 0001, Hang Qiu 0001, Z. Morley Mao, Myungjin Lee
SenSys4
2024 On Data Fabrication in Collaborative Vehicular Perception: Attacks and Countermeasures
Qingzhao Zhang 0001, Shuowei Jin, Ruiyang Zhu, Xumiao Zhang, Qi Alfred Chen, Z. Morley Mao
USENIX Security Symposium5
2024 QUIC is not Quick Enough over Fast Internet
abstract
QUIC is expected to be a game-changer in improving web application performance. In this paper, we conduct a systematic examination of QUIC's performance over high-speed networks. We find that over fast Internet, the UDP+QUIC+HTTP/3 stack suffers a data rate reduction of up to 45.2% compared to the TCP+TLS+HTTP/2 counterpart. Moreover, the performance gap between QUIC and HTTP/2 grows as the underlying bandwidth increases. We observe this issue on lightweight data transfer clients and major web browsers (Chrome, Edge, Firefox, Opera), on different hosts (desktop, mobile), and over diverse networks (wired broadband, cellular). It affects not only file transfers, but also various applications such as video streaming (up to 9.8% video bitrate reduction) and web browsing. Through rigorous packet trace analysis and kernel- and user-space profiling, we identify the root cause to be high receiver-side processing overhead, in particular, excessive data packets and QUIC's user-space ACKs. We make concrete recommendations for mitigating the observed performance issues.
Xumiao Zhang, Shuowei Jin, Yi He 0015, Ahmad Hassan 0004, Z. Morley Mao, Feng Qian 0001, Zhi-Li Zhang
WWW1
2023 Poster: QUIC is not Quick Enough over Fast Internet
abstract
QUIC is a multiplexed transport-layer protocol over UDP and comes with enforced encryption. It is expected to be a game-changer in improving web application performance. Together with the network layer and layers below, UDP, QUIC, and HTTP/3 form a new protocol stack for future network communication, whose current counterpart is TCP, TLS, and HTTP/2. In this study, to understand QUIC's performance over high-speed networks and its potential to replace the TCP stack, we carry out a series of experiments to compare the UDP+QUIC+HTTP/3 (QUIC) stack and the TCP+TLS+HTTP/2 (HTTP/2) stack. Preliminary measurements on file download reveal that QUIC suffers from a data rate reduction compared to HTTP/2 across different hosts.
Xumiao Zhang, Shuowei Jin, Yi He 0015, Ahmad Hassan 0004, Z. Morley Mao, Feng Qian 0001, Zhi-Li Zhang
IMC1
2023 Robust Real-time Multi-vehicle Collaboration on Asynchronous Sensors
abstract
Cooperative perception significantly enhances the perception performance of connected autonomous vehicles. Instead of purely relying on local sensors with limited range, it enables multiple vehicles and roadside infrastructures to share sensor data to perceive the environment collaboratively. Through our study, we realize that the performance of cooperative perception systems is limited in real-world deployment due to (1) out-of-sync sensor data during data fusion and (2) inaccurate localization of occluded areas. To address these challenges, we develop RAO, an innovative, effective, and lightweight cooperative perception system that merges asynchronous sensor data from different vehicles through our novel designs of motion-compensated occupancy flow prediction and on-demand data sharing, improving both the accuracy and coverage of the perception system. Our extensive evaluation, including real-world and emulation-based experiments, demonstrates that RAO outperforms state-of-the-art solutions by more than 34% in perception coverage and by up to 14% in perception accuracy, especially when asynchronous sensor data is present. RAO consistently performs well across a wide variety of map topologies and driving scenarios. RAO incurs negligible additional latency (8.5 ms) and low data transmission overhead (10.9 KB per frame), making cooperative perception feasible.
Qingzhao Zhang 0001, Xumiao Zhang, Ruiyang Zhu, Fan Bai 0002, Mohammad Naserian, Z. Morley Mao
MobiCom2
2021 EMP: edge-assisted multi-vehicle perception
abstract
Connected and Autonomous Vehicles (CAVs) heavily rely on 3D sensors such as LiDARs, radars, and stereo cameras. However, 3D sensors from a single vehicle suffer from two fundamental limitations: vulnerability to occlusion and loss of details on far-away objects. To overcome both limitations, in this paper, we design, implement, and evaluate EMP, a novel edge-assisted multi-vehicle perception system for CAVs. In EMP, multiple nearby CAVs share their raw sensor data with an edge server which then merges CAVs' individual views to form a more complete view with a higher resolution. The merged view can drastically enhance the perception quality of the participating CAVs. Our core methodological contribution is to make the sensor data sharing scalable, adaptive, and resource-efficient over oftentimes highly fluctuating wireless links through a series of novel algorithms, which are then integrated into a full-fledged cooperative sensing pipeline. Extensive evaluations demonstrate that EMP can achieve real-time processing at 24 FPS and end-to-end latency of 93 ms on average. EMP reduces the end-to-end latency by 49% to 65% compared to the traditional vehicle-to-vehicle (V2V) sharing approach without edge support. Our case studies show that cooperative sensing powered by EMP can detect hazards such as blind spots faster by 0.5 to 1.1 seconds, compared to a single vehicle's perception.
Xumiao Zhang, Anlan Zhang, Xiao Zhu 0001, Yihua Guo, Feng Qian 0001, Z. Morley Mao
MobiCom1
2021 A variegated look at 5G in the wild: performance, power, and QoE implications
abstract
Motivated by the rapid deployment of 5G, we carry out an in-depth measurement study of the performance, power consumption, and application quality-of-experience (QoE) of commercial 5G networks in the wild. We examine different 5G carriers, deployment schemes (Non-Standalone, NSA vs. Standalone, SA), radio bands (mmWave and sub 6-GHz), protocol configurations (_e.g._ Radio Resource Control state transitions), mobility patterns (stationary, walking, driving), client devices (_i.e._ User Equipment), and upper-layer applications (file download, video streaming, and web browsing). Our findings reveal key characteristics of commercial 5G in terms of throughput, latency, handover behaviors, radio state transitions, and radio power consumption under the above diverse scenarios, with detailed comparisons to 4G/LTE networks. Furthermore, our study provides key insights into how upper-layer applications should best utilize 5G by balancing the critical tradeoff between performance and energy consumption, as well as by taking into account the availability of both network and computation resources. We have released the datasets and tools of our study at https://github.com/SIGCOMM21-5G/artifact.
Arvind Narayanan, Xumiao Zhang, Ruiyang Zhu, Ahmad Hassan 0004, Shuowei Jin, Xiao Zhu 0001, Denis Rybkin, Zhengxuan Yang, Z. Morley Mao, Feng Qian 0001, Zhi-Li Zhang
SIGCOMM2
2020 MPBond: efficient network-level collaboration among personal mobile devices
abstract
MPBond is an efficient system allowing multiple personal mobile devices to collaboratively fetch content from the Internet. For example, a smartwatch can assist its paired smartphone with downloading data. Inspired by the success of MPTCP, MPBond applies the concept of distributed multipath transport where multiple subflows can traverse different devices. We develop a cross-device connection management scheme, a buffering strategy, a packet scheduling algorithm, and a policy framework tailored to MPBond's architecture. We implement MPBond on commodity mobile devices such as Android smartphones and smartwatches. Our real-world evaluations using different workloads under various network conditions demonstrate the efficiency of MPBond. Compared to state-of-the-art collaboration frameworks, MPBond reduces file download time by 5% to 46%, and improves the video streaming bitrate by 2% to 118%. Meanwhile, it improves the energy efficiency by 10% to 57%.
Xiao Zhu 0001, Xumiao Zhang, Yihua Guo, Feng Qian 0001, Z. Morley Mao
MobiSys3
2020 MPBond: efficient network-level collaboration among personal mobile devices
abstract
We demo MPBond, a novel multipath transport system allowing multiple personal mobile devices to collaboratively fetch content from the Internet. Inspired by the success of MPTCP, MPBond applies the concept of distributed multipath transport where multiple subflows can traverse different devices. Other key design aspects of MPBond include a device/connection management scheme, a buffering strategy, a packet scheduling algorithm, and a policy framework tailored to MPBond's architecture. We install MPBond on commodity mobile devices and show how easy it is to configure the usage of MPBond for unmodified apps. We visualize the runtime behavior of MPBond to further illustrate its design. We also demonstrate the download time and energy reduction of file download, as well as the video streaming QoE improvement with MPBond.
Xiao Zhu 0001, Xumiao Zhang, Yihua Guo, Feng Qian 0001, Z. Morley Mao
MobiSys3
2017 Demo: The Sound of Silence: End-to-End Sign Language Recognition Using SmartWatch
abstract
Sign Language is a natural and fully-fledged communication method for deaf and hearing-impaired people. In this demo, we propose the first SmartWatch-based American sign language (ASL) recognition system, which is more comfortable, portable and user-friendly and offers accessibility anytime, anywhere. This system is based on the intuitive idea that each sign has its specific motion pattern which can be transformed into unique gyroscope and accelerometer signals and then analyzed and learned by using Long-Short term memory recurrent neural network (LSTM-RNN) trained with connectionist temporal classification (CTC). In this way, signs and context information can be correctly recognized based on an off-the-shelf device (eg. SmartWatch, Smartphone). The experiments show that, in the Known user split task, our system reaches an average word error rate of 7.29% to recognize 73 sentences formed by 103 ASL signs and achieves detection ratio up to 93.7% for a single sign. The result also shows our system has a good adaptation, even including new users, it can achieve an average word error rate of 21.6% at the sentence level and reach an average detection ratio of 79.4%. Moreover, our system performs real time ASL translation, outputting the speech within 1.69 seconds for a sentence of 12 signs in average.
Qian Dai, Jiahui Hou, Panlong Yang, Xiang-Yang Li 0001, Fei Wang 0063, Xumiao Zhang
MobiCom6