Shuai Wang 0028

dblp:42/1503-28 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-7814-240XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 22 · 5 first-author · 17 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 OptiFlow: Towards LLM-Driven Optimization of Collective Communication Algorithms
Ziyue Yang 0002, Kaihui Gao, Shuai Wang 0028, Li Chen 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Dan Li 0001
APNet4
2026 A Large-Scale IPv6-Based Measurement of the Starlink Network
abstract
Low Earth Orbit (LEO) satellite networks have attracted considerable attention for their ability to deliver global, low-latency broadband Internet services. In this paper, we present a large-scale measurement study of the Starlink network, the largest LEO satellite constellation to date. We first propose an efficient method for discovering active Starlink user routers, identifying approximately 5.98 million IPv6 addresses across 208 regions in 165 countries. Compared to general-purpose IPv6 target generation algorithms, our router-centric approach achieves near-complete coverage and, to the best of our knowledge, yields the most comprehensive known set of active IPv6 addresses for Starlink user routers. Based on the discovered user routers, we further propose an efficient method for mapping the Starlink backbone network and uncover a topology consisting of 49 Points of Presence (PoPs) interconnected by 98 links. We conduct a detailed statistical analysis of active Starlink user routers and PoPs, and further characterize the IPv6 address assignment strategy adopted by the Starlink network. Finally, we analyze the latency of Starlink user routers, propose a method to distinguish different types of users within the same region using outside-in measurement, and identify the ongoing V2 Mini satellite deployment as a potential driver of the performance improvements. The dataset of the Starlink backbone network is publicly available at https://ki3.org.cn/#/starlink-network.
Bingsen Wang, Shuai Wang 0028, Li Chen 0008, Jinwei Zhao, Dan Li 0001, Yong Jiang 0001
INFOCOM3
2026 OSAVRoute: Advancing Outbound Source Address Validation Deployment Detection with Non-Cooperative Measurement
Shuai Wang 0028, Li Chen 0008, Dan Li 0001, Lancheng Qin
NDSS1
2026 BayWatch: Practical Internet-Scale Topology Monitoring with Dynamic Bayesian Estimation
Zhongxu Guan, Shuai Wang 0028, Li Chen 0008, Zhaoteng Yan, Jiaye Lin, Dan Li 0001, Yong Jiang 0001, Yingxin Wang
NSDI2
2026 Networked Agent Memory and Causality Representation: Experiences towards Interpretable Cloud-Scale Root-Causing
Yanyu Ren, Xianshang Lin, Chenxu Wang 0007, Li Chen 0008, Shuai Wang 0028, Kaihui Gao, Dan Li 0001, Chen Tian 0001, Yunguang Li, Ennan Zhai
SIGCOMM5
2025 LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs - No Silver Bullet for LC or RAG Routing
abstract
As Large Language Model (LLM) context windows expand, the necessity of Retrieval-Augmented Generation (RAG) for integrating external knowledge is debated. Existing RAG vs. long-context (LC) LLM comparisons are often inconclusive due to benchmark limitations. We introduce LaRA, a novel benchmark with 2326 test cases across four QA tasks and three long context types, for rigorous evaluation. Our analysis of eleven LLMs reveals the optimal choice between RAG and LC depends on a complex interplay of model capabilities, context length, task type, and retrieval characteristics, offering actionable guidelines for practitioners. Our code and dataset is provided at:https://github.com/Alibaba-NLP/LaRA
Kuan Li, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Shuai Wang 0028, Minhao Cheng
ICML6
2025 Measuring the Time Source Vulnerabilities in the NTP Ecosystem
abstract
Precise timekeeping is crucial for the dependable functioning and security of multiple Internet infrastructures, such as TLS certificates. Although the Network Time Protocol (NTP) is widely used for time synchronization across devices, it has several security vulnerabilities. Network Time Security (NTS) offers server authentication and integrity verification to protect against man-in-the-middle attacks. However, NTS does not address issues related to erroneous time sources.
Zhentian Huang, Shuai Wang 0028, Li Chen 0008, Dan Li 0001
IMC2
2025 SAIP: Accurate Detection of Anycast Servers with the Rise of Regional Anycast
Shuai Wang 0028, Li Chen 0008, Dan Li 0001
INFOCOM2
2025 Discovering Millions of New Nodes and Links in the Internet by Challenging the Uniformity Assumption in Multipath Detection
abstract
Multipath Detection Algorithms (MDAs) are proposed to discover Internet topology in the presence of load balancing (LB). Existing methods assume uniformity in the load-balancing responses (LBR), i.e., responses from the successors of a LB router. However, we reveal that only 20% of the cases exhibit uniformity in the Internet. This finding significantly challenges the completeness of the Internet topology discovered using current MDAs. In this paper, we propose a novel system BayMuDA, that can estimate LBR distributions and calculate the minimum number of probes needed to statistically discover all nodes and links within a given hop. The validation on controlled topologies shows that BayMuDA discovers at least 85%/73% of nodes/links in ~90% of the cases. Our Internet-wide measurement results indicate that BayMuDA can discover millions of Internet nodes and links obscured by the state-of-the-art MDA algorithm, D-Miner, due to uneven responses.
Zhongxu Guan, Shuai Wang 0028, Li Chen 0008, Zhaoteng Yan, Jiaye Lin, Dan Li 0001, Yong Jiang 0001, Yingxin Wang
SIGCOMM2
2025 Your Shield is My Sword: A Persistent Denial-of-Service Attack via the Reuse of Unvalidated Caches in DNSSEC Validation
Shuai Wang 0028, Li Chen 0008, Dan Li 0001
USENIX Security Symposium2
2024 Robust or Risky: Measurement and Analysis of Domain Resolution Dependency
abstract
DNS relies on domain delegation for good scalability, where domains delegate their resolution service to authoritative nameservers. However, such delegations lead to complex inter-dependencies between DNS zones. While a complex dependency might improve the robustness of domain resolution, it could also introduce security risks unexpectedly. In this work, we perform a large-scale measurement on nearly 217M domains to analyze their resolution dependencies at both zone level and infrastructure level. According to our analysis, domains under country-code TLDs and new generic TLDs generally present more complicated dependency relationships. For robustness consideration, popular domains prefer to configure more complex dependencies. However, the centralization of nameserver hosting and the silent outsourcing of DNS providers could lead to severe false redundancy at infrastructure level. Worse, considerable domain configurations in the wild are "not robust but risky": a more complex dependency may also indicate more vulnerabilities, e.g., domains with a 2× higher dependency complexity have a 2.87× larger proportion suffering from the hijacking risk brought by lame delegation.
Shuai Wang 0028, Dan Li 0001
INFOCOM2
2023 sRDMA: A General and Low-Overhead Scheduler for RDMA
abstract
Remote Direct Memory Access (RDMA) has been widely deployed in data centers to improve application performance. However, the characteristic of RDMA to deliver messages in order cannot meet the emerging requirements of applications for scheduling messages within an RDMA connection, making RDMA unable to be fully utilized. Some works try to schedule the data to be transferred in specific applications before delivering to RDMA, or distribute messages to different connections. However, these approaches tightly couple scheduling logic with application logic and may result in high scheduling overhead.
Xizheng Wang, Shuai Wang 0028, Dan Li 0001
APNet2
2023 Impact of International Submarine Cable on Internet Routing
Honglin Ye, Shuai Wang 0028, Dan Li 0001
INFOCOM2
2023 Poster: Q-Scanner: A Fast Scanning Tool for Large-Scale SSL/TLS Configurations Measurement
abstract
Secure Sockets Layer (SSL) and Transport Layer Security (TLS) protocols are used to encrypt data, protect privacy, and authenticate. However, the security of SSL/TLS itself depends on its configurations. While some scanning tools are used to measure SSL/TLS configurations, their performance is far from meeting the requirement of large-scale measurements. In this paper, we propose a fast SSL/TLS configuration scanning tool, Q-Scanner, which can generate a lightweight scanning solution based on the characteristics of the configurations to be scanned. The experiment shows Q-Scanner achieves a speedup of over 30,000 times compared to SSL Pulse without loss of accuracy.
Shuai Wang 0028, Dan Li 0001
SIGCOMM2
2023 Buffer-Based High-Coverage and Low-Overhead Request Event Monitoring in the Cloud
abstract
Request latency directly affects the performance of modern cloud applications. Due to various causes in hosts and networks, requests can suffer from request latency anomalies (RLAs), which may violate the Service-Level Agreement. However, existing performance monitoring tools have incomplete coverage and inconsistent semantics for monitoring requests and cannot accurately diagnose RLAs. This paper presentsBufScope, a high-coverage and low-overhead request event monitoring system, which monitorsbuffersto capture most RLA-related abnormal events with consistent request-level semantics in the end-to-end datapath of request. First,BufScopemodels the datapath of request as a buffer chain and defines events based on three properties of buffers, so as toend-to-end monitorthe root causes of RLA. Then, to achieveconsistent semanticsfor captured events,BufScopedesigns a request-level semantics injection mechanism to make events captured in networks have the victim requests’ ID. Finally,BufScopeoffloads the semantics operations and event collection in software to SmartNICs forlow CPU overhead. We have implementedBufScopeon commodity SmartNICs and programmable switches. Evaluation results show thatBufScopecan diagnose 98% RLAs with < 0.08% network bandwidth overhead and 0.6% application throughput decline.
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005, Lu Lu 0016
IEEE/ACM Trans. Netw.3
2023 Dependable Virtualized Fabric on Programmable Data Plane
abstract
In modern multi-tenant data centers, each tenant desires reassuring dependability from the virtualized network fabric – bandwidth guarantee with work conservation, bounded tail latency and resilient reachability. However, the slow convergence of prior works under network dynamics and uncertainties can hardly provide the dependability for tenants. Further, state-of-the-art load balance schemes are guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. In this paper, we propose vFab, a dependable virtualized fabric framework which can (1) quickly detect network failure in data plane, (2) explicitly select proper paths for all flows, and (3) converge to ideal bandwidth allocation at sub-millisecond. The core idea of vFab is to leverage the programmable data plane to build a fusion of an active edge (e.g., NIC) and an informative core (e.g., switch), where the core sends link status and tenant information to the edge via telemetry to help the latter make a timely and accurate decision on path selection and traffic admission. We fully implement vFab with commodity SmartNICs and programmable switches. Extensive evaluations show that vFab can keep bandwidth guarantee with high bandwidth utilization, low and bounded latency, and resilient reachability under various network scenarios with limited overhead. Application-level experiments show that vFab can improve QPS by$2.4\times $and cut tail latency by$10\times $compared to the alternatives.
Kaihui Gao, Shuai Wang 0028, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Tao Sun 0010
IEEE/ACM Trans. Netw.2
2022 Bandwidth-efficient Microburst Measurement in Large-scale Datacenter Networks
abstract
Microburst measurement is essential for diagnosing and mitigating performance problems in datacenter networks. The key is to efficiently identify the flows that contribute the most to queue buildup. However, because the existing microburst measurement systems capture packet-level information, they incur significant bandwidth overhead. We present BurstScope, a bandwidth-efficient microburst measurement system that can profile the microburst characteristics and the contributing flows. BurstScope detects the microburst-involved packets in egress pipeline, then aggregates the measurement granularity from packet level to flow level by an invertible sketch. Finally, by carefully partitioning the measurement and statistic tasks between the data and control plane, we generate only one telemetry packet for each microburst. We have implemented BurstScope on Barefoot Tofino switches. Testbed-based evaluations show that BurstScope keeps low bandwidth overhead (< 0.02%) and high identification accuracy (> 97%). Compared with the state-of-the-art system, BurstScope can reduce 60 × bandwidth overhead.
Kaihui Gao, Dan Li 0001, Shuai Wang 0028
APNet3
2022 Buffer-based End-to-end Request Event Monitoring in the Cloud
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005
NSDI3
2022 Predictable vFabric on informative data plane
abstract
In multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works.
Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005
SIGCOMM1
2022 Impact of Synchronization Topology on DML Performance: Both Logical Topology and Physical Topology
abstract
To tackle the increasingly larger training data and models, researchers and engineers resort to multiple servers in a data center for distributed machine learning (DML). On one hand, DML enables us to leverage the computation power of multiple servers, which can effectively accelerate those computation-intensive tasks. On the other hand, DML also incurs significant communication cost due to parameter synchronization among these servers. In this paper, we want to explore the impact of synchronization topology, including both logical topology and physical topology, on the DML performance. First, we revisit the existing logical topologies, e.g., parameter server and ring allreduce, for parameter synchronization, and we find that theseflatsynchronization topologies is inefficient when running a large-scale DML training. Therefore, we propose a hierarchical parameter synchronization topology, called HiPS, which can achieve efficient parameter synchronization even on a large scale. Then, we compare two representative physical network topologies, namely, Fat-Tree and BCube. Based on our analyses, BCube has many advantages over Fat-Tree, e.g., higher bandwidth, better load balance, and lower hardware cost. The simulation results also show that BCube is more friendly to RDMA. Relying on the advantages of HiPS and BCube, the GST of “HiPS+BCube” is 12% ~ 70% lower than other combinations. Moreover, when the cluster size increases from 16 to 1024, the performance of “HiPS+BCube” only drops by 6.5%, while the performance of “Ring+BCube” drops by 44.6%. Hence, we believe “HiPS+BCube” is the optimal solution to benefit DML in large scale.
Shuai Wang 0028, Jinkun Geng, Dan Li 0001
IEEE/ACM Trans. Netw.1
2020 CEFS: compute-efficient flow scheduling for iterative synchronous applications
abstract
Iterative Synchronous Applications (ISApps) are popular in today's data centers, represented by distributed deep learning (DL) training. In ISApps, multiple nodes carry out the computing task iteratively, with globally synchronizing the results in each iteration. To increase the scaling efficiency of ISApps, in this paper we propose a new flow scheduling approach, called CEFS. CEFS saves the waiting time of computing nodes from two aspects. For a single node, flows with data which can trigger earlier computation at the node are assigned with higher priority; among nodes, flows towards slower nodes are assigned with higher priority.
Shuai Wang 0028, Dan Li 0001, Jiansong Zhang 0001, Wei Lin 0016
CoNEXT1
2020 Fela: Incorporating Flexible Parallelism and Elastic Tuning to Accelerate Large-Scale DML
abstract
Distributed machine learning (DML) has become the common practice in industry, because of the explosive volume of training data and the growing complexity of training model. Traditional DML follows data parallelism but causes significant communication cost, due to the huge amount of parameter transmission. The recently emerging model-parallel solutions can reduce the communication workload, but leads to load imbalance and serious straggler problems. More importantly, the existing solutions, either data-parallel or model-parallel, ignore the nature of flexible parallelism for most DML tasks, thus failing to fully exploit the GPU computation power. Targeting at these existing drawbacks, we propose Fela, which incorporates both flexible parallelism and elastic tuning mechanism to accelerate DML. In order to fully leverage GPU power and reduce communication cost, Fela adopts hybrid parallelism and uses flexible parallel degrees to train different parts of the model. Meanwhile, Fela designs token-based scheduling policy to elastically tune the workload among different workers, thus mitigating the straggler effect and achieve better load balance. Our comparative experiments show that Fela can significantly improve the training throughput and outperforms the three main baselines (i.e. dataparallel, model-parallel, and hybrid-parallel) by up to 3.23×, 12.22×, and 1.85× respectively.
Jinkun Geng, Dan Li 0001, Shuai Wang 0028
ICDE3
2020 Geryon: Accelerating Distributed CNN Training by Network-Level Flow Scheduling
abstract
Increasingly rich data sets and complicated models make distributed machine learning more and more important. However, the cost of extensive and frequent parameter synchronizations can easily diminish the benefits of distributed training across multiple machines. In this paper, we present Geryon, a network-level flow scheduling scheme to accelerate distributed Convolutional Neural Network (CNN) training. Geryon leverages multiple flows with different priorities to transfer parameters of different urgency levels, which can naturally coordinate multiple parameter servers and prioritize the urgent parameter transfers in the entire network fabric. Geryon requires no modification in CNN models and does not affect the training accuracy. Based on the experimental results of four representative CNN models on a testbed of 8 GPU servers, Geryon achieves up to 95.7% scaling efficiency even with 10GbE bandwidth. In contrast, for most models, the scaling efficiency of vanilla TensorFlow is no more than 37% and that of TensorFlow with parameter partition and slicing is around 80%. In terms of training throughput, Geryon enhanced with parameter partition and slicing achieves up to 4.37x speedup, where the flow scheduling algorithm itself achieves up to 1.2x speedup over parameter partition and slicing.
Shuai Wang 0028, Dan Li 0001, Jinkun Geng
INFOCOM1
2020 A Scalable, High-Performance, and Fault-Tolerant Network Architecture for Distributed Machine Learning
abstract
In large-scale distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a scalable, high-performance and fault-tolerant DML network architecture on top of Ethernet and commodity devices. BML builds on BCube topology, and runs a fully-distributed gradient synchronization algorithm. Compared to a Fat-Tree network with the same size, a BML network is expected to take much less time for gradient synchronization, for both low theoretical synchronization time and its benefit to RDMA transport. With server/link failures, the performance of BML degrades in a graceful way. Experiments of MNIST and VGG-19 benchmarks on a testbed with 9 dual-GPU servers show that, BML reduces the job completion time of DML training by up to 56.4% compared with Fat-Tree running state-of-the-art gradient synchronization algorithm.
Dan Li 0001, Jinkun Geng, Yanshu Wang, Shuai Wang 0028, Shutao Xia
IEEE/ACM Trans. Netw.6
2019 Accelerating Distributed Machine Learning by Smart Parameter Server
abstract
Parameter Server (PS)-based architecture is widely applied in distributed machine learning (DML), but it is still an open issue how to improve the DML performance in this frame-work. Existing works mainly focus on the view of workers. In this paper, we tackle this problem from another perspective, by leveraging the central control on the PS. Specifically, we propose SmartPS, which transforms the passive role of PS in traditional DML and fully exploits the intelligence of PS. Firstly, the PS holds the global view of parameter dependency, facilitating it to update workers' parameters selectively and proactively. Secondly, the PS records the workers' speeds, and prioritizes parameter transmission to narrow the gap between stragglers and fast workers. Thirdly, the PS considers the parameter dependency in consecutive training iterations, and opportunistically blocks unnecessary pushes from workers. We conduct comparative experiments with two typical benchmarks, Matrix Factorization (MF) and PageRank (PR). The experimental results prove that, compared with all the baseline algorithms (i.e. standard BSP, ASP and SSP), SmartPS can reduce the overall training time by 65.7%~84.9%, with the same training accuracy.
Jinkun Geng, Dan Li 0001, Shuai Wang 0028
APNet3
2019 Rima: An RDMA-Accelerated Model-Parallelized Solution to Large-Scale Matrix Factorization
abstract
Matrix factorization (MF) is a fundamental technique in machine learning and data mining, which gains wide application in many fields. When the matrix becomes large, MF cannot be processed on a single machine. Considering this, many distributed SGD algorithms (e.g. DSGD) have been developed to solve large-scale MF on multiple machines in a model-parallel way. Existing distributed algorithms are primarily implemented under Map/Reduce or PS (parameter server)-based architectures, which incur significant communication overheads. Besides, existing solutions cannot well embrace the benefit of RDMA/RoCE transport and suffer from scalability problems. Targeting at these drawbacks, we propose Rima, which uses ring-based model parallelism to solve large-scale MF with higher communication efficiency. Compared with PS-based SGD algorithms, Rima also consumes less queue pairs (QPs) and can thus better leverage the power of RDMA/RoCE to accelerate the training speed. Our experiment shows that, compared with PS-based DSGD when solving 1M × 1M MF, Rima achieves comparable convergence performance after equal number of iterations, but reduces the training time by 68.7% and 85.4% via TCP and RDMA respectively.
Jinkun Geng, Dan Li 0001, Shuai Wang 0028
ICDE3
2019 Impact of Network Topology on the Performance of DML: Theoretical Analysis and Practical Factors
abstract
To deal with the increasingly larger input data and model sizes, it has become necessary to scale the training of machine learning models to multiple nodes, even a server cluster, which we call distributed machine learning, or DML. However, DML utilizes more computation power at the cost of high communication overhead, which may limit the overall performance in turn. In this paper, we study the impact of network topology on the DML performance both in theory and in practice. We compare two representative network topologies, namely, Fat-Tree which is widely-used in modern data centers, and BCube, which is a low-cost and server-centric network topology, both running on top of RDMA. The results show that Fat-Tree not only has theoretically higher global synchronization time (GST) than BCube, but its practical GST (by NS-3 based simulation) is also considerably larger than the theoretical one. By analyzing the large-scale simulation traces, we find that the root cause for the gap in Fat-Tree comes from the load imbalance among the multiple parallel paths as well as the inevitable PFC frames, both of which do not appear in BCube. For a cluster of around 250 servers, BCube achieves 53%\sim 70% lower GST than Fat-Tree from the simulation. As a result, we suggest using server-centric network topology such as BCube, instead of the common Fat-Tree network, to build a special-purpose DML cluster, due to its parallel synchronization, RDMA friendliness, natural load balance, as well as low economical cost.
Shuai Wang 0028, Dan Li 0001, Jinkun Geng
INFOCOM1
2019 HiPower: A High-Performance RDMA Acceleration Solution for Distributed Transaction Processing
Runhua Zhang 0002, Jinkun Geng, Shuai Wang 0028, Kaihui Gao, Guowei Shen
NPC4
2018 BML: A High-performance, Low-cost Gradient Synchronization Algorithm for DML Training
abstract
In distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a new gradient synchronization algorithm with higher network performance and lower network cost than the current practice. BML runs on BCube network, instead of using the traditional Fat-Tree topology. BML algorithm is designed in such a way that, compared to the parameter server (PS) algorithm on a Fat-Tree network connecting the same number of server machines, BML achieves theoretically 1/k of the gradient synchronization time, with k/5 of switches (the typical number of k is 2∼4). Experiments of LeNet-5 and VGG-19 benchmarks on a testbed with 9 dual-GPU servers show that, BML reduces the job completion time of DML training by up to 56.4%.
Dan Li 0001, Jinkun Geng, Yanshu Wang, Shuai Wang 0028, Shutao Xia
NeurIPS6