Xinle Du

dblp:204/1946 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0001-8918-9580ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 15 · 4 first-author · 15 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Concord: Airtime-Aware Contention Control for Taming Tail Latency from Wi-Fi Frame Bursting
abstract
In congested Wi-Fi, sending less does not guarantee lower latency. Based on measurements on commodity Wi-Fi routers in the wild, sparse microflows can suffer bulk-like tail latency even at negligible load. This latency is driven by MAC-level contention dynamics rather than a flow's sending rate, rendering rate-based congestion control ineffective. We present Concord, a Wi-Fi MAC mechanism that makes burst airtime an explicit control signal and penalizes excessive medium holding. Concord operates entirely within the Wi-Fi driver, requires no flow classification, no client or protocol changes, and incurs only O(1) work per burst. Concord shows that controlling medium-holding time, rather than transmission rate, is key to tail latency in WLANs. With four saturated downlink contenders, Concord reduces the 99.9th-percentile enqueue-to-ACK latency of 100 B microflows from 298/461 ms (IEEE baseline / vendor bursting) to 42 ms without sacrificing bulk throughput. For interactive workloads (cloud gaming), it cuts 99.9th-percentile latency from 231/441 ms to 92 ms and reduces starvation by up to 10× versus the default IEEE stack.
Fengqian Guo, Sihao Miao, Xinle Du, Hancheng Lu
SIGCOMM4
2026 AutoRec: Accelerating Loss Recovery for Live Streaming in a Multi-Supplier Market
abstract
Due to the limited permissions for upgrading dual-side (i.e., server-side and client-side) loss tolerance schemes from the perspective of CDN vendors in a multi-supplier market, modern large-scale live streaming services are still using the automatic-repeat-request (ARQ) based paradigm for loss recovery, which only requires server-side modifications. In this paper, we first conduct a large-scale measurement study with up to 50 million live streams. We find that loss showsdynamicsand live streaming contains frequenton-off mode switchingin the wild. We further find that the recovery latency, enlarged by the ubiquitous retransmission loss, is a critical factor affecting live streaming’s client-side QoE (e.g., video freezing). We then propose an enhanced recovery mechanism called AutoRec, which can transform the disadvantages of on-off mode switching into an advantage for reducing loss recovery latency without any modifications on the client side. AutoRec allows users to customize overhead tolerance and recovery latency tolerance and adaptively adjusts strategies as the network environment changes to ensure that recovery latency meets user demands whenever possible while keeping overhead under control. We implement AutoRec upon QUIC and evaluate it via testbed and real-world commercial services deployments. The experimental results demonstrate the practicability and profitability of AutoRec.
Tong Li 0014, Bo Wu 0002, Fuyu Wang 0006, Jiuxiang Zhu, Haoyi Fang, Xinle Du, Ke Xu 0002
IEEE Trans. Netw.8
2025 FlexSpark: Robust and Efficient Multi-Device Collaborative Inference over Wireless Network
Yiyang Shao, Shuihai Hu, Xinle Du, Jingbin Zhou
APNet4
2025 PRED: Performance-oriented Random Early Detection for Consistently Stable Performance in Datacenters
Xinle Du, Tong Li 0014, Guangmeng Zhou, Zhuotao Liu, Hanlin Huang, Mowei Wang, Kun Tan 0002, Ke Xu 0002
NSDI1
2025 Secure Fault Localization in Path Aware Networking
abstract
Secure data forwarding is critical for users to meet their requirements. In this paper, we propose D3 (Demon Detector in Data Plane), a source-driven, secure fault localization mechanism, which empowers the source to localize faulty link in Path Aware Networking, thus circumventing faulty link to guarantee secure data forwarding. D3 utilizes the source to instruct the on-path routers, thus empowering it to detect whether the on-path routers forward the packet as expected. Compared with existing schemes that are difficult to be deployed in practice due to the heavy storage, computation, and communication overhead, D3 offloads most of the on-path router's storage and computation overhead, thus dramatically improving the deployment efficiency. Particularly, the length of the additional packet header in D3 is 2-5 times less than the state-of-the-art mechanisms, thus having a low communication overhead. Besides that, the destination in D3 could keep stateless processing, thus having backward compatibility and eliminating the opportunity for DoS attacks toward a stateful destination. The BMv2 and Barefoot Tofino hardware evaluations show that D3 could achieve high fault localization accuracy and process the packet at line rate.
Songtao Fu, Qi Li 0002, Xiaoliang Wang 0004, Su Yao, Xuewei Feng, Xinle Du, Kao Wan, Ke Xu 0002
IEEE Trans. Dependable Secur. Comput.7
2025 DiffECN: Differential ECN Marking for Datacenter Networks
abstract
ECN marking has been integrated into datacenter switches to enable high-throughput and low-latency transport. We observe that current marking schemes are coarse-grained: they blindly mark all flows when congestion occurs, causing large flows to occupy undeserved bandwidth and preventing newly arriving small flows from finishing quickly. In this paper, we propose DiffECN, a differential marking strategy that marks only the flows that are the culprits of congestion and protects the remaining flows from being limited. We have implemented it in the Barefoot Tofino switch and performed extensive evaluations via both physical testbed and large-scale simulations. The results show that DiffECN can restrain flows responsible for congestion successfully while providing desirable network performance. For instance, compared to the legacy way of ECN marking, DiffECN achieves up to 32.5% (40.1%) lower average (99th percentile) flow completion time (FCT) for small flows while delivering similar FCT for large flows under production workloads.
Hanlin Huang, Ke Xu 0002, Tong Li 0014, Zhuotao Liu, Xinle Du
IEEE Trans. Netw.5
2025 Revisiting Random Early Detection Tuning for High-Performance Datacenter Networks
abstract
Random Early Detection (RED) has been integrated into datacenter switches as a fundamental Active Queue Management (AQM) for decades. The accurate configuration of RED parameters is crucial to achieving high throughput and low latency. However, due to the highly dynamic nature of workloads in datacenter networks, maintaining consistently high performance with statically configured RED thresholds poses a challenge. Prior work applies reinforcement learning to predict proper thresholds, but their real-world deployment has been hindered by poor tail performance caused by instability. In this paper, we propose$\textsf {PRED}$, a novel system that enables automatic and stable RED parameter adjustment in response to traffic dynamics. Specifically, the system employs a Multiplicative-Increase Multiplicative-Decrease (MIMD) strategy to dynamically adapt to flow concurrency while utilizing an Additive-Increase Additive-Decrease (AIAD) mechanism to adapt to flow distribution. We perform extensive evaluations on our physical testbed and large-scale simulations. The results demonstrate that$\textsf {PRED}$can keep up with the real-time network dynamics generated by realistic workloads. For instance, compared with the static-threshold-based methods,$\textsf {PRED}$keeps 66% shorter switch queue length and obtains up to 80% lower Flow Completion Time (FCT). Compared with the state-of-the-art learning-based method,$\textsf {PRED}$reduces the tail FCT by 34%.
Tong Li 0014, Xinle Du, Guangmeng Zhou, Hanlin Huang, Zhuotao Liu, Mowei Wang, Kun Tan 0002, Ke Xu 0002
IEEE Trans. Netw.2
2024 Revisiting Congestion Control for WiFi Networks
abstract
WiFi networks, widely utilized by wireless devices, have become increasingly complex and congested environments, leading to noticeable delays, jitter, and throughput degradation for end-to-end network flows in today’s Internet. Through detailed experimental observations, we identified the TCP victim problem in WiFi, where TCP erroneously detects congestion and interacts with WiFi routers, resulting in a significant throughput decrease for certain hosts. In this paper, we introduce Cupid, a novel congestion control algorithm that relies on receiver-side WiFi physical layer measurements. Cupid accurately assesses congestion by measuring parameters such as airtime utilization, concurrency, and rates, allowing for precise rate adjustments. Our results demonstrate that whether employed alone or alongside other congestion controls, Cupid effectively mitigates the TCP victim problem. Furthermore, when used independently, Cupid can also reduce latency.
Xinle Du, Yiyang Shao, Wei Wang 0505, Shuihai Hu, Jingbin Zhou, Kun Tan 0003
APNet1
2024 Performant TCP over Wi-Fi Direct
abstract
Wi-Fi Direct has been serving a progressively wide range of applications such as device-to-device file sharing, face-to-face interactive gaming, and wireless projection. However, when TCP meets Wi-Fi Direct, we find that two independent control loops exist, i.e., the transport-layer control loop and the link-layer control loop. First, these functionally redundant loops result in spectrum inefficiency. Second, the lack of effective information interaction between layers results in local optimal. To tackle these issues, this paper proposes Wi-Fi Direct TCP (WDTCP), a performant TCP that provides a full protocol design of the acknowledgment de-redundancy and explicit-capacity-based congestion control. WDTCP tightly couples the two control loops by capturing the WiFi Direct’s key feature of one-hop communication. Evaluation results demonstrate that WDTCP can maximize bandwidth utilization while keeping low latency. For instance, compared to legacy TCP, WDTCP improves throughput by up to 49.2% and reduces average and 95th latency by up to 32.4% and 50.7%, respectively.
Hanlin Huang, Ke Xu 0002, Xinle Du, Yiyang Shao, Tong Li 0014
IWQoS3
2024 Toward Timeliness-Enhanced Loss Recovery for Large-Scale Live Streaming
abstract
Due to the limited permissions for upgrading dual-side (i.e., server-side and client-side) loss tolerance schemes from the perspective of CDN vendors in a multi-supplier market, modern large-scale live streaming services are still using the automatic-repeat-request (ARQ) based paradigm for loss recovery, which only requires server-side modifications. In this paper, we first conduct a large-scale measurement study with up to 50 million live streams. We find that loss shows dynamics and live streaming contains frequent on-off mode switching in the wild. We further find that the recovery latency, enlarged by the ubiquitous retransmission loss, is a critical factor affecting live streaming's client side QoE (e.g., video freezing). We then propose an enhanced recovery mechanism called AutoRec, which can transform the disadvantages of on-off mode switching into an advantage for reducing loss recovery latency without any modifications on the client side. AutoRec also adopts an online learning-based policy to fit the dynamics of loss, balancing the tradeoff between the recovery latency and the incurred overhead. We implement AutoRec upon QUIC and evaluate it via both testbed and real-world commercial services deployments. The experimental results demonstrate the practicability and profitability of AutoRec, in which the average times and duration of client-side video freezing can be lowered by 11.4% and 5.2%, respectively.
Bo Wu 0002, Tong Li 0014, Fuyu Wang 0006, Xinle Du, Ke Xu 0002
ACM Multimedia6
2024 Re-Architecting Buffer Management in Lossless Ethernet
abstract
Converged Ethernet employs Priority-based Flow Control (PFC) to provide a lossless network. However, issues caused by PFC, including victim flow, congestion spreading, and deadlock, impede its large-scale deployment in production systems. The fine-grained experimental observations on switch buffer occupancy find that the root cause of these performance problems is a mismatch of sending rates between end-to-end congestion control and hop-by-hop flow control. Resolving this mismatch requires the switch to provide an additional buffer, which is not supported by the classic dynamic threshold (DT) policy in current shared-buffer commercial switches. In this paper, we propose Selective-PFC (SPFC), a practical buffer management scheme that handles such mismatch. Specifically, SPFC incrementally modifies DT by proactively detecting port traffic and adjusting buffer allocation accordingly to trigger PFC PAUSE frames selectively. Extensive case studies demonstrate that SPFC can reduce the number of PFC PAUSEs on non-bursty ports by up to 69.0%, and reduce the average flow completion time by up to 83.5% for large victim flows.
Hanlin Huang, Xinle Du, Tong Li 0014, Ke Xu 0002, Mowei Wang, Huichen Dai
IEEE/ACM Trans. Netw.2
2024 Stable Byzantine Fault Tolerance in Wide Area Networks With Unreliable Links
abstract
With the increasing demand for blockchain technology in various industry sectors, there has been a growing interest in the Byzantine Fault Tolerance (BFT) consensus that is the backbone of most of these blockchains. However, many state-of-the-art algorithms that require reliable connections can only offer limited throughput in wide-area networks (WANs), where participants are connected over long distances and may experience unpredictable network failures. The partially-connected BFTs are designed for unreliable and highly dynamic networks yet impose exponential communication complexity. This paper proposes Stable Byzantine Fault Tolerance (SBFT), a BFT communication abstraction that can sustain high throughput and low latency in WAN. SBFT separates the leader from consensus in pipelined BFT consensus and uses an adaptive consensus mechanism to resist dynamic faulty links, maintaining consensus efficiency when network connectivity is high while adapting to dynamic networks with low connectivity. We implemented a prototype of SBFT and tested it on the WAN. The results demonstrate that SBFT has a throughput similar to HotStuff in a fault-free environment but can reduce about 80% of consensus latency. Besides, SBFT retains 40% of the original throughput when the link failure probability is 0.4, while the baseline HotStuff retains less than 40% when the link failure probability is only 0.1.
Sitong Ling, Zhuotao Liu, Qi Li 0002, Xinle Du, Ke Xu 0002
IEEE/ACM Trans. Netw.4
2023 R-AQM: Reverse ACK Active Queue Management in Multitenant Data Centers
abstract
TCP incast has become a practical problem for high-bandwidth, low-latency transmissions, resulting in throughput degradation of up to 90% and delays of hundreds of milliseconds, severely impacting application performance. However, in virtualized multi-tenant data centers, host-based advancements in the TCP stack are hard to deploy from the operators’ perspective. Operators only provide infrastructure in the form of virtual machines, in which only tenants can directly modify the end-host TCP stack. In this paper, we present R-AQM, a switch-powered reverse ACK active queue management (R-AQM) mechanism for enhancing ACK-clocking effects through assisting legacy TCP. Specifically, R-AQM proactively intercepts ACKs and paces the ACK-clocked in-flight data packets, preventing TCP from suffering incast collapse. We implement and evaluate R-AQM in NS-3 simulation and NetFPGA-based hardware switch. Both simulation and testbed results show that R-AQM greatly improves TCP performance under heavy incast workloads by significantly lowering packet loss rate, reducing retransmission timeouts, and supporting 16 times (i.e., 60 to 1000) more senders. Meanwhile, the forward queuing delays are also reduced by 4.6 times.
Xinle Du, Ke Xu 0002, Lei Xu 0019, Kai Zheng 0003, Meng Shen 0001, Bo Wu 0002, Tong Li 0014
IEEE/ACM Trans. Netw.1
2023 MASK: Practical Source and Path Verification Based on Multi-AS-Key
abstract
The source and path verification in Path-Aware Networking considers the two critical issues: (1) end hosts could verify that the network follows their forwarding decisions, and (2) both on-path routers and destination host could authenticate the source of packets and filter the malicious traffic. Unfortunately, the state-of-the-art mechanisms require heavy communication overhead in the network and computation overhead in the router; moreover, it is difficult to meet the dynamic requirements of the end host. We propose a user-driven mechanism, source and path verification based on Multi-AS-Key (MASK). MASK decreases the communication overhead by a short additional packet header and reduces the computation overhead by separating the control and data plane in terms of the cryptographic operation. Furthermore, it utilizes the stateful user to instruct the stateless routers to process the packet with a user-driven policy, thus satisfying the user’s requirements such as detecting the packet drop and replay attack. With the plausible design, the communication overhead for realistic path lengths is 1/2 to 1/10 compared with the state-of-the-art mechanisms. We implement MASK in the BMv2 environment and commodity Barefoot Tofino programmable switch, testify that MASK introduces significantly less overhead than the state-of-the-art mechanisms, and demonstrate that MASK could achieve the verification in the programmable switch at line rate.
Songtao Fu, Qi Li 0002, Xiaoliang Wang 0004, Su Yao, Yangfei Guo, Xinle Du, Ke Xu 0002
IEEE/ACM Trans. Netw.7
2022 D3: Lightweight Secure Fault Localization in Edge Cloud
abstract
In pursuit of high-performance applications, the cloud is moving out of the data center and towards the edge. Secure data forwarding is critical for the users between the edge and the remote cloud. In this paper, we propose D3 (Demon Detector in Data Plane), a lightweight, secure fault localization mechanism, which can enable the users in the edge cloud to localize faulty links and thus avoid the faulty links to guarantee secure data forwarding along the path to the remote cloud. D3 utilizes the user to instruct the transit routers, thus empowering the user to detect whether the transit routers forward the packet as expected. Compared with existing schemes that are difficult to be deployed in practice due to the incurred heavy storage, computation, and communication overhead, D3 offloads most of the transit router’s storage and computation overhead, thus dramatically improving the deployment efficiency. Particularly, the length of the additional packet header in D3 is 2-5 times less than the state-of-the-art mechanisms, and the extra control packet overhead is ten times less while keeping a little constant storage overhead in the data plane. The evaluations in BMv2 and Barefoot Tofino hardware show that D3 could achieve high fault localization accuracy and efficiency.
Songtao Fu, Qi Li 0002, Xiaoliang Wang 0004, Su Yao, Xuewei Feng, Xinle Du, Kao Wan, Ke Xu 0002
ICDCS7
2022 WIP: When RDMA Meets Wireless
abstract
The emerging applications including AR/VR inter-active gaming, ultra-high-definition live streaming, 4K wireless projection, Metaverse, etc. imply the demand for ultra-low latency and ultra-high bandwidth wireless transmission. The legacy kernel TCP stack is not fully satisfactory because it induces the CPU bottleneck on hosts. In this paper, we propose Wireless-RDMA (W-RDMA) that enables RDMA in wireless networks to tackle the CPU bottleneck issue on wireless hosts. The feasibility of W-RDMA is demonstrated through testbed experiments. Technical challenges and future opportunities are further discussed. We believe it is a small but crucial step for enabling RDMA for wireless transmission.
Tong Li 0014, Ke Xu 0002, Hanlin Huang, Xinle Du, Kai Zheng 0003
WoWMoM4
2021 R-AQM: Reverse ACK Active Queue Management in Multi-tenant Data Centers
abstract
TCP incast has become a practical problem for high-bandwidth, low-latency transmissions, resulting in throughput degradation of up to 90% and delays of hundreds of milliseconds, severely impacting application performance. However, in virtualized multi-tenant data centers, host-based advancements in the TCP stack are hard to deploy from the operators perspective. Operators only provide infrastructure in the form of virtual machines, in which only tenants can directly modify the end-host TCP stack. In this paper, we present R-AQM, a switch-powered reverse ACK active queue management (R-AQM) mechanism for enhancing ACK-clocking effects through assisting legacy TCP. Specifically, R-AQM proactively intercepts ACKs and paces the ACK-clocked in-flight data packets, preventing TCP from suffering incast collapse. We implement and evaluate R-AQM in NS-3 simulation and NetFPGA-based hardware switch. Both simulation and testbed results show that R-AQM greatly improves TCP performance under heavy incast workloads by significantly lowering packet loss rate, reducing retransmission timeouts, and supporting 16 times (i.e., 60 → 1000) more senders. Meanwhile, the forward queuing delays are also reduced by 4.6 times.
Xinle Du, Tong Li 0014, Lei Xu 0019, Kai Zheng 0003, Meng Shen 0001, Bo Wu 0002, Ke Xu 0002
ICNP1
2021 MASK: Practical Source and Path Verification based on Multi-AS-Key
abstract
The source and path verification in path-aware Internet consider the two critical issues: (1) end hosts could verify that their forwarding decisions followed by the network, (2) both intermediate routers and destination host could authenticate the source of packets and filter the malicious traffic. Unfortunately, the current verification mechanism requires validation operations in each router on the path in an inter-domain environment, thus requiring high communication and computation overhead, reducing its usefulness; besides, it is also difficult to meet the dynamic requirements of the end host. Ideally, the verification should be secure and provide the customized capability to meet the end host’s requirements. We propose a new mechanism called source and path verification based on Multi-AS-Key (MASK). Instead of each packet verified and marked at each router on the path, MASK improves the verification by empowering the end hosts to instruct the routers to achieve the verification, thus decreasing the router’s overhead while ensuring security performance to meet the end host’s requirements. With the plausible design, the communication overhead for realistic path lengths is 3–8 times smaller than the state-of-the-art mechanisms. The computation overhead in the routers is 2-5 times smaller. We implement our design in the BMv2 environment and commodity Barefoot Tofino programmable switch, demonstrating that MASK introduces significantly less overhead than the existing mechanisms.
Songtao Fu, Ke Xu 0002, Qi Li 0002, Xiaoliang Wang 0004, Su Yao, Yangfei Guo, Xinle Du
IWQoS7
2019 SmartCrowd: Decentralized and Automated Incentives for Distributed IoT System Detection
abstract
Internet of Things (IoT) devices achieve the rapid development and have been widely deployed recently. Meanwhile, inherent vulnerabilities of IoT systems (including firmware and software) have been continually uncovered and thus the systems are always exposed to various attacks. The root cause of the issue is that IoT systems always have design flaws and implementation bugs. In particular, the released systems (e.g., by third-party marketplaces and IoT vendors) may be maliciously repackaged with malware. Unfortunately, IoT consumers are not able to effectively capture such vulnerabilities because of the limited detection capabilities. In this paper, we propose SmartCrowd, a blockchain-based platform that aims to outsource security detection of IoT systems to distributed detectors with strong detection incentives. SmartCrowd enables built-in accountability for IoT providers and authoritative references of detection results for IoT consumers. By building smart contracts, we can incentivize the efficient and high-coverage security detection of IoT systems, while providing decentralized and automated incentives for both IoT providers releasing secure IoT systems and detectors uncovering vulnerabilities. We present the security and theoretical analysis that demonstrates the security of SmartCrowd and the incentives for participators. We prototype SmartCrowd by using Ethereum and the experimental results show that SmartCrowd has both technical feasibility and financial benefits, which can be applied to build a secure IoT ecosystem.
Bo Wu 0002, Ke Xu 0002, Qi Li 0002, Zhuotao Liu, Yih-Chun Hu, Xinle Du, Bingyang Liu, Shoushou Ren
ICDCS7