VLDB 2026 Research / reviewers in the wild / expert
Jianer Zhou
dblp:137/5221
· DBLP profile ↗
26ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-0639-9978ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 20 · 7 first-author · 16 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling Low‑Altitude 5G Performance: Linking Key Influencing Factors with UAV Flight Parameters
Jianer Zhou, Xiaoyong Ni, Ke Luo 0001, Zhenyu Li 0001, Xiaofeng Tao 0001, Weichao Li 0001 |
SIGCOMM | 2 |
| 2026 | Traffic-Aware Design for Multi-Dimensional Lookup and Forwarding: From IP Routing to Packet ClassificationabstractPacket processing in modern routers and switches relies on rule matching, primarily performed by two core modules: IP prefix lookup for next-hop determination and packet classification for multi-field policy enforcement. However, most existing algorithms are rule-centric and assume uniform rule access, overlooking the highly skewed nature of real-world network traffic. Such mismatch between static rule organization and dynamic traffic behavior leads to inefficiency in both lookup and classification. To address this limitation, we propose a Traffic-aware Lookup and Forwarding (TLF) framework that leverages traffic measurement with lookup operations, enabling online adaptation to dynamic traffic patterns and frequent rule updates. Experimental results demonstrate that TLF provides 1.04×–3.37× speedups for lookup and forwarding over state-of-the-art algorithms, while substantially reducing both memory overhead and construction time. Furthermore, integrating TLF into Vector Packet Processor (VPP) and Open vSwitch (OVS) results in throughput improvements of 2.61× and 4.88×, respectively. Xinyi Zhang 0004, Qianrui Qiu, Peng He 0003, Guangxing Zhang, Luyiyun Li, Jianer Zhou, Kavé Salamatian, Gaogang Xie |
IEEE Trans. Netw. | 7 |
| 2025 | From Limited Resources to Powerful Insights: Empowering Low-Cost Cameras for Efficient Retrospective QueryingabstractUploading videos from low-cost cameras to the cloud for retrospective analysis presents challenges in privacy, network, and computation. To address these issues and achieve low latency, we propose READY, a novel client-cloud collaborative system. READY aims to enhance the quality of uploaded frames by selectively uploading only the frames relevant to queries. To achieve this, READY establishes an index during video capture, recording object categories and probabilities for each frame. READY adopts an innovative semi-supervised approach for frame indexing, wherein frames are indexed through a continuously updated feature distribution space constructed by k-nearest neighbors (KNN). This enables resource-constrained low-cost cameras to independently establish long-term frame indexes. Additionally, READY utilizes progressively improving operators (lightweight classification models) dispatched by the cloud to optimize the upload order of frames, prioritizing positive frames. By sharing the backbone of low-performance operators, high-performance operators can be efficiently executed, significantly enhancing the camera’s frame processing capability. The established frame index also enables efficient multiple consecutive queries on different classes. Over 110 h of diverse queries across 11 videos, READY outperformed competing alternative designs by achieving an average response time of 67.8% and reducing the proportion of uploaded videos by an average of 79.8% (compared to CloudOnly). Qiaodi Wen, Jianer Zhou, Ziqi Luo, Gareth Tyson, Weichao Li 0001, Jinfan Wang |
IEEE Internet Things J. | 2 |
| 2025 | Efficient Data Center Network Monitoring and Troubleshooting With LMon: Leveraging ECMP Hashing Linearity and Lightweight ProbingabstractNetwork performance monitoring and troubleshooting are crucial yet challenging tasks in datacenter management. Despite the numerous solutions that have been proposed in recent years, their efforts are often hindered by high costs and unreliable failure localization, making it difficult to deploy them in real-world environments. In this paper, we presentLMon, a highly reliable and efficient system for monitoring and troubleshooting in datacenter networks. LMon utilizes the characteristic of ECMP hashing linearity to control probe packet routing, enabling the monitoring of targeted paths without any modification of underlying protocols and devices. Additionally, LMon leverages a lightweight probing technique to reduce monitoring overhead, as well as integrates the improved LASSO regression and hypothesis testing for higher accuracy and faster processing in link failure localization. We evaluate the performance of LMon in our testing environment. Compared to the monitoring system Pingmesh, LMon generates only one-third probes while maintaining 99% accuracy and 1% false negatives. Qinglin Xun, Weichao Li 0001, Jianer Zhou, Jingpu Duan, Yi Wang 0004, Xiaofeng Tao 0001, Jinbei Zhang |
IEEE Trans. Netw. | 3 |
| 2025 | Safety in DRL-Based Congestion Control: A Framework Empowered by Expert RefinementabstractDeep reinforcement learning (DRL) has been used in congestion control algorithms (CCAs) for its ability to adapt to different network environments. However, its effectiveness is often hindered by the limited availability of training data and constrained training scales. While it has been proved that combining rule-based (expert) CCAs as a guide for DRL (namely hybrid CCAs) can address this limitation, we show through experimental measurements that rule-based CCAs potentially restrict action exploration of DRL models and may cause the DRL models to overly rely on them for higher reward gains. To address this gap, this paper proposes Marten, a framework that improves the effectiveness of rule-based CCAs for DRL. Marten’s key innovations include an entropy-based dynamic exploration scheme that expands the exploration of DRL, and a reward adjustment scheme to prevent the DRL models’ over-reliance on experts in hybrid CCAs. We have implemented Marten in both simulation platform OpenAI Gym and deployment platform QUIC. Experimental results in both emulated and production networks demonstrate Marten can improve throughput by 0.31% and reduce latency by 12.69% on average compared to the state-of-the-art hybrid CCAs. Compared to BBR, Marten achieves a 2.79% increase in throughput and an 11.73% reduction in latency on average. Jianer Zhou, Zhiyuan Pan, Zhenyu Li 0001, Gareth Tyson, Weichao Li 0001, Xinyi Qiu, Xinyi Zhang 0004, Gaogang Xie |
IEEE Trans. Netw. | 1 |
| 2024 | Patronum: In-network Volumetric DDoS Detection and Mitigation with Programmable Switches
Penglai Cui, Jianer Zhou, Peng He 0003, Yanbiao Li 0001, Zhenyu Li 0001, Gaogang Xie |
ESORICS (4) | 5 |
| 2024 | Roundabout: Solving PFC Deadlocks With Distributed Detection and Buffer CollaborationabstractRDMA over Converged Ethernet (RoCEv2) employs Priority-based Flow Control (PFC) for a lossless fabric to maintain high performance. However, PFC can cause Deadlocks, which pauses traffic and potentially leads to severe exceptions for applications. Existing solutions solve deadlocks at a considerable cost, resulting in degradation of end-to-end network performance.We present Roundabout, a data plane scheme designed to detect and resolve deadlocks with minimal side effects. We first analyze how switches in different states contribute to deadlocks. Based on the analysis, we design an election-based distributed detection scheme that efficiently and robustly identifies deadlocks. By exploiting buffer configuration redundancy, we develop an innetwork collaborative packet scheduling scheme that forwards deadlocked packets to their destinations in a lossless manner, facilitating natural deadlock resolution. Additionally, we implement a barrier mechanism to ensure in-order packet delivery to the receiver. Both analysis and experiments demonstrate that Roundabout effectively detects and resolves deadlocks while minimizing side effects to the network, making it an ideal enhancement for PFC switches. Chengjun Jia, Jianer Zhou, Yanbiao Li 0001, Zhenyu Li 0001, Gaogang Xie |
ICNP | 6 |
| 2024 | VAKY: Scheduling In-network Aggregation for Distributed Deep Training AccelerationabstractDistributed machine learning (DML) has recently experienced widespread application. A major performance bottleneck is the costly communication for gradients synchronization. Recently, researchers have explored the use of programmable switches for in-network synchronous aggregation of gradients to mitigate the communication overhead. Nevertheless, the performance of in-network synchronous aggregation is significantly impacted by the stragglers. Unfortunately, the schedulers in existing DML systems are no longer effective in dealing with stragglers because of the ignorance of the aggregation progress that is offloaded from the parameter servers to the programmable switches. To address this gap, this paper presents VAKY, an adaptive scheduler specifically designed for in-network aggregation. At the heart of VAKY is the variable K-block sync method, where the aggregators stop waiting for updates from more workers once having received updates from the fastest K workers for each block of gradients. We propose an efficient solution that can dynamically choose the optimal values of K during the training process, in order to minimize the expected training completion time. We have integrated VAKY into PyTorch, and our experiments show that compared to the state-of-the-art in-network aggregation systems, VAKY improves the aggregation throughput by up to $40 \%$ and reduces the training time by $25 \%$. Penglai Cui, Jianer Zhou, Qinghua Wu 0004, Zhaohua Wang, Zhenyu Li 0001 |
ICPADS | 3 |
| 2023 | Disco: A Framework for Dynamic Selection of Multipath Congestion Control AlgorithmsabstractMany mobile devices are usually equipped with multiple interfaces, providing the opportunity of using multipath transport protocols such as Multipath TCP (MPTCP) to boost performance. The multipath congestion control algorithm (CCA) in MPTCP plays a vital role in achieving high performance and multipath fairness in mobile environments where the paths are often heterogeneous and dynamic. Such environments are very challenging for existing one-size-fits-all CCAs to achieve high performance while ensuring multipath fairness. In this paper, we present a novel framework, Disco, to dynamically select the most appropriate CCAs for MPTCP subflows at runtime according to the perceived network condition. Extensive experiments show that compared with existing multipath CCAs, the proposed solution can improve the average throughput by 19% – 25% and reduce the average queuing delay by up to 21 % while it barely does harm to multipath fairness. Furong Yang, Zhenyu Li 0001, Jianer Zhou, Xinyi Zhang 0004, Qinghua Wu 0004, Giovanni Pau 0001, Gaogang Xie |
ICNP | 3 |
| 2023 | Hawkeye: A Dynamic and Stateless Multicast Mechanism with Deep Reinforcement LearningabstractMulticast traffic is growing rapidly due to the development of multimedia streaming. Lately, stateless multicast protocols, such as BIER, have been proposed to solve the excessive routing states problem of traditional multicast protocols. However, the high complexity of multicast tree computation and the limited scalability for concurrent requests still pose daunting challenges, especially under dynamic group membership. In this paper, we propose Hawkeye, a dynamic and stateless multicast mechanism with deep reinforcement learning (DRL) approach. For real-time responses to multicast requests, we leverage DRL enhanced by a temporal convolutional network (TCN) to model the sequential feature of dynamic group membership and thus is able to build multicast trees proactively for upcoming requests. Moreover, an innovative source aggregation mechanism is designed to help the DRL agent converge when faced with a large amount of multicast requests, and relieve ingress routers from excessive routing states. Evaluation with real-world topologies and multicast requests demonstrates that Hawkeye adapts well to dynamic multicast: it reduces the variation of path latency by up to 89.5% with less than 12% additional bandwidth consumption compared with the theoretical optimum. Lie Lu, Qing Li 0006, Dan Zhao 0003, Yuan Yang 0001, Zeyu Luan, Jianer Zhou, Yong Jiang 0001, Mingwei Xu 0001 |
INFOCOM | 6 |
| 2023 | Marten: A Built-in Security DRL-Based Congestion Control Framework by Polishing the ExpertabstractDeep reinforcement learning (DRL) has been proved to be an effective method to improve the congestion control algorithms (CCAs). However, the lack of training data and training scale affect the effectiveness of DRL model. Combining rule-based CCAs (such as BBR) as a guide for DRL is an effective way to improve learning-based CCAs. By experiment measurement, we find that the rule-based CCAs limit the action exploration and even cause DRL’s excessive dependence to gain higher DRL’s reward gain. To overcome the constraints, we propose Marten, a framework which improves the effectiveness of rule-based CCAs for DRL. Marten uses entropy as the degree of exploration and uses it to expand the exploration of DRL. Furthermore, Marten introduces the shielding mechanism to avoid wrong DRL actions. We have implemented Marten in both simulation platform OpenAI Gym and deployment platform QUIC. The experimental results in production network demonstrate Marten can improve throughput by 0.36% and reduce latency by 14.89% on average compared with Eagle, and improve throughput by 2.79% and reduce latency by 11.73% on average compared with BBR. Zhiyuan Pan, Jianer Zhou, Xinyi Qiu, Weichao Li 0001 |
INFOCOM | 2 |
| 2023 | Enabling Reliable and Efficient Performance Monitoring and Troubleshooting in Datacenter NetworksabstractNetwork performance monitoring and troubleshooting is a crucial but challenging task in datacenter management. Despite the numerous solutions that have been proposed in recent years, their efforts are often hindered by high costs and unreliable fault localization, making it difficult to deploy them in real-world environments. In this paper, we present LMon, a highly reliable and efficient system for monitoring and troubleshooting in datacenter networks. LMon utilizes the characteristic of ECMP hashing linearity to control the packet routing without any modification of the underlying protocols. Additionally, LMon leverages a lightweight probing technique to reduce monitoring overhead. Furthermore, the system integrates improved LASSO regression and statistical hypothesis testing for higher accuracy and faster processing in link failure localization. The effectiveness of LMon is demonstrated through its implementation and evaluation in ns-3 simulation. The results validate the reliability and efficiency of the system, making it a promising option for ensuring long-term network maintenance in datacenters. Qinglin Xun, Weichao Li 0001, Haorui Guo, Qianyi Huang, Jianer Zhou, Jingpu Duan, Yi Wang 0004, Jinbei Zhang |
IWQoS | 5 |
| 2023 | A Scalable Asynchronous Traffic Shaping Mechanism for TSN with Time Slot and PollingabstractThe IEEE 802.1 Time Sensitive Networking (TSN) task group is devoted to improving deterministic delay during data communication. To schedule traffic in TSN, Asynchronous Traffic Shaping (ATS) has been introduced to guarantee bounded maximum delays without complicated time synchronization, and Urgency Based Scheduler (UBS) becomes the default implementation for ATS. However, due to the traffic fluctuation, UBS cannot achieve best delay bounds if some parameters are not configured appropriately in real time. What is more, UBS needs to consumes excessive butter resources, which prevents TSN from being deployed in real world. To solve these problems, we propose Time Sensitive Queuing (TSQ), a novel ATS mechanism. TSQ lowers the difficulty of TSN deployment by removing the parameter that need to be configured in real time. In addition, to reduce the butter consumption, TSQ applies time-slot-based packet allocation mechanism as the enqueue strategy, and time-slot-based packet polling mechanism as the dequeue strategy, respectively. Our Network Calculus analysis shows that TSQ can provide bounded delays and up to 40% butter resource reduction compared to UBS. The extensive experiments implemented in ns-3 show that TSQ consumes 33% less butter resource compared to UBS. Haorui Guo, Weichao Li 0001, Jiashuo Lin, Jianer Zhou, Qinglin Xun, Shuangping Zhan, Yi Wang 0004, Qingsha Cheng |
NOMS | 4 |
| 2023 | Cable: A framework for accelerating 5G UPF based on eBPF
Jianer Zhou, Zengxie Ma, Weijian Tu, Xinyi Qiu, Jingpu Duan, Zhenyu Li 0001, Qing Li 0006, Xinyi Zhang 0004, Weichao Li 0001 |
Comput. Networks | 1 |
| 2023 | DiffTREAT: Differentiated Traffic Scheduling Based on RNN in Data CentersabstractTransmission schemes in data centers are supposed to accurately distinguish flow types for different scheduling. However, prior efforts failed to meet the needs at all levels in a cost-effectively way. Nor the existing schemes proved applicable to all the diverse scenarios or dynamic traffic patterns. Therefore, we proposedDifferentiatedTraffic schEduling in dAta cenTers (DiffTREAT) based the Recurrent Neural Network (RNN), aiming to simplify the transmission in the dynamic and diverse network scenarios. First, DiffTREAT utilizes deep learning methods for traffic classification and flow size prediction. Second, according to the classified results of flows, DiffTREAT adopts multilevel priority queues to ensure the preferential transmission of latency-sensitive flows while optimizing the overall average flow completion time (FCT). Third, DiffTREAT employs the network cache to increase the capacity of data center networks (DCN), which effectively fights against the traffic burst and improves the throughput of latency-insensitive flows. DiffTREAT has been tested in different topologies in the contexts of diverse network loads and real-world workloads. Experiment results showed that compared with state-of-the-art schemes, DiffTREAT yielded both the lower average flow completion time for latency-sensitive flows and the higher throughput for latency-insensitive flow. Ziqi Wei 0004, Qing Li 0006, Keke Zhu, Jianer Zhou, Longhao Zou, Yong Jiang 0001, Xi Xiao 0001 |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | A Machine Learning-Based Framework for Dynamic Selection of Congestion Control AlgorithmsabstractMost congestion control algorithms (CCAs) are designed for specific network environments. As such, there is no known algorithm that achieves uniformly good performance in all scenarios for all flows. Rather than devising a one-size-fits-all algorithm (which is a likely impossible task), we propose a system to dynamically switch between the most suitable CCAs for specific flows in specific environments. This raises a number of challenges, which we address through the design and implementation of Antelope, a system that can dynamically reconfigure the stack to use the most suitable CCA for individual flows. We build a machine learning model to learn which algorithm works best for individual conditions and implement kernel-level support for dynamically switching between CCAs. The framework also takes application requirements of performance into consideration to fine-tune the selection based on application-layer needs. Moreover, to reduce the overhead introduced by machine learning on individual front-end servers, we (optionally) implement the CCA selection process in the cloud, which allows the share of models and the selection among front-end servers. We have implemented Antelope in Linux, and evaluated it in both emulated and production networks. The results demonstrate the effectiveness of Antelope via dynamic adjusting the CCAs for individual flows. Specifically, Antelope achieves an average 16% improvement in throughput compared with BBR, and an average 19% improvement in throughput and 10% reduction in delay compared with CUBIC. Jianer Zhou, Xinyi Qiu, Zhenyu Li 0001, Qing Li 0006, Gareth Tyson, Jingpu Duan, Yi Wang 0004, Qinghua Wu 0004 |
IEEE/ACM Trans. Netw. | 1 |
| 2022 | APS: Adaptive Packet Sizing for Efficient End-to-End Network TransmissionabstractMuch effort has been devoted to improving the performance of network transmission. Yet, the impact of packet size which is limited by the 1500-byte maximum transmission unit (MTU) has not received adequate attention. Through comprehensive experiments, we find that jumbo frames which are commonly used as an alternate do not always yield the best performance under different transmission situations.In this paper, we elaborate on the limitations of the regular and jumbo frames and analyze how packet sizes affect network performance. Based on these, we present Adaptively Packet Sizing (APS), a dynamic packet size adjustment method that can be easily integrated into existing window-based congestion control algorithms. APS utilizes a machine learning method to predict the optimal packet size, which can minimize flow completion time (FCT) according to the instantaneous network condition. Besides, a packet size based priority mechanism is proposed to further improve the performance. We implement APS in both simulation and testbed environments. APS reduces the FCT by up to 50% and gains better performance in scenarios with various loss rates. Feixue Han, Qing Li 0006, Jianer Zhou, Hong Xu 0001, Yong Jiang 0001 |
IWQoS | 3 |
| 2022 | ADRIoT: An Edge-Assisted Anomaly Detection Framework Against IoT-Based Network AttacksabstractInternet of Things (IoT) has entered a stage of rapid development and increasing deployment. Meanwhile, these low-power devices typically cannot support complex security mechanisms and, thus, are highly susceptible to malware. This article proposes ADRIoT, an anomaly detection framework for IoT networks, which leverages edge computing to uncover potential threats. An edge is empowered with an anomaly detection module, which consists of a traffic capturer, a traffic preprocessor, and a collection of anomaly detectors dedicated to each type of device. Each detector is constructed by an LSTM autoencoder in an unsupervised manner that requires no labeled attack data and is able to handle emerging zero-day attacks. When a device connects to the edge, the edge will fetch the corresponding detector from the cloud and execute it locally. Another problem is the resource constraint of a single edge device like a home router hinders the deployment of such a detection module. To mitigate this problem, we design a multiedge collaborative mechanism that integrates the resource of multiple edges in a local network to increase the overall load capacity. The evaluation demonstrates that ADRIoT can detect various IoT-based attacks effectively and efficiently, showing that ADRIoT can feasibly help build a more secure IoT environment. Ruoyu Li 0003, Qing Li 0006, Jianer Zhou, Yong Jiang 0001 |
IEEE Internet Things J. | 3 |
| 2022 | Enabling In-Network Floating-Point Arithmetic for Efficient Computation OffloadingabstractProgrammable switches are recently used for accelerating data-intensive distributed applications. Some computational tasks, traditionally performed on servers in data centers, are offloaded into the network on programmable switches. These tasks may require the support of on-the-fly floating-point operations. Unfortunately, programmable switches are restricted to simple integer arithmetic operations. Existing systems circumvent this restriction by converting floats to integers or relying on local CPUs of switches, incurring extra processing delayed and accuracy loss. To address this gap, we propose NetFC, a table-lookup method to achieve on-the-fly in-network floating-point arithmetic operations nearly without accuracy loss. Specifically, NetFC utilizes logarithm projection and transformation to convert the original huge table enumerating all operands and results into several much smaller tables that can fit into the data plane of programmable switches. To cope with the table inflation problem on 32-bit floats, we also propose an approximation method that further breaks the large tables into smaller ones. In addition, NetFC leverages two optimizations to improve accuracy and reduce on-chip memory consumption. We use both synthetic and real-life datasets to evaluate NetFC. The experimental results show that the average accuracy of NetFC is above 99.9% with only 448KB memory consumption for 16-bit floats and 99.1% with 496KB memory consumption for 32-bit floats. Furthermore, we integrate NetFC into two distributed applications and two in-network telemetry systems to show its effectiveness in further improving the performance. Penglai Cui, Zhenyu Li 0001, Penghao Zhang, Tianhao Miao, Jianer Zhou, Hongtao Guan, Gaogang Xie |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | Antelope: A Framework for Dynamic Selection of Congestion Control AlgorithmsabstractMost congestion control mechanisms are designed for specific network environments. Hence, there is no known algorithm that achieves uniformly good performance in all scenarios for all flows. Rather than devising such a one-size-fits-all algorithm, we propose a system to dynamically switch between the most suitable congestion control mechanisms for specific flows in specific environments. This raises a number of challenges, which we address through the design and implementation of Antelope, a system that can dynamically reconfigure to use the most suitable congestion control mechanism for an individual flow. We build a machine learning approach to learn which algorithm works best for individual conditions and implement kernel-level support for dynamically adjusting congestion control algorithms. We have implemented Antelope in Linux, and evaluated it in both emulated and production networks. We show that in WAN, DCN, and cellular networks, Antelope achieves an average 16% improvement in throughput compared with BBR; compared with Cubic, Antelope achieves an average 19% improvement in throughput and 10% reduction in delay. Jianer Zhou, Xinyi Qiu, Zhenyu Li 0001, Gareth Tyson, Qing Li 0006, Jingpu Duan, Yi Wang 0004 |
ICNP | 1 |
| 2021 | LAFS: Learning-Based Application-Agnostic Flow Scheduling for DatacentersabstractMany cloud applications in modern datacenters have very demanding latency requirements, making flow completion time (FCT) an important metric for evaluating the network performance. Existing network flow scheduling methods either base on pre-known information or have poor performance. Therefore, we present LAFS, an efficient learning-based flow scheduling approach which minimizes the FCT with estimated information of flows. LAFS combines system call monitoring and learning methods to learn the flow size and implements the Shortest Remaining Processing Time (SRPT) principle with in-network priorities. Moreover, LAFS adopts flowlets to alleviate the packets disorder problem in fine-grained flow scheduling. Our theoretical analysis and extensive simulations show that LAFS is a practical design and significantly outperforms other information-agnostic designs like DCTCP and PIAS under diverse workloads. Feixue Han, Qing Li 0006, Keke Zhu, Jianer Zhou, Yong Jiang 0001, Zhuyun Qi, Fuliang Li |
IPCCC | 4 |
| 2019 | TCP Stalls at the Server Side: Measurement and MitigationabstractTCP is an important factor affecting user-perceived performance of Internet applications. Diagnosing the causes behind TCP performance issues in the wild is essential for better understanding the current shortcomings in TCP. This paper presents a TCP flow performance analysis framework that classifies causes of TCP stalls. The framework forms the basis of a tool that we use to analyze packet-level traces of three services (cloud storage, software download, and web search) deployed by a popular service provider. We find that as many as 20% of the flows are stalled for half of their lifetime. Network-related causes, especially timeout retransmissions, dominate the stalls. A breakdown of the causes for timeout retransmission stalls reveals that double retransmission and tail retransmission are among the top contributors. The importance of these causes depends however on the specific service. Based on these observations, we propose smart-retransmission time out (S-RTO), a mechanism that mitigates timeout retransmission stalls through careful and gentle aggression for retransmission. S-RTO is evaluated in a controlled network and also in a production network. The results consistently show that it is effective at improving TCP performance, especially for short flows. Jianer Zhou, Zhenyu Li 0001, Qinghua Wu 0004, Peter Steenkiste, Steve Uhlig, Jun Li 0002, Gaogang Xie |
IEEE/ACM Trans. Netw. | 1 |
| 2016 | An Empirical Analysis of a Large-scale Mobile Cloud Storage Service
Zhenyu Li 0001, Xiaohui Wang 0012, Ningjing Huang, Mohamed Ali Kâafar, Zhenhua Li 0001, Jianer Zhou, Gaogang Xie, Peter Steenkiste |
Internet Measurement Conference | 6 |
| 2015 | Demystifying and mitigating TCP stalls at the server sideabstractTCP is an important factor affecting user-perceived performance of Internet applications. Diagnosing the causes behind TCP performance issues in the wild is essential for better understanding the current shortcomings in TCP. This paper presents a TCP flow performance analysis framework that classifies causes of TCP stalls. The framework forms the basis of a tool that is publicly available to the research community. We use our tool to analyze packet-level traces of three services (cloud storage, software download and web search) deployed by a popular Chinese service provider. We find that as many as 20% of the flows are stalled for half of their lifetime. Network-related causes, especially timeout retransmission, dominate the stalls. A breakdown of the causes for timeout retransmission stalls reveals that double retransmission and tail retransmission are among the top contributors. The importance of these causes depends however on the specific service. We also propose S-RTO, a mechanism that mitigates timeout retransmission stalls. S-RTO has been deployed on production front-end servers and results show that it is effective at improving TCP performance, especially for short flows. Jianer Zhou, Qinghua Wu 0004, Zhenyu Li 0001, Steve Uhlig, Peter Steenkiste, Gaogang Xie |
CoNEXT | 1 |
| 2015 | A proactive transport mechanism with Explicit Congestion Notification for NDNabstractNamed Data Networking (NDN) shifts the communication paradigm from the quest of where the content is to what content is to be consumed. In such a new Internet architecture, transmission control mechanisms are of particular importance and have to be carefully designed to enable efficient data transmission. Existing work advocates the use of TCP-like reactive mechanisms for NDN transmission control. In this paper, we show that the statefull and adaptive forwarding properties of NDN makes proactive and efficient mechanisms for transmission control possible. We achieve this by using Explicit Congestion Notifications (ECN), which explicitly notify content consumers about network conditions through the communication path. Specifically, we propose an ECN-based proactive interest-sending rate control mechanism, which aims to achieve a high link utilisation for fast data transmission as well as a low packet dropping rate. To have a globally optimal data transmission, we further propose a smart forwarding mechanism, which locally utilises network-wide information to select the forwarding paths for individual flows. Extensive packet-level simulations in ndnSIM demonstrate that the ECN-based approach, coupled with smart forwarding, outperforms TCP-like reactive mechanisms in terms of link utilisation, packet dropping rate and flow completion time. Jianer Zhou, Qinghua Wu 0004, Zhenyu Li 0001, Mohamed Ali Kâafar, Gaogang Xie |
ICC | 1 |
| 2015 | The role of innovation in inventory turnover performance
Hsiao-Hui Lee, Jianer Zhou, Po-Hsuan Hsu |
Decis. Support Syst. | 2 |