Fengyuan Ren

dblp:05/6995 · DBLP profile ↗
← Back
177ranked-venue papers
15as first author
51since 2021 · last 2026
0000-0001-6526-3889ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 127 · 10 first-author · 34 since 2021Systems, architecture and hardware · 36 · 4 first-author · 14 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Congestion Quarantine in Lossless Ethernet
abstract
Lossless Ethernet uses hop-by-hop backpressure to prevent buffer overflow and has become the mainstream choice for running Remote Direct Memory Access (RDMA) in AI and cloud data centers. Despite preventing congestion-induced drops, lossless networks introduce congestion contagion, which causes head-of-line blocking, congestion spreading, and deadlocks. Congestion control schemes have been introduced to mitigate the drawbacks of lossless Ethernet. However, congestion control mechanisms struggle with bursty traffic, face a dilemma, and can be sidelined by backpressure. In this paper, we propose congestion quarantine (CQ) as a complementary congestion management mechanism for lossless Ethernet. CQ uses a separate queue to quarantine congested flows, preventing congestion contagion, resolving the CC dilemma, and preventing CC from being sidelined. Results show that congestion quarantine can eliminate head-of-line blocking in scenarios with multiple congestion trees. Large-scale simulations demonstrate that CQ reduces the average and 99th percentile FCT slowdown of normal flows by 25–86% and 33.4–90%, respectively, with negligible impact on bursty traffic.
Dongkang Hu, Ran Shu 0001, Wenxue Cheng, Fengyuan Ren
APNet4
2026 Low-Compilation-Cost Register Allocation in LLVM-Based Binary Translation
abstract
Efficiently allocating guest architecture registers to host registers is a crucial technique for improving performance in dynamic binary translation. Translators typically enlarge code regions to reduce context-switching overhead and expose more optimization opportunities. Traditional approaches load guest registers upon entering a code region and save them upon exiting, aiming to retain guest registers in host registers for as long as possible. However, as the size of code region increases, the growing number of intermediate variables and the increasing complexity of control flow lead to higher computational costs for register allocation, thereby increasing compilation overhead.
Wei Li 0262, Fengyuan Ren
EuroSys4
2026 Come Hell or Still Water: Alleviating Tail Latency in Cloud Block Store
Chaolei Hu, Kun Qian 0004, Erci Xu, Xue Li 0024, Yuesheng Gu, Lingjun Zhu, Fengyuan Ren, Ennan Zhai
NSDI9
2026 CCC: Re-architecting Delay-based Congestion Control in Datacenter Networks
Wanchun Jiang, Haoyang Li 0006, Danfeng Shan, Fengyuan Ren, Jiawei Huang 0001, Jianxin Wang 0001
NSDI7
2026 Enabling Bounded Delay in TSN Using Per-Flow Hierarchical Scheduling
abstract
Time-Sensitive Networking (TSN) enables deterministic transmission of Scheduled Traffic (ST) through mechanisms such as the Time-Aware Shaper (TAS) and Cyclic Queuing and Forwarding (CQF). However, both rely on high-precision global clock synchronization and adopt a per-class queuing paradigm, making them susceptible to clock synchronization errors. Asynchronous Traffic Shaping (ATS) ensures bounded delay without requiring global time synchronization by assigning an eligibility time to each flow during per-flow shaping, yet often exhibits larger delay jitter. In this paper, we propose a novel scheduling paradigm for TSN by introducing per-flow queues and present the Per-Flow Hierarchical Scheduling (PFHS) mechanism. PFHS divides a hyperperiod into equal-length time slots, which are then assigned to different levels. ST flows are also assigned to these levels, and each flow is scheduled to transmit using only the time slots assigned to its level. By partitioning the bandwidth among distinct ST flows’ per-flow queues in this hierarchical manner, PFHS achieves fine-grained bandwidth allocation. We conduct a theoretical worst-case delay analysis of PFHS, demonstrating its capability to ensure deterministic transmission and tolerate bounded clock synchronization errors. We evaluate the performance of PFHS across various scenarios by extensive simulations on the OMNeT++ platform. Experimental results show that PFHS guarantees the end-to-end delay of ST flows. Moreover, the proposed hierarchical division algorithm successfully schedules 93% of 1000 flows.
Tong Zhang 0018, Xiaoqin Feng, Hao Yang 0064, Fengyuan Ren
IEEE Internet Things J.8
2026 Enhancing CQF Robustness to Time Synchronization Errors Using Shadow Queues in Time-Sensitive Network
Hao Yang 0064, Tong Zhang 0018, Xiaoqin Feng, Wenxue Wu, Fengyuan Ren
IEEE Internet Things J.9
2026 MigrRDMA: Enabling RDMA Live Migration in the Software
abstract
Live migration is critical to ensure services are not interrupted during host maintenance in data centers. On the other hand, RDMA has been widely adopted in data centers, and has attracted both academia and industry for years. However, live migration of RDMA is not supported in today’s data centers. Although modifying RDMA NICs (RNICs) to be aware of live migration has been proposed for years, it relies on extra hardware support. This paper proposes MigrRDMA, a software-based RDMA live migration system. MigrRDMA provides a software indirection layer to achieve transparent switching to new RDMA communications. Unlike previous RDMA virtualization that provides sharing and isolation, MigrRDMA’s indirection layer focuses on keeping the RDMA states on the migration source and destination identical from the perspective of applications. We implemented the MigrRDMA prototype over Mellanox RNICs. Our evaluation shows that MigrRDMA adds little downtime when migrating a container with live RDMA connections running at line rate. Besides, the MigrRDMA virtualization layer only adds 2% ∼ 9% extra overhead in the data path. When migrating Hadoop tasks, MigrRDMA only incurs an extra 3-second job completion time.
Weizhe Zhang, Fengyuan Ren
IEEE Trans. Netw.3
2026 Fast and Accurate Software Traffic Shaping With Inter-Flow Batching
Danfeng Shan, Shihao Hu, Hao Li 0011, Yazhe Tang, Peng Zhang 0011, Wanchun Jiang, Fengyuan Ren
IEEE Trans. Netw.8
2026 Compass: Congestion Control for Disobedient Traffic
abstract
Increasingly stringent service-level objectives demand fast and accurate congestion control (CC) in data center networks. We observe that the feedback loop used by existing data center sender-driven CC schemes is inherently limited as it suffers from an inevitable delay. Consequently, a portion of traffic, which we call disobedient traffic, could finish before the feedback loop, thereby escaping the control of these schemes and exerting a negative impact on their performance. Furthermore, as link speeds continue to climb, the proportion and impact of disobedient traffic are concurrently escalating, exacerbating these issues. In this paper, we propose Compass, a solution implemented entirely on switch data plane to mitigate the impact of disobedient traffic. Compass utilizes sketching techniques to efficiently estimate the sending rate of disobedient traffic, and seamlessly integrates into HPCC and PowerTCP through incorporating the sending rate of disobedient traffic into their high-precision rate control algorithms. Such integration enhances network performance without incurring additional bandwidth overhead or modifications to host-side logic. Additionally, Compass supports incremental brownfield deployment, and all these properties make it highly practical for production deployment. Extensive simulations show that Compass improves both throughput and latency. For instance, Compass reduces tail flow completion times of medium and large flows by up to 35% for PowerTCP and 17% for HPCC.
Kaicheng Yang 0001, Tianbao Zhou, Hengyang Zhou, Kaitai Zhang, Yikai Zhao 0001, Yuanpeng Li 0002, Yuhan Wu 0001, Zili Meng, Fengyuan Ren, Tong Yang 0003
IEEE Trans. Netw.9
2025 Introspective Congestion Control for Consistent High Performance
abstract
The congestion control (CC) algorithm is expected to achieve consistent high performance under different network environments. Traditionally, classic CCs are designed with the methodology of inferring path conditions to guide the rate adjustment. However, this methodology suffers from wrong path condition inferences in certain cases, which mislead the rate adjustment and lead to performance degradation. To avoid wrong path condition inferences, we develop the projection-based introspective method and design the introspective congestion control (ICC) algorithm in this paper. Specifically, the rate adjustment rules are designed to possess a specialized profile such that the projection of the profile can be distinguished under unchanged path conditions. In this way, the projection, which can be distinguished from the time series of delay signals in the frequency domain, facilitates ICC to extract more information for path condition inferences. Consequently, with the introspection on the projection, ICC can avoid being misled by wrong path condition inferences and thus achieve consistent high performance under different conditions. The advantages of ICC are confirmed through extensive experiments conducted on various locally emulated scenarios, global testbeds over the Internet, and the Alipay platform.
Wanchun Jiang, Haoyang Li 0006, Jia Wu 0011, Fengyuan Ren, Jianxin Wang 0001
EuroSys5
2025 Occamy: A Preemptive Buffer Management for On-chip Shared-memory Switches
abstract
Today's high-speed switches employ an on-chip shared packet buffer. The buffer is becoming increasingly insufficient as it cannot scale with the growing switching capacity. Nonetheless, the buffer needs to face highly intense bursts and meet stringent performance requirements for datacenter applications. This imposes rigorous demand on the Buffer Management (BM) scheme, which dynamically allocates the buffer across queues. However, the de facto BM scheme, designed over two decades ago, is ill-suited to meet the requirements of today's network.
Danfeng Shan, Yunguang Li, Jinchao Ma, Xinyu Wen, Hao Li 0011, Wanchun Jiang, Nan Li 0047, Fengyuan Ren
EuroSys10
2025 Collaborative Edge-Device DNN Inference with Dynamic Model Partitioning, Data Compression and Resource Allocation
abstract
Collaborative edge-device inference is a promising way to empower resource-constrained mobile devices to execute deep neural network (DNN)-based applications with heavy computational workloads. In particular, a DNN model is partitioned into two parts that are executed on the mobile device and the edge server, respectively. However, offline model partitioning methods suffer from poor adaptability to real computing environments, while online methods have the problem of delayed feedback. In addition, model partitioning inevitably incurs large transmission overheads of DNN's intermediate data. To tackle these challenges, we propose a collaborative edge-device inference optimization algorithm (JPCA) with Joint DNN Partitioning, data Compression, and resource Allocation. Our goal is to maximize the inference accuracy of all tasks while satisfying inference latency and energy requirements in a dynamically changing environment. In JPCA, the joint selection of model partition point and data quantization bit-width is first extracted from the original problem and we propose an improved deep reinforcement learning (DRL)-based algorithm to learn joint decisions. Optimal schemes under different bandwidth conditions are recorded, enabling mobile devices to adjust joint decisions in response to significant bandwidth changes. Furthermore, we design a dynamic edge resource allocation algorithm that makes edge resource allocation decisions for tasks arriving in real time, thereby accelerating DNN inference. The results of the testbed experiments affirm the effectiveness of our proposed algorithms in terms of inference accuracy.
Yufan Tang, Tong Zhang 0018, Kun Zhu 0001, Fengyuan Ren
HPCC4
2025 ScalaTap: Scalable Outbound Rate Limiting in Public Cloud
Zhongjie Chen, Yingchen Fan, Kun Qian 0017, Qingkai Meng 0001, Ran Shu 0001, Bo Wang 0066, Wei Li 0262, Fengyuan Ren
INFOCOM10
2025 Low-Latency Microsecond Message Scheduling with Global Consistent Priorities in RDMA Networks
abstract
With the increase in computing speed and network bandwidth in data centers, microsecond-level tail latency has become a key metric for internet-based online services. However, the tail latency of messages in data centers is mainly determined by queuing delay, which is usually much larger than pure message transmission time. Existing traffic scheduling mechanisms fail to effectively coordinate end-side and in-network resources, leading to head-of-line (HOL) blocking for microsecond-level messages at both ends and switches, which severely impacts the tail latency. To address this issue, this paper proposes a network-wide Global Priority based Multi-Path message scheduling mechanism GPMP. It assigns the highest priority to microsecond-level messages at both ends and swtiches to ensure such messages can quickly acquire resources and complete quickly. Furthermore, GPMP also optimizes the transmission of low-priority large messages by introducing a multi-path method, which effectively reduces transmission time and improves the bandwidth utilization. Extensive simulation results show that GPMP not only meets the strict latency requirements of microsecond-level messages, but also significantly reduces the completion time of long messages, leading to an overall improvement in bandwidth utilization.
Qiuyu Yu, Tong Zhang 0018, Kun Zhu 0001, Fengyuan Ren, Yufan Tang, Xiaoxiang Hua
IWQoS4
2025 eTran: Extensible Kernel Transport with eBPF
Zhongjie Chen, Qingkai Meng 0001, ChonLam Lao, Fengyuan Ren, Minlan Yu, Yang Zhou 0008
NSDI5
2025 Software-based Live Migration for RDMA
abstract
Live migration is critical to ensure services are not interrupted during host maintenance in data centers. On the other hand, RDMA has been widely adopted in data centers, and has attracted both academia and industry for years. However, live migration of RDMA is not supported in today's data centers. Although modifying RDMA NICs (RNICs) to be aware of live migration has been proposed for years, there is no sign of supporting it on commodity RNICs. This paper proposes MigrRDMA, a software-based RDMA live migration that does not rely on any extra hardware support. MigrRDMA provides a software indirection layer to achieve transparent switching to new RDMA communications. Unlike previous RDMA virtualization that provides sharing and isolation, MigrRDMA's indirection layer focuses on keeping the RDMA states on the migration source and destination identical from the perspective of applications. We implemented MigrRDMA prototype over Mellanox RNICs. Our evaluation shows that MigrRDMA adds little downtime when migrating a container with live RDMA connections running at line rate. Besides, the MigrRDMA virtualization layer only adds 3% ~ 9% extra overheads in the data path. When migrating Hadoop tasks, MigrRDMA only incurs an extra 3-second job completion time.
Ran Shu 0001, Yongqiang Xiong, Fengyuan Ren
SIGCOMM4
2025 Improving Robustness of Time-Aware Shaper in Time-Sensitive Networking
abstract
Deterministic delivery of scheduled traffic (ST) is critical in time-sensitive networking (TSN). The time-aware shaper (TAS) defined by IEEE 802.1Qbv is the enabler to ensure deterministic end-to-end delays of ST flows. However, TAS does not consider the emergency sporadic flows that commonly exist in automotive and industrial control applications. Therefore, TAS cannot resist the interference of emergency sporadic traffic on normal scheduled traffic. In this paper, we improve the TAS’s time slot allocation and queue occupancy rule and propose a combined scheduling strategy composed of dynamic local regulation and static global planning. Dynamic local regulation executes the earliest deadline first (EDF) discipline at switching nodes to adjust ST frames’ transmission dynamically. Static global planning calculates the offset of ST flows at the source and guarantees bounded delay and jitter. Furthermore, we formalize the offset optimizing problem with considering the EDF, delay, and jitter constraints. Finally, we implement our prototype on OMNet++ 6.0.1. Simulation results further demonstrate that the combined strategy has bounded latency and jitter in the presence of emergency sporadic flows. Besides, the comparison experiments show that our design is superior to eTAS in terms of the jitter and end-to-end delay.
Tong Zhang 0018, Xiaoqin Feng, Hao Yang 0064, Fengyuan Ren
IEEE Internet Things J.6
2025 Absorbing Time Synchronization Errors Using Shadow Queues in Time-Sensitive Networking
abstract
Time-sensitive networking (TSN) is widely used in industrial automation and automotive applications due to its ability to provide deterministic transmission. To meet the stringent deterministic requirements of time-sensitive traffic, the mainstream traffic management mechanism time-aware shaper (TAS) relies on time synchronization among network devices. TAS schedules periodic traffic transmission using preallocated time windows. However, in real networks, time-sensitive traffic may not be precisely forwarded as scheduled, thus failing to achieve the expected performance. A fundamental reason is that the statically planned transmission schemes cannot accommodate dynamic time synchronization errors. To address this problem, we propose a novel TSN traffic scheduling strategy called robust TAS (RTAS), which utilizes a dual-queue gating structure to define a new scheduling rule for absorbing time synchronization errors to guarantee the real-time performance of scheduled traffic (ST). We conducted extensive simulations on OMNeT++ to study the performance of RTAS in industrial automation scenarios. The results show that RTAS enables deterministic transmission with imperfectly synchronized clocks and minimizes the impact of time synchronization errors on ST.
Hao Yang 0064, Tong Zhang 0018, Xiaoqin Feng, Fengyuan Ren
IEEE Internet Things J.6
2025 Dynamic Per-Flow Queues in Shared Buffer TSN Switches
abstract
Time-Sensitive Networking (TSN), as an enhancement based on Ethernet, can ensure deterministic traffic transmission with low delays and minimal jitters. However, TSN switches have only eight priority queues inherited from Ethernet at each egress port, which limits the flexibility and efficiency of traffic scheduling, as well as the support for developing traffic management mechanisms. Although per-flow queues boost scheduling and Quality of Service (QoS), static per-flow hardware queues in switches are considered unpractical due to resource limits. In this article, we leverage the limitation of buffer size on the number of concurrent flows in shared buffer TSN switches to design Dynamic Per-Flow Queues (DFQ). DFQ only maintains a fixed number of virtual queues determined by the buffer size and dynamically manages the mapping between virtual queues and active flows to provide the capability of per-flow queuing. By constructing Flow Mapping Table (FMT) with content-addressable memory (or hash bucket), DFQ can quickly match, create, and recycle queues to multiplex limited switch resource. We prototype DFQ on an FPGA switch and evaluate its performance in different scenarios. Experimental results show that DFQ can decrease the overhead of per-flow isolation with minimal impact on delay and throughput, indicating that DFQ is an effective per-flow queues solution.
Wenxue Wu, Tong Zhang 0018, Xiaoqin Feng, Fengyuan Ren
ACM Trans. Design Autom. Electr. Syst.6
2025 Fault-Tolerant Cyclic Queuing and Forwarding with Fast ACK in Time-Sensitive Networking
abstract
TSN is widely used in industrial automation networks because it can provide deterministic transmission services for critical data. Cyclic Queuing and Forwarding (CQF) is used to shape critical data. However, unexpected data errors may occur due to transient failures like electromagnetic interference. IEEE 802.1CB provides a solution to tolerate such failures by transmitting multiple replicas of data over disjoint paths. However, this solution introduces network resources wastage. Compared to redundant transmission, retransmission can reduce resource waste, but may violate the determinism in TSN. To address this issue, we propose a fault-tolerant mechanism for CQF that supports retransmission, called fault-tolerant CQF (FT-CQF). FT-CQF adopts the Go-Back-N concept to resist failure. Therefore, it does not violate the original transmission sequence of frames. On the basis of standard CQF, FT-CQF occupies an additional queue to cache replicas of Time-Trigger (TT) flows and reserves time slots to forward them. FT-CQF will forward or remove these replicas based on the ACK information. Non-TT flows can use this time slot to transmit when replicas are removed. We implemented FT-CQF on OMNeT++ and verified the performance of FT-CQF. Simulation experiments show that FT-CQF is effective in terms of reliability, bandwidth consumption, and delay.
Tong Zhang 0018, Xiaoqin Feng, Hao Yang 0064, Fengyuan Ren
ACM Trans. Design Autom. Electr. Syst.6
2025 ACK-Driven Congestion Control for Lossless Ethernet
abstract
Congestion control is a key enabler for lossless Ethernet at scale. In this paper, we revisit this classic topic from a new perspective, i.e., understanding and exploiting the intrinsic properties of the underlying lossless network. We experimentally and analytically find that the intrinsic properties of lossless networks, such as packet conservation, can indeed provide valuable implications in estimating pipe capacity and the precise number of excessive packets. Besides, we derive principles on how to treat congested flows and victim flows individually to handle HoL blocking efficiently. Then, we propose ACK-driven congestion control (ACC) for lossless Ethernet, which simply resorts to the knowledge of ACK time series (supports ACK coalescing) to exert a temporary halt to exactly drain out excessive packets of congested flows and then match its rate to pipe capacity. Testbed and large-scale simulations demonstrate that ACC ameliorates fundamental issues in lossless Ethernet (e.g., congestion spreading, HoL blocking, and deadlock) and achieves excellent low latency and high throughput performance. For instance, compared with existing schemes, ACC improves the average and 99th percentile FCT performance of small flows by$1.3\sim 3.3\times $and$1.4\sim 11.5\times $, respectively.
Qingkai Meng 0001, Chaolei Hu, Shangguang Wang, Fengyuan Ren
IEEE Trans. Netw.6
2024 Software-based Live Migration for Containerized RDMA
abstract
Container live migration is critical to ensure services are not interrupted during host maintenance in data centers. On the other hand, RDMA containerization has attracted both academia and industry for years. However, live migration for containerized RDMA is not supported in today’s data centers. Although modifying RDMA NICs (RNICs) to be aware of live migration has been proposed for years, there is no sign of supporting it on commodity RNICs. This paper proposes MigrRDMA, a software-based RDMA live migration for containers, which does not rely on any extra hardware support. MigrRDMA provides a minimum virtualization layer inside the RDMA library loaded in applications, which achieves transparent switching to new RDMA communications. Unlike previous RDMA virtualization that provides sharing and isolation, MigrRDMA’s virtualization layer focuses on keeping the RDMA states on the migration source and destination the same from the perspective of applications. Our evaluation shows that MigrRDMA only adds 0.7 ∼ 12.1 ms downtime to migrate a container with live RDMA connections running at line rate. Besides, the MigrRDMA virtualization layer only adds 3% ∼ 9% overheads in the data path operations.
Ran Shu 0001, Yongqiang Xiong, Fengyuan Ren
APNet4
2024 Enabling Low Latency for ECQF based Flow Aggregation Scheduling in Time-Sensitive Networking
abstract
Cyclic Queuing and Forwarding (CQF) configures the same cycle length on the flow path, resulting in certain flows unschedulable. Enhanced CQF (ECQF) based flow aggregation utilizes variable cycle length to address this issue. However, it remains a conceptual model without a concrete implementation. In this paper, we propose a jointly optimize aggregation cycle and flows' offsets (JACO) mechanism to achieve ECQF-based flow aggregation. We also design an incremental heuristic algorithm for JACO. Finally, we evaluate the performance of JACO in different scenarios using OMNet++ simulation platform. Compared with ECQF, the results show that JACO reduces latency and improves resource utilization.
Tong Zhang 0018, Xiaoqin Feng, Fengyuan Ren
DAC5
2024 Dynamic Per-Flow Queues for TSN Switches
abstract
Dynamic Per-Flow Queues (DFQ) extend queues from per-class to per-flow in Time-Sensitive Networking (TSN) switches that overcome large resource consumption by dynamically mapping a fixed number of physical queues to active flows. It can implement per-flow queuing with much less on-chip resource. Compared to brute-force hardware queues, DFQ prototyped on an FPGA, can effectively manage more per-flow queues, allowing for improved priority scheduling with minimal throughput and latency impact.
Wenxue Wu, Tong Zhang 0018, Xiaoqin Feng, Xuelong Qi, Fengyuan Ren
DATE7
2024 Fault- Tolerant Cyclic Queuing and Forwarding in Time-Sensitive Networking
abstract
Time-sensitive networking (TSN) provides determin-istic time-sensitive transmission services for critical data at the link layer. Cyclic Queuing and Forwarding (CQF) defined by IEEE 802.1Qch is used for critical data transmission. However, unexpected data errors may occur due to transient faults like electromagnetic interference. At present, the solution to such faults defined in the IEEE TSN standards is to transmit multiple data copies on redundant paths, which introduces network resources wastage. Compared to redundant transmission, retransmission can reduce resource waste, but may violate the deterministic transmission guarantee in TSN. To tackle with this issue, we propose a time-redundant fault-tolerant mechanism for CQF, called fault-tolerant CQF (FT-CQF). On the basis of standard CQF, FT-CQF occupies an additional queue to cache copies of Time- Trigger (TT) flows and reserves time slots to forward them. According to the returned CRC-related messages, FT-CQF will decide whether to forward these copies. Non- Ttflows can also be transmitted during this time when copies are not required to be forwarded. We implement FT-CQF in OMNeT++, and verify the performance of FT-CQF in typical network scenarios. The extensive simulation experiments show that FT-CQF is effective in terms of fault-tolerant effects, consumed resources, delay, and jitter.
Tong Zhang 0018, Wenxue Wu, Xiaoqin Feng, Guoxi Lin, Fengyuan Ren
DATE6
2024 DockRDMA: Hybrid RDMA Virtualization for Containerized Clouds
abstract
Containers have become the de facto choice for major cloud services. Meanwhile, with demands for extremely high performance, data centers have widely adopted RDMA for their online services. RDMA virtualization is the critical technology that enables RDMA for containers. Hybrid RDMA virtualization leverages the software flexibility in the control path, and keeps the native performance in the data path. Thus, it is the best choice for RDMA virtualization. State-of-the-art hybrid RDMA virtualization cannot address containerspecific problems. This paper proposes DockRDMA, the first hybrid RDMA virtualization solution for containerized clouds. DockRDMA develops several mechanisms, including embedding physical addresses in virtual ones to provide efficient address translation, hybrid network policy enforcement at scale, a general virtual RDMA NIC initialization method to be compatible with all container platforms, and namespace checking to protect the RDMA NIC instances. Evaluation results show that DockRDMA provides bare-metal RDMA performance in the data path, and almost native communication setup time in the control path. Compared with the state-of-the-art hybrid virtualization technology, DockRDMA reduces Hadoop job completion time by 6%. It offers seamless integration with existing container platforms, protects critical information of RDMA NIC instances, and exhibits excellent scalability to meet diverse network policies required by different containers.
Ran Shu 0001, Zhongjie Chen, Xiaohui Luo, Bo Wang 0066, Qingkai Meng 0001, Fengyuan Ren
ICNP8
2024 BCC: Re-architecting Congestion Control in DCNs
abstract
The nature of datacenter traffic is a high volume of bursty tiny flows and standing long flows, which forms the coexistence of transient and persistent congestion. Traditional congestion control (CC) algorithms have inherent limitations in reconciling fast response and high efficiency towards transients with stability and fairness during persistence. In this paper, we provide an insight that re-architects CC with two control laws, tailored to transient and persistent concerns, respectively. Armed with this key insight, we propose bimodal congestion control (BCC), which is founded on two core ideas: (i) Quaternary network state detection, which further distinguishes transient and persistent states in switches, and (ii) Bimodal control law, which is manifested as the transient controller and persistent controller at sources. The transient controller employs a precise control paradigm that pauses flows to drain backlogged packets and ramps down/up flow rates to bottleneck bandwidth directly, striving for high efficiency. The persistent controller grounds itself in traditional CC algorithms, inheriting stability and fairness. We implement BCC in the Linux kernel and P4-programmable switch. In our evaluation, compared to DCQCN, HPCC, PowerTCP, and Swift, BCC reduces flow completion times by 14% ~ 99%.
Qingkai Meng 0001, Shan Zhang 0001, Zhiyuan Wang 0004, Tao Tong, Chaolei Hu, Hongbin Luo, Fengyuan Ren
INFOCOM7
2024 Explicit Dropping Notification in Data Centers
abstract
Datacenter applications increasingly demand microsecond-scale latency and tight tail latency. Despite recent advances in datacenter transport protocols, we notice that the timeout caused by packet loss is the killer of microsecond-scale latency. Moreover, refining the RTO setting is impractical due to the significant fluctuations in RTT. In this paper, we propose explicit dropping notification (EDN) to avoid timeouts. EDN rekindles ICMP Source Quench, where the switch notifies the source of precise packet loss information. Then the source can rapidly pinpoint dropped packets for fast retransmission instead of waiting for timeouts. More importantly, fast retransmission does not mean immediate retransmission which is prone to aggravate congestion and deteriorate latency. In light of this, we suggest finessing the timing and sending rate of retransmission. Specifically, as a reward of the paradigm shift to explicit notification, the source can pause for the queue draining time piggybacked on EDN messages and estimate connection capacity to figure out a proper sending rate, thus avoiding congestion aggravation. We implement EDN on the P4-programmable switching ASIC and Linux kernel. Evaluations show that, compared with state-of-the-art loss recovery schemes, EDN reduces the latency by up to 4.1× on average and 3.6× at the 99th-percentile.
Qingkai Meng 0001, Chaolei Hu, Bo Wang 0066, Fengyuan Ren
INFOCOM5
2024 Revisiting Congestion Control for Lossless Ethernet
Qingkai Meng 0001, Chaolei Hu, Fengyuan Ren
NSDI4
2024 CrossMapping: Harmonizing Memory Consistency in Cross-ISA Binary Translation
Wei Li 0262, Jinhui Lai, Fengyuan Ren
USENIX ATC6
2024 Unmanned Aerial Vehicle-enabled grassland restoration with energy-sensitive of trajectory design and restoration areas allocation via a cooperative memetic algorithm
Dongbin Jiao, Peng Yang 0008, Weibo Yang, Zhanhuan Shang, Fengyuan Ren
Eng. Appl. Artif. Intell.7
2024 Designing transport scheme of 3D naked-eye system
Xiaoqin Feng, Fengyuan Ren
J. Netw. Comput. Appl.3
2024 Switch-Assistant Loss Recovery for RDMA Transport Control
abstract
RoCEv2 (RDMA over Converged Ethernet version 2) is the canonical method for deploying RDMA in Ethernet-based datacenters. Traditionally, RoCEv2 runs over the lossless network which is in turn achieved by enabling Priority Flow Control (PFC) within the network. However, as the scale of the datacenter increases, PFC’s side effects, such as head-of-line blocking, congestion spreading, and pause frame storm, are amplified. Datacenter operators can no longer tolerate these problems. In hence, they are seeking PFC alternatives for RDMA networks. Rather than aiming at the lossless RDMA network, we instead handle packet loss effectively to support RDMA over Ethernet. In this paper, we propose Switch-assistant Loss Recovery (SLR), a switch building block to enhance RoCEv2’s loss recovery. Specifically, SLR-enabled switches send loss notifications to request fast retransmissions. To cooperate with go-back-N retransmission, SLR generates loss notifications only when expected packets (i.e., in-order packets expected by receivers) are dropped and then filters out unexpected packets, which can avoid timeouts and prevent exacerbating congestion. Further, we adapt SLR to multi-bottleneck scenarios by inferring expected packets among multiple switch views. We implement SLR prototypes on commodity programmable switches. Evaluations show that SLR reduces the 99.9th-percentile FCT slowdown by up to 21.6$\times$compared to PFC and other state-of-the-arts.
Qingkai Meng 0001, Shan Zhang 0001, Zhiyuan Wang 0004, Tong Zhang 0018, Hongbin Luo, Fengyuan Ren
IEEE/ACM Trans. Netw.7
2024 Enforcing Fairness in the Traffic Policer Among Heterogeneous Congestion Control Algorithms
abstract
Traffic policing is widely used by ISPs to limit their customers’ traffic rates. It has long been believed that a well-tuned traffic policer offers a satisfactory performance for TCP. However, we find this belief breaks with the emergence of new congestion control (CC) algorithms: flows using new CC algorithms can easily occupy the majority of bandwidth, starving traditional TCP flows. We confirm this problem with experiments and reveal its root cause as follows. Without a buffer in traffic policers, congestion only causes packet losses, while new CC algorithms are loss-resilient. When being policed, they will not reduce the sending rate until an unacceptable loss ratio for TCP is reached, resulting in low throughput for competing TCP flows. Simply adding a buffer to the traffic policer improves fairness but incurs high latency. To this end, we propose FairPolicer, which can achieve fair bandwidth allocation without sacrificing latency. FairPolicer regards a token as a basic unit of bandwidth and fairly allocates tokens to active flows in a round-robin manner. To avoid bandwidth waste when flows come and go, FairPolicer puts all available tokens in a global bucket and maintains the amount of residual bucket space rather than the number of available tokens. To scale to massive concurrent flows, FairPolicer uses a Count-Min Sketch structure to maintain per-flow data with a small memory footprint. Testbed experiments show that FairPolicer can allocate bandwidth in a max-min fair manner and achieve much lower latency than other kinds of rate limiters.
Danfeng Shan, Linbing Jiang, Peng Zhang 0011, Wanchun Jiang, Hao Li 0011, Yazhe Tang, Fengyuan Ren
IEEE/ACM Trans. Netw.7
2024 Enhancing Low Latency Adaptive Live Streaming Through Precise Bandwidth Prediction
abstract
To ensure high performance for HTTP adaptive streaming (HAS), it is critical to provide accurate prediction of end-to-end network bandwidth. Low Latency Live Streaming (LLLS), which has been gaining popularity, faces even greater challenges in this regard. Unlike Video-on-Demand (VOD) streaming, which only needs long-term bandwidth prediction and can tolerate some prediction errors, LLLS demands precise short-term bandwidth predictions. These challenges are amplified by the fact that short-term bandwidth experiences both large abrupt changes and uncertain fluctuations. Furthermore, obtaining valid bandwidth measurement samples in LLLS poses difficulties due to the on-off traffic pattern. In this work, we present DeeProphet, a system designed to enhance the performance of LLLS by achieving accurate bandwidth prediction. DeeProphet collects valid bandwidth samples by identifying intervals of packet continuous sending leveraging TCP state information, estimates the segment-level bandwidth robustly by filtering out noisy samples, and predicts both significant changes and uncertain fluctuations in future bandwidth by combining both time series and learning-based models. Experimental results demonstrate that DeeProphet effectively enhances the overall Quality of Experience (QoE) by 39.5% to 464.6% compared to state-of-the-art LLLS Adaptive Bitrate (ABR) algorithms.
Bo Wang 0066, Muhan Su, Wufan Wang, Bingyang Liu, Fengyuan Ren, Mingwei Xu 0001, Jiangchuan Liu
IEEE/ACM Trans. Netw.6
2023 weBurst can be Harmless: Achieving Line-rate Software Traffic Shaping by Inter-flow Batching
abstract
Traffic shaping is a common function at end hosts. Compared with hardware ones, software shapers are more flexible to be developed and deployed, and thus are very attractive. Nevertheless, software approaches are still unsatisfactory as they struggle to saturate 40Gbps and higher speed.While much effort has been made to reduce the intrinsic overhead of software traffic shaping, we find that it is the extrinsic overhead, such as PCIe communications and interrupts, that hinders shaping from achieving 40Gbps - 100Gbps speed. Batching is an effective way to amortize these overheads. However, blindly batching can degrade the network performance, as it introduces bursts into the network. Diving into the dilemma, we find that intra-flow burst is to blame for harming the network performance, while inter-flow burst, consisting of packets from different flows, can be naturally demultiplexed in the network.Based on the insight, we present FlowBundler, which can achieve efficient traffic shaping by inter-flow batching. Testbed experiments show that FlowBundler can achieve an accurate shaping of 98Gbps with a single CPU core, which is 2.6× better than state-of-the-art approaches. Large-scale simulations show that FlowBundler can batch packet transmissions without harming the network performance.
Danfeng Shan, Shihao Hu, Wanchun Jiang, Hao Li 0011, Peng Zhang 0011, Yazhe Tang, Huanzhao Wang, Fengyuan Ren
INFOCOM9
2023 DeeProphet: Improving HTTP Adaptive Streaming for Low Latency Live Video by Meticulous Bandwidth Prediction
abstract
The performance of HTTP adaptive streaming (HAS) depends heavily on the prediction of end-to-end network bandwidth. The increasingly popular low latency live streaming (LLLS) faces greater challenges since it requires accurate, short-term bandwidth prediction, compared with VOD streaming which needs long-term bandwidth prediction and has good tolerance against prediction error. Part of the challenges comes from the fact that short-term bandwidth experiences both large abrupt changes and uncertain fluctuations. Additionally, it is hard to obtain valid bandwidth measurement samples in LLLS due to its inter-chunk and intra-chunk sending idleness. In this work, we present DeeProphet, a system for accurate bandwidth prediction in LLLS to improve the performance of HAS. DeeProphet overcomes the above challenges by collecting valid measurement samples using fine-grained TCP state information to identify the packet bursting intervals, and by combining the time series model and learning-based model to predict both large change and uncertain fluctuations. Experiment results show that DeeProphet improves the overall QoE by 17.7%-359.2% compared with state-of-the-art LLLS ABR algorithms, and reduces the median bandwidth prediction error to 2.7%.
Bo Wang 0066, Wufan Wang, Fengyuan Ren
WWW5
2023 Towards Impact of Chunk-Level Characteristics on Mobile Live Streaming Performance
abstract
Today, mobile live streaming is gaining a rapid growth in use, which refers to watching the media content recorded and broadcast in real time on mobile devices. In live streaming process, each video segment must go through recording, encoding, uploading, transcoding, publishing, downloading, decoding before playback. The ingest algorithm inside the streamer decides the upload bitrate, while the adaptive bitrate (ABR) algorithm inside the player determines the download bitrate. Thanks to the chunked CMAF standard, each segment is split into smaller chunks that can be independently encoded, transferred, decoded and played. It is of great help to quantify the impact of chunk-level characteristics on mobile live streaming performance. In this paper, we establish a tandem queuing model to describe the whole streaming system. Based on the model, we respectively characterize rebuffering probability, rebuffering count, streaming latency, and average bitrate, analyzing the impact of chunk upload and download rates, upload and download time variances, startup threshold and chunk length on them. From analysis results, we propose insights and recommendations for bitrate adaptation in mobile live streaming and design simple heuristic ingest and ABR algorithms leveraging them. Extensive simulations verify the insights as well as effectiveness of designed algorithms.
Tong Zhang 0018, Zhewei Tang, Jiakun Bao, Fengyuan Ren
IEEE Trans. Mob. Comput.4
2023 Revisiting Congestion Detection in Lossless Networks
abstract
Congestion detection is the cornerstone of end-to-end congestion control. Through in-depth observations and understandings, we reveal that existing congestion detection mechanisms in mainstream lossless networks (i.e., Converged Enhanced Ethernet and InfiniBand) are improper, due to failing to cognize the interaction between hop-by-hop flow controls and congestion detection behaviors in switches. We define the ternary states of switch ports and present Ternary Congestion Detection (TCD) for mainstream lossless networks. TCD utilizes the ON-OFF sending pattern and the feature of queue length evolutions to detect the transitions among ternary states. We also enable TCD under the practical multiple queues scenario by TCD-MQ. Testbed and extensive simulations demonstrate that TCD can detect congestion ports accurately and identify flows contributing to congestion as well as flows only affected by hop-by-hop flow controls. Meanwhile, we shed light on how to incorporate TCD with rate control. Case studies show that existing congestion control algorithms can achieve$3.3\times $and$2.0\times $better median and 99th-percentile FCT slowdown by combining with TCD.
Qingkai Meng 0001, Fengyuan Ren
IEEE/ACM Trans. Netw.4
2022 CrossDBT: An LLVM-Based User-Level Dynamic Binary Translation Emulator
Wei Li 0262, Xiaohui Luo, Qingkai Meng 0001, Fengyuan Ren
Euro-Par5
2022 Demystifying and Mitigating TCP Capping
abstract
Today’s Internet user experience greatly depends on some user-perceived network metrics, such as throughput and latency. To improve these metrics, many Internet content providers build the content delivery network (CDN) to provide their services. Generally, CDNs adopt TCP as their transport protocol. A recent line of work improves TCP by proposing novel congestion control algorithms. However, we measure TCP performance in the production CDN and identify an interesting phenomenon termed TCP capping. When the flows experience TCP capping, the fixed-size receive window (rwnd) restricts these flows from fully utilizing network bandwidth. Through in-depth analysis, we demystify that the root cause of TCP capping is an inappropriate constraint on rwnd due to not considering the receiver’s processing capability. To mitigate it, this paper proposes a server-side scheme Apollo and a client-side scheme Artemis for Internet content providers and users, respectively. Apollo probes the receiver’s processing capability and assists the sender in packet sending. And Artemis adjusts the receive buffer in light of the receiver’s processing capability. In our evaluation, compared to vanilla TCP, TCP (w/ Apollo) and TCP (w/ Artemis) shorten flow completion time by up to 91.8% and 94.9%, respectively.
Qingkai Meng 0001, Fengyuan Ren, Tong Zhang 0018, Danfeng Shan, Yajun Yang
IWQoS2
2022 Cratus: A Lightweight and Robust Approach for Mobile Live Streaming
abstract
Live video applications are getting popular, and content providers widely use adaptive bitrate (ABR) streaming to improve QoE while maintaining low latency. However, users’ increasing preference to watch videos on mobile devices poses great challenges for ABR algorithm due to the dramatically varying cellular network. Existing learn-based ABR algorithms face difficulties to generalize to various network conditions because of their reliance on training traces, and model/rule-based ABR schemes suffer from rebuffering under low latency constraint since they cannot robustly control the buffer occupancy within a small range. To address it, this work proposes Cratus, a lightweight and robust ABR algorithm for mobile live streaming, which achieves high QoE and low latency by accurately regulating the buffer at a small level. To enhance the control ability, Cratus controls the buffer dynamic behavior rather than the buffer occupancy. By using sliding mode control approach, Cratus robustly controls the buffer dynamic and ensures that the buffer occupancy is bounded around the target level regardless of network uncertainties. Trace-driven experiments show that Cratus outperforms existing ABRs: average QoE is increased by 12.3 to 28.6 percent, and rebuffering time is limited within 0.8$s$on average, which is reduced by 53.5 to 92.3 percent.
Bo Wang 0066, Mingwei Xu 0001, Fengyuan Ren, Chao Zhou 0003
IEEE Trans. Mob. Comput.3
2022 Improving Robustness of DASH Against Unpredictable Network Variations
abstract
Most video players use adaptive bitrate (ABR) algorithms to provide good quality-of-experience (QoE) in dynamic network conditions. To deal with the adaptation challenges, many ABR algorithms select bitrate by optimizing a defined QoE function. Within the framework, various algorithms mainly differ in how the optimization problem is solved, including prediction-based approaches and learn-based approaches. However, these algorithms suffer from limited performance in the current popular mobile streaming which has limited resources and rapidly changing link rates. Existing machine-learning approaches face deployment difficulties on mobile devices, and prediction-based approaches that rely on throughput prediction experience large buffer occupancy variations in cellular networks, resulting in rebuffering frequently. To provide a robust and lightweight ABR algorithm for mobile streaming, this work improves the robustness of prediction-based scheme against unpredictable network variations and develops RBC (Robust Bitrate Controller) algorithm. Rather than optimizing QoE over the entire buffer capacity, RBC creates buffer margins to absorb the impact of throughput jitters and solves QoE maximization on the narrowed buffer range. The amount of buffer margin is dynamically adjusted based on the real-time throughput fluctuation to ensure sufficient de-jitter space. For online lightweight deployment, RBC provides a closed-form solution of the desired bitrate with small computation complexity by using adaptive control approach. Trace-driven experiments and real-world tests show that RBC effectively reduces the playback freezing and gains an improvement in overall QoE.
Bo Wang 0066, Mingwei Xu 0001, Fengyuan Ren
IEEE Trans. Multim.3
2021 Receiver-Driven Congestion Control for InfiniBand
abstract
InfiniBand (IB) has become one of the most popular high-speed interconnects in High Performance Computing (HPC). The backpressure effect of credit-based link-layer flow control in IB introduces congestion spreading, which increases queueing delay and hurts application completion time. IB congestion control (IB CC) has been defined in IB specification to address the congestion spreading problem. Nowadays, HPC clusters are increasingly being used to run diverse workloads with a shared network infrastructure. The coexistence of messages transfers of different applications imposes great challenges to IB CC. In this paper, we re-exam IB CC through fine-grained experimental observations and reveal several fundamental problems. Inspired by our understanding and insights, we present a new receiver-driven congestion control for InfiniBand (RR CC). RR CC includes two key mechanisms: receiver-driven congestion identification and receiver-driven rate regulation, which empower eliminating both in-network congestion and endpoint congestion in one control loop. RR CC has much fewer parameters and requires no modifications to InfiniBand switches. Evaluations show that RR CC achieves better average/tail message latency and link utilization than IB CC under various scenarios.
Kun Qian 0017, Fengyuan Ren
ICPP3
2021 Towards the Fairness of Traffic Policer
abstract
Traffic policing is widely used by ISPs to limit their customers' traffic rates. It has long been believed that a well-tuned traffic policer offers a satisfactory performance for TCP. However, we find this belief breaks with the emergence of new congestion control (CC) algorithms like BBR: flows using these new CC algorithms can easily occupy the majority of the bandwidth, starving traditional TCP flows. We confirm this problem with experiments and reveal its root cause as follows. Without buffer in traffic policers, congestion only causes packet losses, while new CC algorithms are loss-resilient, i.e. they adjust the sending rate based on other network feedback like delay. Thus, when being policed they will not reduce the sending rate until an unacceptable loss ratio for TCP is reached, resulting in low throughput for TCP. Simply adding buffer to the traffic policer improves fairness but incurs high latency. To this end, we propose FairPolicer, which can achieve fair bandwidth allocation without sacrificing latency. FairPolicer regards token as a basic unit of bandwidth and fairly allocates tokens to active flows in a round-robin manner. Testbed experiments show that FairPolicer can significantly improve the fairness and achieve much lower latency than other kinds of rate-limiters.
Danfeng Shan, Peng Zhang 0011, Wanchun Jiang, Hao Li 0011, Fengyuan Ren
INFOCOM5
2021 RBA: Adaptive TCP Receive Buffer Sizing
abstract
With the rapid growth of hardware devices, a single host may have simultaneous connections that vary in network bandwidth and CPU processing capability as several orders of magnitude. State-of-art flow control mechanism, i.e., TCP auto-tuning, still needs to configure the maximum receive buffer, which cannot be applied to all connections in one host. In this paper, we reveal that improper receive buffer restrained by this configuration either (i) underutilizes the available network and CPU resources or (ii) occupies too much memory and then causes overall throughput collapse. To fully utilize resources with less memory occupancy, we present Receive Buffer Adaptive-regulating (RBA) algorithm, which regulates receive buffer according to the estimation of network bandwidth and receiver's processing capability. Testbed experiments show that RBA adapts to different scenarios and brings substantial performance improvement compared to TCP auto-tuning.
Qingkai Meng 0001, Kun Qian 0017, Wenxue Cheng, Fengyuan Ren
ISCC4
2021 Lightning: A Practical Building Block for RDMA Transport Control
abstract
RoCEv2 (RDMA over Converged Ethernet version 2) is the canonical method for deploying RDMA in Ethernet-based datacenters. Traditionally, RoCEv2 runs over the lossless network which is in turn achieved by enabling Priority Flow Control (PFC) within the network. However, with the scale of data center increases, PFC’s side effects, such as head-of-line blocking, congestion spreading, and PFC storms, are amplified. Datacenter operators can no longer tolerate these problems. They are seeking PFC alternatives for RDMA networks. Rather than aim at the lossless RDMA network, we instead handle packet loss effectively to support RDMA over Ethernet.In this paper, we propose Lightning, a switch building block to enhance RoCE’s simple loss recovery. Lightning enhances the switches to send loss notifications directly to the sources with high priority, thus informing sources as quickly as possible. Then, sources can retransmit packets sooner. By addressing challenges such as that shared buffer status is not available at ingress in modern switches, Lightning generates loss notification only when the expected packet is dropped and filters other unexpected packets at ingress, so as to avoid timeouts and prevent unnecessary congestion from unexpected packets. We implement Lightning on commodity programmable switches. In our evaluation, Lightning achieves up to 16.08× reduction of 99.9th percentile flow completion time compared to PFC, IRN and other alternatives.
Qingkai Meng 0001, Fengyuan Ren
IWQoS2
2021 Congestion detection in lossless networks
abstract
Congestion detection is the cornerstone of end-to-end congestion control. Through in-depth observations and understandings, we reveal that existing congestion detection mechanisms in mainstream lossless networks (i.e., Converged Enhanced Ethernet and InfiniBand) are improper, due to failing to cognize the interaction between hop-by-hop flow controls and congestion detection behaviors in switches. We define ternary states of switch ports and present Ternary Congestion Detection (TCD) for mainstream lossless networks. Testbed and extensive simulations demonstrate that TCD can detect congestion ports accurately and identify flows contributing to congestion as well as flows only affected by hop-by-hop flow controls. Meanwhile, we shed light on how to incorporate TCD with rate control. Case studies show that existing congestion control algorithms can achieve 3.3x and 2.0x better median and 99th-percentile FCT slowdown by combining with TCD.
Qingkai Meng 0001, Fengyuan Ren
SIGCOMM4
2021 Optimizing the Response Time of Memcached Systems via Model and Quantitative Analysis
abstract
Memcached is a widely used in-memory caching solution in large-scale searching scenarios. The most crucial metric of Memcached systems is the response time, which is affected by various factors such as workload, service rate, unbalanced load distribution, and cache miss ratio. This article aims to quantify the influence of each factor on the response time of Memcached systems. First, we establish a theoretical model for Memcached systems that captures their main features, including burst and concurrent key arrival, unbalanced load distribution, and cache miss process. By solving this model using queuing and stochastic theories, we obtain an estimate of the response time in Memcached systems. Intensive experiments based on real-world components demonstrate that the estimate always matches perfectly with the actual value. Furthermore, we obtain a comprehensive and quantitative understanding of all factors. The main insights are threefold. 1) There exists an optimum range of utilization at Memcached servers in which the response time is kept at a low level with a small penalty. 2) The influence of the cache miss ratio on the response time is logarithmic rather than linear. 3) The number of keys generated from an end-user request has the greatest impact in Memcached systems.
Wenxue Cheng, Fengyuan Ren, Wanchun Jiang, Tong Zhang 0018
IEEE Trans. Computers2
2021 Improving the Performance of Online Bitrate Adaptation with Multi-Step Prediction Over Cellular Networks
abstract
Video streaming over mobile is flourishing, and most commercial players use adaptive bitrate (ABR) streaming to deliver video in varying network conditions. Using network capacity and buffer occupancy as system states, ABR algorithms adjust bitrate based on the instantaneous system states, which is able to adapt to network changes in real-time and ensure high quality of experience (QoE). However, they are incapable of providing good QoE over mobile. Due to the high dynamic characteristics of cellular network, the system states change rapidly over time. The instantaneous state-based adaptation can induce significant video quality fluctuation which greatly degrades QoE. In this paper, we propose an online ABR algorithm called MSPC to provide good QoE in cellular network. To balance the conflict between rapid adaptation and smooth bitrate, MSPC utilizes the multi-step prediction of future system states to select bitrates instead of the instantaneous current states. At the same time, it controls the buffer occupancy to eliminate the impact of prediction error on performance. We implement MSPC on a reference video player with performance evaluated based on realistic cellular traces. Experimental results show that MSPC reduces the bitrate change of existing online algorithms by 62.4 percent on average while maintaining high bitrates and achieving zero rebuffering over 97.83 percent of all tested sessions.
Bo Wang 0066, Fengyuan Ren, Jiahai Yang 0001, Chao Zhou 0003
IEEE Trans. Mob. Comput.2
2021 Minimizing Coflow Completion Time in Optical Circuit Switched Networks
abstract
Nowadays, optical circuit switching is becoming an increasingly favored technology in scaling data center networks for its definitive advantages in data rate, power consumption, and device cost. Concurrently, reducing coflow completion time (CCT) is of great significance for improving application-level performance. However, minimizing CCT in circuit switched networks is totally different from that in traditional packet switched networks due to port constraints and circuit reconfiguration delays. To address this issue, this article proposes Grouped Optimization-based Scheduling (GOS), a CCT minimization algorithm for circuit switched networks integrating circuit and coflow scheduling. We first formalize the CCT minimization problem into a 0-1 programming problem, then relax and solve the problem in 2 steps to obtain the coflow order and flow grouping decisions on each circuit. Thus intra-group reconfiguration delays are saved, and small coflows can be prioritized at the group level. Theoretical analysis proves GOS is a 4-approximation algorithm in average CCT. To reduce computing overheads, we further propose a heuristic approximation algorithm. Extensive simulations show that the heuristic algorithm has satisfactory CCT performance (0.12× Varys, 0.36× Sunflow) as well as high throughput (16.74× Varys, 1.32× Sunflow), and well adapts to a wide range of reconfiguration delays and algorithm decision time.
Tong Zhang 0018, Fengyuan Ren, Jiakun Bao, Ran Shu 0001, Wenxue Cheng
IEEE Trans. Parallel Distributed Syst.2
2020 One Rein to Rule Them All: A Framework for Datacenter-to-User Congestion Control
abstract
Today, considerable Internet traffic is sent from datacenter and heads for users. The network characteristics of connections served by servers in datacenters are usually diverse. As a result, a specific congestion control algorithm hardly accommodates the heterogeneity and performs well in various scenarios. In this work, we present Rein — a novel framework for Internet congestion control. With Rein, diverse congestion control algorithms can be assigned purposely to connections in one server to adapt to heterogeneity. We design and implement Rein in Linux, and the experiments validate that Rein is capable of smoothly switching among various candidate algorithms on the fly to achieve potential performance gain. Meanwhile, the overheads introduced by Rein are moderate and acceptable.
Danfeng Shan, Xiaohui Luo, Tong Zhang 0018, Yajun Yang, Fengyuan Ren
APNet6
2020 Modeling and Analyzing Live Streaming Performance
abstract
Today, live streaming is gaining a rapid growth in use, which refers to streaming the media content recorded and broadcast in real time. In live streaming, latency is of utmost importance since smaller latency means higher user engagement. HTTP adaptive streaming (HAS) is now the most popular live streaming technology, where the video client sends HTTP requests to server to download video segments. The bitrate adaptation (ABR) algorithm inside the client determines bitrate level for every segment. It is of great help for ABR algorithm to quantify the influence of different HAS factors on streaming performance. However, existing work mainly focuses on video on demand (VoD) streaming rather than live streaming. In this paper, we theoretically analyze live streaming performance. We first establish a queuing model to describe playout buffer evolution. Based on the model, we respectively characterize rebuffering probability, rebuffering count and streaming latency, and analyze the effects of chunk arrival rate, arrival interval fluctuation, startup threshold and video skipping on them. From analysis results, we propose insights and recommendations for bitrate adaptation in live streaming and design a simple heuristic ABR algorithm leveraging them. Extensive simulations verify the insights as well as effectiveness of the designed algorithm.
Tong Zhang 0018, Fengyuan Ren, Bo Wang 0066
IWQoS2
2020 Re-architecting Congestion Management in Lossless Ethernet
Wenxue Cheng, Kun Qian 0017, Wanchun Jiang, Tong Zhang 0018, Fengyuan Ren
NSDI5
2020 Towards Influence of Chunk Size Variation on Video Streaming in Wireless Networks
abstract
In recent years, the growth in popularity of mobile video streaming services is unbroken. There are tremendous demands for video streaming over wireless networks. Currently, most video streaming is over HTTP. Up to now, HTTP-based adaptive video streaming is standardized as DASH, where a client-side video player can dynamically pick the bitrate level according to the perceived network conditions. Actually, not only the available bandwidth drastically varies due to wireless network properties, but also the chunk sizes in the same bitrate level significantly fluctuate, which also influences the bitrate adaptation. However, existing bitrate adaptation algorithms mostly focus on available bandwidth but do not involve chunk size variation, leading to performance losses. In this paper, we theoretically analyze the influence of chunk size variation on bitrate adaptation performance in wireless networks. Based on DASH system features, we build a general model describing playback buffer evolution. Applying stochastic theories, we respectively analyze the influence of the chunk size variation on rebuffering probability, average bitrate, and bitrate switching interval. Furthermore, based on theoretical insights, we provide several suggestions for algorithm designing and rate encoding, and also design a simple bitrate adaptation algorithm. Extensive simulations verify our insights, suggestions, and designed algorithm effectiveness.
Tong Zhang 0018, Fengyuan Ren, Wenxue Cheng, Xiaohui Luo, Ran Shu 0001
IEEE Trans. Mob. Comput.2
2020 Observing and Mitigating Micro-Burst Traffic in Data Center Networks
abstract
Micro-burst traffic is not uncommon in data centers. It can cause packet dropping, which may result in serious performance degradation (e.g., Incast problem). However, current approaches to mitigate micro-burst is usually ad-hoc and not based on a principled understanding of the underlying behaviors. On the other hand, traditional studies focus on traffic burstiness in a single flow, while micro-burst traffic in the data centers could occur with highly fan-in communication pattern, and its dynamic behavior is still unclear. To this end, in this paper, we re-examine the micro-burst traffic in typical data center scenarios. We find that the evolution of micro-burst is determined by both TCP's self-clocking mechanism and congestion control algorithm. Besides, dynamic behaviors of micro-burst under various scenarios can all be described by the time derivative of queue length evolution.Our observations also implicate that conventional solutions like absorbing and pacing are ineffective to mitigate micro-burst traffic.Instead, senders need to rapidly respond to some explicit signals of the queue buildup caused by the micro-burst traffic rather than independently and ineffectually pacing themselves in isolation. Inspired by the findings and insights from experimental observations, we propose Micro-burst-Aware Transport Control Protocol (MATCP), which leverages characteristic behaviors of micro-burst traffic derived from the time derivative of the queue occupancy. MATCP can suppress the sharp queue length increment by over 2x and reduce the tail query completion time by up to 84.4%.
Danfeng Shan, Fengyuan Ren, Peng Cheng 0005, Ran Shu 0001, Chuanxiong Guo
IEEE/ACM Trans. Netw.2
2020 Towards Power Efficient High Performance Packet I/O
abstract
Recently, high performance packet I/O frameworks continue to flourish for their ability to process packets from high-speed links. To achieve high throughput and low latency, high performance packet I/O frameworks usually employ busy polling. As busy polling will burn all CPU cycles even if there's no packet to process, these frameworks are quite power inefficient. However, exploiting power management techniques such as DVFS and LPI in the frameworks is challenging, because neither the OS nor the frameworks can provide information (e.g., actual CPU utilization, available idle period, or the target frequency) required by these techniques. In this article, we establish a model that can formulate the packet processing flow of high performance packet I/O to help and address the above challenges. From the model, we can deduce the information needed for power management techniques, and gain the insights to balance the power and latency. After suggesting to use pause instruction to reduce CPU power within short idle period, we propose two approaches to conduct power conservation for high performance packet I/O: one with the aid of traffic information and the other without. Experiments with Intel DPDK show that both approaches can achieve significant power reduction with little latency increase.
Wenxue Cheng, Tong Zhang 0018, Fengyuan Ren, Bailong Yang
IEEE Trans. Parallel Distributed Syst.4
2019 FlexGate: High-performance Heterogeneous Gateway in Data Centers
abstract
Large-scale data centers support various applications and process/issue terabits per second traffic from/to Internet. On the boundary of data center, the gateway needs to execute a series of network functions for each incoming packet. The Network Function Virtualization (NFV) technology leverages commodity servers to flexibly implement network functions. This solution provides satisfying processing and storage capability. However, state-of-the-art NFV platforms can merely process network functions at the line rate of 10~40Gbps. Supporting throughput of terabits per second requires dozens or even hundreds of servers operating exclusively for network functions, which is not only expensive but also difficult to maintain. On the other hand, programmable packet processing hardwares proposed in recent years offer a new platform for implementing network functions. They can execute user-defined packet processing logics at ultra-high line rate while containing limited processing and storage resources.
Kun Qian 0017, Mao Miao, Jianyuan Lu, Tong Zhang 0018, Peilong Wang, Fengyuan Ren
APNet8
2019 Improving Robustness of DASH Against Network Uncertainty
abstract
Most video players use adaptive bitrate (ABR) algorithms to ensure good quality-of-experience (QoE) across diverse network conditions. To balance conflicting QoE factors, state-of-the-art ABR algorithms select bitrate by optimizing a defined QoE function. However, this scheme relies on throughput prediction that is sensitive to network conditions, so the achieved QoE can be poor in unstable networks. In this paper, we propose a robust ABR algorithm called RBC to avoid the impact of prediction error on QoE. RBC controls the buffer occupancy within a safe range while maximizing the QoE, and employs an adaptive bitrate controller to ensure a good control performance in various network condition. Trace-driven experiments show that RBC achieves much less playback freezing and an improvement of 13.5% on average QoE over the best approach RobustMPC, which confirms the effectiveness of buffer control and adaptive controller in improving system robustness against network uncertainty.
Bo Wang 0066, Fengyuan Ren
ICME2
2019 Hybrid Control-Based ABR: Towards Low-Delay Live Streaming
abstract
Video content providers are increasingly interested in interactive live streaming since user engagement increases their revenues. To provide high quality of experience (QoE), it is critical to design a low delay adaptive bitrate (ABR) algorithm, but which is lacked in existing studies. The low delay constraint poses much more challenges to achieve high bitrate and low rebuffering. For example, low delay requires the player to maintain a small playback buffer, which, however, increases the risk of rebuffering. This work designs a low delay ABR algorithm called HCA which provides good QoE by controlling the buffer occupancy at a low but non-empty level. To achieve accurate control, HCA uses the hybrid of feedback and feedforward control to regulate the buffer dynamic (buffer occupancy and its variation) based on predictions of future network condition (throughput and its variation). Trace-driven experiments show that HCA achieves zero rebuffering for 98% of all traces while ensuring high bitrate.
Bo Wang 0066, Fengyuan Ren, Chao Zhou 0003
ICME2
2019 BDAC: A Behavior-aware Dynamic Adaptive Configuration on DHCP in Wireless LANs
abstract
DHCP is widely used to dynamically allocate IP addresses to the devices on local area networks, but the explosive increases of WiFi devices and their frequent mobility pose great challenges on DHCP performance in wireless LANs. In this paper, by analyzing large scale real network traces, we observe that the dynamic WiFi user behavior (e.g., online time pattern and spatio-temporal mobility pattern) leads to the poor DHCP performance. The IP pools in some VLANs have been exhausted in rush hours although the total IP utilization in WLAN is only 24%. Therefore, we have to configure IP lease times and IP pools dynamically and make sure that they are adaptive to the WiFi user behavior. In order to achieve this goal, we characterize and model the user behavior across online time pattern and spatiotemporal mobility pattern. Then we propose BDAC, a behaviour-aware dynamic adaptive configuration, which is combined of two strategies: adaptive IP lease time configuration and dynamic IP pool configuration. The former is to set adaptive lease times across user roles and area types based on online time pattern to reclaim IP addresses in time and reduce the peak IP usage, while the latter dynamically migrates the IP addresses across VLANs based on spatio-temporal mobility correlation to save the IP addresses. Using the real network traces of a different week, we conduct experiments to evaluate the performance of BDAC. Results show that BDAC can save up to 60% of IP addresses and the actual IP utilization rises from 24% to 59%. Furthermore, BDAC maintains high IP utilization when the number of VLANs in a WLAN increases.
Congcong Miao, Jilong Wang 0001, Tianying Ji, Hui Wang 0011, Chao Xu 0015, Fengyuan Ren
ICNP7
2019 Active and Adaptive Application-Level Flow Control for Latency Sensitive RPC Applications
abstract
The Remote Procedure Call (RPC) frameworks are widely deployed in industry. Applications supported by RPC frameworks are often latency-sensitive which strictly require to be responded before the deadline. For meeting this requirement, RPC frameworks adopt the application-level flow control mechanism. This mechanism gives an appropriate threshold determining the number of RPC requests that the server can process, thus avoids missing the deadline. However, this threshold at the application-level is a fixed empirical value so that it is hard to obtain respectable performance because an endpoint's processing capacity can take a huge quantity of values by varying workload and different hardware configurations. While other methods based on specialized transport protocols are adaptive, they will introduce extra costs for message reporting from server to client. Furthermore, adopting specialized transport protocols will also introduce extra transplanting efforts for TCP-based applications. In this paper, we provide an active and adaptive application level flow control mechanism at the client side. We first design an algorithm to find the appropriate threshold to achieve the desired response time. Then based on this algorithm, we control the threshold to bound the response time as expected. We implement our flow control mechanism using a memcached testbed. Experiments prove that our mechanism can accurately reduce the mean and 99th percentile response time by at least 71.3% and 69.4% respectively, while keeping a relatively high QPS. Furthermore, compared to static-threshold mechanism, our flow control mechanism is more efficient under low latency constraints.
Jing Xie 0005, Wenxue Cheng, Tong Zhang 0018, Qingkai Meng 0001, Fengyuan Ren
ICPADS7
2019 Comparing Busy Poll Socket and NAPI
abstract
Low response time is the requirement for high performance computing applications. The network latency is an important influence factor on response time. Busy poll socket (BPS) is a new mechanism that can reduce network latency by introducing extra energy cost to busily poll RX queue. Unlike hardware dependent solutions such as RDMA, busy poll socket can improve performance without specialized hardware and transplanting applications. BPS now becomes mature and is supported by many off-the-shelf network interface cards. In this paper, we compare the mechanisms of BPS and NAPI. We make a comprehensive analysis about both BPS and NAPI mechanisms. We analyze which factors would influence latencies in BPS and NAPI modes respectively. We find out that BPS does not definitely outperform NAPI, and the reasons are explored in detail. Our findings can provide suggestions on the future deployment of BPS based applications.
Jing Xie 0005, Qingkai Meng 0001, Xunli Fan, Niu Bo, Fengyuan Ren
ICPADS6
2019 Gentle flow control: avoiding deadlock in lossless networks
abstract
Many applications in distributed systems rely on underlying lossless networks to achieve required performance. Existing lossless network solutions propose different hop-by-hop flow controls to guarantee zero packet loss. However, another crucial problem called network deadlock occurs concomitantly. Once the system traps in a deadlock, a large part of network would be disabled. Existing deadlock avoidance solutions focus all their attentions on breaking the cyclic buffer dependency to eliminate circular wait (one necessary condition of deadlock). These solutions, however, impose many restrictions on network configurations and side-effects on performance.
Kun Qian 0017, Wenxue Cheng, Tong Zhang 0018, Fengyuan Ren
SIGCOMM4
2019 Distributed Bottleneck-Aware Coflow Scheduling in Data Centers
abstract
With the booming development of data parallel frameworks, the coflow abstraction has been greatly favored by data center transport designs, for its prominent ability in capturing application-level semantics. To accelerate job completion, coflow completion time (CCT) is a most important metric, and coflow scheduling is the most effective and widely-adopted means of optimizing CCT. However, most existing coflow scheduling mechanisms neglect the ubiquitous in-network bottlenecks and schedule coflows based on non-blocking giant switch hyperthesis. Such a practice is likely to result in undesired link contention inside the fabric, finally impairing CCT performance. To address this problem, we propose the Distributed Bottleneck-Aware coflow scheduling algorithm called DBA, which approximates the minimum remaining time first (MRTF) heuristic on all fabric-wide links. In this way, core link bandwidths are allocated to coflows as expected and the CCT performance will not be violated. As an evolutionary algorithm, DBA enhances the traditional dual decomposition method thus converges to the optimal bandwidth allocation very fast. Extensive simulations verify DBA's outstanding CCT performance as well as high link utilization. Furthermore, DBA introduces very little overhead and is robust to routing strategies, parameter variations and computation delays.
Tong Zhang 0018, Ran Shu 0001, Zhiguang Shan, Fengyuan Ren
IEEE Trans. Parallel Distributed Syst.4
2018 Estimating Short Connection Capacity on High Performance User Level Network Stack
abstract
Short connections are generally used to transfer small-size messages, which contribute a large part of workload in modern applications. The maximum sustainable short connection rate, which is called short connection capacity, is an important index for admission control, Web QoS control, and energy saving. A capacity estimation mechanism aims to find the workload just saturating the server, and it relies on both workload information and system information. Past researches point out that kernel space network stack becomes the bottleneck when a huge number of concurrent short connections coexist. On the other hand, high performance user level network stacks have been proved to eliminate such bottleneck, thus become a hot research topic in both academia and industry. However, they also bring challenges for estimating short connection capacity, making traditional methods ineffective. Therefore, it is important to find a new method to estimate short connection capacity on high performance user level network stacks. In this paper, we prove that the effective CPU utilization is an adaptive index to different workload patterns and application complexities, which can reflect the server state. Then we design and implement an online capacity estimator on the Seastar platform. We conduct experiments to verify the effectiveness of our online capacity estimator. The results show that our estimator can actually estimate the capacity online. When the server is near saturated, the 90th percentile relative estimating error is no more than 9.18%. Furthermore, our capacity estimator only introduces no more than 1.38% of capacity loss in our experiments.
Jing Xie 0005, Wenxue Cheng, Tong Zhang 0018, Danfeng Shan, Fengyuan Ren
ICCCN5
2018 Micro-Burst in Data Centers: Observations, Analysis, and Mitigations
abstract
Micro-burst traffic is not uncommon in data centers. It can cause packet dropping, which results in serious performance degradation (e.g., Incast problem). However, current solutions that attempt to suppress micro-burst traffic are extrinsic and ad hoc, since they lack the comprehensive and essential understanding of micro-burst's root cause and dynamic behavior. On the other hand, traditional studies focus on traffic burstiness in a single flow, while in data centers micro-burst traffic could occur with highly fan-in communication pattern, and its dynamic behavior is still unclear. To this end, in this paper, we re-examine the microburst traffic in typical data center scenarios. We find that evolution of micro-burst is determined by both TCP's self-clocking mechanism and bottleneck link. Besides, dynamic behaviors of micro-burst under various scenarios can all be described by the slope of queue length evolution. Our observations also implicate that conventional solutions like absorbing and pacing are ineffective to mitigate micro-burst traffic. Instead, senders need to slow down as soon as possible. Inspired by the findings and insights from experimental observations, we propose S-ECN policy, which is an ECN marking policy leveraging the slope of queue length evolution. Transport protocols utilizing S-ECN policy can suppress the sharp queue length increment by over 2×, and reduce the average query completion time by ~12-27%.
Danfeng Shan, Fengyuan Ren, Peng Cheng 0005, Ran Shu 0001, Chuanxiong Guo
ICNP2
2018 Power Efficient High Performance Packet I/O
abstract
Recently, high performance packet I/O frameworks are expected an extensive application for their ability to process packets from 10Gbps or higher speed links. To achieve high throughput and low latency, high performance packet I/O frameworks usually employ busy polling technique. As busy polling will burn all CPU cycles even if there's no packet to process, these frameworks are quite power inefficient. Meanwhile, exploiting power management techniques such as DVFS and LPI in high performance packet I/O frameworks is challenging, because neither the OS nor the frameworks can provide information (e.g., the actual CPU utilization, available idle period, or the target frequency) required by power management techniques. In this paper, we establish an analytical model that can formulate the packet processing flow of high performance packet I/O to help address the above challenges. From the analytical model, we can deduce the actual CPU utilization and average idle period in different traffic load, and gain the insight to choose CPU frequency that can appropriately balance the power consumption and packet latency. Then, we propose two simple but effective approaches to conduct power conservation for high performance packet I/O: one with the aid of traffic information and the other without. Experiments with Intel DPDK show that both approaches can achieve significant power reduction (35.90% and 34.43% on average respectively) while incurring < 1 μs of latency increase.
Wenxue Cheng, Tong Zhang 0018, Jing Xie 0005, Fengyuan Ren, Bailong Yang
ICPP5
2018 High Performance Userspace Networking for Containerized Microservices
Xiaohui Luo, Fengyuan Ren, Tong Zhang 0018
ICSOC2
2018 mTSL: Making mTCP Stack Transparent to Network Applications
abstract
Network applications are widely distributed nowadays, most of which have steep demand on response time. Deploying multi-threaded design on multicore systems is beneficial of scaling applications' performance, but also requires an efficient TCP stack to support. mTCP is a highly scalable userlevel TCP stack fruitful in promoting scalability and improving performance, therefore adopted by more and more applications. However, the original mTCP APIs are not compatible with the in-kernel function calls in form, thus impeding the transparent employment as well as the convenient transplant of mTCP stack for users. To overcome the deficiency, we propose a transparent socket layer for mTCP (mTSL), which overrides the native mTCP APIs and redirects original system calls to our customized versions. Finally, mTSL not only achieves mTCP stack's thorough transparency to applications, but maintains the high performance of mTCP and outperforms Linux kernel stack by 8.9× with respect to throughput on a message benchmark as well as 1.44×~6.49× in terms of transaction rate for a real application.
Xiaohui Luo, Fengyuan Ren
ISCC3
2018 Scheduling Coflows with Incomplete Information
abstract
In recent years, the coflow abstraction has received significant attentions, for its prominent ability to capture application semantics. On this basis, multiple coflow scheduling mechanisms have been proposed to minimize the coflow completion time (CCT). Currently, existing coflow scheduling mechanisms mainly belong to two categories: information-omniscient and information-agnostic. However, in data center applications, there are still quite a few cases in between where incomplete coflow information is known, and such incomplete information makes great contributions to improving the CCT performance. To address such cases, we propose IICS, a coflow scheduling algorithm based on incomplete coflow information. IICS leverages information of a coflow's arrived parts to deduce the coflow's remaining transmission time, and uses it to approximate the Minimum Remaining Time First (MRTF) heuristic. Besides, IICS allocates bandwidth by monopolization and in a maximal manner, which achieves high bandwidth utilization. Extensive simulations under realistic settings show that IICS achieves the average CCT comparable to that of the information-omniscient algorithm and the 99th percentile CCT much smaller than both information-omniscient and information-agnostic algorithms. Furthermore, IICS holds observably higher throughput and is robust to algorithm parameters.
Tong Zhang 0018, Fengyuan Ren, Ran Shu 0001, Bo Wang 0066
IWQoS2
2018 A contention-oriented node sleeping MAC protocol for WBAN
abstract
The wireless body area network (WBAN) is a new-type wireless sensor network which has a steep demand for improving energy efficiency and reducing packet delay. However, in a multi-priority environment, the current IEEE Std. 802.15.6 MAC protocol for WBAN may result in excess transmission delay and power consumption due to the selfishness of high-priority sensor nodes. To overcome the deficiency, in this paper, a contention-oriented node sleeping MAC protocol is proposed. The MAC protocol utilizes a contention orientation mechanism between different contention levels to achieve a fair resource allocation. Furthermore, the sleeping scheme of redundant nodes yields energy efficiency. Finally, simulation results show that our proposed protocol outperforms 802.15.6 MAC, AD-MAC as well as DTD-MAC protocols in terms of both packet delay and energy efficiency.
Jingjing Wang 0001, Chunxiao Jiang, Fengyuan Ren, Yong Ren 0001
WCNC4
2018 Analysing and improving convergence of quantized congestion notification in Data Center Ethernet
Ran Shu 0001, Fengyuan Ren, Jiao Zhang 0002, Tong Zhang 0018, Chuang Lin 0002
Comput. Networks2
2018 ECN Marking With Micro-Burst Traffic: Problem, Analysis, and Improvement
Danfeng Shan, Fengyuan Ren
IEEE/ACM Trans. Netw.2
2018 Towards Stable Flow Scheduling in Data Centers
abstract
At present, soft real-time data center applications are in a booming development and impose stringent delay requirements on internal data transfers. In this context, many recently proposed data center transport protocols share a common goal of minimizing Flow Completion Time (FCT), and the Shortest Remaining Processing Time (SRPT) scheduling algorithm has attracted widespread attentions for its superior performance in average FCT. However, SRPT suffers from the instability problem, incurring more and more flows left uncompleted even if the traffic load is within the fabric capacity, which implies unnecessary bandwidth waste. To solve the problem, this paper proposes a backlog-aware flow scheduling algorithm (BASRPT) for both giant switch and general topologies. Because of taking into account queue backlogs other than flow sizes at scheduling, we prove that BASRPT is stable and still maintains good FCT performance. To overcome the huge computation overhead and enable distributed implementation, a fast and practical approximation algorithm called fast BASRPT is also developed. Extensive flow-level simulations show that fast BASRPT indeed stabilizes the queue length and obtains a higher throughput while being able to push the FCT arbitrarily close to the optimal value in the condition of feasible traffic loads.
Tong Zhang 0018, Fengyuan Ren, Ran Shu 0001
IEEE Trans. Parallel Distributed Syst.2
2017 SoftRDMA: Rekindling High Performance Software RDMA over Commodity Ethernet
abstract
Recent academic and industrial work is exploring the challenges of using RDMA over Ethernet, to support highly reliable, latency-sensitive services in today's datacenters. Previous work on the high-speed packet I/O like netmap, DPDK, etc., and high-performance user-level stacks like mTCP, IX etc., rekindles our inspirations to implement a high-performance software RDMA over commodity Ethernet devices.
Mao Miao, Fengyuan Ren, Xiaohui Luo, Jing Xie 0005, Qingkai Meng 0001, Wenxue Cheng
APNet2
2017 Improving Optimization-Based Rate Adaptation in DASH System
abstract
More and more commercial video players use bitrate adaptation to adjust video quality according to varying network conditions. Optimization-based approaches are widely used for bitrate adaptation in Dynamic Adaptive Streaming over HTTP (DASH). Essentially, the optimization problem is solved based on the prediction of buffer dynamics. However, stochastic chunk size deviates observably the buffer occupancy from the expected value, making the evolution hard to predict. In order to get rid of this effect and improve the prediction accuracy for buffer occupancy, we propose an algorithm based on markov decision process with incorporating chunk size information so that only the network capacity variation need to be considered in the decision-making process. Experiment results show that our solution can effectively eliminate performance oscillation induced by variable chunk size and achieve a good QoE.
Bo Wang 0066, Xiaohui Luo, Fengyuan Ren
ICCCN4
2017 Modeling and Analyzing Latency in the Memcached system
abstract
Memcached is a widely used in-memory caching solution in large-scale searching scenarios. The most pivotal performance metric in Memcached is latency, which is affected by various factors including the workload pattern, the service rate, the unbalanced load distribution and the cache miss ratio. To quantitate the impact of each factor on latency, we establish a theoretical model for the Memcached system. Specially, we formulate the unbalanced load distribution among Memcached servers by a set of probabilities, capture the burst and concurrent key arrivals at Memcached servers in form of batching blocks, and add a cache miss processing stage. Based on this model, algebraic derivations are conducted to estimate latency in Memcached. The latency estimation is validated by intensive experiments. Moreover, we obtain a quantitative understanding of how much improvement of latency performance can be achieved by optimizing each factor and provide several useful recommendations to optimal latency in Memcached.
Wenxue Cheng, Fengyuan Ren, Wanchun Jiang, Tong Zhang 0018
ICDCS2
2017 Improving ECN marking scheme with micro-burst traffic in data center networks
abstract
In data centers, micro-burst is a common traffic pattern. The packet dropping caused by it usually leads to serious performance degradations. Therefore, much attention has been paid to avoiding buffer overflow caused by micro-burst traffic. In particular, ECN is widely used in data centers to keep persistent queue occupancy low, so that enough buffer space can be available as headroom to absorb micro-burst traffic. However, we find that instantaneous-queue-length-based ECN may cause problems in another direction - buffer underflow. Specifically, current ECN marking scheme in data centers is easy to trigger spurious congestion signals, which may result in overreaction of senders and queue length oscillations in switches. Since ECN threshold is low, the buffer may underflow and link capacity is not fully used. In this paper, we reveal this problem by experiments and simulations. Besides, we theoretically deduce the amplitude of queue length oscillations. The analysis result shows that overreaction of senders is caused by ECN mis-marking. Therefore, we propose Combined Enqueue and Dequeue Marking (CEDM), which can mark packets more accurately. Through simulations, we show that CEDM can greatly reduce throughput loss and improve flow completion time.
Danfeng Shan, Fengyuan Ren
INFOCOM2
2017 Modeling and analyzing the influence of chunk size variation on bitrate adaptation in DASH
abstract
Recently, HTTP-based adaptive video streaming has been widely adopted in the Internet. Up to now, HTTP-based adaptive video streaming is standardized as Dynamic Adaptive Streaming over HTTP (DASH), where a client-side video player can dynamically pick the bitrate level according to the perceived network conditions. Actually, not only the available bandwidth is varying, but also the chunk sizes in the same bitrate level significantly fluctuate, which also influences the bitrate adaptation. However, existing bitrate adaptation algorithms do not accurately involve the chunk size variation, leading to performance losses. In this paper, we theoretically analyze the influence of chunk size variation on bitrate adaptation performance. Based on DASH system features, we build a general model describing the playback buffer evolution. Applying stochastic theories, we respectively analyze the influence of the chunk size variation on rebuffering probability and average bitrate level. Furthermore, based on theoretical insights, we provide several recommendations for algorithm designing and rate encoding, and also propose a simple bitrate adaptation algorithm. Extensive simulations verify our insights as well as the efficiency of the proposed recommendations and algorithm.
Tong Zhang 0018, Fengyuan Ren, Wenxue Cheng, Xiaohui Luo, Ran Shu 0001
INFOCOM2
2017 Renovate high performance user-level stacks' innovation utilizing commodity network adaptors
abstract
Today's data center servers are equipped with high speed and complex network adaptors, featuring an array of functions, e.g. hardware TX/RX queues, packet filters, rate limiters, etc. Recent work like IX, Arrakis, MultiStack has made us rekindle the user-level network stacks' innovation utilizing these commodity network adaptors. In this paper, we revisit the idea to move stacks' design from in-kernel shared space into user-level application-specific dedicated one, for high performance and ease of development and deployment. We provide an unified control plane TAPM to exploit and manage the hardware adaptors' resources, and a dedicated data plane hwTAP to support different user-level stacks. TAPM and hwTAP highlight the utilization of hardware features from commodity network adaptors, to support the innovation of different user-level stacks. Experiments show that the hardware switching module can keep the input rate without any overheads and costs. TAPM could configure the hwTAP dynamically. Our run-to-completion user-level stack also achieves high throughput and low latency.
Mao Miao, Xiaohui Luo, Fengyuan Ren, Wenxue Cheng, Jing Xie 0005
ISCC3
2017 Congestion control in Converged Ethernet with heterogeneous and time-varying delays
abstract
Congestion control is an indispensable mechanism in the new trend of enhanced Ethernet as a unified fabric for traditional LAN, SAN, and high-performance computing networks. A congestion management framework for Converged Ethernet (CE) networks has been standardized by IEEE 802.1 Qau work group, and QCN is recommended as the congestion control scheme in the standard draft. QCN is heuristically designed for 1/10Gbps Ethernet without considering the impact of delays. Recent work find that QCN will encounter stability issues with feedback delays, and these issues will be more serious as Ethernet extends to 40/100Gbps and the delays become heterogeneous and time-varying. This work aims to mitigate the negative impact of delays on congestion control scheme in CE. Specially, considering the delays are heterogeneous and time-varying, we build a model for Converged Ethernet with the standard congestion management framework. The model provides a new congestion detector to estimate the real congestion status under the impact of delays and regards the heterogeneous and time-varying feature as disturbances. Leveraging the new congestion detector and tolerating the disturbance through the sliding mode control method, we design the Delay-tolerant Sliding Mode (DSM) congestion control scheme. Extensive simulations show that DSM outperforms other congestion control schemes when the Ethernet ranges from 1Gbps to 100Gbps and the delays are heterogeneous and time-varying.
Wenxue Cheng, Wanchun Jiang, Tong Zhang 0018, Bo Wang 0066, Kun Qian 0017, Fengyuan Ren
IWQoS6
2017 XpressEth: Concise and efficient converged real-time Ethernet
abstract
Owing to Ethernet's low cost, high bandwidth and architecture openness, much attention has been paid to develop converged Ethernet to support both time-critical services and conventional communication services on a unified network infrastructure. The greatest challenge here is providing low and deterministic latency for time-critical packets. Recently, the IEEE time sensitive networking task group is launched to address it. However, their framework is complex and unsuitable for commodity switch architecture. In this paper, we propose a concise and efficient converged real-time Ethernet framework called XpressEth, which leverages Dual Preemption mechanism to minimize the delay of time-critical packets, and employs a lightweight Slot Assignment Scheduler to minimize the conflicts among time-critical packets at sources. XpressEth cuts off great burden from both forwarding and scheduling. The simulation results verify that XpressEth can provide ultra-low and deterministic latency for time-critical packets (1.024μ s per hop and zero jitter in 1Gbps network), which is 13× better than time sensitive networking solution, and the side-effect on conventional communication traffic is negligible.
Kun Qian 0017, Fengyuan Ren, Danfeng Shan, Wenxue Cheng, Bo Wang 0066
IWQoS2
2017 Performance analysis of randomized data fetching in cluster computing
abstract
The shuffle transfer pattern is widely adopted in today's cluster computing applications and the completion time of each group of transmissions directly affects application performance. Because of the restriction on the number of concurrent threads and the TCP Incast problem, the randomized data fetching strategy is widely employed in this kind of communication in practice. In this paper, to assess the performance of randomized data fetching, we build a general analytical model and define two metrics - link overload probability and K-deviation load balancing probability - to evaluate the degree of link overload and load balancing respectively, since they are closely related to the transfer completion time. Leveraging our model, we theoretically analyze the transfer performance in three typical scenarios and provide recommendations for setting the number of concurrent connections per receiver. Finally, we validate the theoretical analysis as well as the recommendations through extensive simulations.
Tong Zhang 0018, Peng Cheng 0005, Wenxue Cheng, Bo Wang 0066, Fengyuan Ren
IWQoS5
2017 Towards Forward-looking Online Bitrate Adaptation for DASH
abstract
Many commercial video players rely on bitrate adaptation algorithm to adapt video bitrate to dynamic network condition. To achieve a high quality of experience, bitrate adaptation algorithm is required to strike a balance between response agility and video quality stability. Existing online algorithms select bitrates according to instantaneous throughput and buffer occupancy, achieving an agile reaction to changes but inducing video quality fluctuations due to the high dynamic of reference signals. In this paper, the idea of multi-step prediction is proposed to guide a better tradeoff, and the bitrate selection is formulated as a predictive control problem. With it, a generalized predictive control based approach is developed to calculate the optimal bitrate by minimizing the cost function over a moving look-ahead horizon. Finally, the proposed algorithm is implemented on a reference video player with performance evaluations conducted using realistic bandwidth traces. Experimental results show that the multi-step predictive control adaptation algorithm can achieve zero rebuffer event and 63.3% of reduction in bitrate switch.
Bo Wang 0066, Fengyuan Ren
ACM Multimedia2
2017 Awakening Power of Physical Layer: High Precision Time Synchronization for Industrial Ethernet
abstract
High-precision time synchronization is critical for nowadays industrial Ethernet systems. Most existing time synchronization mechanisms are implemented based on packet communication. This interaction pattern, however, greatly limits their synchronizing frequency. In order to achieve microsecond-level synchronization precision, expensive high-quality oscillator is necessary for maintaining low clock skew under this long synchronization period (usually several seconds). Furthermore, packet processing introduces many nondeterministic variances (e.g. network stack overhead), which needs to be carefully eliminated. In this paper, we propose the brand-new Industrial Time Protocol (ITP). We deploy the entire ITP in the physical layer, so it eliminates most time uncertainties caused by network stack processing. Furthermore, ITP leverages the InterFrame Gap (IFG), which is the inherent interval between any two Ethernet frames, to carry the synchronization message. With this novel design, ITP can synchronize peer devices at very high frequency without degrading the goodput. The accuracy of ITP is bounded by 16ns for two adjacent devices with only intrinsic cheap oscillator. Furthermore, our theoretical analysis deduces that ITP guarantees 16N-nanosecond accuracy for N-hop network. We implement ITP design with NetFPGA. Experiments show that ITP can provide about 76-nanosecond accuracy for #hops=16 network under severe congestions. In addition, the design of ITP is scalable. It only consumes about 0.67% of logic cells in the low-end FPGA for supporting every ITP-aware port increase.
Kun Qian 0017, Tong Zhang 0018, Fengyuan Ren
RTSS3
2017 Modeling and understanding burst transmission for energy efficient ethernet
Jinli Meng, Fengyuan Ren, Chuang Lin 0002
Comput. Networks2
2017 Throughput optimization of TCP incast congestion control in large-scale datacenter networks
Lei Xu 0019, Ke Xu 0002, Yong Jiang 0001, Fengyuan Ren
Comput. Networks4
2017 Analyzing and Enhancing Dynamic Threshold Policy of Data Center Switches
abstract
Today's data center switches usually employ on-chip shared memory; buffer management policy in them is essential to ensure fair sharing of memory among all ports. Among various polices, Dynamic Threshold (DT) policy is widely used by switch vendors. Meanwhile, in data centers, distributed applications such as MapReduce often introduce micro-burst traffic into network and the packet dropping caused by micro-burst usually leads to serious performance degradation. When micro-burst traffic arrives at switches, DT is unable to fully utilize the buffer to absorb it. Therefore, in this paper, we theoretically deduce the sufficient conditions for packet dropping caused by micro-burst traffic, and quantitatively estimate the free buffer size when packets are dropped. The results show that the free buffer size can be very large when the number of overloaded ports is small. What's worse, to ensure fair sharing of memory among output ports, packets from micro-burst traffic may be dropped even when the traffic size is much smaller than the buffer size. In light of these results, we propose the Enhanced Dynamic Threshold (EDT) policy, which can alleviate packet dropping caused by micro-burst traffic through fully utilizing the switch buffer and temporarily relaxing the fairness constraint. The simulation results show that EDT can absorb more micro-burst traffic than DT.
Danfeng Shan, Wanchun Jiang, Fengyuan Ren
IEEE Trans. Parallel Distributed Syst.3
2016 TFC: token flow control in data center networks
abstract
Services in modern data center networks pose growing performance demands. However, the widely existed special traffic patterns, such as micro-burst, highly concurrent flows, on-off pattern of flow transmission, exacerbate the performance of transport protocols. In this work, an clean-slate explicit transport control mechanism, called Token Flow Control (TFC), is proposed for data center networks to achieve high link utilization, ultra-low latency, fast convergence, and rare packets dropping. TFC uses tokens to represent the link bandwidth resource and define the concept of effective flows to stand for consumers. The total tokens will be explicitly allocated to each consumer every time slot. TFC excludes in-network buffer space from the flow pipeline and thus achieves zero-queueing. Besides, a packet delay function is added at switches to prevent packets dropping with highly concurrent flows. The performance of TFC is evaluated using both experiments on a small real testbed and large-scale simulations. The results show that TFC achieves high throughput, fast convergence, near zero-queuing and rare packets loss in various scenarios.
Jiao Zhang 0002, Fengyuan Ren, Ran Shu 0001, Peng Cheng 0005
EuroSys2
2016 Backlog-Aware SRPT Flow Scheduling in Data Center Networks
abstract
The rapidly developing soft real-time data center applications impose stringent delay requirements on internal data transfers. Therefore many recently emerged network protocols in data center share a common goal of decreasing Flow Completion Time (FCT), in which case the Shortest Remaining Processing Time (SRPT) scheduling discipline has attracted widespread attentions. However, SRPT suffers the instability issue, incurring more and more flows left uncompleted even when traffic load is within network capacity, which implies unnecessary bandwidth waste. To solve the problem, this paper proposes a backlog aware scheduling algorithm (BASRPT) that stabilizes queue length while maintaining relatively low FCT based on Lyapunov optimization. To overcome the huge computational overhead, a fast and practical approximation algorithm called fast BASRPT is also developed. Extensive flow-level simulations show that fast BASRPT indeed stabilizes switch queue and obtains a higher throughput while being able to push FCT arbitrarily close to the optimal value in the condition of feasible traffic load.
Tong Zhang 0018, Fengyuan Ren, Ran Shu 0001
ICDCS2
2016 Monitoring-Based Task Scheduling in Large-Scale SaaS Cloud
Puheng Zhang, Chuang Lin 0002, Xiao Ma 0009, Fengyuan Ren
ICSOC4
2016 Guaranteeing Delay of Live Virtual Machine Migration by Determining and Provisioning Appropriate Bandwidth
abstract
The proliferation of cloud services makes virtualization technology more important. One important feature of virtualization is live Virtual Machine (VM) migration. Two main metrics of evaluating a live VM migration mechanism are total migration time and downtime. Most existing literature on live VM migration focus on designing migration mechanisms to shorten the two metrics or making a tradeoff between them. Few of them can be applied to applications with delay requirements, such as a VM backup process that needs to be done in a specific time. This will negatively impact the user experiences and reduce the profit of cloud service providers. Besides, the frequently varied bandwidth required by the widely used pre-copy mechanism is difficult to be provided by current network technologies. In this work, we theoretically analyze how much bandwidth is required to guarantee the total migration time and downtime of a live VM migration, and then propose a novel transport control mechanism to guarantee the computed bandwidth. The experimental results demonstrate that the bandwidth obtained from the proposed reciprocal-based model guarantees the expected total migration time and downtime, and the proposed transport control mechanism ensures that the live VM migration flow obtains the expected bandwidth even if there are background flows.
Jiao Zhang 0002, Fengyuan Ren, Ran Shu 0001, Tao Huang 0005, Yunjie Liu 0001
IEEE Trans. Computers2
2016 An Energy Efficiency Perspective on Rate Adaptation for 802.11n NIC
abstract
Rate adaptation (RA) has been traditionally used to achieve high goodput. In this work, we design RA for 802.11n NICs from an energy-efficiency perspective. We show that current MIMO RA algorithms are not energy efficient for NICs despite ensuring high throughput. The fundamental problem is that, the high-throughput setting is not equivalent to the energy-efficient one. Marginal throughput gain may be realized at high energy cost. We then propose EERA and EERA+, two energy-based RA schemes that trade off goodput for energy savings at NICs. EERA applies multidimensional ternary search and simultaneous pruning to speed up its runtime convergence in single-client operations, and uses fair airtime sharing to handle multiple-client operations. EERA+ further searches for multiple, staged rates to yield more energy savings over EERA. Our experiments have confirmed their effectiveness in various scenarios.
Chi-Yu Li 0001, Chunyi Peng 0001, Peng Cheng 0005, Songwu Lu, Xinbing Wang, Fengyuan Ren, Tao Wang 0004
IEEE Trans. Mob. Comput.6
2015 Slowing Little Quickens More: Improving DCTCP for Massive Concurrent Flows
abstract
DCTCP is a potential TCP replacement to satisfy the requirements of data center network. It receives wide concerns in both academic and industrial circles. However, DCTCP could only support tens of concurrent flows well and suffers timeouts and throughput collapse facing numerous concurrent flows. This is far from the requirement of data center network. Data centers employing partition/aggregation pattern usually involve hundreds of concurrent flows. In this paper, after tracing DCTCP's dynamic behavior through experiments, we explored two roots for DCTCP's failure under the high fan-in traffic pattern: (1) The regulation mechanism of sending window is ineffective when cwnd is decreased to the minimum size, (2) The bursts induced by synchronized flows with small cwnd cause fatal packet loss leading to severe timeouts. We enhance DCTCP to support massive concurrent flows by regulating the sending time interval and desynchronizing the sending time in particular conditions. The new protocol called DCTCP+ outperforms DCTCP when the number of concurrent flows increases to several hundreds. DCTCP+ can normally work to effectively support the short concurrent query responses in the benchmark from real production clusters, and keep the same good performance with the mixture of background traffic.
Mao Miao, Peng Cheng 0005, Fengyuan Ren, Ran Shu 0001
ICPP3
2015 Comprehensive understanding of TCP Incast problem
abstract
Since TCP Incast has been identified as a catastrophic problem in many typical data center applications, a lot of efforts have been made to analyze or solve it. The analysis work intends to model Incast problem from certain perspective, and the solutions try to solve the problem through designing enhanced mechanisms or algorithms. However, the proposed models are either closely coupled with particular protocol version or dependent on empirical observations, and the solutions cannot eliminate Incast problem entirely because the underlying issues are not identified completely. There is little work which attempts to close the gap between “analyzing” and “solving”, and present a comprehensive understanding. In this paper, we provide an in-depth understanding of how TCP Incast problem happens. We build up an interpretive model which emphasizes particularly on describing qualitatively how various factors, including system parameters and mechanism variables, affect network performances in Incast traffic pattern, but not on calculating the accurate throughput. With this model, we give plausible explanations why the various solutions for TCP Incast problem can help, but do not solve it entirely.
Wen Chen 0026, Fengyuan Ren, Jing Xie 0005, Chuang Lin 0002, Kevin Yin, Fred Baker
INFOCOM2
2015 Absorbing micro-burst traffic by enhancing dynamic threshold policy of data center switches
abstract
In data center networks, micro-burst is a common traffic pattern and the packet dropping caused by it usually leads to serious performance degradation. Meanwhile, most of the current commodity switches employ on-chip shared memory, and the buffer management policies of them ensure fair sharing of memory among all ports. Among various polices, Dynamic Threshold (DT) is widely used by switch vendors. However, because DT needs to reserve a fraction of switch buffer, there is free buffer space while packets from micro-burst traffic are dropped. In this paper, we theoretically deduce the sufficient conditions for packet dropping caused by micro-burst traffic, and estimate the corresponding free buffer size. The results show that the free buffer size is very large when the number of overloaded ports is small. What's worse, to ensure fair sharing of memory among output ports, packets from micro-burst traffic may be dropped even when the traffic size is much smaller than the buffer size. In light of these results, we propose Enhanced Dynamic Threshold (EDT) policy, which can alleviate packet dropping caused by micro-burst traffic through fully utilizing the switch buffer and temporarily relaxing the fairness constraint. The simulation results show that EDT can absorb more micro-burst traffic than DT.
Danfeng Shan, Wanchun Jiang, Fengyuan Ren
INFOCOM3
2015 More load, more differentiation - A design principle for deadline-aware congestion control
abstract
Data center network has become an important facility for hosting various online services and applications, and thus its performance and underlying technologies are attracting more and more interests. In order to achieve better network performance, recent studies have proposed to tailor data center network traffic management in different aspects, devising various routing and transport schemes. In particular, for applications that must serve users in a timely manner, strict deadlines for their internal traffic flows should be met, and are explicitly taken into consideration in some latest flow rate control or scheduling algorithms in data center networks. In this paper, we advocate that when designing such deadline-aware rate control schemes, a simple principle should be followed: flows with different deadlines should be differentiated in their bandwidth allocation/occupation, and the more traffic load, the more differentiation should be made. We derive sufficient and necessary conditions for a flow rate control scheme to follow this principle, and present a simple congestion control algorithm called Load Proportional Differentiation (LPD) as its application. We have evaluated LPD under different topologies and load scenarios, both by simulation and in real testbed. Our results show that LPD nearly always outperforms D2TCP, a latest deadline-aware rate control scheme, and often reduces the number of flows missing their deadlines by more than 50%. We also give some other applications of this principle, for example, in reducing flow completion time.
Han Zhang 0009, Xingang Shi, Xia Yin 0001, Fengyuan Ren
INFOCOM4
2015 Enhancing TCP Incast congestion control over large-scale datacenter networks
abstract
Many-to-one traffic pattern in datacenter networks introduces the problem of Incast congestion for Transmission Control Protocol (TCP) and puts unprecedented pressure to the cloud service providers. To address heavy Incast, we present an Receiver-oriented Datacenter TCP (RDTCP). The proposal is motivated by oscillatory queue size when handling heavy Incast traffic and substantial potential of receiver in congestion control. Finally, RDTCP adopts both open- and closed-loop congestion controls. We provide a systematic discussion on its design issues and implement a prototype to examine its performance. The evaluation results indicate that RDTCP has an average decrease of 47.5% in the mean queue size, 51.2% in the 99th-percentile latency in the increasingly heavy Incast over TCP, and 43.6% and 11.7% over Incast congestion Control for TCP (ICTCP).
Lei Xu 0019, Ke Xu 0002, Yong Jiang 0001, Fengyuan Ren
IWQoS4
2015 Congestion-aware adaptive forwarding in datacenter networks
Jiao Zhang 0002, Fengyuan Ren, Tao Huang 0005, Yunjie Liu 0001
Comput. Commun.2
2015 Sliding Mode Congestion Control for Data Center Ethernet Networks
abstract
Recently, Ethernet is enhanced as the unified switch fabric of data centers, called data center Ethernet. One of the indispensable enhancements is end-to-end congestion management, and currently quantized congestion notification (QCN) has been ratified as the corresponding standard. However, our experiments show that QCN suffers from large oscillations of the queue length at the bottleneck link such that the buffer is emptied frequently and accordingly the link utilization degrades, with certain system parameters and network configurations. This phenomenon is corresponding to our theoretical analysis result that QCN fails to enter into the sliding mode motion (SMM) pattern with certain system parameters and network configurations. Knowing the drawbacks of QCN and realizing the advantage that congestion management system is insensitive to the changes of parameters and network configurations in the SMM pattern, we present sliding mode congestion control (SMCC), which can enter into the SMM pattern under any conditions. SMCC is simple, stable, fair, has short response time, and can be easily used to replace QCN because both of them follow the framework developed by the IEEE 802.1Qau work group. Experiments on the NetFPGA platform show that SMCC is superior to QCN, especially when traffic pattern and network states are variable.
Wanchun Jiang, Fengyuan Ren, Ran Shu 0001, Yongwei Wu 0001, Chuang Lin 0002
IEEE Trans. Computers2
2015 Dynamic Routing for Data Integrity and Delay Differentiated Services in Wireless Sensor Networks
abstract
Applications running on the same Wireless Sensor Network (WSN) platform usually have different Quality of Service (QoS) requirements. Two basic requirements are low delay and high data integrity. However, in most situations, these two requirements cannot be satisfied simultaneously. In this paper, based on the concept ofpotentialin physics, we propose IDDR, a multi-path dynamic routing algorithm, to resolve this conflict. By constructing a virtual hybrid potential field, IDDR separates packets of applications with different QoS requirements according to the weight assigned to each packet, and routes them towards the sink through different paths to improve the data fidelity for integrity-sensitive applications as well as reduce the end-to-end delay for delay-sensitive ones. Using the Lyapunov drift technique, we prove that IDDR is stable. Simulation results demonstrate that IDDR provides data integrity and delay differentiated services.
Jiao Zhang 0002, Fengyuan Ren, Hongkun Yang, Chuang Lin 0002
IEEE Trans. Mob. Comput.2
2015 Phase Plane Analysis of Quantized Congestion Notification for Data Center Ethernet
abstract
Currently, Ethernet is being enhanced to become the unified switch fabric in data centers. With the unified switch fabric, the cost on redundant devices is reduced, while the design and management of data center networks are simplified. Congestion management is one of the indispensable enhancements on Ethernet, and Quantized Congestion Notification (QCN) has just been ratified as the formal standard. Though QCN has been investigated for several years, there exist few in-depth theoretical analyses on QCN. The most possible reason is that QCN is heuristically designed and involves the property of variable structure. The classic linear analysis method is incapable of handling the segmented nonlinearity of the variable structure system. In this paper, we use the phase plane method, which is suitable for systems of segmented nonlinearity, to analyze the QCN system. The overall dynamic behaviors of the QCN system are presented, and the sufficient conditions for the stable QCN system are deduced. These sufficient conditions serve as guidelines toward proper parameters setting. Moreover, we find that the stability of QCN is mainly promised by the sliding mode motion, which is the underlying reason for QCN's stable queue shown in numerous simulations and experiments. Experiments on the NetFPGA platform verify that the analytical results can explain the complex behaviors of QCN.
Wanchun Jiang, Fengyuan Ren, Chuang Lin 0002
IEEE/ACM Trans. Netw.2
2015 Modeling and Solving TCP Incast Problem in Data Center Networks
abstract
TCP Incast problem attracts much attention due to the catastrophic goodput drop. In this paper, a goodput model of the problem is built to understand why goodput collapse occurs and a solution to the problem based on the theoretical analysis is proposed. We found that the TCP Incast goodput deterioration is mainly caused by two types of timeouts, one happens at the tail of data blocks and dominates the goodput when the number of senders is small, while the other one at the head of data blocks and governs the goodput when the number of senders is large. The proposed model describes the relationship between these two types of timeouts and the Incast communication pattern, block size, bottleneck buffer size, and so on. The simulation results indicate that the model well characterizes the features of the TCP Incast problem. Enlightened by the analysis, a PRiority-based solution to the TCP INcast problem (PRIN) is proposed, which avoids timeouts at the head of blocks by reducing TCP send window and prevents timeouts at the tail of blocks by leveraging priority technology. The experimental results show that PRIN solves the TCP Incast problem.
Jiao Zhang 0002, Fengyuan Ren, Chuang Lin 0002
IEEE Trans. Parallel Distributed Syst.2
2015 Relocation routing for energy balancing in mobile sensor networks
abstract
Abstract Wireless sensor networks (WSNs) have been widely investigated in the past decades because of its applicability in various extreme environments. As sensors use battery, most works on WSNs focus on energy efficiency issues (e.g., local energy balancing problems) in statically deployed WSNs. Few works have paid attention to the global energy balancing problem for the scenario that mobile sensor nodes can move freely. In this paper, we propose a new routing protocol called global energy balancing routing protocol (GEBRP) based on an active network framework and node relocation in mobile sensor networks. This protocol achieves global energy efficiency by repairing coverage holes and replacing invalid nodes dynamically. Simulation and experiment results demonstrate that the proposed GEBRP achieves superior performance over the existing scheme. In addition, we analyze the delay performance of GEBRP and study how the delay performance is affected by various system parameters.Copyright © 2013 John Wiley & Sons, Ltd.
Yiping Deng, Chuang Lin 0002, Dapeng Oliver Wu, Fengyuan Ren
Wirel. Commun. Mob. Comput.4
2014 Delay guaranteed live migration of Virtual Machines
abstract
The proliferation of cloud services makes virtualization technology more important. One important feature of virtualization is live Virtual Machine (VM) migration, which can be employed to facilitate load balancing, fault management and server maintenance etc. Two main metrics of evaluating a live VM migration mechanism are total migration time and downtime. The existing literature on live VM migration mainly focus on designing migration mechanisms to shorten these two metrics or making a tradeoff between them. Few of them can be applied to the applications with delay requirements, such as, delay-sensitive web services or a VM backup process that needs to be done in a specific time. This will not only negatively impact the user experiences, but also reduce the profit of cloud service providers. Besides, the frequently varied bandwidth required by the widely used pre-copy mechanism is difficult to be provided by current network technologies. In this work, we theoretically analyze how much bandwidth is required to guarantee the total migration time and downtime of a live VM migration. We first propose a deterministic-based model as a simple example, then assume that the dirtying frequency of each page obeys the bernoulli distribution. At last, we analyze the statistic features of the typical workload running in a VM and build a reciprocal-based workload model, and theoretically give the required bandwidth value to satisfy the performance metrics of a live VM migration. The experimental results demonstrate that the bandwidth obtained from the reciprocal-based model can guarantee the expected total migration time and downtime.
Jiao Zhang 0002, Fengyuan Ren, Chuang Lin 0002
INFOCOM2
2014 Analysing convergence of Quantized Congestion Notification in Data Center Ethernet
abstract
Enhancing Ethernet as the unified data center fabric to concurrently handle the traffic of Local Area Network (LAN), Storage Area Network (SAN), and High Performance Computing (HPC) has attracted much attention. Congestion management is one critical enhancement to fill the performance gap between traditional Ethernet and the unified data center fabric. Currently, Quantized Congestion Notification (QCN) has been approved as the standard congestion management mechanism. However, lots of work pointed out that QCN suffers from the problem of unfairness among different flows. In this paper, we found that QCN could achieve fairness, merely the convergence time to fairness is quite long. Thus, we build a convergence time model to investigate the reasons of the slow convergence process of QCN. The model indicates that the convergence time of QCN can be decreased if RPs have the same rate increase probability or the rate increase step becomes larger at steady state. We validate the precise of our model by comparing with experimental data on the NetFPGA platform. The results show that it well characterizes the convergence time to fairness of QCN. Based on the proposed model, the impact of QCN parameters, network parameters, and QCN variants on the convergence time is analysed. Finally, enlightened by the analysis, we proposed a mechanism, called QCN-T, which replaces the Byte Counter and Timer at sources with a single modified Timer, to reduce the convergence time of QCN.
Ran Shu 0001, Jiao Zhang 0002, Fengyuan Ren, Chuang Lin 0002
IWQoS3
2014 Catch the Whole Lot in an Action: Rapid Precise Packet Loss Notification in Data Center
Peng Cheng 0005, Fengyuan Ren, Ran Shu 0001, Chuang Lin 0002
NSDI2
2014 HEIR: Heterogeneous interference recognition for wireless sensor networks
abstract
With the rapid development of wireless communication technology, a large number of wireless networks and devices that have different PHY and MAC layers coexist with each other. 2.4GHz Industrial, Scientific and Medical (ISM) band is becoming increasingly crowded. Wireless Sensor Networks (WSNs) which use low-power communication standard IEEE802.15.4 share the unlicensed spectrum with a plethora of other devices and technologies, such as WiFi systems underlying IEEE802.11, Bluetooth systems underlying IEEE802.15.1, and even non-communication appliance like microwave ovens. The ability to detect what radios are operating in the neighborhood is a fundamental need of WSNs, ranging from network management to network security. Since there are no explicit mechanisms to recognize such heterogeneous interference sources, WSNs often have no reasonable way to guard against them. In this paper, we describe the main working principle of heterogeneous interference sources, and present HEIR, a detector that is able to accurately identify heterogeneous interference sources. HEIR builds on the insight that there are hidden repeating patterns of signals which can be used to construct unique signatures and identify different types in most wireless protocols. The method can be implemented through signal sampling, interference estimation, feature extraction, and device classification. We show the experimental evaluation in an indoor testbed that HEIR is accurate in several different scenarios, and can live on sensor nodes' hardware. Since no any channel changes, the network topology is not interrupted, and the stable communication in real-time is ensured then.
Meng Hou, Fengyuan Ren, Chuang Lin 0002, Mao Miao
WoWMoM2
2014 Sharing Bandwidth by Allocating Switch Buffer in Data Center Networks
abstract
In today's data centers, the round trip propagation delay is quite small. Therefore, switch buffer sizes are much larger than the Bandwidth Delay Product (BDP). Based on this observation, in this paper we introduce a new transport protocol which provides bandwidth Sharing by Allocating switch Buffer (SAB) for data centers. SAB sets the congestion windows for flows based on the buffer size of the switches along the path. On one hand, as long as the total buffer allocated to all the flows is larger than the BDP, the network bandwidth can be fully utilized. On the other hand, since SAB only allocates the buffer space to flows, the totally injected traffic will not exceed the network capacity. Thus, SAB rarely loses packets. SAB also reduces flow completion time by allowing flows to reach their fair share of bandwidth quickly. The results of a series of experiments and simulations demonstrate that SAB has the features of fast convergence and rare packet loss. It reduces the latency of short flows and solves theTCP Incast and TCP Outcast problems.
Jiao Zhang 0002, Fengyuan Ren, Xin Yue, Ran Shu 0001, Chuang Lin 0002
IEEE J. Sel. Areas Commun.2
2014 Analysis of Backward Congestion Notification with Delay for Enhanced Ethernet Networks
abstract
At present, companies and standards organizations are enhancing Ethernet as the unified switch fabric for all of the TCP/IP traffic, the storage traffic and the high performance computing traffic in data centers. Backward congestion notification (BCN) is the basic mechanism for the end-to-end congestion management enhancement of Ethernet. To fulfill the special requirements of the unified switch fabric, i.e., losslessness and low transmission delay, BCN should hold the buffer occupancy around a target point tightly. Thus, the stability of the control loop and the buffer size are critical to BCN. Currently, the impacts of delay on the performance of BCN are unidentified. When the speed of Ethernet increases to 40 Gbps or 100 Gbps in the near future, the number of on-the-fly packets becomes the same order with the buffer size of switch. Accordingly, the impacts of delay will become significant. In this paper, we analyze BCN, paying special attention on the delay. We model the BCN system with a set of segmented delayed differential equations, and then deduce sufficient condition for the uniformly asymptotic stability of BCN. Subsequently, the bounds of buffer occupancy are estimated, which provides direct guidelines on setting buffer size. Finally, numerical analysis and experiments on the NetFPGA platform verify our theoretical analysis.
Wanchun Jiang, Fengyuan Ren, Yongwei Wu 0001, Chuang Lin 0002, Ivan Stojmenovic
IEEE Trans. Computers2
2014 Optimal power scheduling in 802.11n wireless networks for real-time services
abstract
ABSTRACT The growing popularity of mobile devices in our daily life demands higher throughput of wireless networks. The new communication standard 802.11n has significantly improved throughput because of the use of advanced technologies such as the multiple‐input multiple‐output communication technique. Because mobile devices are usually battery‐operated, power efficiency is critical; on the other hand, delay performance can be improved by transmitting at high power. To address the conflicting requirement of power saving and small delay, power scheduling is needed. In the past, many approaches to power scheduling have been proposed for real‐time applications, but few of them have considered complicated modes of channel state information(CSI) in multiple‐input multiple‐output. In this paper, we study this and classify the CSI into four types, namely, constant, slow fading, fast fading, and unknown. For known CSI, we propose an optimal algorithm for power scheduling. For unknown CSI, we propose an approximate algorithm based on some heuristics. To improve resource utilization, a stochastic delay‐bound method is proposed for fast‐fading condition. Simulation results demonstrate that the performance achieved by the optimal and heuristic algorithms agrees well with the analysis. Copyright © 2012 John Wiley & Sons, Ltd.
Yiping Deng, Chuang Lin 0002, Fengyuan Ren, Dapeng Oliver Wu
Wirel. Commun. Mob. Comput.3
2013 Selecting a preferable access point with more available bandwidth
abstract
It is a common problem for Wireless LAN (WLAN) users when they face more than one available Access Point (AP): which one may serve them better? In the real world, most WLAN devices and users select the AP by Received Signal Strength (RSS), which doesn't consider the load of APs. Therefore, the RSS-based scheme may result in load imbalance and utilization degradation. This paper proposes a simple scheme to inform WLAN users the available bandwidth when selecting a specific AP. Unlike the previous work, the new scheme concerns about the aggregated traffic patterns in AP, which affects the amount of the available bandwidth. Considering the number of active stations and influence of collision, the available bandwidth is calculated and regarded as a selection metric. Common users can acquire this metric through Service Set IDentifier (SSID) without any modifications on their devices, and users who accept modifications can select the preferable AP dynamically and automatically. This scheme is very simple to be implemented and deployed with negligible overhead to network. We conduct simulations to verify this scheme, which can improve by 200% the user's throughput in some situations.
Shibo Xu, Fengyuan Ren, Yinsheng Xu, Chuang Lin 0002
ICC2
2013 Ease the Queue Oscillation: Analysis and Enhancement of DCTCP
abstract
Because of the terrible performance of TCP protocol in data center environment, DCTCP has been proposed as a TCP replacement, which uses a simple marking mechanism at switches and a few amendments at end hosts to adjust congestion window based on the extent of the congestion in networks. Thus, DCTCP can make a proper tradeoff between high throughput and low latency. However, through our observation, we discover that DCTCP causes severe oscillation of queue under some parameters and network configuration. Our perceptual analysis concludes that the rough single-threshold marking mechanism may be the essential reason. Therefore, we propose Double-Threshold DCTCP as an improvement of DCTCP. Then, by applying describing function method in nonlinear control theory, we analyze the stability of both DCTCP and Double-Threshold DCTCP, and theoretically explain why Double-Threshold DCTCP is more stable than DCTCP. At last, we validate theoretical analysis and conclude that the Double- Threshold DCTCP can achieve smaller queue, and the queue length of Double-Threshold DCTCP is less sensitive to the growing number of flows. Further, Double-Threshold DCTCP can postpone the throughput collapse caused by Incast traffic and reduce the tail latency in completion time experiment.
Wen Chen 0026, Peng Cheng 0005, Fengyuan Ren, Ran Shu 0001, Chuang Lin 0002
ICDCS3
2013 Taming TCP incast throughput collapse in data center networks
abstract
The TCP incast problem attracts a lot of attention due to its wide existence in cloud services and catastrophic performance degradation. Some effort has been made to solve it. However, the industry is still struggling with it, such as Facebook. Based on the investigation that the TCP incast problem is mainly caused by the TimeOuts (TOs) occurring at the boundary of the stripe units, this paper presents a simple and effective TCP enhanced mechanism, called GIP (Guarantee Important Packets), for the applications with the TCP incast problem. The main idea is making TCP aware of the boundaries of the stripe units, and reducing the congestion window of each flow at the start of each stripe unit as well as redundantly transmitting the last packet of each stripe unit. GIP modifies TCP a little at the end hosts, thus it can be easily implemented. Also, it poses no impact on the other TCP-based applications. The results of both experiments on our testbed and simulations on the ns-2 platform demonstrate that TCP with GIP can avoid almost all of the TOs and achieve high goodput for applications with the incast communication pattern.
Jiao Zhang 0002, Fengyuan Ren, Chuang Lin 0002
ICNP2
2013 Modeling and understanding burst transmission algorithms for energy efficient ethernet
abstract
Recently, the energy consumption of Ethernet has become one of the hottest topics focused by both academic committee and industry, especially with the increase of the link speed from 1Gbps to 10Gbps nowadays or even 40/100/200Gbps in the near future. To save the energy consumed by the Ethernet, the Energy Efficient Ethernet (EEE) is developed and standardized by the IEEE 802.3az work group. When there is no incoming traffic, the EEE can saves 90 % of its energy consumption by entering into the Low Power Idle (LPI) mode. To maximize the energy saving of Ethernet, the Burst TRansmission (BTR) algorithm, which defines a new way to utilize the LPI mode, is developed as a policy for EEE. Prior work theoretically shows that the BTR algorithm makes a tradeoff between the energy saving and the queuing delay. However, the traffic pattern, on which the performance of EEE greatly depends, is assumed to be deterministic in their analyses. Besides, their models made estimation for many situations. In this paper, assuming that the arrival time of packets can be modeled by Poisson process, we build Markov model for EEE with the BTR algorithm and provide analytical understanding on the BTR algorithm. We propose two actual models: one focuses on the buffer size limit, the other concentrates on tolerable packet delay additionally. We draw some guidelines of parameter selection and policy design for EEE from combination of theory conclusions and simulation results. The results show that the saved energy can be constrained by link occupancy even though the buffer size is variational. The other policy buffer full triggered wake-up can achieve ideal ratio of energy consumption and arrival rate within the scope of the buffer as well. However, the tolerable delay can not be guaranteed by any policies. The buffer size is even fixed, which affects the flexibility of demanded delay for different business. The policy considering tolerable delay is supposed to be a little better than the other policy, with a little more complicated design. Thus we design an adaptive policy: detect the load utilization, apply the buffer full triggered wake-up policy for higher load utilization link, while applying the buffer full and timeout triggered wakeup policy for the delay sensitive business and tiny arrival rate.
Jinli Meng, Fengyuan Ren, Wanchun Jiang, Chuang Lin 0002
IWQoS2
2013 Frequency Domain Packet Scheduling with Stability Analysis for 3GPP LTE Uplink
abstract
In this paper, we investigate the Frequency Domain Packet Scheduling (FDPS) problem for 3GPP Long Term Evolution (LTE) Uplink (UL). Instead of studying a specific scheduling policy, we provide a unified approach to tackle this issue. First, we formalize a general LTE UL FDPS problem, which is suitable for various scheduling policies. Then, we prove that the problem is MAX SNP-hard, which implies that approximation algorithms with constant approximation ratios are the best that we can hope for. Therefore, we design two approximation algorithms, both of which have polynomial runtime. The first algorithm is based on a simple greedy method. The second one is based on the Local Ratio (L-R) technique and it can approximately solve the LTE UL FDPS problem with an approximation ratio of 2. To further analyze the stability of the 2-approximation L-R algorithm, we derive a specific FDPS problem, which incorporates the queue length and channel quality information. We utilize the Lyapunov Drift to prove the L-R algorithm is stable for any $((\omega_0, \epsilon_0))$-admissible LTE UL systems. The simulation results indicate good performance of the L-R scheduler.
Fengyuan Ren, Yinsheng Xu, Hongkun Yang, Jiao Zhang 0002, Chuang Lin 0002
IEEE Trans. Mob. Comput.1
2013 Real-time routing in wireless sensor networks: A potential field approach
abstract
Wireless Sensor Networks (WSNs) are embracing an increasing number of real-time applications subject to strict delay constraints. Utilizing the methodology of potential field in physics, in this article we effectively address the challenges of real-time routing in WSNs. In particular, based on a virtual composite potential field, we propose the Potential-based Real-Time Routing (PRTR) protocol that supports real-time routing using multipath transmission. PRTR minimizes delay for real-time traffic and alleviates possible congestions simultaneously. Since the delay bounds of real-time flows are extremely important, the end-to-end delay bound for a single flow is derived based on the Network Calculus theory. The simulation results show that PRTR minimizes the end-to-end delay for real-time routing, and also guarantees a tight bound on the delay.
Yinsheng Xu, Fengyuan Ren, Tao He 0008, Chuang Lin 0002, Canfeng Chen, Sajal K. Das 0001
ACM Trans. Sens. Networks2
2013 Attribute-Aware Data Aggregation Using Potential-Based Dynamic Routing in Wireless Sensor Networks
abstract
The resources especially energy in wireless sensor networks (WSNs) are quite limited. Since sensor nodes are usually much dense, data sampled by sensor nodes have much redundancy, data aggregation becomes an effective method to eliminate redundancy, minimize the number of transmission, and then to save energy. Many applications can be deployed in WSNs and various sensors are embedded in nodes, the packets generated by heterogenous sensors or different applications have different attributes. The packets from different applications cannot be aggregated. Otherwise, most data aggregation schemes employ static routing protocols, which cannot dynamically or intentionally forward packets according to network state or packet types. The spatial isolation caused by static routing protocol is unfavorable to data aggregation. To make data aggregation more efficient, in this paper, we introduce the concept of packet attribute, defined as the identifier of the data sampled by different kinds of sensors or applications, and then propose an attribute-aware data aggregation (ADA) scheme consisting of a packet-driven timing algorithm and a special dynamic routing protocol. Inspired by the concept of potential in physics and pheromone in ant colony, a potential-based dynamic routing is elaborated to support an ADA strategy. The performance evaluation results in series of scenarios verify that the ADA scheme can make the packets with the same attribute spatially convergent as much as possible and therefore improve the efficiency of data aggregation. Furthermore, the ADA scheme also offers other properties, such as scalable with respect to network size and adaptable for tracking mobile events.
Fengyuan Ren, Jiao Zhang 0002, Yongwei Wu 0001, Tao He 0008, Canfeng Chen, Chuang Lin 0002
IEEE Trans. Parallel Distributed Syst.1
2013 Frequency Domain Packet Scheduling with MIMO for 3GPP LTE Downlink
abstract
In this paper, we formalize a general Frequency Domain Packet Scheduling (FDPS) problem for 3GPP LTE Downlink (DL). The DL FDPS problem incorporates the SingleUser Multiple Input Multiple Output (SU-MIMO) technique, and can express various scheduling policies, including the Proportional-Fair metric, the MaxWeight scheduling, etc. For LTE DL SU-MIMO, the constraint of selecting only one MIMO mode (transmit diversity or spatial multiplexing) per user in each transmission time interval (TTI) increases the hardness of the FDPS problem. We prove the problem is MAX SNP-hard, which implies approximation algorithms with constant approximation ratios are the best we can expect. Subsequently, we propose an approximation algorithm of polynomial runtime. The solution is based on a greedy method for maximizing a non-decreasing submodular function over a matroid. The algorithm can solve the general DL FDPS problem with an approximation ratio of 4. We implement the proposed algorithm and compare its performance with other well-known schedulers.
Yinsheng Xu, Hongkun Yang, Fengyuan Ren, Chuang Lin 0002, Xuemin Shen
IEEE Trans. Wirel. Commun.3
2012 Analysis of backward congestion notification with delay for enhanced ethernet networks
abstract
Recently, companies and standards organizations are enhancing Ethernet as the unified switch fabric for all of the TCP/IP traffic, the storage traffic and the interprocess communication(IPC) traffic in Data Center Networks(DCNs). Backward Congestion Notification(BCN) is the basic mechanism for the end-to-end congestion management enhancement. To fulfill the special requirements of the unified switch fabric that being lossless and of extremely low latency, BCN should hold the queue length around a target point tightly. Thus, the stability of the control loop and the buffer size are critical to BCN. Currently, the impacts of delay on the performance of BCN are unidentified. When the link capacity increases to 40Gbps or 100Gbps in the near future, the number of on-the-fly packets becomes the same order with the shallow buffer size of switches. Thus, the impacts of delay on the performance of BCN will become significant. In this paper, we analyze BCN, paying special attention on the delay. Firstly, we model the BCN system with a set of segmented delayed differential equations. Then, the sufficient condition for the uniformly asymptotic stability of the BCN system is deduced. Subsequently, the bound of buffer occupancy under this sufficient condition are estimated, which provides guidelines on setting buffer size. Finally, the numerical analysis and the experiments on the NetFPGA platform verify the theoretical analysis.
Wanchun Jiang, Fengyuan Ren, Chuang Lin 0002, Ivan Stojmenovic
INFOCOM2
2012 Sliding Mode Congestion Control for data center Ethernet networks
abstract
Recently, Ethernet is being enhanced as the unified switch fabric of data centers, called Data Center Ethernet. The end-to-end congestion management is one of the indispensable enhancements, and Quantized Congestion Notification (QCN) has been ratified to be the standard. Our experiments show that QCN suffers from the oscillation of the queue at the bottleneck link. With the changes of system parameters and network configurations, the oscillation may become so serious that the queue is emptied frequently. As a result, the utilization of the bottleneck link degrades. Theoretical analysis shows that QCN approaches to the equilibrium point mainly through the sliding mode motion. But whether QCN enters into the sliding mode motion also depends on both system parameters and network configurations. Hence, we present the Sliding Mode Congestion Control (SMCC) scheme, which can drive the system into the sliding mode motion under any conditions. SMCC benefits from the advantage that the sliding mode motion is insensitive to system parameters and external disturbances. Moreover, SMCC is simple, stable and has short response time. QCN can be replaced by SMCC easily since both of them follow the framework developed by the IEEE 802.1 Qau work group. Experiments on the NetFPGA platform show that SMCC is superior to QCN, especially in the condition that the traffic pattern and the network state are variable.
Wanchun Jiang, Fengyuan Ren, Ran Shu 0001, Chuang Lin 0002
INFOCOM2
2012 Investigating the interacting two-way tcp connections over 3GPP LTE networks
abstract
This paper investigates the interactions between two-way TCP connections over 3GPP LTE networks. In the LTE network, the two-way TCP flows share buffers on a common bottleneck, i.e., the radio access links. The behaviors of TCPs significantly influence the others in the opposite direction. Specifically, the radio links of LTE are asymmetric, which may induce drastic interactions of TCPs and rapid draining of downlink buffer. The periodic idleness of downlink is a huge waste of the precious radio bandwidth and results in considerable performance degradation. In the viewpoint of Coupled Queues, we thoroughly understand the interacting TCPs and explain the reason for performance degradation. Based on a straightforward modeling procedure, we formalize the evolution of two-way TCPs and model the bottleneck queue size in every slot. The model indicates the queues are close coupled, which is verified with simulations on NS2. If the uplink (queue) is fully utilized, the downlink (queue) will always be underutilized even idle, and vice versa. Furthermore, an effective solution called Preemptive ACK Queueing (PAQ) is designed to decouple the queues, which improves the performance of two-way TCPs over LTE networks.
Yinsheng Xu, Fengyuan Ren, Shibo Xu, Chuang Lin 0002, Sajal K. Das 0001
MSWiM2
2012 Retransmission or redundancy: Transmission reliability study in wireless sensor networks
Hao Wen 0014, Chuang Lin 0002, Fengyuan Ren, Yao Yue, Xiaomeng Huang
Sci. China Inf. Sci.3
2011 Stochastic Delay Bound for Heterogeneous Aggregation in Sensor Networks
abstract
Strict delay performance guarantees are required by many applications in wireless sensor networks. Different from traditional approaches, the data aggregation and the stochastic characteristic need to be considered for the delay analysis in sensor networks. In this paper, the problem that how to calculate the stochastic delay bound under different aggregation schemes is solved with the stochastic network calculus. Meanwhile, to support multifunction in sensor networks, the impact of the heterogeneous services is brought into the analysis of the stochastic delay bound. From some numerical evaluations and the simulation, it is shown that a good tradeoff between performance and implementation can be achieved with a comprehensive aggregation scheme.
Yiping Deng, Chuang Lin 0002, Fengyuan Ren
GLOBECOM3
2011 Modeling and understanding TCP incast in data center networks
abstract
Recently, TCP incast problem attracts increasing attention since the receiver suffers drastic goodput drop when it simultaneously strips data over multiple servers. Lots of attempts have been made to address the problem through experiments and simulations. However, to the best of our knowledge, few solutions can solve it fundamentally at low cost. In this paper, a goodput model of TCP incast is built to understand why goodput collapse occurs. We conclude that TCP incast goodput deterioration is mainly caused by two types of timeouts, one happens at the tail of a data block and dominates the goodput when the number of senders is small, while the other one at the head of a data block and governs the goodput when the number of senders is large. The proposed model describes the causes of these two types of timeouts which are related to the incast communication pattern, block size, bottleneck buffer and so on. We validate the proposed model by comparing with simulation data, finding that it can well characterize the features of TCP incast. We also discuss the impact of most parameters on the goodput of TCP incast.
Jiao Zhang 0002, Fengyuan Ren, Chuang Lin 0002
INFOCOM2
2011 A-ADHOC: An Adaptive Real-time Distributed MAC Protocol for Vehicular Ad Hoc Networks
Jia Liu 0024, Fengyuan Ren, Limin Miao, Chuang Lin 0002
Mob. Networks Appl.2
2011 Modeling and Improving TCP Performance over Cellular Link with Variable Bandwidth
abstract
To facilitate a viable evolution of cellular networks toward extensive packet data traffic, the High Speed Downlink Packet Access (HSDPA) technology is introduced. The various link adaptation techniques employed by HSDPA augment the bandwidth variation, which is identified as one of the most important factors resulting in the deterioration of TCP performance. In this paper, we firstly build an analytical model of TCP throughput to explain why the bandwidth variation degrades the TCP performance. Subsequently, a split-connection Window Adaptation TCP Proxy is proposed to improve the TCP throughput in HSDPA networks. To use the precious cellular link resources sufficiently, the length of the queue in Node-B is intentionally kept around the reference value through adaptively adjusting the sending window size of TCP proxy based on the dynamic values of varying bandwidth. Since both the disturbance caused by bandwidth variation and the feedback delay are prone to lead an unstable queue system, the robust sliding mode variable structure control theory is employed to design the proper control law to weaken the impact of noise and delay on the stability of the queue system in Node-B. The theoretical analysis and the enhanced scheme are verified through simulation experiments. The simulation results show that our TCP proxy is able to resist against bandwidth oscillation and improve the cellular link utilization.
Fengyuan Ren, Chuang Lin 0002
IEEE Trans. Mob. Comput.1
2011 Traffic-Aware Dynamic Routing to Alleviate Congestion in Wireless Sensor Networks
abstract
The congestion problem in Wireless Sensor Networks (WSNs) is quite different from that in traditional networks. Most current congestion control algorithms try to alleviate the congestion by reducing the rate at which the source nodes inject packets into the network. However, this traffic control scheme always decreases the throughput so as to violate fidelity level required by the applications. In this paper, we present a solution that sufficiently exerts the idle or underloaded nodes to alleviate congestion and improve the overall throughput in WSNs. To achieve this goal, a traffic-aware dynamic routing (TADR) algorithm is proposed to route packets around the congestion areas and scatter the excessive packets along multiple paths consisting of idle and underloaded nodes. Utilizing the concept of potential in classical physics, our TADR algorithm is designed through constructing a hybrid virtual potential field using depth and normalized queue length to force the packets to steer clear of obstacles created by congestion and eventually move toward the sink. The simulation results show that the proposed solution improves the overall throughput by around 370 percent as compared to MintRoute, which is one of benchmark routing protocols. Furthermore, TADR scheme has low overhead suitable for large-scale, dense sensor networks.
Fengyuan Ren, Tao He 0008, Sajal K. Das 0001, Chuang Lin 0002
IEEE Trans. Parallel Distributed Syst.1
2011 EBRP: Energy-Balanced Routing Protocol for Data Gathering in Wireless Sensor Networks
abstract
Energy is an extremely critical resource for battery-powered wireless sensor networks (WSN), thus making energy-efficient protocol design a key challenging problem. Most of the existing energy-efficient routing protocols always forward packets along the minimum energy path to the sink to merely minimize energy consumption, which causes an unbalanced distribution of residual energy among sensor nodes, and eventually results in a network partition. In this paper, with the help of the concept of potential in physics, we design an Energy-Balanced Routing Protocol (EBRP) by constructing a mixed virtual potential field in terms of depth, energy density, and residual energy. The goal of this basic approach is to force packets to move toward the sink through the dense energy area so as to protect the nodes with relatively low residual energy. To address the routing loop problem emerging in this basic algorithm, enhanced mechanisms are proposed to detect and eliminate loops. The basic algorithm and loop elimination mechanism are first validated through extensive simulation experiments. Finally, the integrated performance of the full potential-based energy-balanced routing algorithm is evaluated through numerous simulations in a random deployed network running event-driven applications, the impact of the parameters on the performance is examined and guidelines for parameter settings are summarized. Our experimental results show that there are significant improvements in energy balance, network lifetime, coverage ratio, and throughput as compared to the commonly used energy-efficient routing algorithm.
Fengyuan Ren, Jiao Zhang 0002, Tao He 0008, Chuang Lin 0002, Sajal K. Das 0001
IEEE Trans. Parallel Distributed Syst.1
2010 Effective Data Aggregation Supported by Dynamic Routing in Wireless Sensor Networks
abstract
Data aggregation is an main method to conserve energy in wireless sensor network (WSN). Prior work on data aggregation protocols are generally based on static routing schemes, such as tree-based, cluster-based or chain-based routing schemes. Although they can save energy to some extent, in dynamic scenarios where the source nodes are changing frequently, they will not only incur high overhead to continuously reconstruct the routing but also can not reduce the communication overhead effectively. Our work aims to design an effective data aggregation mechanism supported by dynamic routing (DASDR) which can adapt to different scenarios without incurring much overhead. Enlightened by the concept of potential field in the discipline of physics, the dynamic routing in DASDR is designed based on two potential fields: depth potential field which guarantees packets reaching the sink at last and queue potential field which makes packets more spatially convergent and thus data aggregation will be more efficient. Simulation results show that DASDR is more effective in energy savings as well as scales well with regard to the network size.
Jiao Zhang 0002, Qian Wu 0001, Fengyuan Ren, Tao He 0008, Chuang Lin 0002
ICC3
2010 Exploiting Gaps of Real-Time Traffic to Improve Handover Performance in Wireless Networks
abstract
The handover mechanism in wireless networks has been intensively investigated in the past years. The traditional solutions are to customize various protocols in different layers to provide uninterrupted communication, which is particularly important for real time traffic. In this paper, we intend to exploit silence gap, inherently existing in realtime traffic, to improve handover performance. The proposed strategy can be appended to most of handover mechanisms without extra overhead. We theoretically analyze the relationship between gap distribution and handover performance under the assumption of different statistical distribution of gaps on a case-by-case basis. Subsequently, we insert an additional module into the WiMAX network and conduct the simulations using real voice traces. The results confirm the analytical conclusions, showing that the gap-exploiting approach can obtain significant gain in terms of packet loss particularly in low-speed mobile environment.
Jia Liu 0024, Chuang Lin 0002, Fengyuan Ren
ICCCN3
2010 Phase Plane Analysis of Congestion Control in Data Center Ethernet Networks
abstract
Ethernet has some attractive properties for network consolidation in the data center, but needs further enhancement to satisfy the additional requirements of unified network fabrics. Congestion management is introduced in Ethernet networks to avoid dropping packets due to congestion. The BCN (Backward Congestion Notification) mechanism is a basic element of several standard drafts, and its stability underlies normal network operations. Because the linear stability analysis method is incapable of handling the nonlinearity of the variable structure control employed by the BCN mechanism, some particular phenomena are unexposed and the insights are insufficient. In this paper, we propose the concept of strong stability of the queuing system to satisfy the requirements of no dropped packets in the data center, and build a fluid-flow analytical model for the BCN congestion control system. Considering the nonlinearity involved in the rate regulation laws, we classify the system into different categories according to the shapes of phase trajectories, and conduct a nonlinear stability analysis using phase plane analysis techniques on a case by case basis. The analysis details can provide a comprehensive understanding about the behaviors of the overall congestion control system. Finally, we also deduce an explicit stability criterion presenting the parameters constraints for the strongly stable BCN system, which can provide straightforward guidelines for proper parameter settings.
Fengyuan Ren, Wanchun Jiang
ICDCS1
2010 Frequency-Domain Packet Scheduling for 3GPP LTE Uplink
abstract
In this paper, we investigate the frequency-domain packet scheduling (FDPS) problem for 3GPP LTE Uplink (UL). Instead of studying a specific scheduling policy, we provide a unified approach to tackle this issue. First we formalize a general LTE UL FDPS problem which is suitable for various scheduling policies. Then we prove that the problem is MAX SNP-hard, which implies that approximation algorithms with constant approximation ratios are the best that we can hope for. Therefore we design two approximation algorithms, both of which have polynomial runtime. Subsequently, we analyze the two algorithms and find their approximation ratios. The first algorithm is easy to follow, since it is based on a simple greedy method. The second one is based on the local ratio technique and it can approximately solve the LTE UL FDPS problem with a approximation ratio of 2.
Hongkun Yang, Fengyuan Ren, Chuang Lin 0002, Jiao Zhang 0002
INFOCOM2
2010 Measurements, analysis and modeling of private tracker sites
abstract
BitTorrent plays a very important role in the current Internet content distribution. When BitTorrent public tracker sites are suffering from free-riding problem, private tracker sites (PTs) work very well because of Share Ratio Enforcement (SRE) which is an auxiliary effective incentive mechanism. Understanding PTs is essential to content distribution. We have crawled and traced 15 tracker sites with over 3.5 million torrents for 7 months. We first provide taxonomy of PTs, and then present measurement study on the characteristics of PTs from the user viscosity, torrents evolution, user behaviors, and content distribution. Some of the features are apparently different from public trackers. Furthermore, we analyze SRE mechanism and auxiliary credit system, and use game theory to study effectiveness of SRE mechanism. There exists “uploading starvation” phenomenon in private trackers. We model SRE mechanism and propose an improved SRE mechanism to further incent the users and enhance the performance of private trackers.
Xiaowei Chen 0001, Xiaowen Chu 0001, Yixin Jiang, Fengyuan Ren
IWQoS4
2010 Energy efficient cooperation in underwater sensor networks
abstract
Energy efficient communication is a key requirement of energy-constrained underwater sensor networks (UWSNs). In this paper, we show that the cooperative diversity, which is conventionally utilized to improve reliability in UWSNs, can be employed to reduce energy consumption and preserve a reasonable level of data reliability and communication delay. We first elucidate in what circumstances the cooperative diversity saves energy compared to the non-cooperative diversity. We show that this largely depends on parameters such as distance between the source and the destination node, the location of potential partner node, and the requirements of reliability and communication delay. Second, we propose a simple but effective cooperation scheme to take advantage of the cooperative diversity. The cooperation diversity can achieve a near-optimal energy saving in some circumstances, and it is no worse than the non-cooperative diversity in all cases. Our work provides instructive theoretical guidelines for designing practical UWSNs.
Hongkun Yang, Fengyuan Ren, Chuang Lin 0002, Bin Liu 0004
IWQoS2
2010 Building a potential field to provide real-time transmission in wireless sensor network
abstract
The Wireless Sensor Network (WSN) is embracing an increasing number of real-time applications subject to strict delay constraints. Utilizing the methodology of potential field in physics, we present an effective way to address the challenges in real-time transmission. We propose the Potential based Real-Time Routing (PRTR) protocol, which provides real-time transmission using multi-path routing algorithm based on a composite potential field. PRTR features a delay-minimized real-time routing as well as alleviating congestion simultaneously.
Yinsheng Xu, Fengyuan Ren, Tao He 0008, Chuang Lin 0002, Sajal K. Das 0001
MSWiM2
2010 Attribute-aware data aggregation using dynamic routing in wireless sensor networks
abstract
Data aggregation has been widely recognized as an efficient method to reduce energy consumption in wireless sensor networks, which can support a wide range of applications such as monitoring temperature, humidity, level, speed etc. The data sampled by the same kind of sensors have much redundancy since the sensor nodes are usually quite dense in wireless sensor networks. To make data aggregation more efficient, the packets with the same attribute, defined as the identifier of different data sampled by different sensors such as temperature sensors, humidity sensors, etc., should be gathered together. However, to the best of our knowledge, present data aggregation mechanisms did not take packet attribute into consideration. In this paper, we take the lead in introducing packet attribute into data aggregation and propose an Attribute-aware Data Aggregation mechanism using Dynamic Routing (ADADR) which can make packets with the same attribute convergent as much as possible and therefore improve the efficiency of data aggregation. This goal cannot be achieved by present static routing schemes employed in most of data aggregation mechanisms since they construct routes before transmitting the sampled data and thus can not dynamically forward packets in response to the variation of packets at intermediate nodes. Hence, we present a potential-based dynamic routing scheme which employs the concept of potential in physics and pheromone in ant colony to achieve our goal. The results of simulations in series of scenarios show that ADADR indeed conserve energy by reducing the average number of transmissions each packet needs to reach the sink and is scalable with regard to the network size.
Jiao Zhang 0002, Fengyuan Ren, Tao He 0008, Chuang Lin 0002
WOWMOM2
2010 Analysis of efficient and fair explicit congestion control protocol with feedback delay: Stability and convergence
Hangxing Wu, Fengyuan Ren, Wenping Pan, Yadong Zhai
Comput. Commun.2
2009 Optimization of Energy Efficient Transmission in Underwater Sensor Networks
abstract
Underwater communication is a challenging topic due to its singular channel characteristics. Most protocols used in terrestrial wireless communication can not be directly applied in the underwater world. In this paper, we focus on the issue of energy efficient transmission in underwater sensor networks (UWSNs) and analyze this problem in a rigorous and theoretical way. We formalize an optimization problem which aims to minimize energy consumption and simultaneously accounts for other performance metrics such as the data reliability and the communication delay. With the help of Karush-Kuhn-Tucker conditions (KKT conditions), we derive a simple and explicit, but nevertheless accurate, approximate solution under reasonable assumptions. This approximate solution provides theoretical guidelines for designing durable and reliable UWSNs. Our result also shows that reliability and communication delay are crucial factors to the energy consumption for transmission.
Hongkun Yang, Bin Liu 0004, Fengyuan Ren, Hao Wen 0014, Chuang Lin 0002
GLOBECOM3
2009 Modeling the Effects of Variable Bandwidth on TCP Throughput
abstract
Links of an end-to-end long lived TCP connection may provide bandwidth with significant variance. In this work, we model and study the effects of variable bandwidth on TCP throughput, taking into account various influencing factors such as random loss of links, buffer size and characteristics of the variable bandwidth in the bottleneck link. Result shows that the variance of bandwidth can severely deteriorate TCP throughput, even making it less than one thirds of the one without bandwidth variance in many cases and arbitrarily low if situations are bad enough, and that under the circumstance of variable bandwidth, a buffer size larger than the traditional recommended bandwidth delay product can greatly improve the throughput. Under the circumstance of variable bandwidth, our model can accurately model the TCP throughput, with an average error around 5%, and a maximum error less than 15%, while traditional formula based TCP throughput prediction can be significantly inaccurate, even with an error larger than 200%.
Fengyuan Ren, Chuang Lin 0002
ICCCN2
2009 Saturation aware TCP throughput prediction
abstract
With the ever growing network traffic, capacity of paths in the networks nowadays has become increasingly easy to saturate. This brings new challenge to TCP throughput prediction. The main problem is loss rate and RTT during the TCP flow are often expected as inputs for the throughput models used in the prediction; however, only loss rate and RTT before the flow are available and are used instead. If the flow itself causes significant changes in loss rate and RTT, e.g., when the TCP flow attempts to saturate the underlying available bandwidth, the prediction error can be unacceptably large. Though new prediction approaches are being proposed, they basic require record of previous TCP transfers and are applicable only when TCP transfers are performed repeatedly, which limits their application. In this work, by properly using a measurement of the underlying available bandwidth, we develop an analytical TCP throughput model which can explicitly capture changes in loss rate and RTT caused by the target TCP flow and hence, can largely improve the prediction accuracy by making the prediction aware of capacity saturation of paths while at the same does not require any history record of previous TCP transfers as current newly proposed works do. Results show that when the changes in loss rate and RTT are large, the errors by traditional models can be as large as over 200%, whereas the error by the proposed model is usually very small, e.g., with the average error below 10% and the maximum error below 20% for general settings, and is bounded by the measurement error of available bandwidth in the worse case; when the changes in loss rate and RTT are small, even a very rough estimation of available bandwidth, e.g., with an error of around 50%, can lead to very accurate prediction by the proposed model.
Fengyuan Ren, Chuang Lin 0002
IPCCC2
2009 RENA: region-based routing in intermittently connected mobile network
abstract
Considering the constraint brought by mobility and resources, it is important for routing protocols to efficiently deliver data in Intermittently Connected Mobile Network (ICMN). Different from previous works that use the knowledge of previous encounters to predict the future contact, we propose a storagefriendly REgioN-bAsed protocol, namely, RENA, in this paper. Instead of using temporal information, RENA builds routing tables based on regional movement history, which avoids excessive storage for tracking encounter history. We validate the generality of RENA through time-variant community mobility model with parameters extracted from the MIT WLAN trace, and the vehicular network based on 8 bus routes of the city of Helsinki. The comprehensive simulation results show that RENA is not only storage-friendly but also more efficient than the epidemic routing, the restricted replication protocol SNW and the encounter-based protocol RAPID under various conditions.
Hao Wen 0014, Jia Liu 0024, Chuang Lin 0002, Fengyuan Ren, Pan Li 0001, Yuguang Fang
MSWiM4
2009 Idle Detection Based Optimal Throughput Rate Adaptation in Multi-Rate WLANs
abstract
Multiple transmission rates are supported in the current 802.11 WLANs, which allows stations to exploit different rates in an adaptive manner to cope with the variability of wireless channels and achieve the higher system throughput. Many efforts have been made at rate adaptation in the literature. The design of rate adaptation algorithms needs to take four basic issues into account. In this paper, by properly addressing the four issues, we present a novel rate adaptation algorithm, idle detection based optimal throughput rate adaptation (ITRA). We apply a comprehensive approach to designing the rate adaptation mechanism and attempt to maximize the overall system throughput. ITRA is furnished with two instructive features. First, ITRA modifies the conventional exponential backoff rule to a constant one. With the help of constant contention window (CW), ITRA can discriminate collisions from channel errors, and does not need RTS/CTS handshakes which are adopted by many existing works. We derive an explicit relation between the throughput and the CW size, and choose the CW size to optimize the throughput as well as improve the fairness. Secondly, ITRA directly estimates the current network throughput and changes transmission rate according to the estimation. Since the fundamental purpose of rate adaptation is to improve the network throughput, using the throughput as the metric is more straightforward than applying indirect criteria like frame transmission statistics or PHY metrics. We evaluate our proposed scheme by simulation, and the result shows that ITRA achieves satisfactory performance.
Bin Liu 0004, Fengyuan Ren, Hongkun Yang, Chuang Lin 0002
SECON2
2009 An efficient and fair explicit congestion control protocol for high bandwidth-delay product networks
Hangxing Wu, Fengyuan Ren, Xianwu Gong
Comput. Commun.2
2008 The Redeployment Issue in Underwater Sensor Networks
abstract
The mobility of underwater sensor nodes makes the network topology inconveniently controlled and slowly changed. Thus, in order to enable underwater sensor networks to work more effectively, it is necessary for us to periodically detect the coverage rate and redeploy nodes to non-coverage areas. In this paper, we take the lead in introducing the redeployment issue in underwater sensor networks. In our opinion, the key point of the redeployment issue is coverage. For this special coverage topic, we first propose a coverage rate definition scheme. Then along with the definitions, two redeployment algorithms are introduced, of which one is based on adding new nodes while the other one is by the means of moving redundant ones. By modeling the mobility behavior of underwater nodes with three-dimensional random walks, we employ simulation experiments to verify our ideas, the results of which show the importance of redeployment in the underwater environment.
Bin Liu 0004, Fengyuan Ren, Chuang Lin 0002, Yaqin Yang, Rongfei Zeng, Hao Wen 0014
GLOBECOM2
2008 Analyzing the Reliability of Group Transmission in Wireless Sensor Network
abstract
Most previous models about wireless channels are mainly adopted to obtain average packet error rates or to generate artificial network traces. However, when we combine packet group transmission with error correcting mechanisms, the steady packet error rate (PER) is not accurate enough to depict the short-term error event. In this paper, we propose a Markov chain model for group transmission to indicate influences of group length N and initial channel state on transmission reliability. Based on the model, we prove that the difference between steady PER and packet error rate within a group (PERG) is O(1/N), which can be used to estimate group length under a certain error rate constraint. Finally, we apply the model to compare the multipath transmission with the single-path transmission, and investigate an extreme case where the multipath way is less reliable than the single-path way.
Hao Wen 0014, Hongkun Yang, Chuang Lin 0002, Fengyuan Ren, Yao Yue
GLOBECOM4
2008 Performance Analysis of Sleep Scheduling Schemes in Sensor Networks using Stochastic Petri Net
abstract
As the most important issue in wireless sensor networks, power saving is catching researchers' great attentions all the time. Among several power saving strategies, one of the best methods is to make unused components inactive whenever possible, i.e. applying sleep scheduling schemes to sensors. In this paper, by analyzing and concluding existing works, we propose four sleep scheduling schemes (ACAA, SRAA, SAA and NCAA) for different kinds of application environments, then analyze each of them by stochastic petri net (SPN). We can easily get the average power consumptions and event delays of sensor nodes by using the steady state probability matrix of the SPN models. Moreover, the numeric results show that this concise graphic analysis method is suitable for analyzing sleep scheduling schemes.
Bin Liu 0004, Fengyuan Ren, Chuang Lin 0002
ICC2
2008 Performance Analysis of Retransmission and Redundancy Schemes in Sensor Networks
abstract
In this paper, by establishing the probability models, we systematically and comprehensively analyze the roles of the packet retransmission, the block retransmission, and the erasure coding in the reliable transport of wireless sensor networks. And as well as the three kinds of packet level schemes, we also consider the effect of two kinds of bit level strategies, CRC and FEC. At last, based on the numeric results, the appropriate schemes for different BER (high, medium, and low) are determined, and we also present some principles that reveal profound insights in designing reliable protocols and mechanisms in wireless sensor networks.
Bin Liu 0004, Fengyuan Ren, Chuang Lin 0002, Ying Ouyang
ICC2
2008 End-to-End Congestion Control for High Speed Networks Based on Population Ecology Models
abstract
Since TCP congestion control is ill-suited for high speed networks, designing a replacement for TCP has become a challenge. To address this problem, we extend the population ecology theory to design a novel congestion control algorithm. We treat the network flows as the species in nature, the throughput of the flows as the population number, and the bottleneck bandwidth as the food resources. Then we use the key idea of constructing population ecology models to develop a novel congestion control model, and implement the corresponding end-to-end transport protocol through measurement, which called Population Ecology TCP (PE-TCP). The theoretical analysis and simulation results validate that PE-TCP achieves high utilization, fast convergence, fair bandwidth allocation, and near-zero packet drops. These qualities are desirable for high speed networks.
Xiaomeng Huang, Fengyuan Ren, Guangwen Yang 0002, Yongwei Wu 0001, W. Zhen, Chuang Lin 0002
ICDCS2
2008 Joint Adaptive Redundancy and Partial Retransmission for Reliable Transmission in Wireless Sensor Networks
abstract
As a potentially competitive technique, erasure coding has been employed in wireless sensor networks (WSNs) to enhance transmission reliability. In this paper, to design a practical and efficient redundancy mechanism in WSNs, we firstly provide a theoretical study of packet delivery probability and average energy consumption for retransmission and redundancy mechanisms. The theoretical results indicate that in WSNs, the adaptive redundancy coding mechanism is more energy efficient than retransmission while keeping the same level of reliability in most scenarios. Then based on the mapping table and two basic design principles obtained from the theoretical analysis, we propose a reliable transport protocol ARRTP which combines the adaptive redundancy mechanism and partial retransmission. The simulation and trace-driven results both verify that when the loss probability varies from low to high, our protocol provides reliability with comparably lower energy consumption. Furthermore, compared with fixed redundancy degrees, the robustness of the adaptive mechanism are also evaluated.
Hao Wen 0014, Chuang Lin 0002, Fengyuan Ren, Hongkun Yang, Tao He 0008, Eryk Dutkiewicz
IPCCC3
2008 Alleviating Congestion Using Traffic-Aware Dynamic Routing in Wireless Sensor Networks
abstract
The congestion problem in wireless sensor networks (WSNs) is quite different from that in traditional networks. Most current congestion control algorithms try to alleviate the congestion by reducing the rate at which the source nodes inject packets into the network. However, this traffic control scheme always decreases the throughput so as to violate fidelity level required by the application. In this paper, we present a solution that sufficiently exert the idle or under-loaded nodes to alleviate congestion and improve the overall throughput. To achieve this goal, a traffic-aware dynamic routing(TADR) algorithm is proposed to route packets around the congestion areas and scatter the excessive packets along multiple paths consisting of idle and under-loaded nodes. Enlightened by the concept of potential in common physics, our TADR algorithm is designed through constructing a mixed potential field using depth and normalized queue length to force the packets to steer clear of obstacles created by congestion and eventually move towards the sink. The simulation results show that our solution achieves its objectives and improves the overall throughput by around 370% as compared to the benchmark routing protocol. Furthermore, our TADR has low overhead suitable for large scale dense sensor networks.
Tao He 0008, Fengyuan Ren, Chuang Lin 0002, Sajal K. Das 0001
SECON2
2008 A Study of Forward Error Correction Schemes for Reliable Transport in Underwater Sensor Networks
abstract
Underwater communications is a very challenging topic due to its singular channel characteristics. Most protocols used in terrestrial wireless communications can not be directly applied in the underwater world. A high bit error rate and low propagation delay make the design of reliable transport protocols especially awkward. In this paper, we first propose four schemes that combine forward error correction mechanisms at the bit and/or packet level to increase the reliability in a non-cooperative scenario. The broadcast property of the underwater environment allows us to extend them to a cooperative setting. Based on our analyses, we introduce ADELIN: an adaptive reliable transport protocol for underwater sensor networks. We suggest an architecture for implementation and compare our protocol to other schemes. We show that it succeeds in a better probability and energy tradeoff for both single- and multi-hop communications.
Bin Liu 0004, Florent Garcin, Fengyuan Ren, Chuang Lin 0002
SECON3
2008 Performance analysis of reliable transport schemes joint with the optimal frequency and optimal packet length in underwater sensor networks
abstract
In underwater sensor networks (UWSNs), high-delay acoustic channels are used for communications, which introduces new challenges for analyzing and choosing proper reliable transport mechanisms. In this paper, we investigate four key parameters (node distance, communication frequency, packet length and SNR) influencing the transport of UWSNs. Since treating all the four factors equally will make the analysis almost intractable, we propose a four-step analytical framework. In this framework, by optimizing both the communication frequency and packet length, we firstly show that the average energy consumption and the average communication delay of different transport schemes are only related to node distance and SNR. Then when considering practical scenarios, we further demonstrate that the scheme performances are mainly determined by SNR, which rather simplifies the analysis process. In addition, an important metric PEDP is proposed, which is only related to SNR but has the capability of evaluating the performance of transport schemes roundly and effectively. Finally, the numeric results offer us some profound insights for choosing proper transport schemes in UWSNs.
Bin Liu 0004, Hao Wen 0014, Fengyuan Ren, Chuang Lin 0002
WOWMOM3
2008 Design and analysis of an ONOFF variable structure controller for AQM routers supporting TCP flows
Fengyuan Ren, Yunhe Yin, Chuang Lin 0002
Sci. China Ser. F Inf. Sci.1
2008 Improving TCP Throughput over HSDPA Networks
abstract
The various link adaptation techniques employed by High Speed Downlink Packet Access (HSDPA) in the third generation (3G) networks augment the bandwidth oscillation, which is identified as one of the most important factors resulting in the throughput deterioration of Transmission Control Protocol (TCP). In this paper, we firstly explain why the bandwidth oscillation degrades the TCP performance through a special simulation experiment. Subsequently, a split connection Window Adaptation TCP Proxy is proposed to improve the TCP throughput over HSDPA networks. In this solution, the built-in attributes of HSDAP system are sufficiently utilized. In order to effectively use the precious cellular link resources, the length of the queue connected with it is intentionally kept around the reference value through adjusting the sending window size of TCP proxy based on the dynamic values of varying bandwidth. A discrete-time stochastic state space model is formulated to analyze the system stability. The validity of enhanced scheme is verified through simulation experiments. The performance of TCP proxy is compared with the standard TCP protocol. The numerical results show that our TCP proxy is able to keep the cellular link utilization over 90%, and to improve TCP throughput by 100% under most conditions.
Fengyuan Ren, Xiaomeng Huang, Chuang Lin 0002
IEEE Trans. Wirel. Commun.1
2007 Improving the Convergence and Stability of Congestion Control Algorithm
abstract
The traditional TCP congestion control is inefficient for high speed networks and it is a challenge to design a high speed replacement for TCP. By simulating some existing high speed protocols, we find that these high speed protocols have limitations in convergence and stability. To address these problems, we apply a population ecology model to design a novel congestion control algorithm-Coupling Logistic TCP(CLTCP). It is based on bandwidth pre-assignment that is similar to XCP and MaxNet. The pre-assignment rate factor is computed in the routers based on the information of the router capacity, the aggregate incoming traffic and the queue length. Then the senders adjust the sending rate according to the pre-assignment rate factor which carries by the packet to strengthen the convergence and stability of transport protocol. The theoretical analysis and simulation results show that CLTCP provides not only fast convergence and strong stability, but also high utilization and fair bandwidth allocation regardless of round trip time.
Xiaomeng Huang, Chuang Lin 0002, Fengyuan Ren, Guangwen Yang 0002, Peter D. Ungsunan, Yuanzhuo Wang
ICNP3
2007 A Simple Active Congestion Control in Wireless Sensor Network
abstract
More attention has been paid to congestion control in the emerging area of wireless sensor network (WSN). However, most research works in the past stayed at the level of algorithm design or modification, and seldom sought solutions on the viewpoint of architecture. In this paper, Active Networking (AN) technology is used to make congestion control more responsive to detect/recover congestion in WSN. We design a simple Active Backpressure (BP) mechanism to allocate bandwidth Proportional to the Size of tree (ABPS). ABPS introduces programs in each data packet that tell nodes how to react to congestion, and quickly converges to a fair and efficient rate. Finally, we evaluate ABPS extensively on a 50-node wireless sensor network. Simulation results validate the effectiveness of our ABPS.
Ying Ouyang, Fengyuan Ren, Chuang Lin 0002, Tao He 0008, Yada Hu, Hao Wen 0014
MASS2
2007 Retransmission or Redundancy: Transmission Reliability in Wireless Sensor Networks
abstract
As an application-driven network, wireless sensor network generally requires high data reliability to maintain detection and response capabilities. Although two approaches, which are retransmission and redundancy, have been proposed to enhance data reliability, the theoretical work is required to evaluate their impact on transmission reliability and energy efficiency. In this paper, we offer a comprehensive theoretical study on the packet arrival probability and average energy consumption for both approaches. Our analysis indicates that when loss probability remains low or moderate, erasure coding, a scheme based on redundancy, is more reliable and energy efficient than retransmission. However, the performance of erasure coding would largely deteriorate under high packet loss condition. We also demonstrate that its resistance capability against packet loss weakens as hop number increases. Furthermore, with the increase in redundancy, erasure coding has to sacrifice the advantage of energy efficiency for reliability.
Hao Wen 0014, Chuang Lin 0002, Fengyuan Ren, Yao Yue, Xiaomeng Huang
MASS3
2007 Design and Analysis of a Backpressure Congestion Control Algorithm in Wireless Sensor Network
abstract
More attention has been paid to congestion control in the emerging area of wireless sensor network (WSN). However, most research works in the past stayed at the level of the current algorithms design or modification, and seldom sought solutions on the viewpoint of architecture. In this paper, Backpressure(BP) under Active Network(AN) architecture is used to make congestion control more responsive to detect/recover congestion in WSN. We design a simple Active Backpressure mechanism to allocate bandwidth Proportional to the Size of tree (ABPS), and we present a fluid-based analytical model of ABPS using stochastic differential equations. ABPS introduces programs in each data packet that tell nodes how to react to congestion, and quickly converge to a fair and efficient rate. We demonstrate a deterministic approach to analyse the stochastic model, in which we obtain a set of ordinary differential equations from our model, and we derive the average behavior of queue length and flow throughput from the ordinary differential equations. Finally, we evaluate ABPS extensively on a 50-node wireless sensor network. Simulation results validate the effectiveness of our ABPS and match well with the theoretic analysis.
Ying Ouyang, Chuang Lin 0002, Fengyuan Ren, Hongkun Yang, Xiaomeng Huang
PDCAT3
2007 A novel high speed transport protocol based on explicit virtual load feedback
Xiaomeng Huang, Chuang Lin 0002, Fengyuan Ren
Comput. Networks3
2006 AntiWorm NPU-based Parallel Bloom filters in Giga-Ethernet LAN
abstract
In this paper, an AntiWorm system based on the Intel IXP Network Processor was implemented using the Parallel Bloom filters technique. The AntiWorm system consists of two components: Bloom filters and Exact Matching engines. The Parallel Bloom filters can identify the suspicious traffic quickly and effectively, and then dispatch them to Exact Matching engines for further investigation. Both the principles and the implementation of the AntiWorm system are introduced in detail. With the consideration of the system performance parameters, two feasible implementation solutions are investigated and the advantages and disadvantages are also compared. The selections of configuration parameters of the AntiWorm system are also discussed. A hash scheme based on MD5's function is proposed for implementing fast hash functions. To test the performance of the AntiWorm system, such as throughput and delay, some experiments are carried out with different simulated traffic condition. The internal statistics of IXP network processor are also collected and analyzed for optimizing the system performance. To demonstrate the operation of the AntiWorm system, assaults by Worm Blaster are used in the test bed, and the experimental results prove the effectiveness of the AntiWorm system. The Software Package WormDetector1.0 is also provided as a software release from the research.
Zhen Chen 0001, Chuang Lin 0002, Jia Ni, Dong-Hua Ruan, Bo Zheng 0007, Zhangxi Tan, Yixin Jiang, Xuehai Peng, An'an Luo, Yao Yue, Yang Wang 0018, Peter D. Ungsunan, Fengyuan Ren
ICC14
2005 Dynamic channel allocation for mobile cellular systems using a control theoretical approach
abstract
The guard channel scheme in wireless mobile networks has attracted and is still drawing research interest owing to easy implementation and flexible control. However guard channel schemes can not adapt to changing traffic loads because of static reserved guard channels. Therefore dynamic guard channel schemes have been proposed in the literature to adapt to varying traffic load. This paper presents a novel control-theoretic approach to dynamically reserve guard channels called PI-guard channel (PI-GC) controller. Experiments show that our proposed scheme can maintain the handoff blocking probability (HBP) to a predefined value while it still improves the channel resource utilization.
Yaya Wei, Chuang Lin 0002, Raad Raad, Fengyuan Ren
GLOBECOM4
2005 Class-Based Latency Assurances for Web Servers
Yaya Wei, Chuang Lin 0002, Xiaowen Chu 0001, Zhiguang Shan, Fengyuan Ren
HPCC5
2005 Fuzzy Control for Guaranteeing Absolute Delays in Web Servers
abstract
This paper presents a fuzzy control approach that guarantees absolute delays in Web servers. Previous work has proposed the use of classical PI controllers for delay guarantees. However, a disadvantage of the classical PI controller is that the system model, which is obtained by system identification, mismatches the real system and inevitably degrades the performance of the Web system. In contrast with classical PI controllers, fuzzy controllers are nonlinear and therefore independent of the accurate model of the plant, i.e. the controlled system. Hence, fuzzy controllers seem to be very suitable for Web servers. Our experiments show that fuzzy controllers indeed perform better than PI controllers presented in earlier papers.
Yaya Wei, Fengyuan Ren, Chuang Lin 0002, Thiemo Voigt
QSHINE2
2005 A nonlinear control theoretic analysis to TCP-RED system
Fengyuan Ren, Chuang Lin 0002
Comput. Networks1
2005 A robust active queue management algorithm in large delay networks
Fengyuan Ren, Chuang Lin 0002
Comput. Commun.1
2005 Design a congestion controller based on sliding mode variable structure control
Fengyuan Ren, Chuang Lin 0002, Xunhe Yin
Comput. Commun.1
2005 Modeling and Inference of Extended Interval Temporal Logic for Nondeterministic Intervals
abstract
Extended interval temporal logic (EITL), an extension of the traditional point-interval temporal logic (PITL), is proposed. In contrast to PITL that represents the dynamic aspects of deterministic intervals, EITL can model and reason about the temporal relations among nondeterministic intervals in discrete-event systems, in which the duration of an event is indeterminate and only the lower bound and upper bound of the ending time can be predicted in advance. Time Petri nets (TPNs) are used for modeling EITL, for they give a straightforward view of temporal relations between the extended intervals and also provide a number of theoretical and practical analysis methods. An inference engine based on the TPN modeling complemented with algebraic inequalities is proposed to construct an analytical representation of the EITL relations and solve qualitative temporal reasoning problems. Linear inference mechanism based on TPN reduction rules is used to infer new temporal relations and handle quantitative temporal reasoning problems with linear time complexity, as our example shows.
Chuang Lin 0002, Zhiguang Shan, Fengyuan Ren
IEEE Trans. Syst. Man Cybern. Part A5
2004 Design a robust controller for active queue management in large delay networks
abstract
AQM (active queue management) can maintain the smaller queuing delay and higher throughput by the purposefully dropping the packets at the intermediate nodes. Almost all the existed schemes for AQM neglect the impact of large delay on performance. In this study, we firstly verify a fact through simulation experiments, which is the queues controlled by the several popular AQM schemes, including RED, PI controller and REM, appear the dramatic oscillations in large delay networks, which decreases the utilization of the bottleneck link and introduces the avoidable delay jitter. After some appropriate model approximation, we design a robust AQM controller to compensate the delay using the principle of internal mode compensation in control theory. The novel algorithm restrains the negative effect on queue stability caused by large delay. The simulation results show that the integrated performance of the proposed algorithm is obviously superior to that of several well-known schemes when the connections have large delay, at the same time, the buffers keep at small queue length.
Fengyuan Ren, Chuang Lin 0002
ISCC1
2004 Using fuzzy-PI controller in active queue management
abstract
In this paper we propose using fuzzy logic to improve the performance of PI controller in the design of AQM (active queue management). With the introduction of integral factor in PI controller, the steady state error in proportional controller (such as RED) is eliminated. However, the response speed is slowed down. We design a fuzzy-PI (FPl) controller to solve this problem. FPI controller combines the advantages of fuzzy control while maintaining the simplicity and robustness of a conventional PI controller. The performance of FPI is verified and compared with Pl controller using ns-2 simulations. It is suggested that FPI is superior to Pl in response speed. Thus FPI is more robust in the presence of disturbance caused by the non-responsive UDP flows and the changing number of active flows.
Fengyuan Ren, Ke Xu 0002
ISCC2
2004 Dynamic handoff scheme in differentiated QoS wireless multimedia networks
Yaya Wei, Chuang Lin 0002, Fengyuan Ren, Raad Raad, Eryk Dutkiewicz
Comput. Commun.3
2003 Dynamic Priority Handoff Scheme in Differentiated QoS Wireless Multimedia Networks
abstract
Handoff is one of the key elements in ensuring quality of service (QoS) in mobile wireless networks. Handoff connections generally have higher priority than new connections. Traditional reservation policies that reserve some channels for handoff connections are not adaptive to traffic load changes. This paper proposes a new dynamic guard channel scheme (DGCS), which 1) adapts to various traffic loads; 2) combines differentiated QoS service model and priority handoff mechanism; 3) provides fairness for differentiated QoS services; 4) does not need to exchange state information among different cells, so it is easy to be implemented and is simple enough to be used in real time environments; 5) and utilizes network resources efficiently and puts a bound on each service blocking probability. The simulation results show that the ratios among different QoS service probabilities are guaranteed to be predefined values and system utilization is improved greatly.
Yaya Wei, Chuang Lin 0002, Fengyuan Ren, Raad Raad, Eryk Dutkiewicz
ISCC3
2003 Design a PID Controller for Active Queue Management
abstract
As an enhancement mechanism for the end-to-end congestion control, active queue management (AQM) can keep smaller queuing delay and higher throughput by purposefully dropping the packets at the intermediate nodes. Comparing with RED algorithm, although the PI (proportional-integral) controller for AQM designed by C. Hollot improves the stability, the transient performance of the PI controller is not perfect, such as the regulating time is too long. In order to overcome this drawback, in this paper, the PID (proportional-integral-differential) controller is proposed to speed up the responsiveness of AQM system. The controller parameters are tuned based on the determined gain and phase margins. The simulation results show that the integrated performance of the PID controller is obviously superior to that of the PI controller.
Yanfei Fan, Fengyuan Ren, Chuang Lin 0002
ISCC2
2002 A Robust Active Queue Management Algorithm Based on Sliding Mode Variable Structure Control
abstract
As an effective mechanism acting on the intermediate nodes to support end-to-end congestion control, active queue management (AQM) takes a trade-off between link utilization and delay experienced by data packets. Most of the existing AQM algorithms are heuristic, and a lack systematic and theoretical design and analysis approach. From the viewpoint of control theory, it is rational to regard AQM as a typical regulating system. Although the PI controller for AQM outperforms the RED algorithm, the mismatches in the simplified TCP flow model inevitably degrade the performance of a controller designed with classical control theory. In this paper, a robust SMVS controller for AQM is put forward based on sliding mode variable structure control (SMVS), its superiority is insensitive to the noise and variance of the parameters, thus it very suitable to a time-varying network system. The principle and guidelines on design of a SMVS controller are presented in detail. The integrated performance is evaluated using ns simulations. The results show that SMVS is very responsive and robust against disturbances. At the same time, a complete comparison between the SMVS controller and PI controller is made. The conclusion is that both the transient and steady performance of the SMVS controller is superior to that of the PI controller, thus the SMVS controller is in favor of the achievement of AQM objectives.
Fengyuan Ren, Chuang Lin 0002, Ying Xunhe, Xiuming Shan, Wang Fubao
INFOCOM1
2002 Design of a fuzzy controller for active queue management
Fengyuan Ren, Yong Ren 0001, Xiuming Shan
Comput. Commun.1
2001 Analysis and Improvement of the EFCI Algorithm
abstract
ATM networks oriented connections provide pure QoS (quality of service) for diversified services through a series of traffic management mechanisms, the ABR (available bit rate) flow control is especially important among these approaches. In the binary flow control scheme, the cell rate and queue length may oscillate with great magnitude to reduce link utilization, so that the EFCI (explicit forward congestion indication) algorithm is regard as ineffective, however its simplicity is attractive to high performance switch design. In this paper, the EFCI algorithm is analyzed based on classical control theory. It is found that the nonlinear hysteresis determines the congestion, and is the dominant reason that causes the oscillation. Then, a probability congestion detection approach, called p-EFCI is put forward. Numerical results show that the improved algorithm deeply constrains the oscillation magnitude, and the queue length is controlled in limited scope to guarantee cell zero loss.
Fengyuan Ren, Yong Ren 0001, Xiuming Shan, Wang Fubao
ISCC1