VLDB 2026 Research / reviewers in the wild / expert
Ren Wang 0001
dblp:29/50-1
· DBLP profile ↗
47ranked-venue papers
6as first author
14since 2021 · last 2026
0009-0005-2937-5804ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 21 · 6 first-authorSystems, architecture and hardware · 20 · 13 since 2021Software engineering, systems software and programming languages · 10 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiLo: Harnessing the on-Chip Accelerators in Intel CPUs for Compressed LLM Inference AccelerationabstractThe ever-growing sizes of large language models (LLMs) introduce significant infrastructure challenges due to their immense memory capacity demands. While the de facto approach has been to deploy multiple high-end GPUs, each with a limited memory capacity, the prohibitive cost of such systems has become a major barrier to the widespread deployment of frontier LLMs. As a result, CPU-based inference has become an appealing and cost-efficient alternative, since a CPU can offer an order of magnitude larger memory capacity at a fraction of the cost while providing competitive throughput for matrixvector multiplication with the latest Advanced Matrix Extensions (AMX). It not only broadens accessibility for users without multiGPU setups but also enables hyperscalers to leverage underutilized CPU servers to accommodate temporarily surging inference demand. Nevertheless, even CPU's large memory capacity has become insufficient to serve LLMs with hundreds of billions of parameters. Under the memory capacity constraint, we may offload parameters to storage devices and fetch them on demand, but doing so significantly degrades inference performance due to the high latency and low bandwidth of storage devices. To address this challenge, we propose LILO, an LLM inference framework that leverages In-memory Analytics Accelerator (IAA) in the latest Intel CPUs, to accelerate inference under memory capacity constraints. By storing model parameters in a compressed format and decompressing them on demand using IAA, LILO enables significantly reduced storage access during inference under memory capacity constraints while preserving the model accuracy and behavior. LILO orchestrates the concurrent execution of on-chip accelerators, i.e., IAA, Advanced Vector Extensions (AVX), and AMX, to facilitate high-throughput decompression alongside inference computation. Furthermore, LILO implements selective compression, a Mixture-of-Expert (MoE)-aware optimization that reduces the decompression overhead by up to 1.9×. We demonstrate that LILO reduces inference latency by up to 4.9× and 4.3× for Llama3-405B and DeepSeekR1, respectively, under memory capacity constraints compared to the baseline inference solely relying on storage-offloading without compression. Hyungyo Kim, Qirong Xia, Jinghan Huang 0001, Nachuan Wang, Younjoo Lee 0001, Jung Ho Ahn, Wajdi K. Feghali, Ren Wang 0001, Nam Sung Kim |
HPCA | 8 |
| 2025 | M5: Mastering Page Migration and Memory Management for CXL-based Tiered Memory SystemsabstractCXL has emerged as a promising memory interface that can cost-effectively expand the capacity and bandwidth of a memory system, complementing the traditional DDR interface. However, CXL DRAM presents 2-3x longer access latency than DDR DRAM, forming a tiered-memory system that demands an effective and efficient page-migration solution. Although many page-migration solutions have been proposed for past tiered-memory systems, they have achieved limited success. To tackle the challenge of managing tiered-memory systems, this work first presents a CXL-driven profiling solution to precisely and transparently count the number of accesses to every 4KB page and 64B word in CXL DRAM. Second, using the profiling solution, this work uncovers that (1) widely used CPU-driven page-migration solutions often identify warm pages as hot pages, and (2) certain applications have sparse hot pages, where only a small percentage of words in each of these pages are frequently accessed. Besides, this work demonstrates that the performance overhead of identifying hot pages is sometimes high enough to degrade application performance. Lastly, this work presents M5, a platform designed to facilitate the development of effective CXL-driven page-migration solutions, providing hardware-based hot-page and hot-word trackers in the CXL controller. On average, M5 can identify 47% hotter pages and offer 14% higher performance than the best CPU-driven page-migration solution, even with a simple policy. Jongyul Kim 0001, Zeduo Yu, Jiyuan Zhang 0003, Siyuan Chai 0001, Michael Jaemin Kim, Hwayong Nam, Jaehyun Park 0006, Eojin Na, Ren Wang 0001, Jung Ho Ahn, Tianyin Xu, Nam Sung Kim |
ASPLOS (2) | 11 |
| 2025 | Dynamic Load Balancer in Intel Xeon Scalable Processor: Performance Analyses, Enhancements, and GuidelinesabstractThe rapid increase in inter-host networking speed has challenged host processing capabilities, as bursty traffic and uneven load distribution among host CPU cores give rise to excessive queuing delays and service latency variances.To cost-efficiently tackle this challenge, the latest Intel Xeon Scalable Processor has integrated an on-chip accelerator, named Dynamic Load Balancer (DLB).It consists of hardware queues, arbiters, and priority-based Quality of Service (QoS) features to maximize load-balancing performance while minimizing host CPU cycle consumption.In this work, we first compare the performance of DLB against popular softwarebased load balancers with a microbenchmark suite, demonstrating that DLB significantly outperforms them.Yet, DLB still consumes a significant number of host CPU cycles to prepare and enqueue work descriptors for received packets.Second, to eliminate the consumption of host CPU cycles in load balancing, we propose a system architecture/software co-design solution, AccDirect.Specifically, AccDirect leverages the Peer-to-Peer (P2P) communication capability of PCIe devices to enable a direct communication path between DLB and NIC, through which NIC directly enqueues the work descriptors.Our evaluation shows that AccDirect-based DLB offers practically the same performance as conventional DLB while reducing the system-wide power consumption of host CPU by 10%.Compared to the throughput of a commodity hardware-based load balancer for an end-to-end application, that of AccDirect-based DLB is 14-50% higher with comparable p99 latency.Lastly, we provide guidelines to make the best use of DLB, which is riddled with a vast configuration space and advanced features, after conducting a comprehensive evaluation. Jiaqi Lou, Srikar Vanavasam, Ren Wang 0001, Nam Sung Kim |
ISCA | 4 |
| 2025 | Re-architecting End-host Networking with CXL: Coherence, Memory, and Offloading
Houxiang Ji, Yang Zhou 0050, Ipoom Jeong, Ren Wang 0001, Saksham Agarwal, Nam Sung Kim |
MICRO | 5 |
| 2025 | RACER: Avoiding End-to-End Slowdowns in Accelerated Chip Multi-ProcessorsabstractRecent chip multiprocessors incorporate several on-chip accelerators, marking the beginning of the Accelerated Chip Multi-Processor (XMP) era in datacenters. Despite the close proximity of accelerators and general-purpose cores, offloading functions to accelerators may not always be beneficial. Offloading to hardware accelerators can introduce several end-to-end overheads that can negate the speedup of the accelerable function. In this article, we design RACER, a hardware architecture and runtime system that evades the danger of end-to-end slowdowns when using hardware acceleration. RACER leverages a low-overhead interface between general-purpose cores and on-chip accelerators, fine-grained context switching, accelerator-initiated preemption, and seamless data motion between general-purpose cores and accelerators to improve the performance of workloads that use on-chip accelerators. We evaluate RACER on five representative request processing workloads featuring diverse memory access patterns, accelerable functions, and compute intensities. RACER improves the performance of hardware acceleration on a real XMP by an average of 1.31× on a range of diverse workloads and guarantees that accelerator offloads never cause slowdowns. Ren Wang 0001, Mohammad Alian |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable ProcessorsabstractAs semiconductor power density is no longer constant with the technology process scaling down, we need different solutions if we are to continue scaling application performance. To this end, modern CPUs are integrating capable data accelerators on the chip, aiming to improve performance and efficiency for a wide range of applications and usages. One such accelerator is the Intel® Data Streaming Accelerator (DSA) introduced since Intel® 4th Generation Xeon® Scalable CPUs (Sapphire Rapids). DSA targets data movement operations in memory that are common sources of overhead in datacenter workloads and infrastructure. In addition, it supports a wider range of operations on streaming data, such as CRC32 calculations, computation of deltas between data buffers, and data integrity field (DIF) operations. This paper aims to introduce the latest features supported by DSA, dive deep into its versatility, and analyze its throughput benefits through a comprehensive evaluation with both microbenchmarks and real use cases. Along with the analysis of its characteristics and the rich software ecosystem of DSA, we summarize several insights and guidelines for the programmer to make the most out of DSA, and use an in-depth case study of DPDK Vhost to demonstrate how these guidelines benefit a real application. Reese Kuper, Ipoom Jeong, Ren Wang 0001, Narayan Ranganathan, Philip Lantz, Nam Sung Kim |
ASPLOS (2) | 4 |
| 2024 | Intel Accelerators Ecosystem: An SoC-Oriented Perspective : Industry ProductabstractA growing demand for hyperscale services has compelled hyperscalers to deploy more compute resources at an unprecedented pace, further accelerated by the demise of Dennard scaling. Meanwhile, a considerable portion of the compute resources are consumed to execute common functions present across the hyperscale services, i.e., datacenter taxes. These challenges motivated many to explore specialized accelerators for these functions. Leading such a technology trend, Intel has integrated diverse on-chip accelerators into its recent flagship datacenter CPU products. Furthermore, to support the easy and efficient use of these accelerators for successful deployment in production hyperscale services, Intel has developed a hardware/software ecosystem. In this paper, we first focus on Intel’s holistic efforts to build the hardware/software ecosystem, presenting key SoC-level features that facilitate efficient CPU-accelerator interaction, effortless programming and use, and scalable accelerator sharing and virtualization. Next, we delve into the functions, microarchitectures, and software stacks of three new on-chip accelerators: Data Streaming Accelerator (DSA), In-memory Analytics Accelerator (IAA), and Dynamic Load Balancer (DLB). Lastly, we demonstrate that Intel’s on-chip accelerators can not only significantly reduce the datacenter taxes but also accelerate data-intensive applications essential for hyperscale services, with little effort to use the accelerators. Ren Wang 0001, Narayan Ranganathan, Philip Lantz, Vivekananthan Sanjeepan, Jorge Cabrera 0003, Atul Kwatra, Rajesh Sankaran, Ipoom Jeong, Nam Sung Kim |
ISCA | 2 |
| 2024 | Demystifying a CXL Type-2 Device: A Heterogeneous Cooperative Computing PerspectiveabstractCXL is the latest interconnect technology built on PCIe, providing three protocols to facilitate three distinct types of devices, each with unique capabilities. Among these devices, a CXL Type-2 device has become commercially available, followed by CXL Type-3 devices. Therefore, it is timely to understand capabilities and characteristics of the CXL Type-2 device, as well as explore suitable applications. In this work, first, we delve into three key features of a CXL Type-2 device: cache-coherent device accelerator to host memory, device accelerator to device memory, and host CPU to device memory accesses. Second, using microbenchmarks, we comprehensively characterize the latency and bandwidth of these memory accesses with a CXL Type-2 device, and then compare them with those of equivalent memory accesses with comparable devices, such as emulated CXL Type-2, CXL Type-3, and PCIe devices. Lastly, as applications that exploit the unique capabilities of a CXL Type-2 device, we propose two CXL-based Linux memory optimization features: compressed RAM cache for swap (zswap) and memory deduplication (ksm). Our evaluation shows that Redis, when running with traditional CPU-based zswap and ksm, suffers from a tail latency increase of 4.5-10.3× compared to Redis running alone. While PCIe-based zswap and ksm still experience a tail latency increase of up to 8.1×, CXL-based zswap and ksm practically eliminate the tail latency increase with faster and more efficient host-device communication than PCIe-based zswap and ksm. Houxiang Ji, Srikar Vanavasam, Yang Zhou 0050, Qirong Xia, Jinghan Huang 0001, Ren Wang 0001, Pekon Gupta, Bhushan Chitlur, Ipoom Jeong, Nam Sung Kim |
MICRO | 7 |
| 2024 | Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration
Lingfeng Xiang, Weishu Deng, Hui Lu 0001, Jia Rao, Ren Wang 0001 |
OSDI | 7 |
| 2023 | Rambda: RDMA-driven Acceleration Framework for Memory-intensive µs-scale Datacenter ApplicationsabstractResponding to the "datacenter tax" and "killer microseconds" problems for memory-intensive datacenter applications, diverse solutions including Smart NIC-based ones have been proposed. Nonetheless, they often suffer from high overhead of communications over network and/or PCIe links. To tackle the limitations of the current solutions, this paper proposes RAMBDA, a holistic network and architecture co-design solution that leverages current RDMA and emerging cache-coherent off-chip interconnect technologies. Specifically, RAMBDA consists of four hardware and software components: (1) unified abstraction of inter- and intra-machine communications synergistically managed by one-sided RDMA write and cache-coherent memory write; (2) efficient notification of requests to accelerators assisted by cache coherence; (3) cache-coherent accelerator architecture directly interacting with NIC; and (4) adaptive device-to-host data transfer for modern server memory systems comprising both DRAM and NVM exploiting state-of-the-art features in CPUs and PCIe. We prototype RAMBDA with a commercial system and evaluate three popular datacenter applications: (1) in-memory key-value store, (2) chain replication-based distributed transaction system, and (3) deep learning recommendation model inference. The evaluation shows that RAMBDA provides 30.1~69.1% lower latency, 0.2~2.5× throughput, and ~ 3× higher energy efficiency than the current state-of-the-art solutions, including Smart NIC. For those cases where Rambda performs poorly, we also envision future architecture to improve it. Jinghan Huang 0001, Jacob Nelson 0001, Dan R. K. Ports, Yipeng Wang 0002, Ren Wang 0001, Tsung-Yuan Charlie Tai, Nam Sung Kim |
HPCA | 8 |
| 2023 | Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesabstractThe ever-growing demands for memory with larger capacity and higher bandwidth have driven recent innovations on memory expansion and disaggregation technologies based on Compute eXpress Link (CXL). Especially, CXL-based memory expansion technology has recently gained notable attention for its ability not only to economically expand memory capacity and bandwidth but also to decouple memory technologies from a specific memory interface of the CPU. However, since CXL memory devices have not been widely available, they have been emulated using DDR memory in a remote NUMA node. In this paper, for the first time, we comprehensively evaluate a true CXL-ready system based on the latest 4th-generation Intel Xeon CPU with three CXL memory devices from different manufacturers. Specifically, we run a set of microbenchmarks not only to compare the performance of true CXL memory with that of emulated CXL memory but also to analyze the complex interplay between the CPU and CXL memory in depth. This reveals important differences between emulated CXL memory and true CXL memory, some of which will compel researchers to revisit the analyses and proposals from recent work. Next, we identify opportunities for memory-bandwidth-intensive applications to benefit from the use of CXL memory. Lastly, we propose a CXL-memory-aware dynamic page allocation policy, Caption to more efficiently use CXL memory as a bandwidth expander. We demonstrate that Caption can automatically converge to an empirically favorable percentage of pages allocated to CXL memory, which improves the performance of memory-bandwidth-intensive applications by up to 24% when compared to the default page allocation policy designed for traditional NUMA systems. Zeduo Yu, Reese Kuper, Chihun Song, Jinghan Huang 0001, Houxiang Ji, Siddharth Agarwal, Jiaqi Lou, Ipoom Jeong, Ren Wang 0001, Jung Ho Ahn, Tianyin Xu, Nam Sung Kim |
MICRO | 11 |
| 2022 | IDIO: Network-Driven, Inbound Network Data Orchestration on Server ProcessorsabstractHigh-bandwidth network interface cards (NICs), each capable of transferring 100s of Gigabits per second, are making inroads into the servers of next-generation datacenters. Such unprecedented data delivery rates impose immense pressure, especially on the server’s memory subsystem, as NICs first transfer network data to DRAM before processing. To alleviate the pressure, the cache hierarchy has evolved, supporting a direct data I/O (DDIO) technology to directly place network data in the last-level cache (LLC). Subsequently, various policies have been explored to manage such LLC and have proven to effectively reduce service latency and memory bandwidth consumption of network applications. However, the more recent evolution of the cache hierarchy decreased the size of LLC per core but significantly increased that of midlevel cache (MLC) with a non-inclusive policy. This calls for a re-examination of the aforementioned DDIO technology and management policies. In this paper, first, we identify three shortcomings of the current static data placement policy placing network data to LLC first and the non-inclusive policy with a commercial server system: (1) ineffectively using large MLC, (2) suffering from high rates of writebacks from MLC to LLC, and (3) breaking the isolation between application and network data enforced by limiting cache ways for DDIO. Second, to tackle the three shortcomings, we propose an intelligent direct I/O (IDIO) technology that extends DDIO to MLC and provides three synergistic mechanisms: (1) self-invalidating I/O buffer, (2) network-driven MLC prefetching, and (3) selective direct DRAM access. Our detailed experiments using a full-system simulator — capable of running modern DPDK userspace network functions while sustaining 100Gbps + network bandwidth — show that IDIO significantly reduces data movement (up to 84% MLC and LLC writeback reduction), provides LLC isolation (up to 22% performance improvement), and improves tail latency (up to 38% reduction in 99thlatency) for receive-intensive network applications. Mohammad Alian, Siddharth Agarwal, Daehoon Kim 0001, Ren Wang 0001, Nam Sung Kim |
MICRO | 7 |
| 2021 | QEI: Query Acceleration Can be Generic and Efficient in the CloudabstractData query operations of different data structures are ubiquitous and critical in today's data center infrastructures and applications. However, query operations are not always performance-optimal to be executed on general-purpose CPU cores. These operations exhibit insufficient memory-level parallelism and frontend bottlenecks due to unstructured control flow. Furthermore, the data access patterns are not cache- or prefetch-friendly. Based on our performance analysis on a commodity server, query operations can consume a large percentage of the CPU cycles in various modern cloud workloads. Existing accelerator solutions for query operations do not strike a balance between their generality, scalability, latency, and hardware complexity. In this paper, we propose QEI, a generic, integrated, and efficient acceleration solution for various data structure queries. We first abstract the query operations to a few regular steps and map them to a simple and hardware-friendly configurable finite automaton model. Based on this model, we develop the QEI architecture that allows multiple query operations to execute in parallel to maximize throughput. We also propose a novel way to integrate the accelerator into the CPU that balances performance, latency, and hardware cost. QEI keeps the main control logic near the L2 cache to leverage existing hardware resources in the core while distributing the data-intensive comparison logic to each last-level cache slice for higher parallelism. Our results with five representative data center workloads show that QEI can achieve 6.5× ~11.2× performance improvement in various scenarios with low overhead. Yipeng Wang 0002, Ren Wang 0001, Rangeen Basu Roy Chowdhury, Tsung-Yuan Charlie Tai, Nam Sung Kim |
HPCA | 3 |
| 2021 | Don't Forget the I/O When Allocating Your LLCabstractIn modern server CPUs, last-level cache (LLC) is a critical hardware resource that exerts significant influence on the performance of the workloads, and how to manage LLC is a key to the performance isolation and QoS in the cloud with multi-tenancy. In this paper, we argue that in addition to CPU cores, high-speed I/O is also important for LLC management. This is because of an Intel architectural innovation – Data Direct I/O (DDIO) – that directly injects the inbound I/O traffic to (part of) the LLC instead of the main memory. We summarize two problems caused by DDIO and show that (1) the default DDIO configuration may not always achieve optimal performance, (2) DDIO can decrease the performance of non-I/O workloads that share LLC with it by as high as 32%.We then present, the first LLC management mechanism that treats the I/O as the first-class citizen. Iat monitors and analyzes the performance of the core/LLC/DDIO using CPU’s hardware performance counters and adaptively adjusts the number of LLC ways for DDIO or the tenants that demand more LLC capacity. In addition, Iat dynamically chooses the tenants that share its LLC resource with DDIO to minimize the performance interference by both the tenants and the I/O. Our experiments with multiple microbenchmarks and real-world applications demonstrate that with minimal overhead, Iat can effectively and stably reduce the performance degradation caused by DDIO. Mohammad Alian, Yipeng Wang 0002, Ren Wang 0001, Ilia Kurakin, Tsung-Yuan Charlie Tai, Nam Sung Kim |
ISCA | 4 |
| 2020 | Data Direct I/O Characterization for Future I/O System ExplorationabstractI/O performance plays a critical role in the overall performance of modern servers. The emergence of ultra high-speed I/O devices makes the data movement between processors, main memory, and devices a major performance bottleneck. Conventionally, the main memory is used as an intermediate buffer between the processor and I/O devices and I/O devices cannot directly access processor side caches. Data Direct I/O (DDIO) technology aims to reduce the memory bandwidth utilization by enabling the I/O devices to leverage Last Level Cache (LLC) as the intermediate buffer. Our experimental results show that DDIO can completely eliminate memory bandwidth utilization while running network-or storage-intensive applications. However, when modeling the I/O subsystem using architectural simulators, DDIO is often ignored, which can result in inaccurate assessments about the I/O and memory sub system of emerging and future large-scale computer systems. In this paper, we provide a detailed background on DDIO technology in Intel server processors. Then we present our cycle-accurate I/O subsystem model in gem5 simulator that can be configured to model DDIO. We verify our model against baseline gem5 and validate it by comparing its results against a physical computer system. Mohammad Alian, Jie Zhang 0048, Ren Wang 0001, Myoungsoo Jung, Nam Sung Kim |
ISPASS | 4 |
| 2020 | RLDRM: Closed Loop Dynamic Cache Allocation with Deep Reinforcement Learning for Network Function VirtualizationabstractNetwork function virtualization (NFV) technology attracts tremendous interests from telecommunication industry and data center operators, as it allows service providers to assign resource for Virtual Network Functions (VNFs) on demand, achieving better flexibility, programmability, and scalability. To improve server utilization, one popular practice is to deploy best effort (BE) workloads along with high priority (HP) VNFs when high priority VNF's resource usage is detected to be low. The key challenge of this deployment scheme is to dynamically balance the Service level objective (SLO) and the total cost of ownership (TCO) to optimize the data center efficiency under inherently fluctuating workloads. With the recent advancement in deep reinforcement learning, we conjecture that it has the potential to solve this challenge by adaptively adjusting resource allocation to reach the improved performance and higher server utilization. In this paper, we present a closed-loop automation system RLDRM11RLDRM: Reinforcement Learning Dynamic Resource Management to dynamically adjust Last Level Cache allocation between HP VNFs and BE workloads using deep reinforcement learning. The results demonstrate improved server utilization while maintaining required SLO for the HP VNFs. Bin Li 0018, Yipeng Wang 0002, Ren Wang 0001, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Zhu Zhou, Andrew Herdrich, Ameer Haj-Ali, Ion Stoica, Krste Asanovic |
NetSoft | 3 |
| 2019 | Adaptive Cloud Application Tuning with Enhanced Structural Bayesian OptimizationabstractBayesian optimization has been widely applied on the performance tuning for workloads such as: large web applications running on Java Virtual Machine (JVM), parallel databases such as Hive / HBase, and Hive on Spark. However, with the rapidly expanding search space of these complex and large applications on cloud, the evaluation phase takes inordinate amount of time, rendering Bayesian optimization ineffective. In this paper, we propose a novel Bayesian optimization based framework, called, Adaptive Cloud Application Optimization Framework (ACAOF) to efficiently and optimally tune the performance of cloud workloads via significantly pruning the search space. We conducted extensive evaluations on ACAOF to compare with non-optimized Bayesian optimization on multiple categories of cloud workloads. The results demonstrate that the ACAOF outperforms approximately by up to 218%. The comparison with other machine learning techniques such as Random Search, Neural Network, Genetic Algorithm and Hill Climb also shows the significant effectiveness of ACAOF. Yuankun Shi, Ziyang Peng, Ren Wang 0001, Zhaojuan Bian |
CloudCom | 3 |
| 2019 | HALO: accelerating flow classification for scalable packet processing in NFVabstractNetwork Function Virtualization (NFV) has become the new standard in the cloud platform, as it provides the flexibility and agility for deploying various network services on general-purpose servers. However, it still suffers from sub-optimal performance in software packet processing. Our characterization study of virtual switches shows that the flow classification is the major bottleneck that limits the throughput of the packet processing in NFV, even though a large portion of the classification rules can be cached in the last level cache (LLC) in modern servers. Yipeng Wang 0002, Ren Wang 0001, Jian Huang 0006 |
ISCA | 3 |
| 2018 | A Network-Centric Hardware/Algorithm Co-Design to Accelerate Distributed Training of Deep Neural NetworksabstractTraining real-world Deep Neural Networks (DNNs) can take an eon (i.e., weeks or months) without leveraging distributed systems. Even distributed training takes inordinate time, of which a large fraction is spent in communicating weights and gradients over the network. State-of-the-art distributed training algorithms use a hierarchy of worker-aggregator nodes. The aggregators repeatedly receive gradient updates from their allocated group of the workers, and send back the updated weights. This paper sets out to reduce this significant communication cost by embedding data compression accelerators in the Network Interface Cards (NICs). To maximize the benefits of in-network acceleration, the proposed solution, named INCEPTIONN (In-Network Computing to Exchange and Process Training Information Of Neural Networks), uniquely combines hardware and algorithmic innovations by exploiting the following three observations. (1) Gradients are significantly more tolerant to precision loss than weights and as such lend themselves better to aggressive compression without the need for the complex mechanisms to avert any loss. (2) The existing training algorithms only communicate gradients in one leg of the communication, which reduces the opportunities for in-network acceleration of compression. (3) The aggregators can become a bottleneck with compression as they need to compress/decompress multiple streams from their allocated worker group. To this end, we first propose a lightweight and hardware-friendly lossy-compression algorithm for floating-point gradients, which exploits their unique value characteristics. This compression not only enables significantly reducing the gradient communication with practically no loss of accuracy, but also comes with low complexity for direct implementation as a hardware block in the NIC. To maximize the opportunities for compression and avoid the bottleneck at aggregators, we also propose an aggregator-free training algorithm that exchanges gradients in both legs of communication in the group, while the workers collectively perform the aggregation in a distributed manner. Without changing the mathematics of training, this algorithm leverages the associative property of the aggregation operator and enables our in-network accelerators to (1) apply compression for all communications, and (2) prevent the aggregator nodes from becoming bottlenecks. Our experiments demonstrate that INCEPTIONN reduces the communication time by 70.9~80.7% and offers 2.2~3.1x speedup over the conventional training system, while achieving the same level of accuracy. Youjie Li, Jongse Park, Mohammad Alian, Zheng Qu 0002, Peitian Pan, Ren Wang 0001, Alexander G. Schwing, Hadi Esmaeilzadeh, Nam Sung Kim |
MICRO | 7 |
| 2018 | Is cloud storage ready? Performance comparison of representative IP-based storage systems
Zhonghong Ou, Meina Song, Zhen-Huan Hwang, Antti Ylä-Jääski, Ren Wang 0001, Yong Cui 0001, Pan Hui 0001 |
J. Syst. Softw. | 5 |
| 2017 | Optimizing Open vSwitch to Support Millions of FlowsabstractSoftware switch has emerged as a critical component in software defined networking and network virtualization areas. Open vSwitch (OvS) is a widely used software switch which uses tuple space search algorithm for packet classification, and an exact match cache (EMC) for caching most frequently used flows. In this paper, we propose two new optimizations for OvS to further improve its performance and scalability. First aims to completely remove the sequential search overhead of the tuple space search layer of OvS, and second is a dynamic cache insertion optimization for the EMC to improve EMC effectiveness. We show that the optimizations can improve OvS's throughput by up to 3.5X for millions of active flows. Yipeng Wang 0002, Tsung-Yuan Charlie Tai, Ren Wang 0001, Sameh Gobriel, Janet Tseng, James Tsai |
GLOBECOM | 3 |
| 2017 | Understanding I/O Performance Behaviors of Cloud Storage from a Client's PerspectiveabstractCloud storage has gained increasing popularity in the past few years. In cloud storage, data is stored in the service provider’s data centers, and users access data via the network. For such a new storage model, our prior wisdom about conventional storage may not remain valid nor applicable to the emerging cloud storage. In this article, we present a comprehensive study to gain insight into the unique characteristics of cloud storage and optimize user experiences with cloud storage from a client’s perspective. Unlike prior measurement work that mostly aims to characterize cloud storage providers or specific client applications, we focus on analyzing the effects of various client-side factors on the user-experienced performance. Through extensive experiments and quantitative analysis, we have obtained several important findings. For example, we find that (1) a proper combination of parallelism and request size can achieve optimized bandwidths, (2) a client’s capabilities and geographical location play an important role in determining the end-to-end user-perceivable performance, and (3) the interference among mixed cloud storage requests may cause performance degradation. Based on our findings, we showcase a sampling- and inference-based method to determine a proper combination for different optimization goals. We further present a set of case studies on client-side chunking and parallelization for typical cloud-based applications. Our studies show that specific attention should be paid to fully exploiting the capabilities of clients and the great potential of cloud storage services. Binbing Hou, Feng Chen 0005, Zhonghong Ou, Ren Wang 0001, Michael P. Mesnier |
ACM Trans. Storage | 4 |
| 2016 | CAF: Core to Core Communication Acceleration FrameworkabstractAs the number of cores in a multicore system increases, core-to-core (C2C) communication is increasingly limiting the performance scaling of workloads that share data frequently. The traditional way cores communicate is by using shared memory space between them. However, shared memory communication fundamentally involves coherence invalidations and cache misses, which cause large performance overheads and incur a high amount of network traffic. Many important workloads incur significant C2C communication and are affected significantly by the costs, including pipelined packet processing which is widely used in software-based networking solutions. In these workloads, threads run on different cores and pass packets from one core to another for different stages of processing using software queues. Yipeng Wang 0002, Ren Wang 0001, Andrew Herdrich, James Tsai, Yan Solihin |
PACT | 2 |
| 2016 | Understanding I/O performance behaviors of cloud storage from a client's perspectiveabstractCloud storage has gained increasing popularity in the past few years. In cloud storage, data is stored in the service provider's data centers, and users access data via the network. For such a new storage model, our prior wisdom about conventional storage may not remain valid nor applicable to the emerging cloud storage. In this paper, we present a comprehensive study and attempt to gain insight into the unique characteristics of cloud storage, primarily from the client's perspective. Through extensive experiments and quantitative analysis, we have acquired several interesting, and in some cases unexpected, findings. (1) Parallelizing I/Os and increasing request sizes are keys to improving the performance, but optimal bandwidth may only be achieved with a proper combination of parallelism and request size. (2) Client capabilities, including CPU, memory, and storage, play an unexpectedly important role in determining the achievable performance. (3) A geographically long distance affects client-perceived performance but does not always result in lower bandwidth and longer latency. Based on our experimental studies, we further present a case study on appropriate chunking and parallelization in a cloud storage client. Our studies show that specific attention should be paid to fully exploiting the capabilities of clients and the great potential of cloud storage services. Binbing Hou, Feng Chen 0005, Zhonghong Ou, Ren Wang 0001, Michael P. Mesnier |
MSST | 4 |
| 2015 | Scaling Up Clustered Network Appliances with ScaleBricksabstractThis paper presents ScaleBricks, a new design for building scalable, clustered network appliances that must "pin" flow state to a specific handling node without being able to choose which node that should be. ScaleBricks applies a new, compact lookup structure to route packets directly to the appropriate handling node, without incurring the cost of multiple hops across the internal interconnect. Its lookup structure is many times smaller than the alternative approach of fully replicating a forwarding table onto all nodes. As a result, ScaleBricks is able to improve throughput and latency while simultaneously increasing the total number of flows that can be handled by such a cluster. This architecture is effective in practice: Used to optimize packet forwarding in an existing commercial LTE-to-Internet gateway, it increases the throughput of a four-node cluster by 23%, reduces latency by up to 10%, saves memory, and stores up to 5.7x more entries in the forwarding table. Dong Zhou 0006, Hyeontaek Lim, David G. Andersen, Michael Kaminsky, Michael Mitzenmacher, Ren Wang 0001, Ajaypal Singh |
SIGCOMM | 7 |
| 2015 | Utilize Signal Traces from Others? A Crowdsourcing Perspective of Energy Saving in Cellular Data CommunicationabstractWith the tremendous growth in wireless network deployment and increasing use of mobile devices, e.g., smartphones and tablets, improving energy efficiency in such devices, especially with communication driven workloads, is critical to providing a satisfactory user experience. Studies show that signal strength plays an important role on energy consumption of cellular data communications. While energy consumption can be minimized by accurately predicting signal strengths and reacting to it in real-time, the dynamic nature of wireless environments makes signal strengths highly unpredictable. In this paper, after analyzing in detail the signal strength variation and its impact on energy consumption, we propose to use crowdsourcing approach to optimize mobile devices' energy efficiency by utilizing signal strength traces reported/shared by other users/devices in cellular networks. Via a comprehensive measurement study, we observe that signal strength traces collected from different devices are pseudo-identical, and they even exhibit similar threshold-based behaviors in the relationship between signal strength and device power consumption. Based on our observations, we propose a predictive scheduling algorithm that: (i) selects the right set of signal strength traces based on its location, (ii) applies a filter to smooth out signal strengths and hide abrupt changes, (iii) digitizes the signal strength to “good” and “bad” areas, and (iv) schedules transmissions based on power-throughput characteristics to optimize the transmission energy efficiency. To demonstrate the efficacy of the proposed algorithms, we prototype the crowdsourcing-based predicative scheduling algorithm on Android-based smartphones. Our experiment results from real-life driving tests demonstrate that, by leveraging others' signal traces, mobile devices can save energy up to 35 percent compared to the conventional opportunistic scheduling, i.e., schedule transmissions only based on instantaneous channel conditions. Zhonghong Ou, Antti Ylä-Jääski, Pan Hui 0001, Ren Wang 0001, Alexander W. Min |
IEEE Trans. Mob. Comput. | 7 |
| 2014 | Simulating Hive Cluster for Deployment Planning, Evaluation and OptimizationabstractIn the era of big data, Hive has quickly gained popularity for its superior capability to manage and analyze very large datasets, both structured and unstructured, residing in distributed storage systems. However, great opportunity comes with great challenges: Hive query performance is impacted by many factors which makes capacity planning and tuning for Hive cluster extremely difficult. These factors include system software stacks (Hive, MapReduce framework, JVM and OS), cluster hardware configurations (processor, memory, storage, and network) and HIVE data models and distributions. Current planning methods are mostly trial-and-error or very high-level estimation based. These approaches are far from efficient and accurate, especially with the increasing software stack complexity, hardware diversity, and unavoidable data skew in distributed database system. In this paper, we propose a Hive simulation framework based on CSMethod, which simulates the whole hive query execution life cycle, including query plan generation and MapReduce task execution. The framework is validated using typical query operations with varying changes in hardware, software and workload parameters, showing high accuracy and fast simulation speed. We also demonstrate the application of this framework with two real-world use cases: helping customers to perform capacity planning and estimate business query response time before system provisioning. Kebing Wang, Zhaojuan Bian, Qian Chen 0031, Ren Wang 0001, Gen Xu |
CloudCom | 4 |
| 2014 | Joint optimization of DVFS and low-power sleep-state selection for mobile platformsabstractTo provide the ultimate mobile user experience, extended battery life is critical to small form-factor mobile platforms such as smartphones and tablets. Dynamic voltage and frequency scaling (DVFS) and low-power CPU/platform sleep states are commonly used power management features, as they allow dynamic control of power and performance to the time-varying needs of workloads. Despite the potential power saving benefit from synergistic integration of DVFS and sleep-state selection, it is challenging to optimize them jointly for mobile workloads (e.g., video streaming), and most existing work considers them only individually. To address this problem, we study joint optimization of CPU frequency (a.k.a. CPU P-states) and CPU/platform sleep-state selections to reduce energy consumption in mobile platforms. This joint optimization becomes feasible with advanced power management techniques and power aware software development methodologies that regulate (e.g., coalesce/align) system activities, making workload characteristics and system idle duration more deterministic and predictable. We then analyze the optimal operating state that minimizes the expected platform energy consumption based on workload characteristics, and present an algorithm to adapt to it at run time. Our evaluation results on mobile workloads show that the proposed scheme can reduce system power consumption by up to 24%, compared to the conventional CPU-utilization-based approach, which seeks mainly to minimize processor energy. Alexander W. Min, Ren Wang 0001, James Tsai, Tsung-Yuan Charlie Tai |
ICC | 2 |
| 2013 | Energy-efficient interconnect via Router ParkingabstractThe increase in on-chip core counts in Chip Multi Processors (CMPs) has led to the adoption of interconnects such as Mesh and Torus, which consume an increasing fraction of the chip power. Moreover, as technology and voltage continue to scale down, static power consumes a larger fraction of the total power; reducing it is increasingly important for energy proportional computing. Currently, processor designers strive to send under-utilized cores into deep sleep states in order to reduce idling power and improve overall energy efficiency. However, even in state-of-the-art CMP designs, when a core goes to sleep the router attached to it remains active in order to continue packet forwarding. In this paper, we propose Router Parking - selectively power-gating routers attached to parked cores. Router Parking ensures that network connectivity is maintained, and limits the average interconnect latency impact of packet detouring around parked routers. We present two Router Parking algorithms - an aggressive approach to park as many routers as possible, and a conservative approach that parks a limited set of routers in order to keep the impact on latency increase minimal. Further, we propose an adaptive policy to choose between the two algorithms at run-time. We evaluate our algorithms using both synthetic traffic as well as real workloads taken from SPEC CPU2006 and PARSEC 2.1 benchmark suites. Our evaluation results show that Router Parking can achieve significant savings in the total interconnect energy (average of 32%, 40% and 41% for the synthetic, SPEC CPU2006, and PARSEC 2.1 workloads, respectively). Ahmad Samih, Ren Wang 0001, Anil Krishna, Christian Maciocco, Tsung-Yuan Charlie Tai, Yan Solihin |
HPCA | 2 |
| 2012 | Evaluating Dynamics and Bottlenecks of Memory Collaboration in Cluster SystemsabstractWith the fast development of highly-integrated distributed systems (cluster systems), designers face interesting memory hierarchy design choices while attempting to avoid the notorious disk swapping. Swapping to the free remote memory through Memory Collaboration has demonstrated its cost-effectiveness compared to over provisioning the cluster for peak load requirements. Recent memory collaboration studies propose several ways on accessing the under-utilized remote memory in static system configurations, without detailed exploration of the dynamic memory collaboration. Dynamic collaboration is an important aspect given the run-time memory usage fluctuations in clustered systems. Further, as the interest in memory collaboration grows, it is crucial to understand the existing performance bottlenecks, overheads, and potential optimization. In this paper we address these two issues. First, we propose an Autonomous Collaborative Memory System (ACMS) that manages memory resources dynamically at run time to optimize performance. We implement a prototype realizing the proposed ACMS, experiment with a wide range of real-world applications, and show up to 3× performance speedup compared to a non-collaborative memory system without perceivable performance impact on nodes that provide memory. Second, we analyze, in depth, the end-to-end memory collaboration overhead and pinpoint the corresponding bottlenecks. Ahmad Samih, Ren Wang 0001, Christian Maciocco, Tsung-Yuan Charlie Tai, Ronghui Duan, Jiangang Duan, Yan Solihin |
CCGRID | 2 |
| 2011 | Reducing Power Consumption for Mobile Platforms via Adaptive Traffic CoalescingabstractBattery life remains to be a critical competitive metric for today's mobile platforms that offer ubiquitous connectivity through their wireless communication interfaces. With most usage models being driven by always-on communication activities, e.g, Internet video streaming, web browsing, etc., it is imperative to understand the impact of network activities on the overall platform power, and optimize power consumption for such activities. As shown by our investigation, various real-world network-driven workloads exhibit bursty and random behavior, which motivates our work on regulating and coalescing incoming packets to reduce platform wake events. To understand the performance impact of packet coalescing, we conduct an extensive investigation to study how coalescing may affect the throughput and user experience. Armed with the deep understandings, we propose, implement and evaluate an Adaptive Traffic Coalescing (ATC) scheme that monitors the incoming traffic at the Network Interface Card (NIC), and adaptively coalesces the packets for a limited duration in the NIC buffer, thus requiring no network or eco-system support. The proposed ATC scheme effectively reduces platform wake events, and enables the platform to enter and stay in the low-power state longer for energy efficiency. We have implemented the scheme in commercial wireless NICs. Using various mobile platforms, we evaluate the power savings and performance impact of the proposed ATC scheme. Experiments show that ATC achieves significant power saving for major platform components, around 20% for real-world Internet workloads, without impacting performance and user experience. Ren Wang 0001, James Tsai, Christian Maciocco, Tsung-Yuan Charlie Tai, Jackie Wu |
IEEE J. Sel. Areas Commun. | 1 |
| 2009 | IDC: An Energy Efficient Communication Scheme for Connected Mobile PlatformsabstractMobile platforms (e.g. laptops) offer ubiquitous network connectivity through its wireless communication interfaces, with most of usage models driven by always-on communication activities. However this in turn creates significant power challenges as the battery life is a critical resource for all mobile platforms. While the communication device itself consumes a relatively small portion of the total power, the impact of communication on the overall platform power is significant, due to the non-deterministic nature of the network traffic, which tends to keep the platform busy more than necessary. In this paper, we extensively investigate characteristics of various real- world network-centric workloads, e.g., VoIP, web browsing, etc., and their impact on the mobile platform in terms of power consumption and energy. Based on the understandings, we propose, implement and evaluate an Interrupt /DMA (Direct Memory Access) Coalescing (IDC) scheme at the wireless network interface card (NIC). This scheme effectively reduces platform wakeups due to incoming network traffic, Another advantage of the scheme is that it requires only NIC modification at the receiving node, and is transparent to the user and the network. Using a commercial NIC with a laptop testing platform, we evaluate the power savings and performance impact of our proposed scheme. The measurements show that our scheme achieves significant platform power saving, around 25% of platform components, without impacting user experience. Ajay Kulkarni, Ren Wang 0001, Christian Maciocco, Sanjay Bakshi, James Tsai |
ICC | 2 |
| 2006 | Fluid-flow analysis of TCP Westwood with RED
Fernando Paganini, M. Y. Sanadidi, Ren Wang 0001, Mario Gerla |
Comput. Networks | 4 |
| 2005 | TCP bulk repeat
Guang Yang 0001, Ren Wang 0001, Mario Gerla, M. Y. Sanadidi |
Comput. Commun. | 2 |
| 2005 | TCP with sender-side intelligence to handle dynamic, large, leaky pipesabstractTransmission control protocol Westwood (TCPW) has been shown to provide significant performance improvement over high-speed heterogeneous networks. The key idea of TCPW is to use eligible rate estimation (ERE) methods to intelligently set the congestion window (cwnd) and slow-start threshold (ssthresh) after a packet loss. ERE is defined as the efficient transmission rate eligible for a sender to achieve high utilization and be friendly to other TCP variants. This work presents TCP Westwood with agile probing (TCPW-A), a sender-side only enhancement of TCPW, that deals well with highly dynamic bandwidth, large propagation time/bandwidth, and random loss in the current and future heterogeneous Internet. TCPW-A achieves this goal by adding the following two mechanisms to TCPW. 1) When a connection initially begins or restarts after a timeout, instead of exponentially expanding cwnd to an arbitrary preset sthresh and then going into linear increase, TCPW-A uses agile probing, a mechanism that repeatedly resets ssthresh based on ERE and forces cwnd into an exponential climb each time. The result is fast convergence to a more appropriate ssthresh value. 2) In congestion avoidance, TCPW-A invokes agile probing upon detection of persistent extra bandwidth via a scheme we call persistent noncongestion detection (PNCD). While in congestion avoidance, agile probing is actually invoked under the following conditions: a) a large amount of bandwidth that suddenly becomes available due to change in network conditions; b) random loss during slow-start that causes the connection to prematurely exit the slow-start phase. Experimental results, both in ns-2 simulation and lab measurements using actual protocols implementation, show that TCPW-A can significantly improve link utilization over a wide range of bandwidth, propagation delay, and dynamic network loading. Ren Wang 0001, Kenshin Yamada, M. Y. Sanadidi, Mario Gerla |
IEEE J. Sel. Areas Commun. | 1 |
| 2004 | TCP westwood with agile probing: dealing with dynamic, large, leaky pipesabstractTCP westwood (TCPW) has been shown to provide significant performance improvement over high-speed heterogeneous networks. The key idea of TCPW is to use eligible rate estimation (ERE) methods, to set the congestion window (cwnd) and slow start threshold (ssthresh) after a packet loss. ERE is defined as the transmission rate a sender ought to use to achieve high utilization and remain friendly to other TCP variants. This paper presents TCP westwood with agile probing (TCPW-A), a sender-side only enhancement of TCPW. TCPW-A perform well when faced with highly dynamic bandwidth, large propagation time/bandwidth, and random loss in the current and future heterogeneous Internet. TCPW-A achieves its goal by incorporating the following, two mechanisms: 1) when a connection initially begins or re-starts after a timeout, instead of exponentially expanding cwnd to an arbitrary preset ssthresh and then going into linear increase. TCPW-A uses agile probing, a mechanism that repeatedly resets ssthresh based on ERE and forces cwnd into an exponential climb each time. The result is fast convergence to a more appropriate ssthresh value. 2) In congestion avoidance, TCPW-A invokes agile probing upon indication of unused extra bandwidth via a scheme we call load gauge (LG). Experimental results, both in Ns-2, and in measurements using FreeBSD implementation, show that TCPW-A can significantly improve link utilization over a wide range of bandwidth, propagation delay and dynamic network loading. Kenshin Yamada, Ren Wang 0001, M. Y. Sanadidi, Mario Gerla |
ICC | 2 |
| 2004 | TCP Start up Performance in Large Bandwidth Delay Networksabstract.4brtroct- Nest generation nehvorlis with large bandwidth and long drlay pose a major challenge to TCP performance, especially during the startup period. In this paper we evaluate the performance of TCP RenaiNcwrcno. Vegns and Hoe's modification in large bandwidth delay nrhvork. We propose n modified Slow-start mechanism, rnllcd Adaptive Start (Astart), to improve the startup performance in such networks. When a connection initially begins or re-starts after a coarse timrout,- Astart ndaptivcly and repentedly resets the Slow-start Threshold (suthreslr) based on an cligihlr sending I'iitr estimation mrchanisrn proposrd in TCP Westwond. By iidapting to network conditions during the startup phase. it wndw is able to grow the congestion window (ocnh fast without incurring risk of huNw owrflow and multiple Iossrs. Simulation rxpcrirncnts show that Astart can significantly improve the link utiliiation under various bandwidth, buNrr sur;~nd round-trip propagation timrs. The mrthud avoids both under-utiliriition dur to prrmature Slowstart termination, as wcll ils multiple I~XII~S due to initinlly setting srrlire.sli too high, or. increaing nmd tin) fiat. Experiments also show that Astart uchiews good fttirnrss rind fricndlincss toward TCP NewReno. Lab measuremrnts using a FrreBSD Astart implementation are also reported in this paper, providing futrhcr evidence of the gains nchirvahlr via Astart. Kqw-orr%r-congesrionn control;.sIow-.start; rate estimution. large bundwidth ddq nehvorks I. Ren Wang 0001, Giovanni Pau 0001, Kenshin Yamada, M. Y. Sanadidi, Mario Gerla |
INFOCOM | 1 |
| 2004 | TCP Westwood with adaptive bandwidth estimation to improve efficiency/friendliness tradeoffs
Mario Gerla, Bryan K. F. Ng, M. Y. Sanadidi, Massimo Valla, Ren Wang 0001 |
Comput. Commun. | 5 |
| 2003 | Fluid-flow analysis of TCP Westwood with REDabstractThe paper concerns TCP Westwood, a recently-developed modification of TCP, in combination with RED queue management. We develop a fluid-flow model of the protocol, and use it to study both equilibrium and dynamic features. On the equilibrium side, we identify the scaling of the congestion window with loss-probability, and compare it to TCP NewReno. We also use the model to find the boundary of stability, beyond which we see large oscillations; we find that the stable region of TCP Westwood is enhanced with respect to TCP NewReno. Furthermore we show preliminary evidence that oscillations, when they occur, have a limited impact on network throughput. Fernando Paganini, Ren Wang 0001, M. Y. Sanadidi, Mario Gerla |
GLOBECOM | 3 |
| 2003 | TCPW with bulk repeat in next generation wireless networksabstractWireless links are error-prone. In next generation wireless networks, link capacities will still grow. It is known that in such configurations the performance of current TCP variants degrades severely. In this paper we propose a new transmission strategy for TCP in high error situations- bulk repeat (BR). We apply BR to TCP westwood (TCPW). BR has only three sender-side modifications to TCP, i.e. bulk retransmission, fixed retransmission timeout and intelligent window adjustment. BR permits efficient recovery from multiple losses in the same congestion window. To discriminate error from congestion loss, a loss discrimination algorithm (LDA), based on spike and rate gap threshold, is used. Simulation results in wireless network scenarios show that TCPW BR improves throughput performance up to an order of magnitude over TCPW and newreno when the error rate is high (>5%). The results also show that TCPW BR has satisfactory fairness and friendliness to TCP newreno. Guang Yang 0001, Ren Wang 0001, M. Y. Sanadidi, Mario Gerla |
ICC | 2 |
| 2002 | Adaptive bandwidth share estimation in TCP WestwoodabstractTCP Westwood (TCPW) is a recently proposed sender side modification of TCP congestion control. TCPW relies an bandwidth share estimation techniques to enhance congestion control over high speed and/or wireless networks. The bandwidth share estimation methods turn out to be critical to guarantee both throughput improvement and friendliness towards widely used TCP protocols such as NewReno. In this paper we propose a new bandwidth share estimation technique, called "adaptive bandwidth share estimation", or ABSE. ABSE adapts to changing network congestion level, round trip times, and other relevant network conditions, as well as to the rate at which such changes occur. We compare the throughput gain and friendliness of ABSE with that of NewReno and previous estimation methods used for TCPW. We test the new technique using different simulation scenarios, including RED, to show the benefit of our proposed ABSE estimation method. Ren Wang 0001, Massimo Valla, M. Y. Sanadidi, Mario Gerla |
GLOBECOM | 1 |
| 2002 | Enhancing TCP performance in networks with small buffersabstractTCP performance can be significantly affected when the buffer capacity at routers is small. This is possible when either many flows share the network or the bandwidth-delay product is large (e.g. satellite links). The behavior of various versions of TCP with respect to buffer capacity issues has not been studied in much detail. We investigate the behavior and performance of different TCP variants under small buffer capacity conditions. We recognize TCP pacing as a potential solution. However, instead of using TCP's sending rate as the dictating metric, we make use of the bandwidth-share estimate (BSE), maintained by TCP Westwood, to set the pacing interval. We call this newly proposed protocol paced-Westwood. We also show the need to scale BSE further to mitigate the effects of positive feedback in BSE. For this, we propose a further enhancement that we call /spl alpha/-paced Westwood that uses a scaling parameter /spl alpha/ to enforce convergence of BSE and the pacing interval. The proposed /spl alpha/-paced Westwood uses its BSE to space the packet bursts during the slow-start phase, resulting in a superior throughput in the troublesome low buffer capacity cases. With the help of simulations, we show that our enhanced TCP Westwood outperforms both unpaced as well as paced TCP NewReno under low buffer capacity networks. Ashu Razdan, Alok Nandan, Ren Wang 0001, M. Y. Sanadidi, Mario Gerla |
ICCCN | 3 |
| 2002 | Using Adaptive Rate Estimation to Provide Enhanced and Robust Transport over Heterogeneous NetworksabstractThe rapid advancement in wireless communication technology has spurred significant interest in the design and development of enhanced TCP protocols. Among them, TCP Westwood (TCPW) is a sender side only modification to improve TCP performance particularly over heterogeneous networks. The key idea of TCPW is to use rate estimation methods to set the congestion window and slow start threshold after a packet loss. When packet losses are not only due to buffer overflow, but random errors as well, TCPW estimation methods have been shown to provide significant performance improvement. The earliest estimation method, called bandwidth estimation (BE), however, may result in over-estimation under certain circumstances, and thus may be unfriendly toward non-TCPW traffic. TCPW CRB (combined rate and bandwidth estimation) and TCPW ABSE (adaptive bandwidth share estimation), have been later introduced to address this concern. The schemes provide better control of the tradeoffs among efficiency, friendliness, and implementation complexity. CRB may slightly sacrifice the efficiency gain to ensure friendliness. ABSE adaptivity mechanisms are more sophisticated and provide both better efficiency and friendliness.We summarize ABSE, which adapts to congestion level, as well as round drip time, and other network dynamics, thus providing enhanced and robust performance under various network conditions. Extensive experiments show that TCPW ABSE is able to enhance TCP performance significantly over "large leaky pipes", while maintaining friendliness toward TCP NewReno. We show that TCPW ABSE is robust to packet and ACK compression due to cross traffic on forward and backward paths. We also show that ABSE is robust to buffer size variations, which are inevitable in today's networks. Ren Wang 0001, Massimo Valla, M. Y. Sanadidi, Mario Gerla |
ICNP | 1 |
| 2002 | Efficiency/friendliness tradeoffs in TCP WestwoodabstractWe propose a refinement of TCP Westwood which allows management of the efficiency/friendliness-to-NewReno tradeoff. We show that the refined TCP Westwood is able to achieve higher efficiency yet at the same time maintain friendliness. TCP Westwood (TCPW) implements a novel window congestion control algorithm based on available bandwidth estimation. The performance of TCPW has been promising, exceeding that of TCP NewReno in high speed and/or wired/wireless networks. However, under certain circumstances, TCP NewReno may experience some performance degradation because TCPW possesses more information and thus can take better advantage of available bandwidth. We propose combining the original TCPW sampling strategy that produces available bandwidth estimates (BE), with a new strategy that produces rate estimates (RE). Our studies show that RE works best when packet losses are mostly due to congestion. If on the other hand, the packet losses are mostly due to link errors, BE gives better performance. To achieve the "best of all worlds", we introduce a method we call combined rate and bandwidth estimation (CRB). A connection first infers the predominant cause of packet losses, and then uses the most appropriate estimation method. We also introduce the efficiency/friendliness tradeoff graph that provides better tradeoff visualization. In our experiments, we found that CRB provides a better compromise between efficiency and friendliness, and the means to manage such a tradeoff. Ren Wang 0001, Massimo Valla, M. Y. Sanadidi, Bryan K. F. Ng, Mario Gerla |
ISCC | 1 |
| 2002 | TCP Westwood: End-to-End Congestion Control for Wired/Wireless Networks
Claudio Casetti, Mario Gerla, Saverio Mascolo, M. Y. Sanadidi, Ren Wang 0001 |
Wirel. Networks | 5 |
| 2001 | TCP Westwood: congestion window control using bandwidth estimationabstractWe study the performance of TCP Westwood (TCPW), a new TCP protocol with a sender-side modification of the window congestion control scheme. TCP Westwood controls the window using end-to-end rate estimation in a way that is totally transparent to routers and to the destination. Thus, it is compatible with any network and TCP implementation. The key innovative idea is to continuously estimate, at the TCP sender, the packet rate of the connection by monitoring the ACK reception rate. The estimated connection rate is then used to compute congestion window and slow start threshold settings after a congestion episode. Resetting the window to match available bandwidth makes TCPW more robust to sporadic losses due to wireless channel problems. These often cause conventional TCP to overreact, leading to unnecessary window reduction. Experimental studies of TCPW show significant improvements in throughput performance over Reno and SACK, particularly in mixed wired/wireless networks over high-speed links. The contributions of this paper include a model for fair and friendly sharing of the bottleneck link and a Markov Chain performance model in presence of link errors/loss. TCPW performance is compared to that of TCP Reno, and analytic results are validated against simulation results. Internet and laboratory measurements using a Linux TCPW implementation are also reported, providing further evidence of the gains achievable via TCPW. Mario Gerla, M. Y. Sanadidi, Ren Wang 0001, Andrea Zanella, Claudio Casetti, Saverio Mascolo |
GLOBECOM | 3 |
| 2001 | TCP westwood: Bandwidth estimation for enhanced transport over wireless linksabstractTCP Westwood (TCPW) is a sender-side modification of the TCP congestion window algorithm that improves upon the performance of TCP Reno in wired as well as wireless networks. The improvement is most significant in wireless networks with lossy links, since TCP Westwood relies on end-to-end bandwidth estimation to discriminate the cause of packet loss (congestion or wireless channel effect) which is a major problem in TCP Reno. An important distinguishing feature of TCP Westwood with respect to previous wireless TCP “extensions” is that it does not require inspection and/or interception of TCP packets at intermediate (proxy) nodes. Rather, it fully complies with the end-to-end TCP design principle. The key innovative idea is to continuously measure at the TCP source the rate of the connection by monitoring the rate of returning ACKs. The estimate is then used to compute congestion window and slow start threshold after a congestion episode, that is, after three duplicate acknowledgments or after a timeout. The rationale of this strategy is simple: in contrast with TCP Reno, which “blindly” halves the congestion window after three duplicate ACKs, TCP Westwood attempts to select a slow start threshold and a congestion window which are consistent with the effective bandwidth used at the time congestion is experienced. We call this mechanism faster recovery. The proposed mechanism is particularly effective over wireless links where sporadic losses due to radio channel problems are often misinterpreted as a symptom of congestion by current TCP schemes and thus lead to an unnecessary window reduction. Experimental studies reveal improvements in throughput performance, as well as in fairness. In addition, friendliness with TCP Reno was observed in a set of experiments showing that TCP Reno connections are not starved by TCPW connections. Most importantly, TCPW is extremely effective in mixed wired and wireless networks where throughput improvements of up to 550% are observed. Finally, TCPW performs almost as well as localized link layer approaches such as the popular Snoop scheme, without incurring the O/H of a specialized link layer protocol. Saverio Mascolo, Claudio Casetti, Mario Gerla, M. Y. Sanadidi, Ren Wang 0001 |
MobiCom | 5 |